Abstract: In one possible implementation, a user's voice is represented by a textual input signal when they desire to have a conversation in Sanskrit. This opens up the possibility of a Sanskrit conversation between the user and another person who doesn't speak Sanskrit. Phonemes may be identified in the received textual input and used to produce a preliminary speech output in certain embodiments. The phonemes might be translated into a preliminary set of visemes in certain embodiments. Some implementations may include creating a speech from the phonemes and visemes found. One such embodiment is assigning a unique animated face to each of the visemes. In certain embodiments, the created animated face representations are shown on a screen for the user to practise with.
1. A method for enabling user to enable practice to speak Sanskrit efficiently, the method comprising: receiving a textual input signal representing a speech of a user who wants to engage in a Sanskrit dialogue; analysing the received textural input to identify phonemes to generate a first speech result; converting the phonemes into a first set of visemes; generating a speech based on the identified phonemes and visemes; generating a set of animated facial representations to each of the visemes; and displaying the generated animated facial representations to a display screen to provide training to a user.
2. The method of claim 1 further comprising segmenting the received textual input to create a corpus of a first plurality of words.
3. The method of claim 1, wherein the phonemes and visemes are generated using a machine learning approach.
5. The method of claim 1, further comprising segmenting the received textual input to create a corpus of a first plurality of words to facilitate phoneme identification.
6. The method of claim 1, wherein the machine learning approach is selected from Linear regression, Logistic regression, Decision tree, SVM algorithm, Naive Bayes algorithm, KNN algorithm, K-means, Random forest algorithm.
7. A system for making Sanskrit more accessible to a speaker and listener comprising: a touch screen to receive a textual input signal representing a speech of a user who wants to engage in a Sanskrit dialogue; and a control unit is arranged to analyse the received textural input to identify phonemes to generate a first speech result; convert the phonemes into a first set of visemes; generate a speech based on the identified phonemes and visemes; generate a set of animated facial representations to each of the visemes; and display the generated animated facial representations to a display screen to provide training to a user.
8. The system of claim 7, wherein control unit is configured to segment the received textual input to create a corpus of a first plurality of words.
9. The system of claim 7, wherein the phonemes and visemes are generated using a machine learning approach.
10. The system of claim 7, wherein the machine learning approach is selected from Linear regression, Logistic regression, Decision tree, SVM algorithm, Naive Bayes algorithm, KNN algorithm, K-means, Random Forest algorithm. ELECTRONIC FRAMEWORK FOR SPEAKING AND UNDERSTANDING SANSKRIT Abstract In one possible implementation, a user's voice is represented by a textual input signal when they desire to have a conversation in Sanskrit. This opens up the possibility of a Sanskrit conversation between the user and another person who doesn't speak Sanskrit. Phonemes may be identified in the received textual input and used to produce a preliminary speech output in certain embodiments. The phonemes might be translated into a preliminary set of visemes in certain embodiments. Some implementations may include creating a speech from the phonemes and visemes found. One such embodiment is assigning a unique animated face to each of the visemes. In certain embodiments, the created animated face representations are shown on a screen for the user to practise with s:
1. A method for enabling user to enable practice to speak Sanskrit efficiently, the method comprising: receiving a textual input signal representing a speech of a user who wants to engage in a Sanskrit dialogue; analysing the received textural input to identify phonemes to generate a first speech result; converting the phonemes into a first set of visemes; generating a speech based on the identified phonemes and visemes; generating a set of animated facial representations to each of the visemes; and displaying the generated animated facial representations to a display screen to provide training to a user.
2. The method of claim 1 further comprising segmenting the received textual input to create a corpus of a first plurality of words.
3. The method of claim 1, wherein the phonemes and visemes are generated using a machine learning approach.
5. The method of claim 1, further comprising segmenting the received textual input to create a corpus of a first plurality of words to facilitate phoneme identification.
6. The method of claim 1, wherein the machine learning approach is selected from Linear regression, Logistic regression, Decision tree, SVM algorithm, Naive Bayes algorithm, KNN algorithm, K-means, Random forest algorithm.
7. A system for making Sanskrit more accessible to a speaker and listener comprising: a touch screen to receive a textual input signal representing a speech of a user who wants to engage in a Sanskrit dialogue; and a control unit is arranged to analyse the received textural input to identify phonemes to generate a first speech result; convert the phonemes into a first set of visemes; generate a speech based on the identified phonemes and visemes; generate a set of animated facial representations to each of the visemes; and display the generated animated facial representations to a display screen to provide training to a user.
8. The system of claim 7, wherein control unit is configured to segment the received textual input to create a corpus of a first plurality of words.
9. The system of claim 7, wherein the phonemes and visemes are generated using a machine learning approach.
10. The system of claim 7, wherein the machine learning approach is selected from Linear regression, Logistic regression, Decision tree, SVM algorithm, Naive Bayes algorithm, KNN algorithm, K-means, Random Forest algorithm.
Field of the Invention
[0001] The present invention relates generally to system and method for enabling user to enable practice to speak and understand Sanskrit efficiently.
Background
[0002] The background description includes information that may be useful in understanding the present invention. It is not an admission that any of the information provided herein is prior art or relevant to the presently claimed invention, or that any publication specifically or implicitly referenced is prior art.
[0003] Sanskrit is an ancient language filled with rich literature and a wide variety of form. The accessibility of Sanskrit language to the speaker can decode the various formulae for the developments in different fields of science and technology from ancient times to modern times. Sanskrit is the primary liturgical language of Hinduism and the predominant language of most works of Hindu philosophy as well as some of the principal texts of Buddhism and Jainism.
[0004] Speech processing have been employed for processing of digital speech signal in conjunction with natural language. Speech has always been the basis of textual language, which contrasts in its vocabulary and phonetics from its spoken one. Parsing is the process of analysing a string of symbols either in natural language or computer languages according to the rule of formal grammar. None of the prior art approaches were able to determine the functions of words in the input sentence, especially for engaging in Sanskrit language.
[0005] The phonetic sounds of alphabets in foreign language and their counterparts in Sanskrit Varna-mala is different. Varna-mala is the Sanskrit corpus of alphabets. As an epitome of illustration, in Italian language the sounds of the alphabets clash with their counterparts in quite a few occasions. The range of foreign language vocabulary may also not be sufficient since the phonetics of a lot of alphabets do not map to individual phonetics of the Sanskrit Varnas.
[0006] There are certain patent documents which cater similar segment and filed in different jurisdictions. Few exemplary patent documents are illustrated below.
[0007] US20230016962A1 (by SERVICENOW INC) - A specification of a first natural language understanding (NLU) machine learning model for a first human communication language is received. The specification specifies a language content associated with one or more intents of the first NLU machine learning model in the first human communication language. An identification of an association between the first NLU machine learning model and a second NLU machine learning model for a second human communication language is received. The first NLU machine learning model and the second NLU machine learning model are managed together. This includes detecting a change to the first NLU machine learning model in the first human communication language and in response automatically assisting in maintaining consistency in the second NLU machine learning model in the second human communication language with respect to the detected change.
[0008] US10891435B1 (by INTERACTIONS LLC) - Machine translation is used to leverage the semantic properties (e.g., intent) already known for one natural language for use in another natural language. In a first embodiment, the corpus of a first language is translated to each other language of interest using machine translation, and the corresponding semantic properties are transferred to the translated corpuses. Semantic models can then be generated from the translated corpuses and the transferred semantic properties. In a second embodiment, given a first language for which there is a semantic model, if a query is received in a second, different language lacking its own semantic model, machine translation is used to translate the query into the first language. Then, the semantic model for the first language is applied to the translated query, thereby obtaining the semantic properties for the query, even though no semantic model existed for the language in which the query was specified.
[0009] JP2004101727A (by - NIPPON TELEGRAPH & TELEPHONE) - PROBLEM TO BE SOLVED: To provide a method, a device, and a program for multilingual speech recognition and a method, a device, and a program for multilingual speaker adaptation that secure high speech recognition precision and operate at a high speed even when a non-native speaker speaks words of one language. SOLUTION: There are provided the method, device, and program for multilingual speech recognition which are equipped with pronunciation dictionaries 102 of a plurality of languages to be recognized and corresponding sound models 103, one multilingual grammar 104 for recognition generated by using words defined in the pronunciation dictionaries 102 of the plurality of languages to be recognized, and one speech analysis part 101 and a search part 105 independent of the languages; and the search part 105 uses the sound models of the plurality of languages to be recognized for feature parameters extracted by the speech analysis part 101 from an input speech and collates them with a word string represented by the grammar for recognition to output the work string having the largest collation score as a recognition result.
[00010] Thus, there exists a need in the art for a method for enhancing accessibility of Sanskrit language to a speaker and listener.
[00011] All references, including publications, patent applications, and patents, cited herein are hereby incorporated by reference to the same extent as if each reference were individually and specifically indicated to be incorporated by reference and were set forth in its entirety herein.
Summary of the Invention
[00012] Various objects, features, and advantages of the disclosed subject matter can be more fully appreciated with reference to the following detailed description of the disclosed subject matter when considered in connection with the following drawings, in which like reference numerals identify like elements.
[00013] The present invention relates generally to system and method for enabling user to enable practice to speak and understand Sanskrit efficiently.
[00014] Embodiments of the present disclosure may include a method for making Sanskrit more accessible to a speaker and listener, the method including receiving a textual input signal representing a speech of a user who wants to engage in a Sanskrit dialogue. Embodiments may also include analysing the received textural input to identify phonemes to generate a first speech result.
[00015] Embodiments may also include converting the phonemes into a first set of visemes. Embodiments may also include generating a speech based on the identified phonemes and visemes. Embodiments may also include generating a set of animated facial representations to each of the visemes. Embodiments may also include displaying the generated animated facial representations to a display screen to provide training to a user.
[00016] In some embodiments, the method may include segmenting the received textual input to create a corpus of a first plurality of words. In some embodiments, the phonemes and visemes may be generated using a machine learning approach. In some embodiments, the method may include segmenting the received textual input to create a corpus of a first plurality of words to facilitate phoneme identification. In some embodiments, the machine learning approach may be selected from Linear regression, Logistic regression, Decision tree, SVM algorithm, Naive Bayes algorithm, KNN algorithm, K-means, Random forest algorithm.
[00017] Embodiments of the present disclosure may also include a system for making Sanskrit more accessible to a speaker and listener including a touch screen to receive a textual input signal representing a speech of a user who wants to engage in a Sanskrit dialogue. Embodiments may also include a control unit may be arranged to. Embodiments may also include analyse the received textural input to identify phonemes to generate a first speech result.
[00018] Embodiments may also include convert the phonemes into a first set of visemes. Embodiments may also include generate a speech based on the identified phonemes and visemes. Embodiments may also include generate a set of animated facial representations to each of the visemes. Embodiments may also include display the generated animated facial representations to a display screen to provide training to a user.
[00019] Embodiments may also include control unit may be configured to segment the received textual input to create a corpus of a first plurality of words. In some embodiments, the phonemes and visemes may be generated using a machine learning approach. In some embodiments, the machine learning approach may be selected from Linear regression, Logistic regression, Decision tree, SVM algorithm, Naive Bayes algorithm, KNN algorithm, K-means, Random Forest algorithm.
[00020] Various objects, features, aspects and advantages of the inventive subject matter will become more apparent from the following detailed description of preferred embodiments, along with the accompanying drawing figures in which like numerals represent like components.
Brief Description of the Drawings
[00021] The features and advantages of the present disclosure would be more clearly understood from the following description taken in conjunction with the accompanying drawings in which:
[00022] FIG. 1 is a flowchart illustrating a method enabling user to speak and understand Sanskrit efficiently, according to some embodiments of the present disclosure.
[00023] FIG. 2 is a block diagram illustrating a system to speak and understand Sanskrit efficiently, according to some embodiments of the present disclosure.
Detailed Description
[00024] The following is a detailed description of exemplary embodiments to illustrate the principles of the invention. The embodiments are provided to illustrate aspects of the invention, but the invention is not limited to any embodiment. The scope of the invention encompasses numerous alternatives, modifications and equivalent; it is limited only by the claims.
[00025] In view of the many possible embodiments to which the principles of the present discussion may be applied, it should be recognized that the embodiments described herein with respect to the drawing figures are meant to be illustrative only and should not be taken as limiting the scope of the claims. Therefore, the techniques as described herein contemplate all such embodiments as may come within the scope of the following claims and equivalents thereof.
[00026] The detailed description is described with reference to the accompanying figures. In the figures, the left-most digit(s) of a reference number identifies the figure in which the reference number first appears. The use of the same reference numbers in different instances in the description and the figures may indicate similar or identical items.
[00027] The present invention relates generally to system and method for enabling user to enable practice to speak and understand Sanskrit efficiently.
[00028] FIG. 1 is a flowchart that describes a method for making Sanskrit, according to some embodiments of the present disclosure. In some embodiments, at 110, the method may include receiving a textual input signal representing a speech of a user who wants to engage in a Sanskrit dialogue. At 120, the method may include analysing the received textural input to identify phonemes to generate a first speech result. At 130, the method may include converting the phonemes into a first set of visemes. At 140, the method may include generating a speech based on the identified phonemes and visemes. At 150, the method may include generating a set of animated facial representations to each of the visemes. At 160, the method may include displaying the generated animated facial representations to a display screen to provide training to a user.
[00029] In some embodiments, the method may include segmenting the received textual input to create a corpus of a first plurality of words. In some embodiments, the phonemes and visemes may be generated using a machine learning approach. In some embodiments, the method may include segmenting the received textual input to create a corpus of a first plurality of words to facilitate phoneme identification. In some embodiments, the machine learning approach may be selected from Linear regression, Logistic regression, Decision tree, SVM algorithm, Naive Bayes algorithm, KNN algorithm, K-means, Random forest algorithm.
[00030] FIG. 2 is a block diagram that describes a system 200, according to some embodiments of the present disclosure. In some embodiments, the system 200 may include a touch screen 210 to receive a textual input signal representing a speech of a user who wants to engage in a Sanskrit dialogue and a control unit 220 may be arranged to. Analyse the received textural input to identify phonemes to generate a first speech result. Convert the phonemes into a first set of visemes. Generate a speech based on the identified phonemes and visemes. Generate a set of animated facial representations to each of the visemes. Display the generated animated facial representations to a display screen to provide training to a user.
[00031] In some embodiments, control unit may be configured to segment the received textual input to create a corpus of a first plurality of words. In some embodiments, the phonemes and visemes may be generated using a machine learning approach. In some embodiments, the machine learning approach may be selected from Linear regression, Logistic regression, Decision tree, SVM algorithm, Naive Bayes algorithm, KNN algorithm, K-means, Random Forest algorithm.
[00032] In one possible implementation, a user's voice is represented by a textual input signal when they desire to have a conversation in Sanskrit. This opens up the possibility of a Sanskrit conversation between the user and another person who doesn't speak and understand Sanskrit. Phonemes may be identified in the received textual input and used to produce a preliminary speech output in certain embodiments. The phonemes might be translated into a preliminary set of visemes in certain embodiments. Some implementations may include creating a speech from the phonemes and visemes found. One such embodiment is assigning a unique animated face to each of the visemes. In certain embodiments, the created animated face representations are shown on a screen for the user to practise with. Segmenting the incoming textual input to generate a corpus of a first plurality of words is an option in certain implementations of the approach. A machine learning technique may be used to produce phonemes and visemes in certain implementations. Segmenting the incoming textual input to generate a corpus of a first plurality of words is one possible implementation of the approach that may aid in phoneme identification. Some examples of appropriate machine learning techniques are linear regression, logistic regression, decision trees, support vector machines, naive bayes, k-nearest neighbours, k-means, and random forests. A touch screen may be used to accept a textual input signal representing a speech of a user who wishes to participate in a Sanskrit conversation, which is one embodiment of a system for making Sanskrit more accessible to a speaker and listener. A control unit may be provided in certain embodiments as well. Some embodiments may additionally include analysing the received textual input to identify phonemes in order to provide a preliminary voice output. In certain implementations, the phonemes are translated into an initial set of visemes. In certain implementations, the phonemes and visemes are used to form a speech. One possible embodiment is to give each of the visemes an animated face. Displaying the produced animated face representations on a display screen for the purpose of instructing a user is another possible embodiment. The control unit of certain embodiments is programmed to divide the incoming text into smaller pieces, which are then used to form a corpus of a first set of words. A machine learning technique may be used to produce phonemes and visemes in certain implementations. Linear regression, logistic regression, decision trees, support vector machines, naive bayes, k-nearest neighbours, and random forests are some examples of machine learning techniques that may be used in various implementations.
[00033] Further, while operations are depicted in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Likewise, while several specific implementation details are contained in the above discussions, these should not be construed as limitations on the scope of the subject matter described herein, but rather as descriptions of features that may be specific to particular embodiments. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable sub-combination. Moreover, although features may be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a sub-combination or variation of a sub-combination.
Claims
I/We claims:
1. A method for enabling user to enable practice to speak Sanskrit efficiently, the method comprising:
receiving a textual input signal representing a speech of a user who wants to engage in a Sanskrit dialogue;
analysing the received textural input to identify phonemes to generate a first speech result; converting the phonemes into a first set of visemes;
generating a speech based on the identified phonemes and visemes;
generating a set of animated facial representations to each of the visemes; and
displaying the generated animated facial representations to a display screen to provide training to a user.
2. The method of claim 1 further comprising segmenting the received textual input to create a corpus of a first plurality of words.
3. The method of claim 1, wherein the phonemes and visemes are generated using a machine learning approach.
5. The method of claim 1, further comprising segmenting the received textual input to create a corpus of a first plurality of words to facilitate phoneme identification.
6. The method of claim 1, wherein the machine learning approach is selected from Linear regression, Logistic regression, Decision tree, SVM algorithm, Naive Bayes algorithm, KNN algorithm, K-means, Random forest algorithm.
7. A system for making Sanskrit more accessible to a speaker and listener comprising:
a touch screen to receive a textual input signal representing a speech of a user who wants to engage in a Sanskrit dialogue; and
a control unit is arranged to
analyse the received textural input to identify phonemes to generate a first speech result;
convert the phonemes into a first set of visemes;
generate a speech based on the identified phonemes and visemes;
generate a set of animated facial representations to each of the visemes; and
display the generated animated facial representations to a display screen to provide training to a user.
8. The system of claim 7, wherein control unit is configured to segment the received textual input to create a corpus of a first plurality of words.
9. The system of claim 7, wherein the phonemes and visemes are generated using a machine learning approach.
10. The system of claim 7, wherein the machine learning approach is selected from Linear regression, Logistic regression, Decision tree, SVM algorithm, Naive Bayes algorithm, KNN algorithm, K-means, Random Forest algorithm.
ELECTRONIC FRAMEWORK FOR SPEAKING AND UNDERSTANDING SANSKRIT
Abstract
In one possible implementation, a user's voice is represented by a textual input signal when they desire to have a conversation in Sanskrit. This opens up the possibility of a Sanskrit conversation between the user and another person who doesn't speak Sanskrit. Phonemes may be identified in the received textual input and used to produce a preliminary speech output in certain embodiments. The phonemes might be translated into a preliminary set of visemes in certain embodiments. Some implementations may include creating a speech from the phonemes and visemes found. One such embodiment is assigning a unique animated face to each of the visemes. In certain embodiments, the created animated face representations are shown on a screen for the user to practise with
I/We claims:
1. A method for enabling user to enable practice to speak Sanskrit efficiently, the method comprising:
receiving a textual input signal representing a speech of a user who wants to engage in a Sanskrit dialogue;
analysing the received textural input to identify phonemes to generate a first speech result; converting the phonemes into a first set of visemes;
generating a speech based on the identified phonemes and visemes;
generating a set of animated facial representations to each of the visemes; and
displaying the generated animated facial representations to a display screen to provide training to a user.
2. The method of claim 1 further comprising segmenting the received textual input to create a corpus of a first plurality of words.
3. The method of claim 1, wherein the phonemes and visemes are generated using a machine learning approach.
5. The method of claim 1, further comprising segmenting the received textual input to create a corpus of a first plurality of words to facilitate phoneme identification.
6. The method of claim 1, wherein the machine learning approach is selected from Linear regression, Logistic regression, Decision tree, SVM algorithm, Naive Bayes algorithm, KNN algorithm, K-means, Random forest algorithm.
7. A system for making Sanskrit more accessible to a speaker and listener comprising:
a touch screen to receive a textual input signal representing a speech of a user who wants to engage in a Sanskrit dialogue; and
a control unit is arranged to
analyse the received textural input to identify phonemes to generate a first speech result;
convert the phonemes into a first set of visemes;
generate a speech based on the identified phonemes and visemes;
generate a set of animated facial representations to each of the visemes; and
display the generated animated facial representations to a display screen to provide training to a user.
8. The system of claim 7, wherein control unit is configured to segment the received textual input to create a corpus of a first plurality of words.
9. The system of claim 7, wherein the phonemes and visemes are generated using a machine learning approach.
10. The system of claim 7, wherein the machine learning approach is selected from Linear regression, Logistic regression, Decision tree, SVM algorithm, Naive Bayes algorithm, KNN algorithm, K-means, Random Forest algorithm.
| # | Name | Date |
|---|---|---|
| 1 | 202311014538-REQUEST FOR EARLY PUBLICATION(FORM-9) [03-03-2023(online)].pdf | 2023-03-03 |
| 2 | 202311014538-POWER OF AUTHORITY [03-03-2023(online)].pdf | 2023-03-03 |
| 3 | 202311014538-FORM-9 [03-03-2023(online)].pdf | 2023-03-03 |
| 4 | 202311014538-FORM FOR SMALL ENTITY(FORM-28) [03-03-2023(online)].pdf | 2023-03-03 |
| 5 | 202311014538-FORM 1 [03-03-2023(online)].pdf | 2023-03-03 |
| 6 | 202311014538-EVIDENCE FOR REGISTRATION UNDER SSI(FORM-28) [03-03-2023(online)].pdf | 2023-03-03 |
| 7 | 202311014538-EVIDENCE FOR REGISTRATION UNDER SSI [03-03-2023(online)].pdf | 2023-03-03 |
| 8 | 202311014538-EDUCATIONAL INSTITUTION(S) [03-03-2023(online)].pdf | 2023-03-03 |
| 9 | 202311014538-DRAWINGS [03-03-2023(online)].pdf | 2023-03-03 |
| 10 | 202311014538-DECLARATION OF INVENTORSHIP (FORM 5) [03-03-2023(online)].pdf | 2023-03-03 |
| 11 | 202311014538-COMPLETE SPECIFICATION [03-03-2023(online)].pdf | 2023-03-03 |