Abstract: System and method to improve a spoken Hindi language skill of a speaker Abstract A method for assessing spoken language abilities may be included in certain embodiments of the present disclosure. This method may involve receiving an audio signal from a user that includes spoken language as part of the signal. Doing a spectrum analysis on the audio signal in order to acquire a collection of spectral characteristics is another option that may be included in certain embodiments. A set of linguistic characteristics may be obtained from an embodiment by first applying a trained acoustic model to the set of spectral features, which is another possible embodiment. In certain embodiments, there is also the possibility of conducting an analysis of the collection of linguistic characteristics in order to arrive at a measurement of the user's proficiency in spoken language.
1. A method for evaluating spoken language skills, comprising: receiving an audio signal comprising spoken language from a user; performing a spectral analysis on the audio signal to obtain a set of spectral features; applying a trained acoustic model to the set of spectral features to obtain a set of linguistic features; and analysing the set of linguistic features to determine a measure of spoken language skill for the user.
2. The method of claim 1, wherein the trained acoustic model is a deep neural network.
3. The method of claim 1, wherein the set of linguistic features comprises phoneme-level information.
4. The method of claim 1, further comprising comparing the measure of spoken language skill for the user to a predetermined threshold.
5. The method of claim 4, further comprising providing feedback to the user based on the comparison of the measure of spoken language skill to the predetermined threshold. 6 The method of claim 1, wherein the measure of spoken language skill is based on one or more of fluency, pronunciation, and grammar.
7. A system for evaluating spoken language skills, comprising: a microphone for receiving an audio signal comprising spoken language from a user; a processor configured to perform a spectral analysis on the audio signal to obtain a set of spectral features, and to apply a trained acoustic model to the set of spectral features to obtain a set of linguistic features; a memory for storing the trained acoustic model and instructions for analysing the set of linguistic features to determine a measure of spoken language skill for the user. 8: The system of claim 7, further comprising (d) a user interface for providing feedback to the user based on the measure of spoken language skill determined by the processor.
9. The system of claim 7, wherein the trained acoustic model is a deep neural network.
10. The system of claim 7, wherein the set of linguistic features comprises phoneme-level information. System and method to improve a spoken Hindi language skill of a speaker Abstract A method for assessing spoken language abilities may be included in certain embodiments of the present disclosure. This method may involve receiving an audio signal from a user that includes spoken language as part of the signal. Doing a spectrum analysis on the audio signal in order to acquire a collection of spectral characteristics is another option that may be included in certain embodiments. A set of linguistic characteristics may be obtained from an embodiment by first applying a trained acoustic model to the set of spectral features, which is another possible embodiment. In certain embodiments, there is also the possibility of conducting an analysis of the collection of linguistic characteristics in order to arrive at a measurement of the user's proficiency in spoken language. , Claims:Claims :
1. A method for evaluating spoken language skills, comprising: receiving an audio signal comprising spoken language from a user; performing a spectral analysis on the audio signal to obtain a set of spectral features; applying a trained acoustic model to the set of spectral features to obtain a set of linguistic features; and analysing the set of linguistic features to determine a measure of spoken language skill for the user.
2. The method of claim 1, wherein the trained acoustic model is a deep neural network.
3. The method of claim 1, wherein the set of linguistic features comprises phoneme-level information.
4. The method of claim 1, further comprising comparing the measure of spoken language skill for the user to a predetermined threshold.
5. The method of claim 4, further comprising providing feedback to the user based on the comparison of the measure of spoken language skill to the predetermined threshold. 6 The method of claim 1, wherein the measure of spoken language skill is based on one or more of fluency, pronunciation, and grammar.
7. A system for evaluating spoken language skills, comprising: a microphone for receiving an audio signal comprising spoken language from a user; a processor configured to perform a spectral analysis on the audio signal to obtain a set of spectral features, and to apply a trained acoustic model to the set of spectral features to obtain a set of linguistic features; a memory for storing the trained acoustic model and instructions for analysing the set of linguistic features to determine a measure of spoken language skill for the user. 8: The system of claim 7, further comprising (d) a user interface for providing feedback to the user based on the measure of spoken language skill determined by the processor.
9. The system of claim 7, wherein the trained acoustic model is a deep neural network.
10. The system of claim 7, wherein the set of linguistic features comprises phoneme-level information.
Description:System and method to improve a spoken Hindi language skill of a speaker
Field of the Invention
[0001] The present invention relates generally to methodology of speech processing. More particularly to a system and technique to assist a speaker/user to improve a spoken Hindi language skill.
Background
[0002] The background description includes information that may be useful in understanding the present invention. It is not an admission that any of the information provided herein is prior art or relevant to the presently claimed invention, or that any publication specifically or implicitly referenced is prior art.
[0003] Spoken language skill evaluation is the process of assessing an individual's ability to communicate effectively in a spoken language. It is an important aspect of language learning, teaching, and assessment, as well as for various applications, such as language training, hiring decisions, and immigration policies Automated spoken language skill evaluation has many potential benefits, such as improving the accuracy and consistency of evaluations, reducing costs and time required for evaluations, and providing objective feedback to language learners. However, there are also concerns about the potential limitations and biases of automated systems, as well as privacy and security concerns related to the use of speech data.
[0004] Various technological solutions (e.g., interactive web interface for interactive voice response, recording apparatus, etc.) for evaluating spoken language skills are disclosed in patent literature.
[0005] The US8223929B2 (By- Avaya Inc) relates to a system and method for establishing a communication session between an end user and a communication destination of an interactive voice response (IVR) system. A click-to-call server transmits a web page over a network for display in a browser window at a client system. The web page presents a web menu corresponding to an IVR audio menu of the IVR system and has a link by which the end user of the client system can submit a click-to-call request. The click-to-call server receives a click-to-call request from the client system over the network in response to activation of the link in the web page. The click-to-call server sets up a communication session between a communication device identified by the end user and a communication destination of the IVR system in response to receiving the click-to-call request.
[0006] The JP5417957B2 (By- Seiko Epson Corp) relates to a recording apparatus that includes a driving unit that drives the transporting unit; a phase detecting unit that detects a phase origin of the transport roller rotating in accordance with the driving of the driving unit and a rotational phase indicating an amount of rotation from the phase origin; a recording unit that performs recording on the sheet transported with the transporting unit; and a control unit that controls the driving unit such that the transport roller rotates within a certain transport range based on a preliminarily set reference phase, forms a first pattern by controlling the recording unit at a first rotation phase at a control start point of the driving unit, forms a second pattern by controlling the recording unit at a second rotation phase at a control end point of the driving unit, and then forms a correction pattern including the first and second patterns.
[0007] However, these evaluations can be time-consuming, expensive, and prone to bias. Therefore, there is a growing interest in developing automated systems for evaluating spoken language skills.
Summary
[0008] The present invention relates generally to methodology of speech processing. More particularly to a system and technique to assist a speaker/user to improve a spoken Hindi language skill.
[0009] The following presents a simplified summary of various aspects of this disclosure in order to provide a basic understanding of such aspects. This summary is not an extensive overview of all contemplated aspects, and is intended to neither identify key or critical elements nor delineate the scope of such aspects. Its purpose is to present some concepts of this disclosure in a simplified form as a prelude to the more detailed description that is presented later.
[00010] The following paragraphs provide additional support for the claims of the subject application.
[00011] Embodiments of the present disclosure may include a method for evaluating spoken language skills, including receiving an audio signal including spoken language from a user. Embodiments may also include performing a spectral analysis on the audio signal to obtain a set of spectral features. Embodiments may also include applying a trained acoustic model to the set of spectral features to obtain a set of linguistic features. Embodiments may also include analysing the set of linguistic features to determine a measure of spoken language skill for the user.
[00012] In some embodiments, the trained acoustic model may be a deep neural network. In some embodiments, the set of linguistic features may include phoneme-level information. In some embodiments, the method may include comparing the measure of spoken language skill for the user to a predetermined threshold.
[00013] In some embodiments, the method may include providing feedback to the user based on the comparison of the measure of spoken language skill to the predetermined threshold. Embodiments may also include measurement of spoken language skill may be based on one or more of fluency, pronunciation, and grammar.
[00014] Embodiments of the present disclosure may also include a system for evaluating spoken language skills, including a microphone for receiving an audio signal including spoken language from a user. Embodiments may also include a processor configured to perform a spectral analysis on the audio signal to obtain a set of spectral features, and to apply a trained acoustic model to the set of spectral features to obtain a set of linguistic features.
[00015] Embodiments may also include a memory for storing the trained acoustic model and instructions for analysing the set of linguistic features to determine a measure of spoken language skill for the user. Embodiments may also include a user interface for providing feedback to the user based on the measure of spoken language skill determined by the processor. In some embodiments, the trained acoustic model may be a deep neural network. In some embodiments, the set of linguistic features may include phoneme-level information.
Brief Description of the Drawings
[00016] The features and advantages of the present disclosure would be more clearly understood from the following description taken in conjunction with the accompanying drawings in which:
[00017] FIG. 1 is a flowchart illustrating a method for evaluating spoken language skills, according to some embodiments of the present disclosure.
[00018] FIG. 2 is a detailed flowchart further illustrating the method for evaluating spoken language skills from FIG. 1, according to some embodiments of the present disclosure.
[00019] FIG. 3 is a block diagram illustrating a system for evaluating spoken language skills, according to some embodiments of the present disclosure.
Detailed Description
[00020] In the following detailed description of the invention, reference is made to the accompanying drawings that form a part hereof, and in which is shown, by way of illustration, specific embodiments in which the invention may be practiced. In the drawings, like numerals describe substantially similar components throughout the several views. These embodiments are described in sufficient detail to claim those skilled in the art to practice the invention. Other embodiments may be utilized and structural, logical, and electrical changes may be made without departing from the scope of the present invention. The following detailed description is, therefore, not to be taken in a limiting sense, and the scope of the present invention is defined only by the appended claims and equivalents thereof.
[00021] The use of the terms a and an and the and at least one and similar referents in the context of describing the invention (especially in the context of the following claims) are to be construed to cover both the singular and the plural, unless otherwise indicated herein or clearly contradicted by context. The use of the term at least one followed by a list of one or more items (for example, at least one of A and B) is to be construed to mean one item selected from the listed items (A or B) or any combination of two or more of the listed items (A and B), unless otherwise indicated herein or clearly contradicted by context. The terms comprising, having, including, and containing are to be construed as open-ended terms (i.e., meaning including, but not limited to,) unless otherwise noted. Recitation of ranges of values herein are merely intended to serve as a shorthand method of referring individually to each separate value falling within the range, unless otherwise indicated herein, and each separate value is incorporated into the specification as if it were individually recited herein. All methods described herein can be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by context. The use of any and all examples, or exemplary language (e.g., such as) provided herein, is intended merely to better illuminate the invention and does not pose a limitation on the scope of the invention unless otherwise claimed. No language in the specification should be construed as indicating any non-claimed element as essential to the practice of the invention.
[00022] The present invention relates generally to methodology of speech processing. More particularly to a system and technique to assist a speaker/user to improve a spoken Hindi language skill.
[00023] FIG. 1 is a flowchart that describes a method for evaluating spoken language skills, according to some embodiments of the present disclosure. In some embodiments, at 110, the method may include receiving an audio signal comprising spoken language from a user. At 120, the method may include performing a spectral analysis on the audio signal to obtain a set of spectral features. At 130, the method may include applying a trained acoustic model to the set of spectral features to obtain a set of linguistic features. At 140, the method may include analysing the set of linguistic features to determine a measure of spoken language skill for the user. In some embodiments, the trained acoustic model may be a deep neural network. In some embodiments, the set of linguistic features may comprise phoneme-level information.
[00024] FIG. 2 is a detailed flowchart that further describes the method for evaluating spoken language skills (from FIG. 1), according to some embodiments of the present disclosure. In some embodiments, at 210, the method may include comparing the measure of spoken language skill for the user to a predetermined threshold. In some embodiments, at 220, the method may include providing feedback to the user based on the comparison of the measure of spoken language skill to the predetermined threshold. In an embodiment, measurement of spoken language skill may be based on one or more of fluency, pronunciation, and grammar.
[00025] FIG. 3 is a block diagram that describes a system 300 for evaluating spoken language skills, according to some embodiments of the present disclosure. In some embodiments, the system 300 may include a microphone 310 for receiving an audio signal, a memory 330 for storing the trained acoustic model and instructions for analysing the set of linguistic features to determine a measure of spoken language skill for the user, anda user interface 340 for providing feedback to the user based on the measure of spoken language skill determined by a processor 320. The system 300 may also include theprocessor 320 configured to perform a spectral analysis on the audio signal to obtain a set of spectral features, and to apply a trained acoustic model to the set of spectral features to obtain a set of linguistic features. The microphone 310 may include spoken language 312 from a user. In some embodiments, the trained acoustic model may be a deep neural network. In some embodiments, the set of linguistic features may include phoneme-level information.
[00026] Embodiments of the present disclosure may feature a method for assessing spoken language abilities, comprising receiving an audio signal containing spoken language from the user. Embodiments may further comprise running a spectrum analysis on the audio signal to get a set of spectral characteristics. Embodiments may additionally comprise applying a trained acoustic model to the set of spectral data to get a set of linguistic features. Embodiments may additionally comprise assessing the collection of linguistic characteristics to derive a measure of spoken language proficiency for the user.
[00027] In certain examples, the learned audio model may be a deep neural network. In certain instances, the collection of linguistic traits may comprise phoneme-level information. In certain implementations, the procedure may comprise comparing the measure of spoken language proficiency for the user to a specified threshold.
[00028] In other examples, the process may involve delivering feedback to the user based on the comparison of the measure of spoken language competence to the predefined threshold. Embodiments may also include assessment of spoken language proficiency may be based on one or more of fluency, pronunciation, and grammar.
[00029] Embodiments of the present disclosure may additionally comprise a system for assessing spoken language abilities, wherein the system comprises the microphone 310 for receiving an audio signal containing spoken language from a user. Embodiments may additionally contain the processor configured to execute a spectrum analysis on the audio signal to acquire a set of spectral characteristics, and to apply the trained acoustic model to the set of spectral data to generate the set of linguistic features.
[00030] Embodiments may further comprise the memory for storing the trained acoustic model and instructions for evaluating the collection of linguistic characteristics to derive the measure of spoken language proficiency for the user. Embodiments may further comprise the user interface for giving feedback to the user depending on the measure of spoken language proficiency calculated by the processor. In certain examples, the learned audio model may be a deep neural network. In certain instances, the collection of linguistic traits may comprise phoneme-level information.
[00031] Example embodiments herein have been described above with reference to block diagrams and flowchart illustrations of methods and apparatuses. It will be understood that each block of the block diagrams and flowchart illustrations, and combinations of blocks in the block diagrams and flowchart illustrations, respectively, can be implemented by various means including hardware, software, firmware, and a combination thereof. For example, in one embodiment, each block of the block diagrams and flowchart illustrations, and combinations of blocks in the block diagrams and flowchart illustrations can be implemented by computer program instructions. These computer program instructions may be loaded onto a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions which execute on the computer or other programmable data processing apparatus create means for implementing the functions specified in the flowchart block or blocks.
[00032] Throughout the present disclosure, the term ‘processing means’ or ‘microprocessor’ or ‘processor’ or ‘processors’ includes, but is not limited to, a general purpose processor (such as, for example, a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a microprocessor implementing other types of instruction sets, or a microprocessor implementing a combination of types of instruction sets) or a specialized processor (such as, for example, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), or a network processor).
[00033] The term non-transitory storage device or storage or memory, as used herein relates to a random access memory, read only memory and variants thereof, in which a computer can store data or software for any duration.
[00034] Operations in accordance with a variety of aspects of the disclosure is described above would not have to be performed in the precise order described. Rather, various steps can be handled in reverse order or simultaneously or not at all.
[00035] While several implementations have been described and illustrated herein, a variety of other means and/or structures for performing the function and/or obtaining the results and/or one or more of the advantages described herein may be utilized, and each of such variations and/or modifications is deemed to be within the scope of the implementations described herein. More generally, all parameters, dimensions, materials, and configurations described herein are meant to be exemplary and that the actual parameters, dimensions, materials, and/or configurations will depend upon the specific application or applications for which the teachings is/are used. Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific implementations described herein. It is, therefore, to be understood that the foregoing implementations are presented by way of example only and that, within the scope of the appended claims and equivalents thereto, implementations may be practiced otherwise than as specifically described and claimed. Implementations of the present disclosure are directed to each individual feature, system, article, material, kit, and/or method described herein. In addition, any combination of two or more such features, systems, articles, materials, kits, and/or methods, if such features, systems, articles, materials, kits, and/or methods are not mutually inconsistent, is included within the scope of the present disclosure.
Claims
I/We Claim:
1. A method for evaluating spoken language skills, comprising:
receiving an audio signal comprising spoken language from a user;
performing a spectral analysis on the audio signal to obtain a set of spectral features;
applying a trained acoustic model to the set of spectral features to obtain a set of linguistic features; and
analysing the set of linguistic features to determine a measure of spoken language skill for the user.
2. The method of claim 1, wherein the trained acoustic model is a deep neural network.
3. The method of claim 1, wherein the set of linguistic features comprises phoneme-level information.
4. The method of claim 1, further comprising comparing the measure of spoken language skill for the user to a predetermined threshold.
5. The method of claim 4, further comprising providing feedback to the user based on the comparison of the measure of spoken language skill to the predetermined threshold.
6 The method of claim 1, wherein the measure of spoken language skill is based on one or more of fluency, pronunciation, and grammar.
7. A system for evaluating spoken language skills, comprising:
a microphone for receiving an audio signal comprising spoken language from a user;
a processor configured to perform a spectral analysis on the audio signal to obtain a set of spectral features, and to apply a trained acoustic model to the set of spectral features to obtain a set of linguistic features;
a memory for storing the trained acoustic model and instructions for analysing the set of linguistic features to determine a measure of spoken language skill for the user.
8: The system of claim 7, further comprising (d) a user interface for providing feedback to the user based on the measure of spoken language skill determined by the processor.
9. The system of claim 7, wherein the trained acoustic model is a deep neural network.
10. The system of claim 7, wherein the set of linguistic features comprises phoneme-level information.
System and method to improve a spoken Hindi language skill of a speaker
Abstract
A method for assessing spoken language abilities may be included in certain embodiments of the present disclosure. This method may involve receiving an audio signal from a user that includes spoken language as part of the signal. Doing a spectrum analysis on the audio signal in order to acquire a collection of spectral characteristics is another option that may be included in certain embodiments. A set of linguistic characteristics may be obtained from an embodiment by first applying a trained acoustic model to the set of spectral features, which is another possible embodiment. In certain embodiments, there is also the possibility of conducting an analysis of the collection of linguistic characteristics in order to arrive at a measurement of the user's proficiency in spoken language. , Claims:Claims
I/We Claim:
1. A method for evaluating spoken language skills, comprising:
receiving an audio signal comprising spoken language from a user;
performing a spectral analysis on the audio signal to obtain a set of spectral features;
applying a trained acoustic model to the set of spectral features to obtain a set of linguistic features; and
analysing the set of linguistic features to determine a measure of spoken language skill for the user.
2. The method of claim 1, wherein the trained acoustic model is a deep neural network.
3. The method of claim 1, wherein the set of linguistic features comprises phoneme-level information.
4. The method of claim 1, further comprising comparing the measure of spoken language skill for the user to a predetermined threshold.
5. The method of claim 4, further comprising providing feedback to the user based on the comparison of the measure of spoken language skill to the predetermined threshold.
6 The method of claim 1, wherein the measure of spoken language skill is based on one or more of fluency, pronunciation, and grammar.
7. A system for evaluating spoken language skills, comprising:
a microphone for receiving an audio signal comprising spoken language from a user;
a processor configured to perform a spectral analysis on the audio signal to obtain a set of spectral features, and to apply a trained acoustic model to the set of spectral features to obtain a set of linguistic features;
a memory for storing the trained acoustic model and instructions for analysing the set of linguistic features to determine a measure of spoken language skill for the user.
8: The system of claim 7, further comprising (d) a user interface for providing feedback to the user based on the measure of spoken language skill determined by the processor.
9. The system of claim 7, wherein the trained acoustic model is a deep neural network.
10. The system of claim 7, wherein the set of linguistic features comprises phoneme-level information.
| # | Name | Date |
|---|---|---|
| 1 | 202311019722-REQUEST FOR EARLY PUBLICATION(FORM-9) [22-03-2023(online)].pdf | 2023-03-22 |
| 2 | 202311019722-POWER OF AUTHORITY [22-03-2023(online)].pdf | 2023-03-22 |
| 3 | 202311019722-FORM-9 [22-03-2023(online)].pdf | 2023-03-22 |
| 4 | 202311019722-FORM FOR SMALL ENTITY(FORM-28) [22-03-2023(online)].pdf | 2023-03-22 |
| 5 | 202311019722-FORM 1 [22-03-2023(online)].pdf | 2023-03-22 |
| 6 | 202311019722-EVIDENCE FOR REGISTRATION UNDER SSI(FORM-28) [22-03-2023(online)].pdf | 2023-03-22 |
| 7 | 202311019722-EVIDENCE FOR REGISTRATION UNDER SSI [22-03-2023(online)].pdf | 2023-03-22 |
| 8 | 202311019722-EDUCATIONAL INSTITUTION(S) [22-03-2023(online)].pdf | 2023-03-22 |
| 9 | 202311019722-DRAWINGS [22-03-2023(online)].pdf | 2023-03-22 |
| 10 | 202311019722-DECLARATION OF INVENTORSHIP (FORM 5) [22-03-2023(online)].pdf | 2023-03-22 |
| 11 | 202311019722-COMPLETE SPECIFICATION [22-03-2023(online)].pdf | 2023-03-22 |