Abstract: System and method for Identifying language in song Abstract A system for determining the language used in a song is included in certain embodiments of the present disclosure. This system may comprise an audio input device that is able to receive the audio file that contains the song. A feature extraction module that is set up to extract features from an audio file is another component that may be included in embodiments. A language identification module that is able to receive the extracted characteristics and determine the language of the song based on a trained language identification model is another component that may be included in certain embodiments. Additionally, a display device that is capable of displaying the language that has been detected.
1. A system for identifying the language in a song comprising: an audio input device configured to receive the audio file of the song; a feature extraction module configured to extract features from the audio file; a language identification module configured to receive the extracted features and identify the language of the song based on a trained language identification model; and a display device configured to display the identified language.
2. The system of claim 1, wherein the audio input device comprises a microphone, a music player, or any other suitable audio acquisition device.
3. The system of claim 1, wherein the feature extraction module comprises Mel-frequency cepstral coefficients (MFCC), spectral centroid, or any other suitable feature extraction algorithm.
4. The system of claim 1, wherein the language identification module comprises a machine learning model trained on a dataset of audio files in different languages.
5. A method for identifying the language in a song comprising: receiving the audio file of the song; extracting features from the audio file; identifying the language of the song based on a trained language identification model; displaying the identified language.
6. The method of claim 5, wherein the audio file of the song is received from a microphone, a music player, or any other suitable audio acquisition device.
7. The method of claim 5, wherein extracting features from the audio file comprises using Mel-frequency cepstral coefficients (MFCC), spectral centroid, or any other suitable feature extraction algorithm.
8. The method of claim 5, wherein identifying the language of the song comprises using a machine learning model trained on a dataset of audio files in different languages.
9. The method of claim 5, wherein the identified language is displayed on a screen or output to an audio device.
10. The method of claim 5, wherein the system is integrated into a music player or mobile device for automatic language identification of songs. System and method for Identifying language in song Abstract A system for determining the language used in a song is included in certain embodiments of the present disclosure. This system may comprise an audio input device that is able to receive the audio file that contains the song. A feature extraction module that is set up to extract features from an audio file is another component that may be included in embodiments. A language identification module that is able to receive the extracted characteristics and determine the language of the song based on a trained language identification model is another component that may be included in certain embodiments. Additionally, a display device that is capable of displaying the language that has been detected. , Claims:Claims :
1. A system for identifying the language in a song comprising: an audio input device configured to receive the audio file of the song; a feature extraction module configured to extract features from the audio file; a language identification module configured to receive the extracted features and identify the language of the song based on a trained language identification model; and a display device configured to display the identified language.
2. The system of claim 1, wherein the audio input device comprises a microphone, a music player, or any other suitable audio acquisition device.
3. The system of claim 1, wherein the feature extraction module comprises Mel-frequency cepstral coefficients (MFCC), spectral centroid, or any other suitable feature extraction algorithm.
4. The system of claim 1, wherein the language identification module comprises a machine learning model trained on a dataset of audio files in different languages.
5. A method for identifying the language in a song comprising: receiving the audio file of the song; extracting features from the audio file; identifying the language of the song based on a trained language identification model; displaying the identified language.
6. The method of claim 5, wherein the audio file of the song is received from a microphone, a music player, or any other suitable audio acquisition device.
7. The method of claim 5, wherein extracting features from the audio file comprises using Mel-frequency cepstral coefficients (MFCC), spectral centroid, or any other suitable feature extraction algorithm.
8. The method of claim 5, wherein identifying the language of the song comprises using a machine learning model trained on a dataset of audio files in different languages.
9. The method of claim 5, wherein the identified language is displayed on a screen or output to an audio device.
10. The method of claim 5, wherein the system is integrated into a music player or mobile device for automatic language identification of songs.
Description:System and method for Identifying language in song
Field of the Invention
[0001] The present invention relates generally to smart audio player, and more particularly to providing a digital assistant's auditory-based interface for detection of language of song.
Background
[0002] The background description includes information that may be useful in understanding the present invention. It is not an admission that any of the information provided herein is prior art or relevant to the presently claimed invention, or that any publication specifically or implicitly referenced is prior art.
[0003] Audio analysis has many applications, such as speech recognition, music information retrieval, speaker identification, and audio scene analysis. The goal of audio file analysis is to extract meaningful information from audio signals and use it for various purposes.
[0004] In recent years, research in audio file analysis has focused on developing systems that can analyze audio files in real-time.
[0005] Various technological solutions (e.g., neural latent variable model for spoken language understanding, Exploitation of language identification of media file data in speech dialog systems, etc.) for identifying the language in a song are disclosed in patent literature.
[0006] The EP1909263A1 (By- Harman Becker Automotive Systems GmbH) relates to a method for outputting a synthesized speech signal corresponding to an orthographic string stored in a media file comprising audio data, comprising the steps of analyzing the audio data to determine at least one candidate for a language of the orthographic string, estimating a phonetic representation of the orthographic string based on the determined at least one candidate for a language and synthesizing a speech signal based on the estimated phonetic representation of the orthographic string. The invention also relates to a media player incorporating such a method for a estimating phonetic representation for song and album titles as well as artists' names for speech recognition. Furthermore, the invention relates to the choice of an appropriate speech recognizer for automatically transcribing the lyrics of songs by using audio-based language estimates.
[0007] The US8666727B2 (By- Harman Becker Automotive Systems GmbH) relates to a voice-controlled data system may include a data storage unit including media files having associated file identification data, and a vocabulary generating unit generating phonetic data corresponding to the file identification data, the phonetic data being supplied to a speech recognition unit as a recognition vocabulary, where one of the media files may be selected according to a recognized speech control command on the basis of the generated phonetic data, where the file identification data include a language identification part for identifying the language of the file identification data, and where the vocabulary generating unit generates the phonetic data for the file identification data of a media file based on its language identification part.
[0008] The US9911413B1 (By- Amazon Technologies Inc) relates to a linguist classifier, for instance intent or slot classifier, is updated using data with only partial annotation indicating overall correctness rather that specific correct intent or slot values, which are treated as “latent” (i.e., unknown) variables. Full annotation of the data is not required. A small amount of fully annotated data may be combined with a substantially larger amount of partially annotated data to update the linguistic classifier. In a specific implementation, the linguistic classifier is a neural network and the weights are trained using a reinforcement learning approach.
[0009] One of the primary challenges in audio file analysis is the complexity of audio signals. Audio signals can contain a wide range of frequencies and complex patterns that can be difficult to analyze. Additionally, audio signals can be affected by environmental noise, distortion, and other factors that can affect the accuracy of analysis. Thus, there is need of technical advancement in this domain.
Summary
[00010] The present invention relates generally to smart audio player, and more particularly to providing a digital assistant's auditory-based interface for detection of language of song.
[00011] The following presents a simplified summary of various aspects of this disclosure in order to provide a basic understanding of such aspects. This summary is not an extensive overview of all contemplated aspects, and is intended to neither identify key or critical elements nor delineate the scope of such aspects. Its purpose is to present some concepts of this disclosure in a simplified form as a prelude to the more detailed description that is presented later.
[00012] The following paragraphs provide additional support for the claims of the subject application.
[00013] Embodiments of the present disclosure may include a system for identifying the language in a song including an audio input device configured to receive the audio file of the song. Embodiments may also include a feature extraction module configured to extract features from the audio file. Embodiments may also include a language identification module configured to receive the extracted features and identify the language of the song based on a trained language identification model. Embodiments may also include a display device configured to display the identified language.
[00014] In some embodiments, the audio input device may include a microphone, a music player, or any other suitable audio acquisition device. In some embodiments, the feature extraction module may include Mel-frequency cepstral coefficients (MFCC), spectral centroid, or any other suitable feature extraction algorithm. In some embodiments, the language identification module may include a machine learning model trained on a dataset of audio files in different languages.
[00015] Embodiments may also include identifying the language of the song may include using a machine learning model trained on a dataset of audio files in different languages. In some embodiments, the identified language may be displayed on a screen or output to an audio device. In some embodiments, the system may be integrated into a music player or mobile device for automatic language identification of songs.
Brief Description of the Drawings
[00016] The features and advantages of the present disclosure would be more clearly understood from the following description taken in conjunction with the accompanying drawings in which:
[00017] FIG. 1 is a block diagram illustrating a system for identifying the language in a song, according to some embodiments of the present disclosure.
[00018] FIG. 2 is a detailed block diagram further illustrating the system (from FIG. 1) for identifying the language in a song, according to some embodiments of the present disclosure.
[00019] FIG. 3 is a modified block diagram further illustrating the system (from FIG. 1) for identifying the language in a song, according to some embodiments of the present disclosure.
[00020] FIG. 4 is a flowchart illustrating a method for identifying the language in a song, according to some embodiments of the present disclosure.
Detailed Description
[00021] In the following detailed description of the invention, reference is made to the accompanying drawings that form a part hereof, and in which is shown, by way of illustration, specific embodiments in which the invention may be practiced. In the drawings, like numerals describe substantially similar components throughout the several views. These embodiments are described in sufficient detail to claim those skilled in the art to practice the invention. Other embodiments may be utilized and structural, logical, and electrical changes may be made without departing from the scope of the present invention. The following detailed description is, therefore, not to be taken in a limiting sense, and the scope of the present invention is defined only by the appended claims and equivalents thereof.
[00022] The use of the terms a and an and the and at least one and similar referents in the context of describing the invention (especially in the context of the following claims) are to be construed to cover both the singular and the plural, unless otherwise indicated herein or clearly contradicted by context. The use of the term at least one followed by a list of one or more items (for example, at least one of A and B) is to be construed to mean one item selected from the listed items (A or B) or any combination of two or more of the listed items (A and B), unless otherwise indicated herein or clearly contradicted by context. The terms comprising, having, including, and containing are to be construed as open-ended terms (i.e., meaning including, but not limited to,) unless otherwise noted. Recitation of ranges of values herein are merely intended to serve as a shorthand method of referring individually to each separate value falling within the range, unless otherwise indicated herein, and each separate value is incorporated into the specification as if it were individually recited herein. All methods described herein can be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by context. The use of any and all examples, or exemplary language (e.g., such as) provided herein, is intended merely to better illuminate the invention and does not pose a limitation on the scope of the invention unless otherwise claimed. No language in the specification should be construed as indicating any non-claimed element as essential to the practice of the invention.
[00023] The present invention relates generally to smart audio player, and more particularly to providing a digital assistant's auditory-based interface for detection of language of song.
[00024] FIG. 1 is a block diagram that describes a system 100, according to some embodiments of the present disclosure. In some embodiments, the system 100 may include an audio input device 110 configured to receive the audio file of the song, a feature extraction module 120 configured to extract features from the audio file, a language identification module 130 configured to receive the extracted features and identify the language of the song based on a trained language identification model, and a display device 140 configured to display the identified language. In some embodiments, the language identification module 130 may include a machine learning model trained on a dataset of audio files in different languages.
[00025] FIG. 2 is a detailed block diagram that further describes the system 100 from FIG. 1, according to some embodiments of the present disclosure. In some embodiments, the audio input device 110 may include a microphone 212, a music player 214, and a suitable audio acquisition device 216.
[00026] FIG. 3 is a modified block diagram that further describes the system 100 from FIG. 1, according to some embodiments of the present disclosure. In some embodiments, the feature extraction module 120 may include any other suitable feature extraction 324 algorithm. The feature extraction module 120 may also include Mel-frequency cepstral coefficients 322 (MFCC), spectral centroid, etc.
[00027] FIG. 4 is a flowchart that describes a method for identifying the language in a song, according to some embodiments of the present disclosure. In some embodiments, at 410, the method may include receiving the audio file of the song. At 420, the method may include extracting features from the audio file. At 430, the method may include identifying the language of the song based on a trained language identification model. At 440, the method may include displaying the identified language. In some embodiments, the audio file of the song may be received from a microphone, a music player, or any other suitable audio acquisition device.
[00028] In some embodiments, features can be extracted from the audio file comprises using Mel-frequency cepstral coefficients (MFCC), spectral centroid, or any other suitable feature extraction algorithm. In some embodiments, identifying the language of the song includes utilization of a machine learning model trained on a dataset of audio files in different languages. In some embodiments, the identified language may be displayed on a screen or output to an audio device. In some embodiments, the system may be integrated into a music player or mobile device for automatic language identification of songs.
[00029] A system for determining the language used in a song may be included in certain embodiments of the current disclosure. This system may comprise an audio input device that is able to receive the audio file that contains the song. A feature extraction module that is set up to extract features from an audio file is another component that may be included in embodiments. A language identification module that is able to receive the extracted characteristics and determine the language of the song based on a trained language identification model is another component that may be included in certain embodiments. In certain embodiments, there is additionally a display device that is capable of displaying the language that has been detected.
[00030] In some implementations, the audio input device may take the form of a microphone, a music player, or any other audio acquisition device that is deemed appropriate. Mel-frequency cepstral coefficients (MFCC), also known as spectral centroids, or any other feature extraction technique that is deemed appropriate may be included in the feature extraction module of some implementations. A machine learning model that has been trained on a dataset consisting of audio recordings in a variety of languages could be included in the language identification module of some implementations.
[00031] A method for determining the language used in a song may also be included in certain embodiments of the present disclosure. This method may involve the step of receiving the song's audio file. It's possible that certain embodiments will include extracting features from an audio file. Identifying the music's language using a trained language identification model is another possible embodiment. Displaying the language that was discovered is another possible aspect of embodiments.
[00032] In some implementations, the audio file of the song might be obtained through a microphone, a music player, or any other audio acquisition device that is deemed appropriate. In certain embodiments, there is additionally a step that involves extracting features from an audio file. This step may include using Mel-frequency cepstral coefficients (MFCC), spectral centroid, or any other feature extraction technique that is deemed appropriate.
[00033] A machine learning model that has been trained on a dataset including audio files in a variety of languages may be used in some embodiments to determine the language of the song. It is possible that the recognised language will be shown on a screen or sent to an audio output device in some implementations. For the purpose of performing automated language recognition of songs, the system may, in certain implementations be included into a music player or a mobile device.
[00034] Example embodiments herein have been described above with reference to block diagrams and flowchart illustrations of methods and apparatuses. It will be understood that each block of the block diagrams and flowchart illustrations, and combinations of blocks in the block diagrams and flowchart illustrations, respectively, can be implemented by various means including hardware, software, firmware, and a combination thereof. For example, in one embodiment, each block of the block diagrams and flowchart illustrations, and combinations of blocks in the block diagrams and flowchart illustrations can be implemented by computer program instructions. These computer program instructions may be loaded onto a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions which execute on the computer or other programmable data processing apparatus create means for implementing the functions specified in the flowchart block or blocks.
[00035] Throughout the present disclosure, the term ‘processing means’ or ‘microprocessor’ or ‘processor’ or ‘processors’ includes, but is not limited to, a general purpose processor (such as, for example, a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a microprocessor implementing other types of instruction sets, or a microprocessor implementing a combination of types of instruction sets) or a specialized processor (such as, for example, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), or a network processor).
[00036] The term non-transitory storage device or storage or memory, as used herein relates to a random access memory, read only memory and variants thereof, in which a computer can store data or software for any duration.
[00037] Operations in accordance with a variety of aspects of the disclosure is described above would not have to be performed in the precise order described. Rather, various steps can be handled in reverse order or simultaneously or not at all.
[00038] While several implementations have been described and illustrated herein, a variety of other means and/or structures for performing the function and/or obtaining the results and/or one or more of the advantages described herein may be utilized, and each of such variations and/or modifications is deemed to be within the scope of the implementations described herein. More generally, all parameters, dimensions, materials, and configurations described herein are meant to be exemplary and that the actual parameters, dimensions, materials, and/or configurations will depend upon the specific application or applications for which the teachings is/are used. Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific implementations described herein. It is, therefore, to be understood that the foregoing implementations are presented by way of example only and that, within the scope of the appended claims and equivalents thereto, implementations may be practiced otherwise than as specifically described and claimed. Implementations of the present disclosure are directed to each individual feature, system, article, material, kit, and/or method described herein. In addition, any combination of two or more such features, systems, articles, materials, kits, and/or methods, if such features, systems, articles, materials, kits, and/or methods are not mutually inconsistent, is included within the scope of the present disclosure.
Claims
I/We Claim:
1. A system for identifying the language in a song comprising:
an audio input device configured to receive the audio file of the song;
a feature extraction module configured to extract features from the audio file;
a language identification module configured to receive the extracted features and identify the language of the song based on a trained language identification model; and
a display device configured to display the identified language.
2. The system of claim 1, wherein the audio input device comprises a microphone, a music player, or any other suitable audio acquisition device.
3. The system of claim 1, wherein the feature extraction module comprises Mel-frequency cepstral coefficients (MFCC), spectral centroid, or any other suitable feature extraction algorithm.
4. The system of claim 1, wherein the language identification module comprises a machine learning model trained on a dataset of audio files in different languages.
5. A method for identifying the language in a song comprising:
receiving the audio file of the song;
extracting features from the audio file;
identifying the language of the song based on a trained language identification model;
displaying the identified language.
6. The method of claim 5, wherein the audio file of the song is received from a microphone, a music player, or any other suitable audio acquisition device.
7. The method of claim 5, wherein extracting features from the audio file comprises using Mel-frequency cepstral coefficients (MFCC), spectral centroid, or any other suitable feature extraction algorithm.
8. The method of claim 5, wherein identifying the language of the song comprises using a machine learning model trained on a dataset of audio files in different languages.
9. The method of claim 5, wherein the identified language is displayed on a screen or output to an audio device.
10. The method of claim 5, wherein the system is integrated into a music player or mobile device for automatic language identification of songs.
System and method for Identifying language in song
Abstract
A system for determining the language used in a song is included in certain embodiments of the present disclosure. This system may comprise an audio input device that is able to receive the audio file that contains the song. A feature extraction module that is set up to extract features from an audio file is another component that may be included in embodiments. A language identification module that is able to receive the extracted characteristics and determine the language of the song based on a trained language identification model is another component that may be included in certain embodiments. Additionally, a display device that is capable of displaying the language that has been detected. , Claims:Claims
I/We Claim:
1. A system for identifying the language in a song comprising:
an audio input device configured to receive the audio file of the song;
a feature extraction module configured to extract features from the audio file;
a language identification module configured to receive the extracted features and identify the language of the song based on a trained language identification model; and
a display device configured to display the identified language.
2. The system of claim 1, wherein the audio input device comprises a microphone, a music player, or any other suitable audio acquisition device.
3. The system of claim 1, wherein the feature extraction module comprises Mel-frequency cepstral coefficients (MFCC), spectral centroid, or any other suitable feature extraction algorithm.
4. The system of claim 1, wherein the language identification module comprises a machine learning model trained on a dataset of audio files in different languages.
5. A method for identifying the language in a song comprising:
receiving the audio file of the song;
extracting features from the audio file;
identifying the language of the song based on a trained language identification model;
displaying the identified language.
6. The method of claim 5, wherein the audio file of the song is received from a microphone, a music player, or any other suitable audio acquisition device.
7. The method of claim 5, wherein extracting features from the audio file comprises using Mel-frequency cepstral coefficients (MFCC), spectral centroid, or any other suitable feature extraction algorithm.
8. The method of claim 5, wherein identifying the language of the song comprises using a machine learning model trained on a dataset of audio files in different languages.
9. The method of claim 5, wherein the identified language is displayed on a screen or output to an audio device.
10. The method of claim 5, wherein the system is integrated into a music player or mobile device for automatic language identification of songs.
| # | Name | Date |
|---|---|---|
| 1 | 202311019726-REQUEST FOR EARLY PUBLICATION(FORM-9) [22-03-2023(online)].pdf | 2023-03-22 |
| 2 | 202311019726-POWER OF AUTHORITY [22-03-2023(online)].pdf | 2023-03-22 |
| 3 | 202311019726-FORM-9 [22-03-2023(online)].pdf | 2023-03-22 |
| 4 | 202311019726-FORM FOR SMALL ENTITY(FORM-28) [22-03-2023(online)].pdf | 2023-03-22 |
| 5 | 202311019726-FORM 1 [22-03-2023(online)].pdf | 2023-03-22 |
| 6 | 202311019726-EVIDENCE FOR REGISTRATION UNDER SSI(FORM-28) [22-03-2023(online)].pdf | 2023-03-22 |
| 7 | 202311019726-EVIDENCE FOR REGISTRATION UNDER SSI [22-03-2023(online)].pdf | 2023-03-22 |
| 8 | 202311019726-EDUCATIONAL INSTITUTION(S) [22-03-2023(online)].pdf | 2023-03-22 |
| 9 | 202311019726-DRAWINGS [22-03-2023(online)].pdf | 2023-03-22 |
| 10 | 202311019726-DECLARATION OF INVENTORSHIP (FORM 5) [22-03-2023(online)].pdf | 2023-03-22 |
| 11 | 202311019726-COMPLETE SPECIFICATION [22-03-2023(online)].pdf | 2023-03-22 |