Abstract: MACHINE TRANSLATION TO ENABLE CROSS-LINGUAL COMMUNICATION BETWEEN DHUNDHARI AND HINDI Abstract A machine translator enables indigenous language data as source to be translated in Hindi. The machine translation process involves analysis and comparison against prestored data in a database that includes data terms and definitions of Hindi and Dhundhari. The translation utilizes unsupervised stemmer, rule based stemmer and lemmatizer. The lemmatizer implements two rounds of suffix removal. The two stemmers are rule-based stemmer and unsupervised stemmer. The rule-based stemmer comprises stemming rules and unsupervised stemmer extract stems from the indigenous language data.
1. A machine translator system for translating indigenous language data into Hindi, comprising a database of Hindi and Dhundhari terms and definitions.
2. The machine translator system of claim 1, utilizing unsupervised stemmer, rule-based stemmer, and lemmatizer for the translation process.
3. The machine translator system of claim 2, wherein the lemmatizer implements two rounds of suffix removal to improve translation accuracy.
4. The machine translator system of claim 2, wherein the rule-based stemmer comprises a set of stemming rules for Hindi and Dhundhari.
5. The machine translator system of claim 2, wherein the unsupervised stemmer extracts stems from the indigenous language data to improve translation accuracy.
6. The machine translator system of claim 1, wherein the database includes a comprehensive list of Hindi and Dhundhari terms and definitions.
7. The machine translator system of claim 2, utilizing a comparison process against prestored data in the database to improve translation accuracy.
8. A method for translating indigenous language data into Hindi using a machine translator system, comprising utilizing unsupervised stemmer, rule-based stemmer, and lemmatizer.
9. The method of claim 8, wherein rule-based stemmer identifies patterns in the indigenous language text and reduce inflections to their root form.
10. The method of claim 8, wherein the identified root forms and base forms of the indigenous language text with the data terms and definitions in the database are utilized to determine the translation in Hindi. MACHINE TRANSLATION TO ENABLE CROSS-LINGUAL COMMUNICATION BETWEEN DHUNDHARI AND HINDI Abstract A machine translator enables indigenous language data as source to be translated in Hindi. The machine translation process involves analysis and comparison against prestored data in a database that includes data terms and definitions of Hindi and Dhundhari. The translation utilizes unsupervised stemmer, rule based stemmer and lemmatizer. The lemmatizer implements two rounds of suffix removal. The two stemmers are rule-based stemmer and unsupervised stemmer. The rule-based stemmer comprises stemming rules and unsupervised stemmer extract stems from the indigenous language data. , Claims:Claims :
1. A machine translator system for translating indigenous language data into Hindi, comprising a database of Hindi and Dhundhari terms and definitions.
2. The machine translator system of claim 1, utilizing unsupervised stemmer, rule-based stemmer, and lemmatizer for the translation process.
3. The machine translator system of claim 2, wherein the lemmatizer implements two rounds of suffix removal to improve translation accuracy.
4. The machine translator system of claim 2, wherein the rule-based stemmer comprises a set of stemming rules for Hindi and Dhundhari.
5. The machine translator system of claim 2, wherein the unsupervised stemmer extracts stems from the indigenous language data to improve translation accuracy.
6. The machine translator system of claim 1, wherein the database includes a comprehensive list of Hindi and Dhundhari terms and definitions.
7. The machine translator system of claim 2, utilizing a comparison process against prestored data in the database to improve translation accuracy.
8. A method for translating indigenous language data into Hindi using a machine translator system, comprising utilizing unsupervised stemmer, rule-based stemmer, and lemmatizer.
9. The method of claim 8, wherein rule-based stemmer identifies patterns in the indigenous language text and reduce inflections to their root form.
10. The method of claim 8, wherein the identified root forms and base forms of the indigenous language text with the data terms and definitions in the database are utilized to determine the translation in Hindi.
Description:MACHINE TRANSLATION TO ENABLE CROSS-LINGUAL COMMUNICATION BETWEEN DHUNDHARI AND HINDI
Field of the Invention
[0001] This invention is in the field of indigenous language data collection and prospecting. More specially towards system and method pertaining to machine translation process to enable cross-lingual communication between Dhundhari and Hindi.
Background
[0002] The background description includes information that may be useful in understanding the present invention. It is not an admission that any of the information provided herein is prior art or relevant to the presently claimed invention, or that any publication specifically or implicitly referenced is prior art.
[0003] Dhundhari is a language spoken by approximately 2 million people in the Dhundhari region of Rajasthan, India. While Dhundhari has its own unique grammar and vocabulary, it is closely related to the Hindi language and shares many similarities with it.
[0004] Machine translation (MT) is the process of using advance computing techology to automatically translate text from one language to another. MT has become an important tool for communication and information sharing in a globalized world, particularly in domains such as business, diplomacy, and academia. However, the quality of MT output can vary widely depending on the complexity of the languages involved and the algorithms and models used for translation. Patent literature disclosed technological solutions related to language translation. Few of exemplary documents are cited below.
[0005] EP3161667A2 (by Google LLC) - A computer-implemented technique can include receiving, at a server from a mobile computing device, the server having one or more processors, an image including a text. The technique can include obtaining, at the server, optical character recognition (OCR) text corresponding to the text, the OCR text having been obtained by performing OCR on the image. The technique can include identifying, at the server, non-textual context information from the image, the non-textual context information (i) representing context information other than the text itself and (ii) being indicative of a context of the image. The technique can include based on the non-textual context information, obtaining, at the server, a translation of the OCR text to a target language to obtain a translated OCR text. The technique can include outputting, from the server to the mobile computing device, the translated OCR text.
[0006] US20220318523 (by IBM) - A computer system extracts clauses using machine translation. An input sentence in a source language is translated into a translated sentence in a target language using a trained machine translation model, wherein the trained machine translation model inserts a grammatical indicator into a position of the translated sentence that identifies a dependent clause. The input sentence and the translated sentence are aligned to determine a position in the input sentence that corresponds to the position of the grammatical indicator in the translated sentence. The dependent clause is extracted, in the source language, from the input sentence based on the determined position in the input sentence. Embodiments of the present invention further include a method and program product for clause extraction using machine translation in substantially the same manner described above.
[0007] CN215219699 (by SHENZHEN WOOASK TECH CO LTD) - The utility model provides a multi-language small translator, which relates to the technical field of translators, and comprises a machine body, a display screen, a clamping groove, a mounting plate, a side plate, a clamping block, a fixing component and a supporting component, the display screen is fixedly mounted on the side wall of one side of the machine body, the clamping groove is formed in the side wall of one side of the machine body, and the mounting plate is slidably connected to the side wall of one side, far away from the display screen, of the machine body; a clamping block is fixedly mounted on the side wall of one side of the side plate, and a fixing assembly is slidably connected to the side wall of one side of the mounting plate. By arranging the machine body, the clamping grooves, the mounting plates, the side plates and the clamping blocks, the function of protecting the translating machine is achieved, the translating machine is prevented from being damaged due to collision, economic losses are reduced, meanwhile, the translating machine can be rapidly and conveniently fixed by arranging the fixing blocks, the connecting blocks, the connecting grooves, the limiting blocks and the limiting grooves, and the translating machine is convenient to use. The translation machine is prevented from being separated from the protection device, the installation speed of the protection device is increased, and the working efficiency of installation personnel is improved.
[0008] The task of translating Dhundhari to Hindi poses several challenges for MT systems. One of the main challenges is the lack of resources, such as parallel corpora, for training and evaluating MT models. Another challenge is the syntactic and morphological differences between the two languages, which can lead to errors and inconsistencies in the translation output. Thus, there is remain need of newer technology to overcome limitations of current techniques.
Summary
[0009] The following presents a simplified summary of various aspects of this disclosure in order to provide a basic understanding of such aspects. This summary is not an extensive overview of all contemplated aspects and is intended to neither identify key or critical elements nor delineate the scope of such aspects. Its purpose is to present some concepts of this disclosure in a simplified form as a prelude to the more detailed description that is presented later.
[00010] This invention is in the field of indigenous language data collection and prospecting. More specially towards system and method pertaining to machine translation process to enable cross-lingual communication between Dhundhari and Hindi.
[00011] The following paragraphs provide additional support for the claims of the subject application.
[00012] Embodiments of the present disclosure may include a machine translator system for translating indigenous language data into Hindi, including a database of Hindi and Dhundhari terms and definitions. In some embodiments, the machine translator system, utilizing unsupervised stemmer, rule-based stemmer, and lemmatizer for the translation process. In some embodiments, the lemmatizer implements two rounds of suffix removal to improve translation accuracy.
[00013] In some embodiments, the rule-based stemmer may include a set of stemming rules for Hindi and Dhundhari. In some embodiments, the unsupervised stemmer extracts stems from the indigenous language data to improve translation accuracy. In some embodiments, the machine translator system, utilizing a comparison process against prestored data in the database to improve translation accuracy. In some embodiments, the database includes a comprehensive list of Hindi and Dhundhari terms and definitions.
[00014] Embodiments of the present disclosure may also include a method for translating indigenous language data into Hindi using a machine translator system, including utilizing unsupervised stemmer, rule-based stemmer, and lemmatizer. Embodiments may also include rule-based stemmer identifies patterns in the indigenous language text and reduce inflections to their root form. In some embodiments, the identified root forms and base forms of the indigenous language text with the data terms and definitions in the database may be utilized to determine the translation in Hindi.
Brief Description of the Drawings
[00015] The features and advantages of the present disclosure would be more clearly understood from the following description taken in conjunction with the accompanying drawings in which:
[00016] FIG. 1 is a block diagram illustrating a machine translator system, according to some embodiments of the present disclosure.
[00017] FIG. 2 is a flowchart illustrating a method for translating indigenous language data to Hindi, according to some embodiments of the present disclosure.
Detailed Description
[00018] In the following detailed description of the invention, reference is made to the accompanying drawings that form a part hereof, and in which is shown, by way of illustration, specific embodiments in which the invention may be practiced. In the drawings, like numerals describe substantially similar components throughout the several views. These embodiments are described in sufficient detail to claim those skilled in the art to practice the invention. Other embodiments may be utilized and structural, logical, and electrical changes may be made without departing from the scope of the present invention. The following detailed description is, therefore, not to be taken in a limiting sense, and the scope of the present invention is defined only by the appended claims and equivalents thereof.
[00019] The use of the terms “a” and “an” and “the” and “at least one” and similar referents in the context of describing the invention (especially in the context of the following claims) are to be construed to cover both the singular and the plural, unless otherwise indicated herein or clearly contradicted by context. The use of the term “at least one” followed by a list of one or more items (for example, “at least one of A and B”) is to be construed to mean one item selected from the listed items (A or B) or any combination of two or more of the listed items (A and B), unless otherwise indicated herein or clearly contradicted by context. The terms “comprising,” “having,” “including,” and “containing” are to be construed as open-ended terms (i.e., meaning “including, but not limited to,”) unless otherwise noted. Recitation of ranges of values herein are merely intended to serve as a shorthand method of referring individually to each separate value falling within the range, unless otherwise indicated herein, and each separate value is incorporated into the specification as if it were individually recited herein. All methods described herein can be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by context. The use of any and all examples, or exemplary language (e.g., “such as”) provided herein, is intended merely to better illuminate the invention and does not pose a limitation on the scope of the invention unless otherwise claimed. No language in the specification should be construed as indicating any non-claimed element as essential to the practice of the invention.
[00020] This invention is in the field of indigenous language data collection and prospecting. More specially towards system and method pertaining to machine translation process to enable cross-lingual communication between Dhundhari and Hindi.
[00021] In accordance with some implementations of the current disclosure, the machine translator system 100 seen in FIG. 1 is shown in the form of a block diagram. A database 110 including terminology and meanings written in Hindi and Dhundhari may be included in some implementations of the machine translator system 100. In some implementations, the translation process is carried out by making use of an unsupervised stemmer, a rule-based stemmer, and a lemmatizer. To achieve a higher level of precision in translation, the lemmatizer, in certain implementations, may use a two-stage process of suffix elimination. The rule-based stemmer may, in some implementations, consist of a collection of stemming rules for the languages of Hindi and Dhundhari. In some implementations, the unsupervised stemmer extracts may be derived from the data of the indigenous language in order to achieve more accuracy in translating. In some implementations, the accuracy of the translation is improved by using a process of comparison against data that has been pre-stored in the database 110. The database 110 may, according to certain implementations, include a complete collection of Hindi and Dhundhari words together with their respective meanings.
[00022] A method for translating data in an indigenous language is shown in flowchart form in FIG. 2, which presents the technique in accordance with various implementations of the present disclosure. Using an unsupervised stemmer, a rule-based stemmer, and a lemmatizer are all possible steps that might be included in the technique at step 210, depending on the specific embodiment. The Rule-based stemmer may, in certain implementations, recognise patterns in the text of the indigenous language and reduce inflections to their basic form. It is possible that the translation into Hindi may be determined by combining the indigenous language text's recognised root forms and base forms with the data words and definitions included in the database. This is one of the many possible implementations.
[00023] A machine translation system for translating data from an indigenous language into Hindi may be included in certain embodiments of the present disclosure. Such a system may comprise a database of terminology and meanings in both Hindi and Dhundhari. In certain implementations of the machine translator system, the translation process is carried out with the assistance of an unsupervised stemmer, a rule-based stemmer, and a lemmatizer. To achieve a higher level of accuracy in translation, the lemmatizer, in certain implementations, includes not one but two rounds of suffix elimination.
[00024] The rule-based stemmer may, in some implementations, consist of a collection of stemming rules for the languages of Hindi and Dhundhari. In certain implementations, an unsupervised stemmer is used to increase the accuracy of translation by extracting stems from data pertaining to the native language. The accuracy of the translation may be improved in certain implementations of the machine translator system by using a comparison procedure against data that has been pre-stored in the database. In some implementations, the database has a complete collection of Hindi and Dhundhari words together with their respective meanings.
[00025] A method for translating data from an indigenous language into Hindi using a machine translator may also be included in certain embodiments of the current disclosure. This method may include the use of an unsupervised stemmer, a rule-based stemmer, and a lemmatizer, among other components. In other embodiments, a rule-based stemmer may be used to recognise patterns in the text of the indigenous language and reduce inflections to their basic form. It is possible that the translation into Hindi may be determined by combining the indigenous language text's recognised root forms and base forms with the data words and definitions included in the database. This is one of the many possible implementations.
[00026] Example embodiments herein have been described above with reference to block diagrams and flowchart illustrations of methods and apparatuses. It will be understood that each block of the block diagrams and flowchart illustrations, and combinations of blocks in the block diagrams and flowchart illustrations, respectively, can be implemented by various means including hardware, software, firmware, and a combination thereof. For example, in one embodiment, each block of the block diagrams and flowchart illustrations, and combinations of blocks in the block diagrams and flowchart illustrations can be implemented by computer program instructions. These computer program instructions may be loaded onto a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions which execute on the computer or other programmable data processing apparatus create means for implementing the functions specified in the flowchart block or blocks.
[00027] Throughout the present disclosure, the term ‘processing means’ or ‘microprocessor’ or ‘processor’ or ‘processors’ includes, but is not limited to, a general purpose processor (such as, for example, a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a microprocessor implementing other types of instruction sets, or a microprocessor implementing a combination of types of instruction sets) or a specialized processor (such as, for example, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), or a network processor).
[00028] The term “non-transitory storage device” or “storage” or “memory,” as used herein relates to a random access memory, read only memory and variants thereof, in which a computer can store data or software for any duration.
[00029] Operations in accordance with a variety of aspects of the disclosure is described above would not have to be performed in the precise order described. Rather, various steps can be handled in reverse order or simultaneously or not at all.
[00030] While several implementations have been described and illustrated herein, a variety of other means and/or structures for performing the function and/or obtaining the results and/or one or more of the advantages described herein may be utilized, and each of such variations and/or modifications is deemed to be within the scope of the implementations described herein. More generally, all parameters, dimensions, materials, and configurations described herein are meant to be exemplary and that the actual parameters, dimensions, materials, and/or configurations will depend upon the specific application or applications for which the teachings is/are used. Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific implementations described herein. It is, therefore, to be understood that the foregoing implementations are presented by way of example only and that, within the scope of the appended claims and equivalents thereto, implementations may be practiced otherwise than as specifically described and claimed. Implementations of the present disclosure are directed to each individual feature, system, article, material, kit, and/or method described herein. In addition, any combination of two or more such features, systems, articles, materials, kits, and/or methods, if such features, systems, articles, materials, kits, and/or methods are not mutually inconsistent, is included within the scope of the present disclosure.
Claims
I/We Claim:
1. A machine translator system for translating indigenous language data into Hindi, comprising a database of Hindi and Dhundhari terms and definitions.
2. The machine translator system of claim 1, utilizing unsupervised stemmer, rule-based stemmer, and lemmatizer for the translation process.
3. The machine translator system of claim 2, wherein the lemmatizer implements two rounds of suffix removal to improve translation accuracy.
4. The machine translator system of claim 2, wherein the rule-based stemmer comprises a set of stemming rules for Hindi and Dhundhari.
5. The machine translator system of claim 2, wherein the unsupervised stemmer extracts stems from the indigenous language data to improve translation accuracy.
6. The machine translator system of claim 1, wherein the database includes a comprehensive list of Hindi and Dhundhari terms and definitions.
7. The machine translator system of claim 2, utilizing a comparison process against prestored data in the database to improve translation accuracy.
8. A method for translating indigenous language data into Hindi using a machine translator system, comprising utilizing unsupervised stemmer, rule-based stemmer, and lemmatizer.
9. The method of claim 8, wherein rule-based stemmer identifies patterns in the indigenous language text and reduce inflections to their root form.
10. The method of claim 8, wherein the identified root forms and base forms of the indigenous language text with the data terms and definitions in the database are utilized to determine the translation in Hindi.
MACHINE TRANSLATION TO ENABLE CROSS-LINGUAL COMMUNICATION BETWEEN DHUNDHARI AND HINDI
Abstract
A machine translator enables indigenous language data as source to be translated in Hindi. The machine translation process involves analysis and comparison against prestored data in a database that includes data terms and definitions of Hindi and Dhundhari. The translation utilizes unsupervised stemmer, rule based stemmer and lemmatizer. The lemmatizer implements two rounds of suffix removal. The two stemmers are rule-based stemmer and unsupervised stemmer. The rule-based stemmer comprises stemming rules and unsupervised stemmer extract stems from the indigenous language data. , Claims:Claims
I/We Claim:
1. A machine translator system for translating indigenous language data into Hindi, comprising a database of Hindi and Dhundhari terms and definitions.
2. The machine translator system of claim 1, utilizing unsupervised stemmer, rule-based stemmer, and lemmatizer for the translation process.
3. The machine translator system of claim 2, wherein the lemmatizer implements two rounds of suffix removal to improve translation accuracy.
4. The machine translator system of claim 2, wherein the rule-based stemmer comprises a set of stemming rules for Hindi and Dhundhari.
5. The machine translator system of claim 2, wherein the unsupervised stemmer extracts stems from the indigenous language data to improve translation accuracy.
6. The machine translator system of claim 1, wherein the database includes a comprehensive list of Hindi and Dhundhari terms and definitions.
7. The machine translator system of claim 2, utilizing a comparison process against prestored data in the database to improve translation accuracy.
8. A method for translating indigenous language data into Hindi using a machine translator system, comprising utilizing unsupervised stemmer, rule-based stemmer, and lemmatizer.
9. The method of claim 8, wherein rule-based stemmer identifies patterns in the indigenous language text and reduce inflections to their root form.
10. The method of claim 8, wherein the identified root forms and base forms of the indigenous language text with the data terms and definitions in the database are utilized to determine the translation in Hindi.
| # | Name | Date |
|---|---|---|
| 1 | 202311020095-REQUEST FOR EARLY PUBLICATION(FORM-9) [22-03-2023(online)].pdf | 2023-03-22 |
| 2 | 202311020095-POWER OF AUTHORITY [22-03-2023(online)].pdf | 2023-03-22 |
| 3 | 202311020095-FORM-9 [22-03-2023(online)].pdf | 2023-03-22 |
| 4 | 202311020095-FORM FOR SMALL ENTITY(FORM-28) [22-03-2023(online)].pdf | 2023-03-22 |
| 5 | 202311020095-FORM 1 [22-03-2023(online)].pdf | 2023-03-22 |
| 6 | 202311020095-EVIDENCE FOR REGISTRATION UNDER SSI(FORM-28) [22-03-2023(online)].pdf | 2023-03-22 |
| 7 | 202311020095-EVIDENCE FOR REGISTRATION UNDER SSI [22-03-2023(online)].pdf | 2023-03-22 |
| 8 | 202311020095-EDUCATIONAL INSTITUTION(S) [22-03-2023(online)].pdf | 2023-03-22 |
| 9 | 202311020095-DRAWINGS [22-03-2023(online)].pdf | 2023-03-22 |
| 10 | 202311020095-DECLARATION OF INVENTORSHIP (FORM 5) [22-03-2023(online)].pdf | 2023-03-22 |
| 11 | 202311020095-COMPLETE SPECIFICATION [22-03-2023(online)].pdf | 2023-03-22 |