Sign In to Follow Application
View All Documents & Correspondence

Automatic Machine Translation Accuracy And Fluency Evaluation Results Analysis System.

Abstract: The present invention discloses a system (10) for automatic determination of quality of machine translation and a method thereof. The system (10) comprising an input module (20) configured to input data. The data comprising a source data in any language, a target data in another language, wherein the target data are machine translated from the source data and a reference data obtained by human translation manually, wherein the reference data human translation of the source data. The system (10) further comprising an evaluation processor (35) configured to evaluate the quality of the target data, wherein the evaluation processor (35) comprises of at least four evaluation engines, a user interface module configured to enable a user to select the at least one evaluation engine, a memory module configured to store the data and the evaluation score by indexing and a processor module (40) configured to automatically determine the quality machine translation using the evaluation of the evaluation processor. The system (10) further comprising a human evaluation module (12) to manually evaluate the quality of the translation on the basis of at least ten predefined criteria by which a human translator can efficiently evaluate machine translation.

Get Free WhatsApp Updates!
Notices, Deadlines & Correspondence

Patent Information

Application #
Filing Date
25 February 2020
Publication Number
36/2021
Publication Type
INA
Invention Field
COMPUTER SCIENCE
Status
Email
drjainbharti@gmail.com
Parent Application

Applicants

BANASTHALI VIDYAPITH
Banasthali, Newai, Tonk, Rajasthan – 304022 India

Inventors

1. Kuldeep Kumar Yogi
712 Ramanujan Banasthali Vidyapith Newai

Claims

1. A system (10) for automatic determination of quality of machine translation, the system (10) comprising: - an input module (20) configured to input data, wherein the data comprising: a source data in any language; a target data in another language, wherein the target data are machine translated from the source data ; and a reference data obtained by human translation manually, wherein the reference data is the human translation of the source data; - a translation processor (30) configured to translate the source data into the target data; - an evaluation processor (35) configured to evaluate the quality of the target data, wherein the evaluation processor (35) comprises of at least four evaluation engines; - a user interface module configured to enable a user to select the at least one evaluation engine; - a memory module configured to store the data and the evaluation score by indexing; - a processor module (40) configured to automatically determine the quality machine translation using the evaluation of the evaluation processor; and - a human evaluation module (12) to manually evaluate the quality of the translation on the basis of at least ten predefined criteria by which a human translator can efficiently evaluate machine translation.

2. The system (10) for automatic evaluation of quality of machine translation as claimed in claim 1, wherein the translation processor (30) comprising TAG, SMT, EBMT and ANAL-GEN engines.

3. The system (10) for automatic evaluation of quality of machine translation as claimed in claim 1, wherein evaluation processor (35) is configured to use at least four automatic tools of evaluation.

4. The system (10) for automatic evaluation of quality of machine translation as claimed in claim 3, wherein the at least four automatic tools of evaluation comprising BLEU, Meteor, F-measure and TER, and wherein each of the BLEU, Meteor, F-measure and TER outputs a value.

5. The system (10) for automatic evaluation of quality of machine translation as claimed in claim 4, wherein the processor module (40) calculates an average score of the value outputted by each of the BLEU, Meteor, F-measure and TER in real-time and/or by the human evaluation (12), and if the average score is greater than a predetermined value, then the processor module (40) determines and the user interface module outputs that the quality of the automatic translation and/or manual translation is high, and if the average score is equal to or less than the predetermined value, then the processor module (40) determines and the user interface module outputs that the quality of the automatic translation and/or manual translation is low.

6. The system (10) for automatic evaluation of quality of machine translation as claimed in claim 5, wherein the predetermined value is 0.5.

7. A method for automatic determination of quality machine translation, the method comprising the steps of: - inputting data (S1) by an input device, wherein the data comprising: a source data in any language; a target data in another language, wherein the target data are machine translated from the source data obtained by any one or more of four translation engines comprising TAG, SMT, EBMT and ANAL-GEN engines; and a reference data obtained by human translation manually, wherein the reference data is the human translation of the source data; - automatically evaluating (S2), by an evaluation processor, the quality of the target data using each of the at least four evaluation engines, said engines comprising at least four automatic tools comprising BLEU, Meteor, F-measure and TER; - selecting (S3), by a user interface, the at least one evaluation engine; - storing (S4), in a memory, the data and the evaluation score by indexing; and - determining (S5), by a processor, the quality of machine translation automatically using the evaluation of the evaluation processor (35) and on the basis of the human evaluation (12)

8. The method for automatic evaluation of quality of machine translation as claimed in claim 7 further comprising outputting a value for each of the BLEU, Meteor, F-measure and TER.

9. The method for automatic evaluation of quality of machine translation as claimed in claim 7 further comprising calculating (S6), by the processor, an average score of the value outputted by each of the BLEU, Meteor, F-measure and TER in real-time, and if the average score of the automatic evaluation and/or the human evaluation is greater than a predetermined value, then determining (S7) by the processor and outputting by the user interface that the quality of the machine translated data is high, and if the average score of the automatic evaluation and/or the human evaluation is equal to or less than the predetermined value, then determining (S8) by the processor and outputting by the user interface that the quality of the machine translated data is low.

10. The method for automatic evaluation of quality of machine translation as claimed in claim 9 further comprising manually evaluating the quality of translation on the basis of at least ten criteria by which a human translator can efficiently evaluate machine translation manually during human evaluation (12).

Specification

The present invention relates to the field of evaluating the quality of machine translation. More particularly, the present invention relates to a system for automatic determination of quality of machine translation and a method thereof.

BACKGROUND OF THE INVENTION
Background description includes information that may be useful in understanding the present invention. It is not an admission that any of the information provided herein is prior art or relevant to the presently claimed invention, or that any publication specifically or implicitly referenced is prior art.

Assessing the quality of a translated sentence in any language by Machine Translation (MT) is an important task before using the sentence at work. At the core of the technology, machine translation substitutes words in one natural language for words in another language. However, this alone cannot produce an understandable translation of text because recognition of whole phrases and their closest counterparts in the target language are needed. Machine translation has improved to translate words, phrases, and sentence structure through gathering data for both the source language and the target language and using this data to generate more linguistically accurate translations. However, evaluating the quality of a machine -translated work is a difficult task.

MT evaluation task was originated in 1956 by ALPAC. 144 Russian book’s sentences were taken for experiment. Some sentences were translated in English by human translators and some from machine. The evaluation result of these sentences found that MT outputs are not fluent and accurate as human translation. Practically, the quality of machine translation can be estimate or evaluate. If the task is short and we want to get 100% accuracy then mostly we do it manually but if it is in bulk then we prefer performing it using a machine. In fact, many machine translation (MT) evaluation algorithms or metrics are developed for checking MT accuracy and fluency. However, the drawbacks of the prior art of conventionally known software are that none of the metrics is good enough to give a reliable output.

Another effort was taken by Skype in 1979 for MT evaluation. He evaluated SYSTRAN system. The objective of evaluation was not set to correctness of translation where as the aim was to find-out cost of corrections. This task has changed the views of peoples about MT evaluation. They were started to think that MT output can be make useful with some minor changes (post-edit). One more experiment performed by Falkedal, 1991 for English-French MT system. MT output and post-edited MT-output compared with human translators and result shown that MT post-edited output were better than MT-outputs.

Church and Hovy tried a different trick in 1993 for MT-evaluation. They gave some MT-outputs to human evaluator and asked them some multiple choice questions related to MT-output’ information. If the human evaluator is able to answering all questions it means the translation is correct or it is less correct. If a translation is expressing the right meaning of source sentence then it is good translation.

Thus, one can only rely on a manual MT-evaluation because it involves a human translator having domain knowledge and skills to understand the language of a source sentence, i.e., sentence or data to be translated, and the language of a target sentence, i.e., language or data of the translated sentence. However, the drawbacks of the conventional way of translation, i.e., manual MT-evaluation, is that it is very costly and time consuming for translation industries, and it involves a lot of man powers, like human translators, from many different domains to produce a reliable output. Thus, the automatic MT-evaluation becomes an essential part of machine translation process and there is a need to improve the results of the automatic machine translations so that the translation can be used at a large scale. The results of machine translation can be improved.

Conventionally, there are two major types of machine translation amongst many other types. The first major type of machine translation is rule-based machine translation, which uses a combination of language and grammar rules plus dictionaries for common words to translate the corpora from one language to another. A corpora is the plural version of corpus which is a collection of written texts, especially the entire works of a particular author or a body of writing on a particular subject. The other major type of machine translation is statistical machine translation. This type of machine translation has no knowledge of language rules. Instead, statistical machine translation “learns” to translate by analyzing large amounts of data for the source and target language including and not limited to, bilingual and monolingual text. Both major types of machine translation produce similar results; however, statistical machine translation delivers more fluent sounding translation, but at the cost of being less consistent with the translations. by developing a tool that can evaluate the quality of a machine-translated sentence. However, there is a need for a tool to decide automatically the quality of machine translation. In sum, there is a need of a tool to decide which machine translation is good enough to produce a reliable output.

In view of the above problems, efficient and effective development, selection, and/or maintenance of machine translation systems require some quantitative measure for evaluating their performance relative to a reference translation. Such a quantitative measure of machine translation performance may be used either by a system developer for tuning a machine translation system being developed or maintained, by a user that requires some measure of performance for choosing between existing machine translators, or by a machine translation system to self-tune its internal system parameters and thereby improve future translation performance.
The conventional tools available for detecting the quality of machine translation are not reliable. Conventionally, using MT-evaluation metric, it is possible to check that translation the quality of machine-translated sentence, i.e., whether the translation is accurate and fluent enough or not. However, each available MT-metric has their own advantages and limitations. Therefore, it is a complex task to decide which metric will be good enough for this purpose.

Prior art US7587307B2 discloses a method and an apparatus for evaluating machine translation quality. The method for computing machine translation performance comprising receiving a sequence of natural language data in a first language, translating the sequence of natural language data to a second language to define a machine translation of the sequence of natural language data comprising symbols, receiving a reference translation of the sequence of natural language data in the second language comprising symbols, computing a sequence kernel that provides a similarity measure between the machine translation and the reference translation based on occurrences of sub sequences that are shared by the machine translation and the reference translation, outputting a signal indicating the similarity measure. The similarity measure accounts for non-contiguous occurrences of subsequences of the selected subsequence length that are shared between the machine translation and the reference translation, in which the non-contiguous subsequences share symbols and comprise a gap of at least one symbol which has been determined not to match a symbol in the gap of the other.

Prior art US9940324B2 discloses a method for evaluating performance of machine translation. The method comprising receiving a first document in a source, translating the first document in the source language to a second document in a target language and performing post processing to obtain a second quantity of information. The method comprising evaluating the second document in the target language and determining whether the comparison of the one or more aspects of the translation of the first document in the source language to the second document in the target language is greater than the predetermined threshold. The method comprising determining that no new models are created and that no degradation has taken place on a translation performance and comparing a first model used for a translation that failed to meet the predetermined threshold to a second model used for a translation that meets the predetermined threshold and determining that the first model and the second model exceeds a percentage of similarity.

Thus, with the known systems and methods of evaluating the quality of machine translation, as mentioned above, a large amount of time, effort, energy, inaccuracy and inefficiency are involved. No prior art considers the use of manual and automatic translations for evaluating the quality of machine translation in accordance to domain information. Thus, the known systems and methods will not produce accurate and reliable result.

In particular, there is not any system and a method, which provides a user-friendly environment in which MT-evaluation can perform on a large data set using various MT-metrics. Further, there is not any system and a method that gives a reliable evaluation of the machine translation by using a combination of manual and automatic translation. Further, there is not any system and a method that uses human translations for automatic evaluation. Thus, there is a need to overcome the drawbacks of the prior art. To solve the problems of prior art, we propose a system for automatic determination of quality (high, good, acceptable to the user or bad) of machine translation and a method thereof.

OBJECTS OF THE INVENTION
The main object of the present invention is to provide a system for automatic determination of quality (high, good, acceptable to the user or bad) of machine translation and a method thereof.
Another object of the present invention is to provide a system and a method for human evaluation of the machine translation manually.

Another object is to provide a system and a method to that provides a user-friendly environment in which MT-evaluation can perform on a large data set using various MT-metrics.

Yet another object of the present invention is to provide a system and a method that gives a reliable evaluation of the machine translation by using a combination of manual and automatic translation.

Yet another object of the present invention is to provide a system and a method that uses human translations for automatic evaluation.

Yet another object of the present invention is to provide a system and a method that uses human translations for automatic evaluation on real-time basis.

Yet another object of the present invention is to provide a system for automatic determination of quality of machine translation and a method thereof at an effective cost.

Yet another object of the present invention is to provide a system for automatic determination of quality of machine translation and a method thereof that provides manual MT-evaluation based on various evaluation criteria.

SUMMARY OF THE INVENTION
This summary is provided to introduce a selection of concepts in a simplified format that are further described in the detailed description of the present disclosure. This summary is not intended to identify key or essential inventive concepts of the present disclosure, nor is it intended for determining the scope of the present disclosure.

The prior art fails to disclose an efficient and reliable system and a method system for automatic determination of quality (high, good, acceptable to the user or bad) of machine translation. Estimating the quality of machine translation is not an easy task to do. Even if one can estimate the quality, none of the prior art suggests how to evaluate the quality of the machine translation that gives a reliable output by using a combination of manual and automatic translation.

The present invention provides the solution. The present invention has solved the aforesaid problems of the prior art in various embodiments and aspects defined hereinafter. In particular, according to an aspect of the present invention, a system for automatic determination of quality of machine translation is disclosed. The system comprising an input module configured to input data. The data comprising a source data in any language, a target data in another language, wherein the target data are machine translated from the source data and a reference data obtained by human translation manually. The reference data to the human translation of the source data. The system further comprising a translation processor configured to translate the source data into the target data. The system further comprising an evaluation processor configured to evaluate the quality of the target data, wherein the evaluation processor comprises of at least four evaluation engines, a user interface module configured to enable a user to select the at least one evaluation engine, a memory module configured to store the data and the evaluation score by indexing and a processor module configured to automatically determine the quality (high, good, acceptable to the user or bad) machine translation using the evaluation of the evaluation processor. The system further comprising a human evaluation module to manually evaluate the quality of the translation on the basis of at least ten predefined criteria by which a human translator can efficiently evaluate machine translation.
According to an embodiment of the invention, the translation processor comprising TAG, SMT, EBMT and ANAL-GEN engines.

According to an embodiment of the invention, the evaluation processor uses at least four automatic tools of evaluation.

According to an embodiment of the invention, the at least four automatic tools of evaluation comprising BLEU, Meteor, F-measure and TER.

According to an embodiment of the invention, each of the BLEU, Meteor, F-measure and TER outputs a value.

According to an embodiment of the invention, the processor module calculates an average score of the value outputted by each of the BLEU, Meteor, F-measure and TER and/or by the human evaluation (12) in real-time, and if the average score is greater than a predetermined value, then the processor module determines and the user interface module outputs that the quality of the automatic translation and/or manual translation is high (good or acceptable), and if the average score is less than the predetermined value, then the processor module determines and the user interface module outputs that the quality of the automatic translation and/or manual translation is low.

According to an embodiment of the invention, the predetermined value is 0.5.

According to an embodiment of the invention, the source data, the target data and the reference data are comprising of any one or any combination of at least one sentence each or at least one file each in any format or any real-time data inputted from an external device.

According to an embodiment of the invention, a method for automatic determination of quality machine translation is disclosed. The method comprising the steps of inputting data by an input device. The data comprising a source data in any language, a target data in another language, wherein the target data are machine translated from the source data obtained by any one or more of four translation engines comprising TAG, SMT, EBMT and ANAL-GEN and a reference data obtained by human translation manually, wherein the reference data correspond to the human translation of the source data. The method further comprises automatically evaluating, by an evaluation processor, the quality of the target data using each of the at least four evaluation engines, said engines comprising at least four automatic tools comprising BLEU, Meteor, F-measure and TER. The method further comprises selecting, by a user interface, the at least one evaluation engine. The method further comprises storing, in a memory, the data and the evaluation score by indexing and determining, by a processor, the quality machine translation automatically using the evaluation of the evaluation processor and on the basis of the human evaluation.

According to an embodiment of the invention, the method further comprising outputting a value for each of the BLEU, Meteor, F-measure and TER.

According to an embodiment of the invention, the method further comprising calculating, by the processor, an average score of the value outputted by each of the BLEU, Meteor, F-measure and TER in real-time, and if the average score of the automatic evaluation and/or the human evaluation is greater than a predetermined value, then determining by the processor and outputting by the user interface that the quality of the automatic evaluation and/or the human evaluation is high (good or acceptable), and if the average score is equal or less than the predetermined value, then determining by the processor and outputting by the user interface that the quality of the automatic evaluation and/or the human evaluation data is low.
According to an embodiment of the invention, the method further comprising manually evaluating the quality of translation on the basis of at least ten criteria by which a human translator can efficiently evaluate machine translation manually during human evaluation .

BRIEF DESCRIPTION OF DRAWINGS:

Fig. 1 illustrates a block diagram of the system for automatic determination of quality of machine translation according to an aspect of the invention.

Fig. 2 illustrates a flow diagram of a method for automatic determination of quality of machine translation.

DETAILED DESCRIPTION OF THE INVENTION WITH NON-LIMITING EMBODIMENTS, EXAMPLES AND ACCOMPANYING DRAWINGS
For the purpose of promoting an understanding of the principles of the invention, reference will now be made to the embodiment illustrated in the drawings, and specific language will be used to describe the same. It will nevertheless be understood that no limitation of the scope of the invention is thereby intended. Such alterations and further modifications in the illustrated system, and such further applications of the principles of the invention as illustrated therein would be contemplated as would normally occur to one skilled in the art to which the invention relates. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skilled in the art. The system, methods, and examples provided herein are illustrative only and are not intended to be limiting.

The term “some” as used herein is to be understood as “none or one or more than one or all.” Accordingly, the terms “none,” “one,” “more than one,” “more than one, but not all” or “all” would all fall under the definition of “some.” The term “some embodiments” may refer to no embodiments or to one embodiment or to several embodiments or to all embodiments, without departing from the scope of the present disclosure.

The terminology and structure employed herein is for describing, teaching, and illuminating some embodiments and their specific features. It does not in any way limit, restrict or reduce the spirit and scope of the claims or their equivalents.

More specifically, any terms used herein such as but not limited to “includes,” “comprises,” “has,” “consists,” and grammatical variants thereof do not specify an exact limitation or restriction and certainly do not exclude the possible addition of one or more features or elements, unless otherwise stated, and furthermore must not be taken to exclude the possible removal of one or more of the listed features and elements, unless otherwise stated with the limiting language “must comprise” or “needs to include.”

Whether or not a certain feature or element was limited to being used only once, either way, it may still be referred to as “one or more features” or “one or more elements” or “at least one feature” or “at least one element.” Furthermore, the use of the terms “one or more” or “at least one” feature or element do not preclude there being none of that feature or element, unless otherwise specified by limiting language such as “there needs to be one or more . . . ” or “one or more element is required.”

Unless otherwise defined, all terms, and especially any technical and/or scientific terms, used herein may be taken to have the same meaning as commonly understood by one having ordinary skills in the art.

Reference is made herein to some “embodiments.” It should be understood that an embodiment is an example of a possible implementation of any features and/or elements presented in the attached claims. Some embodiments have been described for the purpose of illuminating one or more of the potential ways in which the specific features and/or elements of the attached claims fulfill the requirements of uniqueness, utility and non-obviousness.

Use of the phrases and/or terms including, but not limited to, “a first embodiment,” “a further embodiment,” “an alternate embodiment,” “one embodiment,” “an embodiment,” “multiple embodiments,” “some embodiments,” “other embodiments,” “further embodiment”, “furthermore embodiment”, “additional embodiment” or variants thereof do not necessarily refer to the same embodiments. Unless otherwise specified, one or more particular features and/or elements described in connection with one or more embodiments may be found in one embodiment, or may be found in more than one embodiment, or may be found in all embodiments, or may be found in no embodiments. Although one or more features and/or elements may be described herein in the context of only a single embodiment, or alternatively in the context of more than one embodiment, or further alternatively in the context of all embodiments, the features and/or elements may instead be provided separately or in any appropriate combination or not at all. Conversely, any features and/or elements described in the context of separate embodiments may alternatively be realized as existing together in the context of a single embodiment.

Any particular and all details set forth herein are used in the context of some embodiments and therefore should not be necessarily taken as limiting factors to the attached claims. The attached claims and their legal equivalents can be realized in the context of embodiments other than the ones used as illustrative examples in the description below.

Embodiments of the present invention will be described below in detail with reference to the accompanying drawings.

The figures and the forgoing description give examples of embodiments. Those skilled in the art will appreciate that one or more of the described elements may well be combined into a single functional element. Alternatively, certain elements may be split into multiple functional elements. Elements from one embodiment may be added to another embodiment. For example, orders of processes described herein may be changed and are not limited to the manner described herein. Moreover, the actions of any flow diagram need not be implemented in the order shown; nor do all the acts necessarily need to be performed. Also, those acts that are not dependent on other acts may be performed in parallel with the other acts. The scope of the embodiments is by no means limited by these specific examples. Numerous variations, whether explicitly given in the specification or not, such as differences in structure, dimension, and use of material, are possible.

According to an embodiment of the invention, a system (10) for automatic determination of quality of machine translation is disclosed. The system (10) comprising an input module (20), a translation processor (30), an evaluation processor (35), a user interface module, a memory module and a processor module (40).

In an embodiment of the invention, the system (10) may be configured with any one of or a combination of the following language models, but not limited to, N-Gram or Markov chain or bigram or trigram or smoothing models, Google translator, Microsoft MT-Engine, Babylon MT Engine, Anusaaraka MT-Engine, Moses MT-Toolkit, Joshua MT- Toolkit, TAG, SMT, EBMT and ANAL-GEN engines, BLEU (Bilingual Evaluation Understudy Score), Meteor, F-measure and TER.

According to an embodiment of the invention, the input module (20) is configured to input data. The input module (20) may be any input device connected to the system (10). In some embodiments, the input module (20) may be any conventional inputting device connected to the system (10). The input module (20) may be an inputting system and/or any input processing device. The input module (20) may be configured to inputting data in the system (10) from an external source, like a communication medium between two devices or a data receiver unit for the system (10). The input module (20) may comprise of a processor, a computer or any processing unit capable of enabling input of a plurality of sentences, text files in any format, like PDF, an image, .doc file etc., and/or an image to text converter to extract text from the image. In some embodiments of the present invention, the input device (20) may be connected to another device, like a processor or an output device of any external system, to provide automatic input to the input device (20) of the system (10).

The input module (20) may be configured to receive input data in any readable form. According to an embodiment of the invention, the input data may comprise of a source data in any language, a target data in another language, wherein the target data are machine translated from the source data and a reference data obtained by human translation manually. The reference data may be a human translation of the source data, which may correspond identically to the target data. According to an embodiment of the invention, the source data, the target data and the reference data are comprising of any one or any combination of at least one sentence each or at least one file each in any format or any real-time data inputted from an external device. In one embodiment, the source data, the target data and the reference data are comprising of any one or any combination of at least one sentence each or at least one file each in any format or any real-time data inputted from an external device. The input module (20) may be configured to receive data from the external source on real-time basis.

According to an embodiment of the invention, the source data may comprise of source sentences in any language. The source data may be one sentence or a plurality of sentences to be translated. The source data may further comprise of text files having source sentences. According to an embodiment of the invention, the target data may be comprised of a data obtained after translation of the source data in another language. The target data is a machine translation (MT) of the source data in another language, as required by a user. According to an embodiment of the invention, the reference data may be comprised of a data obtained after translation of the source data in another language. The reference data may be a translated data obtained after translation of the source data manually by human. The reference data may be a human translation of the source data, which may correspond identically to the target data.

According to an embodiment of the invention, the system (10) may input the source data, the target data and the reference data directly from the input module (20). According to some embodiments of the invention, the system (10) may input the source data and provide the automatic machine translation of the source data in real-time basis, as the target data. The automatic machine translation may be provided on real-time basis using any in-house or external translation tool, like a Google translator, Microsoft MT-Engine, Babylon MT Engine, Anusaaraka MT-Engine, Moses MT-Toolkit, Joshua MT- Toolkit, any open source machine translator or any in-house translator specifically created for the purpose of the invention. It may be comprised of one or more than one of any machine translation (MT) engine(s) available for online and offline translation.

The input module (20) may record the reference data for determining the quality of the automatic machine translation accordingly. The system (10) may use any conventional tool for automatic machine translation of the source data.

According to an embodiment of the invention, the system (10) may comprise a translation processor (30). The translation processor (30) may be any computing hardware or software or a combination of hardware software module known in the state of the art. The translation processor (30) may be configured to translate the source data in a second language. According to an embodiment of the invention, the translation processor (30) may receive the input data from the input module (20). In some embodiments of the invention, the translation engine (30) may comprise TAG ( Tree Adjoining Grammar), SMT (Statistical machine translation), EBMT (Example-based machine translation) and ANAL-GEN (Analytical Generator) engines. In some embodiments, the translation processor (30) may be configured with more than one translation engines and preferably with any one or more of the TAG, SMT, EBMT and ANAL-GEN engines. The output of the translation processor (30) is used as the input data to perform evaluation automatically. For example, the system (10) may comprise of an evaluation processor (35) for providing the quality of translation automatically and the evaluation processor (35) may be configured with at least four evaluation engines. The four evaluation engines may be comprise of evolution matrics, like BLEU, METEOR, F-measure, TER etc.

According to an embodiment of the invention, the system (10) further comprises a user interface module (not shown) configured to enable the user to select the at least one evaluation engine for assessing quality of the machine translated target data. The user interface module may be any interface enabled hardware or software or a combination of hardware software module known in the state of the art. In some embodiments, the system (10) may be configured with any such kind of user interface module. The user may select any evaluation engines, as configured in the evaluation processor (35), for evaluating the machine translation of the target data, and preferably with any one of the BLEU (Bilingual Evaluation Understudy Score), Meteor, F-measure and TER (Translation Error Rate).
According to an embodiment of the invention, the evaluation processor (35) comprising at least four evaluation engines may use at least four automatic tools of evaluation. According to an embodiment of the invention, the four automatic tools of evaluation may be offline or online evaluation metrics based on BLEU (Bilingual Evaluation Understudy Score), Meteor, F-measure and TER (Translation Error Rate). According to an embodiment of the invention, each of the evaluation metrics and/or evaluation tools outputs a value. The evaluation processor (35) is configured to provide the outputted values to the user interface module to display to the user. The evaluation processor (35) is configured to provide the outputted values to the memory unit and the processor module (40) for further processing.

In some embodiments of the invention, the evaluation processor (35) may comprise of any conventional processing device or a computing device, or a software tool or a combination of the hardware and software for determining the quality.

According to an embodiment of the invention, the evaluation processor (35) comprising a memory module (not shown) configured to store the data and/or the evaluation score by indexing. The memory module may be an external storage unit or an internal storage unit for storing the data. The system (10) may be configured to communication with the external storage unit in any conventional way, if required. The parameters, like the storage capacity etc., of the memory module may be configured as per the requirements of the invention.

According to an embodiment of the invention, the system (10) for automatic evaluation of a quality of machine translation comprises a processor module (40) configured to automatically determine quality machine translation using the evaluation of the evaluation processor (35) and/or on the basis of human evaluation (12).

According to an embodiment of the present invention, the system (10) further comprises a human evaluation module (12) for manual evaluation of the translated sentence. The human evaluation (12) or manual evaluation (12) is performed on data on the basis of predefined list of criteria. The human evaluation (12) then provides a score, which may be less than or greater than or equal to a predetermined value, like the values of the evaluation engines for automatic evaluation. In some embodiments, the predetermined value is calculated on the basis of at least ten criteria by which a human translator can efficiently evaluate machine translation manually. Each of the predefined criteria is given some weights and the average value of all the criteria that comes after evaluation is taken to check the quality of the translation. The ideal average value may also be used as the predetermined criteria, like the value of 0.5 for ten criteria listed in the preceding sections.
In some embodiments, the reference data is mapped on the target data for determining the quality of the machine translation in the automatic evaluation and/or in the manual evaluation.

Accordingly, the system (10) is configured to provide both the human evaluation (12) and automatic evaluation (35). In human evaluation (12), the system (10) uses source sentence (like English) and target sentence (like machine translated in Hindi) and in the automatic evaluation (35), the system (10) uses one source sentence, one target sentence and at least one or more reference sentences for evaluation.

For example, a person having a good knowledge of source and target language can be a human evaluator. Since, evaluation task is domain specific, if a domain expert or linguistic is doing this task then he/she will give more accurate result. Accordingly, human evaluates a sentence as he/she thinks about its translation. If he/she has been evaluating sentences continuously since a long time then the result will not be similar. In sum, if the human evaluated a sentence with poor score, he may, after some time, give high weight to same structure’s sentence. So, to get similar results, the invention proposes to define the criteria on which the sentence can be evaluated.

According to an embodiment of the invention, the human evaluation (12) provides following 10 different criteria and four weights (.25, .50, .75, 1) for the following criteria and the overall score is taken as the predetermined value:
1. Translation of Gender and Number of the Noun/s.
2. Translation of tense in the source sentence.
3. Translation of voice in the source sentence.
4. Identification of the proper nouns.
5. Use of adjectives and adverbs corresponding to the nouns and verbs in the source sentence.
6. Selection of proper words/synonyms.
7. Usage of phrases and clauses in the translation.
8. Use of punctuation signs in the translation.
9. Maintaining the stress on the significant part in the source sentence in the translation.
10. Maintaining the semantics of the source sentence in the transition.
11. Overall Score (Predetermined value).

According to some embodiments of the invention, the user may change the above-defined criteria, weights and the number of criteria to be considered as per the requirements of the system (10). Thus, the system (10) is configured to enable such modifications in the operational parameters.

According to an embodiment of the invention, the processor module (40) calculates an average score of the value outputted by each of the BLEU, Meteor, F-measure and TER in real-time. The system (10) is enabled with manual and automatic machine translation evaluation and evaluation result analysis for future machine translation improvement research. Accordingly, if the average score of the value outputted by each of the BLEU, Meteor, F-measure and TER is greater than the predetermined value, then the processor module (40) is configured to then determine that the quality of the target data is high (good or acceptable) and/or the user interface module is configured to output the determined result. If the average score of the value outputted by each of the BLEU, Meteor, F-measure and TER is equal to or less than the predetermined value, the processor module (40) is configured to then determine that the quality of the target data is low and/or the user interface module is configured to output the determined result. In some embodiments of the invention, the processor module (40) may comprise of any conventional processing device or a computing device, or a software tool or a combination of the hardware and software for determining the quality.

As per an exemplary embodiment of the invention, the system (10), to perform the automatic Machine Translation valuation, is configured to allow the user to select the text files of source sentences, target sentences (MT-translations) and their corresponding reference sentences (human translations) of any language. The system (10) provides four choices of evaluation metrics to the user and the user may select any one from the metric’s list. Upon feeding the required information in the system (10), all sentences or input data (10) will be stored in database or the memory unit. The evaluation processor (35) will then start the evaluation process and generate the evaluation score, which will be stored in the database. Since, each evaluation metrics give different score, the system (10) takes the average score of all metrics as a final score. The processor module (40) is configured to perform the above actions of calculating the final score. The database is also provided with the predetermined evaluation of the human evaluation (12). Then the processor module (40) compares the final score of the automatic evaluation (12) and the human evaluation (12) with the predetermined score. If the final score of the automatic evaluation (12) and the human evaluation (12) is greater than the predetermined value, i.e., 0.5, then the human translation and/or the automatic translation is of good quality. Otherwise, the human translation and/or the automatic translation is of bad quality. In particular, if the predetermined value of the human evaluation (12) on the basis of the above-listed 10 criteria is 0.5, then the processor module (40) compares the final score with 0.5. If the final score is greater than 0.5, then the quality of the machine translation is good enough. However, if the final score is less than 0.5, then the quality of the machine translation is bad. Stored results in database compare for checking the efficiency of MT-engine-outputs domain wise and file wise.

Fig. 2 illustrates a flow diagram of a method for automatic determination of quality of machine translation. The features of the method described below are analogous to the features of the system (10) described above. According to an embodiment of the invention, the method for automatic determination of quality (good or acceptable to the user) of machine translation is disclosed. The method comprising the step of inputting data (S1), to the system (10), by an input device (20). The data comprising a source data in any language, a target data in another language and/or a reference data obtained by human translation manually.

The method further comprises the next step of translating (S2), by a translation processor (30), the source sentence in to the target sentence by machine translation and human translation. For machine translation, the system (10) uses at least four translation engines comprising TAG, SMT, EBMT and ANAL-GEN engines. The results of the machine translation is used as an input for automatic evaluation. The invention uses at least four evaluation engines use or at least four automatic tools comprising BLEU, Meteor, F-measure and TER. The method further comprises outputting a value for each of the BLEU, Meteor, F-measure and TER.

The method further comprises the next step of selecting (S3), by a user interface, at least one evaluation engine. The user can select any one of the evaluation engines for obtaining the quality of machine translation using the automatic tools of BLEU, Meteor, F-measure and TER.

The method further comprises the next step of storing (S4), in a memory, the data and the evaluation score by indexing. The method further comprises the next step of determining (S5), by a processor, the quality (high, good, bad or acceptable to the user) machine translation automatically using the evaluation of the evaluation processor (35).

The method further comprises the next step of calculating (S6), by a processor (40), an average score of the value outputted by each of the BLEU, Meteor, F-measure and TER in real-time, and if the average score is greater than a predetermined value, then determining (S7) by the processor (40) and outputting by the user interface that the quality of the machine translated data is high (good or acceptable to the user), and if the average score is equal or less than the predetermined value, then determining (S8) by the processor (40) and outputting by the user interface that the quality of the machine translated data is low or bad.

The method further comprises the next step of manually evaluating the translation on the basis of at least ten criteria by which a human translator can efficiently evaluate machine translation manually (or human evaluation (12)). The method further calculates the predetermined value.

According to an embodiment of the invention, the human evaluation (12) provides following 10 different criteria and four weights (.25, .50, .75, 1) for the following criteria. The human evaluation then calculates the overall score using the 10 different criteria and four weights (.25, .50, .75, 1) on the following criteria. The calculated overall score is taken as the predetermined value.
1. Translation of Gender and Number of the Noun/s.
2. Translation of tense in the source sentence.
3. Translation of voice in the source sentence.
4. Identification of the proper nouns.
5. Use of adjectives and adverbs corresponding to the nouns and verbs in the source sentence.
6. Selection of proper words/synonyms.
7. Usage of phrases and clauses in the translation.
8. Use of punctuation signs in the translation.
9. Maintaining the stress on the significant part in the source sentence in the translation.
10. Maintaining the semantics of the source sentence in the transition.
11. Overall Score (Predetermined value).

According to an exemplary embodiment of the invention, during human evaluation (12), the user provides one weight out of .25, .50, .75, 1 to each of the above-listed criteria. Then the average of all the given weights are calculated for the following criteria. The average of the weights is used to calculate the predetermined value. According to some embodiments of the invention, the user may change the above-defined criteria, weights, the number of criteria to be considered and the calculation of the predetermined value as per the requirements of the system (10). Thus, the system (10) is configured to enable such modifications in the operational parameters.

The figures and the forgoing description give examples of embodiments. Those skilled in the art will appreciate that one or more of the described elements may well be combined into a single functional element. Alternatively, certain elements may be split into multiple functional elements. Elements from one embodiment may be added to another embodiment. For example, orders of processes described herein may be changed and are not limited to the manner described herein. Moreover, the actions of any flow diagram need not be implemented in the order shown; nor do all the acts necessarily need to be performed. Also, those acts that are not dependent on other acts may be performed in parallel with the other acts. The scope of the embodiments is by no means limited by these specific examples. Numerous variations, whether explicitly given in the specification or not, such as differences in structure, dimension, and use of material, are possible.

We claim:

1. A system (10) for automatic determination of quality of machine translation, the system (10) comprising:
- an input module (20) configured to input data, wherein the data comprising:
a source data in any language;
a target data in another language, wherein the target data are machine translated from the source data ; and
a reference data obtained by human translation manually, wherein the reference data is the human translation of the source data;
- a translation processor (30) configured to translate the source data into the target data;
- an evaluation processor (35) configured to evaluate the quality of the target data, wherein the evaluation processor (35) comprises of at least four evaluation engines;
- a user interface module configured to enable a user to select the at least one evaluation engine;
- a memory module configured to store the data and the evaluation score by indexing;
- a processor module (40) configured to automatically determine the quality machine translation using the evaluation of the evaluation processor; and
- a human evaluation module (12) to manually evaluate the quality of the translation on the basis of at least ten predefined criteria by which a human translator can efficiently evaluate machine translation.

2. The system (10) for automatic evaluation of quality of machine translation as claimed in claim 1, wherein the translation processor (30) comprising TAG, SMT, EBMT and ANAL-GEN engines.

3. The system (10) for automatic evaluation of quality of machine translation as claimed in claim 1, wherein evaluation processor (35) is configured to use at least four automatic tools of evaluation.

4. The system (10) for automatic evaluation of quality of machine translation as claimed in claim 3, wherein the at least four automatic tools of evaluation comprising BLEU, Meteor, F-measure and TER, and wherein each of the BLEU, Meteor, F-measure and TER outputs a value.

5. The system (10) for automatic evaluation of quality of machine translation as claimed in claim 4, wherein the processor module (40) calculates an average score of the value outputted by each of the BLEU, Meteor, F-measure and TER in real-time and/or by the human evaluation (12), and if the average score is greater than a predetermined value, then the processor module (40) determines and the user interface module outputs that the quality of the automatic translation and/or manual translation is high, and if the average score is equal to or less than the predetermined value, then the processor module (40) determines and the user interface module outputs that the quality of the automatic translation and/or manual translation is low.

6. The system (10) for automatic evaluation of quality of machine translation as claimed in claim 5, wherein the predetermined value is 0.5.

7. A method for automatic determination of quality machine translation, the method comprising the steps of:
- inputting data (S1) by an input device, wherein the data comprising:
a source data in any language;
a target data in another language, wherein the target data are machine translated from the source data obtained by any one or more of four translation engines comprising TAG, SMT, EBMT and ANAL-GEN engines; and
a reference data obtained by human translation manually, wherein the reference data is the human translation of the source data;
- automatically evaluating (S2), by an evaluation processor, the quality of the target data using each of the at least four evaluation engines, said engines comprising at least four automatic tools comprising BLEU, Meteor, F-measure and TER;
- selecting (S3), by a user interface, the at least one evaluation engine;
- storing (S4), in a memory, the data and the evaluation score by indexing; and
- determining (S5), by a processor, the quality of machine translation automatically using the evaluation of the evaluation processor (35) and on the basis of the human evaluation (12)

8. The method for automatic evaluation of quality of machine translation as claimed in claim 7 further comprising outputting a value for each of the BLEU, Meteor, F-measure and TER.

9. The method for automatic evaluation of quality of machine translation as claimed in claim 7 further comprising calculating (S6), by the processor, an average score of the value outputted by each of the BLEU, Meteor, F-measure and TER in real-time, and if the average score of the automatic evaluation and/or the human evaluation is greater than a predetermined value, then determining (S7) by the processor and outputting by the user interface that the quality of the machine translated data is high, and if the average score of the automatic evaluation and/or the human evaluation is equal to or less than the predetermined value, then determining (S8) by the processor and outputting by the user interface that the quality of the machine translated data is low.

10. The method for automatic evaluation of quality of machine translation as claimed in claim 9 further comprising manually evaluating the quality of translation on the basis of at least ten criteria by which a human translator can efficiently evaluate machine translation manually during human evaluation (12).

Documents

Application Documents

# Name Date
1 202011007908-FORM 1 [25-02-2020(online)].pdf 2020-02-25
2 202011007908-FIGURE OF ABSTRACT [25-02-2020(online)].jpg 2020-02-25
3 202011007908-DRAWINGS [25-02-2020(online)].pdf 2020-02-25
4 202011007908-COMPLETE SPECIFICATION [25-02-2020(online)].pdf 2020-02-25
5 202011007908-CLAIMS UNDER RULE 1 (PROVISIO) OF RULE 20 [25-02-2020(online)].pdf 2020-02-25
6 202011007908-PA [22-02-2021(online)].pdf 2021-02-22
7 202011007908-FORM-26 [22-02-2021(online)].pdf 2021-02-22
8 202011007908-CORRESPONDENCE-OTHERS [22-02-2021(online)].pdf 2021-02-22
9 202011007908-ASSIGNMENT DOCUMENTS [22-02-2021(online)].pdf 2021-02-22
10 202011007908-8(i)-Substitution-Change Of Applicant - Form 6 [22-02-2021(online)].pdf 2021-02-22
11 202011007908-POA [23-02-2021(online)].pdf 2021-02-23
12 202011007908-MARKED COPIES OF AMENDEMENTS [23-02-2021(online)].pdf 2021-02-23
13 202011007908-FORM 13 [23-02-2021(online)].pdf 2021-02-23
14 202011007908-Annexure [23-02-2021(online)].pdf 2021-02-23
15 202011007908-AMENDED DOCUMENTS [23-02-2021(online)].pdf 2021-02-23
16 202011007908-DRAWING [24-02-2021(online)].pdf 2021-02-24
17 202011007908-CORRESPONDENCE-OTHERS [24-02-2021(online)].pdf 2021-02-24
18 202011007908-COMPLETE SPECIFICATION [24-02-2021(online)].pdf 2021-02-24
19 202011007908-FORM 18 [10-09-2021(online)].pdf 2021-09-10
20 abstract.jpg 2021-10-18
21 202011007908-Power of Attorney-010321.pdf 2021-10-18
22 202011007908-OTHERS-010321.pdf 2021-10-18
23 202011007908-Form 5-010321.pdf 2021-10-18
24 202011007908-Correspondence-010321.pdf 2021-10-18
25 202011007908-FER.pdf 2022-03-14
26 202011007908-MARKED COPIES OF AMENDEMENTS [24-06-2022(online)].pdf 2022-06-24
27 202011007908-FORM 13 [24-06-2022(online)].pdf 2022-06-24
28 202011007908-FER_SER_REPLY [24-06-2022(online)].pdf 2022-06-24
29 202011007908-CLAIMS [24-06-2022(online)].pdf 2022-06-24
30 202011007908-AMMENDED DOCUMENTS [24-06-2022(online)].pdf 2022-06-24
31 202011007908-ABSTRACT [24-06-2022(online)].pdf 2022-06-24
32 202011007908-FORM-8 [31-10-2024(online)].pdf 2024-10-31

Search Strategy

1 202011007908E_11-03-2022.pdf