Abstract: The present invention discloses an automated system and a method to determine a requirement to edit at least one translated sentence. The system comprising an input module (04) to input at least one sentence in a first language, a translator module (06) to translate the at least one sentence into the at least one translated sentence in a second language, a probability processing module (10) to calculate a probability score of the at least one translated sentence to estimate quality of the at least one translated sentence and a processor module (12) to determine the requirement to pre-edit (12b) or post-edit (12a) the at least one translated sentence on the basis of the estimated quality of the at least one translated sentence.
1. An automated system to determine a requirement to edit at least one translated sentence, said system comprising: an input module to input at least one sentence in a first language; a translator module to translate the at least one sentence into the at least one translated sentence in a second language; a probability processing module to calculate a probability score of the at least one translated sentence to estimate quality of the at least one translated sentence; and a processor module to determine the requirement to pre-edit or post-edit the at least one translated sentence on the basis of the estimated quality of the at least one translated sentence.
2. The automated system to determine a requirement to edit at least one translated sentence as claimed in claim 1, wherein the processor module compares the probability score with a pre-defined score (T).
3. The automated system to determine a requirement to edit at least one translated sentence as claimed in claim 2, wherein the processor module outputs a post-edit command if the probability score is greater than the pre-defined score (T).
4. The automated system to determine a requirement to edit at least one translated sentence as claimed in claim 2, wherein the processor module outputs a pre-edit command if the probability score is less than the pre-defined score (T).
5. The automated system to determine a requirement to edit at least one translated sentence as claimed in claim 2, wherein the pre-defined score (T) is calculated by formula (1): T_£P=1£jg1p(w°) n+m Where as :-n=no. of sentences m=no. of MT-Engine
6. The automated system to determine a requirement to edit at least one translated sentence as claimed in claim 1, wherein the quality of the at least one translated sentence is estimated on a real-time basis.
7. The automated system to determine a requirement to edit at least one translated sentence as claimed in claim 1, wherein the first language is English and the second language is Hindi.
8. The automated system to determine a requirement to edit at least one translated sentence as claimed in claim 1, wherein the probability processing module is trained on a language model.
9. The automated system to determine a requirement to edit at least one translated sentence as claimed in claim 8, wherein the language model comprising a N-Gram or Markov chain or bigram or trigram or smoothing models or any combination thereof.
10. The automated system to determine a requirement to edit at least one translated sentence as claimed in claim 9, wherein the probability processing module calculates the probability score by any one of the language models for the at least one translated sentence or by averaging probability scores calculated by any one of the language models for all the translated sentences or by averaging all the probability scores calculated by all the language models for the at least one translated sentence or all the translated sentences, or by any combination thereof.
11. A method for automated determination of a requirement to edit at least one translated sentence, said method comprising: inputting, by an input device, at least one sentence in a first language; - translating, by a translator, the at least one sentence into the at least one translated sentence in a second language; calculating, by a probability processing module, a probability score of the at least one translated sentence to estimate quality of the at least one translated sentence; and determining, by a processor, the requirement to post-edit or pre-edit the at least one translated sentence on the basis of the estimated quality of the at least one translated sentence.
12. The method for automated determination of a requirement to edit at least one translated sentence as claimed in claim 11, wherein the processor compares the probability score with a pre-defined score (T).
13. The method for automated determination of a requirement to edit at least one translated sentence as claimed in claim 12, wherein the processor outputs a pre-edit command if the probability score is less than the pre-defined score (T).
14. The method for automated determination of a requirement to edit at least one translated sentence as claimed in claim 12, wherein the processor outputs a post-edit command if the probability score is greater than the pre-defined score (T).
15. The method for automated determination of a requirement to edit at least one translated sentence as claimed in claim 12, wherein the pre-defined score (T) is calculated by formula (1): T_£P=1£jg1p(wg) n+m Where as :-n=no. of sentences m=no. of MT-Engine
16. The method for automated determination of a requirement to edit at least one translated sentence as claimed in claim 11, wherein the first language is English and the second language is Hindi.
17. The method for automated determination of a requirement to edit at least one translated sentence as claimed in claim 11, wherein the quality of the at least one translated sentence is estimated on a real-time basis.
18. The method for automated determination of a requirement to edit at least one translated sentence as claimed in claim 11, wherein the probability processing module is trained on a language model to calculate the probability score of the at least one translated sentence.
19. The method for automated determination of a requirement to edit at least one translated sentence as claimed in claim 18, wherein the language model comprises of a N-Gram or Markov chain or bigram or trigram or smoothing models or any combination thereof.
20. The method for automated determination of a requirement to edit at least one translated sentence as claimed in claim 19, wherein the probability processing module calculates the probability score by any one of the language models for the at least one translated sentence or by averaging probability scores calculated by any one of the language models for all the translated sentences or by averaging all the probability scores calculated by all the language models for the at least one translated sentence or all the translated sentences, or by any combination thereof.
FIELD OF THE INVENTION
The present invention relates to a system and a method to determine a requirement to edit a translated sentence and a method thereof. More particularly, the present invention relates to an automated system and a method to determine a requirement to edit at least one translated sentence.
BACKGROUND OF THE INVENTION
Background description includes information that may be useful in understanding the present invention. It is not an admission that any of the information provided herein is prior art or relevant to the presently claimed invention, or that any publication specifically or implicitly referenced is prior art.
India being a linguistically rich country has around 780 different languages and 22 languages among them are constitutional languages. The Indian languages are scripted in 86 scripts. The official language of the country is "Hindi" and the associated official language is English. However, only 40% of the population in India uses Hindi for communication; most of the information written and made available on the Internet is in English language. Thus, there is a requirement for a linguistic bridge between both the languages. In fact, this problem is faced at a much larger scale in the world. The total number of different languages around the world is more than 6500 and only one language, English, is taken up as bridging language to enable communication among the people speaking in hundreds of different languages. To close the gap of different languages and bring the translation to more accuracy, many researchers have been pursuing research in Machine Translation (MT). The success report of the machine translation is around 70% but it is not sufficient for publications in official communication. Moreover, the machine translations are most often inaccurate and give confusing results. Therefore, the drawbacks of the machine translations need modifications for proper understanding of literature.
The demand for translation services is increasing beyond the rate of growth of world trade, which is in turn higher than the growth rate of the world economy. More than half of all Internet traffic is now in a language other than English, and the evidence is that the trend towards domination by English in commercial life more generally is slowing down. Recruitment to the translation profession, though increasing, is still not adequate to meet demand. Meanwhile, new technologies in the processing of natural language are raising the prospect of ever greater involvement of the computer in the handling of translation.
Conventionally, in a machine translation if the accuracy is insufficient, post-editing is performed to improve the translation accuracy. The post-editing work may be performed manually or mechanically. Human translators could overcome the aforesaid drawbacks of machine translation. However, if we take all MT-outputs for modification (post-editing), it would involve extreme human effort for correction; highly time consuming, expensive and ineffective at large. The ineffectiveness arises because the output of machine translation is not considered suitable and good for use in official communication. Sometimes the output may be accurate, however, the reliability on the output of machine translation is highly doubtful. This results in the necessity of post-editing a machine translated text, in order to improve the quality to acceptable standards. To overcome this drawback, the prior art provides the following options, which are however, inefficient: The prior art JP 2004 318 424A discloses a post-translation editing method that automatically supports post-editing using a document converted into electronic data. For example, there is a technique that supports post-editing by performing automatic post-editing and displaying the translated text to which post-editing has been applied and the translated text to which post-editing has not been applied separately.
Another prior art US 2009 0076 792A1 discloses a text editing apparatus for the editing of text translated from at least a first language to a second language. The apparatus comprising user input means; and a controller
adapted to identify the language of the pre-translated text and/or post-translated text, and to use said identification of the language(s) to automatically select and/or verify selection of post-editing processes for post¬editing of the translated text. The controller may be configured to identify a sequence of translated languages used to translate said text from at least a first to a second to a third language, and to use said sequence for selection or verification of the selection of post-editing processes.
Yet another prior art US 2003 0236 658A1 discloses a system, method and computer program product for translating information from a first language into a second language. Information in a first language is edited to prepare the information for translation into a second language utilizing a first editing module. The information is translated into the second language utilizing a translation module. The translated information is translated to conform the translated information with grammatical rules of the second language utilizing a second editing module.
Thus, with the present machine translation systems of the prior art, as mentioned above, a large amount of time, effort, energy and inefficiency are involved. There is still a need to overcome the drawbacks of prior art. The prior art systems fail to provide the output that could convert the output of the machine translations into human-quality translation. Thus, there is a requirement to confirm if a machine translated sentence is suitable enough for publication or to be used without further ado by humans. There is a further requirement to overcome the drawbacks of the prior art. To solve the problems of prior art, we propose an automated system and a method to determine a requirement to edit a translated sentence.
It is, therefore, an object of the invention to provide an automated system and a method to determine a requirement to edit at least one translated sentence.
OBJECTS OF THE INVENTION
The main object of the present invention is to provide an automated system to determine a requirement to edit at least one translated sentence.
Another object is to provide a method to determine a requirement to edit at least one translated sentence.
Yet another object of the present invention is to provide an automated efficient system that determines a requirement to edit at least one translated sentence. Yet another object of the present invention is to provide an automated system that determines a requirement to edit at least one translated sentence on real¬time basis.
Yet another object of the present invention is to provide an automated system that determines a requirement to edit at least one translated sentence to produce a human quality translation.
Yet another object of the present invention is to provide an automated system that determines a requirement to edit at least one translated sentence at an effective cost.
SUMMARY OF THE INVENTION
This summary is provided to introduce a selection of concepts in a simplified format that are further described in the detailed description of the present disclosure. This summary is not intended to identify key or essential inventive concepts of the present disclosure, nor is it intended for determining the scope of the present disclosure.
The prior art could teach us to consider only those machine translation outputs that are nearest to human translation or those are near good quality. The good quality translations can be corrected with little human effort. However, the prior art fails to disclose an efficient and reliable system and a method to determine how can we estimate the quality of translation. Estimating the quality of machine translation is not an easy task to do. Even if one can estimate the quality, none of the prior art suggests how to attain the human quality translation from the machine translation using the value of the estimated quality of a machine translation.
According to an embodiment, the present invention has solved this problem for any languages and preferably for English to Hindi language pair by using
various language models. Based on language models scores, the present invention estimates the quality of machine translation in good and bad categories. Accordingly, the translations that fall under the category of good translation can be used for post-editing. The human quality translation is achieved by determining a requirement for post-editing of machine translation. According to an aspect of the present invention, an automated system to determine a requirement to edit at least one translated sentence. The system comprising an input module to input at least one sentence in a first language, a translator module to translate the at least one sentence into the at least one translated sentence in a second language, a probability processing module to calculate a probability score of the at least one translated sentence to estimate quality of the at least one translated sentence and a processor module to determine the requirement to pre-edit or post-edit the at least one translated sentence on the basis of the estimated quality of the at least one translated sentence.
According to an embodiment of the present invention, the processor module compares the probability score with a pre-defined score. According to an embodiment of the present invention, the processor module outputs a post-edit command if the probability score is greater than the pre¬defined score.
According to an embodiment of the present invention, the processor module outputs a pre-edit command if the probability score is less than the pre-defined score.
According to an embodiment of the present invention, the pre-defined score (T) is calculated by formula (1): T_£P=1£{H1P(wg)
n+m
Where, n=number of sentences, and m= number of MT engine
According to an embodiment of the present invention, the quality of the at
least one translated sentence is estimated on a real-time basis.
According to an embodiment of the present invention, the first language is
English and the second language is Hindi.
According to an embodiment of the present invention, wherein the probability
processing module is trained on a language model.
According to an embodiment of the present invention, the language model
comprising a N-Gram or Markov chain or bigram or trigram or smoothing
models or any combination thereof.
According to an embodiment of the present invention, the probability
processing module calculates the probability score by any one of the language
models for the at least one translated sentence or by averaging probability
scores calculated by any one of the language models for all the translated
sentences or by averaging all the probability scores calculated by all the
language models for the at least one translated sentence or all the translated
sentences, or by any combination thereof.
According to an aspect of the present invention, a method for automated
determination of a requirement to edit at least one translated sentence. The
method comprising inputting, by an input device, at least one sentence in a
first language; translating, by a translator, the at least one sentence into the at
least one translated sentence in a second language; calculating, by a
probability processing module, a probability score of the at least one translated
sentence to estimate quality of the at least one translated sentence; and
determining, by a processor, the requirement to post-edit or pre-edit the at
least one translated sentence on the basis of the estimated quality of the at
least one translated sentence.
According to an embodiment of the present invention, the processor compares
the probability score with a pre-defined score (T).
According to an embodiment of the present invention, the processor outputs a
pre-edit command if the probability score is less than the pre-defined score
(T).
According to an embodiment of the present invention, the processor outputs a
post-edit command if the probability score is greater than the pre-defined
score (T).
According to an embodiment of the present invention, the pre-defined score
(T) is calculated by formula (1):
T_£P=1£{H1P(wg)
n+m
Where, n=number of sentences; m= number of MT-Engine
According to an embodiment of the present invention, the first language is
English and the second language is Hindi.
According to an embodiment of the present invention, the quality of the at
least one translated sentence is estimated on a real-time basis.
According to an embodiment of the present invention, the probability
processing module is trained on a language model to calculate the probability
score of the at least one translated sentence.
According to an embodiment of the present invention, the language model
comprises of a N-Gram or Markov chain or bigram or trigram or smoothing
models or any combination thereof.
According to an embodiment of the present invention, the probability
processing module calculates the probability score by any one of the language
models for the at least one translated sentence or by averaging probability
scores calculated by any one of the language models for all the translated
sentences or by averaging all the probability scores calculated by all the
language models for the at least one translated sentence or all the translated
sentences, or by any combination thereof.
BRIEF DESCRIPTION OF DRAWINGS:
Fig. 1 illustrates a flow diagram of generating a machine translation output in a second language for a sentence input by a user in a first language; Fig. 2 illustrates a flow diagram of generating a probability score for a machine translated output in a second language;
Fig. 3 illustrates a flow diagram of an automated system to determine a requirement to edit at least one machine translated sentence; and Fig. 4 illustrates a flowchart of a method for automated determination of a requirement to edit at least one translated sentence.
DETAILED DESCRIPTION OF THE INVENTION WITH NON-LIMITING EMBODIMENTS, EXAMPLES AND ACCOMPANYING DRAWINGS
For the purpose of promoting an understanding of the principles of the invention, reference will now be made to the embodiment illustrated in the drawings, and specific language will be used to describe the same. It will nevertheless be understood that no limitation of the scope of the invention is thereby intended. Such alterations and further modifications in the illustrated system, and such further applications of the principles of the invention as illustrated therein would be contemplated as would normally occur to one skilled in the art to which the invention relates. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skilled in the art. The system, methods, and examples provided herein are illustrative only and are not intended to be limiting.
The term "some" as used herein is to be understood as "none or one or more than one or all." Accordingly, the terms "none," "one," "more than one," "more than one, but not all" or "all" would all fall under the definition of "some." The term "some embodiments" may refer to no embodiments or to one embodiment or to several embodiments or to all embodiments, without departing from the scope of the present disclosure.
The terminology and structure employed herein is for describing, teaching, and illuminating some embodiments and their specific features. It does not in any way limit, restrict or reduce the spirit and scope of the claims or their equivalents.
More specifically, any terms used herein such as but not limited to "includes," "comprises," "has," "consists," and grammatical variants thereof do not specify an exact limitation or restriction and certainly do not exclude the possible addition of one or more features or elements, unless otherwise stated, and furthermore must not be taken to exclude the possible removal of one or more of the listed features and elements, unless otherwise stated with the limiting language "must comprise" or "needs to include." Whether or not a certain feature or element was limited to being used only once, either way, it may still be referred to as "one or more features" or "one or more elements" or "at least one feature" or "at least one element." Furthermore, the use of the terms "one or more" or "at least one" feature or element do not preclude there being none of that feature or element, unless otherwise specified by limiting language such as "there needs to be one or more ..." or "one or more element is required."
Unless otherwise defined, all terms, and especially any technical and/or scientific terms, used herein may be taken to have the same meaning as commonly understood by one having ordinary skills in the art. Reference is made herein to some "embodiments." It should be understood that an embodiment is an example of a possible implementation of any features and/or elements presented in the attached claims. Some embodiments have been described for the purpose of illuminating one or more of the potential ways in which the specific features and/or elements of the attached claims fulfill the requirements of uniqueness, utility and non-obviousness. Use of the phrases and/or terms including, but not limited to, "a first embodiment," "a further embodiment," "an alternate embodiment," "one embodiment," "an embodiment," "multiple embodiments," "some embodiments," "other embodiments," "further embodiment", "furthermore embodiment", "additional embodiment" or variants thereof do not necessarily refer to the same embodiments. Unless otherwise specified, one or more particular features and/or elements described in connection with one or more embodiments may be found in one embodiment, or may be found in more than
one embodiment, or may be found in all embodiments, or may be found in no embodiments. Although one or more features and/or elements may be described herein in the context of only a single embodiment, or alternatively in the context of more than one embodiment, or further alternatively in the context of all embodiments, the features and/or elements may instead be provided separately or in any appropriate combination or not at all. Conversely, any features and/or elements described in the context of separate embodiments may alternatively be realized as existing together in the context of a single embodiment.
Any particular and all details set forth herein are used in the context of some embodiments and therefore should not be necessarily taken as limiting factors to the attached claims. The attached claims and their legal equivalents can be realized in the context of embodiments other than the ones used as illustrative examples in the description below.
Embodiments of the present invention will be described below in detail with reference to the accompanying drawings.
The figures and the forgoing description give examples of embodiments. Those skilled in the art will appreciate that one or more of the described elements may well be combined into a single functional element. Alternatively, certain elements may be split into multiple functional elements. Elements from one embodiment may be added to another embodiment. For example, orders of processes described herein may be changed and are not limited to the manner described herein. Moreover, the actions of any flow diagram need not be implemented in the order shown; nor do all the acts necessarily need to be performed. Also, those acts that are not dependent on other acts may be performed in parallel with the other acts. The scope of the embodiments is by no means limited by these specific examples. Numerous variations, whether explicitly given in the specification or not, such as differences in structure, dimension, and use of material, are possible. According to an embodiment of the invention, an automated system (10) is disclosed. The automated system (01) is used to determine a requirement to
edit at least one translated sentence. The system (01) may comprise of an input module (04) to input at least one sentence in a first language. The system (01) may further comprise of a translator module (06) to translate the at least one sentence into the at least one translated sentence in a second language. The system (01) may further comprise of a probability processing module (10) to calculate a probability score of the at least one translated sentence to estimate quality of the at least one translated sentence. The system (01) may further comprise of a processor module (12) to determine the requirement to pre-edit or post-edit the at least one translated sentence on the basis of the estimated quality of the at least one translated sentence.
Fig. 1 illustrates a flow diagram of generating a machine translation output in a second language for a sentence input by a user (02) in a first language according to an embodiment of the present invention. The user may input at least one sentence in an input device (04) for generating a machine translation of the input sentence in the second language. According to an embodiment of the invention, the user (02) may also input a file comprising of text in a first language to be translated into the second language. The input device (04) may be a processor, a computer or any processing unit capable of enabling input of a plurality of sentences or at least one sentence to be translated, text files in any format, like PDF, .doc etc. to be translated, or an image to text converter. According to an embodiment of the invention, the input device (04) may be connected to another device, like a processor or an output device of any system, to provide automatic input to the input device (04). The input may be at least one sentence in a first language to be translated into the second language. The device may substitute the user (02). The user can select the second language for translation at the time of inputting the text in the first language. According to an embodiment of the invention, the selection of the second language for translation may also take place at the time of operating a translator module (06) for machine translation.
According to an embodiment of the invention, the translator module (06) may be configured to detect the first language inputted by the user (02) for
translation. The translator module (06) may be further configured to predefine the second language for translation for automated translation. In another embodiment of the invention, the user (02) may select the second language for translation in the translator module (06). In another embodiment of the invention, the translator module (06) may receive the input text in the first language directly from the user (02) for translation into the second language. The translator module (06) may provide translation of the first sentence into the second sentence on real-time basis.
According to an embodiment of the invention, the translator module (06) may be any module known in the state of the art. For example, the system (01) may comprise of a translator module like a Google translator, Microsoft MT-Engine, Babylon MT Engine, Anusaaraka MT-Engine, Moses MT-Toolkit, Joshua MT- Toolkit, any open source machine translator or any in-house translator specifically created for the purpose of the invention. The translator module (06) may be comprised of one or more than one of any machine translation (MT) engine(s) available for online and offline translation. According to an embodiment of the invention, the system (01) comprising the input device (04) configured to be operated by the user (02) enables an input of text in the first language into the translator module (06). The translator module (06) generates a MT-output (08), the translated sentence, in the second language. The system (01) may be configured to generate an automated MT-output (08) on real-time basis.
According to an embodiment of the invention, the user (02) may select any language existing in the world as its first language to input text in the input device (04) for machine translation. Similarly, the user (02) may also select any language, other than the first language, as its second language for translating the first language. According to a preferred embodiment of the invention, the sentences input by the user (02) may be in English, as a first language and the user (02) may select Hindi, as the second language or the MT-output (08), for translating the Hindi text.
Fig. 2 illustrates a flow diagram of generating a probability score for a machine translated (MT) output (08) in a second language according to an embodiment of the invention. The system (01) may be configured to generate the probability score of the MT-output (08) using a probability-processing module (10). In an embodiment of the invention, the probability-processing module (10) may comprise of any one of or a combination of the following language models: N-Gram or Markov chain or bigram or trigram or smoothing models.
Below is the description of the following language models: N-Gram Language Model
The N-Gram is a language model, which uses language corpus to predict the Nth word using previous N-l words. The collection of words and texts in the database is called corpus. The N-Gram model may assign a probability to the translated sentence and can also be used to assign a probability to the next word in an incomplete sentence. It is a probabilistic model, which uses probability to determine the next word. If English has 100000 words, the
probability of any word would be or .00001
^ J J 100000
Markov Chain
According to this language model, the assumption for the probability of a word depends only on the previous word. In case, there is a string of words
containing n words, like wl, w2, w3, w4 Wn. The conditional
probability of the word with respect to the previous word(s) is used to predict the upcoming word. In the above string, if we consider each word occurring in its correct location as an independent event then we might represent the probability as follows:-
P(Wi,W2,W3 ,W„-i,w„);
The chain rule of probability can be used to decompose this probability:
P( WJ1 ) = P(Wi) P(W2|Wi) P(W3| Wx2 ) P(W„| Wf"1 )
=n?=1P(Wr|UO
Where P (W2IW1) represents the conditional probability of word W2 with
respect to Wi and P (WslW^) represents the conditional probability of word
W3 with respect to Wi and W2 and so on.
Bigram Language Model
Calculating probabilities in the Markov Chain, like P(Wn|W1n_1), can be
replaced by bigram model for easier calculations. For Example:
My mother gifted me a dress on my " " The predicted word can
be "birthday". P(birthday | My mother gifted me a dress on my). We can
approximate this probability just by P(birthday | my). Hence the bigram model
instead of using the probability of a word given all the previous words
P(Wi|Wi-i). The bigram model can compute probability of a sentence like,
"My favorite subject is NLP": P(My favorite subject is NLP) = P(My | )
P(favorite | my) P(subject | favorite); P(is | subject) P(NLP | is).
Trigram Language Model
A trigram model looks just the same as a bigram model, except that we
condition on the two previous words.
Example: P(My favorite subject is NLP) = P(My | ) P(favorite | My
) P(subject I My favorite) P(is | favorite subject) P(NLP | subject is)
Limitation: The limitation of N-Gram model is that if any of the conditional
probability comes to be zero, the probability of the whole sentence becomes
zero though corpus is a very big collection of words, there can be a possibility
that the conational probability between two words is zero. In that case, the
probability of the whole sentence will become zero.
Smoothing
Smoothing may be introduced to overcome the limitation of N-Gram. When
the result of N-Gram model gives zero probability, we apply smoothing in to
revaluate that part of the N-gram, which gives zero or low probability.
Types of Smoothing
Add-One Smoothing: Add-one smoothing is one of the simplest method of
smoothing in which we add one to all the counts before we normalize them
into probabilities though it is generally not used, this method of smoothing
forms the base for the other smoothing algorithms. Applying add-one smoothing to compute unigram probabilities where the total occurrence(count)
of a word is divided by the total number of word tokens N: P(WX) ~
Eic(Wi)
e(Wx)
N
In case when c(Wx) becomes zero the probability will be zero. To solve this
problem the formula has been modified by adjusting count c*:
C* = (Q + 1) , where one is added to the count of each word type and a
N
normalization factor is introduced V is the vocabulary size i.e. the total
N+V J
number of unique words in the corpus. Since we are adding 1 to the count for each word type, the number of types must increase the total number of tokens. Example: Taking an example from Berkeley Restaurant project corpus of -10000 sentences. "I want to eat Chinese food lunch". The normal bigram
c(Wn— 1 I Wn)
probabilities are computed by the formula: P(Wn| Wn-i) = —: :—
c(Wn-l )
For add-one smoothed bigram counts we need to first augment the unigram count by the number of total word types in the vocabulary V: P* (Wn| Wn-i )
c(Wn-1 Wn)+1 , ,,,. ,. , .... ,. , ,. Tri , „
c(Wn-l)+V ' v,"w^ yin una v^vjipua^ IVJIVJ. IN
the unigram counts:
I 3437+ 1616 = 5053
Want 1215 + 1616 = 2931
To 3256+ 1616 = 4872
Eat 938+ 1616 = 2554
Chinese 213 + 1616= 1829
Food 1506+1616=3122
Lunch 459+1616=2075
Using the formula, the probabilities are as follows:
I .68
Want .42
To .69
Eat .37
Chinese .12
Food .48
Lunch .22
Backoff:
Backoff is based on trigram model. It is a nonlinear method introduced by Katz in 1987. It is based on the concept that if we don't have a particular trigram Wp to help us compute PCWp | W3) we can use bigram probability
Ji(W1 | W2) and it don't have bigram probability we can look to the unigram P(wi) . In Backoff model, we build an N-gram model based on (N-1) gram model. We "back-off to a lower order N-gram (bigram or unigram) if we have zero evidence for a higher-order N-gram (trigram). The trigram version of Backoff is:
P* (Wi| Wi-2 Wi-i )
(P(W; I Wi_2Wi_1)J ifc(Wi_2Wi_1Wi) > 0
and cCWi^Wi) > 0
v P(W;), otherwise
Kneser-Ney Smoothing: Kneser-Ney smoothing is an extension of absolute discounting which Kneser and Ney introduced in 1995. Inventors know that a lower order distribution is very important in the combined model only when few or more counts are present with the higher order distribution. For example, a bigram model say San Francisco, inventors know that FRANCISCO occurs only after a single word SAN. Since C(FRANCISCO) is high, the unigram probability P(FRANCISCO) will be high and an algorithm such as absolute discounting will assign a relatively high probability to the word FRANCISSO. But this probability should not be high since the word FRANCISCO follows a single word SAN. So the word FRANCISCO should receive a low unigram probability. In Kneser-Ney smoothing, we generalize this argument, not setting the unigram probability to be proportional t the number of occurrences of a word, but instead to the number of different words
that it follows. For example, for bigram model: £Wi ^PKntwj.-LWi) =
c(wL)
EvviC(wi)
The left-hand side of this equation is the unigram marginal for WJ of the smoothed bigram distribution Pkn, and the right-hand side is the unigram frequency of W.
Modified Kneser-Ney Smoothing: In the Kneser Ney smoothing, for bigram
model, the equation is:- P^Cw; | wjln+i) =
max{C(wj_n+1-D. 0)} , D ._., , , :_i N _ , , j_i N ,
L,^!) + SWlc(w1_n+1)N1 +M-« > P-(wil^+2 ), where
0<=D<=1
The count or zero whichever is maximum is taken so the original form used is:
PKn (Wi| Wl_n+1) = £wiC(wi-n+l) lf C(Wi-n+l) > °
Y (wj-n+i) PKn (w,| wl-n+2) if C(Wi_n+1) = 0, where Y(wj-n+i) is chosen to make the distribution sum to l.In modified Kneser-Ney smoothing, instead of using a single discount D for all non-zero counts as in Kneser-Ney smooting, three different parameters D1,D2 and D3+ in n-grams with one, two or three or more counts, respectively so the equation is:
P Cw I wi_1 ^ -c(wi-n+i)~D(c(wi-u+i))+A..ri-l \ p (w I wi-l *Kn lwi I wi-n+l) y—7^3 S +(wi-n+l) FKn lwi I wi-n+2)
{
0 ifc=0 Dl ifc = 1 D2 ifc = 2 D3+ ifc > 3.
According to an embodiment of the invention, the sentence translated in the second language by MT-Engines or any MT-Toolkit is used for calculating the probability of the machine-translated sentence in the probability-processing module (10). In an embodiment of the invention, the system (01) may be configured to comprise a MT-toolkit having the MT-engine for translating the sentence and the probability-processing module (10) for calculating the probability of the MT sentence using above described Language Models
(LM). We are using two thresholds for making decision of good and bad translation. In an embodiment of the invention, the probability-processing module (10) calculates the probability score by any one of the language models for the at least one translated sentence or by averaging probability scores calculated by any one of the language models for all the translated sentences or by averaging all the probability scores calculated by all the language models for the at least one translated sentence or all the translated sentences, or by any combination thereof. One value is the average of all translated sentences probability score of any one MT-engine and another one is the average of all MT-engines-translations and MT-Toolkits-translations scores. The system (01) may be configured to generate an automated probability score on real-time basis; the probability-processing module (10) may be configured to work in automated manner and to deliver output on real¬time basis.
Fig. 3 illustrates a flow diagram of an automated system (01) to determine a requirement to edit at least one machine-translated sentence according to an embodiment of the invention. The system may further comprise of a processor module (12) to determine the requirement to pre-edit or post-edit the at least one translated sentence on the basis of the estimated quality of the at least one translated sentence. In an embodiment of the invention, the processor module (12) compares the probability score with a pre-defined score (T). The pre¬defined score may be a threshold value stored in the system (01) or in the processor module (12). In an embodiment of the invention, the threshold values estimated using
£P_ £!" P(Wn)
following formula: T =—^—l— — where n = number of sentences and m
& n+m
= number of MT-Engine.
According to an embodiment of the invention, upon comparing the probability score of the Machine-translated sentence with the threshold value, the system (01) or the processor module (12) comprised in the system may determine which translated sentence is of good quality and which one should be
retranslated because of its bad translation. In one embodiment of the invention, the system (01) compares the machine-translated sentence's trigram or bigram probability score with above threshold, if sentence's score is greater than threshold then the translation is good else it should be retranslated. Formula used is as follows: If P(W1n)>=T, Good Translation, else, Bad translation.
According to an embodiment of the invention, the processor module (12) outputs a post-edit command (12a) if the probability score is greater than the pre-defined score (T), or if the machine-translated sentence falls in the good translation category.
According to an embodiment of the invention, the processor module (12) outputs a pre-edit command (12b) if the probability score is less than the pre-defined score (T) or if the machine-translated sentence falls in the bad translation category.
According to an embodiment of the invention, if the processor module (12) has given the post-edit command (12a), then the user (01) will make slight improvements in the machine-translated sentence and the post-edit sentence or final post-edit text (14) obtained after human translation is of human-translation quality acceptable for publication. According to an embodiment of the invention, if the processor module (12) has given the pre-edit command (12b), then the user (01) will make improvements in the text or the sentence of the first language and then input the pre-edited sentence of the first language in the input device (04) again for machine translation in the translator module (06). The system (01) will then again check the probability score and determine the requirement to post-edit (12b) or pre-edit (12b) the translated sentence. Finally, the good translation will send to post-editor for minor correction and bad translation will again be sent to pre-editor for make changes in their sentence structure and it will be re-translate by MT-engine. According to an embodiment of the invention, the processor module (12) comprised in the system (01) may be configured for automated determination of the requirement to post-edit (12b) or pre-edit (12b) the translated sentence
on real-time basis; the processor module (12) may be configured to work in automated manner and to deliver output on real-time basis. Fig. 4 illustrates a flowchart of a method (SI) for automated determination of a requirement to edit at least one translated sentence according to an embodiment of the invention. The method for automated determination of a requirement to edit at least one translated sentence is disclosed. The method (SI) has a step (S2) of inputting, by an input device (04), at least one sentence in a first language. The next step (S3) is of translating, by a translator module (06), the at least one sentence into the at least one translated sentence in a second language. The machine translated sentence (08) is then used in a step (S4) for calculating, by a probability processing module (10), a probability score of the at least one translated sentence to estimate quality of the at least one translated sentence. Next step (S5) is of determining, by a processor (12), the requirement to post-edit or pre-edit the at least one translated sentence on the basis of the estimated quality of the at least one translated sentence. In an embodiment of the invention, the determination step (S5) may be comprised of comparing the probability score with a pre-defined score (T). The pre-defined score (T) is calculated by formula (1): T
E"= E'= P(W")
= _i_i—j_i —where: n = number of sentences and m = number of MT-
n+m
Engine.
In an embodiment of the invention, if the probability score is greater than the threshold value at the comparison step (S5), the processor (12) outputs (S7) a post-edit command. The post-edit command means a good translation and the translated sentence will be sent to post-editor for minor corrections. According to an embodiment of the invention, if the probability score is less than the threshold value at the comparison step (S5), the processor (12) outputs (S6) a pre-edit command. If the processor module (12) has given the pre-edit command (12b), then the user (01) will make improvements in the text or the sentence of the first language and then input the pre-edited sentence of the first language in the input device (04) again for machine translation in
the translator module (06). The system (01) will then again check the probability score and determine the requirement to post-edit (12b) or pre-edit (12b) the translated sentence. Finally, the good translation will send to post-editor for minor correction and bad translation will again be sent to pre-editor for make changes in their sentence structure and it will be re-translate by MT-engine. For the translation falling under the good category, the user (01) may further make slight improvements in the machine-translated sentence and the post-edit sentence or final post-edit text (14) obtained after human translation is of human-translation quality acceptable for publication. According to an embodiment of the invention, the method (SI) to determine quality of the at least one translated sentence is estimated on a real-time basis. In an embodiment of the invention, the method (SI) determines quality of the at least one translated sentence is an automated method. In an embodiment of the invention, any one or more than one or all of the steps of inputting a sentence (SI), translating the input sentence (S2), calculating the probability score (S3), comparison and determination step (S5), commanding pre-editing (S6) or post-editing (S7) comprised in the method (SI) may be automated and may deliver output on real-time basis.
The figures and the forgoing description give examples of embodiments. Those skilled in the art will appreciate that one or more of the described elements may well be combined into a single functional element. Alternatively, certain elements may be split into multiple functional elements. Elements from one embodiment may be added to another embodiment. For example, orders of processes described herein may be changed and are not limited to the manner described herein. Moreover, the actions of any flow diagram need not be implemented in the order shown; nor do all the acts necessarily need to be performed. Also, those acts that are not dependent on other acts may be performed in parallel with the other acts. The scope of the embodiments is by no means limited by these specific examples. Numerous variations, whether explicitly given in the specification or not, such as differences in structure, dimension, and use of material, are possible.
We claim:
1. An automated system to determine a requirement to edit at least one
translated sentence, said system comprising:
an input module to input at least one sentence in a first language; a translator module to translate the at least one sentence into the at least one translated sentence in a second language; a probability processing module to calculate a probability score of the at least one translated sentence to estimate quality of the at least one translated sentence; and
a processor module to determine the requirement to pre-edit or post-edit the at least one translated sentence on the basis of the estimated quality of the at least one translated sentence.
2. The automated system to determine a requirement to edit at least one translated sentence as claimed in claim 1, wherein the processor module compares the probability score with a pre-defined score (T).
3. The automated system to determine a requirement to edit at least one translated sentence as claimed in claim 2, wherein the processor module outputs a post-edit command if the probability score is greater than the pre-defined score (T).
4. The automated system to determine a requirement to edit at least one translated sentence as claimed in claim 2, wherein the processor module outputs a pre-edit command if the probability score is less than the pre-defined score (T).
5. The automated system to determine a requirement to edit at least one translated sentence as claimed in claim 2, wherein the pre-defined score (T) is calculated by formula (1):
T_£P=1£jg1p(w°)
n+m
Where as :-n=no. of sentences m=no. of MT-Engine
6. The automated system to determine a requirement to edit at least one translated sentence as claimed in claim 1, wherein the quality of the at least one translated sentence is estimated on a real-time basis.
7. The automated system to determine a requirement to edit at least one translated sentence as claimed in claim 1, wherein the first language is English and the second language is Hindi.
8. The automated system to determine a requirement to edit at least one translated sentence as claimed in claim 1, wherein the probability processing module is trained on a language model.
9. The automated system to determine a requirement to edit at least one translated sentence as claimed in claim 8, wherein the language model comprising a N-Gram or Markov chain or bigram or trigram or smoothing models or any combination thereof.
10. The automated system to determine a requirement to edit at least one translated sentence as claimed in claim 9, wherein the probability processing module calculates the probability score by any one of the language models for the at least one translated sentence or by averaging probability scores calculated by any one of the language models for all the translated sentences or by averaging all the probability scores calculated by all the language models for the at least one translated sentence or all the translated sentences, or by any combination thereof.
11. A method for automated determination of a requirement to edit at least one translated sentence, said method comprising:
inputting, by an input device, at least one sentence in a first language; - translating, by a translator, the at least one sentence into the at least one translated sentence in a second language;
calculating, by a probability processing module, a probability score of the at least one translated sentence to estimate quality of the at least one translated sentence; and
determining, by a processor, the requirement to post-edit or pre-edit the at least one translated sentence on the basis of the estimated quality of the at least one translated sentence.
12. The method for automated determination of a requirement to edit at least one translated sentence as claimed in claim 11, wherein the processor compares the probability score with a pre-defined score (T).
13. The method for automated determination of a requirement to edit at least one translated sentence as claimed in claim 12, wherein the processor outputs a pre-edit command if the probability score is less than the pre-defined score (T).
14. The method for automated determination of a requirement to edit at least one translated sentence as claimed in claim 12, wherein the processor outputs a post-edit command if the probability score is greater than the pre-defined score (T).
15. The method for automated determination of a requirement to edit at least one translated sentence as claimed in claim 12, wherein the pre-defined score (T) is calculated by formula (1):
T_£P=1£jg1p(wg)
n+m
Where as :-n=no. of sentences m=no. of MT-Engine
16. The method for automated determination of a requirement to edit at least one translated sentence as claimed in claim 11, wherein the first language is English and the second language is Hindi.
17. The method for automated determination of a requirement to edit at least one translated sentence as claimed in claim 11, wherein the quality of the at least one translated sentence is estimated on a real-time basis.
18. The method for automated determination of a requirement to edit at least one translated sentence as claimed in claim 11, wherein the probability
processing module is trained on a language model to calculate the probability score of the at least one translated sentence.
19. The method for automated determination of a requirement to edit at least one translated sentence as claimed in claim 18, wherein the language model comprises of a N-Gram or Markov chain or bigram or trigram or smoothing models or any combination thereof.
20. The method for automated determination of a requirement to edit at least one translated sentence as claimed in claim 19, wherein the probability processing module calculates the probability score by any one of the language models for the at least one translated sentence or by averaging probability scores calculated by any one of the language models for all the translated sentences or by averaging all the probability scores calculated by all the language models for the at least one translated sentence or all the translated sentences, or by any combination thereof.
| # | Name | Date |
|---|---|---|
| 1 | 202111003242-STATEMENT OF UNDERTAKING (FORM 3) [23-01-2021(online)].pdf | 2021-01-23 |
| 2 | 202111003242-FORM 1 [23-01-2021(online)].pdf | 2021-01-23 |
| 3 | 202111003242-FIGURE OF ABSTRACT [23-01-2021(online)].jpg | 2021-01-23 |
| 4 | 202111003242-DRAWINGS [23-01-2021(online)].pdf | 2021-01-23 |
| 5 | 202111003242-DECLARATION OF INVENTORSHIP (FORM 5) [23-01-2021(online)].pdf | 2021-01-23 |
| 6 | 202111003242-COMPLETE SPECIFICATION [23-01-2021(online)].pdf | 2021-01-23 |
| 7 | 202111003242-Proof of Right [01-02-2021(online)].pdf | 2021-02-01 |
| 8 | 202111003242-FORM-26 [01-02-2021(online)].pdf | 2021-02-01 |
| 9 | 202111003242-ENDORSEMENT BY INVENTORS [01-02-2021(online)].pdf | 2021-02-01 |
| 10 | 202111003242-Power of Attorney-040221.pdf | 2021-10-19 |
| 11 | 202111003242-OTHERS-040221.pdf | 2021-10-19 |
| 12 | 202111003242-Form 5-040221.pdf | 2021-10-19 |
| 13 | 202111003242-Correspondence-040221.pdf | 2021-10-19 |
| 14 | 202111003242-FORM 18 [08-10-2022(online)].pdf | 2022-10-08 |
| 15 | 202111003242-FER.pdf | 2023-01-25 |
| 16 | 202111003242-OTHERS [11-03-2023(online)].pdf | 2023-03-11 |
| 17 | 202111003242-MARKED COPIES OF AMENDEMENTS [11-03-2023(online)].pdf | 2023-03-11 |
| 18 | 202111003242-FORM 13 [11-03-2023(online)].pdf | 2023-03-11 |
| 19 | 202111003242-FER_SER_REPLY [11-03-2023(online)].pdf | 2023-03-11 |
| 20 | 202111003242-EDUCATIONAL INSTITUTION(S) [11-03-2023(online)].pdf | 2023-03-11 |
| 21 | 202111003242-CLAIMS [11-03-2023(online)].pdf | 2023-03-11 |
| 22 | 202111003242-AMMENDED DOCUMENTS [11-03-2023(online)].pdf | 2023-03-11 |
| 23 | 202111003242-FORM-8 [31-10-2024(online)].pdf | 2024-10-31 |
| 1 | postreditingtranslationE_24-01-2023.pdf |