Sign In to Follow Application
View All Documents & Correspondence

Semantic Similarity Detection Between English Textual Information With Other Language

Abstract: SEMANTIC SIMILARITY DETECTION BETWEEN ENGLISH TEXTUAL INFORMATION WITH OTHER LANGUAGE Abstract One embodiment of the current disclosure is a system for identifying semantic similarities between English and non-English text, which may include a processor. In certain implementations, RAM is included. In certain implementations, a natural language processing component is also incorporated. This component would examine English textual material and provide a collection of characteristics based on the semantics of the English textual information. The module for processing natural language would be kept in memory and could be run by the CPU. A translation module may be placed in memory and run by the processor to translate the English textual material into other languages. In certain implementations, this component is also present. Embodiments may further comprise a semantic similarity module that is stored in memory and is executable by the processor. The goal of this section is to examine the translated text and the collection of characteristics created by the natural language processing module to ascertain the degree of semantic similarity between the two sets of text. Fig. 1

Get Free WhatsApp Updates!
Notices, Deadlines & Correspondence

Patent Information

Application #
Filing Date
22 March 2023
Publication Number
19/2023
Publication Type
INA
Invention Field
COMPUTER SCIENCE
Status
Email
Parent Application

Applicants

BANASTHALI VIDYAPITH
BANASTHALI VIDYAPITH, P.O. BANASTHALI, BANASTHALI, RAJASTHAN, INDIA, 304022, JAIPUR

Inventors

1. GEETHA YADAV
BANASTHALI VIDYAPITH, P.O. BANASTHALI, BANASTHALI, RAJASTHAN, INDIA, 304022, JAIPUR
2. DR. SNEHA ASOPA
BANASTHALI VIDYAPITH, P.O. BANASTHALI, BANASTHALI, RAJASTHAN, INDIA, 304022, JAIPUR

Claims

1. A system for semantic similarity detection between English textual information and other languages, comprising: a processor; a memory; a natural language processing module stored in the memory and executable by the processor to analyse the English textual information and generate a set of features based on the semantics of the English textual information; a translation module stored in the memory and executable by the processor to translate the English textual information into other languages; and a semantic similarity module stored in the memory and executable by the processor to analyse the translated text and the set of features generated by the natural language processing module to determine the semantic similarity between the English textual information and the translated text.

2. The system of claim 1, wherein the natural language processing module further comprises: a tokenization module to divide the English textual information into a set of tokens; a part-of-speech tagging module to identify the part of speech of each token; a dependency parsing module to identify the relationships between the tokens; a named entity recognition module to identify named entities in the English textual information; a semantic role labelling module to identify the semantic roles of the tokens.

3. The system of claim 1, wherein the translation module further comprise a machine translation module to translate the English textual information into other languages.

4. The system of claim 1, wherein the semantic similarity module further comprises a semantic vector representation module to generate a semantic vector representation of the English textual information and the translated text, and a similarity calculation module to determine the semantic similarity between the English textual information and the translated text based on the semantic vector representations.

5. A method for semantic similarity detection between English textual information and other languages, comprising: analysing the English textual information using a natural language processing module to generate a set of features based on the semantics of the English textual information; translating the English textual information into other languages using a translation module; and analysing the translated text and the set of features generated by the natural language processing module using a semantic similarity module to determine the semantic similarity between the English textual information and the translated text.

6. The method of claim 5, wherein analysing the English textual information using the natural language processing module further comprises: dividing the English textual information into a set of tokens using a tokenization module; identifying the part of speech of each token using a part-of-speech tagging module; identifying the relationships between the tokens using a dependency parsing module; identifying named entities in the English textual information using a named entity recognition module; identifying the semantic roles of the tokens using a semantic role labelling module. 7.The method of claim 5, wherein translating the English textual information into other languages using the translation module that use a machine translation module to translate the English textual information into other languages.

8. The method of claim 5, wherein analysing the translated text and the set of features generated by the natural language processing module using the semantic similarity module further configured to generate a semantic vector representation of the English textual information and the translated text using a semantic vector representation module, and determine the semantic similarity between the English textual information and the translated text based on the semantic vector representations using a similarity calculation module. SEMANTIC SIMILARITY DETECTION BETWEEN ENGLISH TEXTUAL INFORMATION WITH OTHER LANGUAGE Abstract One embodiment of the current disclosure is a system for identifying semantic similarities between English and non-English text, which may include a processor. In certain implementations, RAM is included. In certain implementations, a natural language processing component is also incorporated. This component would examine English textual material and provide a collection of characteristics based on the semantics of the English textual information. The module for processing natural language would be kept in memory and could be run by the CPU. A translation module may be placed in memory and run by the processor to translate the English textual material into other languages. In certain implementations, this component is also present. Embodiments may further comprise a semantic similarity module that is stored in memory and is executable by the processor. The goal of this section is to examine the translated text and the collection of characteristics created by the natural language processing module to ascertain the degree of semantic similarity between the two sets of text. Fig. 1 , Claims:Claims :

1. A system for semantic similarity detection between English textual information and other languages, comprising: a processor; a memory; a natural language processing module stored in the memory and executable by the processor to analyse the English textual information and generate a set of features based on the semantics of the English textual information; a translation module stored in the memory and executable by the processor to translate the English textual information into other languages; and a semantic similarity module stored in the memory and executable by the processor to analyse the translated text and the set of features generated by the natural language processing module to determine the semantic similarity between the English textual information and the translated text.

2. The system of claim 1, wherein the natural language processing module further comprises: a tokenization module to divide the English textual information into a set of tokens; a part-of-speech tagging module to identify the part of speech of each token; a dependency parsing module to identify the relationships between the tokens; a named entity recognition module to identify named entities in the English textual information; a semantic role labelling module to identify the semantic roles of the tokens.

3. The system of claim 1, wherein the translation module further comprise a machine translation module to translate the English textual information into other languages.

4. The system of claim 1, wherein the semantic similarity module further comprises a semantic vector representation module to generate a semantic vector representation of the English textual information and the translated text, and a similarity calculation module to determine the semantic similarity between the English textual information and the translated text based on the semantic vector representations.

5. A method for semantic similarity detection between English textual information and other languages, comprising: analysing the English textual information using a natural language processing module to generate a set of features based on the semantics of the English textual information; translating the English textual information into other languages using a translation module; and analysing the translated text and the set of features generated by the natural language processing module using a semantic similarity module to determine the semantic similarity between the English textual information and the translated text.

6. The method of claim 5, wherein analysing the English textual information using the natural language processing module further comprises: dividing the English textual information into a set of tokens using a tokenization module; identifying the part of speech of each token using a part-of-speech tagging module; identifying the relationships between the tokens using a dependency parsing module; identifying named entities in the English textual information using a named entity recognition module; identifying the semantic roles of the tokens using a semantic role labelling module. 7.The method of claim 5, wherein translating the English textual information into other languages using the translation module that use a machine translation module to translate the English textual information into other languages.

8. The method of claim 5, wherein analysing the translated text and the set of features generated by the natural language processing module using the semantic similarity module further configured to generate a semantic vector representation of the English textual information and the translated text using a semantic vector representation module, and determine the semantic similarity between the English textual information and the translated text based on the semantic vector representations using a similarity calculation module.

Specification

Description:SEMANTIC SIMILARITY DETECTION BETWEEN ENGLISH TEXTUAL INFORMATION WITH OTHER LANGUAGE
Field of the Invention
[0001] This invention relates to information retrieval and management in particular to a system and method for use of machine learning (ML) and natural language processing (NLP) for determining a semantic similarity between English textual information with other language information.
Background
[0002] The background description includes information that may be useful in understanding the present invention. It is not an admission that any of the information provided herein is prior art or relevant to the presently claimed invention, or that any publication specifically or implicitly referenced is prior art.
[0003] Semantic similarity detection has numerous practical applications, including information retrieval, plagiarism detection, machine translation, and sentiment analysis. For example, search engines use semantic similarity algorithms to retrieve relevant results for a given query, while plagiarism detection systems use semantic similarity algorithms to identify instances of text reuse. The background of research on semantic similarity detection can be traced back to the early days of natural language processing and information retrieval. Researchers have long been interested in developing algorithms and systems that can accurately determine the semantic similarity between two pieces of text.
[0004] In the past, the focus of semantic similarity research was primarily on lexical and syntactic features of the text, such as word overlap, word order, and grammatical structure. However, these approaches often failed to capture the more subtle nuances of meaning and context that are crucial for accurately measuring semantic similarity. Patent literature had disclosed various technique for semantic similarity search. Few of them discussed below.
[0005] The US8170969B2 (By: SIEMENS) - An overall semantic similarity score value between pairs of named entities in a text corpus is obtained by calculating for at least one pair of named entities a plurality of corresponding pair similarity score values according to a first and at least a second classifier using electronic information sources. Each pair similarity score value of the pair of named entities per classifier is normalized by calculating a rank list per classifier, for example, for each named entity. The rank list holds each pair of named entities of the text corpus, wherein a rank of each pair of named entities within the rank list reflects the respective pair similarity score value. Further an arithmetic mean of the normalized pair similarity score value of each pair of named entities is calculated to provide the overall semantic similarity score value.
[0006] The WO2019003069A1 (By: IBM) - Method and system are provided for adaptive evaluation of meta-relationships in semantic graphs. The method includes providing a semantic graph based on a knowledge base in which concepts in the form of graph nodes are linked by semantic relationships in the form of graph edges. Metadata is encoded in the edges and nodes of the semantic graph, of weightings for measuring a meta-relationship, wherein the meta-relationship applies to the concepts of the semantic graph and is independent of the semantic relationship defined by the edges of the semantic graph. A graph activation is carried out for an input context relating to one or more concepts of the semantic graph, wherein the weightings are applied to a spreading activation signal through the semantic graph to produce a measure of the meta-relationship for a sub-set of concepts of the semantic graph.
[0007] The AU2021200962B2 (By: ADOBE SYSTEMS) - Techniques and systems are described for performing semantic text searches. A semantic text-searching solution uses a machine learning system (such as a deep learning system) to determine associations between the semantic meanings of words. These associations are not limited by the spelling, syntax, 5 grammar, or even definition of words. Instead, the associations can be based on the context in which characters, words, and/or phrases are used in relation to one another. In response to detecting a request to locate text within an electronic document associated with a keyword, the semantic text-searching solution can return strings within the document that have matching and/or related semantic meanings or contexts, in addition to exact matches (e.g., string matches) within the document. The semantic text 0 searching solution can then output an indication of the matching strings. Docket No. P9700-US (102582-647310-001400US) 4/12 Semantic Search System Search Input 400 Detection Engine Token Engine Tokens Representation Engine Representations Match Engine Matches Output Engine Search Output
[0008] The WO2021158369A1 (By: MICROSOFT TECHNOLOGY LICENSING) - A technique is described herein for processing a given query item in a latency-efficient and resource-efficient manner. The technique uses a first transformer-based encoder to transform the given query item into an encoded query item. In one case, the given query item is an expression that includes one or more query-expression linguistic tokens. The technique includes a second transformer-based encoder for transforming a given target item into an encoded target item. The given target item may likewise correspond to an expression that includes one or more target-expression linguistic tokens. A similarity-assessing mechanism then assesses the semantic similarity between the given query item and the given target item based on the encoded query item and the encoded target item. Each transformer-based encoder uses one or more self-attention mechanisms. The second transformer-based encoder can optionally perform its work in an offline manner, prior to receipt of the given query item.
[0009] However, these techniques are associated with several limitations such as non-compatibility with other languages, lower efficiency, etc. Thus, there is need of technological advancement in this domain.

Summary
[00010] The following presents a simplified summary of various aspects of this disclosure in order to provide a basic understanding of such aspects. This summary is not an extensive overview of all contemplated aspects, and is intended to neither identify key or critical elements nor delineate the scope of such aspects. Its purpose is to present some concepts of this disclosure in a simplified form as a prelude to the more detailed description that is presented later.
[00011] The following paragraphs provide additional support for the claims of the subject application.
This invention relates to information retrieval and management in particular to a system and method for use of machine learning (ML) and natural language processing (NLP) for determining a semantic similarity between English textual information with other language information.
[00012] Embodiments of the present disclosure may include a system for semantic similarity detection between English textual information and other languages, wherein the system including a processor and a memory. Embodiments may also include a natural language processing module stored in the memory and executable by the processor to analyse the English textual information and generate a set of features based on the semantics of the English textual information.
[00013] Embodiments may also include a translation module stored in the memory and executable by the processor to translate the English textual information into other languages. Embodiments may also include a semantic similarity module stored in the memory and executable by the processor to analyse the translated text and the set of features generated by the natural language processing module to determine the semantic similarity between the English textual information and the translated text.
[00014] In some embodiments, the natural language processing module may include a tokenization module to divide the English textual information into a set of tokens. Embodiments may also include a part-of-speech tagging module to identify the part of speech of each token. Embodiments may also include a dependency parsing module to identify the relationships between the tokens. Embodiments may also include a named entity recognition module to identify named entities in the English textual information. Embodiments may also include a semantic role labelling module to identify the semantic roles of the tokens.
[00015] In some embodiments, the translation module may include a machine translation module to translate the English textual information into other languages. In some embodiments, the semantic similarity module may include a semantic vector representation module to generate a semantic vector representation of the English textual information and the translated text, and a similarity calculation module to determine the semantic similarity between the English textual information and the translated text based on the semantic vector representations.
[00016] Embodiments of the present disclosure may also include a method for semantic similarity detection between English textual information and other languages, including analysing the English textual information using a natural language processing module to generate a set of features based on the semantics of the English textual information. Embodiments may also include translating the English textual information into other languages using a translation module. Embodiments may also include analysing the translated text and the set of features generated by the natural language processing module using a semantic similarity module to determine the semantic similarity between the English textual information and the translated text.
[00017] Embodiments may also include analysing the English textual information using the natural language processing module may include dividing the English textual information into a set of tokens using a tokenization module. Embodiments may also include identifying the part of speech of each token using a part-of-speech tagging module.
[00018] Embodiments may also include identifying the relationships between the tokens using a dependency parsing module. Embodiments may also include identifying named entities in the English textual information using a named entity recognition module. Embodiments may also include identifying the semantic roles of the tokens using a semantic role labelling module.
[00019] Embodiments may also include translating the English textual information into other languages using the translation module that use a machine translation module to translate the English textual information into other languages. Embodiments may also include analysing the translated text and the set of features generated by the natural language processing module using the semantic similarity module further configured to generate a semantic vector representation of the English textual information and the translated text using a semantic vector representation module, and determine the semantic similarity between the English textual information and the translated text based on the semantic vector representations using a similarity calculation module.
Brief Description of the Drawings
[00020] The features and advantages of the present disclosure would be more clearly understood from the following description taken in conjunction with the accompanying drawings in which:
[00021] FIG. 1 is a block diagram illustrating a system for semantic similarity detection between English textual information and other languages, according to some embodiments of the present disclosure.
[00022] FIG. 2 is a modified block diagram further illustrating the system (from FIG. 1) for semantic similarity detection between English textual information and other languages, according to some embodiments of the present disclosure.
[00023] FIG. 3 is a flowchart illustrating a method for semantic similarity detection between English textual information and other languages, according to some embodiments of the present disclosure.
[00024] FIG. 4 is a detailed flowchart further illustrating the method (from FIG. 3) for semantic similarity detection between English textual information and other languages, according to some embodiments of the present disclosure.
Detailed Description
[00025] In the following detailed description of the invention, reference is made to the accompanying drawings that form a part hereof, and in which is shown, by way of illustration, specific embodiments in which the invention may be practiced. In the drawings, like numerals describe substantially similar components throughout the several views. These embodiments are described in sufficient detail to claim those skilled in the art to practice the invention. Other embodiments may be utilized and structural, logical, and electrical changes may be made without departing from the scope of the present invention. The following detailed description is, therefore, not to be taken in a limiting sense, and the scope of the present invention is defined only by the appended claims and equivalents thereof.
[00026] The use of the terms “a” and “an” and “the” and “at least one” and similar referents in the context of describing the invention (especially in the context of the following claims) are to be construed to cover both the singular and the plural, unless otherwise indicated herein or clearly contradicted by context. The use of the term “at least one” followed by a list of one or more items (for example, “at least one of A and B”) is to be construed to mean one item selected from the listed items (A or B) or any combination of two or more of the listed items (A and B), unless otherwise indicated herein or clearly contradicted by context. The terms “comprising,” “having,” “including,” and “containing” are to be construed as open-ended terms (i.e., meaning “including, but not limited to,”) unless otherwise noted. Recitation of ranges of values herein are merely intended to serve as a shorthand method of referring individually to each separate value falling within the range, unless otherwise indicated herein, and each separate value is incorporated into the specification as if it were individually recited herein. All methods described herein can be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by context. The use of any and all examples, or exemplary language (e.g., “such as”) provided herein, is intended merely to better illuminate the invention and does not pose a limitation on the scope of the invention unless otherwise claimed. No language in the specification should be construed as indicating any non-claimed element as essential to the practice of the invention.
[00027] This invention relates to information retrieval and management in particular to a system and method for use of machine learning (ML) and natural language processing (NLP) for determining a semantic similarity between English textual information with other language information.
[00028] FIG. 1 is a block diagram that describes a system 100 for semantic similarity detection between English textual information and other languages, according to some embodiments of the present disclosure. In some embodiments, the system 100 may include a processor 110, a memory 120, a natural language processing module 130 stored in the memory 120 and executable by the processor 110 to analyse the English textual information and generate a set of features based on the semantics of the English textual information, a translation module 140 stored in the memory 120 and executable by the processor 110 to translate the English textual information into other languages, and a semantic similarity module 150 stored in the memory 120 and executable by the processor 110 to analyse the translated text and the set of features generated by the natural language processing module 130 to determine the semantic similarity between the English textual information and the translated text.
[00029] In some embodiments, the translation module 140 may include a machine translation module to translate the English textual information into other languages. In some embodiments, the semantic similarity module 150 may also include a semantic vector representation module to generate a semantic vector representation of the English textual information and the translated text, and a similarity calculation module to determine the semantic similarity between the English textual information and the translated text based on the semantic vector representations.
[00030] FIG. 2 is a modified block diagram that further describes the system 100 (from FIG. 1) for semantic similarity detection between English textual information and other languages, according to some embodiments of the present disclosure. In some embodiments, the natural language processing module 130 can be arranged to analyse the English textual information and generate a set of features based on the semantics of the English textual information.
[00031] FIG. 3 is a flowchart that describes a method for semantic similarity detection between English textual information and other languages, according to some embodiments of the present disclosure. In some embodiments, at 310, the method may include analysing the English textual information using the natural language processing module 130 to generate a set of features based on the semantics of the English textual information. At 320, the method may include translating the English textual information into other languages using the translation module 140. At 330, the method may include analysing the translated text and the set of features generated by the natural language processing module 130 using a semantic similarity module 150 to determine the semantic similarity between the English textual information and the translated text.
[00032] In some embodiments, translating the English textual information into other languages using the translation module 140 that use a machine translation module to translate the English textual information into other languages. In some embodiments, analysing the translated text and the set of features generated by the natural language processing module 130 using the semantic similarity module 150 further configured to generate a semantic vector representation of the English textual information and the translated text using a semantic vector representation module, and determine the semantic similarity between the English textual information and the translated text based on the semantic vector representations using a similarity calculation module.
[00033] FIG. 4 is a detailed flowchart that further describes the method from FIG. 3 for semantic similarity detection between English textual information and other languages, according to some embodiments of the present disclosure. In some embodiments, analysing the English textual information using the natural language processing module 130 further comprises, the method may include 410 to 450. At 410, the method may include; dividing the English textual information; at 420, identifying the part of speech; at 430, identifying the relationships; at 440, identifying named entities in the English textual information; at 450, identifying the semantic roles of the tokens.
[00034] Embodiments of the current disclosure may incorporate the system for semantic similarity identification between English textual material and other languages, wherein the system comprising the processor 110 and the memory. Embodiments may additionally comprise the natural language processing module 130 stored in the memory 120 and executable by the processor 110 to analyse the English textual information and create a set of characteristics based on the semantics of the English textual Information.
[00035] Embodiments may additionally contain the translation module 140 stored in the memory 120 and executable by the processor 110 to convert the English textual material into other languages. Embodiments may also include the semantic similarity module 150 stored in the memory 120 and executable by the processor 110 to analyse the translated text and the set of features generated by the natural language processing module 110 to determine the semantic similarity between the English textual information and the translated text.
[00036] In certain implementations, the natural language processing module 130 may comprise a tokenization module 220 to partition the English textual information into a collection of tokens. Embodiments may additionally contain a part-of-speech tagging module to identify the part of speech of each token. Embodiments may additionally contain a dependency parsing module 230 to determine the links between the tokens. Embodiments may additionally contain a named entity recognition module to recognise named entities in the English textual material. Embodiments may additionally incorporate a semantic role labelling module 240 to determine the semantic roles of the tokens.
[00037] In other examples, the translation module 140 may contain a machine translation module to convert the English textual material into other languages. In some embodiments, the semantic similarity module 150 may include a semantic vector representation module to generate a semantic vector representation of the English textual information and the translated text, and a similarity calculation module to determine the semantic similarity between the English textual information and the translated text based on the semantic vector representations.
[00038] Embodiments of the present disclosure may also include a method for semantic similarity detection between English textual information and other languages, including analysing the English textual information using the natural language processing module 130 to generate a set of features based on the semantics of the English textual information. Embodiments may additionally comprise converting the English textual material into other languages utilising the translation module 140. Embodiments may additionally comprise comparing the translated text and the set of characteristics created by the natural language processing module 130 using the semantic similarity module 150 to identify the semantic similarity between the English textual information and the translated text.
[00039] Analysing the English textual information using the natural language processing module 130 may also be part of certain embodiments. Analysing the English textual information may also entail separating the English textual information into a collection of tokens using the tokenization module 220. A part-of-speech tagging module may also be used in certain embodiments in order to determine the grammatical category that each token belongs to.
[00040] Identifying the links between the tokens by using the dependency parsing module 230 is another possible step that may be included in embodiments. Identifying named entities in the English textual information with the use of a named entity recognition module is another possible aspect of embodiments. Identifying the semantic roles of the tokens using the semantic role labelling module 240 is another possible step that may be included in embodiments.
[00041] In certain embodiments, the English textual information may also be translated into other languages using a machine translation module that is part of the translation module 140 that is used to translate the English textual information into other languages. Analysing the translated text and the set of features generated by the natural language processing module 130 using the semantic similarity module 150 that is further configured to generate a semantic vector representation of the English textual information and the translated text using a semantic vector representation module, and determining the semantic similarity between the English textual information and the translated text based on the semantic vector representation may also be included in some embodiments.
[00042] Example embodiments herein have been described above with reference to block diagrams and flowchart illustrations of methods and apparatuses. It will be understood that each block of the block diagrams and flowchart illustrations, and combinations of blocks in the block diagrams and flowchart illustrations, respectively, can be implemented by various means including hardware, software, firmware, and a combination thereof. For example, in one embodiment, each block of the block diagrams and flowchart illustrations, and combinations of blocks in the block diagrams and flowchart illustrations can be implemented by computer program instructions. These computer program instructions may be loaded onto a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions which execute on the computer or other programmable data processing apparatus create means for implementing the functions specified in the flowchart block or blocks.
[00043] Throughout the present disclosure, the term ‘processing means’ or ‘microprocessor’ or ‘processor’ or ‘processors’ includes, but is not limited to, a general purpose processor (such as, for example, a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a microprocessor implementing other types of instruction sets, or a microprocessor implementing a combination of types of instruction sets) or a specialized processor (such as, for example, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), or a network processor).
[00044] The term “non-transitory storage device” or “storage” or “memory,” as used herein relates to a random access memory, read only memory and variants thereof, in which a computer can store data or software for any duration.
[00045] Operations in accordance with a variety of aspects of the disclosure is described above would not have to be performed in the precise order described. Rather, various steps can be handled in reverse order or simultaneously or not at all.
[00046] While several implementations have been described and illustrated herein, a variety of other means and/or structures for performing the function and/or obtaining the results and/or one or more of the advantages described herein may be utilized, and each of such variations and/or modifications is deemed to be within the scope of the implementations described herein. More generally, all parameters, dimensions, materials, and configurations described herein are meant to be exemplary and that the actual parameters, dimensions, materials, and/or configurations will depend upon the specific application or applications for which the teachings is/are used. Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific implementations described herein. It is, therefore, to be understood that the foregoing implementations are presented by way of example only and that, within the scope of the appended claims and equivalents thereto, implementations may be practiced otherwise than as specifically described and claimed. Implementations of the present disclosure are directed to each individual feature, system, article, material, kit, and/or method described herein. In addition, any combination of two or more such features, systems, articles, materials, kits, and/or methods, if such features, systems, articles, materials, kits, and/or methods are not mutually inconsistent, is included within the scope of the present disclosure.

Claims
I/We Claim:
1. A system for semantic similarity detection between English textual information and other languages, comprising:
a processor;
a memory;
a natural language processing module stored in the memory and executable by the processor to analyse the English textual information and generate a set of features based on the semantics of the English textual information;
a translation module stored in the memory and executable by the processor to translate the English textual information into other languages; and
a semantic similarity module stored in the memory and executable by the processor to analyse the translated text and the set of features generated by the natural language processing module to determine the semantic similarity between the English textual information and the translated text.

2. The system of claim 1, wherein the natural language processing module further comprises:
a tokenization module to divide the English textual information into a set of tokens;
a part-of-speech tagging module to identify the part of speech of each token;
a dependency parsing module to identify the relationships between the tokens;
a named entity recognition module to identify named entities in the English textual information;
a semantic role labelling module to identify the semantic roles of the tokens.

3. The system of claim 1, wherein the translation module further comprise a machine translation module to translate the English textual information into other languages.

4. The system of claim 1, wherein the semantic similarity module further comprises a semantic vector representation module to generate a semantic vector representation of the English textual information and the translated text, and a similarity calculation module to determine the semantic similarity between the English textual information and the translated text based on the semantic vector representations.

5. A method for semantic similarity detection between English textual information and other languages, comprising:
analysing the English textual information using a natural language processing module to generate a set of features based on the semantics of the English textual information;
translating the English textual information into other languages using a translation module; and
analysing the translated text and the set of features generated by the natural language processing module using a semantic similarity module to determine the semantic similarity between the English textual information and the translated text.

6. The method of claim 5, wherein analysing the English textual information using the natural language processing module further comprises:
dividing the English textual information into a set of tokens using a tokenization module;
identifying the part of speech of each token using a part-of-speech tagging module;
identifying the relationships between the tokens using a dependency parsing module;
identifying named entities in the English textual information using a named entity recognition module;
identifying the semantic roles of the tokens using a semantic role labelling module.

7.The method of claim 5, wherein translating the English textual information into other languages using the translation module that use a machine translation module to translate the English textual information into other languages.

8. The method of claim 5, wherein analysing the translated text and the set of features generated by the natural language processing module using the semantic similarity module further configured to generate a semantic vector representation of the English textual information and the translated text using a semantic vector representation module, and determine the semantic similarity between the English textual information and the translated text based on the semantic vector representations using a similarity calculation module.

SEMANTIC SIMILARITY DETECTION BETWEEN ENGLISH TEXTUAL INFORMATION WITH OTHER LANGUAGE
Abstract
One embodiment of the current disclosure is a system for identifying semantic similarities between English and non-English text, which may include a processor. In certain implementations, RAM is included. In certain implementations, a natural language processing component is also incorporated. This component would examine English textual material and provide a collection of characteristics based on the semantics of the English textual information. The module for processing natural language would be kept in memory and could be run by the CPU. A translation module may be placed in memory and run by the processor to translate the English textual material into other languages. In certain implementations, this component is also present. Embodiments may further comprise a semantic similarity module that is stored in memory and is executable by the processor. The goal of this section is to examine the translated text and the collection of characteristics created by the natural language processing module to ascertain the degree of semantic similarity between the two sets of text.
Fig. 1 , Claims:Claims
I/We Claim:
1. A system for semantic similarity detection between English textual information and other languages, comprising:
a processor;
a memory;
a natural language processing module stored in the memory and executable by the processor to analyse the English textual information and generate a set of features based on the semantics of the English textual information;
a translation module stored in the memory and executable by the processor to translate the English textual information into other languages; and
a semantic similarity module stored in the memory and executable by the processor to analyse the translated text and the set of features generated by the natural language processing module to determine the semantic similarity between the English textual information and the translated text.

2. The system of claim 1, wherein the natural language processing module further comprises:
a tokenization module to divide the English textual information into a set of tokens;
a part-of-speech tagging module to identify the part of speech of each token;
a dependency parsing module to identify the relationships between the tokens;
a named entity recognition module to identify named entities in the English textual information;
a semantic role labelling module to identify the semantic roles of the tokens.

3. The system of claim 1, wherein the translation module further comprise a machine translation module to translate the English textual information into other languages.

4. The system of claim 1, wherein the semantic similarity module further comprises a semantic vector representation module to generate a semantic vector representation of the English textual information and the translated text, and a similarity calculation module to determine the semantic similarity between the English textual information and the translated text based on the semantic vector representations.

5. A method for semantic similarity detection between English textual information and other languages, comprising:
analysing the English textual information using a natural language processing module to generate a set of features based on the semantics of the English textual information;
translating the English textual information into other languages using a translation module; and
analysing the translated text and the set of features generated by the natural language processing module using a semantic similarity module to determine the semantic similarity between the English textual information and the translated text.

6. The method of claim 5, wherein analysing the English textual information using the natural language processing module further comprises:
dividing the English textual information into a set of tokens using a tokenization module;
identifying the part of speech of each token using a part-of-speech tagging module;
identifying the relationships between the tokens using a dependency parsing module;
identifying named entities in the English textual information using a named entity recognition module;
identifying the semantic roles of the tokens using a semantic role labelling module.

7.The method of claim 5, wherein translating the English textual information into other languages using the translation module that use a machine translation module to translate the English textual information into other languages.

8. The method of claim 5, wherein analysing the translated text and the set of features generated by the natural language processing module using the semantic similarity module further configured to generate a semantic vector representation of the English textual information and the translated text using a semantic vector representation module, and determine the semantic similarity between the English textual information and the translated text based on the semantic vector representations using a similarity calculation module.

Documents

Application Documents

# Name Date
1 202311019770-REQUEST FOR EARLY PUBLICATION(FORM-9) [22-03-2023(online)].pdf 2023-03-22
2 202311019770-POWER OF AUTHORITY [22-03-2023(online)].pdf 2023-03-22
3 202311019770-FORM-9 [22-03-2023(online)].pdf 2023-03-22
4 202311019770-FORM FOR SMALL ENTITY(FORM-28) [22-03-2023(online)].pdf 2023-03-22
5 202311019770-FORM 1 [22-03-2023(online)].pdf 2023-03-22
6 202311019770-EVIDENCE FOR REGISTRATION UNDER SSI(FORM-28) [22-03-2023(online)].pdf 2023-03-22
7 202311019770-EVIDENCE FOR REGISTRATION UNDER SSI [22-03-2023(online)].pdf 2023-03-22
8 202311019770-EDUCATIONAL INSTITUTION(S) [22-03-2023(online)].pdf 2023-03-22
9 202311019770-DRAWINGS [22-03-2023(online)].pdf 2023-03-22
10 202311019770-DECLARATION OF INVENTORSHIP (FORM 5) [22-03-2023(online)].pdf 2023-03-22
11 202311019770-COMPLETE SPECIFICATION [22-03-2023(online)].pdf 2023-03-22