Sign In to Follow Application
View All Documents & Correspondence

English Dubbing Material Generation For Multimedia Content Display

Abstract: ENGLISH DUBBING MATERIAL GENERATION FOR MULTIMEDIA CONTENT DISPLAY Abstract A method for creating English dubbing material for multimedia content may be included among the embodiments of the current disclosure. This method may involve the step of receiving a source audio track of the multimedia content in a language other than English. Machine translation may also be used to perform an automated translation of the original audio recording into English as part of certain embodiments. The generation of a synthesized English audio track that matches the lip movements and emotional subtleties of the original audio track may also be considered an embodiment. This generation can be accomplished using voice synthesis technology. In certain embodiments, there is also the possibility of producing the synthesized English audio track so that it may be used as material for dubbing the multimedia content. Fig. 1

Get Free WhatsApp Updates!
Notices, Deadlines & Correspondence

Patent Information

Application #
Filing Date
22 March 2023
Publication Number
19/2023
Publication Type
INA
Invention Field
ELECTRONICS
Status
Email
Parent Application

Applicants

BANASTHALI VIDYAPITH
BANASTHALI VIDYAPITH, P.O. BANASTHALI, BANASTHALI, RAJASTHAN, INDIA, 304022, JAIPUR

Inventors

1. GEETHA YADAV
BANASTHALI VIDYAPITH, P.O. BANASTHALI, BANASTHALI, RAJASTHAN, INDIA, 304022, JAIPUR
2. DR. BHARTI NATHANI
BANASTHALI VIDYAPITH, P.O. BANASTHALI, BANASTHALI, RAJASTHAN, INDIA, 304022, JAIPUR

Claims

1. A method for generating English dubbing material for multimedia content, comprising: receiving a source audio track of the multimedia content in a non-English language; automatically translating the source audio track into English using machine translation; generating a synthesized English audio track that matches the lip movements and emotional nuances of the source audio track using voice synthesis technology; and outputting the synthesized English audio track for use as dubbing material for the multimedia content.

2. The method of claim 1, wherein the voice synthesis technology uses a neural network trained on a large dataset of human speech to generate the synthesized English audio track.

3. The method of claim 1, wherein the lip movements of the source audio track are analysed using computer vision techniques to ensure accurate synchronization with the synthesized English audio track.

4. The method of claim 1, further comprising a step of adjusting the pitch, tone, or speed of the synthesized English audio track to match the characteristics of a specific voice actor.

5. The method of claim 1, wherein the emotional nuances of the source audio track are analysed using natural language processing techniques to ensure accurate expression in the synthesized English audio track.

6. The method of claim 1, further comprising a step of selecting and integrating background music or sound effects into the synthesized English audio track to enhance the overall user experience.

7. The method of claim 1, further comprising a step of generating subtitles or closed captions that correspond to the synthesized English audio track.

8. The method of claim 1, further comprising a step of segmenting the source audio track into individual phrases or sentences for more efficient translation and synthesis.

9. The method of claim 1, further comprising a step of assessing the quality of the synthesized English audio track using a feedback mechanism that takes into account user preferences and feedback. 10 A system for generating English dubbing material for multimedia content, comprising: a processor; a memory storing instructions that, when executed by the processor, cause the system to: receive a source audio track of the multimedia content in a non-English language; automatically translate the source audio track into English using machine translation; generate a synthesized English audio track that matches the lip movements and emotional nuances of the source audio track using voice synthesis technology; and output the synthesized English audio track for use as dubbing material for the multimedia content.   ENGLISH DUBBING MATERIAL GENERATION FOR MULTIMEDIA CONTENT DISPLAY Abstract A method for creating English dubbing material for multimedia content may be included among the embodiments of the current disclosure. This method may involve the step of receiving a source audio track of the multimedia content in a language other than English. Machine translation may also be used to perform an automated translation of the original audio recording into English as part of certain embodiments. The generation of a synthesized English audio track that matches the lip movements and emotional subtleties of the original audio track may also be considered an embodiment. This generation can be accomplished using voice synthesis technology. In certain embodiments, there is also the possibility of producing the synthesized English audio track so that it may be used as material for dubbing the multimedia content. Fig. 1 , Claims:Claims :

1. A method for generating English dubbing material for multimedia content, comprising: receiving a source audio track of the multimedia content in a non-English language; automatically translating the source audio track into English using machine translation; generating a synthesized English audio track that matches the lip movements and emotional nuances of the source audio track using voice synthesis technology; and outputting the synthesized English audio track for use as dubbing material for the multimedia content.

2. The method of claim 1, wherein the voice synthesis technology uses a neural network trained on a large dataset of human speech to generate the synthesized English audio track.

3. The method of claim 1, wherein the lip movements of the source audio track are analysed using computer vision techniques to ensure accurate synchronization with the synthesized English audio track.

4. The method of claim 1, further comprising a step of adjusting the pitch, tone, or speed of the synthesized English audio track to match the characteristics of a specific voice actor.

5. The method of claim 1, wherein the emotional nuances of the source audio track are analysed using natural language processing techniques to ensure accurate expression in the synthesized English audio track.

6. The method of claim 1, further comprising a step of selecting and integrating background music or sound effects into the synthesized English audio track to enhance the overall user experience.

7. The method of claim 1, further comprising a step of generating subtitles or closed captions that correspond to the synthesized English audio track.

8. The method of claim 1, further comprising a step of segmenting the source audio track into individual phrases or sentences for more efficient translation and synthesis.

9. The method of claim 1, further comprising a step of assessing the quality of the synthesized English audio track using a feedback mechanism that takes into account user preferences and feedback. 10 A system for generating English dubbing material for multimedia content, comprising: a processor; a memory storing instructions that, when executed by the processor, cause the system to: receive a source audio track of the multimedia content in a non-English language; automatically translate the source audio track into English using machine translation; generate a synthesized English audio track that matches the lip movements and emotional nuances of the source audio track using voice synthesis technology; and output the synthesized English audio track for use as dubbing material for the multimedia content.

Specification

Description:ENGLISH DUBBING MATERIAL GENERATION FOR MULTIMEDIA CONTENT DISPLAY
Field of the Invention
[0001] The present disclosure generally relates to audio data processing methodology. More specifically, the present invention relates to system and method for dubbing or dubbing of information of audio or video material.
Background
[0002] The background description includes information that may be useful in understanding the present invention. It is not an admission that any of the information provided herein is prior art or relevant to the presently claimed invention, or that any publication specifically or implicitly referenced is prior art.
[0003] Dubbing material generation involves the creation of dubbed versions of multimedia content such as movies, TV shows, and video games. Dubbing is the process of replacing the original audio track of a media with a new audio track in a different language or dialect. Dubbing is commonly used to make media accessible to people who do not speak the original language in which the media was produced. Dubbing material generation is a complex process that involves various stages, including translation, script adaptation, voice acting, and audio post-production. The goal is to produce a high-quality dub that accurately conveys the content and intent of the original media while also taking into account cultural and linguistic differences. With the increasing globalization of media, the demand for high-quality dubbing has grown significantly in recent years. Dubbing is not only a means of making media accessible to people in different regions of the world, but it is also a way to reach a wider audience and increase the commercial success of a production. As a result, there is a growing need for efficient and effective dubbing material generation methods that can produce high-quality dubs in a timely and cost-effective manner. Dubbing is the process of replacing the original dialogue of a multimedia content such as a movie, TV show, or video game with a translation in another language. English is a commonly dubbed language due to its widespread use in the entertainment industry and its status as a global lingua franca. Dubbing material generation for multimedia content display involves creating new translations that match the timing, lip movements, and emotional nuances of the original language while retaining accuracy and cultural context.
[0004] Various automated or semi-automated approaches for language related material generation are disclosed in patent literature. Exemplary documents are disclosed.
[0005] The US20060204212A1 (By: PANASONIC) - An information recording and reproducing apparatus is capable of dubbing a video and audio signal onto a digital recording medium in a short time, using a control signal recorded on an analog recording medium together with the video and audio signal. The information recording and reproducing apparatus includes a control signal detecting part for detecting a control signal recorded on the analog recording medium, a control part, and the like. When a dubbing instruction is input, the control part instructs a recording part to start a dubbing operation of recording the video and audio signal reproduced by a reproducing part onto the digital recording medium, and during the dubbing operation, when a non-recorded region where the control signal is not recorded on the analog recording medium is detected, the control part instructs the recording part to temporarily stop recording onto the digital recording medium.
[0006] The JP2675412B2 (By: SANYO ELECTRIC) - PURPOSE: To stop dubbing in the noisy state by switching high-speed dubbing to set-speed dubbing in the case of occurrence of error during high-speed dubbing and stop dubbing in the case of the occurrence of error even after this switching and dubbing the next music. If the error whose correction processing is impossible occurs during high-speed dubbing, an error flag is counted up by one. It is discriminated whether the error flag is '2' or not. In this case, a tape deck 1 is set to the rewinding state by a system control circuit 2 and the tape is rewound to the dubbing recording start position of the music by an inter-music part detecting circuit 7. Meanwhile, the reproducing position of a disk is returned to the start position of the music by a system control circuit 11 of a CD player 10. Thereafter, the tape deck 1 and the CD player 10 are switched to the set-speed dubbing state by circuits 2 and 11. If the error whose correction processing is impossible occurs in the set-speed dubbing state, the deck 1 and the player 10 are set to stand-by states for high-speed recording and high-speed reproducing respectively, and they are switched to the high-speed dubbing state of the next music.
[0007] The KR10-0223179B1 (By: SAMSUNG ELECTRONICS) - PURPOSE: An apparatus for audio dubbing of an 8mm video cassette recorder is provided to conveniently record an explanation with clear quality of sound in the case that the explanation is inserted into a tape by dubbing an audio signal in a PCM block which isn't used in the 8mm video tape. A key input unit(21) receives a determined function key from a user. A microprocessor (22) responds to a signal inputted from the key input unit(21), and controls components in order to perform relevant function. An operation unit(23) for operating a deck unit(24) under the control of the microprocessor(22) is connected to the microprocessor(22). A general video signal processing unit(30), an audio signal processing unit(31) and an A/V amplifying unit(29) are equipped. A microphone(26) receives an audio signal for dubbing. An audio signal for dubbing processing unit(27) processes the signal inputted from the microphone(26). In addition, an audio amplifier for dubbing(28) is equipped. A switching unit(25) switches a connection condition of switches(SW1,SW2) under the control of the microprocessor(22). A switch(SW3) connects one output terminal among output terminals of audio signal for dubbing processing unit(27) and the audio signal processing unit(31) to an audio output terminal under the control of the microprocessor(22).
[0008] FR2768878B1 (By: SONY PICTURES ENTERTAINMENT ) - The present system edits digitized audio/video data. The system comprises a plurality of stations containing the digitized audio/video data and an editing station located remotely from the plurality of stations and coupled in communication with the plurality of stations via a network. The editing station may also include a graphical interface and accept input via the graphical interface wherein the input specifies edit functions to be performed on the digitized audio/video data residing on the plurality of audio stations.
[0009] However, this process can be time-consuming and expensive, requiring skilled translators, voice actors, and audio engineers. Therefore, there is a need for efficient and effective methods to generate high-quality dubbing material for English language content.

Summary
[00010] The following presents a simplified summary of various aspects of this disclosure in order to provide a basic understanding of such aspects. This summary is not an extensive overview of all contemplated aspects, and is intended to neither identify key or critical elements nor delineate the scope of such aspects. Its purpose is to present some concepts of this disclosure in a simplified form as a prelude to the more detailed description that is presented later.
[00011] The following paragraphs provide additional support for the claims of the subject application.
[00012] The present disclosure generally relates to audio data processing methodology. More specifically, the present invention relates to system and method for dubbing or dubbing of information of audio or video material.
[00013] Embodiments of the present disclosure may include a method for generating English dubbing material for multimedia content, including receiving a source audio track of the multimedia content in a non-English language. Embodiments may also include automatically translating the source audio track into English using machine translation. Embodiments may also include generating a synthesized English audio track that matches the lip movements and emotional nuances of the source audio track using voice synthesis technology. Embodiments may also include outputting the synthesized English audio track for use as dubbing material for the multimedia content.
[00014] In some embodiments, the voice synthesis technology uses a neural network trained on a large dataset of human speech to generate the synthesized English audio track. In some embodiments, the lip movements of the source audio track may be analysed using computer vision techniques to ensure accurate synchronization with the synthesized English audio track.
[00015] In some embodiments, the method may include a step of adjusting the pitch, tone, or speed of the synthesized English audio track to match the characteristics of a specific voice actor. In some embodiments, the emotional nuances of the source audio track may be analysed using natural language processing techniques to ensure accurate expression in the synthesized English audio track.
[00016] In some embodiments, the method may include a step of selecting and integrating background music or sound effects into the synthesized English audio track to enhance the overall user experience. In some embodiments, the method may include a step of generating subtitles or closed captions that correspond to the synthesized English audio track.
[00017] In some embodiments, the method may include a step of segmenting the source audio track into individual phrases or sentences for more efficient translation and synthesis. In some embodiments, the method may include a step of assessing the quality of the synthesized English audio track using a feedback mechanism that takes into account user preferences and feedback. Embodiments may also include a system for generating English dubbing material for multimedia content, including a processor.
[00018] Embodiments may also include a memory storing instructions, which when executed by the processor, cause the system to receive a source audio track of the multimedia content in a non-English language. Embodiments may also include automatically translate the source audio track into English using machine translation. Embodiments may also include generatation of a synthesized English audio track that matches the lip movements and emotional nuances of the source audio track using voice synthesis technology. Embodiments may also include output the synthesized English audio track for use as dubbing material for the multimedia content.
Brief Description of the Drawings
[00019] The features and advantages of the present disclosure would be more clearly understood from the following description taken in conjunction with the accompanying drawings in which:
[00020] FIG. 1 is a flowchart illustrating a method for generating English dubbing material for multimedia content, according to some embodiments of the present disclosure.
[00021] FIG. 2 is a flowchart further illustrating the method for generating English dubbing material for multimedia content from FIG. 1, according to some embodiments of the present disclosure.
Detailed Description
[00022] In the following detailed description of the invention, reference is made to the accompanying drawings that form a part hereof, and in which is shown, by way of illustration, specific embodiments in which the invention may be practiced. In the drawings, like numerals describe substantially similar components throughout the several views. These embodiments are described in sufficient detail to claim those skilled in the art to practice the invention. Other embodiments may be utilized and structural, logical, and electrical changes may be made without departing from the scope of the present invention. The following detailed description is, therefore, not to be taken in a limiting sense, and the scope of the present invention is defined only by the appended claims and equivalents thereof.
[00023] The use of the terms “a” and “an” and “the” and “at least one” and similar referents in the context of describing the invention (especially in the context of the following claims) are to be construed to cover both the singular and the plural, unless otherwise indicated herein or clearly contradicted by context. The use of the term “at least one” followed by a list of one or more items (for example, “at least one of A and B”) is to be construed to mean one item selected from the listed items (A or B) or any combination of two or more of the listed items (A and B), unless otherwise indicated herein or clearly contradicted by context. The terms “comprising,” “having,” “including,” and “containing” are to be construed as open-ended terms (i.e., meaning “including, but not limited to,”) unless otherwise noted. Recitation of ranges of values herein are merely intended to serve as a shorthand method of referring individually to each separate value falling within the range, unless otherwise indicated herein, and each separate value is incorporated into the specification as if it were individually recited herein. All methods described herein can be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by context. The use of any and all examples, or exemplary language (e.g., “such as”) provided herein, is intended merely to better illuminate the invention and does not pose a limitation on the scope of the invention unless otherwise claimed. No language in the specification should be construed as indicating any non-claimed element as essential to the practice of the invention.
[00024] The present disclosure generally relates to audio data processing methodology. More specifically, the present invention relates to system and method for dubbing or dubbing of information of audio or video material.
[00025] The technique for producing English dubbing material for multimedia content is shown in flowchart form in FIG. 1, which also provides a description of the process in accordance with various implementations of the present disclosure. Receiving a source audio track of the multimedia material in a language other than English is an optional step that may be included in some implementations of the technique at step 110. In step 120, the technique could incorporate the use of machine translation to carry out an automated translation of the source audio track into English. Using voice synthesis technology, the technique may comprise, at step 130, the generation of a synthesized English audio track that matches the lip motions and emotional subtleties of the original audio recording. At step 140, the technique may comprise the step of producing the synthesized English audio track in order to utilise it as material for dubbing the multimedia content.
[00026] While generating the synthesized English audio track, the voice synthesis technology may, in certain implementations, make use of a neural network that has been trained on a huge dataset consisting of human speech. It is possible, in certain implementations, to analyse the motions of the speaker's lips in the source audio track using computer vision methods in order to guarantee exact synchronisation with the English audio track that was synthesised. A phase that involves modifying the pitch, tone, or speed of the synthesised English audio track in order to match the characteristics of a particular voice actor may be included in some implementations of the approach.
[00027] To guarantee proper expression in the synthesised English audio track, it is possible, according to certain implementations, to conduct an analysis of the emotional subtleties included in the original audio track using natural language processing methods. In some implementations of the approach, one of the steps may include picking background music or sound effects and incorporating them into the synthesized English audio track. This is done with the intention of improving the user experience as a whole. The technique may, in some implementations, include a phase for the generation of subtitles or closed captions that correspond to the synthesized English audio track that is being played. In some implementations of the approach, one of the steps may include segmenting the raw audio recording into individual phrases or sentences so that translation and synthesis may take place with more efficiency.
[00028] The technique for creating English dubbing material for multimedia content shown in FIG. 1 is shown in more detail in the flowchart that may be found in FIG. 2. This approach is consistent with certain implementations of the current disclosure. The method may, in some implementations, comprise, at step 250, an assessment of the quality of the synthesised English audio track using a feedback system that takes into consideration the preferences of the user and feedback received from the user. Receiving a source audio track of the multimedia material in a language other than English may be an option during the assessment phase located at number 210. The process of evaluating may involve, at position 220, the act of automatically translating the original audio recording into English using a machine translation programme. The phase of evaluating may include, at 230, the generation of a synthesized English audio track by employing voice synthesis technology. This English audio track should match the lip movements and emotional subtleties of the original audio recording. The process of evaluating may comprise outputting the synthesized English audio track at position 240 so that it may be used as material for dubbing the multimedia content.The technique for creating English dubbing material for multimedia content may comprise the performance of one or more additional stages. The system is designed to generate English dubbing material for multimedia content. A memory capable of storing instructions, which when carried out by the processor, cause the system to, complete the operation.
[00029] A method for creating English dubbing material for multimedia content may be included among the embodiments of the current disclosure. This method may involve the step of receiving a source audio track of the multimedia content in a language other than English. Machine translation may also be used to perform an automated translation of the original audio recording into English as part of certain embodiments. The generation of a synthesized English audio track that matches the lip movements and emotional subtleties of the original audio track may also be considered an embodiment. This generation can be accomplished using voice synthesis technology. In certain embodiments, there is also the possibility of producing the synthesized English audio track so that it may be used as material for dubbing the multimedia content.
[00030] While generating the synthesized English audio track, the voice synthesis technology may, in certain implementations, make use of a neural network that has been trained on a significant amount of human speech data. It is possible, in certain implementations, to analyse the motions of the speaker's lips in the source audio track using computer vision methods in order to guarantee exact synchronisation with the English audio track that was synthesised.
[00031] A phase that involves modifying the pitch, tone, or speed of the synthesised English audio track in order to match the characteristics of a particular voice actor may be included in some implementations of the approach. To guarantee proper expression in the synthesised English audio track, it is possible, according to certain implementations, to conduct an analysis of the emotional subtleties included in the original audio track using natural language processing methods.
[00032] In some implementations of the approach, one of the steps may include picking background music or sound effects and incorporating them into the synthesized English audio track. This is done with the intention of improving the user experience as a whole. The technique may, in some implementations, include a phase for the generation of subtitles or closed captions that correspond to the synthesized English audio track that is being played.
[00033] In some implementations of the approach, one of the steps may include segmenting the raw audio recording into individual phrases or sentences so that translation and synthesis may take place with more efficiency. A phase that assesses the quality of the synthesised English audio track using a feedback system that takes into account user preferences and feedback may be included in certain implementations of the approach. Moreover, embodiments could comprise a system for creating English dubbing material for multimedia content, wherein the system might contain a processor.
[00034] In some implementations, there is also a memory that stores instructions, which when executed by the processor, cause the system to receive a source audio track of the multimedia content that is spoken in a language other than English. Moreover, embodiments may involve the use of machine translation to automatically convert the original audio recording into English. Voice synthesis technology may also be used in certain embodiments in order to provide a synthesized English audio track that, when compared to the original audio track, more closely replicates the lip movements and emotional subtleties. In certain embodiments, there is also the possibility of outputting the synthesized English audio track in order to utilise it as material for dubbing the multimedia content.
[00035] Example embodiments herein have been described above with reference to block diagrams and flowchart illustrations of methods and apparatuses. It will be understood that each block of the block diagrams and flowchart illustrations, and combinations of blocks in the block diagrams and flowchart illustrations, respectively, can be implemented by various means including hardware, software, firmware, and a combination thereof. For example, in one embodiment, each block of the block diagrams and flowchart illustrations, and combinations of blocks in the block diagrams and flowchart illustrations can be implemented by computer program instructions. These computer program instructions may be loaded onto a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions which execute on the computer or other programmable data processing apparatus create means for implementing the functions specified in the flowchart block or blocks.
[00036] Throughout the present disclosure, the term ‘processing means’ or ‘microprocessor’ or ‘processor’ or ‘processors’ includes, but is not limited to, a general purpose processor (such as, for example, a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a microprocessor implementing other types of instruction sets, or a microprocessor implementing a combination of types of instruction sets) or a specialized processor (such as, for example, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), or a network processor).
[00037] The term “non-transitory storage device” or “storage” or “memory,” as used herein relates to a random access memory, read only memory and variants thereof, in which a computer can store data or software for any duration.
[00038] Operations in accordance with a variety of aspects of the disclosure is described above would not have to be performed in the precise order described. Rather, various steps can be handled in reverse order or simultaneously or not at all.
[00039] While several implementations have been described and illustrated herein, a variety of other means and/or structures for performing the function and/or obtaining the results and/or one or more of the advantages described herein may be utilized, and each of such variations and/or modifications is deemed to be within the scope of the implementations described herein. More generally, all parameters, dimensions, materials, and configurations described herein are meant to be exemplary and that the actual parameters, dimensions, materials, and/or configurations will depend upon the specific application or applications for which the teachings is/are used. Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific implementations described herein. It is, therefore, to be understood that the foregoing implementations are presented by way of example only and that, within the scope of the appended claims and equivalents thereto, implementations may be practiced otherwise than as specifically described and claimed. Implementations of the present disclosure are directed to each individual feature, system, article, material, kit, and/or method described herein. In addition, any combination of two or more such features, systems, articles, materials, kits, and/or methods, if such features, systems, articles, materials, kits, and/or methods are not mutually inconsistent, is included within the scope of the present disclosure.

Claims
I/We Claim:
1. A method for generating English dubbing material for multimedia content, comprising:
receiving a source audio track of the multimedia content in a non-English language;
automatically translating the source audio track into English using machine translation;
generating a synthesized English audio track that matches the lip movements and emotional nuances of the source audio track using voice synthesis technology; and
outputting the synthesized English audio track for use as dubbing material for the multimedia content.

2. The method of claim 1, wherein the voice synthesis technology uses a neural network trained on a large dataset of human speech to generate the synthesized English audio track.

3. The method of claim 1, wherein the lip movements of the source audio track are analysed using computer vision techniques to ensure accurate synchronization with the synthesized English audio track.

4. The method of claim 1, further comprising a step of adjusting the pitch, tone, or speed of the synthesized English audio track to match the characteristics of a specific voice actor.

5. The method of claim 1, wherein the emotional nuances of the source audio track are analysed using natural language processing techniques to ensure accurate expression in the synthesized English audio track.

6. The method of claim 1, further comprising a step of selecting and integrating background music or sound effects into the synthesized English audio track to enhance the overall user experience.

7. The method of claim 1, further comprising a step of generating subtitles or closed captions that correspond to the synthesized English audio track.

8. The method of claim 1, further comprising a step of segmenting the source audio track into individual phrases or sentences for more efficient translation and synthesis.

9. The method of claim 1, further comprising a step of assessing the quality of the synthesized English audio track using a feedback mechanism that takes into account user preferences and feedback.

10 A system for generating English dubbing material for multimedia content, comprising:
a processor;
a memory storing instructions that, when executed by the processor, cause the system to:
receive a source audio track of the multimedia content in a non-English language;
automatically translate the source audio track into English using machine translation;
generate a synthesized English audio track that matches the lip movements and emotional nuances of the source audio track using voice synthesis technology; and
output the synthesized English audio track for use as dubbing material for the multimedia content.

ENGLISH DUBBING MATERIAL GENERATION FOR MULTIMEDIA CONTENT DISPLAY
Abstract
A method for creating English dubbing material for multimedia content may be included among the embodiments of the current disclosure. This method may involve the step of receiving a source audio track of the multimedia content in a language other than English. Machine translation may also be used to perform an automated translation of the original audio recording into English as part of certain embodiments. The generation of a synthesized English audio track that matches the lip movements and emotional subtleties of the original audio track may also be considered an embodiment. This generation can be accomplished using voice synthesis technology. In certain embodiments, there is also the possibility of producing the synthesized English audio track so that it may be used as material for dubbing the multimedia content.

Fig. 1 , Claims:Claims
I/We Claim:
1. A method for generating English dubbing material for multimedia content, comprising:
receiving a source audio track of the multimedia content in a non-English language;
automatically translating the source audio track into English using machine translation;
generating a synthesized English audio track that matches the lip movements and emotional nuances of the source audio track using voice synthesis technology; and
outputting the synthesized English audio track for use as dubbing material for the multimedia content.

2. The method of claim 1, wherein the voice synthesis technology uses a neural network trained on a large dataset of human speech to generate the synthesized English audio track.

3. The method of claim 1, wherein the lip movements of the source audio track are analysed using computer vision techniques to ensure accurate synchronization with the synthesized English audio track.

4. The method of claim 1, further comprising a step of adjusting the pitch, tone, or speed of the synthesized English audio track to match the characteristics of a specific voice actor.

5. The method of claim 1, wherein the emotional nuances of the source audio track are analysed using natural language processing techniques to ensure accurate expression in the synthesized English audio track.

6. The method of claim 1, further comprising a step of selecting and integrating background music or sound effects into the synthesized English audio track to enhance the overall user experience.

7. The method of claim 1, further comprising a step of generating subtitles or closed captions that correspond to the synthesized English audio track.

8. The method of claim 1, further comprising a step of segmenting the source audio track into individual phrases or sentences for more efficient translation and synthesis.

9. The method of claim 1, further comprising a step of assessing the quality of the synthesized English audio track using a feedback mechanism that takes into account user preferences and feedback.

10 A system for generating English dubbing material for multimedia content, comprising:
a processor;
a memory storing instructions that, when executed by the processor, cause the system to:
receive a source audio track of the multimedia content in a non-English language;
automatically translate the source audio track into English using machine translation;
generate a synthesized English audio track that matches the lip movements and emotional nuances of the source audio track using voice synthesis technology; and
output the synthesized English audio track for use as dubbing material for the multimedia content.

Documents

Application Documents

# Name Date
1 202311019766-REQUEST FOR EARLY PUBLICATION(FORM-9) [22-03-2023(online)].pdf 2023-03-22
2 202311019766-POWER OF AUTHORITY [22-03-2023(online)].pdf 2023-03-22
3 202311019766-OTHERS [22-03-2023(online)].pdf 2023-03-22
4 202311019766-FORM-9 [22-03-2023(online)].pdf 2023-03-22
5 202311019766-FORM FOR SMALL ENTITY(FORM-28) [22-03-2023(online)].pdf 2023-03-22
6 202311019766-FORM 1 [22-03-2023(online)].pdf 2023-03-22
7 202311019766-EVIDENCE FOR REGISTRATION UNDER SSI(FORM-28) [22-03-2023(online)].pdf 2023-03-22
8 202311019766-EDUCATIONAL INSTITUTION(S) [22-03-2023(online)].pdf 2023-03-22
9 202311019766-DRAWINGS [22-03-2023(online)].pdf 2023-03-22
10 202311019766-DECLARATION OF INVENTORSHIP (FORM 5) [22-03-2023(online)].pdf 2023-03-22
11 202311019766-COMPLETE SPECIFICATION [22-03-2023(online)].pdf 2023-03-22