Abstract: ABSTRACT METHOD AND SYSTEM FOR IMPROVING ONLINE INTERACTION A method and corresponding system are described for improving the immersive mature of online interaction. The interaction may be between plurality of user or a user and AI. <>
1. A method of improving online interaction, comprising: providing, by a computing entity 300, a plurality of users with an opportunity to select one of a plurality of user representation; receiving, by the computing entity 300, an input from at least a first subset of 5 the plurality of users; transforming, by the computing entity 300, the input according to the selected user representations; and transmitting, by the computing entity 300, the transformed input to at least a second subset of the plurality of the users, wherein, transforming includes generating 10 an audio clip.
2. The method as claimed in claim 1, wherein the computing entity 300 has access to a memory module 320 containing a plurality of voice prints associated with at least some of user representation.
3. The method as claimed in claim 1, wherein the computing entity 300 includes a 15 neural network trained using a database of a plurality of voice prints associated with at least some of user representation.
4. The method as claimed in claim 3, wherein the computing entity 300 uses the artificial intelligence to generate the audio clip.
5. A system 100 enabling online interaction, comprising: 20 a first set of client devices, wherein each of first set of client devices 200 include, an input module 210 for capturing a user’s a plurality of inputs at the client device 200; 14 a client transceiver module 240 for sending the captured plurality of inputs to the server 300 and receiving a bitstream from the server 300; and an output module 220 for presenting the received bitstream to the user; and a server 300, wherein the server 300 includes, 5 a server transceiver module 310 for receiving the captured plurality of inputs and sending the bitstream; a processing module 330; a memory 320 storing a plurality of voice prints, a plurality of user selections and associations between user representations and voice prints; 10 wherein at least one of the captured plurality of inputs indicate a selected user representation from among a plurality of user presentations, wherein at least one of the captured plurality of inputs represents an information to be shared with at least a subset of the first set of client devices, wherein the processing module 330 is capable of selecting a voice print 15 from a plurality of voice prints based on the selected user representation, wherein the processing module 330 is capable of transforming the information to be shared into an audio clip, and wherein the processing module 330 is capable of converting the audio file into the bitstream. 20
6. The system as claimed in claim 5, wherein the input module 210 includes at least two of a pointing device, a microphone, and a keyboard.
7. The system as claimed in claim 5, wherein the output module 220 includes at least a display device and a speaker. 15
8. The system as claimed in claim 5, wherein the processing module 330 uses a neural network to generate the audio clip. Dated this 17th of May, 2021 5 Applicant’s Patent Agent of Effectual Legal Services Pvt. Ltd. , Description:1 FORM – 2 THE PATENTS ACT, 1970 (39 of 1970) & THE PATENTS RULES, 2003 COMPLETE SPECIFICATION (See section 10 and rule 13) METHOD AND SYSTEM FOR IMPROVING ONLINE INTERACTION Applicant: 1. Rajat House no 485, Ward No. 5, Mehrauli, New Delhi 110030, India Inventor: 1) Rajat The following specification particularly describes the invention and the manner in which it is to be performed 2 TECHNICAL FIELD [0001] The present invention is directed towards improving the user experience when interacting online with other user(s) and/or bot(s) online. DESCRIPTION OF THE RELATED ART 5 [0002] The development of internet has improved human communication immensely. Anyone using internet is now capable of interacting with anyone else using internet in near real-time, wherein the time difference between sender transmitting the interaction and receiver experiencing the interaction is imperceivable to humans. The interactions can take one of many forms, e.g. text, voice, video. 10 [0003] These exchanges of interactions are enabled by servers, provided by organizations. A user wishing to interact with other users will connect their client device to the server via internet connections. Upon successfully connecting with server, the user can interact with any other user, or even bot, which are software capable of interacting with humans. 15 [0004] The organizations provide these servers either as standalone service or as a part of bigger service. One example of the being part of bigger service is communication capability between users experiencing a virtual and/or augmented reality or between a user experiencing a virtual and/or augmented reality and a bot which is part of said virtual and/or augmented reality. 20 [0005] Another example is providing a bot as a part of website, hosted on server to answer any queries that a user visiting the website may have. [0006] It is customary for a user entering the virtual and/or augmented reality to select a character to represent the user in the said reality. Generally, there are few phrases pre-recorded in a voice actor’s voice, each corresponding to one character. 25 3 [0007] Generally, these virtual and/or augmented realities also allow a user to send and/or talk to other users who are part of same virtual and/or augmented reality. However, listening to a character speak in a voice actor’s voice some time and in user’s voice rest of the time effects the user experience. Further, some of these realities allow users to interact using text-based chat. However, reading text during performing the 5 tasks in virtual and/or augmented reality makes reduces efficiency of user. [0008] In view of the foregoing, a method and corresponding system is disclosed to address the abovementioned issues in the art. SUMMARY 10 [0009] In view of the foregoing, an embodiment herein provides a method for receiving a character selection from a user, receiving a voice clip and/or text message from the user, converting the voice clip and/or text message into a voice clip in the voice of voice actor corresponding to the selected character. The voice clip in the voice of the voice actor is sent to one or more of users. The set of users receiving the voice 15 clip may or may not include the sender of the voice clip and/or text message. [0010] In some embodiments, conversion takes place using an Artificial Intelligence, sometimes abbreviated as AI, wherein the AI is trained to generated audio clips in the voice actors corresponding to selected characters. [0011] In some embodiments, the method includes a text-to-speech generator. 20 [0012] In some embodiments, the method includes a speech-to-text generator. BRIEF DESCRIPTION OF THE DRAWINGS [0013] The embodiments herein will be better understood from the following detailed description with reference to the drawings, in which: 25 4 [0014] FIG. 1 illustrates a first exemplary environment wherein the invention may be implemented. [0015] FIG. 2 illustrates an exploded view of a client device 200 according to some embodiments herein. [0016] FIG. 3 illustrates an exploded view of a server 300 according to some 5 embodiments herein. [0017] FIG. 4 illustrates a method employed at the server 300 according to some embodiments herein. 10 15 20 5 DETAILED DESCRIPTION OF PREFERRED EMBODIMENTS [0018] The embodiments herein and the various features and advantageous details thereof are explained more fully with reference to the non-limiting embodiments that are illustrated in the accompanying drawings and detailed in the following description. Descriptions of well-known components and processing techniques are omitted so as 5 not to unnecessarily obscure the embodiments herein. The examples used herein are intended merely to facilitate an understanding of ways in which the embodiments herein may be practiced and to further enable those of skill in the art to practice the embodiments herein. Accordingly, the examples should not be construed as limiting the scope of the embodiments herein. 10 [0019] As mentioned, there remains a need for a system and a method of improving the online interaction experiences for users of web services provided by various organizations. Referring now to the drawings, and more particularly to FIGS. 1 through 4, where similar reference characters denote corresponding features consistently throughout the figures, there are shown preferred embodiments. 15 [0020] A "server" refers to a piece of software and/or hardware that allows an organization to provide services to users of the internet. A server can be a combination of multiple machines also. The defining characteristic of the server is that it receives data from a user’s client device and is capable of processing said received data. Further, the server may also be capable of transmitting data to one or more of client devices 20 connected to it. Terms “server” and “computing entity” are used interchangeably in this application. [0021] A "client device" refers to a piece of software and/or hardware that allows a user to receive services made available online by one or more organization. The defining characteristic of the client device is that it sends and/or receives data to/from 25 at least one server and is capable of outputting said data to user. Client device can also 6 be a “thin client”. Terms “client” and “user” are used interchangeably in this application. [0022] A "voice print" refers to measurable qualities of a voice that uniquely identifies a speaker. [0023] A "AI" refers to any software capable of self-learning and decision making. For 5 the purpose of this invention, term “AI” includes all of machine learning software as well as neural network based software, [0024] FIG. 1 illustrates a first exemplary environment diagram for a system and method to improve a user’s experience of interacting with other users and/or bots according to an embodiment herein. The organization offering services to users of internet will upload the services to a server 10 and inform users on how to connect to the host server. The organization may inform users on how to connect to the users explicitly, e.g., by providing a Uniform Resource Locator (URL), or implicitly, e.g., providing the user with a software program wherein the program contains URL and connects to server without informing the user of the URL. 15 [0025] A user of client device 200, desiring to use the services of an organization, may connect to the server hosting the services in many ways. In one embodiment, the user starts an internet browser program installed on the client device and input the URL of the server on which the services are hosted. In another embodiment user starts a program previously installed on the system. The program’s code includes the 20 instructions on how to connect to the server on which the services are hosted. [0026] The client device 200 establishes a link to the server using an internet connection. The link may be direct, e.g., using a VPN tunnel or when user is aware of public IP address and/or port of the server or it maybe indirect wherein the connection passes through a plurality of networking devices, e.g., Domain Name Servers, routers, 25 switches, gateways etc. 7 [0027] FIG. 2 illustrates an exploded view of the client devices 200 for using services hosted on a server across the internet according to an exemplary embodiment of the invention. The client device 200 includes an input module 210, an output module 220, a processing module 230 and a transceiver module 240. Please note that the client device 200 may have even more modules for various other functions. However, to 5 avoid unnecessary details, we will not be explaining those. For example, the client device may also include memory module, not shown in the FIG. 2, to store one or more of softwares, firmwares, operating systems, and data. Further, in one embodiment, one or more of the modules may be combined in one module. [0028] The input module 210 consists of all the submodules that allow a user to input 10 information to the client device. According to certain embodiments, the input module 210 consists of a speech input capture submodule 211, a text input capture submodule 212 and a selection input capture submodule 213. A non-limiting example of the speech input capture submodule 211 is a microphone. The microphone receives a speech spoken by users, generates an analogue signal representing spoken speech and finally, 15 converts the analogue signal into a digital representation. This digital representation is further processed. A non-limiting example of text input capture submodule 212 is a keyboard. The keyboard contains keys that a user presses to input texts, commands, and special characters etc. The Keyboard recognizes what key was pressed and generates a corresponding digital representation. This digital representation is further 20 processed. A non-limiting example of selection input capture submodule 213 is a mouse. The mouse allows a user to control selection and input selection. The mouse generates a digital representation of selection and forwards it for the processing. [0029] The output module 220 consists of all the submodules that allow a user to receive information from the client device 200. According to certain embodiments, the 25 output module 220 consists of a visual output submodule 221 and an audio output submodule 222. A non-limiting example of the visual output submodule 221 is a display. The display receives digital information from the processing module 230 and 8 converts the digital information using a digital-to-analog converter. This analog information is used to drive circuits to present information for the user to see. A non-limiting example of audio output submodule is a speaker. The speaker receives digital information from processing module and converts it into a sound wave, which is outputted to the user’s ears. 5 [0030] The processing module 230 consists of all the submodules that allow a client device to process information. According to certain embodiments, the processing module 230 consists of a central processing submodule 231 and, optionally, a graphics processing submodule 232. The central processing submodule 231 receives digital information from all the modules, processes received information and sends the result 10 of processing to appropriate module. A non-limiting example of the central processing submodule 231 is a CPU. In some embodiments, the central processing submodule 231 is accompanied by a graphics processing submodule 232. A non-limiting example of a graphics processing submodule is a GPU. The graphics processing submodule 232 is a hardware and accompanying firmware that are specifically built for processing visual 15 information. [0031] The transceiver module 240 consists of all the submodules that allow a client device to establish connections with other client devices and/or servers. According to certain embodiments, the transceiver module 240 consists of a receiving submodule 241 and a transmitting submodule 242. In certain embodiments logical receiving 20 submodule 241 and transmitting submodule 242 are implemented in same physical circuits. Few non-limiting examples of the transceiver module 240 are Wi-Fi modules, Ethernet Cards, Cellular Antennas etc. [0032] FIG. 3 illustrates an exploded view of the server 300 used for hosting services according to an exemplary embodiment of the invention. Server includes a transceiver 25 module 310, a processing module 330 and a memory module 320. Please note that the server 300 may have even more modules for various other functions. However, to avoid 9 unnecessary details, we will not be explaining those. Further, in one embodiment, one or more of the modules may be combined in one module. [0033] The transceiver module 310 consists of all the submodules that allow a server 300 to establish connections with other client devices and/or servers. According to certain embodiments, the transceiver module 310 consists of a receiving submodule 5 311 and a transmitting submodule 312. In certain embodiments logical receiving submodule 311 and transmitting submodule 312 are implemented in same physical circuits. Few non-limiting examples of The transceiver module 310 are Wi-Fi modules, Ethernet Cards, Cellular Antennas etc. [0034] The memory module 320 consists of all the submodules that allow the server 10 to store one or more of softwares, firmwares, operating systems, and data. According to certain embodiments, the transceiver module 310 consists of a voice print database submodule 321, a Character-to-Voice Print mapping submodule 322 and an instruction submodule 323. In certain embodiments one or more of the voice print database submodule submodule 321, the Character-to-Voice Print mapping submodule 322 and 15 the instruction submodule 323 may be implemented in same physical circuits. [0035] The voice print database submodule 321 stores voiceprints of all voice actors in whose voice an audio output might be generated. The Character-to-Voice Print mapping submodule 322 stores an association between a selectable character presented to the user and voice print of an actor chosen as the voice of the character. The 20 instruction submodule 323 stores all the instructions that may be executed by processing module 330. [0036] The processing module 330 consists of all the submodules that allow a server to receive data and/or instructions, process the data and/or instructions and output the result of processing. According to certain embodiments, the processing module 330 25 consists of an AI submodule 331, a Text to speech submodule 332 and a Speech Synthesis submodule 333. In certain embodiments one or more of the AI submodule 10 331, the text to speech submodule 332 and the speech synthesis submodule 333 may be implemented in same physical circuits. In certain embodiments the processing module 330 further includes a speech recognition submodule, not shown in the figure. [0037] Fig. 4 illustrates a method of invention, implemented at the server according to one of the embodiments. The method starts after an internet connection has been 5 established between a client device 200 and the server 300. The client device 200 outputs to user of client device 200 a plurality of characters. The user selects one among the presented plurality of characters, to represent the user in the online environment. The selected character may be considered as user representation. As non-limiting example, the client device 200 displays plurality of user avatars to user. As another 10 non-limiting example, client device 200 displays multiple topics and asks user to select the topic on which the user requires help and/or information, wherein each topic has a pre-associated character. User may or may not be aware of the association between the topics and characters. At step 410, the server 300 receives character selection made by user over the established connection. At step 420, the server 300 creates and stores an 15 association between user and character. The server 300 may store this association between user and character either temporarily or permanently. Sometime later, the user input some information at the client devices. As non-limiting example, the user speaks into the microphone of the client device 200 or types a text using keyboard. At step 430, the server 300 receives digital representation of the input from client device 200. 20 At the step 440, the server 300 extracts input from the received digital representation. As non-limiting example, the server 300 uses a speech recognition submodule to recognize what the user spoke or recognizes the text potion from keyboard input. The server 300 also extracts emotions/tones from the input. As a non-limiting example, the server 300 recognizes the user spoke the input in agitated voice. As another non-25 limiting example, the server 300 recognizes that the text was typed in a frustrated manned. At the step 450, the server 300 decides a response. As a non-limiting example, if the user asked a question appropriate, the server 300 finds the answer to the question, 11 either in memory of the server 300 or over the internet. if the server 300 unable to find answer, the server 300 may inform user of the same. As another non-limiting example, If the user provide an instruction, the server 300 may acknowledge the request or if the server is incapable of completing the task, the server 300 may convey the same to the user. As another non-limiting example, the server 300 may filter out some information. 5 The set of client devices may or may not include the client device 200 which submitted the information. At step 460, the server 300 uses an AI-powered speech synthesis system to generate an audio clip in the voice of the voice actor associated with the character selected by the user. The AI-powered speech synthesis is trained using the voice prints of all the voice actors and can generate audio clips in voice of voice actors 10 in a range of emotions. The emotion in which the audio clip is generated is determined based on the emotion extracted in step 440. As a non-limiting example, in an environment where the user has been interacting with the server 300 to seek information and it is sensed that the user has become frustrated, the audio clip may be generated in a calmer or soothing tone. As another non-limiting example, in an 15 environment where user has been interacting with server to communicate with other members of his gaming party, the audio clip may be generated in tone which was extracted at step 440. At step 470, the server converts the audio clip in a format fit for transmission over the internet and transmits the converted audio clip to intended recipients. The intended recipients may simply be the user, or the server 300 may 20 determine based on some rules, e.g., all users in the user’s party, all users who the user specified as recipients, all other users whose characters are within a defined proximity of user’s character in a virtual world, etc. [0038] The foregoing description of the specific embodiments has fully revealed the general nature of the embodiments herein that others can, by applying current 25 knowledge, readily modify and/or adapt for various applications such specific embodiments without departing from the generic concept, and, therefore, such adaptations and modifications should and are intended to be comprehended within the 12 meaning and range of equivalents of the disclosed embodiments. It is to be understood that the phraseology or terminology employed herein is for the purpose of description and not of limitation. Therefore, while the embodiments herein have been described in terms of preferred embodiments, those skilled in the art will recognize that the embodiments herein can be practiced with modification within the spirit and scope. 5 Dated this 17th day of May, 2021 Applicant’s Patent Agent of Effectual Legal Services Pvt. Ltd. 10
Claims:13
I/We claim:
1. A method of improving online interaction, comprising:
providing, by a computing entity 300, a plurality of users with an opportunity to select one of a plurality of user representation;
receiving, by the computing entity 300, an input from at least a first subset of 5 the plurality of users;
transforming, by the computing entity 300, the input according to the selected user representations; and
transmitting, by the computing entity 300, the transformed input to at least a second subset of the plurality of the users, wherein, transforming includes generating 10 an audio clip.
2. The method as claimed in claim 1, wherein the computing entity 300 has access to a memory module 320 containing a plurality of voice prints associated with at least some of user representation.
3. The method as claimed in claim 1, wherein the computing entity 300 includes a 15 neural network trained using a database of a plurality of voice prints associated with at least some of user representation.
4. The method as claimed in claim 3, wherein the computing entity 300 uses the artificial intelligence to generate the audio clip.
5. A system 100 enabling online interaction, comprising: 20
a first set of client devices, wherein each of first set of client devices 200 include,
an input module 210 for capturing a user’s a plurality of inputs at the client device 200;
14
a client transceiver module 240 for sending the captured plurality of inputs to the server 300 and receiving a bitstream from the server 300; and
an output module 220 for presenting the received bitstream to the user; and
a server 300, wherein the server 300 includes, 5
a server transceiver module 310 for receiving the captured plurality of inputs and sending the bitstream;
a processing module 330;
a memory 320 storing a plurality of voice prints, a plurality of user selections and associations between user representations and voice prints; 10
wherein at least one of the captured plurality of inputs indicate a selected user representation from among a plurality of user presentations,
wherein at least one of the captured plurality of inputs represents an information to be shared with at least a subset of the first set of client devices,
wherein the processing module 330 is capable of selecting a voice print 15 from a plurality of voice prints based on the selected user representation,
wherein the processing module 330 is capable of transforming the information to be shared into an audio clip, and
wherein the processing module 330 is capable of converting the audio file into the bitstream. 20
6. The system as claimed in claim 5, wherein the input module 210 includes at least two of a pointing device, a microphone, and a keyboard.
7. The system as claimed in claim 5, wherein the output module 220 includes at least a display device and a speaker.
15
8. The system as claimed in claim 5, wherein the processing module 330 uses a neural network to generate the audio clip.
Dated this 17th of May, 2021 5
Applicant’s Patent Agent
of Effectual Legal Services Pvt. Ltd. , Description:1
FORM – 2
THE PATENTS ACT, 1970
(39 of 1970)
&
THE PATENTS RULES, 2003
COMPLETE SPECIFICATION
(See section 10 and rule 13)
METHOD AND SYSTEM FOR IMPROVING ONLINE INTERACTION
Applicant: 1. Rajat
House no 485, Ward No. 5, Mehrauli, New Delhi 110030, India
Inventor: 1) Rajat
The following specification particularly describes the invention and the manner in which it is to be performed
2
TECHNICAL FIELD
[0001] The present invention is directed towards improving the user experience when interacting online with other user(s) and/or bot(s) online.
DESCRIPTION OF THE RELATED ART 5
[0002] The development of internet has improved human communication immensely. Anyone using internet is now capable of interacting with anyone else using internet in near real-time, wherein the time difference between sender transmitting the interaction and receiver experiencing the interaction is imperceivable to humans. The interactions can take one of many forms, e.g. text, voice, video. 10
[0003] These exchanges of interactions are enabled by servers, provided by organizations. A user wishing to interact with other users will connect their client device to the server via internet connections. Upon successfully connecting with server, the user can interact with any other user, or even bot, which are software capable of interacting with humans. 15
[0004] The organizations provide these servers either as standalone service or as a part of bigger service. One example of the being part of bigger service is communication capability between users experiencing a virtual and/or augmented reality or between a user experiencing a virtual and/or augmented reality and a bot which is part of said virtual and/or augmented reality. 20
[0005] Another example is providing a bot as a part of website, hosted on server to answer any queries that a user visiting the website may have.
[0006] It is customary for a user entering the virtual and/or augmented reality to select a character to represent the user in the said reality. Generally, there are few phrases pre-recorded in a voice actor’s voice, each corresponding to one character. 25
3
[0007] Generally, these virtual and/or augmented realities also allow a user to send and/or talk to other users who are part of same virtual and/or augmented reality. However, listening to a character speak in a voice actor’s voice some time and in user’s voice rest of the time effects the user experience. Further, some of these realities allow users to interact using text-based chat. However, reading text during performing the 5 tasks in virtual and/or augmented reality makes reduces efficiency of user.
[0008] In view of the foregoing, a method and corresponding system is disclosed to address the abovementioned issues in the art.
SUMMARY 10
[0009] In view of the foregoing, an embodiment herein provides a method for receiving a character selection from a user, receiving a voice clip and/or text message from the user, converting the voice clip and/or text message into a voice clip in the voice of voice actor corresponding to the selected character. The voice clip in the voice of the voice actor is sent to one or more of users. The set of users receiving the voice 15 clip may or may not include the sender of the voice clip and/or text message.
[0010] In some embodiments, conversion takes place using an Artificial Intelligence, sometimes abbreviated as AI, wherein the AI is trained to generated audio clips in the voice actors corresponding to selected characters.
[0011] In some embodiments, the method includes a text-to-speech generator. 20
[0012] In some embodiments, the method includes a speech-to-text generator.
BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The embodiments herein will be better understood from the following detailed description with reference to the drawings, in which: 25
4
[0014] FIG. 1 illustrates a first exemplary environment wherein the invention may be implemented.
[0015] FIG. 2 illustrates an exploded view of a client device 200 according to some embodiments herein.
[0016] FIG. 3 illustrates an exploded view of a server 300 according to some 5 embodiments herein.
[0017] FIG. 4 illustrates a method employed at the server 300 according to some embodiments herein.
10
15
20
5
DETAILED DESCRIPTION OF PREFERRED EMBODIMENTS
[0018] The embodiments herein and the various features and advantageous details thereof are explained more fully with reference to the non-limiting embodiments that are illustrated in the accompanying drawings and detailed in the following description. Descriptions of well-known components and processing techniques are omitted so as 5 not to unnecessarily obscure the embodiments herein. The examples used herein are intended merely to facilitate an understanding of ways in which the embodiments herein may be practiced and to further enable those of skill in the art to practice the embodiments herein. Accordingly, the examples should not be construed as limiting the scope of the embodiments herein. 10
[0019] As mentioned, there remains a need for a system and a method of improving the online interaction experiences for users of web services provided by various organizations. Referring now to the drawings, and more particularly to FIGS. 1 through 4, where similar reference characters denote corresponding features consistently throughout the figures, there are shown preferred embodiments. 15
[0020] A "server" refers to a piece of software and/or hardware that allows an organization to provide services to users of the internet. A server can be a combination of multiple machines also. The defining characteristic of the server is that it receives data from a user’s client device and is capable of processing said received data. Further, the server may also be capable of transmitting data to one or more of client devices 20 connected to it. Terms “server” and “computing entity” are used interchangeably in this application.
[0021] A "client device" refers to a piece of software and/or hardware that allows a user to receive services made available online by one or more organization. The defining characteristic of the client device is that it sends and/or receives data to/from 25 at least one server and is capable of outputting said data to user. Client device can also
6
be a “thin client”. Terms “client” and “user” are used interchangeably in this application.
[0022] A "voice print" refers to measurable qualities of a voice that uniquely identifies a speaker.
[0023] A "AI" refers to any software capable of self-learning and decision making. For 5 the purpose of this invention, term “AI” includes all of machine learning software as well as neural network based software, [0024] FIG. 1 illustrates a first exemplary environment diagram for a system and method to improve a user’s experience of interacting with other users and/or bots according to an embodiment herein. The organization offering services to users of internet will upload the services to a server 10 and inform users on how to connect to the host server. The organization may inform users on how to connect to the users explicitly, e.g., by providing a Uniform Resource Locator (URL), or implicitly, e.g., providing the user with a software program wherein the program contains URL and connects to server without informing the user of the URL. 15
[0025] A user of client device 200, desiring to use the services of an organization, may connect to the server hosting the services in many ways. In one embodiment, the user starts an internet browser program installed on the client device and input the URL of the server on which the services are hosted. In another embodiment user starts a program previously installed on the system. The program’s code includes the 20 instructions on how to connect to the server on which the services are hosted.
[0026] The client device 200 establishes a link to the server using an internet connection. The link may be direct, e.g., using a VPN tunnel or when user is aware of public IP address and/or port of the server or it maybe indirect wherein the connection passes through a plurality of networking devices, e.g., Domain Name Servers, routers, 25 switches, gateways etc.
7
[0027] FIG. 2 illustrates an exploded view of the client devices 200 for using services hosted on a server across the internet according to an exemplary embodiment of the invention. The client device 200 includes an input module 210, an output module 220, a processing module 230 and a transceiver module 240. Please note that the client device 200 may have even more modules for various other functions. However, to 5 avoid unnecessary details, we will not be explaining those. For example, the client device may also include memory module, not shown in the FIG. 2, to store one or more of softwares, firmwares, operating systems, and data. Further, in one embodiment, one or more of the modules may be combined in one module.
[0028] The input module 210 consists of all the submodules that allow a user to input 10 information to the client device. According to certain embodiments, the input module 210 consists of a speech input capture submodule 211, a text input capture submodule 212 and a selection input capture submodule 213. A non-limiting example of the speech input capture submodule 211 is a microphone. The microphone receives a speech spoken by users, generates an analogue signal representing spoken speech and finally, 15 converts the analogue signal into a digital representation. This digital representation is further processed. A non-limiting example of text input capture submodule 212 is a keyboard. The keyboard contains keys that a user presses to input texts, commands, and special characters etc. The Keyboard recognizes what key was pressed and generates a corresponding digital representation. This digital representation is further 20 processed. A non-limiting example of selection input capture submodule 213 is a mouse. The mouse allows a user to control selection and input selection. The mouse generates a digital representation of selection and forwards it for the processing.
[0029] The output module 220 consists of all the submodules that allow a user to receive information from the client device 200. According to certain embodiments, the 25 output module 220 consists of a visual output submodule 221 and an audio output submodule 222. A non-limiting example of the visual output submodule 221 is a display. The display receives digital information from the processing module 230 and
8
converts the digital information using a digital-to-analog converter. This analog information is used to drive circuits to present information for the user to see. A non-limiting example of audio output submodule is a speaker. The speaker receives digital information from processing module and converts it into a sound wave, which is outputted to the user’s ears. 5
[0030] The processing module 230 consists of all the submodules that allow a client device to process information. According to certain embodiments, the processing module 230 consists of a central processing submodule 231 and, optionally, a graphics processing submodule 232. The central processing submodule 231 receives digital information from all the modules, processes received information and sends the result 10 of processing to appropriate module. A non-limiting example of the central processing submodule 231 is a CPU. In some embodiments, the central processing submodule 231 is accompanied by a graphics processing submodule 232. A non-limiting example of a graphics processing submodule is a GPU. The graphics processing submodule 232 is a hardware and accompanying firmware that are specifically built for processing visual 15 information.
[0031] The transceiver module 240 consists of all the submodules that allow a client device to establish connections with other client devices and/or servers. According to certain embodiments, the transceiver module 240 consists of a receiving submodule 241 and a transmitting submodule 242. In certain embodiments logical receiving 20 submodule 241 and transmitting submodule 242 are implemented in same physical circuits. Few non-limiting examples of the transceiver module 240 are Wi-Fi modules, Ethernet Cards, Cellular Antennas etc.
[0032] FIG. 3 illustrates an exploded view of the server 300 used for hosting services according to an exemplary embodiment of the invention. Server includes a transceiver 25 module 310, a processing module 330 and a memory module 320. Please note that the server 300 may have even more modules for various other functions. However, to avoid
9
unnecessary details, we will not be explaining those. Further, in one embodiment, one or more of the modules may be combined in one module.
[0033] The transceiver module 310 consists of all the submodules that allow a server 300 to establish connections with other client devices and/or servers. According to certain embodiments, the transceiver module 310 consists of a receiving submodule 5 311 and a transmitting submodule 312. In certain embodiments logical receiving submodule 311 and transmitting submodule 312 are implemented in same physical circuits. Few non-limiting examples of The transceiver module 310 are Wi-Fi modules, Ethernet Cards, Cellular Antennas etc.
[0034] The memory module 320 consists of all the submodules that allow the server 10 to store one or more of softwares, firmwares, operating systems, and data. According to certain embodiments, the transceiver module 310 consists of a voice print database submodule 321, a Character-to-Voice Print mapping submodule 322 and an instruction submodule 323. In certain embodiments one or more of the voice print database submodule submodule 321, the Character-to-Voice Print mapping submodule 322 and 15 the instruction submodule 323 may be implemented in same physical circuits.
[0035] The voice print database submodule 321 stores voiceprints of all voice actors in whose voice an audio output might be generated. The Character-to-Voice Print mapping submodule 322 stores an association between a selectable character presented to the user and voice print of an actor chosen as the voice of the character. The 20 instruction submodule 323 stores all the instructions that may be executed by processing module 330.
[0036] The processing module 330 consists of all the submodules that allow a server to receive data and/or instructions, process the data and/or instructions and output the result of processing. According to certain embodiments, the processing module 330 25 consists of an AI submodule 331, a Text to speech submodule 332 and a Speech Synthesis submodule 333. In certain embodiments one or more of the AI submodule
10
331, the text to speech submodule 332 and the speech synthesis submodule 333 may be implemented in same physical circuits. In certain embodiments the processing module 330 further includes a speech recognition submodule, not shown in the figure.
[0037] Fig. 4 illustrates a method of invention, implemented at the server according to one of the embodiments. The method starts after an internet connection has been 5 established between a client device 200 and the server 300. The client device 200 outputs to user of client device 200 a plurality of characters. The user selects one among the presented plurality of characters, to represent the user in the online environment. The selected character may be considered as user representation. As non-limiting example, the client device 200 displays plurality of user avatars to user. As another 10 non-limiting example, client device 200 displays multiple topics and asks user to select the topic on which the user requires help and/or information, wherein each topic has a pre-associated character. User may or may not be aware of the association between the topics and characters. At step 410, the server 300 receives character selection made by user over the established connection. At step 420, the server 300 creates and stores an 15 association between user and character. The server 300 may store this association between user and character either temporarily or permanently. Sometime later, the user input some information at the client devices. As non-limiting example, the user speaks into the microphone of the client device 200 or types a text using keyboard. At step 430, the server 300 receives digital representation of the input from client device 200. 20 At the step 440, the server 300 extracts input from the received digital representation. As non-limiting example, the server 300 uses a speech recognition submodule to recognize what the user spoke or recognizes the text potion from keyboard input. The server 300 also extracts emotions/tones from the input. As a non-limiting example, the server 300 recognizes the user spoke the input in agitated voice. As another non-25 limiting example, the server 300 recognizes that the text was typed in a frustrated manned. At the step 450, the server 300 decides a response. As a non-limiting example, if the user asked a question appropriate, the server 300 finds the answer to the question,
11
either in memory of the server 300 or over the internet. if the server 300 unable to find answer, the server 300 may inform user of the same. As another non-limiting example, If the user provide an instruction, the server 300 may acknowledge the request or if the server is incapable of completing the task, the server 300 may convey the same to the user. As another non-limiting example, the server 300 may filter out some information. 5 The set of client devices may or may not include the client device 200 which submitted the information. At step 460, the server 300 uses an AI-powered speech synthesis system to generate an audio clip in the voice of the voice actor associated with the character selected by the user. The AI-powered speech synthesis is trained using the voice prints of all the voice actors and can generate audio clips in voice of voice actors 10 in a range of emotions. The emotion in which the audio clip is generated is determined based on the emotion extracted in step 440. As a non-limiting example, in an environment where the user has been interacting with the server 300 to seek information and it is sensed that the user has become frustrated, the audio clip may be generated in a calmer or soothing tone. As another non-limiting example, in an 15 environment where user has been interacting with server to communicate with other members of his gaming party, the audio clip may be generated in tone which was extracted at step 440. At step 470, the server converts the audio clip in a format fit for transmission over the internet and transmits the converted audio clip to intended recipients. The intended recipients may simply be the user, or the server 300 may 20 determine based on some rules, e.g., all users in the user’s party, all users who the user specified as recipients, all other users whose characters are within a defined proximity of user’s character in a virtual world, etc.
[0038] The foregoing description of the specific embodiments has fully revealed the general nature of the embodiments herein that others can, by applying current 25 knowledge, readily modify and/or adapt for various applications such specific embodiments without departing from the generic concept, and, therefore, such adaptations and modifications should and are intended to be comprehended within the
12
meaning and range of equivalents of the disclosed embodiments. It is to be understood that the phraseology or terminology employed herein is for the purpose of description and not of limitation. Therefore, while the embodiments herein have been described in terms of preferred embodiments, those skilled in the art will recognize that the embodiments herein can be practiced with modification within the spirit and scope. 5
Dated this 17th day of May, 2021
Applicant’s Patent Agent
of Effectual Legal Services Pvt. Ltd.
10
| Section | Controller | Decision Date |
|---|---|---|
| u/s 15 | Santanu Acharya | 2025-04-11 |
| u/s 15 | Santanu Acharya | 2025-04-11 |
| # | Name | Date |
|---|---|---|
| 1 | 202111022033-STATEMENT OF UNDERTAKING (FORM 3) [17-05-2021(online)].pdf | 2021-05-17 |
| 2 | 202111022033-REQUEST FOR EXAMINATION (FORM-18) [17-05-2021(online)].pdf | 2021-05-17 |
| 3 | 202111022033-PROOF OF RIGHT [17-05-2021(online)].pdf | 2021-05-17 |
| 4 | 202111022033-POWER OF AUTHORITY [17-05-2021(online)].pdf | 2021-05-17 |
| 5 | 202111022033-FORM 18 [17-05-2021(online)].pdf | 2021-05-17 |
| 6 | 202111022033-FORM 1 [17-05-2021(online)].pdf | 2021-05-17 |
| 7 | 202111022033-FIGURE OF ABSTRACT [17-05-2021(online)].jpg | 2021-05-17 |
| 8 | 202111022033-DRAWINGS [17-05-2021(online)].pdf | 2021-05-17 |
| 9 | 202111022033-DECLARATION OF INVENTORSHIP (FORM 5) [17-05-2021(online)].pdf | 2021-05-17 |
| 10 | 202111022033-COMPLETE SPECIFICATION [17-05-2021(online)].pdf | 2021-05-17 |
| 11 | 202111022033-Proof of Right [17-08-2021(online)].pdf | 2021-08-17 |
| 12 | 202111022033-FORM-26 [17-08-2021(online)].pdf | 2021-08-17 |
| 13 | 202111022033-Retyped Pages under Rule 14(1) [06-03-2023(online)].pdf | 2023-03-06 |
| 14 | 202111022033-2. Marked Copy under Rule 14(2) [06-03-2023(online)].pdf | 2023-03-06 |
| 15 | 202111022033-FER.pdf | 2023-09-25 |
| 16 | 202111022033-Power of Authority [26-03-2024(online)].pdf | 2024-03-26 |
| 17 | 202111022033-PETITION u-r 6(6) [26-03-2024(online)].pdf | 2024-03-26 |
| 18 | 202111022033-OTHERS [26-03-2024(online)].pdf | 2024-03-26 |
| 19 | 202111022033-FER_SER_REPLY [26-03-2024(online)].pdf | 2024-03-26 |
| 20 | 202111022033-DRAWING [26-03-2024(online)].pdf | 2024-03-26 |
| 21 | 202111022033-Covering Letter [26-03-2024(online)].pdf | 2024-03-26 |
| 22 | 202111022033-CLAIMS [26-03-2024(online)].pdf | 2024-03-26 |
| 23 | 202111022033-ABSTRACT [26-03-2024(online)].pdf | 2024-03-26 |
| 24 | 202111022033-US(14)-HearingNotice-(HearingDate-09-07-2024).pdf | 2024-06-28 |
| 25 | 202111022033-Correspondence to notify the Controller [06-07-2024(online)].pdf | 2024-07-06 |
| 26 | 202111022033-Annexure [06-07-2024(online)].pdf | 2024-07-06 |
| 27 | 202111022033-RELEVANT DOCUMENTS [08-07-2024(online)].pdf | 2024-07-08 |
| 28 | 202111022033-POA [08-07-2024(online)].pdf | 2024-07-08 |
| 29 | 202111022033-FORM 13 [08-07-2024(online)].pdf | 2024-07-08 |
| 30 | 202111022033-PETITION UNDER RULE 137 [17-07-2024(online)].pdf | 2024-07-17 |
| 31 | 202111022033-FORM 13 [17-07-2024(online)].pdf | 2024-07-17 |
| 32 | 202111022033-AMMENDED DOCUMENTS [17-07-2024(online)].pdf | 2024-07-17 |
| 33 | 202111022033-Written submissions and relevant documents [21-07-2024(online)].pdf | 2024-07-21 |
| 34 | 202111022033-US(14)-ExtendedHearingNotice-(HearingDate-29-08-2024)-1200.pdf | 2024-08-09 |
| 35 | 202111022033-Correspondence to notify the Controller [28-08-2024(online)].pdf | 2024-08-28 |
| 1 | SearchHistory_202111022033E_22-09-2023.pdf |