Sign In to Follow Application
View All Documents & Correspondence

Speech Processing Device, Speech Processing Method, And Recording Medium

Abstract: A call partner identification means identifies a call partner in order to make it possible for a user to easily identify the call partner by only the sense of hearing. A background sound selection means selects a background sound corresponding to the identified call partner. A synthesis means synthesizes a call speech signal and the selected background sound.

Get Free WhatsApp Updates!
Notices, Deadlines & Correspondence

Patent Information

Application #
Filing Date
21 January 2022
Publication Number
12/2022
Publication Type
INA
Invention Field
ELECTRONICS
Status
Email
Parent Application

Applicants

NEC CORPORATION
7-1, Shiba 5-chome, Minato-ku, Tokyo 1088001

Inventors

1. KAWABE Tomoyuki
c/o NEC CORPORATION, 7-1, Shiba 5-chome, Minato-ku, Tokyo 1088001

Specification

Title of the invention: voice processing device, voice processing method, and recording medium
Technical field
[0001]
The present invention relates to a voice processing device, a voice processing method, and a recording medium, and more particularly to a voice processing device, a voice processing method, and a recording medium that synthesize a voice signal of a call with a background sound.
Background technology
[0002]
Related acoustic AR (Augmented Reality) technology allows users to talk to multiple parties at the same time. For example, the related technique described in Patent Document 1 uses a posture information sensor and left and right independent speakers to define a virtual position for the user to perceive that voice is heard from the other party (sound image localization). .. As a result, the user feels as if he / she is sharing a virtual space with a plurality of callees.
[0003]
Patent Document 2 describes that the speaker information (for example, the name or nickname of the other party) heard from the direction specified by the user is output by voice by a predetermined operation of the user.
Prior art literature
Patent documents
[0004]
Patent Document 1: Japanese Patent Application Laid-Open No. 2013-017027
Patent Document 2: Japanese Unexamined Patent Publication No. 2012-103845
Patent Document 3: International Publication No. 2018/1983310
Outline of the invention
Problems to be solved by the invention
[0005]
If there is a display, the user can determine who the other party is by looking at the image of the other party, the name of the other party, and the nickname displayed on the display. However, in the absence of a display, the user must determine who the other party is by hearing alone. However, in the related technique described in Patent Document 1, when the other party is silent, the user cannot know who the other party is by hearing. Further, in the technique described in Patent Document 2, since the user does not notice that there is a silent call partner, there is no motivation to perform a predetermined operation in order to output the speaker information by voice.
[0006]
The present invention has been made in view of the above problems, and an object of the present invention is to provide a voice processing device or the like that enables a user to easily identify a call partner only by hearing.
Means to solve problems
[0007]
The voice processing device according to one aspect of the present invention talks with a call partner identification means for identifying a call partner, a background sound selection means for selecting a background sound corresponding to the identified call partner, and the selected background sound. It is equipped with a synthesis means for synthesizing the voice signal of.
[0008]
The voice processing method according to one aspect of the present invention identifies a call partner, selects a background sound corresponding to the identified call partner, and synthesizes the selected background sound with the voice signal of the call.
[0009]
A program recorded on a computer-readable recording medium according to one aspect of the present invention identifies a call partner, selects a background sound corresponding to the identified call partner, and selects the background sound. , Let the computer perform the synthesis with the voice signal of the call.
The invention's effect
[0010]
According to one aspect of the present invention, the user can easily identify the other party by hearing only.
A brief description of the drawing
[0011]
FIG. 1 is a block diagram showing a configuration of a voice processing device according to the first embodiment.
FIG. 2 is a flowchart showing the operation of the voice processing device according to the first embodiment.
FIG. 3 is a block diagram showing a configuration of a call device on the speaker side according to the second embodiment.
FIG. 4 is a block diagram showing a configuration of a call device on the listener side according to the second embodiment.
FIG. 5 is a diagram showing an example of a setting list showing a correspondence relationship between a caller ID that identifies a call partner, sound image localization information, and background acoustic information.
FIG. 6 is a flowchart showing a flow of operation of the listener-side call device according to the second embodiment.
[Fig. 7] Shows a user who is talking to multiple parties.
FIG. 8 is a block diagram showing a configuration of a call device on the listener side according to the third embodiment.
FIG. 9 is a flowchart showing a flow of operation of the listener-side call device according to the third embodiment.
FIG. 10 is a diagram illustrating the operation of a call device when a call is made between the designated group and the other party when the first group is a designated group.
FIG. 11 is a diagram illustrating the operation of a call device when a call is made between the designated group and the other party when the second group is a designated group.
FIG. 12 is a diagram showing a hardware configuration of an information processing apparatus according to the fifth embodiment.
Embodiment for carrying out the invention
[0012]
[Embodiment 1]
The first embodiment will be described with reference to FIGS. 1 to 2.
[0013]
In the first embodiment, the user is talking to the other party by using the calling device. The other party may be one or more people. The calling device used by the user is hereinafter referred to as the user's calling device. Further, the calling device used by the user's calling party is hereinafter referred to as a calling device of the calling party. The calling device is, for example, a mobile terminal, a computer device, or a telephone. The calling device is also called a voice device.
[0014]
A communication network, for example, the Internet or a telephone line, is connected between the calling device of the other party and the calling device of the user. The user is talking to the other party using the user's calling device.
[0015]
(Voice processing device 1)
FIG. 1 is a block diagram showing a configuration of a voice processing device 1 according to the first embodiment. As shown in FIG. 1, the voice processing device 1 includes a call partner identification unit 11, a background sound selection unit 12, and a synthesis unit 13. The functions of each part of the voice processing device 1 described below may be realized as software by the processor included in the user's telephone device (not shown) executing a program read into the memory.
[0016]
The call partner identification unit 11 identifies the call partner. The other party identification unit 11 is an example of the other party identification means. Specifically, the call partner identification unit 11 receives incidental information from a call device (not shown) of the call partner. Ancillary information includes information that identifies the calling device of the other party.
[0017]
The incidental information is, for example, a terminal ID (Identification) that identifies the calling device of the other party. Alternatively, the incidental information may be the telephone number of the calling device of the other party. However, the incidental information is not limited to the terminal ID and the telephone number. The incidental information may be information that identifies the other party (for example, biometric authentication information).
[0018]
The call partner identification unit 11 identifies the call partner based on the incidental information received from the call device of the call partner.
[0019]
For example, when the incidental information is the terminal ID of the call device of the call partner, the call partner identification unit 11 refers to the data of the owner list stored in the memory (not shown). The data of the owner list (hereinafter, simply referred to as the owner list) is associated with the terminal ID and the information about the owner of the calling device.
[0020]
The call partner identification unit 11 identifies the call partner based on the terminal ID. For example, the call partner identification unit 11 identifies that the owner of the call device identified by the terminal ID is the call partner by referring to the owner list.
[0021]
The call partner identification unit 11 transmits information indicating the identified call partner (for example, a caller ID included in the owner list) to the background sound selection unit 12.
[0022]
The background sound selection unit 12 selects the background sound according to the identified call partner. The background sound selection unit 12 is an example of the background sound selection means. For example, the background sound selection unit 12 selects a background sound according to the other party by referring to a caller-background sound table (not shown) showing a correspondence relationship between the caller and the background sound. The caller-background acoustic table is stored in memory or storage.
[0023]
The background sound is the sound that the user hears so as to overlap with the main voice (here, the voice of the call). In general, the background sound has a low volume that does not prevent the user from listening to the main sound. The background sound is, for example, one of BGM (Back Ground Music), ambient sound (ambient), and sound effect. Alternatively, the background sound may be a combination of these.
[0024]
For example, in the caller-background sound table, the caller ID that identifies the other party is associated with the information indicating the background sound (background sound information). The caller-background acoustic table may be prepared in advance by the user (here, the listener).
[0025]
Alternatively, the background sound selection unit 12 may select a different background sound for each other party (here, the speaker) by any method. For example, the background sound selection unit 12 assigns different background sounds to each other party based on the sensing information acquired as incidental information (Embodiment 4).
[0026]
The synthesis unit 13 synthesizes the voice signal of the call and the selected background sound. The synthesizing unit 13 is an example of synthesizing means. Specifically, the synthesizing unit 13 generates an audio signal (hereinafter, referred to as an audio signal with background acoustic) in which the background acoustic is superimposed on the audio signal of the call.
[0027]
Here, when another existing background sound is already superimposed on the voice signal of the call, the synthesis unit 13 outputs the background sound selected by the background sound selection unit 12 together with the existing other background sound. You may. Alternatively, the synthesizing unit 13 may mute or reduce the volume of other existing background sounds, while the background sound selected by the background sound selection unit 12 may be output at a louder volume than the other existing background sounds. ..
[0028]
Further, the compositing unit 13 performs sound image localization processing on the audio signal with background sound so that the user can hear the audio signal with background sound from a virtual place preset for each other party. May be good.
[0029]
Specifically, in the sound image localization process, the synthesis unit 13 adjusts the volume and presence of the audio signal with background sound. As a result, the synthesis unit 13 defines a virtual position of the sound image of the audio signal with background sound. The user who listens to the audio signal with background sound output from the speaker or the like perceives that the sound image of the audio signal with background sound is at a virtual position defined by the synthesizer 13.
[0030]
Note that the synthesis unit 13 may convert the pitch, intensity, timbre, or a combination thereof with respect to the background sound according to the group to which the call partner identified by the call partner identification unit 11 belongs. The configuration in which the other party is classified into a plurality of groups will be described in the third embodiment.
[0031]
The synthesis unit 13 outputs the audio signal with background sound generated in this way.
[0032]
(Operation of voice processing device 1)
The operation of the voice processing device 1 according to the first embodiment will be described with reference to FIG. FIG. 2 is a flowchart showing a flow of processing executed by the voice processing device 1.
[0033]
As shown in FIG. 2, the call partner identification unit 11 identifies the call partner based on the incidental information received from the call device of the call partner (S1).
[0034]
The call partner identification unit 11 transmits information indicating the identified call partner (for example, a caller ID) to the background sound selection unit 12.
[0035]
The background sound selection unit 12 selects the background sound according to the other party based on the information (for example, the caller ID) received from the other party identification unit 11 (S2).
[0036]
The background sound selection unit 12 transmits background sound information indicating the selected background sound to the synthesis unit 13. The order of steps S1 and S2 may be reversed or parallel.
[0037]
The synthesis unit 13 receives background sound information from the background sound selection unit 12. Further, the synthesis unit 13 receives the voice signal of the call.
[0038]
The synthesis unit 13 has the background sound indicated by the background sound information and the voice signal of the call.And are synthesized (S3). Then, the synthesis unit 13 outputs a voice signal (voice signal with background sound) obtained by synthesizing the background sound and the voice signal of the call.
[0039]
This completes the operation of the voice processing device 1.
[0040]
(Effect of this embodiment)
According to the configuration of the present embodiment, the call partner identification unit 11 identifies the call partner. The background sound selection unit 12 selects the background sound according to the identified call partner. The synthesizing unit 13 synthesizes the voice signal of the call and the selected background sound. In this way, the voice signal with the background sound obtained by synthesizing the voice signal of the call and the background sound is output. Therefore, the user can easily identify the other party only by hearing.
[0041]
[Embodiment 2]
The second embodiment will be described with reference to FIGS. 3 to 7. In the second embodiment, the speaker-side call device and the listener-side call device will be described.
[0042]
(Call device 100 on the speaker side)
FIG. 3 is a block diagram showing a configuration of a call device 100 on the speaker side (hereinafter, simply referred to as a call device 100). As shown in FIG. 3, the telephone device 100 includes a microphone 101, a memory 102, and a communication unit 103.
[0043]
The calling device 100 may further include a sensor or switch for biometrically authenticating the individual speaker. For example, the call device 100 may acquire information about the posture of the speaker wearing the call device 100 by an image sensor (camera), or may further include a switch or a sensor for selecting a call partner. ..
[0044]
The microphone 101 collects the speaker's remarks. The microphone 101 converts the speaker's remark into an audio signal, and transmits the audio signal of the call to the communication unit 103.
[0045]
The memory 102 stores at least the terminal ID for identifying the own device (that is, the calling device 100) and the information indicating the telephone number of the calling device 100.
[0046]
The communication unit 103 acquires information indicating a terminal ID or a telephone number from the memory 102. The communication unit 103 transmits the voice signal of the call received from the microphone 101 to the call device 200 on the listener side, and also the information indicating the terminal ID or the telephone number acquired from the memory 102 as incidental information on the call on the listener side. Send to device 200. In addition, the communication unit 103 may also transmit the sensing information acquired by the call device 100 to the call device 200 on the listener side as incidental information. For example, the incidental information is a terminal ID (Identification) for identifying the calling device of the other party, or a telephone number of the other party.
[0047]
(Call device 200 on the listener side)
FIG. 4 is a block diagram showing a configuration of a call device 200 on the listener side (hereinafter, simply referred to as a call device 200). As shown in FIG. 4, the telephone device 200 includes a receiving unit 24, a voice processing device 2, and an output control unit 26.
[0048]
As shown in FIG. 4, the receiving unit 24 receives an audio signal and incidental information (here, a terminal ID) from the communication unit 103 of the communication device 100. The receiving unit 24 transmits the received voice signal to the synthesis unit 13 of the voice processing device 2. Further, the receiving unit 24 transmits the received incidental information to the call partner identification unit 11 of the voice processing device 2.
[0049]
(Voice processing device 2)
As shown in FIG. 4, the voice processing device 2 according to the second embodiment includes a call partner identification unit 11, a background sound selection unit 12, and a synthesis unit 13. In the first embodiment, the details of the other party identification unit 11, the background sound selection unit 12, and the composition unit 13 have been described. In the second embodiment, each of these parts will be briefly described.
[0050]
The call partner identification unit 11 identifies the call partner. Specifically, the call partner identification unit 11 identifies the call partner based on the incidental information received from the call device of the call partner.
[0051]
The call partner identification unit 11 transmits the identification result of the call partner to the background sound selection unit 12. Specifically, the call partner identification unit 11 transmits the caller ID that identifies the call partner to the background sound selection unit 12.
[0052]
The background sound selection unit 12 selects the sound corresponding to the call partner identified by the call partner identification unit 11. Specifically, the background sound selection unit 12 refers to the setting list (see FIG. 5). When the caller ID of the other party identified by the other party identification unit 11 is registered in the setting list, the background sound selection unit 12 is one BGM associated with the caller ID of the other party in the setting list ( Back Ground Music) data is acquired. BGM is a kind of background sound.
[0053]
On the other hand, when the call partner identified by the call partner identification unit 11 is not yet registered in the setting list, the background sound selection unit 12 refers to the BGM library (not shown) and which call is registered in the setting list. The data of one BGM that is not associated with the person ID is randomly acquired from the BGM library.
[0054]
Further, the background sound selection unit 12 acquires the sound image localization information associated with the caller ID from the setting list. The sound image localization information is information indicating the coordinate position in the virtual space.
[0055]
The background sound selection unit 12 transmits the selected BGM data (background sound information) and the sound image localization information to the synthesis unit 13.
[0056]
The synthesis unit 13 of the call device 200 receives BGM data and sound image localization information from the background sound selection unit 12. Further, the synthesis unit 13 receives the voice signal of the call from the reception unit 24. The synthesizing unit 13 synthesizes the voice signal of the call and the selected background sound. Specifically, the synthesizing unit 13 generates a voice signal with a background sound in which a background sound (here, BGM) is superimposed on the voice signal of the call.
[0057]
Here, when another existing background sound is already superimposed on the voice signal of the call, the synthesis unit 13 may output the BGM selected by the background sound selection unit 12 together with the other existing background sound. good. Alternatively, the synthesizing unit 13 may mute or reduce the volume of the other existing background sound, while the BGM selected by the background sound selection unit 12 may be output at a louder volume than the other existing background sound.
[0058]
Further, the compositing unit 13 performs the above-mentioned sound image localization processing on the audio signal with background sound so that the user can hear the audio signal with background sound from a virtual position preset for each other party. You may go.
[0059]
The synthesis unit 13 transmits an audio signal with background sound to the output control unit 26.
[0060]
As shown in FIG. 4, the output control unit 26 receives an audio signal with background sound from the synthesis unit 13. The output control unit 26 outputs the received audio signal with background sound to a device such as a speaker (not shown). The output control unit 26 may be a component of the voice processing device 2.
[0061]
(Setting list)
FIG. 5 is a table showing an example of the setting list. As shown in FIG. 5, the setting list includes the caller ID as incidental information for identifying the other party. In the setting list, the sound image localization information and the background acoustic information are associated with each caller ID.
[0062]
The sound image localization information indicates the virtual position of the sound image. The background acoustic information allows the user to identify the other party from the virtual position of the sound image only by hearing. The background acoustic information is, for example, BGM, environmental sound, or acoustic effect data associated with a caller ID for identifying an individual caller.
[0063]
(Operation of call device 200)
The operation of the call device 200 according to the second embodiment will be described with reference to FIGS. 6 and 7. FIG. 6 is a flowchart showing the operation flow of the communication device 200.
[0064]
As shown in FIG. 6, the receiving unit 24 of the calling device 200 of the user (A in FIG. 7) receives the voice signal and incidental information of the call from the calling device 100 of the other party (B to F in FIG. 7). (S101).
[0065]
The receiving unit 24 transmits the voice signal of the call to the synthesis unit 13 of the voice processing device 2, and also transmits the incidental information (here, the terminal ID) to the call partner identification unit 11.
[0066]
The call partner identification unit 11 receives incidental information from the reception unit 24. The call partner identification unit 11 identifies the call partner based on the received incidental information (S102). The other party is one or more.
[0067]
The call partner identification unit 11 transmits the identification result of the call partner to the background sound selection unit 12. Specifically, when the call partner identification unit 11 succeeds in identifying the call partner, the call partner identification unit 11 transmits the caller ID that identifies the call partner to the background sound selection unit 12. On the other hand, when the call partner identification unit 11 fails to identify the call partner, the call partner ID (None) is transmitted to the background sound selection unit 12.
[0068]
The background sound selection unit 12 receives the identification result of the other party from the other party identification unit 11. The background sound selection unit 12 determines whether or not the other party has been registered in the setting list (S103). Specifically, the background sound selection unit 12 determines whether or not the caller ID received from the call partner identification unit 11 is included in the setting list.
[0069]
When the other party is already registered in the setting list (Yes in S103), the background sound selection unit 12 selects the background sound (here, BGM) according to the other party (S104A). On the other hand, when the other party is not registered in the setting list (No in S103), the background sound selection unit 12 refers to the BGM library and is not associated with any caller ID registered in the setting list (No). Here, BGM) is randomly selected (S104B).
[0070]
The background sound selection unit 12 transmits the selected BGM data (background sound information) to the synthesis unit 13 together with the sound image localization information corresponding to the caller ID of the other party.
[0071]
The synthesis unit 13 receives the background sound information and the sound image localization information selected according to the other party from the background sound selection unit 12. Further, the synthesis unit 13 receives the voice signal of the call from the reception unit 24.
[0072]
The synthesizing unit 13 synthesizes the voice signal of the received call and the BGM which is the background sound (S105). As a result, the synthesis unit 13 generates an audio signal with background sound.
[0073]
The synthesis unit 13 transmits the generated audio signal with background sound to the output control unit 26.
[0074]
The output control unit 26 receives an audio signal with background sound from the synthesis unit 13. The output control unit 26 outputs the received audio signal with background sound (S106).
[0075]
With the above, the operation of the calling device 200 is completed.
[0076]
(Concrete example)
A user (person shown by A in FIG. 7) is making a call at the same time as a plurality of call parties (persons shown by B to F in FIG. 7) using the call device 200.
[0077]
As long as the line between the other party's calling device 100 (FIG. 3) and the user's calling device 200 (FIG. 4) is not disconnected, the user's calling device 200 identifies the other party even if the other party is silent. The BGM associated with the caller ID is output as the background sound. As a result, the user can intuitively know the existence of the silent call partner only by hearing.
[0078]
When the user presses the call end button (not shown) of the call device 200, the telephone line between the call device 100 and the call device 200 is disconnected. At this time, the call device 200 mutes (mute) the BGM corresponding to the caller ID that identifies the other party. This allows the user to intuitively know that the call has ended.
[0079]
(Modification example)
In one modification, each part of the voice processing device 2 may execute the above-mentioned processing on the network server without using the internal resources of the calling device 200.
[0080] [0080]
(Effect of this embodiment)
According to the configuration of the present embodiment, the call partner identification unit 11 identifies the call partner. The background sound selection unit 12 identifies the call.Select the background sound according to the other party. The synthesizing unit 13 synthesizes the voice signal of the call and the selected background sound. In this way, the voice signal obtained by synthesizing the voice signal of the call and the background sound is output. Therefore, the user can easily identify the other party only by hearing.
[0081]
Further, according to the configuration of the present embodiment, the virtual position of each call partner is defined by the sound image localization information according to the call partner. This allows the user to talk to the other party as if they were sharing a virtual space.
[0082]
[Embodiment 3]
In the third embodiment, a configuration will be described in which a group to which the other party belongs is identified and processing is performed according to the identified group.
[0083]
(Call device 300)
FIG. 8 is a block diagram showing the configuration of the call device 300 according to the third embodiment. As shown in FIG. 8, the communication device 300 includes a voice processing device 3 instead of the voice processing device 2 described in the second embodiment.
[0084]
(Voice processing device 3)
As shown in FIG. 8, the voice processing device 3 according to the third embodiment further includes a group determination unit 14 in addition to the call partner identification unit 11, the background sound selection unit 12, and the synthesis unit 13. That is, the voice processing device 3 is different from the voice processing device 2 according to the second embodiment in that the voice processing device 3 includes the group determination unit 14.
[0085]
In the second embodiment, the details of the other party identification unit 11, the background sound selection unit 12, and the composition unit 13 have been described. In the third embodiment, the basic functions of each of these parts are the same as those in the second embodiment. Therefore, in the third embodiment, the description of each of these parts will be omitted.
[0086]
The call partner identification unit 11 receives the group designation information from the reception unit 24 together with the incidental information. The call partner identification unit 11 identifies the call partner based on the incidental information, as in the first embodiment. Further, the other party identification unit 11 determines whether or not the user belongs to the designated group based on the group designation information.
[0087]
If the user does not belong to the designated group, the other party identification unit 11 instructs the output control unit 26 to mute the output. When the user belongs to the designated group, the other party identification unit 11 transmits the identification result of the other party to the group determination unit 14.
[0088]
The group determination unit 14 receives the identification result of the other party from the other party identification unit 11. The group determination unit 14 determines the group to which the call partner identified by the call partner identification unit 11 belongs. The group determination unit 14 is an example of group determination means. The group determination unit 14 transmits the determination result of the group to which the call partner belongs to the background sound selection unit 12 together with the identification result of the call partner.
[0089]
The background sound selection unit 12 selects a BGM according to the group to which the other party belongs based on the determination result of the group received from the group determination unit 14. It is preferred that the background acoustics (ie, the plurality of BGMs) corresponding to the plurality of callers belonging to the same group have the same characteristics (eg, the same pitch, the same volume, or the same timbre).
[0090]
(Operation of call device 300)
The operation of the call device 300 according to the third embodiment will be described with reference to FIGS. 9 to 11. FIG. 9 is a flowchart showing the operation flow of the communication device 300.
[0091]
As shown in FIG. 9, the receiving unit 24 of the calling device 300 of the user (any of B to E in FIG. 10) receives the voice signal and incidental information of the call from the calling device 100 of the other party (A in FIG. 10). Is received (S201).
[0092]
The receiving unit 24 transmits the voice signal of the call to the synthesis unit 13 of the voice processing device 3, and also transmits the incidental information (for example, the terminal ID or the telephone number) and the group designation information to the call partner identification unit 11. The group designation information is information that specifies which group of listeners should listen to the call. The group that allows listening to the call may be designated by the operation of the other party (A in FIG. 10) with respect to the calling device 100.
[0093]
The call partner identification unit 11 receives incidental information and group designation information from the reception unit 24. The call partner identification unit 11 determines whether or not the user (that is, the owner of the call device 300) belongs to the designated group based on the received group designation information (S202).
[0094]
If the user does not belong to the designated group (No in S202), the other party identification unit 11 instructs the output control unit 26 to mute the output (S208).
[0095]
In the example shown in FIG. 10, when the user belongs to the second group (D or E), the other party identification unit 11 transmits a mute instruction to the output control unit 26. On the other hand, in the example shown in FIG. 11, when the user belongs to the first group (B or C), the other party identification unit 11 transmits a mute instruction to the output control unit 26.
[0096]
On the other hand, when the user belongs to the designated group (Yes in S202), the call partner identification unit 11 identifies the call partner based on the incidental information (S203). The other party identification unit 11 transmits the identification result of the other party to the group determination unit 14.
[0097]
In the example shown in FIG. 10, when the user belongs to the first group (B or C), the other party identification unit 11 transmits the identification result of the other party to the group determination unit 14. On the other hand, in the example shown in FIG. 11, when the user belongs to the second group (D or E), the other party identification unit 11 transmits the identification result of the other party to the group determination unit 14.
[0098]
Specifically, when the call partner identification unit 11 succeeds in identifying the call partner, the call partner identification unit 11 transmits the caller ID that identifies the call partner to the group determination unit 14. On the other hand, when the call partner identification unit 11 fails to identify the call partner, the call partner identification unit 11 transmits an empty caller ID (None) to the group determination unit 14.
[0099]
The group determination unit 14 receives the identification result of the other party from the other party identification unit 11. The other party identification unit 11 determines the group to which the other party identified by the other party identification unit 11 belongs. The group determination unit 14 is an example of group determination means. The group determination unit 14 transmits the determination result of the group to which the call partner belongs to the background sound selection unit 12 together with the identification result of the call partner.
The background sound selection unit 12 receives the identification result of the other party and the determination result of the group to which the other party belongs from the group determination unit 14. The background sound selection unit 12 determines whether or not the other party has been registered in the setting list (S204). Specifically, the background sound selection unit 12 determines from the call partner identification unit 11 whether or not the caller ID received as the identification result of the call partner is included in the setting list.
[0100]
When the other party is already registered in the setting list (Yes in S204), the background sound selection unit 12 selects the background sound (here, BGM) according to the other party (and the group to which the other party belongs) (S205A). On the other hand, when the other party is not registered in the setting list (No in S204), the background sound selection unit 12 randomly selects the background sound (here, BGM) by referring to the BGM library (S205B).
[0101]
The background sound selection unit 12 transmits the selected BGM data (background sound information) to the synthesis unit 13 together with the sound image localization information corresponding to the caller ID of the other party.
[0102]
The synthesis unit 13 receives background sound information selected according to the other party from the background sound selection unit 12. Further, the synthesis unit 13 receives the voice signal of the call from the reception unit 24.
[0103]
The synthesizing unit 13 synthesizes the voice signal of the received call and the BGM which is the background sound (S206). As a result, the synthesis unit 13 generates an audio signal with background sound.
[0104]
The synthesis unit 13 transmits the generated audio signal with background sound to the output control unit 26.
[0105]
The output control unit 26 receives an audio signal with background sound from the synthesis unit 13. The output control unit 26 outputs the received audio signal with background sound from a speaker or the like (not shown) (S207).
[0106]
With the above, the operation of the calling device 300 is completed.
[0107]
(Effect of this embodiment)
According to the configuration of the present embodiment, the call partner identification unit 11 identifies the call partner. The background sound selection unit 12 selects the background sound according to the identified call partner. The synthesizing unit 13 synthesizes the voice signal of the call and the selected background sound. In this way, the voice signal obtained by synthesizing the voice signal of the call and the background sound is output. Therefore, the user can easily identify the other party only by hearing.
[0108]
Further, according to the configuration of the present embodiment, the group determination unit 14 determines the group to which the other party belongs. The background sound selection unit 12 preferably selects BGM corresponding to the group to which the other party belongs as the background sound.
[0109]
This allows the user to easily identify the group to which the other party belongs based on the characteristics of the BGM (for example, pitch, strength, timbre).
[0110]
[Embodiment 4]
In the fourth embodiment, the speaker-side telephone device (the telephone device 100 described in the first embodiment) is a hearable device worn on the ear. Hearable devices are a type of wearable device that applies otoacoustic emission authentication technology. An example of a hearable device is described in Patent Document 3.
[0111]
In the fourth embodiment, the communication device 100 includes a speaker that emits an acoustic signal (inspection signal) toward the inside of the speaker's ear, and a microphone that receives a reverberation signal from the inside of the ear. The call device 100 transmits the echo signal received by the microphone as sensing information to the call device on the listener side (the call devices 200 and 300 described in the second and third embodiments).
[0112]
In the fourth embodiment, the call partner identification unit 11 of the call devices 200 and 300 identifies the call partner based on the sensing information from the call device 100. Specifically, the call party identification unit 11 uses a characteristic echo (acoustic characteristic) determined by the shape of the ear hole of the speaker as sensing information. The background sound selection unit 12 selects the BGM based on the identification result of the other party based on the sensing information. The shape of the ear canal is, for example, a shape extending from the ear canal to the eardrum, or a shape extending to the middle ear or the inner ear.
[0113]
(Modification example)
In one modification, the calling device 100 is a wearable device other than a hearable device. For example, the communication device 100 according to this modification is a glass (glasses) type or wristband type wearable device.
[0114]
(Effect of this embodiment)
According to the configuration of the present embodiment, the call partner identification unit 11 identifies the call partner. The background sound selection unit 12 selects the background sound according to the identified call partner. The synthesizing unit 13 synthesizes the voice signal of the call and the selected background sound. In this way, the voice signal obtained by synthesizing the voice signal of the call and the background sound is output. Therefore, the user can easily identify the other party only by hearing.
[0115]
Further, according to the configuration of the present embodiment, the calling device on the speaker side is a wearable device such as a hearable device. Therefore, the call device on the listener side can identify the call partner (speaker) based on the sensing information acquired from the speaker's body, and can select the background sound corresponding to the identified call partner.
[0116]
[Embodiment 5]
The fifth embodiment will be described below with reference to FIG.
[0117]
(About hardware configuration)
Each component of the voice processing apparatus described in the first to fourth embodiments shows a block of functional units. Some of these components or All are realized by, for example, the information processing apparatus 900 as shown in FIG. FIG. 12 is a block diagram showing an example of the hardware configuration of the information processing apparatus 900.
[0118]
As shown in FIG. 12, the information processing apparatus 900 includes the following configuration as an example.
[0119]
・ CPU (Central Processing Unit) 901
ROM (Read Only Memory) 902
・ RAM (Random Access Memory) 903
・ Program 904 loaded into RAM903
・ Storage device 905 for storing program 904
・ Drive device 907 that reads and writes the recording medium 906
・ Communication interface 908 that connects to the communication network 909
・ Input / output interface 910 for inputting / outputting data
・ Bus 911 connecting each component
Each component of the voice processing apparatus described in the first to fourth embodiments is realized by the CPU 901 reading and executing the program 904 that realizes these functions. The program 904 that realizes the functions of each component is stored in, for example, a storage device 905 or ROM 902 in advance, and the CPU 901 is loaded into the RAM 903 and executed as needed. The program 904 may be supplied to the CPU 901 via the communication network 909, or may be stored in the recording medium 906 in advance, and the drive device 907 may read the program and supply the program to the CPU 901.
[0120]
(Effect of this embodiment)
According to the configuration of the present embodiment, the voice processing device described in the above embodiment is realized as hardware. Therefore, it is possible to obtain the same effect as the effect described in the above embodiment.
[0121]
Although the invention of the present application has been described above with reference to the embodiment, the invention of the present application is not limited to the above embodiment. Various changes that can be understood by those skilled in the art can be made within the scope of the present invention in terms of the configuration and details of the present invention.
[0122]
This application claims priority based on Japanese application Japanese Patent Application No. 2019-135799 filed on July 24, 2019, and incorporates all of its disclosures herein.
Industrial applicability
[0123]
The present invention can be used, for example, in a portable game machine and a VR (Virtual Realty) device.
Code description
[0124]
1 Voice processing device
11 Caller identification unit
12 Background sound selection section
13 Synthetic section
14 Group judgment unit
24 receiver
26 Output control unit
100 Calling device on the speaker side
200 Calling device on the listener side
300 Calling device on the listener side
The scope of the claims
[Claim 1]
The other party identification means for identifying the other party and the other party identification means
Background sound selection means for selecting the background sound according to the identified call partner, and
With a synthesis means that synthesizes the selected background sound with the voice signal of the call
With
Voice processing device.
[Claim 2]
The call partner identification means receives group-designated information that specifies which group the listener belongs to to listen to the call, and outputs an output control means that outputs an audio signal based on the received group-designated information. Mute
The voice processing device according to claim 1, characterized in that.
[Claim 3]
Further equipped with a group determination means for determining the group to which the identified call partner belongs,
The background sound selection means selects the background sound according to the determination result of the group to which the call partner belongs.
The voice processing device according to claim 1 or 2, characterized in that.
[Claim 4]
The background sound selection means defines a virtual position for localizing the sound image of the voice signal of the call according to the identified call partner.
The voice processing device according to any one of claims 1 to 3, characterized in that.
[Claim 5]
The background sound is one of BGM (Back Ground Music), environmental sound, and sound effect.
The voice processing device according to any one of claims 1 to 4, characterized in that.
[Claim 6]
The call partner identification means identifies the call partner based on the sensing information acquired from the body of the call partner.
The voice processing device according to any one of claims 1 to 5, characterized in that.
[Claim 7]
The voice processing device according to any one of claims 1 to 6 and
The receiving means for receiving the voice signal of the call and
Provided with an output control means for outputting a voice signal synthesized by the synthesis means of the voice processing device.
Calling device.
[Claim 8]
Identify the other party to call and
Select the background sound according to the identified call partner,
Synthesize the selected background sound with the voice signal of the call
Voice processing method.
[Claim 9]
Identifying the other party and
Selecting the background sound according to the identified call partner and
Combining the selected background sound with the voice signal of the call
A computer-readable recording medium that records a program that causes a computer to execute.

Documents

Application Documents

# Name Date
1 202217003553.pdf 2022-01-21
2 202217003553-TRANSLATIOIN OF PRIOIRTY DOCUMENTS ETC. [21-01-2022(online)].pdf 2022-01-21
3 202217003553-STATEMENT OF UNDERTAKING (FORM 3) [21-01-2022(online)].pdf 2022-01-21
4 202217003553-REQUEST FOR EXAMINATION (FORM-18) [21-01-2022(online)].pdf 2022-01-21
5 202217003553-PRIORITY DOCUMENTS [21-01-2022(online)].pdf 2022-01-21
6 202217003553-POWER OF AUTHORITY [21-01-2022(online)].pdf 2022-01-21
7 202217003553-NOTIFICATION OF INT. APPLN. NO. & FILING DATE (PCT-RO-105-PCT Pamphlet) [21-01-2022(online)].pdf 2022-01-21
8 202217003553-FORM 18 [21-01-2022(online)].pdf 2022-01-21
9 202217003553-FORM 1 [21-01-2022(online)].pdf 2022-01-21
10 202217003553-DRAWINGS [21-01-2022(online)].pdf 2022-01-21
11 202217003553-DECLARATION OF INVENTORSHIP (FORM 5) [21-01-2022(online)].pdf 2022-01-21
12 202217003553-COMPLETE SPECIFICATION [21-01-2022(online)].pdf 2022-01-21
13 202217003553-CLAIMS UNDER RULE 1 (PROVISIO) OF RULE 20 [21-01-2022(online)].pdf 2022-01-21
14 202217003553-Proof of Right [22-04-2022(online)].pdf 2022-04-22
15 202217003553-FER.pdf 2022-05-10
16 202217003553-PETITION UNDER RULE 137 [10-11-2022(online)].pdf 2022-11-10
17 202217003553-OTHERS [10-11-2022(online)].pdf 2022-11-10
18 202217003553-MARKED COPIES OF AMENDEMENTS [10-11-2022(online)].pdf 2022-11-10
19 202217003553-Information under section 8(2) [10-11-2022(online)].pdf 2022-11-10
20 202217003553-FORM 3 [10-11-2022(online)].pdf 2022-11-10
21 202217003553-FORM 13 [10-11-2022(online)].pdf 2022-11-10
22 202217003553-FER_SER_REPLY [10-11-2022(online)].pdf 2022-11-10
23 202217003553-DRAWING [10-11-2022(online)].pdf 2022-11-10
24 202217003553-COMPLETE SPECIFICATION [10-11-2022(online)].pdf 2022-11-10
25 202217003553-CLAIMS [10-11-2022(online)].pdf 2022-11-10
26 202217003553-AMMENDED DOCUMENTS [10-11-2022(online)].pdf 2022-11-10
27 202217003553-ABSTRACT [10-11-2022(online)].pdf 2022-11-10
28 202217003553-US(14)-HearingNotice-(HearingDate-10-10-2024).pdf 2024-09-23
29 202217003553-Correspondence to notify the Controller [07-10-2024(online)].pdf 2024-10-07

Search Strategy

1 SearchHistoryE_09-05-2022.pdf