Abstract: The present invention implements audio signal processing capable of extracting the voice of each speaker even when a plurality of speakers speaks at the same time. An audio signal processing device 400 is provided with: a determination unit 401 that determines a first voice segment for a target speaker linked to a host device on the basis of an externally acquired first audio signal; a sharing unit 402 that transmits the first audio signal and the first voice segment to another device linked to a non-target speaker and receives a second audio signal and a second voice segment associated with the non-target speaker from the other device; an estimation unit 403 that estimates the voice of the non-target speaker mixed in the first audio signal on the basis of the second audio signal and the second voice segment that are received and an estimation parameter associated with the target speaker that is acquired; and a removal unit 404 that removes the voice of the non-target speaker from the first audio signal to thus create a first voice from which the voice of the non-target speaker is removed.
Title of the invention: sound signal processing device, sound signal processing method and storage medium
Technical field
[0001]
The present invention relates to a sound signal processing device for emphasizing the voice of a specific speaker in the presence of a plurality of speakers.
Background technology
[0002]
Voice is a natural means of communication for humans, and communication not only between people in the same place but also with people in different places is realized through voice as a medium using telephones, web conferencing systems, and the like. In addition, it is becoming possible for systems to understand human voices using voice recognition technology, and voice communication has been realized not only between humans but also between humans and systems.
[0003]
In such communication using voice, a technology has been developed that emphasizes the voice of a specific speaker to make it easier to hear while multiple speakers coexist. This technology can be used in various scenes. For example, in a web conferencing system, the voice of a speaker who is mainly speaking is emphasized to reduce the influence of ambient noise, thereby making it easier to hear the speaker's speech. Further, in the voice recognition system, high-precision voice recognition can be realized by inputting voices separated for each speaker instead of inputting mixed voices. The techniques for emphasizing the voice of a specific speaker include the following.
[0004]
Patent Document 1 discloses a technique of performing sound source localization that estimates the speaker's direction using a plurality of microphones and emphasizing (beam forming processing) the sound coming from the speaker's direction estimated by the sound source localization.
[0005]
In Patent Document 2, an ad hoc network is formed between a plurality of terminals provided with a microphone, sound signals recorded by the plurality of terminals are transmitted and received to each other and shared, and a time lag of the voice recorded by each terminal is recorded. We disclose a technology that emphasizes only the voice of a specific speaker from multiple sound signals by correcting and adding.
In addition, Patent Document 3 discloses a technique for determining a voice section, which is related to the technique.
Prior art literature
Patent documents
[0006]
Patent Document 1: Japanese Patent Application Laid-Open No. 2002-091469
Patent Document 2: Japanese Unexamined Patent Publication No. 2011-254464
Patent Document 3: Japanese Patent No. 5299436
Outline of the invention
Problems to be solved by the invention
[0007]
Since the voice is attenuated as the distance increases, it is desirable that the distance between the speaker's mouth that emits the voice and the microphone that receives it is as close as possible. In particular, it is known that the higher the frequency, the faster the attenuation, and the longer the distance, the more easily the voice is affected by ambient noise, and the frequency characteristics of the voice change.
[0008]
In Patent Document 1, voice is emphasized by using a plurality of microphones (for example, a microphone array device) having a fixed position. However, the microphone cannot be brought close to each speaker and is affected by ambient noise.
[0009]
In Patent Document 2, since an independent terminal equipped with a microphone forms an ad hoc network, it is possible to bring the microphone closer to each speaker. However, in the method disclosed in Patent Document 2, when a plurality of speakers speak at the same time or speak without sufficient space between conversations, the voice of the speaker to be emphasized is mixed with the voice of another speaker. It becomes difficult to separate voices for each person.
[0010]
The present disclosure has been made in view of the above-mentioned problems, and a sound signal processing device or the like capable of extracting the voice of a target speaker even in a situation where a plurality of speakers speak at the same time is provided. One of the purposes is to provide.
Means to solve the problem
[0011]
In view of the above problems, the sound signal processing device, which is the first aspect of the present disclosure, is
Judgment means for determining the first voice section of the target speaker associated with the own device based on the first sound signal acquired from the outside,
The first sound signal and the first voice section are transmitted to another device associated with the non-target speaker, and the second sound signal and the second voice related to the non-target speaker are transmitted from the other device. Sharing means to receive the section and
The non-target speaker mixed in the first sound signal based on the received second sound signal and the second voice section and the acquired estimated parameters related to the target speaker. An estimation method for estimating voice, and
With a removal means that removes the voice of the non-target speaker from the first sound signal and generates the voice after the first non-target removal.
To prepare for.
[0012]
The sound signal processing method, which is the second aspect of the present disclosure, is
Based on the first sound signal acquired from the outside, the first voice section of the target speaker associated with the own device is determined.
The first sound signal and the first voice section are transmitted to another device associated with the non-target speaker, and the second sound signal and the second voice related to the non-target speaker are transmitted from the other device. Receive the section,
The non-target speaker mixed in the first sound signal based on the received second sound signal and the second voice section and the acquired estimated parameters related to the target speaker. Estimate the voice,
The voice of the non-target speaker is removed from the first sound signal to generate the voice after the first non-target removal.
Prepare for that.
[0013]
The sound signal processing program, which is the third aspect of the present disclosure, is
Based on the first sound signal acquired from the outside, the first voice section of the target speaker associated with the own device is determined.
The first sound signal and the first voice section are transmitted to another device associated with the non-target speaker, and the second sound signal and the second voice related to the non-target speaker are transmitted from the other device. Receive the section,
The non-target speaker mixed in the first sound signal based on the received second sound signal and the second voice section and the acquired estimated parameters related to the target speaker. Estimate the voice,
The voice of the non-target speaker is removed from the first sound signal to generate the voice after the first non-target removal.
Make that happen to your computer.
[0014]
The sound signal processing program may be stored in a non-temporary storage medium.
Effect of the invention
[0015]
According to the present disclosure, it is possible to provide a sound signal processing device or the like that can extract the voice of a target speaker even in a situation where a plurality of speakers speak at the same time.
A brief description of the drawing
[0016]
FIG. 1 is a block diagram showing a configuration example of a sound signal processing device according to the first embodiment of the present disclosure.
FIG. 2 is a flowchart showing an operation example of the sound signal processing apparatus according to the first embodiment.
FIG. 3 is a diagram showing details of an operation of non-target voice estimation by the sound signal processing device according to the first embodiment.
FIG. 4 is a diagram showing details of an operation of non-target voice estimation by the sound signal processing device according to the first embodiment.
[Fig. 5] Fig. 5 is a schematic diagram illustrating an implementation status of a sound signal processing device.
FIG. 6 is a schematic diagram illustrating a technique according to Patent Document 2.
FIG. 7 is a schematic diagram illustrating a technique relating to a sound signal processing apparatus according to the first embodiment.
FIG. 8 is a block diagram showing a configuration example of a sound signal processing device according to a second embodiment of the present disclosure.
FIG. 9 is a flowchart showing an operation example of the sound signal processing apparatus according to the second embodiment.
FIG. 10 is a diagram showing details of the operation of the sound signal processing device according to the second embodiment.
FIG. 11 is a diagram showing details of the operation of the sound signal processing apparatus according to the second embodiment.
FIG. 12 is a block diagram showing a configuration example of a sound signal processing device according to a third embodiment.
FIG. 13 is a flowchart showing an example of the operation of the sound signal processing device according to the third embodiment.
FIG. 14 is a block diagram showing a configuration example of a sound signal processing device according to a fourth embodiment.
FIG. 15 is a block diagram showing a configuration example of an information processing apparatus applicable to each embodiment.
Embodiment for carrying out the invention
[0017]
Hereinafter, each embodiment will be described in detail with reference to the drawings. In the description of the following drawings, the same or similar parts are designated by the same or similar reference numerals. However, the drawings schematically show the configuration in the embodiment of the present invention. Further, the embodiment of the present invention described below is an example, and can be appropriately changed as long as the essence thereof is the same.
[0018]
(Sound signal processing device)
Hereinafter, the first embodiment of the present invention will be described with reference to the drawings. FIG. 1 is a block diagram showing a configuration example of the sound signal processing device 100 according to the first embodiment. A plurality of sound signal processing devices 100 may exist, and in the present embodiment, they are referred to as a plurality of signal processing devices 100 and 100a. The plurality of signal processing devices 100 and 100a are the same device and have the same internal configuration. Each sound signal processing device 100 is associated with each of the target speakers. Each of the plurality of speakers may own one sound signal processing device 100. The sound signal processing device 100 may be built in a terminal owned by the user.
[0019]
The sound signal processing device 100 includes a sound signal acquisition unit 101, a voice section determination unit 102, a sound signal and voice section sharing unit 103, a non-target voice estimation unit 104, an estimation parameter storage unit 105, and a non-target voice removal unit. A unit 106 is provided.
[0020]
The estimation parameter storage unit 105 stores in advance the estimation parameters related to the target speaker. The details of the estimation parameters will be described later.
[0021]
The sound signal acquisition unit 101 acquires an ambient sound signal using a microphone. One microphone may be provided per device, or a plurality of microphones may be provided. The sound signal acquisition unit 101 mainly acquires the voice of the speaker who owns the sound signal processing device 100, but the voice of another speaker or ambient noise may be mixed. The sound signal is time-series information, and the sound signal acquisition unit 101 converts the sound signal obtained by the microphone from analog data to digital data, for example, to PCM (Pulse Code Modulation) data having a sampling frequency of 48 kHz and 16 bits. Convert and get. The sound signal acquisition unit 101 transmits the acquired sound signal to the voice section determination unit 102, the sound signal and the voice section sharing unit 103, and the non-target voice removal unit 106.
[0022]
The voice section determination unit 102 determines the voice section (first voice section) of the target speaker associated with the own device based on the sound signal (first sound signal) acquired from the outside. Specifically, the voice section determination unit 102 cuts out a section uttered by the speaker possessing the sound signal processing device 100 from the sound signal acquired from the sound signal acquisition unit 101. For example, the voice section determination unit 102 extracts data from time-series digital data with a window width of 512 points and a shift width of 256 points at short intervals, obtains the sound pressure for each cut-out unit, and determines the sound pressure in advance. The presence or absence of voice is determined according to whether or not the set threshold is exceeded, and the section where the voice is continuous is determined as the voice section. In addition to this, existing methods such as a method using HMM (Hidden Markov Model) and a method using LSTM (Long short-term memory) can be used for determining the voice section. The audio section is, for example, the start time and end time of the speaker's speech in the time from the start to the end of the conference. The time from the start time to the end time of the speaker's speech may be added to this. Alternatively, the start time and end time of the speaker's speech may be expressed in standard time by using the time stamp function of the OS (Operation System) that acquires the standard time. The voice section determination unit 102 transmits the determined voice section to the sound signal and the voice section sharing unit 103.
[0023]
The sound signal and voice section sharing unit 103 transmits the sound signal (first sound signal) of the own device and the voice section (first voice section) of the own device to another device associated with a speaker other than the target. , Receives a sound signal (second sound signal) and a voice section (second voice section) related to a speaker other than the subject from the other device. Specifically, the sound signal and voice section sharing unit 103 communicates with the sound signal and voice section sharing unit 103a of the sound signal processing device 100a other than the own device, and transmits and receives the sound signal and the voice section to each other and shares them. .. Each of the sound signal and the voice section sharing unit 103 may be broadcast broadcast asynchronously, or there may be a sound signal processing device 100 that serves as one hub, and the information collected there may be distributed again. Also, all sound signal processing The device 100 may transmit the sound signal and the voice section to the server, and the plurality of sound signals and the voice section collected on the server side may be delivered to the sound signal processing device 100 again.
[0024]
The non-target voice estimation unit 104 acquires information on the sound signal (second sound signal) and voice section (second voice section) acquired by the other sound signal processing device 100a from the sound signal and the voice section sharing unit 103. .. The non-target voice estimation unit 104 acquires the estimation parameters stored in the estimation parameter storage unit 105. The estimation parameter is, for example, information on the arrival time (time lag) and the amount of attenuation until the voice acquired by the other sound signal processing device 100a arrives at the sound signal processing device 100, which is the own device. The non-target voice estimation unit 104 estimates whether the sound signal and the voice section of the other sound signal processing device 100a are non-target voices by using the estimation parameters. That is, the non-target voice estimation unit 104 estimates whether the sound acquired by the other sound signal processing device 100a is a sound signal mixed with the sound acquired by the sound signal acquisition unit 101. The non-target voice estimation unit 104 transmits the estimated non-target voice (mixed sound signal) to the non-target voice removal unit 106. As a result of estimation, the non-target voice estimation unit 104 may determine whether the sound acquired by the other sound signal processing device 100a matches the sound signal mixed in the sound acquired from the sound signal acquisition unit 101. good. In this embodiment, since it is premised that the speakers a to c are specified as shown in FIG. 5, the mixed voice can be easily predicted from the estimation result.
[0025]
The non-target voice removing unit 106 removes the voice of a speaker who is not the target from the sound signal (first sound signal) acquired by the own device, and generates the voice after the non-target removal (the first voice after the removal of the target). .. Specifically, the non-target voice removal unit 106 acquires the non-target voice estimated from the non-target voice estimation unit 104. The non-target voice removing unit 106 removes the estimated non-target voice from the voice acquired by the sound signal acquisition unit 101. At the time of removal, for example, a spectrum subtraction method in which FFT (Fast Fourier Transform) is performed for a short time, divided into frequency bands in the spectrum region, and subtraction is performed, or a winner that calculates and multiplies the gain for noise suppression. Use an existing method such as the filter method.
[0026]
(Operation of sound signal processing device)
Next, the operation of the sound signal processing devices 100 and 100a in the first embodiment will be described with reference to the flowchart of FIG. Since the sound signal processing devices 100 and 100a execute the same operation with the same configuration, the processing contents of steps S101 to S105 and steps S111 to S115 are the same. Further, in the following description, the sound signal processing devices 100 and 100a will be described as being mounted on the terminal A and the terminal B such as the portable communication terminal held by the speaker, respectively. In the following description, the terminal A may be described as the own terminal A owned by itself, and the terminal B may be described as another terminal B owned by another speaker.
[0027]
First, the sound signal acquisition unit 101 acquires a sound signal using a microphone or the like (step S101). In the subsequent processing, the time series of the sound signal may be cut out at short intervals with, for example, a window width of 512 points and a shift width of 256 points, and the processing after step S102 may be performed. Alternatively, the processing after step S102 may be sequentially performed on the time series of the sound signal every second or the like.
[0028]
Here, n is represented by the sample point (time) of the digital signal, and the sound signal acquired by the terminal A is represented by y_A (n). y_A (n) mainly includes the voice signal x_A (n) of the speaker associated with the terminal A, and the voice signal x_B (n)'of the speaker who is not the target is mixed. Using the following procedure, x_B (n)'is estimated and removed to extract only x_A (n). The same process is performed on the terminal B, and only the speaker's voice x_B (n) associated with the terminal B is extracted.
[0029]
Next, the voice section determination unit 102 cuts out only the section uttered by the speaker holding the terminal A from the acquired sound signal (step S102). FIG. 3 is a schematic view showing the processing of steps S102 and S103 (steps S112 and S113). A specific example of voice section determination by terminal A and terminal B is shown in the upper part of FIG. The terminal A is associated so that the speaker a is the target speaker, the terminal B is associated with the speaker b as the target speaker, the terminal A is the voice section of the speaker a, and the terminal B is the speaker. Determine the voice section of person b. At this time, for example, a section in which the volume is larger than the threshold value is determined to be an audio section, and is represented as a rectangular having a long vertical width as shown in FIG. At this time, the width of the rectangle indicates the length of the utterance. In the upper part of FIG. 3, the voice section of speaker a is clear. However, in reality, the volume of the voice changes from moment to moment depending on the type of sound element, etc., and if it is uniquely judged only by comparing the magnitude with the threshold, it may contain an error. Post-processing such as is required. Here, the voice section is expressed as VAD [y_A (n)]. If the sound signal y_A (n) at time n is voice, it is expressed as VAD [y_A (n)] = 1, and if it is non-voice, it is expressed as VAD [y_A (n)] = 0.
[0030]
Next, the sound signal and voice section sharing unit 103 transmits the acquired sound signal and voice section to another terminal B in the vicinity, and receives the sound signal and voice section acquired by the other terminal B to the own terminal A. As a result, the sound signal and the voice section are shared (step S103). A specific example of sharing the sound signal and the voice section is shown in the lower part of FIG. The lower terminal A acquires the voice acquired by the terminal B and the spoken section of the speaker b in addition to the voice acquired by the own terminal and the spoken section of the speaker a. On the contrary, the terminal B acquires the voice acquired by the terminal A and the spoken section of the speaker a in addition to the voice acquired by the own terminal and the spoken section of the speaker b. The same applies when the number of terminals is large, and the number of shares increases according to the number of terminals. Here, the sound signal and voice section acquired by the terminal A are y_A (n), VAD [y_A (n)], and the sound signal and voice section acquired by the terminal B are y_B (n), VAD [y_B (n)]. It will be expressed as.
[0031]
Next, the non-target voice estimation unit 104 mixes the sound signal and voice section information acquired by the other terminal B and the parameters stored in the estimation parameter storage unit 105 into the voice acquired by the own terminal A. Estimate the non-target voice (step S104). FIG. 4 is a schematic view showing the processing of steps S104 and S105 (steps S114 and S115). A specific example of non-target voice estimation by the terminal A and the terminal B is shown in the upper part of FIG. In the estimation parameter storage unit 105, information on the arrival time (time lag) and the amount of attenuation until the voice acquired by the other terminal B arrives at the own terminal A is stored as the estimation parameter, and this information is used. Estimate the non-target voice mixed in the voice acquired by the own terminal A. For example, time lag and attenuation information can be retained in the form of impulse responses. The impulse response is a response to a pulse signal.
[0032]
In the estimation of the non-target audio signal in the terminal A (here, the audio signal of the terminal B mixed in the audio acquired by the terminal A), first, the shared audio signal y_b (n) of the terminal B and the audio section VAD [y_b From (n)], the effective audio signal y_b (n)'is calculated according to Equation 1.
y_b (n)'= y_b (n) ・ VAD [y_b (n)] ・ ・ ・ (Equation 1)
Here, ・ indicates the product. Perform the product every time n. Next, the non-target voice est_b (n) is estimated by convolving the impulse response h (m). Folding can be performed using Equation 2.
est_b (n) = Σ m h (m) ・ y_b (n-m)'・ ・ ・ (Equation 2)
Here, m represents a time lag. Referring to the upper left of FIG. 4, the voice signal of the own terminal A is mixed in the non-target voice signal estimated here, but even in such a case, the impulse response h (m) is 1. Since the value is smaller than the original signal, the value is sufficiently small compared to the original signal, so that the leakage of the target sound is sufficiently small.
[0033]
Similarly, regarding the non-target audio signal in the terminal B (here, the audio signal of the terminal A mixed in the audio acquired by the terminal B), first, the shared audio signal y_a (n) of the terminal A and the audio section VAD [ From y_a (n)], the effective voice signal y_a (n)'is calculated according to Equation 3.
y_a (n) ’= y_a (n) ・ VAD [y_a (n)] ・ ・ ・ (Equation 3)
Next, the non-target voice est_a (n) is estimated according to Equation 4.
est_a (n) = Σ m h (m) ・ y_a (n-m)'・ ・ ・ (Equation 4)
Next, the non-target sound removing unit 106 removes the estimated non-target sound from the sound acquired by the sound signal acquisition unit 101 (step S105). A specific example of estimating the non-target voice is shown in the lower part of FIG. By removing the non-target voice estimated from the sound signal acquired by the own terminal A, only the voice of the target speaker can be extracted. When the target voice is mixed with the non-target voice estimated as shown in the lower left of FIG. 4, distortion may occur due to excessive subtraction, but it is sufficiently small. This effect can be reduced by, for example, providing a flooring in the amount to be subtracted and not subtracting more than a certain value, or by adding a sufficiently small white noise to the value after subtraction and masking. be. Further, the winner filter method may be used. In this case, the minimum value of the gain is set, and the process is performed so that the suppression is not performed below the value.
[0034]
Here, as an example, a spectrum subtraction method in which FFT is performed for a short time, divided into frequency bands in the spectrum region, and subtraction is performed will be described. Y_a (i, ω) is the voice signal y_a (n) of the terminal A that has been subjected to a short-time FFT, and Est_b [i, ω] is the one that has been subjected to a short-time FFT to the non-target voice signal est_b (n). And. Here, i indicates the index of the short-time window, and ω indicates the index of the frequency. By removing the non-target voice signal est_b (n) from Y_a (i, ω), the speaker's voice X_a (i, ω) associated with the own terminal A is acquired according to Equation 5.
X_a (i, ω) = max [Y_a (i, ω) -Est_b (i, ω), floor] ・ ・ ・ (Equation 5)
Here, max [A, B] indicates an operation that takes the larger value of A and B. floor is the flooring of the amount to be subtracted, and indicates that no more than this value is subtracted.
[0035]
Here, the solution to the problem of Patent Document 2 made by the present invention will be shown. First, the problem of Patent Document 2 can be understood as follows.
[0036]
As shown in FIG. 5, a case where three speakers a, b, and c each own terminals A, B, and C equipped with microphones will be described. FIG. 6 shows the voice extraction process for each speaker in Patent Document 2. As shown in FIG. 6, two speakers, speaker a and speaker b, are speaking at almost no time. In this situation, the voice of the speaker a is recorded on the terminal A larger than that of the other terminals, and then the voice of the speaker b is recorded. The voice of the speaker b is recorded on the terminal B larger than that of other terminals, and then the voice of the speaker a is recorded. Each voice is recorded on the terminal C. As described above, depending on the timing of the voices of the two persons, there may be a terminal that cannot be separated in time and is recorded. In such a situation, if the voices of the speaker a are simply staggered in order to emphasize the voices of the speaker a, the voices of the speaker b are mixed, and the expected effect cannot be obtained.
[0037]
Next, the voice extraction process for each speaker in the first embodiment of the present invention in the situation shown in FIG. 5 will be described with reference to FIG. 7. In the sound signal processing device 100 of the first embodiment, the sound signal obtained from the terminal B is mixed with the sound of the speaker b who is a non-target speaker, instead of emphasizing the sound of the speaker a in the terminal A. And estimate and remove using the information of the voice section. By doing so, it is possible to extract the voices of individual speakers even in a situation where a plurality of speakers are speaking in a short time.
[0038]
Also, here we have shown the separation of the voices of two speakers, but it is the same even if there are three or more speakers.It is possible to extract only the voice of the speaker associated with each terminal by estimating and subtracting multiple non-target voices by taking the same procedure.
[0039]
This completes the description of the operation of the sound signal processing devices 100 and 100a.
[0040]
(Effect of the first embodiment)
According to the sound signal processing device 100 of the present embodiment, the voice of the target speaker can be extracted even in a situation where a plurality of speakers speak at the same time. This is because the sound signal and voice section sharing unit 103 included in the own terminal A and the other terminal B transmit and receive and share the sound signal and the voice section with each other. Further, the non-target voice estimation unit 104 estimates the non-target voice mixed in the voice acquired by the own terminal A by using the sound signal shared with each other and the information of the voice section, and by extension, the estimated non-target voice. Is removed from the target voice and the target voice is emphasized.
[0041]
(Sound signal processing device)
In step S105 described above, when the target voice is mixed with the non-target voice estimated as shown in the lower left of FIG. 4, a smaller distortion may occur in the excessive subtraction and noise may be included. In the second embodiment of the present disclosure, a sound signal processing device that suppresses the generation of this distortion will be described.
[0042]
FIG. 8 is a block diagram showing a configuration example of the sound signal processing device 200 according to the second embodiment. The sound signal processing device 200 includes a sound signal acquisition unit 101, a voice section determination unit 102, a sound signal and a voice section sharing unit 103, a non-target voice estimation unit 104, an estimation parameter storage unit 105, a non-target voice removal unit 106, and a non-target. A post-removal voice sharing unit 201, a second non-target voice estimation unit 202, and a second non-target voice removal unit 203 are provided.
[0043]
The non-target removal voice sharing unit 201 shares the voice after removal of the non-target voice with the voice sharing unit 201a after removal of the target of another sound signal processing device 200a as the first voice after removal of the non-target. The non-target removal voice sharing unit 201 transmits the non-target removal voice (first non-target removal voice) to the other sound signal processing device 200a, and the non-target removal voice (the non-target removal voice) of the other sound signal processing device 200a. The second sound after removal from the target) is received from another sound signal processing device 200a. The non-target removal voice sharing unit 201 transmits the received non-target removal voice to the second non-target voice estimation unit 202.
[0044]
The second non-target voice estimation unit 202 estimates the voice of the non-target speaker based on the non-target removal voice (second non-target removal voice) received from another device and the estimation parameters of the own device. .. Specifically, the second non-target voice estimation unit 202 receives the non-target removal voice (second non-target removal voice) of the other sound signal processing device 200a from the non-target removal voice sharing unit 201. Estimated parameters are acquired from the estimated parameter storage unit 105. The second non-target voice estimation unit 202 estimates the second non-target voice by adjusting the time shift and the attenuation amount of the speech section based on the estimation parameter for the received voice after removal of the target. The second non-target voice estimation unit 202 transmits the estimated second non-target voice to the second non-target voice removal unit 203.
[0045]
When the second non-target sound removal unit 203 acquires the second non-target sound estimated from the second non-target sound estimation unit 202, the second non-target sound removal unit 203 is estimated from the sound acquired by the sound signal acquisition unit 101. Remove out-of-target audio.
[0046]
Other parts are the same as those of the first embodiment shown in FIG.
[0047]
(Sound signal processing method)
An example of the operation of the sound signal processing devices 200 and 200a of the present embodiment will be described with reference to the flowchart shown in FIG.
[0048]
First, steps S101 to S105 (steps S111 to S115) in FIG. 9 are the same as the steps of the first embodiment shown in FIG.
[0049]
Next, the non-target removal voice sharing unit 201 of the own terminal A shares the voice after removal of the non-target voice obtained in step S105 with the other terminal B as the first voice after removal of the non-target (step S201). .. FIG. 10 is a diagram showing the processing of steps S201 and S202 (steps S211 and S212). The upper part of FIG. 10 shows a specific example of sharing the voice after the first non-target removal of the terminal A and the terminal B.
[0050]
Next, the second non-target voice estimation unit 202 estimates the second non-target voice by adjusting the time shift and the attenuation amount with respect to the first non-target removed voice received from the other terminal B. (Step S202). A specific example of the second non-target voice estimation of the terminal A and the terminal B is shown in the lower part of FIG. As an estimation parameter, the estimation parameter storage unit 105 stores information on the arrival time and the amount of attenuation until the voice acquired by the other terminal B arrives at the own terminal A, and the own terminal A uses this information. Estimate the non-target voice mixed in the acquired voice. By estimating the non-target voice mixed in the voice acquired by the own terminal A using the voice after the first non-target removal, the influence of distortion is further smaller than that of the first non-target voice estimation unit 104. Become. This is because the distortion caused by the excessive subtraction is less affected by the time shift and the correction of the attenuation amount.
[0051]
Next, the second non-target sound removing unit 203 removes the estimated second non-target sound from the sound acquired by the sound signal acquisition unit 101 (step S203). FIG. 11 shows a specific example of the second non-target voice removal of the terminal A and the terminal B in step S203. By repeating the estimation process twice as shown in FIG. 11, the influence of strain can be reduced to zero, that is, noise can be removed.
[0052]
This completes the description of the operation of the sound signal processing devices 200 and 200a.
[0053]
(Effect of the second embodiment)
According to the sound signal processing device 200 of the present embodiment, the voice of the target speaker can be accurately extracted even in a situation where a plurality of speakers speak at the same time. In this method, in addition to the estimation by the non-target voice estimation unit 104 according to the first embodiment, the sound after removal from the target is shared with the other terminal B, and the second non-target voice estimation unit 202 is the target of the other terminal B. This is because the distortion (noise) is removed by estimating the second non-target voice by adjusting the time shift of the speech section and the attenuation amount of the voice after the external removal.
[0054]
(Sound signal processing device)
In the sound signal processing devices 100 and 200 according to the first and second embodiments, the estimation parameters stored in advance in the estimation parameter storage unit 105 were used. In the third embodiment of the present disclosure, a sound signal processing device that calculates an estimated parameter and stores it in the estimated parameter storage unit 105 will be described. The sound signal processing device according to the third embodiment is, for example, a scene in which an estimation parameter of a non-target voice is calculated at the beginning of a conference and the target voice is extracted during the conference using the estimation parameter. Can be used in.
[0055]
FIG. 12 is a block diagram showing a configuration example of the sound signal processing device 300. Hereinafter, for the sake of brevity, the sound signal processing apparatus 100 according to the first embodiment of FIG. 1 will be described with the addition of a parameter calculation unit 30 for calculating estimated parameters, but the second embodiment will be described. It can also be applied to the sound signal processing device 200.
[0056]
As shown in FIG. 12, the sound signal processing device 300 includes a sound signal acquisition unit 101, a voice section determination unit 102, a sound signal and voice section sharing unit 103, a non-target voice estimation unit 104, an estimation parameter storage unit 105, and a non-target. It includes a voice removing unit 106 and a parameter calculation unit 30. The parameter calculation unit 30 includes an inspection signal reproduction unit 301 and a non-target voice estimation parameter calculation unit 302.
[0057]
The inspection signal reproduction unit 301 reproduces the inspection signal. The inspection signal is an acoustic signal to be used for the estimation parameter calculation process, and may be reproduced in advance in a memory (not shown) or the like, or may be generated in real time. The accuracy of estimation is improved by reproducing the inspection signal from the same position as each speaker. The non-target voice estimation parameter calculation unit 302 receives the inspection signal reproduced by the inspection signal reproduction unit 301. An inspection microphone may be used for reception, or a microphone connected to the sound signal acquisition unit 101 may be used. The microphone is preferably placed near the position of each speaker. The non-target voice estimation parameter calculation unit 302 receives information that becomes an estimation parameter based on the received inspection signal, for example, until the sound acquired by another sound signal processing device 300a arrives at the sound signal processing device 300, which is its own device. Information on the arrival time (time lag) and the amount of attenuation is calculated. The calculated estimated parameter is stored in the estimated parameter storage unit 105.
[0058]
Other parts are the same as in the first embodiment.
[0059]
(Parameter calculation method)
FIG. 13 is a flowchart showing an example of estimation parameter calculation processing of the sound signal processing devices 300 and 300a. As with the sound signal processing device 100, a plurality of sound signal processing devices 300 may also exist, and the following description will be made assuming that the own terminal A includes the sound signal processing device 300 and the other terminal B includes the sound signal processing device 300a. conduct. In FIG. 13, steps S301 and S302 are the same as steps S311 and S312, and steps S101 to S103 are the same as steps S111 to S113.
[0060]
First, the inspection signal reproduction unit 301 reproduces the inspection signal (step S301). The inspection signal is a substitute for the voice of the speaker targeted by the terminal, and the inspection signal reproduction unit 301 reproduces the known signal at a known timing and length. This is to calculate the parameters that enable accurate non-target voice estimation. As the inspection signal, the acoustic signal used as standard for obtaining the impulse response is used. For example, it is conceivable to use an M-series signal, white noise, sweep signal, TSP (Time Stretched Pulse) signal, or the like. It is desirable that a plurality of terminals A and B reproduce known and unique signals. This is because by reproducing a known and unique signal, even if the inspection signal is reproduced at the same time, it can be separated.
[0061]
After that, as in the operation of the first embodiment, the sound signal is acquired (step S101), the voice section is determined (step S102), and the sound signal and the speech section are shared (step S103).
[0062]
Next, the non-target voice estimation parameter calculation unit 302 calculates the non-target voice estimation parameters (step S302). The parameters of non-target voice estimation include time lag and attenuation, and these two quantities can be obtained by calculating the impulse response. As a method for calculating the impulse response, an existing method such as a direct correlation method, a cross spectrum method, or an MLS (Maximum Length Sequence) method is used. Here, an example using the direct correlation method is shown. In the direct correlation method, in a function such as white noise in which the autocorrelation is a delta function, the calculation is performed using the fact that the correlation function is equivalent to the impulse response. Assuming that the time series of the inspection sound is x (n) and the sound signal acquired by a certain terminal is y (n), the mutual correlation function xcorr (m) can be calculated by the following equation 6.
xcorr (m) = (1 / N) ・ Σ n x (n) ・ y (n + m) ・ ・ ・ (Equation 6)
Here, n and m indicate the sample points (time) of the digital signal, and N indicates the number of sample points to be added. The intercorrelation function xcorr (m) represents the magnitude of decay at each time. When the mutual correlation function xcorr (m) is maximized, m represents the magnitude of the time lag. Equation 6 can be calculated for each combination of terminals A and B. Further, the larger the sample score N to be added, the more accurately the mutual correlation function can be obtained. The mutual correlation function can be regarded as the impulse response h (m).
[0063]
Further, it is conceivable not only to calculate the parameters for estimating the non-target voice, but also to calculate the parameters such as the threshold value related to the voice section determination in the voice section determination unit 102. As for the voice section determination unit, the method of the voice detection device described in Patent Document 3 may be used.
[0064]
This completes the description of the operation of the sound signal processing devices 300 and 300a.
[0065]
((Effect of the third embodiment)
According to the sound signal processing device 300 of the present embodiment, as in the first and second embodiments, even in a situation where a plurality of speakers speak at the same time, the voice of the target speaker can be heard. Can be extracted. Further, the sound signal processing device 300 can calculate the estimation parameter of the non-target voice at the beginning of a conference, for example, and extract the target voice during the conference using the calculated estimation parameter, whereby the accuracy is achieved in real time. Good voice extraction is possible.
[0066]
(Variation example)
In the first to third embodiments, it is premised that the parameter for estimating the non-target voice is calculated using the audible sound, but the parameter may be calculated using the non-audible sound. The inaudible sound is a sound signal that cannot be recognized by humans, and it is conceivable to use a sound signal of 18 kHz or more or 20 kHz or more. At the beginning of a meeting, etc., parameters for estimating non-target voice are calculated using both audible and non-audible sounds, and the relationship between the time lag and attenuation for audible sounds and the time lag and attenuation for non-audible sounds. The relationship between the time lag and the attenuation amount for the audible sound and the time lag and the attenuation amount for the non-audible sound is measured by using the non-audible sound during the meeting. From the nature, it is conceivable to predict the time lag and the amount of attenuation with respect to the audible sound and continue updating.
[0067]
For example, at the beginning of a meeting, when the time difference of the audible sound is 0.1 seconds and the attenuation is 0.5 until the test sound played from one terminal is measured by another terminal, it is inaudible. Suppose the time lag is 0.1 seconds, the attenuation is 0.4, the inaudible time lag during the meeting is 0.15 seconds, and the attenuation is 0.2. Since the time lag of audible sound is the same for audible sound and non-audible sound, it can be predicted that the time lag is 0.15 seconds, and the attenuation of audible sound is 5/4 times the attenuation of non-audible sound, so it can be predicted to be 0.25. Actually, since there is a range of frequencies for both audible and non-audible sounds, it is necessary to consider the relationship between multiple frequencies. It is possible to predict the time lag and the amount of attenuation with respect to the audible sound.
[0068]
FIG. 14 shows the sound signal processing device 400 according to the fourth embodiment. The sound signal processing device 400 represents the minimum necessary configuration for realizing the sound signal processing device according to the first to third embodiments. The sound signal processing device 400 has a determination unit 401 that determines a first voice section of a target speaker associated with the own device based on a first sound signal acquired from the outside, a first sound signal, and the first sound signal. 1 A shared unit 402 that transmits a voice section to another device associated with a non-target speaker and receives a second sound signal and a second voice section related to the non-target speaker from the other device. From the estimation unit 403 that estimates the voice of a non-target speaker mixed in the first sound signal based on the received second sound signal and the second voice section and the acquired estimation parameter, and the first sound signal. It is provided with a removing unit 404 that removes the voice of a speaker who is not the target and generates the voice after the first removal of the target.
[0069]
According to the sound signal processing device 400 according to the fourth embodiment, the voice of the target speaker can be extracted even in a situation where a plurality of speakers speak at the same time. This is because the shared unit 402 of the own terminal A and the other terminal B, which together include the sound signal processing device 400, transmits and receives the sound signal and the voice section to each other and shares them. Further, the estimation unit 403 estimates the non-target voice mixed in the voice acquired by the own terminal A by using the sound signal and the information of the voice section shared with each other, and the estimated non-target voice is the target voice. Because it is removed from.
[0070]
(Information processing device)
In each of the above-described embodiments of the present invention, the configuration of a part or all of the components in the sound signal processing apparatus shown in FIGS. 1, 8, 12, etc. is, for example, the information processing apparatus 500 as shown in FIG. It can also be realized by using any combination with the program. As an example, the information processing apparatus 500 includes the following configurations.
[0071]
・ CPU (Central Processing Unit) 501
・ ROM (Read Only Memory) 502
・ RAM (Random Access Memory) 503
・ Storage device 505 that stores program 504 and other data
・ Drive device 507 that reads and writes the recording medium 506
・ Communication interface 508 that connects to the communication network 509
・ Input / output interface 510 for input / output data
・ Bus 511 connecting each component
Each component of the sound signal processing device in each embodiment of the present application is realized by the CPU 501 acquiring and executing the program 504 that realizes these functions. The program 504 that realizes the functions of each component of the sound signal processing device is, for example, stored in the storage device 505 or the RAM 503 in advance, and is read by the CPU 501 as needed. The program 504 may be supplied to the CPU 501 via the communication network 509, or may be stored in the recording medium 506 in advance, and the drive device 507 may read the program and supply the program to the CPU 501. The drive device 507 may be externally attachable to each device.
[0072]
There are various modifications in the method of realizing each device. For example, the sound signal processing device may be realized by any combination of an information processing device and a program that are separate for each component. Further, a plurality of components included in the sound signal processing device may be realized by any combination of one information processing device 500 and a program.
[0073]
Further, a part or all of each component of the sound signal processing device is realized by other general-purpose or dedicated circuits, processors, etc. or a combination thereof. These may be composed of a single chip or may be composed of a plurality of chips connected via a bus.
[0074]
A part or all of each component of the sound signal processing device may be realized by a combination of the above-mentioned circuit or the like and a program.
[0075]
When a part or all of each component of the sound signal processing device is realized by a plurality of information processing devices, circuits, etc., the plurality of information processing devices, circuits, etc. may be centrally arranged or distributed. May be done. For example, the information processing device, the circuit, and the like may be realized as a form in which each is connected via a communication network, such as a client-and-server system and a cloud computing system.
[0076]
Although the present invention has been described above with reference to the present embodiment, the present invention is not limited to the above embodiment. Various changes that can be understood by those skilled in the art can be made within the scope of the present invention in terms of the structure and details of the present invention.
Description of the sign
[0077]
100 Sound signal processing device
100a Sound signal processing device
101 Sound signal acquisition unit
102 Voice section judgment unit
103 Voice section sharing section
103a Audio section sharing section
104 Non-target voice estimation unit
105 Estimated parameter storage unit
106 Excluded voice removal unit
200 Sound signal processing device
200a Sound signal processing device
201 Voice sharing section after removal of non-target
201a Voice sharing section after removal of non-target
202 Second non-target voice estimation unit
203 Second non-target voice removal unit
300 Sound signal processing device
300a sound signal processing device
301 Inspection signal reproduction unit
302 Non-target voice estimation parameter calculation unit
400 Sound signal processing device
401 Judgment unit
402 Shared section
403 Estimator
404 Removal part
500 Information processing device
504 program
505 storage device
506 Recording medium
507 drive device
508 Communication interface
509 Communication network
510 input / output interface
511 bus
The scope of the claims
[Claim 1]
Judgment means for determining the first voice section of the target speaker associated with the own device based on the first sound signal acquired from the outside,
The first sound signal and the first voice section are transmitted to another device associated with the non-target speaker, and the second sound signal and the second voice related to the non-target speaker are transmitted from the other device. Sharing means to receive the section and
The non-target speaker mixed in the first sound signal based on the received second sound signal and the second voice section and the acquired estimated parameters related to the target speaker. An estimation method for estimating voice, and
With a removal means that removes the voice of the non-target speaker from the first sound signal and generates the voice after the first non-target removal.
A sound signal processing device.
[Claim 2]
A second share in which the first non-target removal sound is transmitted to the other device, and the second non-target removal sound from which the target speaker's voice is removed from the second sound signal is received from the other device. Means and
A second estimation means for estimating the voice of the speaker who is not the target based on the received voice after removal of the target and the estimation parameter.
With the second removing means for removing the voice of the non-target speaker estimated by the second estimation means from the first sound signal.
The sound signal processing apparatus according to claim 1.
[Claim 3]
The estimated parameter includes at least one of the time lag and the amount of attenuation until the second sound signal reaches the own device.
The sound signal processing apparatus according to claim 1 or 2.
[Claim 4]
The time lag and the amount of attenuation are calculated based on the impulse response.
The sound signal processing device according to claim 3.
[Claim 5]
Inspection signal reproduction means for reproducing inspection signals and
An estimation parameter calculation means for calculating an estimation parameter for estimating the sound of the other device to be mixed from the inspection signal and the first sound signal.
The sound signal processing apparatus according to claim 1.
[Claim 6]
The estimation parameter calculation means uses an audible sound in the calculation of the estimation parameter.
The sound signal processing device according to claim 5.
[Claim 7]
The estimation parameter calculation means uses an inaudible sound in the calculation of the estimation parameter.
The sound signal processing device according to claim 5.
[Claim 8]
Based on the first sound signal acquired from the outside, the first voice section of the target speaker associated with the own device is determined.
The first sound signal and the first voice section are transmitted to another device associated with the non-target speaker, and the second sound signal and the second voice related to the non-target speaker are transmitted from the other device. Receive the section,
The non-target speaker mixed in the first sound signal based on the received second sound signal and the second voice section and the acquired estimated parameters related to the target speaker. Estimate the voice,
The voice of the non-target speaker is removed from the first sound signal to generate the voice after the first non-target removal.
A sound signal processing method that comprises.
[Claim 9]
Based on the first sound signal acquired from the outside, the first voice section of the target speaker associated with the own device is determined.
The first sound signal and the first voice section are transmitted to another device associated with the non-target speaker, and the second sound signal and the second voice related to the non-target speaker are transmitted from the other device. Receive the section,
The non-target speaker mixed in the first sound signal based on the received second sound signal and the second voice section and the acquired estimated parameters related to the target speaker. Estimate the voice,
The voice of the non-target speaker is removed from the first sound signal to generate the voice after the first non-target removal.
A storage medium that stores a sound signal processing program that makes a computer realize that.
| # | Name | Date |
|---|---|---|
| 1 | 202217016749.pdf | 2022-03-24 |
| 2 | 202217016749-TRANSLATIOIN OF PRIOIRTY DOCUMENTS ETC. [24-03-2022(online)].pdf | 2022-03-24 |
| 3 | 202217016749-STATEMENT OF UNDERTAKING (FORM 3) [24-03-2022(online)].pdf | 2022-03-24 |
| 4 | 202217016749-REQUEST FOR EXAMINATION (FORM-18) [24-03-2022(online)].pdf | 2022-03-24 |
| 5 | 202217016749-POWER OF AUTHORITY [24-03-2022(online)].pdf | 2022-03-24 |
| 6 | 202217016749-NOTIFICATION OF INT. APPLN. NO. & FILING DATE (PCT-RO-105-PCT Pamphlet) [24-03-2022(online)].pdf | 2022-03-24 |
| 7 | 202217016749-FORM 18 [24-03-2022(online)].pdf | 2022-03-24 |
| 8 | 202217016749-FORM 1 [24-03-2022(online)].pdf | 2022-03-24 |
| 9 | 202217016749-DRAWINGS [24-03-2022(online)].pdf | 2022-03-24 |
| 10 | 202217016749-DECLARATION OF INVENTORSHIP (FORM 5) [24-03-2022(online)].pdf | 2022-03-24 |
| 11 | 202217016749-COMPLETE SPECIFICATION [24-03-2022(online)].pdf | 2022-03-24 |
| 12 | 202217016749-Proof of Right [21-04-2022(online)].pdf | 2022-04-21 |
| 13 | 202217016749-Others-130522.pdf | 2022-05-17 |
| 14 | 202217016749-Correspondence-130522.pdf | 2022-05-17 |
| 15 | 202217016749-FER.pdf | 2022-07-28 |
| 16 | 202217016749-OTHERS [09-01-2023(online)].pdf | 2023-01-09 |
| 17 | 202217016749-Information under section 8(2) [09-01-2023(online)].pdf | 2023-01-09 |
| 18 | 202217016749-FORM-26 [09-01-2023(online)].pdf | 2023-01-09 |
| 19 | 202217016749-FORM 3 [09-01-2023(online)].pdf | 2023-01-09 |
| 20 | 202217016749-FER_SER_REPLY [09-01-2023(online)].pdf | 2023-01-09 |
| 21 | 202217016749-COMPLETE SPECIFICATION [09-01-2023(online)].pdf | 2023-01-09 |
| 22 | 202217016749-CLAIMS [09-01-2023(online)].pdf | 2023-01-09 |
| 23 | 202217016749-Response to office action [28-04-2025(online)].pdf | 2025-04-28 |
| 1 | 202217016749E_28-07-2022.pdf |