Sign In to Follow Application
View All Documents & Correspondence

Speech Processing Device, Speech Processing Method, And Recording Medium

Abstract: Provided is a speech processing device, for example, capable of performing processing taking into consideration the influence of noise with respect to speech made by an utterance by a person. A speaker extracting unit extracts a speaker area from an image. A first utterance data generating unit, on the basis of the shape of the lips of the speaker, generates first utterance data indicating the content of the utterance by the speaker. A second utterance data generating unit, on the basis of a speech signal corresponding to the utterance by the speaker, generates second utterance data indicating the content of the utterance by the speaker. A comparison unit compares the first utterance data and the second utterance data with each other.

Get Free WhatsApp Updates!
Notices, Deadlines & Correspondence

Patent Information

Application #
Filing Date
01 February 2022
Publication Number
14/2022
Publication Type
INA
Invention Field
ELECTRONICS
Status
Email
Parent Application

Applicants

NEC CORPORATION
7-1, Shiba 5-chome, Minato-ku, Tokyo 1088001

Inventors

1. SASAKI Kazuyuki
c/o NEC CORPORATION, 7-1, Shiba 5-chome, Minato-ku, Tokyo 1088001

Specification

Technical field
[0001]
This disclosure relates to a voice processing device, a voice processing method, and a recording medium, and particularly to a voice processing device, a voice processing method, and a recording medium for processing a voice corresponding to a statement.
Background technology
[0002]
Patent Document 1 discloses that the content of a statement is reproduced from a voice signal by voice recognition. Specifically, Patent Document 1 describes that a microphone (microphone) is used to collect the remarks of a person and convert the audio signal output from the microphone into text data (character information). ing.
[0003]
Patent Document 2 discloses a technique of collating the lip pattern and voice of a speaker with pre-registered data and outputting specific character information when the collation matches.
[0004]
Further, as a related document, Patent Document 3 discloses a technique for learning the relationship between the shape of the lips and phonemes from a moving image including a voice generated by a speaker.
Prior art literature
Patent documents
[0005]
Patent Document 1: Japanese Unexamined Patent Publication No. 2008-160667
Patent Document 2: Japanese Patent Application Laid-Open No. 08-009254
Patent Document 3: Japanese Unexamined Patent Publication No. 2011-013731
Outline of the invention
Problems to be solved by the invention
[0006]
However, the speech recognition technology described in Patent Documents 1 and 2 is not supposed to consider the magnitude of the influence of noise on the speech (utterance) of a person's speech. In that case, for example, when reproducing the content of the remark from the voice, there is a possibility that the content of the remark cannot be accurately reproduced.
[0007]
This disclosure has been made in view of the above problems, and one of the purposes of the disclosure is to provide a voice processing device or the like that enables processing in consideration of the influence of noise on the voice caused by a person's remark. ..
Means to solve problems
[0008]
The voice processing device according to one aspect of the disclosure is a speaker extraction means for extracting a speaker area from an image, and a first unit showing the content of the speaker's remark based on the shape of the speaker's lip. A first speech data generation means for generating speech data and a second speech data generation means for generating a second speech data indicating the content of the speaker's speech based on the speech signal corresponding to the speaker's speech. And a collation means for collating the first speech data with the second speech data.
[0009]
The voice processing method according to one aspect of this disclosure extracts a speaker area from an image and generates first speech data indicating the content of the speaker's speech based on the shape of the speaker's lip. , A second speech data indicating the content of the speaker's speech is generated based on the voice signal corresponding to the speaker's speech, and the first speech data and the second speech data are collated. Including doing.
[0010]
The recording medium according to one aspect of this disclosure extracts the speaker's area from the image and generates first speech data indicating the content of the speaker's speech based on the shape of the speaker's lip. To generate a second remark data indicating the content of the remark of the speaker based on the voice signal corresponding to the remark of the speaker, the first remark data, and the second remark. Stores a program that causes a computer to collate with speech data.
A brief description of the drawing
[0011]
FIG. 1 is a block diagram showing an example of a configuration of a voice processing device according to the first embodiment.
FIG. 2 is a block diagram showing an example of a configuration of a first speech data generation unit included in the voice processing device according to the first embodiment.
FIG. 3 is a flowchart showing an example of an operation flow of the voice processing device according to the first embodiment.
FIG. 4 is a block diagram showing an example of a configuration of a first speech data generation unit included in the voice processing device according to the second embodiment.
FIG. 5 is a block diagram showing an example of a configuration of a first speech data generation unit included in the voice processing device according to the third embodiment.
FIG. 6 is a block diagram showing an example of the configuration of the voice processing device according to the fourth embodiment.
FIG. 7 is a flowchart showing an example of an operation flow of the voice processing device according to the fourth embodiment.
FIG. 8 is a block diagram showing an example of the configuration of the voice processing device according to the fifth embodiment.
FIG. 9 is a block diagram showing an example of a configuration of a modification of the voice processing device according to the fifth embodiment.
FIG. 10 is a flowchart showing an example of an operation flow of the voice processing device according to the fifth embodiment.
FIG. 11 is a block diagram showing an example of the configuration of the system according to the sixth embodiment.
FIG. 12 is an example of display by a display provided in the system according to the sixth embodiment.
FIG. 13 is a diagram showing an example of a hardware configuration of an information processing apparatus according to the seventh embodiment.
Embodiment for carrying out the invention
[0012]
[Embodiment 1]
The first embodiment will be described with reference to FIGS. 1 to 3.
[0013]
(Voice processing device 1)
The configuration of the voice processing device 1 according to the first embodiment will be described with reference to FIG. FIG. 1 is a block diagram showing an example of the configuration of the voice processing device 1. As shown in FIG. 1, the voice processing device 1 includes a speaker extraction unit 20, a first speech data generation unit 30, a collation unit 40, and a second speech data generation unit 50. The functions of each part of the voice processing device 1 (and the voice processing device according to each embodiment described later) according to the first embodiment are realized as software by executing the program read into the memory by the processor. It may be realized as hardware such as an intelligent camera.
[0014]
The speaker extraction unit 20 extracts the speaker area from the image. The speaker extraction unit 20 is an example of the speaker extraction means.
[0015]
Specifically, the speaker extraction unit 20 acquires time-series image data from a camera or the like (not shown). For example, the time-series image data is an image frame of a moving image for a certain period of time. Alternatively, the time-series image data may be data of a plurality of still images taken at predetermined time intervals. When the function of the speaker extraction unit 20 is realized by the intelligent camera, the intelligent camera itself captures time-series image data.
[0016]
The speaker extraction unit 20 extracts a speaker area from each image data by performing image analysis on the acquired time-series image data. For example, the speaker extraction unit 20 detects a person's area from each image data by using a classifier (also referred to as a trained model) that has learned the characteristics (personality) of the person. The detected area of ​​the person is the area of ​​the image containing at least a part of the person. The area of ​​the person is, for example, a rectangular area surrounding the face portion of the person in the image data.
[0017]
Further, the speaker extraction unit 20 identifies the lip portion of the person from the image data of the detected area of ​​the person. The speaker extraction unit 20 discriminates the same person among time-series image data by, for example, face recognition (verification) or other means. The speaker extraction unit 20 detects a difference (that is, a change) in the shape of the lips of the same person among the time-series image data. When the shape of the lips of the same person changes between the time-series image data, the speaker extraction unit 20 determines that the person is a speaker.
[0018]
The speaker extraction unit 20 uses image data (hereinafter, may be referred to as speaker image data) including an area of ​​a person determined to be a speaker (hereinafter, may be referred to as a speaker area). It is transmitted to the first speech data generation unit 30.
[0019]
The first speech data generation unit 30 generates first speech data indicating the content of the speaker's speech based on the shape of the speaker's lips. The first speech data generation unit 30 is an example of the first speech data generation means. The first remark data generation unit 30 may generate the first remark data by using the technique described in Patent Document 3 (Japanese Patent Laid-Open No. 2011-13731). Alternatively, as described below, the first speech data generation unit 30 can directly generate the first speech data from the image data of the speaker.
[0020]
FIG. 2 is a block diagram showing an example of the configuration of the first remark data generation unit 30. As shown in FIG. 2, the first speech data generation unit 30 includes a mouth shape element identification unit 31 and a mouth shape element-phoneme conversion unit 32.
[0021]
The mouthpiece identification unit 31 receives the image data of the speaker from the speaker extraction unit 20. The mouth shape element identification unit 31 identifies the shape of the speaker's lips based on the speaker's image data. Then, the mouth shape element identification unit 31 identifies the mouth shape element from the shape of the speaker's lips. The mouth shape element identification unit 31 is an example of the mouth shape element identification means.
[0022]
The mouth shape element means the shape of the speaker's lips at the moment when the speaker speaks one phoneme. A phoneme is the smallest unit of speech that a listener can discern in a language. Specifically, a phoneme represents a vowel, consonant, or semivowel discriminated in one language.
[0023]
The mouth shape element identification unit 31 transmits information indicating the mouth shape element to the mouth shape element-phoneme conversion unit 32.
[0024]
The mouth shape element-phoneme conversion unit 32 receives information indicating the mouth shape element from the mouth shape element identification unit 31. The phoneme-phoneme conversion unit 32 converts the information indicating the phoneme into phoneme data to generate first speech data including time-series data of one or a plurality of phonemes. The phoneme-phoneme conversion unit 32 is an example of a phoneme-phoneme conversion means.
[0025]
Specifically, the mouth shape element-phoneme conversion unit 32 was identified from the shape of the speaker's lips with reference to the mouth shape element-phoneme correspondence table (not shown) showing the correspondence relationship between the mouth shape element and the phoneme. Searches and outputs phoneme data corresponding to the mouthpiece. The phoneme-phoneme correspondence table shows a one-to-one correspondence between phonemes and phonemes. In this way, the mouth shape element-phoneme conversion unit 32 executes the conversion from the mouth shape element to the phoneme.
[0026]
The phoneme-phoneme conversion unit 32 provides information indicating the phoneme (phoneme data) corresponding to the phoneme identified from the shape of the speaker's lip and the order of the phonemes (that is, the time series order of the phonemes). As the first speech data, it is transmitted to the collation unit 40. For example, the first speech data has a data structure in which a sequence number (1, 2, 3 ...) is added to one or a plurality of phonemes, respectively.
Alternatively, as described above, the first speech data generation unit 30 may directly generate the first speech data from the image data of the speaker. For example, the first speech data generation unit 30 may train a model (for example, a neural network) so that the corresponding phoneme or audio signal can be discriminated from the image data of the speaker by using the technique of deep learning. .. In this case, the first speech data generation unit 30 inputs the image data of the speaker into the trained model. The trained model discriminates the corresponding phoneme or audio signal from the input speaker image data, and outputs the discriminant result. The first speech data generation unit 30 generates the first speech data based on the output from the trained model.
[0027]
When the speaker extraction unit 20 extracts a plurality of different speaker areas from the time-series image data, the first speech data generation unit 30 first speaks for each speaker area. Generate data. That is, the first speech data generation unit 30 generates a plurality of first speech data corresponding to the plurality of speakers.
[0028]
The first remark data generation unit 30 transmits the generated first remark data to the collation unit 40 shown in FIG.
[0029]
The second speech data generation unit 50 shown in FIG. 1 generates second speech data indicating the content of the speaker's speech based on the voice signal corresponding to the speaker's speech. The second speech data generation unit 50 is an example of the second speech data generation means.
[0030]
Specifically, the voice signal corresponding to the speaker's speech is first input to the second speech data generation unit 50. For example, the second speech data generation unit 50 receives a voice signal corresponding to the speaker's speech collected by the microphone from the microphone. The image of the speaker whose speech is collected by the microphone is included in the time-series image data received by the speaker extraction unit 20.
[0031]
Alternatively, the second speech data generation unit 50 may acquire a pre-recorded audio signal. In this case, the image of the speaker who spoke is included in the recorded time-series image data.
[0032]
No. 1 (2) The speech data generation unit 50 generates a second speech data from the input audio signal.
[0033]
In the first method, the second speech data generation unit 50 uses information indicating the phonemes corresponding to the input audio signal and the order of the phonemes (that is, the time series order of the phonemes) as the second speech data. Generate.
[0034]
In the second method, the second speech data generation unit 50 uses information indicating the single notes included in the input audio signal and the order of the single notes (that is, the time series order of the single notes) as the second speech data. Generate. A single note is a sound that forms a syllable, and is the smallest unit of speech. Generally, a single note is represented by an audio signal having a fundamental frequency and an audio signal that is a multiple of the fundamental frequency.
[0035]
Here, the above-mentioned phoneme is a concept different from a single note. A phoneme is a voice that is discriminated as one single note in one language. Some single notes may be identified as the same phoneme. For example, [sh] and [s] are different single notes (consonants), but they are not distinguished in Japanese, so they are determined to be the same phoneme.
[0036]
In the following description, in order to distinguish between the second speech data generated by the first method and the second speech data generated by the second method, the former is called phoneme data and the latter is single note data. Called.
[0037]
The second remark data generation unit 50 transmits the generated second remark data to the collation unit 40.
[0038]
The collation unit 40 collates the first remark data with the second remark data. The collation unit 40 is an example of collation means.
[0039]
Specifically, the collation unit 40 receives the first remark data from the first remark data generation unit 30. Further, the collation unit 40 receives the second speech data from the second speech data generation unit 50. The collation unit 40 collates the first remark data with the second remark data.
[0040]
When there are a plurality of first remark data, the collation unit 40 collates each of the plurality of first remark data with the second remark data.
[0041]
In the first embodiment, the second speech data may be either phoneme data or single note data described above. First, a case where the second speech data is phoneme data, that is, a case where the second speech data is information indicating the phonemes corresponding to the audio signal and the order of the phonemes will be described below.
[0042]
(Collation method: Example 1)
In Example 1, the collation unit 40 generates a first feature vector in which the features of each phoneme included in the first speech data are arranged according to the sequence number added to the phoneme. Further, the collation unit 40 generates a second feature vector in which the features of each phoneme included in the second speech data are arranged according to the sequence number added to the phoneme. For example, a phoneme feature vector is the amplitude, power, power spectrum, or Mel-Frequency Cepstrum Coefficients (MFCC) of a standard phoneme representing the phoneme. These feature vectors are obtained by performing various transformations on the audio signal representing the phoneme.
[0043]
The collation unit 40 calculates the distance between the first feature vector and the second feature vector. The collation unit 40 calculates the degree of similarity between the first feature vector and the second feature vector based on the calculated distance. The similarity is represented by, for example, a single number between 0 (not quite similar) and 1 (exact match).
[0044]
When the similarity exceeds the threshold value, the collation unit 40 determines that the first remark data and the second remark data are the same (collation success). On the other hand, when the similarity is equal to or less than the threshold value, the collation unit 40 determines that the first remark data and the second remark data are not the same (collation failure).
[0045]
(Collation method: Example 2)
In Example 2, the collation unit 40 performs matching between the individual phonemes included in the first speech data and the individual phonemes included in the second speech data.
[0046]
More specifically, the collation unit 40 includes a phoneme to which the sequence number n (= 1, 2, ...) Be included in the first speech data is added, and the same sequence number n (in the second speech data). = 1, 2, ...) It is determined whether or not the phoneme to which the phoneme is added is the same. When the number of phonemes included in the first speech data is N1 and the number of phonemes included in the second speech data is N2 (≦ N1), the collation unit 40 is from the first to the N1st. It is determined whether or not the phonemes of are the same.
[0047]
The collation unit 40 counts the number of times the matching is successful, that is, the number of phonemes that are the same between the first speech data and the second speech data.
[0048]
When the number of successful matchings exceeds a predetermined number, the collation unit 40 determines that the first remark data and the second remark data are the same (matching success). On the other hand, when the number of successful matchings is less than or equal to a predetermined number, the collation unit 40 determines that the first remark data and the second remark data are not the same (matching failure).
[0049]
Next, a case where the second speech data is single note data, that is, a case where the second speech data is information indicating the single notes included in the audio signal and the order of the single notes will be described below. ..
[0050]
(Collation method: Example 3)
In Example 3, the collation unit 40 considers one or more single notes that are not distinguished from each other in one language to be the same phoneme. The collation unit 40 performs matching between each of one or a plurality of single notes considered to be the same as the individual phonemes contained in the first speech data and the individual single notes included in the second speech data. The matching method in Example 3 is the same as when the second speech data is phoneme data.
[0051]
(Collation method: Example 4)
In Example 4, when the first speech data and the second speech data are audio signals, a method of collating the two will be described.
[0052]
In Example 4, the first speech data generation unit 30 further converts the phoneme converted from the phoneme element by the phoneme-phoneme conversion unit 32 into a voice signal corresponding to the phoneme.
[0053]
For example, the first speech data generation unit 30 converts a phoneme into a corresponding audio signal by referring to a table (not shown) showing the correspondence between the phoneme and the audio signal. The second speech data generation unit 50 transmits the input audio signal itself as the second speech data to the collation unit 40.
[0054]
The collation unit 40 converts the audio signal, which is the first speech data, and the audio signal, which is the second speech data, into spectrograms, respectively. The collation unit 40 performs pattern matching between the first spectrogram representing the first speech data and the second spectrogram corresponding to the second speech data, and the first spectrogram and the second spectrogram are combined. Calculate the similarity of.
[0055]
When the degree of similarity exceeds the threshold value, the collation unit 40 determines that the first remark data and the second remark data match (collation success). On the other hand, when the similarity is equal to or less than the threshold value, the collation unit 40 determines that the first remark data and the second remark data do not match (collation failure).
[0056]
Using any of the methods from Example 1 to Example 4 described above, the collation unit 40 collates the first remark data with the second remark data.
[0057]
The collation unit 40 outputs the collation result of the first remark data and the second remark data. For example, the collation unit 40 outputs information indicating whether or not the collation between the first remark data and the second remark data is successful as a result of the collation.
[0058]
As a result, the voice processing device 1 can perform processing in consideration of the influence of noise on the voice (that is, utterance) caused by the speaker's speech.
[0059]
For example, the collation result by the collation unit 40 is used to associate the speaker with the statement made by the speaker (see the fourth embodiment). Alternatively, the collation result by the collation unit 40 may be used to correct the second remark data by using the first remark data (see the fifth embodiment).
[0060]
(Operation flow of voice processing device 1)
The operation flow of the voice processing device 1 according to the first embodiment will be described with reference to FIG. FIG. 3 is a flowchart showing an example of the operation flow of the voice processing device 1.
[0061]
As shown in FIG. 3, the speaker extraction unit 20 acquires time-series image data and extracts a speaker area from each image data (S101). The speaker extraction unit 20 transmits the extracted image data of the speaker to the first speech data generation unit 30.
[0062]
The first remark data generation unit 30 extracts a part of the speaker's lips from the image data of the speaker, and generates the first remark data based on the shape of the extracted speaker's lips (S102).
[0063]
The first remark data generation unit 30 transmits the generated first remark data to the collation unit 40.
[0064]
The second speech data generation unit 50 acquires a voice signal corresponding to the speaker's speech (S103).
[0065]
The second speech data generation unit 50 generates second speech data indicating the content of the speaker's speech based on the voice signal corresponding to the speaker's speech (S104). The second speech data generation unit 50 transmits the generated second speech data to the collation unit 40.
[0066]
The collation unit 40 receives the first remark data from the first remark data generation unit 30. Further, the collation unit 40 receives the second speech data from the second speech data generation unit 50. Then, the collation unit 40 collates the first remark data with the second remark data (S105).
[0067]
The collation unit 40 outputs the collation result in step S105. For example, the collation unit 40 outputs information indicating whether or not the collation between the first remark data and the second remark data is successful as a result of the collation.
[0068]
This completes the operation of the voice processing device 1.
[0069]
(Effect of this embodiment)
According to the configuration of the present embodiment, the speaker extraction unit 20 extracts the speaker area from the image. The first speech data generation unit 30 generates first speech data indicating the content of the speaker's speech based on the shape of the speaker's lips. The second speech data generation unit 50 generates second speech data indicating the content of the speaker's speech based on the voice signal corresponding to the speaker's speech. The collation unit 40 collates the first remark data with the second remark data.
[0070]
Thereby, for example, when the content of the remark is reproduced from the second remark data, the second remark data can be corrected by using the first remark data that has been successfully collated by the collation unit 40. Specifically, for example, even if the second speech data has noise, the noisy portion can be reproduced based on the first speech data. That is, the voice processing device can reproduce the speech from the voice signal with higher accuracy. Therefore, the voice processing device according to the first embodiment has an effect of making it possible to perform processing in consideration of the influence of noise on the voice caused by the speech of a person.
[0071]
[Embodiment 2]
The second embodiment will be described with reference to FIG. In the second embodiment, when the second speech data is phoneme data (that is, when the second speech data is generated by the first method), the detailed configuration of the second speech data generation unit is provided. explain.
[0072]
The configuration of the voice processing device according to the second embodiment is the same as that of the voice processing device 1 (FIG. 1) described in the first embodiment. The voice processing device according to the second embodiment includes a second speech data generation unit 250 (FIG. 4) instead of the second speech data generation unit 50 (FIG. 1).
[0073]
The second speech data according to the second embodiment is information indicating the phonemes corresponding to the audio signal and the order of the phonemes.
[0074]
(Second remark data generation unit 250)
FIG. 4 is a block diagram showing an example of the configuration of the second remark data generation unit 250 according to the second embodiment. As shown in FIG. 4, the second speech data generation unit 250 includes a feature extraction unit 252 and an audio signal-phoneme conversion unit 253.[0075]
The feature extraction unit 252 performs preprocessing such as sampling (A / D conversion (Analog-Digital Transform)) and filtering on the input audio signal, and then extracts features from the input audio signal. The feature extraction unit 252 is an example of a feature extraction means. The characteristics of the audio signal are, for example, the amplitude of the audio signal, the power of the audio signal at a certain frequency, or the spectrum (spectral envelope).
[0076]
The feature extraction unit 252 transmits information indicating the features extracted from the audio signal to the audio signal-phoneme conversion unit 253.
[0077]
The audio signal-phoneme conversion unit 253 receives information indicating the characteristics of the audio signal from the feature extraction unit 252. The voice signal-phoneme conversion unit 253 inputs the received feature to the trained model.
[0078]
The trained model is a model (for example, a neural network) trained so that phonemes can be discriminated from the characteristics of audio signals. As described above, a phoneme is the smallest unit of speech that a listener can discriminate in one language. The trained model outputs the discrimination result of the phoneme corresponding to the input audio signal.
[0079]
The audio signal-phoneme conversion unit 253 converts the characteristics of the audio signal into the corresponding phonemes based on the output from the trained model, and generates the second speech data including one or a plurality of phonemes. The audio signal-phoneme conversion unit 253 is an example of the audio signal-phoneme conversion means.
[0080] [0080]
The audio signal-phoneme conversion unit 253 transmits information indicating the input audio signal, the corresponding phoneme, and the order of the phonemes to the collation unit 40 as the second speech data. In the data structure of the second speech data, sequence numbers (1, 2, 3 ...) are added to one or a plurality of phonemes corresponding to the audio signal.
[0081]
The collation unit 40 collates the first remark data with the second remark data. Specifically, the collation unit 40 may collate the first remark data with the second remark data by using the method described as Example 1 or Example 2 of the collation method in the first embodiment. good. In the second embodiment, the description of the collation method will be omitted.
[0082]
The collation unit 40 outputs the collation result of the first remark data and the second remark data. For example, the collation unit 40 outputs information indicating whether or not the collation between the first remark data and the second remark data is successful as a result of the collation.
[0083]
(Effect of this embodiment)
According to the configuration of the present embodiment, the speaker extraction unit 20 extracts the speaker area from the image. The first speech data generation unit 30 generates first speech data indicating the content of the speaker's speech based on the shape of the speaker's lips. The second speech data generation unit 250 generates second speech data indicating the content of the speaker's speech based on the voice signal corresponding to the speaker's speech.
[0084]
The second speech data generation unit 250 includes a feature extraction unit 252 and an audio signal-phoneme conversion unit 253. The feature extraction unit 252 extracts features from the audio signal. The audio signal-phoneme conversion unit 253 converts the input audio signal into a phoneme corresponding to the input audio signal. The collation unit 40 collates the first remark data with the second remark data.
[0085]
This makes it possible for the voice processing device according to the second embodiment to perform processing in consideration of the influence of noise on the voice due to the speech of a person. For example, when the collation unit 40 succeeds in collation, the voice processing device according to the second embodiment corrects the second speech data by using the first speech data, so that the speech processing device according to the second embodiment makes the speech higher from the voice signal. It can be reproduced with high accuracy.
[0086]
[Embodiment 3]
The third embodiment will be described with reference to FIG. In the third embodiment, the details of the second speech data generation unit will be described with respect to the case where the second speech data is single note data (that is, the second speech data is generated by the second method). ..
[0087]
The configuration of the voice processing device (not shown) according to the third embodiment is the same as that of the voice processing device 1 (FIG. 1) described in the first embodiment. The voice processing device according to the third embodiment includes a second speech data generation unit 350 instead of the second speech data generation unit 50.
[0088]
The second speech data according to the third embodiment is information indicating the single notes included in the audio signal and the order of the single notes.
[0089]
(Second remark data generation unit 350)
In the third embodiment, the second speech data generation unit 350 generates the second speech data from the input voice signal by the third method described in the first embodiment.
[0090]
FIG. 5 is a block diagram showing an example of the configuration of the second remark data generation unit 350 according to the third embodiment. As shown in FIG. 5, the second speech data generation unit 350 includes a single note extraction unit 351.
[0091]
The single note extraction unit 351 extracts a single note included in the input audio signal and generates a second speech data including one or a plurality of single notes. The single note extraction unit 351 is an example of a single note extraction means. As described above, a single note is represented by an audio signal having a fundamental frequency and an audio signal that is a multiple of the fundamental frequency.
[0092]
The single note extraction unit 351 transmits information indicating the single note included in the input audio signal and the order of the single notes to the collation unit 40 as the second speech data. In the data structure of the second speech data, sequence numbers (1, 2, 3 ...) are added to one or a plurality of single notes corresponding to the audio signal.
[0093]
The collation unit 40 collates the first remark data with the second remark data. Specifically, the collation unit 40 may collate the first remark data with the second remark data by using the collation method described as Example 3 in the first embodiment.
[0094]
The collation unit 40 outputs the collation result of the first remark data and the second remark data. For example, the collation unit 40 outputs information indicating whether or not the collation between the first remark data and the second remark data is successful as a result of the collation.
[0095]
(Effect of this embodiment)
According to the configuration of the present embodiment, the speaker extraction unit 20 extracts the speaker area from the image. The first speech data generation unit 30 generates first speech data indicating the content of the speaker's speech based on the shape of the speaker's lips. The second speech data generation unit 350 generates second speech data indicating the content of the speaker's speech based on the voice signal corresponding to the speaker's speech.
[0096]
The second speech data generation unit 350 includes a single note extraction unit 351 that extracts single notes included in the input audio signal. The single note extraction unit 351 transmits information indicating the single note included in the input audio signal and the order of the single notes to the collation unit 40 as the second speech data. The collation unit 40 collates the first remark data with the second remark data.
[0097]
This makes it possible for the voice processing device according to the third embodiment to perform processing in consideration of the influence of noise on the voice due to the speech of a person. For example, when the collation unit 40 succeeds in collation, the remarks can be reproduced with higher accuracy from the voice signal by correcting the second remark data using the first remark data.
[0098]
[Embodiment 4]
The fourth embodiment will be described with reference to FIGS. 6 and 7. In the fourth embodiment, the speaker is identified from the speaker's speech and the corresponding voice signal and the corresponding second speech data based on the result of collation between the first speech data and the second speech data. The configuration for associating with speaker information will be described.
[0099]
(Voice processing device 4)
The configuration of the voice processing device 4 according to the fourth embodiment will be described with reference to FIG. FIG. 6 is a block diagram showing an example of the configuration of the voice processing device 4.
[0100]
As shown in FIG. 6, the voice processing device 4 includes a speaker extraction unit 20, a first speech data generation section 30, a collation section 40, a second speech data generation section 50, and a matching section 60. That is, the configuration of the voice processing device 4 according to the fourth embodiment is different from the configuration of the voice processing device 1 according to the first embodiment in that the mapping unit 60 is provided.
[0101]
The association unit 60 is connected to the storage unit 300. The storage unit 300 may be connected to the voice processing device 4 via a wireless or wired network. Alternatively, the storage unit 300 may be a part of the voice processing device 4. The storage unit 300 is an example of a storage means.
[0102]
In the fourth embodiment, the speaker extraction unit 20 extracts the speaker area from the time-series image data as described in the first embodiment. Further, the speaker extraction unit 20 generates speaker information that identifies the speaker extracted from the time-series image data.
[0103]
For example, the speaker extraction unit 20 extracts the area of ​​the speaker's face from the time-series image data. Then, the speaker extraction unit 20 generates face image data of the speaker as speaker information. Alternatively, the speaker extraction unit 20 may generate a feature vector representing the features of the speaker's face as speaker information.
[0104]
The speaker information is, for example, at least one of the speaker attribute information, the speaker's position information, the speaker's face image, and the first speech data. However, the speaker information is not limited to these as long as it is information for identifying the speaker.
[0105]
When a plurality of speakers are detected from the time-series image data, the speaker extraction unit 20 generates speaker information for each speaker. The speaker extraction unit 20 transmits the generated speaker information to the matching unit 60.
[0106]
The first remark data generation unit 30 receives the image data of the speaker from the speaker extraction unit 20. The first speech data generation unit 30 generates the first speech data by the image analysis described in the first embodiment based on the received image data. The first speech data generation unit 30 transmits the generated first speech data to the collation unit 40.
[0107]
The second remark data generation unit 50 generates the second remark data as described in the first embodiment. The second speech data generation unit 50 transmits the generated second speech data to the collation unit 40.
[0108]
In the fourth embodiment, the second speech data may be either phoneme data or single note data described above.
[0109]
The collation unit 40 receives the first remark data from the first remark data generation unit 30. Further, the collation unit 40 receives the second speech data from the second speech data generation unit 50. The collation unit 40 collates the first remark data with the second remark data.
[0110]
For example, the collation unit 40 collates the first remark data with the second remark data by using any one of Examples 1 to 4 of the collation method described in the first embodiment.
[0111]
The collation unit 40 transmits the collation result to the collation unit 60. For example, the collation unit 40 transmits, as a result of the collation, information or a flag indicating that the collation between the first remark data and the second remark data was successful or unsuccessful, to the matching unit 60.
[0112]
The association unit 60 associates the speaker information for identifying the speaker in the image with the second speech data based on the collation result. The association unit 60 is an example of the association means.
[0113]
Specifically, the mapping unit 60 receives speaker information from the speaker extraction unit 20. Further, the matching unit 60 receives the collation result described above from the collation unit 40. As described above, the result of the collation is, for example, information or a flag indicating that the collation between the first remark data and the second remark data was successful or unsuccessful.
[0114]
When the collation result indicates that the collation between the first remark data and the second remark data is successful, the matching unit 60 is referred to from the speaker extraction unit 20 based on the collation result by the collation unit 40. The received speaker information is associated with the second speech data received from the second speech data generation unit 50. For example, the association unit 60 assigns an ID (Identification) to a set of speaker information and second speech data.
[0115]
The association unit 60 assigns a set of speaker information and a second remark data to the set.It is stored in the storage unit 300 shown in FIG. 6 together with the ID. Alternatively, the mapping unit 60 may store the associated speaker information and the second speech data in a network server or the like (not shown). Alternatively, the mapping unit 60 converts the second speech data into a voice signal or text data corresponding to the speaker's speech, associates the converted voice signal or text data with the speaker information, and stores the storage unit. It may be stored in 300.
[0116]
(Operation flow of voice processing device 4)
The operation flow of the voice processing device 4 according to the fourth embodiment will be described with reference to FIG. 7. FIG. 7 is a flowchart showing an example of the operation flow of the voice processing device 4. Since steps S101 to S105 shown in FIG. 7 are common to the operation flow described in the first embodiment, the description of S101 to S105 will be omitted in the fourth embodiment. Hereinafter, it will be described from step S206 shown in FIG.
[0117]
As shown in FIG. 7, after the collation unit 40 collates between the second remark data and the first remark data (S105), the collation unit 60 receives the collation result from the collation unit 40. To receive. When the collating unit 40 fails to collate between the first remark data and the second remark data (No in S206), the process returns to the first step S101 of the operation flow shown in FIG.
[0118]
On the other hand, when the collation unit 40 succeeds in collating between the first remark data and the second remark data (Yes in S206), the collation unit 60 receives the speaker information from the speaker extraction unit 20. And the second speech data received from the second speech data generation unit 50 are associated with each other (S207).
[0119]
The association unit 60 stores the associated speaker information and the second speech data in the storage unit 300. This completes the operation of the voice processing device 4 according to the fourth embodiment.
[0120]
(Effect of this embodiment)
According to the configuration of the present embodiment, the speaker extraction unit 20 extracts the speaker area from the image. The first speech data generation unit 30 generates first speech data indicating the content of the speaker's speech based on the shape of the speaker's lips. The second speech data generation unit 50 generates second speech data indicating the content of the speaker's speech based on the voice signal corresponding to the speaker's speech. The collation unit 40 collates the first remark data with the second remark data.
[0121]
This makes it possible to perform processing in consideration of the influence of noise on the voice caused by a person's remark.
[0122]
Further, the matching unit 60 associates the speaker information for identifying the speaker in the image with the second speech data based on the collation result. Thereby, the voice processing device according to the fourth embodiment can easily create, for example, the minutes data describing who said what. Further, the voice processing device according to the fourth embodiment can identify the speaker even when there are a plurality of persons.
[0123]
[Embodiment 5]
The fifth embodiment will be described with reference to FIGS. 8 to 10. In the fifth embodiment, a configuration for correcting the second remark data by using the first remark data will be described.
[0124]
(Voice processing device 5)
The configuration of the voice processing device 5 according to the fifth embodiment will be described with reference to FIG. FIG. 8 is a block diagram showing an example of the configuration of the voice processing device 5. As shown in FIG. 8, the voice processing device 5 further includes a correction unit 70 in addition to the speaker extraction unit 20, the first speech data generation unit 30, the collation unit 40, and the second speech data generation unit 50. There is.
[0125]
That is, the configuration of the voice processing device 5 according to the fifth embodiment is different from the configuration of the voice processing device 1 according to the first embodiment in that the correction unit 70 is provided.
[0126]
In the fifth embodiment, not only the collation unit 40 but also the correction unit 70 receives the second remark data from the second remark data generation unit 50. In addition, the correction unit 70 receives the first speech data from the first speech data generation unit 30.
[0127]
In the fifth embodiment, the second speech data may be either the phoneme data or the single note data described above, as in the first embodiment.
[0128]
The correction unit 70 corrects the second speech data by using the first speech data received from the first speech data generation unit 30. The correction unit 70 is an example of correction means.
[0129]
The correction unit 70 may store the corrected second remark data in a storage unit (not shown), a network server, or both.
[0130]
Specifically, a specific example of a method in which the correction unit 70 corrects the second remark data by using the first remark data will be described below.
[0131]
First, the case where the second speech data is phoneme data, that is, the case where the second speech data is information indicating the phonemes corresponding to the audio signal and the order of the phonemes will be described below.
[0132]
(Correction method: Example 1)
In Example 1, the correction unit 70 compares the phoneme included in the first speech data with the corresponding phoneme included in the phoneme data which is the second speech data. The corresponding phoneme is a phoneme having the same added sequence number. In particular, the correction unit 70 compares the vowel included in the first speech data with the corresponding vowel included in the phoneme data which is the second speech data.
[0133]
When the vowels of both (that is, the first utterance data and the second utterance data) are the same, the correction unit 70 keeps the vowels of the second utterance data as they are.
[0134]
On the other hand, when the two vowels are different, the correction unit 70 replaces the vowel included in the second speech data with the corresponding vowel in the first speech data. In this way, the correction unit 70 corrects the second remark data by using the first remark data.
[0135]
(Correction method: Example 2)
In this example 2, the correction unit 70 selects a phoneme whose SN ratio (S / N) or likelihood is smaller than the threshold value among the phonemes included in the second speech data, and the phoneme corresponding to the first speech data. Replace with.
[0136]
Next, regarding the case where the second speech data is the above-mentioned single note data, that is, the case where the second speech data is information indicating the single notes included in the audio signal and the order of the single notes, the following will be applied. explain.
[0137]
(Correction method: Example 3)
In Example 3, the correction unit 70 adds weights according to the corresponding phonemes of the first speech data to the likelihoods of each of the plurality of single note candidates included in the second speech data, and weights them. Based on the likelihood of making, one of a plurality of candidates for a single note of the second speech data is selected.
[0138]
For example, assume that the likelihood of the first candidate of a certain single note included in the second speech data is a, and the likelihood of the second candidate of the same single note is A. On the other hand, it is assumed that the corresponding phoneme of the first speech data is the same as that of the first candidate. In this case, the correction unit 70 assigns a weight X (> 1) to the likelihood of the first candidate and a weight y (≦ 1) to the likelihood of the second candidate.
[0139]
The correction unit 70 compares the magnitude of the weighted likelihood X × a of the first candidate and the weighted likelihood y × A of the second candidate. The correction unit 70 selects the candidate having the larger weighted likelihood.
[0140]
(Modification example)
A modification of the voice processing device 5 according to the fifth embodiment will be described with reference to FIG. 9. FIG. 9 is a block diagram showing an example of the configuration of the voice processing device 5A according to the modified example.
[0141]
As shown in FIG. 9, the voice processing device 5A includes a first speech data generation unit 30, a collation unit 40, a second speech data generation unit 50, a correction unit 70, and a matching unit 60. That is, the configuration of the voice processing device 5A according to the present modification is different from the configuration of the voice processing device 5 in that the corresponding unit 60 is further provided.
[0142]
The association unit 60 associates the speaker information for identifying the speaker in the image data with the second speech data corrected by the correction unit 70. The association unit 60 is an example of the association means.
[0143]
Specifically, the mapping unit 60 receives speaker information from the speaker extraction unit 20. Further, the association unit 60 receives the corrected second remark data from the correction unit 70. Further, the matching unit 60 receives information or a flag from the collating unit 40 indicating that the collation between the first remark data and the second remark data is successful.
[0144]
When the matching unit 60 receives information or a flag indicating that the matching between the first speech data and the second speech data is successful from the collation unit 40, the speaker information received from the speaker extraction unit 20 And the corrected second remark data received from the correction unit 70 are associated with each other.
[0145]
The association unit 60 associates the corrected second speech data with the speaker information (for example, a speaker's face image) received from the speaker extraction unit 20, and stores the storage unit 300 (see the fourth embodiment). ) Etc.
[0146]
When the collation unit 40 collates the first remark data of a plurality of speakers with the second remark data, the collation unit 60 has the highest degree of similarity only 1 based on the collation result. The pair of the first remark data and the second remark data is specified. Then, the matching unit 60 associates the speaker information about one speaker corresponding to the specified first speech data with the second speech data.
[0147]
The association unit 60 stores the associated speaker information and the second speech data in the storage unit 300 shown in FIG. Alternatively, the mapping unit 60 may store the associated speaker information and the second speech data in a network server or the like (not shown).
[0148]
For example, the mapping unit 60 assigns an ID to each set of speaker information and second speech data. Then, the association unit 60 stores the set of the speaker information and the second speech data in the storage unit 300, the network server, or both together with the ID assigned to the set.
[0149]
In the following, the voice processing device 5 which does not have the mapping unit 60 (that is, which is not the above-mentioned modification) will be described.
[0150]
(Operation flow of voice processing device 5)
The operation flow of the voice processing device 5 according to the fifth embodiment will be described with reference to FIG. FIG. 10 is a flowchart showing an example of the operation flow of the voice processing device 5.
[0151]
In the flowchart shown in FIG. 10, steps S101 to S105 are common to the operation flow described in the first embodiment. Further, the following step S206 is common to the operation flow described in the fourth embodiment.
[0152]
Therefore, in the fifth embodiment, the description of S101 to S105 and S206 will be omitted. Hereinafter, the description will be given immediately after step S206 shown in FIG.
[0153]
As shown in FIG. 10, when the collation unit 40 fails to collate the first remark data with the second remark data (No in S206), the operation flow of the voice processing device 5 returns to step S101. On the other hand, when the collation unit 40 succeeds in collating the first remark data with the second remark data (Yes in S206), the correction unit 70 receives the first remark received from the first remark data generation unit 30. The second speech data received from the second speech data generation unit 50 is corrected by the data (S307).
[0154]
After that, the correction unit 70 may output the corrected second remark data. In the above-described modification, the correction unit 70 transmits the corrected second remark data to the association unit 60. The correspondence unit 60 associates the speaker information (for example, the speaker's face image data) received from the speaker extraction unit 20 with the second speech data corrected by the correction unit 70, and stores the storage unit 300 (for example). (See Fig. 6), etc.
[0155]
This completes the operation of the voice processing device 5 according to the fifth embodiment.
[0156]
(Effect of this embodiment)
According to the configuration of the present embodiment, the speaker extraction unit 20 extracts the speaker area from the image. The first speech data generation unit 30 generates first speech data indicating the content of the speaker's speech based on the shape of the speaker's lips. The second speech data generation unit 50 indicates the content of the speaker's speech based on the audio signal corresponding to the speaker's speech.Generate the second remark data. The collation unit 40 collates the first remark data with the second remark data. This makes it possible to perform processing in consideration of the influence of noise on the voice caused by the speech of a person.
[0157]
Further, when the matching between the first remark data and the second remark data is successful, the correction unit 70 corrects the second remark data by using the first remark data. Therefore, the accuracy of reproducing the content of the statement from the audio signal is improved.
[0158]
[Embodiment 6]
The sixth embodiment will be described with reference to FIGS. 11 and 12. In the sixth embodiment, the configuration of the system including the voice processing device 6 will be described.
[0159]
Even if the system is provided with any one of the voice processing device 1 to the voice processing device 5 described in the first to fifth embodiments in place of the voice processing device 6 according to the sixth embodiment. good.
[0160]
(System configuration)
The configuration of the system according to the sixth embodiment will be described with reference to FIG. FIG. 11 is a block diagram showing an example of a system configuration. As shown in FIG. 11, the system includes a microphone 100, a camera 200, and a display 400 in addition to the voice processing device 6. All or part of the system according to the sixth embodiment may be realized by an intelligent camera (for example, an IP camera or a network camera having an analysis function inside, also referred to as a smart camera).
[0161]
(Voice processing device 6)
The voice processing device 6 according to the sixth embodiment further includes a display control unit 80 in addition to the speaker extraction unit 20, the first speech data generation unit 30, and the collation unit 40. That is, the configuration of the voice processing device 6 according to the sixth embodiment is different from the configuration of the voice processing device 1 according to the first embodiment in that the display control unit 80 is provided.
[0162]
The microphone 100 collects the voice (utterance) of the speaker's remark and generates a voice signal corresponding to the speaker's remark. The microphone 100 includes one or more microphones. The microphone 100 transmits a voice signal corresponding to the voice of the speech to the second speech data generation unit 50.
[0163]
The camera 200 is installed in a place to be photographed (for example, in a conference room). The camera 200 takes a picture of a place to be photographed and a person in the place to be photographed, and image data in a time series (for example, a frame image of a moving image for a certain period of time, or a plurality of images taken at predetermined time intervals). The still image) is transmitted to the speaker extraction unit 20 and the display control unit 80.
[0164]
In the sixth embodiment, the collation unit 40 transmits the collation result of the first remark data and the second remark data to the speaker extraction unit 20.
[0165]
When the speaker extraction unit 20 receives from the collation unit 40 the result that the first remark data and the second remark data are successfully collated, the speaker extraction unit 20 includes image data for superimposition including a figure indicating an area including the speaker. (Hereafter, referred to as sub-image data) is generated. Then, the speaker extraction unit 20 transmits the generated sub-image data to the display control unit 80.
[0166]
Note that the speaker extraction unit 20 does not generate sub-image data when the collation unit 40 receives the result that the collation between the first remark data and the second remark data has failed.
[0167]
The display control unit 80 receives time-series image data from the camera 200. When the display control unit 80 has not received the sub image data from the speaker extraction unit 20, the display control unit 80 converts the time-series image data received from the camera 200 into a format that can be displayed on the display 400 and displays it on the display 400. Display the image.
[0168]
On the other hand, when the display control unit 80 receives the sub-image data from the speaker extraction unit 20, the display control unit 80 superimposes the received sub-image data on the time-series image data and converts the received sub-image data into a format that can be displayed on the display 400. , The superimposed image is displayed on the display 400. The display control unit 80 is an example of display control means. A specific example of the superimposed image will be described below.
[0169]
(Example of superimposed image)
FIG. 12 shows an example of a superimposed image generated by the display control unit 80 and displayed on the display 400.
[0170]
In the superimposed image shown in FIG. 12, a rectangular figure is displayed around the face of the speaker (the person in the upper right in the figure). This rectangular figure is an example of the above-mentioned sub-image data. The user can easily identify the speaker by looking at the superimposed image displayed on the display 400.
[0171]
In the superimposed image, the shape and mode of the figure pointing to the speaker (for example, hatching in the figure or the color, thickness, or type of the line constituting the figure) is not limited to the rectangle as shown in FIG. .. In another example, the figure pointing to the speaker may be an arrow pointing to the speaker.
[0172]
(Effect of this embodiment)
According to the configuration of the present embodiment, the speaker extraction unit 20 extracts the speaker area from the image. The first speech data generation unit 30 generates first speech data indicating the content of the speaker's speech based on the shape of the speaker's lips. The second speech data generation unit 50 generates second speech data indicating the content of the speaker's speech based on the voice signal corresponding to the speaker's speech. The collation unit 40 collates the first remark data with the second remark data.
[0173]
This makes it possible to perform processing in consideration of the influence of noise on the voice caused by a person's remark.
[0174]
Further, the display control unit 80 displays on the display 400 a superimposed image in which a figure indicating an area including a speaker is superimposed on the image. Therefore, the user can easily identify the speaker from the superimposed image on the display 400.
[0175]
[Embodiment 7]
The seventh embodiment will be described below with reference to FIG.
[0176]
(About hardware configuration)
Each component of the voice processing apparatus described in the first to sixth embodiments shows a block of functional units. Some or all of these components are realized by, for example, the information processing apparatus 900 as shown in FIG. FIG. 13 is a block diagram showing an example of the hardware configuration of the information processing apparatus 900. The information processing device 900 shows, for example, the internal configuration of an intelligent camera.
[0177]
As shown in FIG. 13, the information processing apparatus 900 includes the following configuration as an example.
[0178]
・ CPU (Central Processing Unit) 901
ROM (Read Only Memory) 902
・ RAM (Random Access Memory) 903
・ Program 904 loaded into RAM903
・ Storage device 905 for storing program 904
・ Drive device 907 that reads and writes the recording medium 906
・ Communication interface 908 that connects to the communication network 909
・ Input / output interface 910 for inputting / outputting data
・ Bus 911 connecting each component
Each component of the voice processing apparatus described in the first to sixth embodiments is realized by the CPU 901 reading and executing the program 904 that realizes these functions. The program 904 that realizes the functions of each component is stored in, for example, a storage device 905 or ROM 902 in advance, and the CPU 901 is loaded into the RAM 903 and executed as needed. The program 904 may be supplied to the CPU 901 via the communication network 909, or may be stored in the recording medium 906 in advance, and the drive device 907 may read the program and supply the program to the CPU 901.
[0179]
(Effect of this embodiment)
According to the configuration of the present embodiment, the voice processing device described in the above embodiment is realized as hardware. Therefore, it is possible to obtain the same effect as the effect described in the above embodiment.
[0180]
[Appendix]
A part or all of the above embodiments (and examples) may be described as the following appendices, but the following appendices are merely examples. Some or all of the above embodiments (and examples) are not limited to the configurations described in the following appendices.
[0181]
(Appendix 1)
A speaker extraction method that extracts the speaker area from the image,
The first speech data generation means for generating the first speech data indicating the content of the speaker's speech based on the shape of the speaker's lips,
A second speech data generation means for generating a second speech data indicating the content of the speaker's speech based on the voice signal corresponding to the speaker's speech, and
As a collation means for collating the first remark data with the second remark data
A voice processing device equipped with.
[0182]
(Appendix 2)
The first remark data generation means is
The mouth shape element identification means for identifying the mouth shape element from the shape of the speaker's lips,
Includes a phoneme-phoneme conversion means that converts the mouth form into a phoneme and generates the first speech data including one or more phonemes.
The voice processing device according to Appendix 1, which is characterized in that.
[0183]
(Appendix 3)
The second remark data generation means is
With feature extraction means to extract features from the input audio signal
Includes an audio signal-phoneme conversion means that converts the characteristics of the audio signal into corresponding phonemes and generates the second speech data including one or more phonemes.
The voice processing device according to Appendix 1 or 2, characterized in that.
[0184]
(Appendix 4)
The second remark data generation means is
Includes a single note extraction means that extracts a single note contained in the input audio signal and generates the second speech data including one or a plurality of single notes.
The voice processing device according to Appendix 1 or 2, characterized in that.
[0185]
(Appendix 5)
The speaker extraction means generates speaker information for identifying the speaker extracted from the image, and the speaker information is generated.
Further provided with a matching means for associating the speaker information with the second remark data based on the result of the collation.
The voice processing device according to any one of Supplementary note 1 to 4, which is characterized by the above.
[0186]
(Appendix 6)
The first speech data generation means generates a plurality of the first speech data based on the shapes of the lips of the plurality of speakers in the image.
The collation means collates each of the plurality of first remark data with the second remark data.
The matching means associates the speaker information regarding any one of the plurality of speakers with the second speech data based on the result of the collation.
The voice processing device according to Appendix 5, which is characterized in that.
[0187]
(Appendix 7)
When the matching between the first remark data and the second remark data is successful, the correction means for correcting the second remark data by using the first remark data is further provided.
The voice processing device according to any one of Supplementary note 1 to 6, which is characterized by the above.
[0188]
(Appendix 8)
Further provided with a display control means for displaying a superimposed image on the image on which a figure indicating an area including the speaker is superimposed on the image.
The voice processing device according to any one of Supplementary note 1 to 7, characterized in that.
[0189]
(Appendix 9)
Extract the speaker's area from the image,
Based on the shape of the speaker's lips, the first speech data indicating the content of the speaker's speech was generated.
Based on the voice signal corresponding to the speaker's remark, a second remark data indicating the content of the speaker's remark was generated.
Collate the first remark data with the second remark data
Voice processing method including that.
[0190]
(Appendix 10)
Extracting the speaker's area from the image and
Based on the shape of the speaker's lips, the first speech data indicating the content of the speaker's speech was generated, and
Based on the voice signal corresponding to the speaker's remark, the second remark data indicating the content of the speaker's remark was generated, and
Collating the first remark data with the second remark data
A recording medium that stores a program for executing a computer.
[0191]
(Appendix 11)
The speaker information is the above At least one of the speaker attribute information, the speaker's position information, the speaker's face image, and the first speech data.
The voice processing device according to Appendix 5 or 6, characterized in that.
[0192]
Although this disclosure has been described above with reference to the embodiments (and examples), the disclosure is not limited to the above embodiments (and examples). Various changes that can be understood by those skilled in the art can be made within the scope of this disclosure in the configuration and details of the above embodiments (and examples).
[0193]
This application claims priority on the basis of Japanese application Japanese Patent Application No. 2019-142951 filed on August 2, 2019, and incorporates all of its disclosures herein.
Code description
[0194]
1 Voice processing device
2 Voice processing device
3, 3A voice processing device
4, 4A voice processing device
5, 5A voice processing device
6 Voice processing device
20 Speaker extraction department
30 First remark data generation unit
31 Mouth shape element identification part
32 Mouth shape element-phoneme conversion unit
40 collation department
50 Second remark data generation unit
60 Correspondence part
70 correction part
80 display control unit
250 Second remark data generation unit
252 Feature extraction unit
253 Audio signal-phoneme conversion unit
300 memory unit
350 Second remark data generation unit
351 single note extraction unit
400 display
The scope of the claims
[Claim 1]
A speaker extraction method that extracts the speaker area from the image,
The first speech data generation means for generating the first speech data indicating the content of the speaker's speech based on the shape of the speaker's lips,
A second speech data generation means for generating a second speech data indicating the content of the speaker's speech based on the voice signal corresponding to the speaker's speech, and
As a collation means for collating the first remark data with the second remark data
A voice processing device equipped with.
[Claim 2]
The first remark data generation means is
The mouth shape element identification means for identifying the mouth shape element from the shape of the speaker's lips,
Includes a phoneme-phoneme conversion means that converts the mouth form into a phoneme and generates the first speech data including one or more phonemes.
The voice processing device according to claim 1, characterized in that.
[Claim 3]
The second remark data generation means is
With feature extraction means to extract features from the input audio signal
Includes an audio signal-phoneme conversion means that converts the characteristics of the audio signal into corresponding phonemes and generates the second speech data including one or more phonemes.
The voice processing device according to claim 1 or 2, characterized in that.
[Claim 4]
The second remark data generation means is
Includes a single note extraction means that extracts a single note contained in the input audio signal and generates the second speech data including one or a plurality of single notes.
The voice processing device according to claim 1 or 2, characterized in that.
[Claim 5]
The speaker extraction means generates speaker information for identifying the speaker extracted from the image, and the speaker information is generated.
Further provided with a matching means for associating the speaker information with the second remark data based on the result of the collation.
The voice processing device according to any one of claims 1 to 4, characterized in that.
[Claim 6]
The first speech data generation means generates a plurality of the first speech data based on the shapes of the lips of the plurality of speakers in the image.
The collation means collates each of the plurality of first remark data with the second remark data.
The matching means associates the speaker information regarding any one of the plurality of speakers with the second speech data based on the result of the collation.
The voice processing device according to claim 5, characterized in that.
[Claim 7]
When the matching between the first remark data and the second remark data is successful, the correction means for correcting the second remark data by using the first remark data is further provided.
The voice processing device according to any one of claims 1 to 6, characterized in that.
[Claim 8]
Further provided with a display control means for displaying a superimposed image on the image on which a figure indicating an area including the speaker is superimposed on the image.
The voice processing device according to any one of claims 1 to 7, characterized in that.
[Claim 9]
The speaker information is at least one of the attribute information of the speaker, the position information of the speaker, the face image of the speaker, and the first speech data.
The voice processing device according to claim 5 or 6, characterized in that.
[Claim 10]
Extract the speaker's area from the image,
Based on the shape of the speaker's lips, the first speech data indicating the content of the speaker's speech was generated.
Based on the voice signal corresponding to the speaker's remark, a second remark data indicating the content of the speaker's remark was generated.
Collate the first remark data with the second remark data
Voice processing method including that.
[Claim 11]
Extracting the speaker's area from the image and
Based on the shape of the speaker's lips, the first speech data indicating the content of the speaker's speech was generated, and
Based on the voice signal corresponding to the speaker's remark, the second remark data indicating the content of the speaker's remark was generated, and
Collating the first remark data with the second remark data
A recording medium that stores a program for executing a computer.

Documents

Application Documents

# Name Date
1 202217005467.pdf 2022-02-01
2 202217005467-TRANSLATIOIN OF PRIOIRTY DOCUMENTS ETC. [01-02-2022(online)].pdf 2022-02-01
3 202217005467-STATEMENT OF UNDERTAKING (FORM 3) [01-02-2022(online)].pdf 2022-02-01
4 202217005467-REQUEST FOR EXAMINATION (FORM-18) [01-02-2022(online)].pdf 2022-02-01
5 202217005467-PRIORITY DOCUMENTS [01-02-2022(online)].pdf 2022-02-01
6 202217005467-POWER OF AUTHORITY [01-02-2022(online)].pdf 2022-02-01
7 202217005467-NOTIFICATION OF INT. APPLN. NO. & FILING DATE (PCT-RO-105-PCT Pamphlet) [01-02-2022(online)].pdf 2022-02-01
8 202217005467-FORM 18 [01-02-2022(online)].pdf 2022-02-01
9 202217005467-FORM 1 [01-02-2022(online)].pdf 2022-02-01
10 202217005467-DRAWINGS [01-02-2022(online)].pdf 2022-02-01
11 202217005467-DECLARATION OF INVENTORSHIP (FORM 5) [01-02-2022(online)].pdf 2022-02-01
12 202217005467-COMPLETE SPECIFICATION [01-02-2022(online)].pdf 2022-02-01
13 202217005467-Proof of Right [04-03-2022(online)].pdf 2022-03-04
14 202217005467-Others-020522.pdf 2022-05-05
15 202217005467-Correspondence-020522.pdf 2022-05-05
16 202217005467-FER.pdf 2022-07-05
17 202217005467-OTHERS [30-12-2022(online)].pdf 2022-12-30
18 202217005467-FORM-26 [30-12-2022(online)].pdf 2022-12-30
19 202217005467-FORM 3 [30-12-2022(online)].pdf 2022-12-30
20 202217005467-FER_SER_REPLY [30-12-2022(online)].pdf 2022-12-30
21 202217005467-COMPLETE SPECIFICATION [30-12-2022(online)].pdf 2022-12-30
22 202217005467-CLAIMS [30-12-2022(online)].pdf 2022-12-30
23 202217005467-US(14)-HearingNotice-(HearingDate-02-12-2024).pdf 2024-10-28
24 202217005467-Correspondence to notify the Controller [14-11-2024(online)].pdf 2024-11-14

Search Strategy

1 ISRREPORTE_04-07-2022.pdf