Sign In to Follow Application
View All Documents & Correspondence

Voice Processing Device, Voice Processing Method, And Non Transitory Computer Readable Medium Having Program Stored Thereon

Abstract: This voice processing device (1) comprises: a first segmenting means (2_1) that divides a first voice into a plurality of first segment voices; a second segmenting means (2_2) which divides a second voice into a plurality of second segment voices; a primary speaker recognition means (3) which calculates a score indicating the degree of similarity between each of a plurality of first and second segment voices; a threshold value calculation means (4) which calculates a threshold value on the basis of a score indicating the degree of similarity between the plurality of first segment voices, from among the plurality of scores calculated by the primary speaker recognition means (3); a speaker clustering means (5) which classifies, into one or a plurality of clusters, a plurality of second segment voices each having a higher degree of similarity than a degree of similarity indicated by a threshold value; and a secondary speaker recognition means (6) which calculates the degree of similarity between the first voice and each of the one or plurality of clusters, and on the basis of the calculation results, determines whether a voice which corresponds to the first voice is contained any of the one or plurality of clusters.

Get Free WhatsApp Updates!
Notices, Deadlines & Correspondence

Patent Information

Application #
Filing Date
06 December 2021
Publication Number
21/2022
Publication Type
INA
Invention Field
ELECTRONICS
Status
Email
archana@anandandanand.com
Parent Application

Applicants

NEC CORPORATION
7-1, Shiba 5-chome, Minato-ku, Tokyo 1088001

Inventors

1. GUO Ling
c/o NEC Corporation, 7-1, Shiba 5-chome, Minato-ku, Tokyo 1088001
2. YAMAMOTO Hitoshi
c/o NEC Corporation, 7-1, Shiba 5-chome, Minato-ku, Tokyo 1088001
3. KOSHINAKA Takafumi
c/o NEC Corporation, 7-1, Shiba 5-chome, Minato-ku, Tokyo 1088001

Specification

[0001]The present disclosure relates to a speech processing device, a speech processing method, and a non-transitory computer readable medium storing a 10 program.
Background Art [0002]
A speech processing device (speaker recognition device) has been known
15 which determines, based on a similarity between speech to be recognized
(recognition speech) input from the outside and registrant's speech (registered speech) stored in advance, whether the recognition speech is a registrant's utterance. [0003]
20 For example, the speech processing device compares an individuality
feature extracted from the recognition speech with an individuality feature extracted from the registered speech to calculate a score indicating a similarity therebetween and to determine based on the calculation result (score) whether the recognition speech is the registrant's utterance.
25 [0004]
Here, the individuality feature is a feature representing the difference between individuals that can be extracted from speech. For example, the individuality feature can be represented by physical quantities such as spectral envelope characteristics reflecting vocal tract information and fundamental
30 frequency characteristics reflecting vocal cord information, and specifically can be represented using mel-frequency cepstrum coefficients (MFCC). Such a speech processing device (speaker recognition device) is expected to be put into practical use in a wide range as biometric authentication having both accuracy and convenience.

3
[0005]
However, the conventional speech processing device may be reduced in performance due to various factors when being used in a real environment. This is because the conventional speech processing device is difficult to extract 5 accurate individuality feature in the real environment. For example,
individuality features extracted in a noisy environment may contain noise components. Further, individuality features extracted from respective types of speech of the same person recorded in different facilities are ideally the same, but may be actually different from each other. Therefore, the conventional speech
10 processing device may have problems that the registered speech and the recognition speech deviate in individuality feature from each other by
comparison, and as a result, speaker recognition cannot be accurately performe d. [0006]
In particular, the recognition speech often makes use of speech recorded
15 under an arbitrary environment, and thus may contain voices of persons other than the speaker of the recognition target. As in the case of a noisy environment, when the speech of the speaker of the recognition target and other voices are mixed, performance of the speaker recognition is significantly reduced. In order to solve such problems, various research institutes are proceeding with research
20 and development of related technologies regarding speaker recognition with high performance in consideration of actual use situations. [0007]
For example, Non Patent Literature 1 discloses a speaker diarization method of classifying one input speech, in which different types of speech of multiple
25 speakers are mixed, for each speaker. In such a method, after the input speech is divided into short speech segments, an individuality feature is extracted from each of the speech segments. Then, using the extracted individuality featu re, speaker clustering of an agglomerative hierarchical clustering (AHC) type is performed. Finally, speaker change point re-detection is performed, and cluster
30 numbers are output into which respective speech segments are classified. [0008]
In addition, Non Patent Literature 2 discloses a technique in which the speaker diarization method disclosed in Non Patent Literature 1 is used as pre -processing of speaker recognition. In Non Patent Literature 2, after speaker

4
diarization is performed on recognition speech in which different types of speech of multiple speakers are mixed, speaker recognition is performed, and thus an effect of improving the speaker recognition using the speaker diarization is studied. From the result of the study, it is concluded that accuracy of speaker 5 recognition is improved for the recognition speech, in which different types of speech of multiple speakers are mixed, using the speaker diarization method as pre-processing of speaker recognition. [0009]
In the speaker recognition method using the speaker diarization disclosed in
10 Non Patent Literature 2 as pre-processing, first, the recognition speech is divided into short speech segments, and an individuality feature is extracted from each of the speech segments. Then, scores indicating similarities between the speech segments are calculated using the extracted individuality features, clustering of an agglomerative hierarchical clustering (AHC) type is performed based on such
15 scores, and cluster numbers are output into which the respective speech segments are classified. Thereafter, an individuality feature of cluster speech obtained by coupling of the speech segments classified into the same cluster number is obtained, and an individuality feature of each cluster speech is compared with an individuality feature of the registered speech to calculate a score indicating a
20 similarity therebetween. Then, it is determined based on the highest score among the calculated scores whether registrant's speech is contained in the recognition speech in which the plurality of types of speech are mixed. In such a method, the individuality feature can be accurately extracted by classification of the recognition speech, in which the plurality of types of speech of multiple
25 speakers are mixed, into a plurality of types of cluster speech of a single speaker, and thus the accuracy of the speaker recognition can be improved. [0010]
Other techniques related to speech recognition are also disclosed in Patent Literatures 1 and 2.
30
Citation List Patent Literature [0011]
Patent Literature 1: Japanese Unexamined Patent Application Publication

5
No. 2019-008131
Patent Literature 2: Japanese Unexamined Patent Application Publication No. H11-249684 Non Patent Literature 5 [0012]
Non Patent Literature 1: Gregory Sell, David Snyder, Alan McCree, Daniel Garcia-Romero, Jesus Villalba, Matthew Maciejewski, Vimal Manohar, Najim Dehak, Daniel Povey, Shinji Watanabe, Sanjeev Khudanpur, "Diarization is Hard: Some Experiences and Lessons Learned for the JHU Team in the Inaugural 10 DIHARD Challenge", Proc. INTERSPEECH, pp. 2808-2812, 2018.
Non Patent Literature 2: David Snyder, Daniel Garcia-Romero, Gregory
Sell, Alan McCree, Daniel Povey, Sanjeev Khudanpur, "Speaker recognition for
multi-speaker conversations using x-vectors", IEEE International Conference on
Acoustics, Speech and Signal Processing (ICASSP), 2019.
15 Non Patent Literature 3: Jitendra Ajmera, Iain McCowan, and Herve
Bourlard, "Robust Speaker Change Detection", IEEE Signal Processing Letters, VOL. 11, NO. 8, pp. 649-651, 2004.
Non Patent Literature 4: Yin Ruiqing, Herve Bredin, and Claude Barras, "Speaker Change Detection in Broadcast TV using Bidirectional Long Short-Term 20 Memory Networks", Proc. INTERSPEECH, pp. 3827-3831, 2017.
Summary of Invention Technical Problem [0013]
25 However, in the speaker clustering executed in the speaker diarization, it is
necessary to determine in advance a threshold value as a criterion for determining whether a certain speech segment is similar to another speech segment. Here, the clustering result by the speaker clustering, particularly, the number of clusters greatly fluctuates depending on the threshold value. Accordingly, the accuracy
30 of speaker recognition also greatly fluctuates. To obtain an optimum threshold value, it is necessary to adjust the threshold value such that a recognition error rate is minimized by preparing a speech corpus containing speech similar to the registered speech and performing speaker recognition after performing speaker diarization on the speech corpus. In this case, since it is necessary to repeat the

6
speaker diarization and the speaker recognition several times, there has been a
problem that calculation costs have become huge.
[0014]
The present disclosure has been made to solve such a problem. In other
5 words, the present disclosure aims to provide a speech processing device, a speech processing method, and a non-transitory computer readable medium storing a program capable of determining with high accuracy whether speech of a specific speaker is contained in a speech signal in which different types of speech of multiple speakers are mixed.
10
Solution to Problem [0015]
A speech processing device according to the present disclosure includes: first segment means for dividing predetermined first speech into a plurality of
15 first speech segments; second segment means for dividing second speech supplied from the outside into a plurality of second speech segments; primary speaker recognition means for calculating scores indicating similarities between the plurality of first speech segments and the plurality of second speech segments; threshold value calculation means for calculating a threshold value based on
20 scores indicating similarities between the plurality of first speech segments
among the scores calculated by the primary speaker recognition means; speaker clustering means for classifying each of the plurality of second speech segments into one or more clusters having a similarity higher than the similarity indicated by the threshold value; and secondary speaker recognition means for calculating a
25 similarity between each of the one or more clusters and the first speech and
determining based on a result of the calculation whether speech corresponding to
the first speech is contained in any of the one or more clusters.
[0016]
A speech processing method according to the present disclosure includes: a
30 step of dividing predetermined first speech into a plurality of first speech
segments; a step of dividing second speech supplied from the outside into a plurality of second speech segments; a step of calculating scores indicating similarities between the plurality of first speech segments and the plurality of second speech segments; a step of calculating a threshold value based on scores

7
indicating similarities between the plurality of first speech segments among the calculated scores; a step of classifying each of the plurality of second speech segments into one or more clusters having a similarity higher than the similarity indicated by the threshold value; and a step of calculating a similarity between 5 each of the one or more clusters and the first speech and determining based on a result of the calculation whether speech corresponding to the first speech is contained in any of the one or more clusters. [0017]
A non-transitory computer readable medium storing a program according to
10 the present disclosure, the program causing a computer to execute: a process of dividing predetermined first speech into a plurality of first speech segments; a process of dividing second speech supplied from the outside into a plurality of second speech segments; a process of calculating scores indicating similarities between the plurality of first speech segments and the plurality of second speech
15 segments; a process of calculating a threshold value based on scores indicating similarities between the plurality of first speech segments among the calculated scores; a process of classifying each of the plurality of second speech segments into one or more clusters having a similarity higher than the similarity indicated by the threshold value; and a process of calculating a similarity between each of
20 the one or more clusters and the first speech and determining based on a result of the calculation whether speech corresponding to the first speech is contained in any of the one or more clusters.
Advantageous Effect of Invention
25 [0018]
According to the present disclosure, it is possible to provide a speech processing device, a speech processing method, and a non-transitory computer readable medium storing a program capable of determining with high accuracy whether speech of a specific speaker is contained in a speech signal in which
30 different types of speech of multiple speakers are mixed.
Brief Description of the Drawings [0019]
[Fig. 1] Fig. 1 is a block diagram showing a configuration example of a

8
speaker recognition device according to a first example embodiment.
[Fig. 2] Fig. 2 is a view showing an example of a method of dividing recognition speech into a plurality of speech segments.
[Fig. 3] Fig. 3 is a view showing another example of a method of dividing 5 recognition speech into a plurality of speech segments.
[Fig. 4] Fig. 4 is a view showing an example of a process in which a score is calculated by a primary speaker recognition unit.
[Fig. 5] Fig. 5 is a view showing a clustering result of the plurality of
speech segments divided and generated from the recognition speech.
10 [Fig. 6] Fig. 6 is a view showing a clustering result of a plurality of speech
segments divided and generated from each of registered speech and recognition speech.
[Fig. 7] Fig. 7 is a view showing a clustering result of a plurality of speech segments divided and generated from each of registered speech and recognition 15 speech.
[Fig. 8] Fig. 8 is a view showing an averaged result of scores indicating
similarities between each of the speech segments divided and generated from the
recognition speech and the speech segments divided and generated from the
registered speech.
20 [Fig. 9] Fig. 9 is a flowchart showing an operation of the speech processing
device according to the first example embodiment.
Description of Embodiments [0020]
25 An example embodiment according to the present disclosure will be
described below with reference to the drawings. In the respective drawings, the same or corresponding components are denoted by the same reference numerals, and will not be described as necessary for the sake of clarity of description. [0021]
30
Fig. 1 is a block diagram showing a configuration example of a speaker recognition device (speech processing device) 1 according to a first example embodiment. The speaker recognition device 1 according to the present example embodiment is a device configured to determine with high accuracy whe ther

9
specific speaker's speech is contained in recognition speech in which a plurality
of types of speech of multiple speakers are mixed. Hereinafter, a description
will be given in detail.
[0022]
5 As shown in Fig. 1, the speaker recognition device 1 includes a segment
unit (first segment means) 2_1, a segment unit (second segment means) 2_2, a primary speaker recognition unit (primary speaker recognition means) 3, a threshold value calculation unit (threshold value calculation means) 4, a speaker clustering unit (speaker clustering means) 5, and a secondary speaker recognition 10 unit (secondary speaker recognition means) 6. [0023] <>
The segment unit 2_1 divides registrant's speech (registered speech) stored in advance into a plurality of speech segments A1 to An (n being an integer of 2 15 or more). In the example of Fig. 1, registered speech fed back from the
secondary speaker recognition unit 6 to be described below is supplied to the
segment unit 2_1.
[0024]
The segment unit 2_2 divides speech to be recognized (recognition speech) 20 supplied from the outside into a plurality of speech segments B1 to Bm (m being an integer of 2 or more). The recognition speech is a single channel speech signal having a length of about 3 to 5 minutes, for example. [0025]
Fig. 2 is a view showing an example of a method of dividing the recognition 25 speech into a plurality of speech segments. In the example of Fig. 2, the recognition speech is divided into m speech segments B1 to Bm every two seconds without a gap. [0026]
Fig. 3 is a view showing another example of a method of dividing the 30 recognition speech into a plurality of speech segments.
In the example of Fig. 3, the recognition speech is divided into m speech segments B1 to Bm every two seconds with an overlap of one second. [0027]
In the examples of Figs. 2 and 3, the segment unit 2_2 divides the

10
recognition speech into the plurality of speech segments B1 to Bm every
predetermined time (for example, every two seconds), but is not limited theret o.
The segment unit 2_2 may divide the recognition speech into a plurality of speech
segments using a change point by a speaker change detection method, for 5 example. The speaker change point detection method is disclosed in Non Patent
Literatures 3 and 4, for example.
[0028]
A method of dividing the registered speech into a plurality of speech
segments A1 to An is the same as the method of dividing the recognition speech 10 into the plurality of speech segments B1 to Bm, and thus will not be described.
[0029]
<>
The primary speaker recognition unit 3 extracts feature representing
individuality (individuality feature) from each of the plurality of speech segments 15 A1 to An divided and generated from the registered speech and the plurality of
speech segments B1 to Bm divided and generated from the recognition speech.
And then, the primary speaker recognition unit 3 performs speaker recognition
using these extracted individuality features. Specifically, the primary speaker
recognition unit 3 calculates scores indicating similarities between the speech 20 segments using these extracted individuality features.
[0030]
Herein, the feature representing the speaker's individuality is an i-vector
calculated using MFCC (Mel-Frequency Cepstrum Coefficients) that is an
acoustic feature. Alternatively, the feature representing the speaker's 25 individuality is an x-vector calculated using DNN (Deep Neural Network). In
addition, the speaker recognition is a process of calculating a score regarding
whether two utterances are made by the same speaker. The score represents a
similarity of an i-vector (or an x-vector) calculated by PLDA (Probabilistic linear
discriminant analysis), for example. 30 [0031]
Fig. 4 is a view showing an example of a process in which a score is
calculated by the primary speaker recognition unit 3.
As shown in Fig. 4, the number of scores calculated by the primary speaker
recognition unit 3 is the square ((n+m)^2) of the sum of the plurality of speech

11
segments A1 to An divided and generated from the registered speech and the plurality of speech segments B1 to Bm divided and generated from the recognition speech. In Fig. 4, among a plurality of scores represented in a matrix, the scores at symmetrical positions indicate the same value. For example, score 1(n+m) = 5 score (n+m)1. [0032] <>
The threshold value calculation unit 4 calculates a threshold value based on the plurality of scores calculated by the primary speaker recognition unit 3.
10 Specifically, the threshold value calculation unit 4 calculates a threshold value based on a score group (a score group of a region P in Fig. 4) indicating the similarity between the speech segments A1 to An divided and generated from the registered speech among the plurality of scores calculated by the primary speaker recognition unit 3. Here, the registered speech is single speaker's speech.
15 Therefore, the score group of the region P is a score representing the same person. The threshold value calculation unit 4 outputs one of an average score of the score group of the region P, a minimum score thereof, an average score of some high ranks, and a minimum score thereof, as the threshold value. [0033]
20 <>
The speaker clustering unit 5 performs clustering of the speech segments, based on the plurality of scores calculated by the primary speaker recognition unit 3. [0034]
25 (First Specific Example of Speaker Clustering Unit 5)
First, a first specific example of the speaker clustering unit 5 will be described as a speaker clustering unit 5a. The speaker clustering unit 5a performs clustering of the speech segments B1 to Bm, based on a score group (a score group of a region S in Fig. 4) indicating similarity between the speech
30 segments B1 to Bn divided and generated from the recognition speech among the plurality of scores calculated by the primary speaker recognition unit 3. For the clustering, for example, the above-described AHC method is used. Here, the speaker clustering unit 5a classifies the score group of the region S in Fig. 4 into one or more clusters having a similarity (score) higher than the similarity

12
indicated by the threshold value calculated by the threshold value calculation unit
4.
[0035]
Fig. 5 is a view showing a clustering result of the speech segments B1 to 5 Bm divided and generated from the recognition speech. Referring to Fig. 5, the speech segments B1 to Bm divided and generated from the recognition speech is classified into clusters C1 to C3 having a similarity (score) higher than the similarity indicated by the threshold value calculated by the threshold value calculation unit 4.
10 [0036]
(Second Specific Example of Speaker Clustering Unit 5)
Next, a second specific example of the speaker clustering unit 5 will be described as a speaker clustering unit 5b. The speaker clustering unit 5b performs, based on all scores (score group in regions P, Q, R, and S in Fig. 4)
15 calculated by the primary speaker recognition unit 3, not only clustering of the speech segments B1 to Bm but also clustering of the speech segments A1 to An. For the clustering, for example, the above-described AHC method is used. Here, the threshold value calculation unit 4 outputs a threshold value such that the speech segments A1 to An divided and generated from the registered speech are
20 classified into one cluster. For example, the threshold value calculation unit 4
sets a score indicating the minimum value (that is, the score indicating the lowest similarity) in the score group indicating the similarity between the speech segments A1 to An divided and generated from the registered speech, as a threshold value.
25 [0037]
Figs. 6 and 7 are views showing clustering results of a plurality of speech segments divided and generated from each of the registered speech and the recognition speech. As shown in Figs. 6 and 7, the speech segments A1 to An divided and generated from the registered speech are classified into one cluster.
30 [0038]
Here, as shown in Fig. 6, among the speech segments B1 to Bm divided and generated from the recognition speech, the speech segments classified into the same cluster as the speech segments A1 to An divided and generated from the registered speech are classified into one cluster C1. By coupling (combining) of

13
the speech segments classified into such a cluster C1, one cluster speech
corresponding to the cluster C1 is generated.
[0039]
As shown in Fig. 7, when there is no speech segment classified into the 5 same cluster as the speech segments A1 to An divided and generated from the
registered speech among the speech segments B1 to Bm, all the speech segments B1 to Bm are classified as one cluster C1. By coupling of the speech segments classified into such a cluster C1, one cluster speech corresponding to the cluster C1 is generated.
10 [0040]
(Third Specific Example of Speaker Clustering Unit 5)
Next, a third specific example of the speaker clustering unit 5 will be described as a speaker clustering unit 5c. The speaker clustering unit 5c performs filtering based on a score group (a score group of a region Q in Fig. 4)
15 indicating a similarity between the speech segments A1 to An divided and
generated from the registered speech and the speech segments B1 to Bm divided
and generated from the recognition speech.
[0041]
Fig. 8 is a view showing an averaged result of scores indicating similarities
20 between each of the speech segments B1 to Bm divided and generated from the
recognition speech and the speech segments A1 to Am divided and generated from the registered speech. Fig. 8 also shows scores before averaging. Speech segments corresponding to the average value indicating a value larger than the threshold value (that is, indicating a high similarity) among these averaged values
25 are classified as one cluster C1. By coupling of the speech segments classified into such a cluster C1, cluster speech corresponding to the cluster C1 is generated. [0042]
When there is no average value indicating the value larger than the
30 threshold value (that is, indicating a similarity higher than the threshold value)
among these averaged values, all of the speech segments B1 to Bm are classified as one cluster C1. By coupling of the speech segments classified into the cluster C1, one cluster speech corresponding to the cluster C1 is generat ed. [0043]

14
For the clustering, another method may be used without being limited to the AHC method. For example, divisive hierarchical clustering of a top-down approach, may be used instead of the AHC method of a bottom-up approach. Alternatively, any clustering method (for example, a Mean-Shift method) 5 controlled by the threshold value similar to these methods may be used. [0044] <>
The secondary speaker recognition unit 6 calculates a similarity between the speech of each cluster (cluster speech) and the registered speech, and determines
10 based on such a calculation result whether any cluster speech contains speech corresponding to the registered speech. Specifically, the secondary speaker recognition unit 6 extracts an individuality feature of each cluster speech, and then compares the individuality feature of each cluster with the individuality feature of the registered speech (that is, recognizes the speaker), thereby
15 calculating a score indicating a similarity therebetween and outputting the score as a final score. Then, the secondary speaker recognition unit 6 determines, based on the highest score among these calculation results (final scores), whether the registrant's speech is contained in the recognition speech in which a plurality of types of speech are mixed.
20 [0045]
<>
Fig. 9 is a flowchart showing an operation of the speaker recognition device 1. [0046]
25 First, the segment unit 2_1 divides registered speech into a plurality of
speech segments A1 to An, and the segment unit 2_2 divides a recognition speech
into a plurality of speech segments B1 to Bm (step S101).
[0047]
Then, the primary speaker recognition unit 3 performs calculation (speaker
30 recognition) of a score indicating a similarity between segments using the speech segments A1 to An divided and generated from the registered speech and the speech segments B1 to Bm divided and generated from the recognition speech (step S102). [0048]

15
Then, the threshold value calculation unit 4 calculates a threshold value based on scores indicating similarities between the speech segments A1 to An divided and generated from the registered speech among the plurality of scores calculated by the primary speaker recognition unit 3 (step S103). 5 [0049]
Thereafter, the speaker clustering unit 5 classifies at least each of the speech segments B1 to Bm divided and generated from the recognition speech into one or more clusters having a similarity higher than the similarity indicated by the threshold value (step S104).
10 [0050]
Then, the secondary speaker recognition unit 6 calculates a similarity between speech of each cluster (cluster speech) and the registered speech, and determines based on such a calculation result whether any cluster speech contains the speech corresponding to the registered speech (step S105).
15 [0051]
As described above, the speaker recognition device 1 according to the present example embodiment can use the threshold value calculation unit 4 to set the threshold value used for clustering to an appropriate value. Thereby, the speaker recognition device 1 according to the present example embodiment can
20 accurately classify the recognition speech, in which the speech of the plurality of speakers are mixed, for each speaker in the speaker clustering unit 5, and thus can determine with high accuracy whether the recognition speech contains the registered speech. In the speaker recognition device 1 according to the present example embodiment, it is not necessary to optimize the threshold value used for
25 clustering at a huge calculation cost. [0052] <>
In Patent Literature 2, the threshold value used for clustering is calculated based on both the registered speech and the input speaker's speech. Therefore, it
30 is necessary to calculate the threshold value again every time when the input speaker's speech is input. As a result, according to Patent Literature 2, calculation costs become huge. On the other hand, in the speaker recognition device 1 according to the present example embodiment, the threshold value used for clustering is calculated based only on the registered speech. Therefore, in

16
the speaker recognition device 1 according to the present example embodiment, unlike Patent Literature 2, it is not necessary to calculate the threshold value again every time when the recognition speech (the input speaker's speech) is input, a huge calculation cost is prevented. 5 [0053]
The present disclosure has been described as a hardware configuration in the above-described example embodiment, but the present disclosure is not limited thereto. The present disclosure can also realize all or part of the processing of the speaker recognition device 1 by causing a CPU (Central
10 Processing Unit) to execute a computer program. [0054]
Further, the above-described program can be supplied to a computer in a state of being stored using various types of non-transitory computer readable media. The non-transitory computer readable medium includes various types of
15 tangible storage media. The non-transitory computer readable medium includes, for example, a magnetic recording medium, a magneto-optical recording medium, a CD-ROM (Read Only Memory), a CD-R, a CD-R/W, and a semiconductor memory. The magnetic recording medium is, for example, a flexible disk, a magnetic tape, or a hard disk drive. The magneto-optical recording medium is,
20 for example, a magneto-optical disk. The semiconductor memory is, for
example, a mask ROM, a PROM (Programmable ROM), an EPROM (Erasable PROM), a flash ROM, or a RAM (Random Access Memory). In addition, the program may be supplied to the computer by various types of transitory computer readable media. Examples of the transitory computer readable media include an
25 electric signal, an optical signal, and an electromagnetic wave. The transitory computer readable medium can supply the program to the computer via a wired communication path such as an electric wire or an optical fiber, or a wireless communication path. [0055]
30 Although the present disclosure has been described above with reference to
the example embodiment, the present disclosure is not limited to the above-described example embodiment. Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present disclosure within the scope of the present invention.

17
Reference Signs List [0056]
1 SPEAKER RECOGNITION DEVICE 5 2_1 SEGMENT UNIT 2_2 SEGMENT UNIT
3 PRIMARY SPEAKER RECOGNITION UNIT
4 THRESHOLD VALUE CALCULATION UNIT 4a THRESHOLD VALUE CALCULATION UNIT
10 4b THRESHOLD VALUE CALCULATION UNIT 4c THRESHOLD VALUE CALCULATION UNIT
5 SPEAKER CLUSTERING UNIT
6 SECONDARY SPEAKER RECOGNITION UNIT

WE CLAIMS

A speech processing device comprising:
first segment means for dividing predetermined first speech into a plurality 5 of first speech segments;
second segment means for dividing second speech supplied from outside into a plurality of second speech segments;
primary speaker recognition means for calculating scores indicating similarities between the plurality of first speech segments and the plurality of 10 second speech segments;
threshold value calculation means for calculating a threshold value based on scores indicating similarities between the plurality of first speech segments among the scores calculated by the primary speaker recognition means;
speaker clustering means for classifying each of the plurality of second 15 speech segments into one or more clusters having a similarity higher than the similarity indicated by the threshold value; and
secondary speaker recognition means for calculating a similarity between each of the one or more clusters and the first speech and determining based on a result of the calculation whether speech corresponding to the first speech is 20 contained in any of the one or more clusters.
2. The speech processing device according to Claim 1, wherein the speaker
clustering means is configured to classify all of the plurality of second speech
segments into one cluster when any of the plurality of second speech segments is
25 not classified into a cluster having the similarity higher than the similarity indicated by the threshold value.
3. The speech processing device according to Claim 1 or 2, wherein the
threshold value calculation means is configured to output one of an average score
30 of all of the scores indicating the similarities between the plurality of first speech segments, a minimum score thereof, an average score of selected ones of all of the scores, and a minimum score thereof, as the threshold value.
4. The speech processing device according to any one of Claims 1 to 3,

19
wherein the speaker clustering means is configured to perform clustering of the plurality of second speech segments based on the scores indicating the similarities between the plurality of second speech segments among the scores calculated by the primary speaker recognition means, and configured to combine the second 5 speech segments classified into each of the clusters to output cluster speech corresponding to the cluster.
5. The speech processing device according to any one of Claims 1 to 3,
wherein
10 the threshold value calculation means is configured to output the threshold
value such that all of the plurality of first speech segments are classified into one
cluster, and
the speaker clustering means is configured to perform clustering of the
plurality of first speech segments and the plurality of second speech segments 15 based on all of the scores calculated by the primary speaker recognition means,
and configured to combine the second speech segments classified into the same
cluster as the plurality of first speech segments among the plurality of second
speech segments to output cluster speech corresponding to the cluster.
20 6. The speech processing device according to Claim 5, wherein the speaker
clustering means is configured to classify all of the plurality of second speech segments into one cluster when no second speech segments are classified into the same cluster as the plurality of first speech segments among the plurality of second speech segments.
25
7. The speech processing device according to any one of Claims 1 to 3,
wherein the speaker clustering means is configured to calculate average values of
scores indicating similarities between each of the second speech segments and the
plurality of first speech segments, and to combine the second speech segments
30 indicating an average value of similarities higher than the similarity indicated by the threshold value among the calculated average values and to output the combined second speech segments as cluster speech.
8. The speech processing device according to Claim 7, wherein the speaker

20
clustering means is configured to combine all the plurality of second speech segments and to output the combined second speech segments as cluster speech when there are no second speech segments indicating the average value of the similarities higher than the similarity indicated by the threshold value among the 5 calculated average values.
9. A speech processing method comprising:
a step of dividing predetermined first speech into a plurality of first speech segments;
a step of dividing second speech supplied from outside into a plurality of second speech segments;
a step of calculating scores indicating similarities between the plurality of first speech segments and the plurality of second speech segments;
a step of calculating a threshold value based on scores indicating similarities between the plurality of first speech segments among the calculated scores;
a step of classifying each of the plurality of second speech segments into one or more clusters having a similarity higher than the similarity indicated by the threshold value; and
a step of calculating a similarity between each of the one or more clusters and the first speech and determining based on a result of the calculation whether speech corresponding to the first speech is contained in any of the one or more clusters.
25 10. A non-transitory computer readable medium storing a program that
causes a computer to execute:
a process of dividing predetermined first speech into a plurality of first speech segments;
a process of dividing second speech supplied from outside into a plurality 30 of second speech segments;
a process of calculating scores indicating similarities between the plurality of first speech segments and the plurality of second speech segments;
a process of calculating a threshold value based on scores indicating similarities between the plurality of first speech segments among the calculated

21
scores;
a process of classifying each of the plurality of second speech segments into
one or more clusters having a similarity higher than the similarity indicated by
the threshold value; and
5 a process of calculating a similarity between each of the one or more
clusters and the first speech and determining based on a result of the calculation whether speech corresponding to the first speech is contained in any of the one or more clusters.

Documents

Application Documents

# Name Date
1 202117056577.pdf 2021-12-06
2 202117056577-TRANSLATIOIN OF PRIOIRTY DOCUMENTS ETC. [06-12-2021(online)].pdf 2021-12-06
3 202117056577-STATEMENT OF UNDERTAKING (FORM 3) [06-12-2021(online)].pdf 2021-12-06
4 202117056577-REQUEST FOR EXAMINATION (FORM-18) [06-12-2021(online)].pdf 2021-12-06
5 202117056577-POWER OF AUTHORITY [06-12-2021(online)].pdf 2021-12-06
6 202117056577-NOTIFICATION OF INT. APPLN. NO. & FILING DATE (PCT-RO-105-PCT Pamphlet) [06-12-2021(online)].pdf 2021-12-06
7 202117056577-FORM 18 [06-12-2021(online)].pdf 2021-12-06
8 202117056577-FORM 1 [06-12-2021(online)].pdf 2021-12-06
9 202117056577-DRAWINGS [06-12-2021(online)].pdf 2021-12-06
10 202117056577-DECLARATION OF INVENTORSHIP (FORM 5) [06-12-2021(online)].pdf 2021-12-06
11 202117056577-COMPLETE SPECIFICATION [06-12-2021(online)].pdf 2021-12-06
12 202117056577-FORM 3 [27-05-2022(online)].pdf 2022-05-27
13 202117056577-FER.pdf 2022-06-03
14 202117056577-Proof of Right [03-08-2022(online)].pdf 2022-08-03
15 202117056577-Others-290822.pdf 2022-09-08
16 202117056577-Correspondence-290822.pdf 2022-09-08

Search Strategy

1 202117056577E_02-06-2022.pdf