Sign In to Follow Application
View All Documents & Correspondence

Information Processing Device, Information Processing Method, And Storage Medium

Abstract: Provided is an information processing device that is equipped with: an acquisition unit that acquires an acoustic characteristic in a frequency domain based on a sound wave propagating through the head of a user; and an extraction unit that generates a first frequency response function including a rational polynomial in which the denominator contains a term indicative of a peak feature of the acoustic characteristic with the numerator containing a term indicative of a notch feature of the acoustic characteristic, and then, on the basis of the first frequency response function, extracts a feature quantity which is used for biometric authentication of the user.

Get Free WhatsApp Updates!
Notices, Deadlines & Correspondence

Patent Information

Application #
Filing Date
07 March 2022
Publication Number
23/2022
Publication Type
INA
Invention Field
COMPUTER SCIENCE
Status
Email
Parent Application

Applicants

NEC CORPORATION
7-1, Shiba 5-chome, Minato-ku, Tokyo 1088001

Inventors

1. ITO, Yoshitaka
c/o NEC Corporation, 7-1, Shiba 5-chome, Minato-ku, Tokyo 1088001
2. ARAKAWA, Takayuki
c/o NEC Corporation, 7-1, Shiba 5-chome, Minato-ku, Tokyo 1088001

Specification

Title of the invention: Information processing apparatus, information processing method and storage medium
Technical field
[0001]
The present invention relates to an information processing apparatus, an information processing method and a storage medium.
Background technology
[0002]
Patent Document 1 discloses an earphone having a personal authentication function based on an acoustic signal propagating on the user's head. The personal authentication device of Patent Document 1 extracts a logarithmic spectrum, a Melkeptrum coefficient, a linear predictive analysis coefficient, etc. as feature quantities from acoustic characteristics and uses them for user authentication.
Prior art literature
Patent documents
[0003]
Patent Document 1: International Publication No. 2018/034178
Patent Document 2: International Publication No. 2018/1983310
Patent Document 3: Japanese Unexamined Patent Publication No. 2005-339265
Outline of the invention
Problems to be solved by the invention
[0004]
The acoustic characteristics acquired by a wearable device as described in Patent Document 1 may change depending on the wearable state. Since changes in acoustic characteristics due to such factors may affect the authentication accuracy, there is a need for a feature extraction method that is robust against differences in wearing conditions.
[0005]
An object of the present invention is to provide an information processing device, an information processing method, and a storage medium capable of extracting a feature amount robust to a difference in wearing state.
Means to solve the problem
[0006]
According to one aspect of the present invention, the denominator includes an acquisition unit for acquiring the acoustic characteristics of the frequency domain based on the sound wave propagating in the user's head, and a term indicating the characteristics of the peak of the acoustic characteristics. A first frequency response function having a rational polypoly in which a term indicating the notch characteristic of acoustic characteristics is included in the molecule is generated, and a feature quantity used for bioauthentication of the user is extracted based on the first frequency response function. An information processing apparatus is provided that comprises an extraction unit.
[0007]
According to another aspect of the present invention, the denominator includes a step of acquiring the acoustic characteristics of the frequency domain based on the sound wave propagating in the user's head, and a term indicating the characteristics of the peak of the acoustic characteristics. A first frequency response function having a rational polypoly in which a term indicating the notch characteristic of the acoustic characteristic is included in the molecule is generated, and a feature quantity used for bioauthentication of the user is calculated based on the first frequency response function. An information processing method is provided that comprises a step of extraction.
[0008]
According to another aspect of the present invention, the denominator includes a step in which the computer acquires the acoustic characteristics of the frequency domain based on the sound wave propagating in the user's head, and a term indicating the characteristics of the peak of the acoustic characteristics. A first frequency response function having a rational polypoly in which a term indicating the notch characteristic of the acoustic characteristic is included in the molecule is generated, and is used for biometric authentication of the user based on the first frequency response function. A storage medium is provided in which a program for executing an information processing method including a step of extracting a feature amount and a program for executing the information processing method is stored.
Effect of the invention
[0009]
According to the present invention, there is provided an information processing device, an information processing method, and a storage medium capable of extracting a feature amount robust to a difference in wearing state.
A brief description of the drawing
[0010]
FIG. 1 is a schematic diagram showing an overall configuration of an information processing system according to a first embodiment.
FIG. 2 is a block diagram showing a hardware configuration example of the earphone control device according to the first embodiment.
FIG. 3 is a block diagram showing a hardware configuration example of the information communication device according to the first embodiment.
FIG. 4 is a functional block diagram of an earphone and an information communication device according to the first embodiment.
FIG. 5 is a flowchart showing an outline of a biometric authentication process performed by the information communication device according to the first embodiment.
FIG. 6 is a flowchart showing a feature quantity extraction process performed by the information communication device according to the first embodiment.
[Fig. 7] Fig. 7 is a graph showing a measurement example of acoustic characteristics.
[Fig. 8] Fig. 8 is a graph showing an example of changes in acoustic characteristics when the earphones are repeatedly attached and detached.
FIG. 9 is a flowchart showing a feature quantity extraction process performed by the information communication device according to the second embodiment.
FIG. 10 is a flowchart showing a feature quantity extraction process performed by the information communication device according to the third embodiment.
FIG. 11 is a flowchart showing a feature quantity extraction process performed by the information communication device according to the fourth embodiment.
FIG. 12 is a functional block diagram of the information processing apparatus according to the fifth embodiment.
Embodiment for carrying out the invention
[0011]
Hereinafter, an exemplary embodiment of the present invention will be described with reference to the drawings. Similar elements or corresponding elements may be designated by the same reference numerals in the drawings, and the description thereof may be omitted or simplified.
[0012]
[First embodiment]
The information processing system according to this embodiment will be described. The information processing system of the present embodiment is a system for performing biometric authentication by a wearable device such as an earphone.
[0013]
FIG. 1 is a schematic diagram showing the overall configuration of the information processing system according to the present embodiment. The information processing system includes an information communication device 1 and an earphone 2 that can be wirelessly connected to each other.
[0014]
The earphone 2 includes an earphone control device 20, a speaker 26, and a microphone 27. The earphone 2 is an acoustic device that can be worn on the head of the user 3, particularly the ear, and is typically a wireless earphone, a wireless headset, or the like. The speaker 26 functions as a sound wave generating unit that emits sound waves toward the external auditory canal of the user 3 when worn, and is arranged on the mounting surface side of the earphone 2. The microphone 27 is arranged on the mounting surface side of the earphone 2 so that it can receive the sound echoed by the external auditory canal of the user 3 at the time of wearing. The earphone control device 20 controls the speaker 26 and the microphone 27 and communicates with the information communication device 1.
[0015]
In the present specification, "sound" such as sound sound and voice includes inaudible sound whose frequency or sound pressure level is out of the audible range.
[0016]
The information communication device 1 is, for example, a computer communicably connected to the earphone 2 and performs biometric authentication based on acoustic information. The information communication device 1 further controls the operation of the earphone 2, transmits voice data for generating a sound wave emitted from the earphone 2, receives voice data obtained from the sound sound received by the earphone 2, and the like. As a specific example, when the user 3 listens to music using the earphone 2, the information communication device 1 transmits the compressed data of the music to the earphone 2. When the earphone 2 is a telephone device for business commands at an event venue, a hospital, or the like, the information communication device 1 transmits voice data of business instructions to the earphone 2. In this case, the voice data of the speech of the user 3 may be further transmitted from the earphone 2 to the information communication device 1.
[0017]
Note that this overall configuration is an example, and for example, the information communication device 1 and the earphone 2 may be connected by wire. Further, the information communication device 1 and the earphone 2 may be configured as an integrated device, or another device may be included in the information processing system.
[0018]
FIG. 2 is a block diagram showing a hardware configuration example of the earphone control device 20. The earphone control device 20 includes a CPU (Central Processing Unit) 201, a RAM (Random Access Memory) 202, a ROM (Read Only Memory) 203, and a flash memory 204. Further, the earphone control device 20 includes a speaker I / F (Interface) 205, a microphone I / F206, a communication I / F207, and a battery 208. Each part of the earphone control device 20 is connected to each other via a bus, wiring, a driving device, etc. (not shown).
[0019]
The CPU 201 is a processor that performs a predetermined calculation according to a program stored in the ROM 203, the flash memory 204, etc., and also has a function of controlling each part of the earphone control device 20. The RAM 202 is composed of a volatile storage medium and provides a temporary memory area necessary for the operation of the CPU 201. The ROM 203 is composed of a non-volatile storage medium and stores necessary information such as a program used for the operation of the earphone control device 20. The flash memory 204 is a storage device composed of a non-volatile storage medium, which temporarily stores data, stores an operation program of the earphone control device 20, and the like.
[0020]
Communication I / F207 is a communication interface based on standards such as Bluetooth (registered trademark) and Wi-Fi (registered trademark), and is a module for communicating with the information communication device 1.
[0021]
The speaker I / F 205 is an interface for driving the speaker 26. The speaker I / F 205 includes a digital-analog conversion circuit, an amplifier and the like. The speaker I / F 205 converts audio data into an analog signal and supplies it to the speaker 26. As a result, the speaker 26 emits a sound wave based on the voice data.
[0022]
The microphone I / F206 is an interface for acquiring a signal from the microphone 27. The microphone I / F 206 includes an analog-digital conversion circuit, an amplifier and the like. The microphone I / F 206 converts an analog signal generated by a sound wave received by the microphone 27 into a digital signal. As a result, the earphone control device 20 acquires voice data based on the received sound wave.
[0023]
The battery 208 is, for example, a secondary battery and supplies the power required for the operation of the earphone 2. As a result, the earphone 2 can operate wirelessly without being connected to an external power source by wire.
[0024]
The hardware configuration shown in FIG. 2 is an example, and devices other than these may be added or some devices may not be provided. Further, some devices may be replaced with other devices having similar functions. For example, the earphone 2 may further include an input device such as a button so that the operation by the user 3 can be received, and further includes a display device such as a display and an indicator lamp for providing information to the user 3. You may. As described above, the hardware configuration shown in FIG. 2 can be changed as appropriate.
[0025]
FIG. 3 is a block diagram showing a hardware configuration example of the information communication device 1. The information communication device 1 includes a CPU 101, a RAM 102, a ROM 103, and an HDD (Hard Disk Drive) 104. Further, the information communication device 1 includes a communication I / F 105, an input device 106, and an output device 107. Each part of the information communication device 1 is connected to each other via a bus, wiring, a driving device, etc. (not shown).
[0026]
In FIG. 3, each part constituting the information communication device 1 is shown as an integrated device, but some of these functions may be provided by an external device. For example, the input device 106 and the output device 107 may be external devices different from the parts constituting the functions of the computer including the CPU 101 and the like.
[0027]
The CPU 101 is a processor that performs predetermined calculations according to programs stored in the ROM 103, HDD 104, etc., and also has a function of controlling each part of the information communication device 1. The RAM 102 is composed of a volatile storage medium and provides a temporary memory area necessary for the operation of the CPU 101. The ROM 103 is composed of a non-volatile storage medium and stores necessary information such as a program used for the operation of the information communication device 1. The HDD 104 is a storage device composed of a non-volatile storage medium, which temporarily stores data transmitted to and received from the earphone 2, stores an operation program of the information communication device 1, and the like.
[0028]
Communication I / F105 is a communication interface based on standards such as Bluetooth (registered trademark) and Wi-Fi (registered trademark), and is a module for communicating with other devices such as earphone 2.
[0029]
The input device 106 is a keyboard, a pointing device, or the like, and is used by the user 3 to operate the information communication device 1. Examples of pointing devices include mice, trackballs, touch panels, pen tablets and the like.
[0030]
The output device 107 is, for example, a display device. The display device is a liquid crystal display, an OLED (Organic Light Emitting Diode) display, etc., and is used for displaying information, a GUI (Graphical User Interface) for operation input, and the like. The input device 106 and the output device 107 may be integrally formed as a touch panel.
[0031]
The hardware configuration shown in FIG. 3 is an example, and devices other than these may be added or some devices may not be provided. Further, some devices may be replaced with other devices having similar functions. Further, a part of the functions of the present embodiment may be provided by another device via a network, or the functions of the present embodiment may be distributed and realized by a plurality of devices. For example, the HDD 104 may be replaced with an SSD (Solid State Drive) using a semiconductor memory, or may be replaced with a cloud storage. As described above, the hardware configuration shown in FIG. 3 can be changed as appropriate.
[0032]
FIG. 4 is a functional block diagram of the earphone 2 and the information communication device 1 according to the present embodiment. The information communication device 1 includes an acoustic characteristic acquisition unit 121, a feature amount extraction unit 122, and a determination unit 123. Since the configuration of the block diagram of the earphone 2 is the same as that of FIG. 2, the description thereof will be omitted. The acoustic characteristic acquisition unit 121 may be more generally referred to as an acquisition unit, and the feature amount extraction unit 122 may be more generally referred to as an extraction unit.
[0033]
The CPU 101 performs a predetermined arithmetic process by loading the program stored in the ROM 103, the HDD 104, etc. into the RAM 102 and executing the program. Further, the CPU 101 controls each part of the information communication device 1 such as the communication I / F 105 based on the program. As a result, the CPU 101 realizes the functions of the acoustic characteristic acquisition unit 121, the feature amount extraction unit 122, and the determination unit 123. The specific contents of the processing performed by each functional block will be described later.
[0034]
Note that, in FIG. 4, some or all of the functions of the functional blocks described in the information communication device 1 may be provided in the earphone control device 20 instead of the information communication device 1. That is, each of the above-mentioned functions may be realized by the information communication device 1, the earphone control device 20, or the information communication device 1 and the earphone control device 20 in cooperation with each other. good. The information communication device 1 and the earphone control device 20 may be more generally referred to as an information processing device. In the following description, unless otherwise specified, as shown in FIG. 4, each functional block related to acquisition and determination of acoustic information is assumed to be provided in the information communication device 1.
[0035]
FIG. 5 is a flowchart showing an outline of the biometric authentication process performed by the information communication device 1 according to the present embodiment. The operation of the information communication device 1 will be described with reference to FIG.
[0036]
The biometric authentication process of FIG. 5 is executed, for example, when the user 3 starts using the earphone 2 by operating the earphone 2. Alternatively, the biometric authentication process of FIG. 5 may be executed every time a predetermined time elapses when the power of the earphone 2 is on.
[0037]
In step S101, the acoustic characteristic acquisition unit 121 gives an instruction to the earphone control device 20 to emit an inspection sound. The earphone control device 20 transmits an inspection signal to the speaker 26, and the speaker 26 emits an inspection sound generated based on the inspection signal to the external auditory canal of the user 3.
[0038]
As the inspection signal, a signal containing a predetermined range of frequency components such as a charp signal, an M series (Maximum Length Sequence) signal, white noise, and an impulse signal can be used. This makes it possible to acquire an acoustic signal including information on frequencies within a predetermined range. The inspection sound may be an audible sound whose frequency and sound pressure level are within the audible range. In this case, by making the user 3 perceive the sound wave at the time of authentication, it is possible to notify the user 3 that the authentication is being performed. Further, the inspection sound may be an inaudible sound whose frequency or sound pressure level is out of the audible range. In this case, the sound wave can be less likely to be perceived by the user 3, and the comfort at the time of use is improved.
[0039]
In step S102, the microphone 27 receives the reverberation sound (ear sound) in the external auditory canal or the like and converts it into an electric signal in the time domain. This electrical signal is sometimes called an acoustic signal. The microphone 27 transmits an acoustic signal to the earphone control device 20, and the earphone control device 20 transmits an acoustic signal to the information communication device 1.
[0040]
In step S103, the acoustic characteristic acquisition unit 121 acquires the acoustic characteristic of the frequency domain based on the sound wave propagating in the user's head. This acoustic characteristic can be, for example, a frequency spectrum obtained by converting an acoustic signal in the time domain into a frequency domain using an algorithm such as fast Fourier transform. The acquired acoustic characteristics are stored in the HDD 104.
[0041]
In step S104, the feature amount extraction unit 122 extracts the feature amount from the acoustic characteristics. The specific content of this process will be described later. The extracted feature amount is stored in the HDD 104 and used for biometric authentication.
[0042]
In step S105, the determination unit 123 determines whether or not the user 3 is a registrant by collating the feature amount extracted by the feature amount extraction unit 122 with the feature amount of the registrant previously recorded in the HDD 104. Is determined. If it is determined that the user 3 is a registrant (YES in step S106), the process proceeds to step S107. When it is determined that the user 3 is not a registrant (NO in step S106), the process proceeds to step S108.
[0043]
In step S107, the information communication device 1 transmits a control signal indicating that the user 3 is permitted to use the earphone 2 to the earphone 2. As a result, the earphone 2 is ready for use by the user 3.
[0044]
In step S108, the information communication device 1 transmits a control signal indicating that the use of the earphone 2 by the user 3 is not permitted to the earphone 2. As a result, the earphone 2 becomes unusable by the user 3. The unusable state may be, for example, a state in which no sound is emitted from the speaker 26 of the earphone 2. The control in steps S107 and S108 may not control the earphone 2 side, but may control the information communication device 1 side. For example, the use permission state and the use disapproval state may be switched by changing the communication connection state between the information communication device 1 and the earphone 2.
[0045]
In step S105, the determination unit 123 may further determine whether or not the earphone 2 is attached to the ear of the user 3 based on the acoustic characteristics or the feature amount. When it is determined that the user 3 is not attached to the ear, a process of disallowing the use of the earphone 2 may be performed as in the case where it is determined that the user 3 is not a registrant.
[0046]
FIG. 6 is a flowchart showing a feature amount extraction process performed by the information communication device 1 according to the present embodiment. The feature amount extraction process in step S104 of FIG. 5 will be described in more detail with reference to FIG.
[0047]
In step S111, the feature amount extraction unit 122 calculates the frequency response function of the rational polynomial model based on the acoustic characteristics acquired by the acoustic characteristic acquisition unit 121. The acoustic characteristics and the frequency response function of the rational polynomial model will be described. The frequency response function of the rational polynomial model is sometimes called the first frequency response function.
[0048]
FIG. 7 is a graph showing a measurement example of acoustic characteristics acquired by the acoustic characteristic acquisition unit 121. The horizontal axis of FIG. 7 shows the frequency, and the vertical axis shows the signal strength in an arbitrary unit. The vertical axis is displayed on a logarithmic scale. This intensity indicates, for example, the intensity of the sound wave received by the microphone 27. In the following, it is assumed that the acoustic characteristics are obtained by standardizing the intensity of the sound wave received by the microphone 27 by the intensity of the sound wave emitted from the speaker 26.
[0049]
As shown in FIG. 7, the acquired acoustic characteristics have a plurality of peaks (maximum points) P1, P2, P3 and a plurality of notches (minimum points) N1, N2. Peaks P1, P2, P3 and notches N1 and N2 indicate resonances that occur in the air column formed by the user 3's external auditory canal, tympanic membrane and earphone 2. Since a plurality of peaks P1, P2, P3 and notches N1 and N2 are observed, it can be seen that there are a plurality of resonance modes. For example, each of the peaks P1, P2, and P3 shows a resonance mode in which the amplitude of the sound wave becomes large at the position of the microphone 27 of the earphone 2. In addition to the resonance in the air column described above, a resonance mode in the earphone 2 can also be observed. Since the properties of the peaks P1, P2, P3 and the notches N1 and N2 depend on the shape of the external auditory canal of the user 3, it is effective for extracting the feature amount for identifying an individual.
[0050]
Such acoustic characteristics can be expressed by the frequency response function based on the rational polynomial model shown in the following equation (1).
[Number 1]

[0051]
Here, ω is the angular frequency. Ω l (ω) is a base function in polynomial expansion, and this base function does not depend on the subject whose acoustic characteristics have been acquired. Note that Ω l (ω) is a complex function. α l and β l are real coefficients of the polynomial, and this term differs depending on the subject. That is, α l and β l are feature quantities indicating the characteristics of the subject. N m is the number of modes assumed in the rational polymorphic model, and l is an argument indicating each mode.
[0052]
The denominator of equation (1) becomes zero at the frequency where the peak occurs, and the molecule of equation (1) becomes zero at the frequency where the notch occurs. Thus, in the frequency response function based on the rational polynomial model, the peak feature is considered in the denominator and the notch feature is considered in the molecule. The feature amount extraction unit 122 determines the coefficients α l and β l so as to approximate the acquired acoustic characteristics by the equation (1), and estimates the frequency response function. For this approximation, an approximate solution method such as a minimum square method or a maximum probability estimation method, which includes an algorithm such as minimization of an error function, can be used.
[0053]
In step S112, the feature amount extraction unit 122 outputs the coefficients α l and β l of the frequency response function by the rational polynomial model as the feature amount indicating the user 3. In the output of this feature amount, the coefficients α l and β l are extracted for a plurality of N m different from each other, and the coefficients α l and β l obtained in the plurality of N m are combined and extracted as a feature amount. May be good. Thereby, the properties of the peak and the notch that can be stably obtained at a plurality of N m can be reflected in the feature amount.
[0054]
The coefficients α l and β l are feature quantities that are robust against differences in wearing conditions because they do not require processing to calculate the logarithm at the time of derivation. Details of this will be described later.
[0055]
The effect of the above-mentioned feature quantity extraction method will be explained. As mentioned above, the acoustic properties include peaks and notches due to resonances that occur in the air column formed by the user 3's external auditory canal, tympanic membrane and earphone 2. Since the frequencies and magnitudes of these peaks and notches depend on the shape of the air column described above, they may change depending on the position of the earphone 2. Therefore, the acoustic characteristics may change depending on the wearing state of the earphone 2.
[0056]
FIG. 8 is a graph showing an example of changes in acoustic characteristics when the earphone 2 is repeatedly attached and detached. FIG. 8 shows that the wearing state is changed by repeating the attachment / detachment to acquire the acoustic characteristics five times, and the acoustic characteristics are superimposed and displayed. Focusing on the peaks P4, P5 and notch N3 in FIG. 8, it can be seen that the variation of the notch N3 is larger than the variation of the peaks P4 and P5.To.
[0057]
The reason why the variation of the notch N3 is larger than the variation of the peaks P4 and P5 will be explained with a simple example. It is assumed that there are only two resonance modes existing in the external auditory canal, and the acoustic characteristics caused by the two resonance modes can be separated into S 1 (f) and S 2 (f). Then, it is assumed that S 1 (f) has a characteristic that a peak occurs at the natural frequency f 1 and S 2 (f) has a characteristic that a peak occurs at the natural frequency f 2. At this time, at the frequency between f 1 and f 2, the acoustic characteristic Si (f) considering the interference between the two modes satisfies the following equation (2), where θ is the phase difference between the modes.
[Number 2]

[0058]
When the wearing state of the earphone 2 changes, | S i (f) | also changes due to the change of the three parameters | S 1 (f) |, | S 2 (f) | and θ. Here, since the influence of S 2 (f) is sufficiently small at the natural frequency f 1, it can be approximated as the following equation (3).
[Number 3]

[0059]
Further, since the influence of S 1 (f) is sufficiently small at the natural frequency f 2, it can be approximated as the following equation (4).
[Number 4]

[0060]
From equations (3) and (4), terms containing θ can be ignored in the vicinity of the natural frequency. Therefore, it can be seen that in the vicinity of the natural frequency, although it is affected by the change in the absolute value (power) due to the change in the mounting state, it is not easily affected by the change in the phase difference due to the change in the mounting state. This corresponds to the relatively small fluctuation in the vicinity of the peaks P4 and P5 in FIG. On the other hand, since the above approximation does not hold at frequencies near the middle of f 1 and f 2, both changes in absolute value and changes in phase difference can be factors that change the acoustic characteristics. This corresponds to the relatively large fluctuation in the vicinity of the notch N3 in FIG.
[0061]
As described in Patent Document 1, conventionally, the Mercepstram coefficient may be used as a feature quantity for certification. The process of deriving the Melspectrum coefficient may include a process of calculating the logarithmic characteristic. Converting acoustic properties to a logarithmic scale emphasizes areas of low intensity.
[0062]
As described above, in the vicinity of the notch, the phase difference θ between adjacent modes is easily affected, so that the fluctuation due to the change in the mounting state is remarkable. In addition to this, the conversion to a logarithmic scale emphasizes the effect of fluctuations near the notch where the intensity is minimal. Due to the superposition of these two effects, the feature quantity extraction using the Melkeptrum coefficient, which includes the process of calculating the logarithm in the derivation process, may be easily affected by the change in the wearing state.
[0063]
On the other hand, the feature amount extraction method of the present embodiment does not require the process of calculating the logarithm at the time of derivation, and is therefore robust against the difference in the mounting state. For the above reasons, in the present embodiment, while paying attention to the features of the notch and the peak, it is possible to extract a feature amount that is robust against the difference in the wearing state of the earphone 2.
[0064]
In this embodiment, the feature amount is extracted from the frequency response function by the rational polynomial model. It will be explained that this process is a feature amount extraction method focusing on the features of notches and peaks. The frequency response function based on the rational polynomial model of equation (1) can be converted into the frequency response function based on the pole / remainder model shown in equation (5) below. This transformation corresponds to the partial fractional expansion of the frequency response function by the rational polynomial model. The frequency response function of the pole / residence model is sometimes called a second frequency response function.
[Number 5]

[0065]
Here, λ k is a complex number indicating the pole of the acoustic characteristic, and R k is a complex number indicating the shape of the pole. λ k and R k are sometimes called poles and distillates, respectively. Note that j is an imaginary unit, C is a remainder term, and "*" is a symbol indicating complex conjugate.
[0066]
Physically, λ k includes information on the natural frequency f k of the pole and the attenuation ratio ζ k of the pole. The natural frequency f k and the pole attenuation ratio ζ k are expressed by the following equations (6) and (7), respectively. Note that Re (λ k) is a real part of λ k, and Im (λ k) is an imaginary part of λ k.
[Number 6]

[Number 7]

[0067]
At the natural frequency, the imaginary part of the term of jω-λ k in the parentheses of equation (5) becomes zero, and the real part of the term of jω-λ k becomes a constant determined by the attenuation ratio ζ k. Therefore, the term in parentheses in Eq. (5) is a function of dampening vibration representing the kth peak. In other words, Eq. (5) expresses the frequency response function by the sum of Nm decay vibrations.
[0068]
In this way, the frequency response function based on the rational polynomial model in Eq. (1) can be converted into the frequency response function based on the pole / distillate model constructed by the sum of decay vibrations. Therefore, the frequency response function based on the rational polynomial model substantially includes the characteristics of the resonance mode, and is suitable for feature quantity extraction focusing on the peak and notch characteristics of the resonance mode.
[0069]
Note that the feature quantity may be extracted from the terms included in the frequency response function of the pole / remainder model instead of the rational polynomial model. However, since each term of the pole / retention model changes depending on the subject, it is difficult to determine the term to be compared when collating the feature quantities, whereas the base function of the rational polynomial model is Since it does not depend on the user, there is an advantage that it is easy to make a comparison when collating feature quantities. Therefore, it is desirable to extract the feature quantity from the term included in the frequency response function of the rational polynomial model.
[0070]
[Second embodiment]
The information processing system of this embodiment is different from the first embodiment in the content of the feature quantity extraction process, but the other parts are the same as those of the first embodiment. Hereinafter, the differences from the first embodiment will be mainly described, and the description of the common parts will be omitted or simplified.
[0071]
FIG. 9 is a flowchart showing a feature amount extraction process performed by the information communication device 1 according to the present embodiment. The present embodiment differs from the first embodiment in that a process of converting a rational polynomial model into a pole / remaining number model and excluding some terms is added.
[0072]
In step S111, the feature amount extraction unit 122 calculates the frequency response function of the rational polynomial model as in the first embodiment.
[0073]
In step S121, the feature quantity extraction unit 122 converts the frequency response function of the rational polynomial model into the frequency response function of the pole / remainder model. Since the content of this process is as described in the first embodiment using the equations (1), (5) and the like, the description thereof will be omitted.
[0074]
In step S122, the feature amount extraction unit 122 determines an insignificant term from each term of the frequency response function of the pole / retention model. The insignificant term is a term in which it is clear that the dampening vibration based on the physical phenomenon of resonance in the external auditory canal is not properly expressed. Specifically, there are terms that clearly contradict the physical phenomenon of resonance in the external auditory canal, such as terms in which the attenuation ratio ζ k is negative and terms in which the natural frequency f k is outside the frequency range in which resonance can occur. Will be.
[0075]
In step S123, the feature amount extraction unit 122 determines an unstable term from each term in the frequency response function of the pole / distillate model. The unstable term is a term that does not appear stably when the process of approximating the acoustic characteristics with the frequency response function is performed several times by changing the calculation conditions such as the approximation conditions. For example, a term that does not have the same natural frequency and the same attenuation ratio when the number N m of modes considered in the model as shown in Eq. (1) or Eq. (5) is changed is an unstable term. .. As a specific example of the method for determining an unstable term, a mode in which the amount of change in the natural frequency or the attenuation ratio exceeds a predetermined error range depending on whether the number of modes to be considered is Na or Na + 1. There is a method of determining that the mode is unstable. Such an unstable term does not represent the physical phenomenon of resonance in the external auditory canal, but is likely to be a pseudo one caused by a factor in arithmetic processing.
[0076]
In step S124, the feature quantity extraction unit 122 excludes some terms from the frequency response function of the pole / residue model and converts it into the frequency response function of the rational polynomial model. Here, the "partial term" may be an insignificant term extracted in step S122 and an unstable term extracted in step S123. This excludes terms that do not properly represent physical phenomena. Further, since the surplus term C does not need to be extracted as a feature quantity indicating the feature of the user 3, the "partial term" may include the surplus term C. The process of excluding some of the terms may be, specifically, a process of changing the value of the distillate R k or the remainder term C of the corresponding term to zero.
[0077]
In step S112, the feature amount extraction unit 122 outputs the coefficient of the frequency response function of the rational polynomial model obtained in step S124 as the feature amount by the same processing as in the first embodiment.
[0078]
According to the method of the present embodiment, it is possible to extract a feature amount that is robust against the difference in the wearing state as in the first embodiment. In addition, since the feature amount is extracted after excluding insignificant terms, unstable terms, etc., it is possible to extract the feature amount that more appropriately reflects the biological features such as the shape of the external auditory canal of the user 3. can.
[0079]
[Third embodiment]
The information processing system of this embodiment is different from the first and second embodiments in the content of the feature quantity extraction process, but the other parts are the same as those of the first and second embodiments. Hereinafter, the differences from the second embodiment will be mainly described, and the description of the common parts will be omitted or simplified.
[0080] [0080]
FIG. 10 is a flowchart showing a feature amount extraction process performed by the information communication device 1 according to the present embodiment. In this embodiment, a process of reconstructing the spectrum from the frequency response function of the rational polynomial model obtained by excluding some terms to calculate the Mel-Frequency Cepstrum Coefficient (MFCC) is added. The point is different from the second embodiment.
[0081]
Since the processes from step S111 to step S124 are the same as those in the second embodiment, the description thereof will be omitted.
[0082]
In step S131, the feature amount extraction unit 122 reconstructs the spectrum of acoustic characteristics from the frequency response function of the rational polynomial model generated in step S124.
[0083]
In step S132, the feature amount extraction unit 122 calculates the mel frequency Kepstram coefficient from the reconstructed spectrum. Hereinafter, an example of the calculation method of the Mel frequency Kepstram coefficient will be described.
[0084]
First, the feature amount extraction unit 122 applies the mel filter bank to the reconstructed spectrum. The mel filter bank is a filter bank containing a plurality of triangular window functions. The plurality of triangular window functions are configured so that the higher the frequency, the wider the frequency width, based on the Mel scale considering the human auditory characteristics. The spectrum obtained by applying the mel filter bank is called the mel spectrum. Further, the mel spectrum in which the intensity (vertical axis) of the mel spectrum is converted into a logarithmic scale is called a mel log spectrum. By converting to a logarithmic scale, a high-order resonance phenomenon with a small gain can be well expressed.
[0085]
Next, the feature quantity extraction unit 122 performs discrete cosine conversion on the mel logarithmic spectrum to calculate the Cepstrum of the quefrency domain. The Mel frequency Keptram coefficient is calculated by extracting the coefficient of the order term in the predetermined range of the Kepstram. In this method, it is possible to obtain features that are efficiently compressed by weighting in consideration of human auditory characteristics.
[0086]
In step S133, the feature amount extraction unit 122 outputs the mel frequency kepstram coefficient acquired by the process of step S132 as the feature amount.
[0087]
Method of this embodiment According to the second embodiment, it is possible to extract a feature amount that more appropriately expresses the features of the living body such as the shape of the external auditory canal of the user 3. In addition, since weighting is performed in consideration of human auditory characteristics in the process of calculating the Mel frequency Kepstram coefficient, it is possible to extract features with efficient compression.
[0088]
[Fourth embodiment]
The information processing system of the present embodiment is different from the first to third embodiments in the content of the feature quantity extraction process, but the other parts are the same as those of the first to third embodiments. Hereinafter, the differences from the third embodiment will be mainly described, and the description of the common parts will be omitted or simplified.
[0089]
FIG. 11 is a flowchart showing a feature amount extraction process performed by the information communication device 1 according to the present embodiment. The present embodiment differs from the second or third embodiment in that both the coefficient of the frequency response function of the rational polynomial model and the mel frequency Keptram coefficient are extracted as feature quantities.
[0090]
Since the processes from step S111 to step S124 are the same as those in the second embodiment, the description thereof will be omitted.
[0091]
In step S141, the feature amount extraction unit 122 outputs the coefficient of the frequency response function of the rational polynomial model obtained in step S124 as the first feature amount by the same processing as in the first or second embodiment.
[0092]
Since the processes of steps S131 and S132 are the same as those of the third embodiment, the description thereof will be omitted.
[0093]
In step S142, the feature amount extraction unit 122 outputs the mel frequency kepstram coefficient acquired by the process of step S132 as the second feature amount.
[0094]
In the determination process after extracting the feature amount (step S105 in FIG. 5), it is determined whether or not the user 3 is a registrant based on both the first feature amount and the second feature amount. Specifically, the first score obtained by the collation using the first feature amount and the second score obtained by the collation using the second feature amount are weighted and added at a predetermined ratio, and after the addition. A method of determining whether or not the user 3 is a registrant based on the score can be used. Further, the collation using the first feature amount and the collation using the second feature amount may be performed separately, and the logical sum or the logical product of the two collation results may be used as the final collation result.
[0095]
According to the method of the present embodiment, since the feature amount extraction method according to the second embodiment and the feature amount extraction method according to the third embodiment are used in combination, both effects can be obtained. Further, in the present embodiment, since the first feature amount obtained by extracting the characteristics of the peak and the notch and the second feature amount obtained mainly by extracting the characteristics of the waveform of the acoustic characteristic can be used in combination, the determination can be made more multifaceted. Highly accurate biometric authentication can be realized in consideration of various information.
[0096]
The system described in the above-described embodiment can also be configured as in the following fifth embodiment.
[0097]
[Fifth embodiment]
FIG. 12 is a functional block diagram of the information processing apparatus 4 according to the fifth embodiment. The information processing apparatus 4 includes an acquisition unit 421 and an extraction unit 422. The acquisition unit 421 acquires the acoustic characteristics of the frequency domain based on the sound wave propagating in the user's head. The extraction unit 422 generates a first frequency response function having a rational polypoly in which a term indicating the characteristic of the peak of the acoustic characteristic is included in the denominator, and the feature used for bioauthentication of the user based on the first frequency response function. Extract the amount.
[0098]
According to the present embodiment, there is provided an information processing apparatus 4 capable of extracting a feature amount robust against a difference in wearing state.
[0099]
[Modified embodiment]
The present invention is not limited to the above-described embodiment, and can be appropriately modified without departing from the spirit of the present invention. For example, an example in which a partial configuration of any one embodiment is added to another embodiment or an example in which a partial configuration of another embodiment is replaced with another embodiment is also an embodiment of the present invention.
[0100]
In the above-described embodiment, the earphone 2 is illustrated as an example of a wearable device, but the earphone is not limited to the earphone mounted in the external ear canal as long as the acoustic information necessary for processing can be acquired. For example, the wearable device may be headphones that cover the entire ear, or may be a bone conduction type acoustic device that indirectly transmits sound waves from the head other than the external ear canal. Further, the wearable device may be composed of two earphones to be worn on both ears, in which case the biometric authentication of the above-described embodiment may be performed on both ears, and the above-mentioned only on one ear. The biometric authentication of the embodiment may be performed.
[0101]
A processing method in which a program for operating the configuration of the embodiment is recorded in a storage medium so as to realize the functions of the above-described embodiment, the program recorded in the storage medium is read out as a code, and the program is executed in a computer is also described in each embodiment. Included in the category. That is, a computer-readable storage medium is also included in the scope of each embodiment. Further, not only the storage medium in which the above-mentioned program is recorded but also the program itself is included in each embodiment. Further, one or more components included in the above-described embodiment are circuits such as ASIC (Application Specific Integrated Circuit) and FPGA (Field Programmable Gate Array) configured to realize the functions of each component. There may be.
[0102]
As the storage medium, for example, a floppy (registered trademark) disk, a hard disk, an optical disk, a photomagnetic disk, a CD (Compact Disk) -ROM, a magnetic tape, a non-volatile memory card, or a ROM can be used. In addition, the program recorded on the storage medium is not limited to the one that executes the processing by itself, but the one that operates on the OS (Operating System) and executes the processing in cooperation with other software and the function of the expansion board. Is also included in the category of each embodiment.
[0103]
The service realized by the functions of each of the above-described embodiments can also be provided to the user in the form of SaaS (Software as a Service).
[0104]
It should be noted that the above-mentioned embodiments are merely examples of embodiment in carrying out the present invention, and the technical scope of the present invention should not be construed in a limited manner by these. That is, the present invention can be implemented in various forms without departing from its technical idea or its main features.
[0105]
A part or all of the above-described embodiment may be described as in the following appendix, but is not limited to the following.
[0106]
(Appendix 1)
The acquisition unit that acquires the acoustic characteristics of the frequency domain based on the sound wave propagating through the user's head,
A term indicating the characteristics of the peak of the acoustic characteristics is included in the denominator, and a term indicating the characteristics of the notch of the acoustic characteristics is contained in the molecule to generate a first frequency response function having a rational polypoly. 1 An extraction unit that extracts features used for biometric authentication of the user based on a frequency response function, and an extraction unit.
Information processing device equipped with.
[0107]
(Appendix 2)
The extraction unit extracts the coefficient of the rational polypoly as the feature quantity.
The information processing device described in Appendix 1.
[0108]
(Appendix 3)
The denominator of the first frequency response function is zero at the frequency corresponding to the peak.
The information processing device according to Appendix 1 or 2.
[0109]
(Appendix 4)
The acoustic characteristics have a plurality of the peaks,
The extraction unit generates a second frequency response function including the sum of a plurality of fractions.
The information processing device according to any one of Supplementary note 1 to 3.
[0110]
(Appendix 5)
The imaginary part of each denominator of the plurality of fractions included in the second frequency response function is zero at the frequency corresponding to one of the plurality of the peaks.
The information processing device described in Appendix 4.
[0111]
(Appendix 6)
The first frequency response function and the second frequency response function can be converted to each other.
The information processing device according to Appendix 4 or 5.
[0112]
(Appendix 7)
The second frequency response function includes a partial fractional expansion of the first frequency response function.
The information processing device described in Appendix 6.
[0113]
(Appendix 8)
The extraction unit generates the first frequency response function by excluding some terms from the plurality of fractions included in the second frequency response function and then converting the term into the first frequency response function.
The information processing device according to any one of Supplementary note 4 to 7.
[0114]
(Appendix 9)
The extraction unit generates the first frequency response function by excluding a physically insignificant term from the plurality of fractions included in the second frequency response function and then converting the term into the first frequency response function. do,
The information processing device according to any one of Supplementary note 4 to 8.
[0115]
(Appendix 10)
The extraction unit removes terms unstable with respect to changes in calculation conditions from the plurality of fractions included in the second frequency response function, and then converts the first frequency response function into the first frequency response function. Generate a frequency response function,
The information processing device according to any one of Supplementary note 4 to 9.
[0116]
(Appendix 11)
The extraction unit generates a spectrum of a frequency domain using the first frequency response function converted from the second frequency response function.
The information processing device according to any one of Supplementary note 8 to 10.
[0117]
(Appendix 12)
The extraction unit extracts the feature amount based on the spectrum.
The information processing device according to Appendix 11.
[0118]
(Appendix 13)
The extraction unit extracts the first feature amount based on the coefficient of the first frequency response function, and extracts the second feature amount based on the spectrum.
The information processing device according to Appendix 11 or 12.
[0119]
(Appendix 14)
The acoustic characteristics are based on sound waves propagating in the external auditory canal of the user acquired by a wearable device worn on the user's head.
The information processing device according to any one of Supplementary note 1 to 13.
[0120]
(Appendix 15)
The step of acquiring the acoustic characteristics of the frequency domain based on the sound wave propagating in the user's head,
A term indicating the characteristics of the peak of the acoustic characteristics is included in the denominator, and a term indicating the characteristics of the notch of the acoustic characteristics is contained in the molecule to generate a first frequency response function having a rational polypoly. 1 The step of extracting the feature amount used for the biometric authentication of the user based on the frequency response function, and
Information processing method with.
[0121]
(Appendix 16)
On the computer
The step of acquiring the acoustic characteristics of the frequency domain based on the sound wave propagating in the user's head,
A term indicating the characteristics of the peak of the acoustic characteristics is included in the denominator, and a term indicating the characteristics of the notch of the acoustic characteristics is contained in the molecule to generate a first frequency response function having a rational polypoly. 1 The step of extracting the feature amount used for the biometric authentication of the user based on the frequency response function, and
A storage medium in which a program for executing an information processing method is stored.
Description of the sign
[0122]
1 Information and communication equipment
2 Earphones
3 user
4 Information processing device
20 Earphone control device
26 speaker
27 microphone
101, 201 CPU
102, 202 RAM
103, 203 ROM
104 HDD
105, 207 Communication I / F
106 Input device
107 Output device
121 Acoustic characteristics acquisition department
122 Feature amount extraction unit
123 Judgment unit
204 Flash memory
205 Speaker I / F
206 Microphone I / F
208 battery
421 Acquisition Department
422 Extraction section
The scope of the claims
[Claim 1]
Frequency based on sound waves propagating in the user's head The acquisition unit that acquires the acoustic characteristics of the domain,
A term indicating the characteristics of the peak of the acoustic characteristics is included in the denominator, and a term indicating the characteristics of the notch of the acoustic characteristics is contained in the molecule to generate a first frequency response function having a rational polypoly. 1 An extraction unit that extracts features used for biometric authentication of the user based on a frequency response function, and an extraction unit.
Information processing device equipped with.
[Claim 2]
The extraction unit extracts the coefficient of the rational polypoly as the feature quantity.
The information processing device according to claim 1.
[Claim 3]
The denominator of the first frequency response function is zero at the frequency corresponding to the peak.
The information processing device according to claim 1 or 2.
[Claim 4]
The acoustic characteristics have a plurality of the peaks,
The extraction unit generates a second frequency response function including the sum of a plurality of fractions.
The information processing device according to any one of claims 1 to 3.
[Claim 5]
The imaginary part of each denominator of the plurality of fractions included in the second frequency response function is zero at the frequency corresponding to one of the plurality of the peaks.
The information processing device according to claim 4.
[Claim 6]
The first frequency response function and the second frequency response function can be converted to each other.
The information processing device according to claim 4 or 5.
[Claim 7]
The second frequency response function includes a partial fractional expansion of the first frequency response function.
The information processing device according to claim 6.
[Claim 8]
The extraction unit generates the first frequency response function by excluding some terms from the plurality of fractions included in the second frequency response function and then converting the term into the first frequency response function.
The information processing device according to any one of claims 4 to 7.
[Claim 9]
The extraction unit generates the first frequency response function by excluding a physically insignificant term from the plurality of fractions included in the second frequency response function and then converting the term into the first frequency response function. do,
The information processing device according to any one of claims 4 to 8.
[Claim 10]
The extraction unit removes terms unstable with respect to changes in calculation conditions from the plurality of fractions included in the second frequency response function, and then converts the first frequency response function into the first frequency response function. Generate a frequency response function,
The information processing device according to any one of claims 4 to 9.
[Claim 11]
The extraction unit generates a spectrum of a frequency domain using the first frequency response function converted from the second frequency response function.
The information processing apparatus according to any one of claims 8 to 10.
[Claim 12]
The extraction unit extracts the feature amount based on the spectrum.
The information processing device according to claim 11.
[Claim 13]
The extraction unit extracts the first feature amount based on the coefficient of the first frequency response function, and extracts the second feature amount based on the spectrum.
The information processing device according to claim 11 or 12.
[Claim 14]
The acoustic characteristics are based on sound waves propagating in the external auditory canal of the user acquired by a wearable device worn on the user's head.
The information processing device according to any one of claims 1 to 13.
[Claim 15]
The step of acquiring the acoustic characteristics of the frequency domain based on the sound wave propagating in the user's head,
A term indicating the characteristics of the peak of the acoustic characteristics is included in the denominator, and a term indicating the characteristics of the notch of the acoustic characteristics is contained in the molecule to generate a first frequency response function having a rational polypoly. 1 The step of extracting the feature amount used for the biometric authentication of the user based on the frequency response function, and
Information processing method with.
[Claim 16]
On the computer
The step of acquiring the acoustic characteristics of the frequency domain based on the sound wave propagating in the user's head,
A term indicating the characteristics of the peak of the acoustic characteristics is included in the denominator, and a term indicating the characteristics of the notch of the acoustic characteristics is contained in the molecule to generate a first frequency response function having a rational polypoly. 1 The step of extracting the feature amount used for the biometric authentication of the user based on the frequency response function, and
A storage medium in which a program for executing an information processing method is stored.

Documents

Application Documents

# Name Date
1 202217012204.pdf 2022-03-07
2 202217012204-STATEMENT OF UNDERTAKING (FORM 3) [07-03-2022(online)].pdf 2022-03-07
3 202217012204-REQUEST FOR EXAMINATION (FORM-18) [07-03-2022(online)].pdf 2022-03-07
4 202217012204-POWER OF AUTHORITY [07-03-2022(online)].pdf 2022-03-07
5 202217012204-NOTIFICATION OF INT. APPLN. NO. & FILING DATE (PCT-RO-105-PCT Pamphlet) [07-03-2022(online)].pdf 2022-03-07
6 202217012204-FORM 18 [07-03-2022(online)].pdf 2022-03-07
7 202217012204-FORM 1 [07-03-2022(online)].pdf 2022-03-07
8 202217012204-DRAWINGS [07-03-2022(online)].pdf 2022-03-07
9 202217012204-DECLARATION OF INVENTORSHIP (FORM 5) [07-03-2022(online)].pdf 2022-03-07
10 202217012204-COMPLETE SPECIFICATION [07-03-2022(online)].pdf 2022-03-07
11 202217012204-MARKED COPIES OF AMENDEMENTS [15-03-2022(online)].pdf 2022-03-15
12 202217012204-FORM 13 [15-03-2022(online)].pdf 2022-03-15
13 202217012204-AMMENDED DOCUMENTS [15-03-2022(online)].pdf 2022-03-15
14 202217012204-FORM 3 [16-05-2022(online)].pdf 2022-05-16
15 202217012204-FER.pdf 2022-07-29
16 202217012204-Proof of Right [01-11-2022(online)].pdf 2022-11-01
17 202217012204-FORM-26 [24-01-2023(online)].pdf 2023-01-24
18 202217012204-FORM 3 [24-01-2023(online)].pdf 2023-01-24
19 202217012204-PETITION UNDER RULE 137 [25-01-2023(online)].pdf 2023-01-25
20 202217012204-OTHERS [25-01-2023(online)].pdf 2023-01-25
21 202217012204-FER_SER_REPLY [25-01-2023(online)].pdf 2023-01-25
22 202217012204-COMPLETE SPECIFICATION [25-01-2023(online)].pdf 2023-01-25
23 202217012204-CLAIMS [25-01-2023(online)].pdf 2023-01-25
24 202217012204-GPA-270223.pdf 2023-03-06
25 202217012204-Correspondence-270223.pdf 2023-03-06
26 202217012204-US(14)-HearingNotice-(HearingDate-05-03-2025).pdf 2025-02-20
27 202217012204-Correspondence to notify the Controller [27-02-2025(online)].pdf 2025-02-27

Search Strategy

1 202217012204_AmendAE_30-07-2023.pdf
2 202217012204E_29-07-2022.pdf