Abstract: The present technology relates to an audio processing device and method, and to a program, which enable audio reproduction with increased efficiency. In a head-related transfer function synthesis unit in the present technology, a diagonalized head-related transfer function matrix is pre-held. The head-related transfer function synthesis unit synthesizes an input signal in an annular harmonic domain for audio reproduction and the diagonalized head-related transfer function matrix pre-held. An annular harmonic inverse transformation unit generates a headphone drive signal in a time-frequency domain by performing, on the basis of an annular harmonic function, annular harmonic inverse transformation on a signal resulting from the synthesis performed by the head-related transfer function synthesis unit. The present technology can be applied to an audio processing device.
Technical field
[0001]
This technology is audio processing apparatus and method, and a program, audio processing apparatus and method capable of reproducing more efficiently audio, and a program.
BACKGROUND
[0002]
In recent years, recording the spatial information from all around in the field of voice, transmission, and the development and dissemination of the system to play has been advanced. For example, in the Super Hi-Vision is planned broadcasting in 3-dimensional multi-channel sound 22.2 channel.
[0003]
In addition to video surrounding the entire circumference in the field of virtual reality, but also to reproduce signals which surrounds the entire periphery in the speech is being to circulate in the world.
[0004]
Called Ambisonics therein, there is method of representing three-dimensional audio information flexibly adaptable to any recording and playback system, has attracted attention. Particularly Ambisonics which orders a quadratic or higher order Ambisonics (HOA (Higher Order Ambisonics)) are called (for example, see Non-Patent Document 1).
[0005]
In the three-dimensional multi-channel sound, the information of the sound is spread to the space axis in addition to the time axis, the frequency transform on the three-dimensional polar angle direction in Ambisonics, i.e. holds information by spherical harmonic transform ing. Moreover, given the horizontal plane only, cyclic harmonic transform is being performed. Spherical harmonic conversion and cyclic harmonic transform can be considered as equivalent to time-frequency transformation on the time axis of the audio signal.
[0006]
The advantage of this method lies in being able to encode and decode the information for any of the speaker array from any microphone array without limiting the number of number and the speaker microphones.
[0007]
On the other hand, factors that prevent the spread of Ambisonics, the speaker array consisting of a large amount of speaker reproduction environment is needed and the range where the sound space can be reproduced (sweet spot) can be cited narrower.
[0008]
For example when an attempt is made to raise the spatial resolution of the sound, it is necessary speaker array of more speakers, making such a system, such as at home is unrealistic. Also, narrow area that can reproduce the sound space in the space such as a movie theater, it is difficult to provide the desired effect for all audiences.
CITATION
Non-patent literature
[0009]
非特許文献1 : Jerome Daniel, Rozenn Nicol, Sebastien Moreau, “Further Investigations of High Order Ambisonics and Wavefield Synthesis for Holophonic Sound Imaging,” AES 114th Convention, Amsterdam, Netherlands, 2003.
Summary of the Invention
Problems that the Invention is to Solve
[0010]
Therefore, it is conceivable to combine the Ambisonics binaural reproduction technology. Binaural reproduction technology, commonly referred to as auditory display (VAD (Virtual Auditory Display)), head-related transfer function (HRTF (Head-Related Transfer Function)) is realized used.
[0011]
Here, the head-related transfer function, in which information from all directions surrounding the human head related transmitted how sound to both ears eardrum expressed as a function of direction of arrival and frequency.
[0012]
If those synthesized HRTFs from a certain direction with respect to the speech to be objective presented by headphones, rather than from the headphones for the listener, the sound from the direction of the head transfer functions using its arrived It is perceived as dolphin. VAD is a system that uses such a principle.
[0013]
If a plurality reproduce a virtual speaker using VAD, the same effect as Ambisonics of the speaker array system comprising a hard number of speakers in reality, it is possible to realize by a headphone presentation.
[0014]
However, in such systems, it can not be reproduced sufficiently efficient speech. For example, when combining the Ambisonics binaural reproduction technology, not only many computation of convolution such HRTF, it becomes many memory usage used for the operation or the like.
[0015]
This technology has been made in view of such circumstances, it is to be able to play more efficiently speech.
Means for Solving the Problems
[0016]
Speech processing apparatus according to an embodiment of the present technology, combines the portions and, diagonalized the head related transfer functions corresponding to the annular conditioning region of the input signal of the input signal or the spherical harmonic region, an annular conditioning region and HRTF synthesis unit, a signal obtained by the synthesis by cyclic conditioner inverse conversion based on the cyclic harmonics, and a cyclic harmonic inverse transformation unit for generating a headphone driving signals in the time-frequency domain.
[0017]
The said head transfer function combining unit, and a diagonal matrix obtained by a matrix comprising a plurality of HRTF diagonalized by an annular harmonic transform, from the input signal corresponding to each order of the cyclic harmonics made by causing sought product of the vector, it is possible to combine the said input signal and said diagonalized the HRTF.
[0018]
Head is the head transfer function combining unit, by using only the order of elements in the time of a predetermined configurable for each frequency of the diagonal elements of the diagonal matrix, which is the diagonalized and the input signal it is possible to perform the synthesis of the parts transfer function.
[0019]
Wherein the diagonal matrix may be the diagonalized the HRTF is used in common by each user to contain as elements.
[0020]
Wherein the diagonal matrix may be so that the diagonalized the HRTF depends on the individual user is included as an element.
[0021]
The speech processing apparatus, constituting the diagonal matrix, as well as previously held the diagonalized the HRTF common to the user, the head transfer function said diagonalized depends on individual users acquires a head transfer function said diagonalized obtained, further providing the matrix generating unit that generates the diagonal matrix from a head transfer function the diagonalized stored in advance can.
[0022]
Wherein the annular conditioner inverse conversion unit, to the annular harmonic function matrix comprising a circular harmonics in each direction are held, on the basis of the row corresponding to the predetermined direction of the annular harmonic function matrix, causing the annular harmonic inverse transform can.
[0023]
The speech processing device, the headphone to listen to the sound based on the driving signal further head direction acquisition unit that acquires a direction of the user's head is provided, on the annular conditioner inverse transform unit, said in the annular harmonic function matrix based on the row corresponding to the direction of the user's head, it is possible to perform the annular harmonic inverse transform.
[0024]
The speech processing device, the user of the further head direction sensor which detects the rotation of the head is provided, wherein the head direction obtaining unit, by acquiring the detection result of the head direction sensor unit, the You can obtain the direction of the user's head.
[0025]
The speech processing device, the headphone drive signal time time-frequency inverse conversion section for frequency inversion may be further provided.
[0026]
Audio processing method or a program according to an embodiment of the present technology, the input signal of the annular conditioning region or the portion corresponding to the annular conditioning region of the input signal of spherical harmonic region, and a head transfer function diagonalized, combined, a signal obtained by the synthesis by cyclic conditioner inverse conversion based on the cyclic harmonics, comprising the steps of generating a headphone driving signals in the time-frequency domain.
[0027]
In one aspect of the present technique, an input signal of the annular conditioning region or the portion corresponding to the annular conditioning region of the input signal of spherical harmonic regions, the diagonalized the head-related transfer function, are synthesized, the synthetic based the obtained signal to the annular harmonic functions by by cyclic harmonic inverse transform, headphone drive signal in the time frequency domain is generated.
Effect of the invention
[0028]
According to one aspect of the present technology, it is possible to reproduce more efficiently speech.
[0029]
Here, the advantages described in the present invention is not necessarily limited, it may be any of the effects described in the present disclosure.
BRIEF DESCRIPTION OF THE DRAWINGS
[0030]
It is a diagram illustrating a simulated stereophonic with [1] HRTF.
2 is a diagram showing the configuration of a typical speech processing unit.
3 is a diagram for describing calculation of the driving signal according to a general method.
Is a diagram showing the configuration of [4] the sound processing apparatus to add the head tracking.
Is a diagram for describing calculation of the drive signals of adding to FIG. 5 head tracking.
Is a diagram for describing calculation of the driving signal according to [6] proposed method.
7 is a diagram for explaining operations at the time of driving signals for calculating the proposed method and expansion techniques.
8 is a diagram showing an example of the configuration of a voice processing apparatus according to the present technology.
9 is a flowchart illustrating a driving signal generation process.
It is a diagram illustrating a calculation amount reduction due to [10] degree truncation.
11 is a diagram illustrating the calculation amount and the memory requirements of the proposed method and general method.
Is a diagram illustrating the generation of [12] HRTF matrix.
13 is a diagram for explaining calculation amount reduction due order truncation.
14 is a diagram for explaining calculation amount reduction due order truncation.
It is a diagram illustrating a configuration example of FIG. 15 speech processing apparatus according to an embodiment of the present technology.
16 is a flowchart illustrating a driving signal generation process.
17 is a diagram for explaining the arrangement of the virtual speakers.
18 is a diagram explaining the arrangement of the virtual speakers.
19 is a diagram for explaining the arrangement of the virtual speakers.
FIG. 20 is a diagram for explaining the arrangement of the virtual speakers.
21 is a diagram showing a configuration example of a computer.
DESCRIPTION OF THE INVENTION
[0031]
Hereinafter, with reference to the accompanying drawings, a description will be given of an embodiment according to the present technology.
[0032]
This technique captures the head related transfer function itself in one plane as a function of the two-dimensional polar coordinates, similarly performing cyclic harmonic transform, spherical harmonic region or cyclic conditioner by performing the combination of the input signal and HRTF in an annular conditioning region without passing through the decoding to the speaker array signal of the input signal is an audio signal region, more efficient regeneration in computation amount and memory usage it is intended to implement the system.
[0033]
For example, spherical harmonic conversion for the function f (theta, phi) on spherical coordinates is expressed by the following formula (1). The annular harmonic transform with respect to function on a two-dimensional polar coordinates f (phi) is represented by the following formula (2).
[0034]
[Number 1]
[0035]
[Number 2]
[0036]
The theta and phi in formula (1), shows the elevation angle and the horizontal angle in the respective spherical coordinate, Y n m (theta, phi) represents the spherical harmonics. Further, spherical harmonics Y n m (theta, phi) on top "-" it is what is written, the spherical harmonics Y n m represents the complex conjugate of (theta, phi).
[0037]
Further, the phi in equation (2) shows a horizontal angle in the two-dimensional polar coordinates, Y m (phi) represents a cyclic harmonics. Cyclic harmonics Y m in (phi) upper "-" is what is written, the cyclic harmonics Y m represents the complex conjugate of (phi).
[0038]
Here spherical harmonics Y n m (theta, phi) is expressed by the following equation (3). The annular harmonics Y m (phi) is expressed by the following equation (4).
[0039]
[Number 3]
[0040]
[Formula 4]
[0041]
N and m in the formula (3) is spherical harmonics Y n m represents the order of the (theta, phi), a -n ≦ m ≦ n. Further, j denotes a pure imaginary number, P n m (x) is the associated Legendre functions represented by the following formula (5). Similarly, in equation (4) m is cyclic harmonics Y m represents the order of the (phi), j represents the pure imaginary number.
[0042]
[Formula 5]
[0043]
Further, spherical harmonics transformed function F n m inverse conversion to the two-dimensional polar coordinates on the function f (phi) is as shown in equation (6). Further cyclic harmonics transformed function F m inverse transform of a function into f (phi) on the two-dimensional polar coordinates from is as shown in equation (7).
[0044]
[Number 6]
[0045]
[Number 7]
[0046]
Is held in the spherical harmonic region From the above, the input signal of the speech after the radial correction D ' n m from (omega), of L arranged on a circle of radius R of each speaker speaker driving signals S (x i conversion to, omega) is shown in the following equation (8).
[0047]
[Number 8]
[0048]
Incidentally, x in formula (8) i denotes the position of the speaker, omega represents time frequency of the sound signal. Input signal D ' n m (omega) is the audio signal corresponding to each order n and degree m of spherical harmonics for a predetermined time-frequency omega, in the calculation of equation (8), the input signal D' n m (omega) of | m | = only the n elements is used. That is, the input signal D ' n m only those corresponding to the annular conditioning region of the (omega) is used.
[0049]
Further, held in an annular conditioning region, the input signal of the speech after the radial correction D ' m from (omega), speaker driving signals of the L each speakers arranged on a circle of radius R S (x i , omega) conversion to is as shown in the following equation (9).
[0050]
[Number 9]
[0051]
Incidentally, x in the formula (9) i denotes the position of the speaker, omega represents time frequency of the sound signal. Input signal D ' m (omega) is an audio signal corresponding to each order m of the annular harmonic function for a given time-frequency omega.
[0052]
The position x in the formula (8) and (9) i is, x i = (Rcosarufa i , Rsinarufa i ) t is, i is shows a speaker index that identifies the speaker. Here, i = 1,2, ..., a L, alpha i represents the horizontal angle indicating the position of the i-th speaker.
[0053]
Conversion indicated by such formula (8) and (9) is a cyclic harmonic inverse transform corresponding to equation (6) and (7). Further, the formula (8) or (9) speaker driving signals S (x i when seeking, omega), and number of speakers L is the number of reproduction loudspeakers, the maximum value of the degree N, that is the order m of the annular harmonics the N must satisfy the relationship expressed by the following equation (10). Incidentally, in the following, the input signal will be described for the case where a signal of cyclic-conditioning region, the input signal is also a signal of the spherical harmonic regions, the input signal D ' n m of the (omega) | m | = only by using the n elements, it is possible to obtain the same effect by the same process. That is, the same argument holds in the case in the input signal of the annular conditioning region for the input signal of the spherical harmonic space.
[0054]
[Formula 10]
[0055]
Meanwhile, common ones as a technique to simulate stereophonic at the ear, the headphone presentation is, for example, a method using a head related transfer function as shown in FIG.
[0056]
In the example shown in FIG. 1, the input Ambisonics signal is decoded, the respective speaker driving signals of the virtual speaker SP11-1 to virtual speaker SP11-8 is generated a plurality of virtual speaker. At this time, the signal to be decoded, for example, the input signal D 'described above n m (omega) and the input signal D' m corresponding to (omega).
[0057]
Here, each of the virtual speaker SP11-1 to virtual speaker SP11-8 are virtually arranged are arranged annularly, speaker driving signals for each virtual speaker, the above-mentioned formula (8) or (9) obtained by calculation. In the following description, when it is particularly necessary to distinguish the virtual speaker SP11-1 to virtual speaker SP11-8, simply and also referred to as a virtual speaker SP11.
[0058]
This way, the speaker drive signal of each virtual speaker SP11 is obtained, for each of those virtual speaker SP11, actually left and right driving signals headphone HD11 for reproducing audio (binaural signal) using HRTF It is generated by the convolution operation. Then, the sum of the drive signals of a headphone HD11 obtained for each virtual speaker SP11 is the final drive signal.
[0059]
Such a technique is, for example, "ADVANCED SYSTEM OPTIONS FOR BINAURAL RENDERING OF AMBISONIC FORMAT (Gerald Enzner et. Al. ICASSP 2013)" are described in detail in such.
[0060]
Left and right drive signal head transfer function H used to generate the headphone HD 11 (x, omega) is the sound source position x in a state where the user's head is present a listener in free space, the user's eardrum position to the transfer characteristic H of 1 (x, omega) and the transfer characteristic H from the sound source position x in a state where the head is not present until the head center O 0 is normalized by (x, omega). That is, the head related transfer function H of the sound source position x (x, omega) is obtained by the following equation (11).
[0061]
[Number 11]
[0062]
Here, the direction of the head related transfer function H (x, ω) convolving any speech signal, by presenting the headphones, though convoluted's HRTF H (x, ω) the listener , in other words it is possible to give the illusion as if the sound from the direction of the sound source position x is heard.
[0063]
In the example shown in FIG. 1, the driving signals of the left and right headphone HD11 with is used this principle is generated.
[0064]
Specifically, the position x of the position of each virtual speaker SP11 i and the speaker driving signals of the virtual speaker SP11 S (x i a, omega).
[0065]
Further, the number of virtual loudspeakers SP11 and L (here L = 8), the final left and right driving signals headphone HD 11, respectively P l and P r to.
[0066]
In this case, the speaker drive signal S (x i , omega) a to simulate headphones HD11 presented, the drive signal P of the left and right headphone HD11 l and the drive signal P r is be determined by calculating the following equation (12) can.
[0067]
[Number 12]
[0068]
In the equation (12), H l (x i , omega) and H r (x i , omega), the position x of the virtual speaker SP11 each i head is normalized to the eardrum of the left and right of the listener from It shows a part transfer function.
[0069]
Such operation, the input signal D 'of the annular conditioning region m the (omega), eventually it is possible to reproduce the headphone presentation. That is, it is possible to achieve the same effect as Ambisonics headphones presented.
[0070]
As described above, general approaches to combine Ambisonics binaural reproduction technology (hereinafter, also general method referred to) by the sound processing device for generating a driving signal of the left and right headphones from the input signal, shown in FIG. 2 It is configured.
[0071]
That is, the sound processing apparatus 11 shown in FIG. 2, an annular conditioning inverse transform unit 21, HRTF synthesis unit 22, and the time-frequency inverse transform unit 23.
[0072]
Cyclic conditioner inverse transform unit 21, the inputted input signal D ' m against (omega), it performs cyclic harmonic inverse transformation by calculating the equation (9), the resulting speaker driving virtual speakers SP11 signal S (x i supplies, omega) to HRTF synthesis unit 22.
[0073]
HRTF synthesis unit 22, loudspeaker driving signals S (x from the annular conditioner inverse transform unit 21 i , omega) and a head transfer previously prepared function H l (x i , omega) and HRTF H r (x i , omega) from the drive signal P of the left and right headphone HD11 by the formula (12) l and the drive signal P r , and outputs.
[0074]
Further, time-frequency inverse conversion unit 23, the driving signal P is a signal of the output time-frequency domain from the head transfer function combining unit 22 l and the drive signal P r with respect to, performs time frequency inversion, as a result driving signal is a signal of the time domain obtained p l (t) and the drive signal p r a (t), to reproduce the sound by supplying the headphone HD 11.
[0075]
In the following, the driving signal P for the time-frequency omega l and the drive signal P r when it is not necessary to distinguish the, simply referred to as a drive signal P (omega), the driving signal p l (t) and the drive signal p r If it is unnecessary to distinguish between a (t), simply referred to as drive signals p (t). Further, the head related transfer function H l (x i , omega) and HRTF H r (x i , omega) when it is not necessary to distinguish the just HRTF H (x i referred, omega) and also .
[0076]
In the sound processing apparatus 11, 1 × 1, that is in order to obtain the first row and first column of the drive signal P (omega), calculation is performed as shown in FIG. 3, for example.
[0077]
In Figure 3, H (omega) is, L pieces of the head related transfer function H (x i represents, omega) 1 × L vector (matrix) composed of. Further, D '(omega) is the input signal D' m represents a vector of (omega), the input signal D bin of the time-frequency omega ' m the number of (omega) When K, vector D' (omega ) is a K × 1. Furthermore Y alpha , each order cyclic harmonics Y m (alpha i represents the matrix of), the matrix Y alpha is the matrix of L × K.
[0078]
Accordingly, the audio processor 11, L × K matrix Y alpha and K × 1 vector D '(omega) and the matrix S obtained from the matrix calculation is required, further vectors of the matrix S and 1 × L (matrix ) matrix operation with H (omega) is performed, so that one driving signal P (omega) is obtained.
[0079]
Further, the listener's head wearing the headphone HD11 is movable in a given direction phi represented by the horizontal angle of the two-dimensional polar coordinate j when rotated in the direction of, for example, the left headphone of the headphone HD11 drive signal P l (phi j , omega) is shown in the following equation (13).
[0080]
[Formula 13]
[0081]
In the equation (13), the drive signal P l (phi j , omega) is the drive signal P described above l shows, where the position, i.e. the direction phi j drive signal in order to clarify the time-frequency omega P L (Fai J , Omega) are marked. Furthermore, the matrix in equation (13) u (phi j ), the angle phi j is a rotation matrix for rotating only. Thus, for example, the predetermined angle phi j When = theta, matrix u (phi j ), i.e. the matrix u (theta) is the rotation matrix for rotating angle theta, expressed by the following equation (14).
[0082]
[Number 14]
[0083]
For general audio processor 11, further example configuration for identifying the direction of rotation of the listener's head, as shown in FIG. 4, i.e. by adding the configuration of the head tracking function, the sound image viewed from the listener position can be fixed in space. Note that portions corresponding to the case in FIG. 2 in FIG. 4 are given the same reference numerals, and description thereof will be omitted as appropriate.
[0084]
In the sound processing apparatus 11 shown in FIG. 4, the configuration shown in FIG. 2, further head direction sensor 51 and the head direction selecting portion 52 is provided.
[0085]
The head direction sensor 51 detects the rotation of the head of the user is a listener, and supplies the detection result to the head direction selection unit 52. The head direction selection unit 52 based on the detection result from the head direction the sensor unit 51, the rotational direction of the listener's head, that is, the direction of the listener's head after rotation direction phi j calculated as the head parts and supplies to the transfer function combining unit 22.
[0086]
In this case, the HRTF synthesis unit 22, direction φ supplied from the head direction selecting portion 52 j on the basis of, among the plurality of head-related transfer function are prepared in advance, as seen from the listener's head relative coordinates u (phi of each virtual speaker SP11 j ) -1 x i using HRTFs calculating a drive signal of the left and right headphone HD 11. Thus, as with the real speaker, even in the case of reproducing the audio by headphone HD 11, a sound image position as viewed from the listener it can be fixed in space.
[0087]
And general technique described in the above, be generated The general procedure headphone driving signal by further technique adds the head tracking function, without using a speaker array, and without the range of the sound space can be reproduced is limited it is possible to obtain the same effect as cyclic placed Ambisonics. However, in these methods, not only many computation of convolution such HRTF, becomes many memory usage used for the operation or the like.
[0088]
Therefore, in this technique, the general procedure of convolution of the head related transfer function which has been performed in the time-frequency domain and to perform in an annular conditioning region. Thus, to reduce the amount of calculation and memory required for convolution can be reproduced more efficiently speech.
[0089]
So, below, we describe a method according to the present technology.
[0090]
For example, focusing on the left headphone, the drive signals P of the left headphone for all the rotational direction of the head of a listener user (listener) l (phi j , omega) consists vector P l (omega) is expressed by the following equation ( It is expressed as shown in 15).
[0091]
[Number 15]
[0092]
In the equation (15), S (ω) is the speaker driving signals S (x i is a vector of, omega), S (omega) = Y alpha is D '(ω). Further, in the equation (15) Y alpha is represented by the following equation (16), the angle alpha of each order and each virtual speaker i cyclic harmonics Y of m (alpha i represents the matrix of). Here, i = 1,2, ..., L, and the maximum value of degree m (the maximum order) is N.
[0093]
D '(omega) is the following represented by equation (17), the input signal D of a sound corresponding to each order' m represents consist (omega) vector (matrix). Each input signal D ' m (omega) is the signal of cyclic-conditioning area.
[0094]
Furthermore, in the equation (15), H (omega) has the following formula represented by (18), the direction of the listener's head direction phi j each virtual speaker when a, viewed from the listener's head relative coordinates u (phi j ) -1 x i HRTF H (u (phi j ) -1 x i represents, omega) consists matrix. In this example, the direction phi 1 to the direction phi M for a total of M direction, the head-related transfer function H of each virtual speaker (u (phi j ) -1 x i , omega) are prepared.
[0095]
[Number 16]
[0096]
[Formula 17]
[0097]
[Equation 18]
[0098]
The head of the listener's direction phi j drive signal P of the left headphone when facing l (phi j , omega) In calculating the among the matrix H (omega) of the head-related transfer function, the listener's head parts orientation direction phi is the j row corresponding to, i.e. HRTF H (u (phi j ) -1 x i , omega) to select the row of the may be performed the calculation of equation (15).
[0099]
In this case, the calculation only the rows needed as shown in FIG. 5 for example is performed.
[0100]
In this example, the head related transfer function are prepared for M each direction, a matrix calculation shown in equation (15) becomes as shown by the arrow A11.
[0101]
In other words, 'the input signal D time-frequency omega m when the number of (omega) and K, the vector D' (omega) is the K × 1, i.e. K rows and one column of the matrix. Further, the matrix Y of cyclic harmonics α is L × K, and the matrix H (omega) is the M × L. Therefore, the calculation of equation (15), the vector P l (omega) is the M × 1.
[0102]
Here, the matrix Y alpha when obtaining the vector S (omega) performs matrix operation between the vector D '(omega) (product-sum operation), the drive signal P l (phi j During calculation of, omega), an arrow A12 of the matrix H (omega), as shown in the direction φ of the listener's head j may select a row corresponding to, reducing the amount of calculation. In Figure 5, hatched applied portion in the matrix H (omega) is the direction phi j represents the row corresponding to, the operation of this line and the vector S (omega) is performed, the left headphone desired drive signal P l (phi j , omega) is calculated.
[0103]
The direction phi 1 to the direction phi M input signal D 'for a total of M number each direction of m a matrix of M × K consisting of the cyclic harmonics corresponding to (omega) Y phi shall be a. In other words, each direction phi 1 to the direction phi M annular harmonics Y for m (phi 1 ) to cyclic harmonics Y m (phi M a) a matrix Y phi and. Further, the matrix Y phi Hermitian transposed matrix of Y phi H and.
[0104]
At this time, when defining a matrix H '(omega), as shown in the following equation (19), the vector P shown in equation (15) l (omega) can be expressed by the following equation (20).
[0105]
[Number 19]
[0106]
[Number 20]
[0107]
In the equation (20), it is a vector B '(ω) = H' (ω) D '(ω).
[0108]
In equation (19), by an annular harmonic transform, HRTF, and more particularly a calculation diagonalizing matrix H composed of head-related transfer function in the time-frequency domain (omega) has been performed. Further, the calculation of equation (20), it can be seen that the convolution of the loudspeaker drive signal and head-related transfer function is performed in an annular conditioning region. Incidentally, the matrix H '(ω) is can be held by calculating in advance.
[0109]
In this case, the head of the listener's direction phi j drive signal P of the left headphone when facing l (phi j , omega) When the calculation of the matrix Y of cyclic harmonics phi of the listener direction phi head j row corresponding to, i.e. annular harmonics Y m (phi j obtained by selecting a row of) that may be performed calculation of equation (20).
[0110]
Here, the diagonal of the matrix H (omega) is possible, i.e. if sufficiently matrix H (omega) is diagonalized by the above-mentioned formula (19), the left headphone drive signal P l (phi j , omega) calculated in calculating is only the calculation shown in the following equation (21). Thus, it is possible to greatly reduce the amount of calculation and memory requirements. In the following, a possible diagonal of matrix H (omega), the matrix H '(ω) is continuously described as being a diagonal matrix.
[0111]
[Number 21]
[0112]
Wherein in (21), H ' m (omega) is the matrix H is a diagonal matrix' (omega) 1 single element, i.e. the matrix H 'direction φ of the head in (omega) of j corresponding component (element) to It shows the HRTF annular conditioning region as the. HRTF H ' m m in (omega) indicates the degree m of the annular harmonics.
[0113]
Similarly Y m (phi j ) is a matrix Y phi direction phi of the head of the j indicates a cyclic harmonics to be one element in the row corresponding to the.
[0114]
In such calculation shown in equation (21), the calculation amount is reduced as shown in FIG. That is, calculation shown in equation (20), the matrix Y of M × K as indicated by an arrow A21 in FIG. 6 phi , K × M matrix Y phi H , M × L matrix H (omega), L × K matrices Y alpha , and the vector D of K × 1 'has a matrix operation (omega).
[0115]
Here, equation (19) to Y as defined phi H H (omega) Y alpha because is a matrix H '(omega), the calculation shown by the arrow A21, eventually, as shown in an arrow A22. In particular, the matrix H 'is calculated to obtain the (omega), offline, that since it is possible to carry out in advance, the matrix H' if and held previously obtained (omega), that much online headphone it is possible to reduce the calculation amount when determining the drive signals.
[0116]
The calculation of the equation (19), that in the calculation for obtaining the matrix H '(omega), the diagonal of the matrix H (omega) is performed. Therefore, as shown by an arrow A22 matrix H '(ω) is a matrix of K × K, by diagonalization, the matrix only the diagonal component represented by shaded area substantially. That is, in the matrix H '(omega), the value of the non-diagonal elements 0, it is possible to greatly reduce the subsequent amount of calculation.
[0117]
With such pre-matrix H '(omega) is obtained, when actually determining the drive signals of the headphones, the calculations indicated by an arrow A22 and the arrow A23, i.e. so that the calculation of the above equation (21) takes place.
[0118]
That is, the matrix H as shown by the arrow A22 'and (omega), inputted input signal D' m 'on the basis of the (omega), the vector B of K × 1 online' (omega) vector D consisting of ( ω) is calculated.
[0119]
Then, a matrix Y as indicated by an arrow A23 phi of the direction phi of the listener's head j is selected row corresponding to the, and the selected row, the matrix operation of the vector B '(omega) , the drive signal P of the left headphone l (phi j , omega) is calculated. In Figure 6, the matrix Y phi hatched decorated portion in the direction phi j represents the row corresponding to the annular harmonic functions Y elements constituting this line is shown in Formula (21) m (phi j ) and a.
[0120]
Now, with reference to FIG. 7, the technique described by method (hereinafter, also our method referred to) at least a was added to the general procedure of head tracking technique ( hereinafter, a comparison of the sum-of-products arithmetic amount and memory requirements of the both extended technique referred to).
[0121]
For example a vector D 'length of (omega) and K, the matrices of the head related transfer function H a (omega) and M × L matrix Y of cyclic harmonics α is L × K, and the matrix Y phi is M × K, and the matrix H '(ω) is a K × K.
[0122]
Here, in the extension scheme, as indicated by an arrow A31 in FIG. 7, the bin of the time-frequency omega respect (hereinafter, also referred to as time-frequency bins omega), to convert the vector D '(omega) the time to the frequency domain process sum of products L × K occurs in only 2L in convolution of the left and right head-related transfer function product-sum operation occurs.
[0123]
Therefore, the total sum of product arithmetic number of cases in the extended approach, the (L × K + 2L).
[0124]
Further, when the respective coefficients of the product-sum operation is to be 1 byte, memory amount required during operation by expansion technique, for each time-frequency bin omega, (Direction HRTF holding) × 2 is a byte, the number of direction head related transfer function to be held, the M × L as indicated by an arrow A31 in FIG. Furthermore, the matrix Y common annular harmonics to all the time-frequency bins omega alpha only L × K bytes for the memory is required.
[0125]
Therefore, when the number of time-frequency bins ω is W, memory requirements in the extended approach, the (2 × M × L × W + L × K) bytes in total.
[0126]
In contrast, in the proposed method, calculation shown by the arrow A32 of Fig. 7 is performed for each time-frequency bin omega.
[0127]
That is, in the proposed method, sum of products of convolution K × K and for each time-frequency bin omega, vector D in the annular conditioning region per ear '(omega) and the matrix H HRTF' (omega) operation occurs, only K further time conversion to the frequency domain product sum calculation is generated.
[0128]
Therefore, the sum of the product sum operation number is the case in the proposed method, the (K × K + K) × 2.
[0129]
However, the matrix H (omega) relative to the diagonal of the head related transfer function as described above is carried out, the product according to the convolution of the vector D '(omega) and the matrix H HRTF' (omega) since the sum operation to be only K per ear, product sum operation number of total becomes 4K.
[0130]
Further, the amount of memory which is needed at the time of operation by the proposed method, for each time-frequency bin omega, the 2K byte since it is only the matrix H 'diagonal elements of (omega) of the head-related transfer function. Further matrices Y common annular harmonics to all the time-frequency bins omega phi memory is needed only M × K bytes for.
[0131]
Therefore, when the number of time-frequency bins ω is W, memory requirements in the proposed method, the (2 × K × W + M × K) bytes in total.
[0132]
Now, when assumed as 12 the maximum degree of cyclic harmonics, and K = 2 × 12 + 1 = 25. The number L of the virtual speaker is assumed to be L = 32 due to the need for greater than K.
[0133]
In this case, the product-sum operation amount of expansion technique whereas a (L × K + 2L) = 32 × 25 + 2 × 32 = 864, product-sum operation of the proposed method requires only 4K = 25 × 4 = 100 because, it can be seen that the amount of greatly calculation is reduced.
[0134]
Further, the amount of memory required at the time of calculation, for example, a W = 100 and M = 100, the extension method is (2 × M × L × W + L × K) = 2 × 100 × 32 × 100 + 32 × 25 = 640800. In contrast, the amount of memory required for calculation of the proposed method, it can be seen that (2 × K × W + M × K) = 2 × 25 × 100 + 100 × 25 = 7500, and the the substantial need memory capacity is reduced.
[0135]
Next, the audio processing apparatus according to an embodiment of the present technology described in the above will be described. Figure 8 is a diagram showing a configuration example of an embodiment of a voice processing apparatus according to the present technology.
[0136]
Audio processing apparatus 81 shown in FIG. 8, the head direction sensor section 91, the head direction selecting portion 92, HRTF synthesis unit 93, an annular conditioning inverse transform unit 94 and the time-frequency inverse conversion unit 95, there. Incidentally, the audio processing apparatus 81 may be incorporated in the headphone, it may be a device different from the headphone.
[0137]
The head direction sensor 91 by, for example, an acceleration sensor or an image sensor attached to the user's head if necessary, by detecting the rotation of the head of the user is a listener (motion), the detection results supplied to the head direction selection unit 92. Note that the user referred to herein is a user who listens to audio reproduced by a headphone based user wearing the headphones, that is, the drive signal of the left and right headphone obtained in the time-frequency inverse conversion unit 95.
[0138]
The head direction selection unit 92 based on the detection result from the head direction the sensor unit 91, the rotational direction of the listener's head, i.e. the direction φ of the listener's head after rotation j seeking, cyclic harmony supplied to the inverse transform unit 94. In other words, the head direction selecting portion 92, by acquiring the detection result from the head direction the sensor unit 91, the direction φ of the user's head j acquired.
[0139]
The HRTF synthesis unit 93, the input signal D 'of each order of annular harmonics for each time-frequency bin omega from the outside is the audio signal of the annular conditioning region m (omega) is supplied. Further, HRTF synthesis unit 93 holds the matrix consisting of head-related transfer function obtained in advance by calculation H '(ω).
[0140]
HRTF synthesis unit 93, the supplied input signal D ' m and (omega), the held matrix H' (omega), i.e. HRTF diagonalized by the above-mentioned formula (19) by performing the convolution operation between the matrix, 'an input signal D in an annular conditioning region m synthesizes the (omega) and head related transfer function, the resulting vector B' (omega) an annular conditioning inverse transform unit It supplies it to the 94. In the following, the vector B '(omega) element B of' m and also referred to as (omega).
[0141]
Cyclic conditioner inverse transform unit 94, a matrix Y that previously made from each direction of the annular harmonic phi holds, the matrix Y phi of the rows constituting the direction supplied from the head direction selecting portion 92 phi j corresponding row, ie cyclic harmonics Y of formula (21) above m (phi j selects consists of) line.
[0142]
Cyclic conditioner inverse transform unit 94, direction phi j the matrix Y selected based on phi annular harmonics constituting the rows of Y m (phi j and), the vector B supplied from the HRTF synthesis unit 93 '( element B 'of omega) m by computing the sum of the product of the (omega), an input signal head related transfer function is synthesized cyclic conditioner inverse transform.
[0143]
Incidentally, the convolution arithmetic HRTF in HRTF synthesis unit 93, an annular conditioning inverse transform the annular conditioner inverse transform unit 94 is performed for each left and right headphones. Thus, the annular conditioner inverse transform unit 94, the drive signal P of the left headphone of the time-frequency domain l (phi j , omega) and the drive signal P of the right headphone in the time-frequency domain r (phi j , omega) and the time obtained for each frequency bin ω.
[0144]
Cyclic conditioner inverse transform unit 94, the drive signal P of the left and right headphone obtained by cyclic conditioner inverse transform l (phi j , omega) and the drive signal P r (phi j , omega) the time supplied to the frequency inverse conversion unit 95 to.
[0145]
Time frequency inverse transform unit 95, one for each cylinder of the headphones, by performing time-frequency inverse conversion with respect to the drive signal supplied time-frequency domain from the annular conditioner inverse transform unit 94, a drive signal of the left headphone of the time domain p l (phi j , t) and the driving signal p to the right headphone time domain r (phi j , t) and obtains and outputs their driving signals to the subsequent stage. Subsequent headphones, headphones including earphones more, the reproducing apparatus for reproducing audio in 2-channel, sound is reproduced based on the driving signal outputted from the time frequency inverse transform unit 95.
[0146]
Subsequently, with reference to the flowchart of FIG. 9, the driving signal generation processing will be described which is performed by the sound processing apparatus 81. The drive signal generation processing, the input signal D 'from the outside m is initiated and (omega) is supplied.
[0147]
In step S11, the head direction sensor 91 detects the rotation of the head of the user is a listener, and supplies the detection result to the head direction selection unit 92.
[0148]
In step S12, the head direction selection unit 92 based on the detection result from the head direction the sensor unit 91, the direction φ of the listener's head j seeking, supplies the annular conditioner inverse transform unit 94.
[0149]
In step S13, HRTF synthesis unit 93, the supplied input signal D ' m (omega) relative advance the held matrix H' HRTF constituting the (omega) H ' m ( convolution omega), and supplies the resultant vector B 'to (omega) annularly conditioner inverse transform unit 94.
[0150]
In step S13, the annular conditioning region, the head related transfer function H ' m a matrix consisting of (ω) H' (ω) , the input signal D ' m the calculation of the product of (omega) vector D consisting of' and (omega) , i.e. H 'of the above formula (21) m (omega) D' m is calculated to obtain the (omega) is carried out.
[0151]
In step S14, cyclic conditioner inverse transform unit 94, a matrix Y stored in advance φ and the supplied direction φ from head direction selecting portion 92 j on the basis of the supplied from the HRTF synthesis unit 93 It was performed annular harmonic inverse transform to the vector B '(ω), and generates a driving signal of the left and right headphones.
[0152]
That is, an annular conditioning inverse transform unit 94, a matrix Y phi direction phi from j to select the row corresponding to the annular harmonic function Y constituting the selected row m (phi j and), the vector B 'to (omega) element B 'constituting m by calculating (omega) Tokara equation (21), the drive signal P of the left headphone l (phi j , omega) is calculated. The annular conditioner inverse transform unit 94 performs the same operation as in the left headphone also right headphone, drive signal P of the right headphone r (phi j , omega) is calculated.
[0153]
Cyclic conditioner inverse transform unit 94, thus the drive signal P of the left and right headphone obtained by l (phi j , omega) and the drive signal P r (phi j supplies, omega) in the time-frequency inverse conversion unit 95 .
[0154]
In step S15, the time-frequency inverse conversion unit 95, one for each cylinder of the headphones, it performs time frequency inversion to the drive signal supplied time-frequency domain from the annular conditioner inverse transform unit 94, the left headphone driving signal p l (phi j , t), and the drive signal p of the right headphone r (phi j , t) is calculated. For example inverse discrete Fourier transform is performed as a time frequency inversion.
[0155]
Time frequency inverse transform unit 95, thus the drive signal in the time domain which is seeking p l (phi j , t) and the drive signal p r (phi j outputs, t) to the left and right headphone, the drive signal generation process It comes to an end.
[0156]
Audio processing apparatus 81 as described above, the convolution of the head related transfer function to an input signal in an annular conditioning region, the convolution result by performing a cyclic harmonic inverse transform on, calculates the drive signals of the left and right headphones.
[0157]
Thus, by performing the convolution of the head related transfer function in an annular conditioning region, with a calculation amount at the time of generating the driving signal of the headphone can be greatly reduced, even significantly the amount of memory which is needed at the time of operation it can be reduced. In other words, it is possible to reproduce more efficiently speech.
[0158]
Incidentally, the head related transfer constituting the matrix H (omega) function H (u (phi j ) -1 x i , omega) has been found to differ the necessary degree in the annular conditioning region, this is, for example, "Efficient Real Spherical harmonic Representation of Head- Related Transfer Functions (Griffin D. Romigh et. al., 2015) ," etc., wherein It is.
[0159]
For example of the matrix H 'diagonal elements of (omega) of the head-related transfer function, if found necessary degree m = N (omega) at each time-frequency bin omega, for example by calculation of the following equation (22) drive signal P of the left headphone l (phi j , omega) and the like to determine a, it is possible to reduce the amount of calculation. This is also true for the right headphone.
[0160]
[Equation 22]
[0161]
Calculation of equation (22), the place is basically the same as the calculation of equation (21), the range of addition target by Σ is was up to order the formula (21) m = -N to N in equation (22) the order m = -N (omega) to N (omega) (where, N ≧ N (omega)) with the difference that up.
[0162]
In this case, for example, in HRTF synthesis unit 93 as shown in FIG. 10, each element of only a portion of the diagonal elements of the matrix H '(ω), i.e. the order m = -N (ω) to N (omega) only so that is used for convolution operation. Note that portions corresponding to the case in FIG. 8 in FIG. 10 are denoted by the same reference numerals, and a description thereof will be omitted.
[0163]
In Figure 10, the character "H" rectangle labeled the (omega) "is a matrix H for each time-frequency bin omega held in HRTF synthesis unit 93 'represents the diagonal elements of (omega) cage, and represents an element part of order m that require hatched portions thereof diagonals, i.e. orders -N (omega) to order N (omega).
[0164]
In this case, step S13 and step S14 in FIG. 9, convolutional and cyclic harmony inverse transform HRTF by calculation of equation (21) instead of equation (22) takes place.
[0165]
Thus matrix H '(omega) performs convolution operation using only the necessary orders of components (elements) of, for other orders it not to perform the operation, further the amount of computation and memory requirements it is possible to reduce. Incidentally, the necessary degree of the matrix H '(ω) is settable for each time-frequency bin omega, i.e. may also be set for each time-frequency bins omega, the total time-frequency bins omega, common orders may be set as the next required number.
[0166]
Here, a general method, shown in Figure 11 the amount of computation and memory requirements in the case of the proposed method described above, the operation only furthermore required degree m in the proposed method.
[0167]
Column of "degree of annular harmonic" in FIG. 11, the maximum degree of annular harmonics | m | represents the value of = N, the column of "necessary virtual speaker speed" is to reproduce the correct sound field It indicates the number of virtual speakers as a minimum.
[0168]
Further, the column of the column "amount of computation (General Procedure)" generally approach indicates the number of product-sum operations required to generate the driving signal of the headphone, the "computation amount (the proposed method)" is It indicates the number of product-sum operations required to generate the driving signal of the headphone by the proposed method.
[0169]
Furthermore, the column "amount of operation (the proposed method, the degree -2)" is the number of multiply-add operations required to generate the driving signal of the headphone the proposed method, and by calculation using up to order N (omega) the shows. In this example, in particular a case the upper secondary content is not calculated truncated in order m.
[0170]
Here, these general techniques, the proposed method, in each operation of the column in the case of performing calculation using up to order N (omega) in the proposed method, marked the product sum operation times for each time-frequency bin omega there.
[0171]
Further, the column of "memory (General Method)" is a general method shows the amount of memory required to generate a driving signal of the headphone, the column of "memory (the proposed method)" is the proposed method of headphone It indicates the amount of memory required to generate a driving signal.
[0172]
Column addition "memory (the proposed method, the degree -2)" indicates the amount of memory required to generate a driving signal of the headphone the proposed method, and by calculation using up to order N (omega). In this example, in particular orders | has the example upper secondary content is not calculated truncated for | m.
[0173]
In the column the symbol "**" is described in FIG. 11, it is shown that the calculation as the order N = 0 because the order -2 becomes negative has been performed.
[0174]
For example, in the example shown in FIG. 11, focusing on the field of operation amount in order N = 4, the calculation amount of the proposed method has a 36. In contrast, in order N = 4, if necessary orders for the time-frequency bins omega there was N (ω) = 2, using the proposed method, and up to order N (omega) in the calculation calculation amount in the case has a 4K = 4 (2 × 2 + 1) = 20. Therefore, it is found that can be reduced up to 55% the amount of calculation as compared with the case the original order N was 4.
[0175]
for the way, the head-related transfer function, diffraction, such as the listener's head and pinna, a filter formed by reflection, head-related transfer function by listener individual is different. Therefore, optimizing the head related transfer function to the individual is important for binaural reproduction.
[0176]
However, keeping only the listener component that is assumed HRTFs individual unworthy in terms of memory capacity. This HRTFs also apply to the case of holding an annular conditioning region.
[0177]
If when used in personal optimized HRTF reproducing system according to the proposed method, the hourly frequency bins omega, or at all time-frequency bins omega, the order which depends the degree that does not depend on individual Once you have specified in advance, it is possible to reduce the personal dependent parameters required. At the time of estimation of the HRTF of the listener individuals from such body shape is also contemplated objective variable coefficient individuals dependent on the annular conditioner area (HRTF).
[0178]
Here, the degree of individuals dependent transfer characteristic varies widely from individual user, that is the head related transfer function H ' m (omega) is the different orders m for each user. Conversely, the order which does not depend on the individual, the difference in transfer characteristics of the individual small enough HRTF H ' m is the degree m of the (omega).
[0179]
Such a person-dependent not HRTF orders, the example of the speech processing apparatus 81 shown from the HRTF order when generating the matrix H '(omega), for example, in FIG. 8 to individuals dependent so the HRTF orders that depend on individual as shown in FIG. 12 is obtained by any method. Note that portions corresponding to the case in FIG. 8 in FIG. 12 are denoted by identical reference numerals, and description thereof will be omitted as appropriate.
[0180]
In the example of FIG. 12 represents a diagonal of the letter "H '(omega)" is a matrix H of the rectangle marked time-frequency bin ω' (ω), the shaded part of the diagonal components, in advance portions held in the audio processing unit 81, i.e. individuals not dependent order HRTF H ' m represents the portion of the (omega). In contrast, the portion indicated by the arrow A91 of the diagonal elements, the head-related transfer orders to individuals dependent function H ' m represents the portion of the (omega).
[0181]
In this example, it is represented by the shaded portion in diagonal, the following number of head transfer that does not depend on individual function H ' m (omega) is a head-related transfer function used in common in all users. In contrast, as indicated by arrow A91, the head-related transfer orders to individuals dependent function H ' m head (omega) is the like which have been optimized for each individual user, is different for each individual user is used parts is the transfer function.
[0182]
Audio processing unit 81, the character "personalized factor" is represented by a rectangle marked, the head-related transfer orders to individuals dependent function H ' m acquires (omega) from the outside, the acquired head transfer function H ' m and (omega), and not the order of the head related transfer function H depends on the individual stored in advance' m generate diagonal component of the matrix H '(omega) from the (omega), HRTF synthesis supplied to the part 93.
[0183]
Here, the matrix H '(omega) is a head transfer function used in common for all users, although those used for each user will be described an example composed of a different head related transfer function, the matrix H' All non-zero elements in the (omega) may be adapted is different for each user. Further, the same matrix H '(ω) may be used in common in all users.
[0184]
The generated matrix H '(ω) is composed of different elements in each time-frequency bins omega as shown in FIG. 13, different for each element operation is performed the time-frequency bins omega as shown in FIG. 14 it may be. Note that portions corresponding to the case in FIG. 8 in FIG. 14 are denoted by identical reference numerals, and a description thereof will be omitted.
[0185]
In Figure 13, as indicated by respective arrows A101 or arrow A106, the character "H 'rectangles (omega)" is marked with a matrix H of a predetermined time-frequency bins omega' represents the diagonal elements of (omega) . Further, the hatched portion of their diagonal components represent component parts of the order m required.
[0186]
For example, in the example shown by respective arrows A101 or arrows A103, among the diagonal elements of the matrix H '(omega), it has become a necessary orders component parts the portion consists of elements adjacent to each other, in the diagonal component positions of the component parts (regions) has a different position in each case.
[0187]
In contrast, in the example shown by the respective arrows A104 or arrows A106, among the diagonal elements of the matrix H '(omega), it is a plurality of portions that require orders of component parts consisting of elements adjacent to each other there. The number and position of the portion consisting of the required elements in the diagonal in these examples, the size is different each example.
[0188]
The audio processing apparatus 81 as shown in FIG. 14, the database of HRTF diagonalized by an annular harmonic transform, i.e. in addition to the matrix H of the respective time-frequency bin ω '(ω), the temporal frequency will have the information indicating the order m required for each bin ω simultaneously as a database.
[0189]
In Figure 14, the letter "H" rectangle labeled the (omega) "is a matrix H for each time-frequency bin omega held in HRTF synthesis unit 93 'represents the diagonal elements of (omega) cage, the shaded portion of their diagonal components represent component parts of the order m required.
[0190]
In this case, the HRTF synthesis unit 93, for example until -N every time-frequency bins omega (omega) required by the time-frequency bins omega from the next order m = N (ω), head-related transfer function and the input signal D ' m the product of the (omega) is obtained. That, H 'in formula (22) above m (omega) D' m the calculation of (omega) is carried out. Thus, in the HRTF synthesis unit 93, it is possible to reduce unnecessary order calculation.
[0191]
when generating the matrix H '(omega), the audio processing apparatus 81 is configured as shown in FIG. 15 for example. Note that portions corresponding to the case in FIG. 8 in FIG. 15 are denoted by identical reference numerals, and description thereof will be omitted as appropriate.
[0192]
Speech processing apparatus shown in FIG. 15 81, the head direction sensor section 91, the head direction selecting portion 92, a matrix generating unit 201, HRTF synthesis unit 93, an annular conditioning inverse transform unit 94, and the time-frequency inverse conversion unit It has a 95.
[0193]
Configuration of the audio processing apparatus 81 shown in FIG. 15 has a structure in which a further matrix generating unit 201 to the audio processing unit 81 shown in FIG.
[0194]
Matrix generating unit 201, holds in advance HRTFs order which does not depend on individual, to get the HRTF orders that depend on individuals from outside, the head-related transfer function obtained, prestored and which generates a matrix H '(ω) from the order of the HRTF that is independent of the individual, and supplies the HRTF synthesis unit 93.
[0195]
Subsequently, with reference to the flowchart of FIG. 16, the drive signal generation processing will be described which is performed by the sound processing apparatus 81 having the configuration shown in FIG. 15.
[0196]
In step S71, the matrix generation unit 201 performs user settings. For example the matrix generator 201, in response to the input operation by a user or the like performs user setting that identifies the information about the listener who listens to audio currently played.
[0197]
The matrix generation unit 201 according to the user setting, the listener who listens to audio currently being reproduced, that is, the user obtains the head related transfer function of the user's order that individuals depend from an external device or the like. Incidentally, the head related transfer function of the user, for example to the time a user setting may be those designated by the input operation by the user or the like, may be those determined on the basis of the information defined in the user settings.
[0198]
In step S72, the matrix generator 201 generates a matrix H HRTF '(omega), and supplies the HRTF synthesis unit 93.
[0199]
That is, the matrix generation unit 201 obtains the head related transfer function of the following number of individuals dependent matrix from a head transfer function obtained that, the order of the HRTF that is independent of the individual stored in advance H 'generates (omega), and supplies the HRTF synthesis unit 93. In this case, the matrix generator 201, held in advance based on the information indicating the necessary degree m of each time-frequency bin omega are the composed of only the order of the elements required matrix H '(ω), time-frequency bins It is generated for each ω.
[0200]
Then, after that, although the driving signal generation processing process performed by the steps S73 to step S77 is terminated, since these processes are the same as those of steps S11 to S15 in FIG. 9, a description thereof will be omitted. In these steps S73 to step S77, the the head transfer functions convolved with the input signal in an annular conditioning region, the driving signal of the headphone is generated. The generation of the matrix H '(ω) may be performed in advance, the input signal may be performed from the supply.
[0201]
Audio processing apparatus 81 as described above, the convolution of the head related transfer function to an input signal in an annular conditioning region, the convolution result by performing a cyclic harmonic inverse transform on, calculates the drive signals of the left and right headphones.
[0202]
Thus, by performing the convolution of the head related transfer function in an annular conditioning region, with a calculation amount at the time of generating the driving signal of the headphone can be greatly reduced, even significantly the amount of memory which is needed at the time of operation it can be reduced. In other words, it is possible to reproduce more efficiently speech.
[0203]
In particular, the audio processing apparatus 81, since the head-related transfer function of the following number of individuals dependent and to generate a matrix H '(ω) acquired from the outside, it is possible not only to further reduce the amount of memory, it can be reproduced properly sound field using a head related transfer function suitable for the individual user.
[0204]
Incidentally, applying the technique of generating the audio processing unit 81, individuals dependent acquires HRTF orders from the outside consisting of only the following number of elements required matrix H 'the (omega) here example It was described. However, the invention is not limited to such an example, it may not be performed unnecessary reduction in the order.
[0205]
Incidentally, in the discussions went above, annularly with respect to a virtual speaker positioned What plane with respect to the head transfer function and the initial head direction holding It is either placed does not matter.
[0206]
For example, the arrangement position of the virtual speaker for HRTF and the initial head position for holding may be a horizontal plane as indicated by an arrow A111 in FIG. 17, on the median plane as indicated by the arrow A112 it may be, or may even on coronal as indicated by the arrow A113. In other words, what ring around the head center of the listener (hereinafter, referred to as ring A) virtual speaker may be placed on.
[0207]
In the example shown by the arrow A111, the virtual speaker ring RG11 on a horizontal plane around the head of the user U11 are arranged annularly. Further, in the example shown by arrow A 112, a virtual speaker ring RG12 on median plane around the head of the user U11 are arranged annularly, in the example shown by the arrow A113, around the head of the user U11 coronary virtual speakers are arranged annularly in the ring RG13 on the surface.
[0208]
The arrangement position of the virtual speaker for HRTF and the initial head direction to hold, for example, as shown in FIG. 18, in a direction perpendicular to the plane that contains certain ring A, move the ring A and it may be the position. In the following, it will be referred to as those of moving such a ring A and ring B. Note that portions corresponding to the case in FIG. 17 in FIG. 18 are denoted by identical reference numerals, and description thereof will be omitted as appropriate.
[0209]
In the example shown by the arrow A121 in FIG. 18, in the drawing the ring RG11 on a horizontal plane around the head of the user U11, virtual speaker ring RG21 or ring RG22 is moved in the vertical direction are arranged annularly. In this example, ring RG21 or ring RG22 is ring B.
[0210]
Further, in the example shown by arrow A 122, in FIG ring RG12 on median plane around the head of the user U11, virtual speaker ring RG23 or ring RG24 is moved in the depth direction are arranged annularly. In the example shown by the arrow A123, in the drawing the ring RG13 on coronal around the head of the user U11, virtual speaker ring RG25 or ring RG26 is moved in the lateral direction are arranged annularly.
[0211]
Furthermore, the arrangement of virtual speakers for HRTF and the initial head direction to hold, as shown in FIG. 19, if there is an input for each of a plurality of rings arranged in a predetermined direction, for each of the rings it can Crossed aforementioned systems. However, sharable ones sensors and headphones may be properly shared. Note that portions corresponding to the case in FIG. 18 in FIG. 19 are denoted by identical reference numerals, and description thereof will be omitted as appropriate.
[0212]
For example, in the example shown by the arrow A131 in FIG. 19, it is to partner in the figure, the ring RG11 arranged vertically, ring RG21, and the above-described systems in a ring RG22. Similarly, in the example shown in the arrow A 132, in the figure the ring RG12 arranged in the depth direction, ring RG23, and each ring RG24 can Crossed the system described above, in the example shown by the arrow A133, in the figure, in the left-right direction lined ring RG13, ring RG25, and each ring RG26 can Crossed systems described above.
[0213]
Furthermore, as shown in FIG. 20, a group of ring A with a plane including a straight line is passing through the head center of the user U11 is listener (hereinafter, referred to as ring Adi) for, diagonalized head portion matrix of the transfer function H'i the (omega) may be more prepared. Note that portions corresponding to the case in FIG. 19 in FIG. 20 are denoted with the same reference numerals, and description thereof will be omitted as appropriate.
[0214]
In the example shown in FIG. 20, for example, in the example shown in each of the arrows A141 or arrows A143, each of a plurality of circles located around the head of the user U11 represents a respective ring Adi.
[0215]
In this case, the input is a matrix H'i of HRTF regarding any ring Adi to the initial head direction (omega), by a change of the head direction of the user, optimal ring Adi matrix H'i (omega) is the process of choosing and thus be e Kuwawa relative aforementioned systems.
[0216]
The series of processes described above can be executed by hardware or can be executed by software. When executing the series of processing by software, a program constituting the software is installed into a computer. Here, the computer includes a computer incorporated in dedicated hardware, by installing various programs, which can execute various functions include, for example, such as a general-purpose computer.
[0217]
Figure 21 is a block diagram showing a configuration example of hardware of a computer that executes the series of processes described above.
[0218]
In the computer, CPU (Central Processing Unit) 501, ROM (Read Only Memory) 502, RAM (Random Access Memory) 503 are connected to each other via a bus 504.
[0219]
The bus 504 is further output interface 505 is connected. Output interface 505, an input unit 506, output unit 507, recording unit 508, a communication unit 509, and a drive 510 are connected.
[0220]
Input unit 506 includes a keyboard, a mouse, a microphone, made of an imaging device. The output unit 507 includes a display and a speaker. Recording unit 508, a hard disk and a nonvolatile memory. Communication unit 509 including a network interface. Drive 510 drives a magnetic disk, an optical disk, a magneto-optical disk, or a removable recording medium 511 such as a semiconductor memory.
[0221]
Series In the computer configured as described above, CPU 501 is, for example, a program recorded in the recording unit 508 via the input-output interface 505 and the bus 504 and executes the loaded into RAM 503, the above-mentioned processing of is performed.
[0222]
Program computer (CPU 501) is executed, for example, can be provided by being recorded in a removable recording medium 511 as a package medium or the like. Further, the program may be provided via a local area network, the Internet, or digital satellite broadcasting, a wired or wireless transmission medium.
[0223]
In the computer, by mounting the removable recording medium 511 into the drive 510, it can be through the input-output interface 505, installed in the recording unit 508. The program via a wired or wireless transmission medium and received by the communication unit 509, can be installed in the recording unit 508. Alternatively, the program may be in the ROM502 and the recording unit 508 installed in advance.
[0224]
The program which the computer executes may be a program in which processes are performed in time series in the order described herein, at a necessary timing such as when the parallel or call was made processing may be a program to be carried out.
[0225]
Further, embodiments of the present technology is not limited to the embodiments described above, but various modifications are possible without departing from the scope of the present disclosure.
[0226]
For example, the present technology, sharing one function by a plurality of devices via a network, it is possible to adopt a configuration of cloud computing which processes jointly.
[0227]
Further, each step described in the above flowcharts may be executed by one device, it can be performed by allocating a plurality of apparatuses.
[0228]
Further, when a plurality of processes are included in one step, the plurality of processes included in the one step may be executed by one device, it can be performed by allocating a plurality of apparatuses.
[0229]
The effects described herein are not intended to be limited to a merely illustrative, there may be other effects.
[0230]
Additionally, the present technology may also be configured as follows.
[0231]
(1)
input signal of an annular conditioning region or the portion corresponding to the annular conditioning region of the input signal of spherical harmonic region, a head transfer function combining unit for combining the head related transfer function which is diagonalized,
wherein the signal obtained by synthesis by cyclic conditioner inverse conversion based on the cyclic harmonic, an annular conditioning inverse transformation unit for generating a headphone driving signals in the time frequency domain
audio processing apparatus comprising a.
(2)
the head related transfer function combining unit, and the diagonal matrix obtained a matrix of a plurality of HRTF diagonalized by an annular harmonic transform, the input corresponding to each order of the cyclic harmonics by obtaining the product of a vector comprising a signal, it combines the HRTF said input signal and is the diagonalized
speech processing apparatus according to (1).
(3)
the head-related transfer function combining unit, by using only the order of elements in the time of a predetermined configurable for each frequency of the diagonal elements of the diagonal matrix is the diagonalized and the input signal to synthesize the head transfer functions
sound processing apparatus according to (2).
(4)
the diagonal matrix, the diagonalized the HRTF is used in common by each user is included as an element
audio processing apparatus according to (2) or (3).
(5)
the diagonal matrix, said diagonalized the HRTF depends on the individual user is included as an element
(2) to the sound processing apparatus according to any one of (4).
(6)
constituting the diagonal matrix, as well as previously held the diagonalized the HRTF common to the user, and obtains the head transfer function said diagonalized depends on individual users Te, and said diagonalized the HRTF acquisition, further comprising a matrix generator for generating the diagonal matrix from a head transfer function the diagonalized held in advance
(2) or speech processing apparatus according to (3).
(7)
the annular conditioner inverse transform unit holds the annular harmonic function matrix comprising a circular harmonics in each direction, based on the row corresponding to the predetermined direction of the spherical harmonic function matrix, said annular conditioner inverse transform It is carried out
(1) to the sound processing apparatus according to any one of (6).
(8)
further comprising a head direction acquisition unit that acquires the direction of the head of the user who listens to audio based on headphone drive signal,
said annular conditioner inverse transform unit, the head of the user in the annular harmonic function matrix based on the row corresponding to the direction of performing the cyclic harmonic inverse transform
speech processing apparatus according to (7).
(9)
further comprising a head direction sensor which detects the rotation of the user's head,
the head direction obtaining unit, by acquiring the detection result of the head direction sensor unit, the head of the user to obtain the orientation of the
sound processing apparatus according to (8).
(10)
the headphone further includes a drive signal Time Time frequency inverse conversion section for frequency inversion
audio processing apparatus according to any one of (1) to (9).
(11)
an annular input signal conditioning region or the portion corresponding to the annular conditioning region of the input signal of spherical harmonic regions, by combining the head related transfer function which is diagonalized,
signal obtained by the synthesis the by cyclic conditioner inverse conversion based on the cyclic harmonics, and generates a headphone driving signals in the time frequency domain
audio processing method comprising the steps.
(12)
the input signal of the annular conditioning region or the portion corresponding to the annular conditioning region of the input signal of spherical harmonic regions, by combining the diagonalized the HRTF,
the signal obtained by the synthesis the by converting cyclic conditioner contrary based on cyclic harmonics, and generates a headphone driving signals in the time frequency domain
program for executing the processing including a step in the computer.
DESCRIPTION OF SYMBOLS
[0232]
81 audio processor, 91 head direction sensor unit, 92 head direction selection unit, 93 HRTF synthesis unit, 94 an annular conditioning inverse transform unit, 95 hours inverse frequency transformation unit, 201 matrix generator
The scope of the claims
[Requested item 1]
Input signals of the annular conditioning region or the portion corresponding to the annular conditioning region of the input signal of spherical harmonic region, and HRTF synthesis unit for synthesizing the diagonalized the HRTF,,
by the synthesis the resulting signals by cyclic conditioner inverse conversion based on the cyclic harmonic, an annular conditioning inverse transformation unit for generating a headphone driving signals in the time frequency domain
audio processing apparatus comprising a.
[Requested item 2]
The HRTF synthesis unit includes a diagonal matrix obtained by a matrix comprising a plurality of HRTF diagonalized by an annular harmonic transform, comprising the input signal corresponding to each order of the cyclic harmonics by obtaining the product of the vector, combining the said input signal and said diagonalized the HRTF
sound processing apparatus according to claim 1.
[Requested item 3]
Head the head transfer function combining unit, by using only the order of elements in the time of a predetermined configurable for each frequency of the diagonal elements of the diagonal matrix, which is the diagonalized and the input signal to synthesize the transfer function
sound processing apparatus according to claim 2.
[Requested item 4]
The pair is diagonal matrix, commonly the diagonalized the HRTF used by is included as an element in each user
speech processing apparatus according to claim 2.
[Requested item 5]
The pair is diagonal matrix, said diagonalized the HRTF depends on the individual user is included as an element
audio processing apparatus according to claim 2.
[Requested item 6]
Constituting the diagonal matrix, as well as previously held the diagonalized the HRTF common to the user, to obtain an HRTF said diagonalized depends on individual user, acquires from then the head-related transfer function said diagonalized was, with the diagonalized the HRTF stored in advance, further comprising a matrix generator for generating the diagonal matrix
speech according to claim 2 processing apparatus.
[Requested item 7]
Said annular conditioner inverse transform unit holds the annular harmonic function matrix comprising a circular harmonics in each direction, based on the row corresponding to the predetermined direction of the annular harmonic function matrix, it performs the annular harmonic inverse transformation
according speech processing apparatus according to claim 1.
[Requested item 8]
Further comprising a head direction acquisition unit that acquires the direction of the head of the user who listens to audio based on the headphone drive signal,
said annular conditioner inverse transformation unit in the direction of the head of the user in the annular harmonic function matrix based on the corresponding row, performs said annular harmonic inverse transform
sound processing apparatus according to claim 7.
[Requested item 9]
Further comprising a head direction sensor which detects the rotation of the head of the user,
the head direction obtaining unit, by acquiring the detection result of the head direction sensor section, the direction of the head of the user acquiring
speech processing apparatus according to claim 8.
[Requested item 10]
Further comprising the time the headphone drive signal to the time frequency inversion frequency inversion unit
sound processing apparatus according to claim 1.
[Requested item 11]
Input signals of the annular conditioning region or the portion corresponding to the annular conditioning region of the input signal of spherical harmonic region, and a head transfer function diagonalized, synthesized,
cyclic harmonic signals obtained by the synthesis by cyclic conditioner inverse transformation based on a function, it generates a headphone driving signals in the time frequency domain
audio processing method comprising the steps.
[Requested item 12]
Input signals of the annular conditioning region or the portion corresponding to the annular conditioning region of the input signal of spherical harmonic region, and a head transfer function diagonalized, synthesized,
cyclic harmonic signals obtained by the synthesis by cyclic conditioner inverse transformation based on a function, it generates a headphone driving signals in the time frequency domain
program for executing the processing including a step in the computer.
| # | Name | Date |
|---|---|---|
| 1 | 201817024324-TRANSLATIOIN OF PRIOIRTY DOCUMENTS ETC. [29-06-2018(online)].pdf | 2018-06-29 |
| 2 | 201817024324-STATEMENT OF UNDERTAKING (FORM 3) [29-06-2018(online)].pdf | 2018-06-29 |
| 3 | 201817024324-PROOF OF RIGHT [29-06-2018(online)].pdf | 2018-06-29 |
| 4 | 201817024324-PRIORITY DOCUMENTS [29-06-2018(online)].pdf | 2018-06-29 |
| 5 | 201817024324-POWER OF AUTHORITY [29-06-2018(online)].pdf | 2018-06-29 |
| 6 | 201817024324-FORM 1 [29-06-2018(online)].pdf | 2018-06-29 |
| 7 | 201817024324-DRAWINGS [29-06-2018(online)].pdf | 2018-06-29 |
| 8 | 201817024324-DECLARATION OF INVENTORSHIP (FORM 5) [29-06-2018(online)].pdf | 2018-06-29 |
| 9 | 201817024324-COMPLETE SPECIFICATION [29-06-2018(online)].pdf | 2018-06-29 |
| 10 | 201817024324-OTHERS-060718.pdf | 2018-07-09 |
| 11 | 201817024324-Correspondence-060718.pdf | 2018-07-09 |
| 12 | abstract.jpg | 2018-08-03 |
| 13 | 201817024324.pdf | 2018-09-25 |
| 14 | 201817024324-FORM 3 [15-10-2018(online)].pdf | 2018-10-15 |
| 15 | 201817024324-FORM 18 [20-12-2019(online)].pdf | 2019-12-20 |
| 16 | 201817024324-OTHERS [11-06-2021(online)].pdf | 2021-06-11 |
| 17 | 201817024324-FER_SER_REPLY [11-06-2021(online)].pdf | 2021-06-11 |
| 18 | 201817024324-DRAWING [11-06-2021(online)].pdf | 2021-06-11 |
| 19 | 201817024324-CORRESPONDENCE [11-06-2021(online)].pdf | 2021-06-11 |
| 20 | 201817024324-COMPLETE SPECIFICATION [11-06-2021(online)].pdf | 2021-06-11 |
| 21 | 201817024324-CLAIMS [11-06-2021(online)].pdf | 2021-06-11 |
| 22 | 201817024324-FER.pdf | 2021-10-18 |
| 23 | 201817024324-US(14)-HearingNotice-(HearingDate-07-08-2023).pdf | 2023-07-19 |
| 24 | 201817024324-Correspondence to notify the Controller [07-08-2023(online)].pdf | 2023-08-07 |
| 25 | 201817024324-Correspondence to notify the Controller [07-08-2023(online)]-1.pdf | 2023-08-07 |
| 1 | 2021-02-2416-25-25E_24-02-2021.pdf |