Abstract: The present technology relates to a microphone array, a recording device and method, and a program which enable wide-band sound field recording to be performed at low cost. The microphone array is for sound field recording configured from a plurality of sub-arrays. Further, each sub-array is configured from a plurality of microphones and has a shape with discrete rotational symmetry having a prescribed radius, and when the radii of the plurality of sub-arrays are arranged as a progression, the progression is a generalized arithmetic progression. The present technology is applicable to microphone arrays and recording devices, for example.
Title of Invention: Microphone Array, Recording Device and Method, and Program
Technical field
[0001]
The present technology relates to microphone arrays, recording devices and methods, and programs, and more particularly to microphone arrays, recording devices and methods, and programs that enable wideband sound field recording at low cost.
Background technology
[0002]
In recent years, recording and reproduction of sound wave surfaces has become common in the audio industry. Wave field synthesis / restoration technology enables sound image localization and spatial noise canceling of objects placed in space, and can provide a more realistic acoustic experience compared to conventional multi-channel playback methods. it can.
[0003]
For example, an open annular microphone array composed of omnidirectional microphones is used in various applications.
[0004]
However, the design of the microphone arrangement of such an annular microphone array is not suitable for recording a wave surface (sound field) in a wide frequency band. The reason is that when the annular microphone array is used, the mode function known as the Bessel function for obtaining the spherical harmonic coefficient of the wave surface of the recorded sound becomes zero in a predetermined frequency band.
[0005]
Therefore, for example, a plurality of microphones are arranged in a double or more multiple ring, a cardioid directional microphone is used (see, for example, Non-Patent Document 1), or a rigid body baffle is used. It is conceivable to reduce the area where becomes zero.
[0006]
In addition, there are some array recording methods based on omnidirectional microphones (see, for example, Non-Patent Document 2, Non-Patent Document 3, and Patent Documents 1 to 3).
Prior art literature
Non-patent literature
[0007]
Non-Patent Document 1: G. Huang, "Design of robust concentric circular differential microphone arrays", The Journal of the Acoustical Society of America, 2017.
Non-Patent Document 2: Z. Prime and C. Doolan, "A comparison of popular beamforming" arrays ", Proceedings of Acoustics 2013 Victor Harbor: Science Technology and Amenity, Annual Conference of the Australian Acoustical Society, 2013.
Non-Patent Document 3: D. Mandal, SP Ghoshal and AK Bhattacharjee," Concentric circular antenna array synthesis using Particle Swarm Optimization with Concentricion Factor Approach ", Indian Antenna Week: A Workshop on Advanced Antenna Technology, 2010.
Patent documents
[0008]
Patent Document 1: US Pat. No. 6,205,224
Patent Document 2: Japanese
Patent Application Laid-Open No. 2005-521283 Patent Document 3: Japanese Patent Application Laid-Open No. 2011-15050
Outline of the invention
Problems to be solved by the invention
[0009]
However, with the above-mentioned technique, it is difficult to record a wide band sound field at low cost.
[0010]
For example, the approach of arranging multiple microphones in multiple rings, using cardioid directional microphones, or using a rigid baffle can provide sound field recording in a sufficiently wide frequency band in many situations. It may not be possible, or it may be difficult to realize due to cost and physical constraints.
[0011]
Further, the techniques described in Non-Patent Document 2 and Patent Documents 1 to 3 are for reducing the side lobe for beamforming, and the techniques described in Non-Patent Document 3 are not intended for sound. is there. Therefore, these array recording methods are not suitable for recording for wave surface reproduction.
[0012]
This technology was made in view of such a situation, and enables low-cost, wide-band sound field recording.
Means to solve problems
[0013]
The microphone array of the first aspect of the present technology is a microphone array for sound field recording composed of a plurality of sub-arrays, and the sub-array is a discrete array composed of a plurality of microphones and having a predetermined radius. The shape is rotationally symmetric, and when the radii of the plurality of subarrays are arranged to form a sequence, the sequence becomes a generalized arithmetic progression.
[0014]
In the first aspect of the present technology, the microphone array is a microphone array for sound field recording composed of a plurality of sub-arrays, and the sub-array is a discrete array composed of a plurality of microphones and having a predetermined radius. The shape is rotationally symmetric, and when the radii of the plurality of subarrays are arranged to form a sequence, the sequence becomes a generalized arithmetic progression.
[0015]
The recording device of the second aspect of the present technology calculates the spherical harmony coefficient based on the multi-channel signal obtained by collecting the sound from the microphone array for sound field recording composed of a plurality of sub-arrays. The sub-array includes a coefficient calculation unit, and the sub-array is a discrete rotationally symmetric shape having a predetermined radius while being composed of a plurality of microphones. When the radii of the plurality of sub-arrays are arranged in a sequence The sequence is a generalized arithmetic progression.
[0016]
The recording method or program of the second aspect of the present technology is a recording method or program corresponding to the recording device of the second aspect of the present technology.
[0017]
In the second aspect of the present technology, the spherical harmonics are calculated based on the multi-channel signal obtained by collecting the sound from the microphone array for sound field recording composed of a plurality of sub-arrays. Further, the sub-array is composed of a plurality of microphones and has a discrete rotationally symmetric shape having a predetermined radius, and when the radii of the plurality of sub-arrays are arranged to form a sequence, the sequence is generally used. It is made to be an arithmetic progression.
The invention's effect
[0018]
According to the first aspect and the second aspect of the present technology, it is possible to perform wideband sound field recording at low cost.
[0019]
The effects described here are not necessarily limited, and may be any of the effects described in the present disclosure.
A brief description of the drawing
[0020]
[Fig. 1] Fig. 1 is a diagram for explaining the value of a mode function according to the arrangement of microphones.
[Fig. 2] Fig. 2 is a diagram for explaining the value of the mode function according to the arrangement of microphones.
[Fig. 3] Fig. 3 is a diagram showing a configuration example of a microphone array to which this technology is applied.
[Fig. 4] Fig. 4 is a diagram illustrating the arrangement of microphones.
[Fig. 5] Fig. 5 is a diagram showing a configuration example of a microphone array to which this technology is applied.
[Fig. 6] Fig. 6 is a diagram showing a configuration example of a microphone array to which this technology is applied.
[Fig. 7] Fig. 7 is a diagram showing a configuration example of a microphone array to which this technology is applied.
[Fig. 8] Fig. 8 is a diagram showing a configuration example of a microphone array to which this technology is applied.
[Fig. 9] Fig. 9 is a diagram for explaining the value of the mode function according to the arrangement of the microphones.
[Fig. 10] Fig. 10 is a diagram illustrating the number of conditions according to the arrangement of microphones.
[Fig. 11] Fig. 11 is a diagram illustrating the number of conditions according to the arrangement of microphones.
[Fig. 12] Fig. 12 is a diagram illustrating the number of conditions according to the arrangement of microphones.
[Fig. 13] Fig. 13 is a diagram showing a configuration example of a recording system and a playback system to which this technology is applied.
FIG. 14 is a flowchart illustrating a recording process.
FIG. 15 is a flowchart illustrating a reproduction process.
[Fig. 16] Fig. 16 is a diagram showing a configuration example of a computer.
Mode for carrying out the invention
[0021]
The
present technology enables the recording and reproduction of a planar sound field to be realized over a wide frequency band by the geometric arrangement of the microphone array.
[0022]
In this technique, the arrangement of each microphone, that is, the microphone unit in the microphone array can be determined parametrically. The arrangement parameters that determine the arrangement of the microphone units may be appropriately determined according to various use cases. For example, a microphone array is composed of a plurality of discrete rotationally symmetric sub-arrays composed of a plurality of microphones, and the shapes of the sub-arrays are similar.
[0023]
According to this technology, the robustness against errors such as microphone installation error and manufacturing variation is improved, and sound field recording and reproduction, that is, sound wave surface recording and reproduction is performed for a wider frequency band. Can be done. In addition, microphone unit performance requirements such as microphone cost and signal-to-noise ratio (SN ratio) can be easily satisfied.
[0024]
Hereinafter, embodiments to which the present technology is applied will be described with reference to the drawings.
[0025]
The microphone array for sound field recording to which this technology is applied is basically a substantially annular microphone array in which each microphone is arranged on a two-dimensional plane so as to surround the center of the microphone array. However, the present invention is not limited to this, and a microphone array may be used for recording a three-dimensional sound field in which microphones are arranged in a three-dimensional space.
[0026]
That is, when the microphones are arranged in a three-dimensional space, for example, in the microphone array to which the present technology is applied, a substantially spherical microphone in which each microphone is arranged in the three-dimensional space so as to surround the center of the microphone array. It can be an array or the like.
[0027]
In the following, the microphone array to which the present technology is applied will be described as having a structure obtained by arranging each microphone on a two-dimensional plane.
[0028]
If there is a zero point of the Bessel function when converting the signal of the wave surface (sound field) of the sound recorded by the microphone array into the signal of the spherical harmonic region, a frequency band in which the conversion cannot be performed accurately occurs.
[0029]
For example, when the microphone array is a single ring or a double ring, a frequency band in which the value of the mode function, that is, the Bessel function becomes zero occurs as shown in FIG.
[0030]
In FIG. 1, the horizontal axis represents the wave number, and the vertical axis represents the order of the spherical harmonic region. Further, the shading in FIG. 1 indicates the value of the Bessel function, and in particular, the black portion indicates the region where the value of the Bessel function is 0 (zero).
[0031]
More specifically, the value of the Bessel function shown in FIG. 1 is the maximum value among the values of the Bessel function for each microphone constituting the microphone array. The value of the Bessel function of each microphone varies depending on the distance from the center of the microphone array to the microphone.
[0032]
In FIG. 1, the portion indicated by the arrow Q11 shows the value of the Bessel function in each region corresponding to the wave number and the order when the microphone array is a single ring. In this example, it can be seen that the Bessel function is zero in many regions, such as the region indicated by arrow A11, and a frequency band in which the wave surface cannot be accurately recorded and reproduced occurs.
[0033]
On the other hand, the part indicated by the arrow Q12 shows the value of the Bessel function in each region corresponding to the wave number and the order when the microphone array is a double ring. In this example, it can be seen that the region where the Bessel function becomes zero is smaller than the example shown by arrow A11. However, there are many small values where the Bessel function is close to zero, which can adversely affect the recording and reproduction of the wave front.
[0034]
Similarly, for example, as shown in FIG. 2, by using a cardioid directional microphone as the microphone constituting the microphone array, the region where the Bessel function becomes zero can be reduced, but the cost increases.
[0035]
In FIG. 2, the horizontal axis represents the wave number, and the vertical axis represents the order of the spherical harmonic region. Further, the shading in FIG. 2 indicates the value of the mode function, that is, the Bessel function, and the black portion particularly indicates the region where the value of the Bessel function is 0 (zero). More specifically, the value of the Bessel function shown in FIG. 2 is the maximum value of the values of the Bessel function for each microphone constituting the microphone array.
[0036]
In the example shown in FIG. 2, at each wave number below a certain order, there is almost no region where the Bessel function becomes zero as compared with the example shown in FIG. 1. However, when a cardioid directional microphone is used, The cost will be high.
[0037]
In addition, some array recording methods based on omnidirectional microphones have been proposed so far, but these methods are not suitable for recording for sound wave surface reproduction.
[0038]
On the other hand, there is also a way to simply avoid the zero point of the Bessel function. For example, when using a double annular microphone array, if the Bessel function value is zero in one annular microphone array, but the Bessel function value is not zero in the other annular microphone array, it is zero. It is also possible to use the value of the Bessel function that is not. However, this method cannot obtain a signal in the spherical harmonic region with sufficient accuracy.
[0039]
In general, sensor noise and environmental noise inherent in microphones cannot be avoided. Furthermore, due to microphone installation errors and manufacturing variations, it is difficult to accurately match the actual placement position of the microphone with the placement position indicated by the theoretically designed coordinates.
[0040]
These errors such as noise, installation error, and manufacturing variation are magnified by division by a small value of the Bessel function when the process of reproducing the recorded wave surface is performed, which adversely affects the numerical calculation. Therefore, it is important not only to avoid the zero point of the Bessel function, but also to optimize and analyze the error tolerance when designing the microphone arrangement.
[0041]
In other words, more accurate recording and reproduction of wave surfaces requires high error tolerance, that is, robustness against errors. In particular, considering the cost, it is necessary to design a microphone arrangement with high error tolerance on the premise of using the minimum necessary omnidirectional microphone in consideration of physical restrictions and ease of signal processing.
[0042]
Here, the wave surface of the sound, that is, the recording and reproduction of the sound field will be described. In the following, the microphones constituting the microphone array will also be referred to as a microphone unit.
[0043]
For example, recording and reproduction of a wave surface of sound can be realized by obtaining a spherical harmonic coefficient of the wave surface.
[0044]
Specifically, when using an annular microphone array for recording the wave surface, if the sampling theorem is satisfied, the spherical harmony coefficient a mn (k) is the sound pressure of the wave surface p k (r, θ q , φ q ). Obtained by sampling at each of the Q points.
[0045]
The component of the ring-shaped microphone array that depends on the radius r contained in the sound pressure p k (r, θ q , φ q ) is divided by b n (kr), which is a component that depends on the radius r . Will be removed.
[0046]
That is, the spherical harmonic coefficient a mn (k) can be obtained by the following equation (1).
[0047]
[Number 1]
[0048]
In Eq. (1), n and m indicate the order of the spherical harmonic region, and q is an index indicating each of the Q points for sampling the sound pressure (however, q = 0, ..., Q-1). In the following, the sampling point indicated by the index q will also be referred to as the point q.
[0049]
Further, k indicates the wave number, and r indicates the radius of the annular microphone array, that is, the distance from the center position of the annular microphone array to the microphone unit. θ q and φ q indicate the elevation angle and the azimuth angle indicating the direction in which the microphone unit located at the point q is directed, respectively.
[0050]
Furthermore, in equation (1), f indicates the frequency, c s indicates the speed of sound, b n (kr) indicates the mode function, and Y * m n (θ q , φ q ) is It shows a spherical harmonic basis. In particular, when the annular microphone array is composed of omnidirectional microphones, the mode function b n (kr) becomes a sphere Bessel function. In the following, b n (kr) will also be referred to simply as the Bessel function. In addition, "*" in the spherical harmonic basis Y * m n (θ q , φ q ) indicates the complex conjugate.
[0051]
An example of the case where the annular microphone array is composed of omnidirectional microphones is described in detail in, for example, "B. Rafaely, Fundamentals of Spherical Array Processing, Springer, 2015." (hereinafter, also referred to as Reference 1). Has been done.
[0052]
Further, the arithmetic processing of the equation (1), particularly the division processing by the Bessel function b n (kr) in the equation (1), may be cited as mode compensation. The mode compensation is described in detail in, for example, "DP Jarrett, EA Habets and PA Naylor, Theory and Applications of Spherical Microphone Array Processing, Springer, 2017." (hereinafter, also referred to as Reference 2).
[0053]
When the sound is picked up at each point q by the annular microphone array and the sound wave surface is recorded by obtaining the sound pressure p k (r, θ q , φ q ) at those points q , the obtained sound pressure p k ( From r, θ q , φ q ), the spherical harmony coefficient a mn (k) can be obtained by Eq. (1) . Further, if the spherical harmonic coefficient a mn (k) obtained in this manner is transmitted to the reproduction system, the reproduction system reproduces the wave surface (sound field) of the sound by using the spherical harmonic coefficient a mn (k). be able to.
[0054]
By the way, when the value of the Bessel function b n (kr ) in the equation (1) is close to zero, a numerical problem called the Bessel zero problem occurs. That is, as explained below, when the value of the Bessel function b n (kr) approaches zero, the number of conditions of the transformation matrix for obtaining the spherical harmonic coefficient a mn (k) increases, and the accurate spherical harmonic coefficient increases. You will not be able to get a mn (k).
[0055]
If the microphone units that make up the annular microphone array are not on the same ring, that is, if the radii r q of the microphone units placed at each point q are different, the sound pressure sampled at the point q is p k (r q , θ q , φ q ) is expressed by the following equation (2). The radius r q of the microphone unit is the distance from the center of the annular microphone array to the microphone unit, that is, the distance to the point q.
[0056]
[Number 2]
[0057]
In this case, the spherical harmonic coefficient a mn (k) of each order and wave number is the distribution of the sound pressure p k (r q , θ q , φ q ) obtained at each point q , that is, the sound pressure p k (r q , It can be obtained by multiplying the vector p k consisting of θ q and φ q ) by B + k , which is the generalized inverse matrix of the transformation matrix B k .
[0058]
That is, for example, by performing the calculation of the following equation (3), a vector a (k) consisting of a spherical harmonic coefficient a mn (k) with respect to the wave number k can be obtained.
[0059]
[Number 3]
[0060]
In Eq. (3), B + k indicates the generalized inverse of the transformation matrix B k . The vector p k is a vector consisting of the sound pressure p k (r l , θ l , φ l ) at each point l as shown in the following equation (4) , where l = 0, ..., L. , L = Q-1. That is, in the equation (4), the index indicating the sampling point of the sound pressure is l, and this l corresponds to q described above.
[0061]
The transformation matrix B k is the Bessel function b n (kr l ) and the spherical harmonic function Y m for the degree n (where 0 ≤ n ≤ N) of each point l, as shown in the following equation (5). It is a matrix whose elements are the products of n (θ l , φ l ).
[0062]
[Number 4]
[0063]
[Number 5]
[0064]
In the above equation (3), in the division of mode compensation, Y * m n (θ q , φ q ) / b n (kr) shown in equation (1) is replaced by the generalized inverse matrix B + k. There is. In order to obtain an accurate spherical harmonic coefficient a mn (k ) from this equation (3) , the transformation matrix B k must be lossless and adverse conditions must be avoided. Here, whether the transformation matrix B k is a good condition or a bad condition can be evaluated by , for example , the number of conditions for the transformation matrix B k .
[0065]
Assuming that the minimum singular value and the maximum singular value of the transformation matrix B k are σ min (B k ) and σ max (B k ), respectively, the conditional number X (k) of the transformation matrix B k is calculated by the following equation (6). Obtainable.
[0066]
[Number 6]
[0067]
When calculating Eq. (3), when the observed vector, that is, here the vector p k , contains an error, the error increases by X (k) times when the conditional number is X (k). ..
[0068]
Therefore, the conditional number X (k) of the transformation matrix B k should be small, and a small conditional number X (k) means that it has strong error tolerance, that is, it has high robustness against errors. Empirically, if the number of conditions exceeds 100, depending on the application, the matrix becomes a bad condition. The error tolerance analysis of the annular microphone array and the spherical microphone array based on the conditional number is described in detail in, for example, Reference Document 1 described above.
[0069]
As described above, the spherical harmonics a mn (k) for reproducing the wave surface of the sound can be obtained by the calculation of the equation (3), but the arrangement of each microphone unit constituting the microphone array and the spherical harmonic region By appropriately setting the maximum value (maximum order) of the order n of, a transformation matrix B k with good conditions can be obtained.
[0070]
Therefore, in this technology, by appropriately setting the arrangement of the microphone units and the maximum order of the spherical harmonic region, a small number of omnidirectional microphone units can be used, that is, at low cost and high noise in a wide frequency band. It has become possible to realize resistance (error resistance).
[0071]
In particular, the recording and reproduction (reproduction) of the wave surface to which the present technology is applied is realized by the parametric microphone array design having the features described below and the frequency-dependent spatial resolution control.
[0072]
For example, the microphone array to which the present technology is applied has the following features F1 to F3. That is, the microphone array to which the present technology is applied is designed based on the following features F1 to F3.
[0073]
(Feature F1) The
microphone array is composed of a plurality of geometrically similar sub-arrays, and each sub-array has discrete rotational symmetry.
(Feature F2) The
microphone units are distributed at an equal angle when viewed from the center of the microphone array. and are
(wherein F3)
a generalized arithmetic progression when the sequence side by side radius of each sub-array
[0074]
The microphone array to which this technology is applied is composed of a plurality of sub-arrays, and each sub-array is composed of a plurality of microphone units.
[0075]
The microphone array may be composed of one sub-array, or the sub-array may be composed of one microphone unit.
[0076]
Further, although all the microphone units constituting the microphone array are basically omnidirectional microphones, some microphone units may be non-omnidirectional microphones.
[0077]
The feature F1 described above is that when the microphone array is composed of a plurality of sub-arrays, all the sub-arrays have a geometrically similar shape (microphone unit arrangement). Here, the geometrical similarity of the sub-arrays means that the arrangement of the microphone units constituting the sub-array is similar.
[0078]
For example, two subarrays that are geometrically similar means that one subarray is scaled, rotated, and inverted when at least one operation is performed on the other subarray. Say that.
[0079]
"Matching" here means that when an operation such as enlargement is performed on one sub-array, the arrangement position of each microphone unit of the sub-array after the operation matches the arrangement position of each microphone unit constituting the other sub-array. It is to be. In this case, the center position of each sub-array coincides with the center position of the microphone array.
[0080]
Further, each sub-array is an array having a discrete rotationally symmetric shape. That is, the sub-array is not continuously rotationally symmetric so that it always has the same shape even if it is rotated by an arbitrary angle, but when it is rotated by a predetermined angle around the center position of the sub-array, that is, the center position of the microphone array. , The sub-arrays before and after rotation have the same discrete rotational symmetry. In the microphone array, since each sub-array is discretely rotationally symmetric, the frequency characteristics can be flattened.
[0081]
Further, each sub-array is an array having a predetermined radius. In particular, here, it is assumed that the radii of all the microphone units constituting the sub-array are the same, and the radius is the radius of the sub-array. The radius of the microphone unit is the center position of the sub-array, that is, the distance from the center position of the microphone array to the microphone unit.
[0082]
Therefore, each of the plurality of microphone units constituting the sub-array is arranged at a position at a distance of the radius of the sub-array from the center position of the microphone array, that is, the center position of the sub-array.
[0083]
Feature F2 is a microphone unit that is projected on one ring centered on the center position of the microphone array, that is, on the circumference when all the microphone units constituting the microphone array are projected radially on the ring. Is evenly distributed. That is, the projected microphone units are arranged at equal intervals on the ring.
[0084]
Here, the projection position on the ring of the microphone unit is the intersection position of the straight line connecting (passing) the center position of the microphone unit and the microphone array and the ring (circle) to be the projection destination. That is, the position of the microphone unit on the ring when viewed from the center position of the microphone array is the position of the projection destination of the microphone unit.
[0085]
By having such a feature F2, it is not necessary to perform complicated signal processing after recording the wave surface. The omission of complicated signal processing due to such a property is described in detail in, for example, Reference Document 1 described above.
[0086]
Further, the feature F3 is a sequence in which some of the plurality of sub-arrays constituting the microphone array have different radii from each other, and the radii values of all the sub-arrays constituting the microphone array are arranged in ascending or descending order to form a sequence. Is a generalized arithmetic progression.
[0087]
In other words, feature F3 is that the microphone units are arranged at intervals corresponding to the tolerance of the generalized arithmetic progression in the direction toward the outside from the center of the microphone array, that is, the direction away from the center.
[0088]
For example, "Z. Prime and C. Doolan," A comparison of popular beamforming arrays ", Proceedings of Acoustics 2013 Victor Harbor: Science Technology and Amenity, Annual Conference of the Australian Acoustical Society, 2013." and US Pat. No. 6,205,224. Discloses a method of arranging a microphone unit with a radius determined by a logarithmic or geometrical sequence of numbers.
[0089]
However, when spatial resolution control is performed for each frequency, it is better to determine the radius of the subarray by the generalized arithmetic progression as in this technique than in these methods, in the region where the value of the Bessel function is zero or almost zero. The reduction effect is higher. That is, the conditional number X (k) of the transformation matrix B k becomes smaller.
[0090]
Further, if the microphone array is designed to have the feature F1 and the feature F2, by using some sub-arrays, it is possible to use the microphone array in a scalable manner according to the demand.
[0091]
It is assumed that several geometrically similar sub-arrays are used as the sub-arrays constituting the microphone array. In this case, for example, when the number of usable microphone units is sufficiently large, the microphone array is configured by three sub-arrays, and when the number of usable microphone units is small, the microphone array is configured by two sub-arrays. It is possible to use it.
[0092]
Further, the transformation matrix B k of the microphone array depends on the frequency, that is, the wave number k, and the spatial resolution of the transformation is appropriately set for each frequency in the operating frequency band in order to obtain accurate sound field information.
[0093]
For example, when trying to obtain the spherical harmonics a mn (k) by the calculation of Eq. (3) , in general, it is possible to obtain a more accurate spherical surface with higher spatial resolution by calculating up to a term of a larger order n. The harmonic coefficient a mn (k) can be obtained. However, for a component whose order n is equal to or higher than a predetermined order determined by the arrangement of microphone units, the value of the Bessel function is 0 or a value close to 0.
[0094]
Therefore, in the present technology, for each degree n having a predetermined size or more, the transformation matrix B is performed by performing a process of excluding (removing) the rows corresponding to the order n from the transformation matrix B k as spatial resolution control. The number of conditions for k was improved. In other words, the degree n used for the operation is limited, that is, the number of rows of the transformation matrix B k is limited.
[0095]
The advantage of this technology is that a wide band sound field (wave surface) can be recorded with high error tolerance, especially by using the minimum number of omnidirectional microphone units required.
[0096]
Spatial resolution control can not only improve error tolerance but also reduce the amount of calculation.
[0097]
Further, by configuring the microphone array to be composed of a plurality of sub-arrays, it is possible to increase the sampling density in the angular direction without using a small microphone unit. This is because, for example, by arranging a plurality of sub-arrays, the microphone units are projected radially on the ring centered on the center position of the microphone array rather than when the microphone units are arranged in a single ring. This is because the density of the projected microphone unit in the above can be increased.
[0098]
Furthermore, the microphone array to which this technology is applied has a self-similarity, that is, a fractal shape. Therefore, this technology has the scalability to configure a microphone array even when only a small number of microphone units can be used. In other words, it can be used in a scalable manner as described above.
[0099]
Next, a more specific configuration example of the microphone array to which the present technology is applied will be described. FIG. 3 is a diagram showing a configuration example of an embodiment of a microphone array to which the present technology is applied.
[0100]
The microphone array MA11 shown in FIG. 3 is a vortex-shaped microphone array composed of a plurality of omnidirectional microphone units. In FIG. 3, each point represents one microphone unit.
[0101]
In this example, the microphone array MA11 is composed of 128 microphone units, and these microphone units are arranged in a vortex shape.
[0102]
In the microphone array MA11, one sub-array is composed of 16 microphone units. That is, the microphone array MA11 is composed of eight sub-arrays having different radii from each other, and these eight sub-arrays are arranged concentrically.
[0103]
For example, a portion composed of 16 microphone units arranged in a ring shape is regarded as one sub-array SA11, and a portion composed of 16 microphone units arranged in a ring shape is regarded as one sub-array SA12.
[0104]
Further, the microphone array MA11 has the above-mentioned features F1 to F3.
[0105]
For example, each sub-array constituting the microphone array MA11 has a shape different only in scale and rotation angle. Specifically, for example, when the sub-array SA11 is enlarged and rotated by a predetermined angle, it matches the sub-array SA12.
[0106]
Further, in each sub-array, microphone units are arranged in an annular shape around the center position O11, whereby the sub-array has a discrete rotationally symmetric shape.
[0107]
A part of the microphone array MA11 is enlarged as shown in FIG. In FIG. 4, each circle represents one microphone unit. Further, in FIG. 4, the same numbers are shown in the circles representing the microphone units constituting the same sub-array.
[0108]
In the example shown in FIG. 4, for example, the microphone unit marked with the number “1” constitutes the sub-array SA11 shown in FIG. 3, and the microphone unit marked with the number “8” is shown in FIG. It constitutes the sub-array SA12.
[0109]
In particular, in this example, it can be seen that the sub-arrays are arranged adjacent to each other, and the sequence consisting of the radii of those sub-arrays is a generalized arithmetic progression. That is, for any sub-array, the difference in radius between adjacent sub-arrays is one of several predetermined values corresponding to the tolerance.
[0110]
In the following, the microphone array MA11 shown in FIG. 3 will also be referred to as a vortex-shaped microphone array.
[0111]
Further, here, an example in which the microphone array is composed of eight sub-arrays and the sub-array is composed of 16 microphone units has been described. However, in addition, for example, the microphone array may be composed of four sub-arrays and each sub-array may be composed of 32 microphone units, or the microphone array may be composed of two sub-arrays and each sub-array may be 64. It may be composed of an individual microphone unit.
[0112]
Further, the microphone array to which the present technology is applied is not limited to the one shown in FIG. 3 as long as it has the above-mentioned features F1 to F3, and may be any kind.
[0113]
Specifically, for example, the microphone array may have the configuration shown in FIG.
[0114]
That is, in the portion shown by the arrow Q31 in FIG. 5, a microphone array MA21 formed by arranging a plurality of omnidirectional microphone units in the shape of the outline of a flower is shown. In the part indicated by the arrow Q31, each point represents one microphone unit.
[0115]
The microphone array MA21 is composed of eight sub-arrays, and each sub-array is composed of 16 microphone units arranged in a ring shape.
[0116]
When a part of this microphone array MA21 is enlarged, it is shown by arrow Q32. In the portion indicated by the arrow Q32, each circle represents one microphone unit, and the same number is written in the circle representing the microphone units constituting the same sub-array.
[0117]
In this example, it can be seen that the eight sub-arrays constituting the microphone array MA21 are arranged concentrically, and each sub-array is arranged adjacent to each other.
[0118]
In particular, the sub-array consisting of the microphone unit marked with the number "2" and the sub-array consisting of the microphone unit marked with the number "8" are in the rotation angle, that is, in the rotation direction when the center position of the microphone array MA21 is centered. It can be seen that the microphone units are arranged in different positions, but the radii of their sub-arrays are the same.
[0119]
Similarly, the sub-array consisting of the microphone unit marked with the number "3" and the sub-array consisting of the microphone unit marked with the number "7" have the same radius although the rotation angles are different. Further, the sub-array consisting of the microphone unit marked with the number "4" and the sub-array consisting of the microphone unit marked with the number "6" also have the same radius although the rotation angle is different.
[0120]
Such a microphone array MA21 has a configuration having the above-mentioned features F1 to F3. Hereinafter, the microphone array MA21 will also be referred to as a flower-shaped microphone array.
[0121]
Further, the microphone array to which the present technology is applied may have the configurations shown in FIGS. 6, 7, and 8, for example.
[0122]
That is, for example, in the portion shown by the arrow Q41 in FIG. 6, a microphone array MA31 formed by arranging a plurality of omnidirectional microphone units so as to form a substantially vortex shape is shown. In the part indicated by the arrow Q41, each point represents one microphone unit.
[0123]
The microphone array MA31 is composed of eight sub-arrays, and each sub-array has the above-mentioned feature F1. In addition, each sub-array is composed of 16 microphone units arranged in a ring shape.
[0124]
When a part of the microphone array MA31 is enlarged, it is shown by arrow Q42. In the portion indicated by the arrow Q42, each circle represents one microphone unit, and the same number is written in the circle representing the microphone units constituting the same sub-array.
[0125]
In this example, the eight sub-arrays constituting the microphone array MA31 are arranged concentrically, and the rotation angle of the sub-arrays when arranging each sub-array is randomly determined.
[0126]
Further, for example, in the portion shown by the arrow Q51 in FIG. 7, a microphone array MA41 formed by arranging a plurality of omnidirectional microphone units so as to form a substantially vortex shape is shown. In the part indicated by the arrow Q51, each point represents one microphone unit.
[0127]
The microphone array MA41 is composed of eight sub-arrays, and each sub-array is composed of 16 microphone units arranged in a ring shape.
[0128]
When a part of this microphone array MA41 is enlarged, it is shown by arrow Q52. In the portion indicated by the arrow Q52, each circle represents one microphone unit, and the same number is written in the circle representing the microphone units constituting the same sub-array.
[0129]
In this example, the eight sub-arrays constituting the microphone array MA41 are arranged concentrically, and the rotation angle of the sub-arrays when arranging each sub-array is randomly determined.
[0130]
In the microphone array shown in FIGS. 6 and 7, the rotation angle of each sub-array is randomly determined, but the microphone array has the above-mentioned features F1 to F3. Hereinafter, such a microphone array will also be referred to as a microphone array having a random shape.
[0131]
Further, for example, FIG. 8 shows a microphone array MA51 formed by arranging a plurality of omnidirectional microphone units so as to form a triple ring. In FIG. 8, each point represents one microphone unit.
[0132]
The microphone array MA51 is composed of three sub-arrays, and each sub-array is composed of 43 microphone units arranged in a ring shape.
[0133]
In particular, in this example, the three sub-arrays constituting the microphone array MA51 are arranged concentrically, and when any sub-array is expanded or contracted and further rotated, it matches the other sub-arrays.
[0134]
According to the microphone array having the features F1 to F3 described above, the region where the value of the Bessel function becomes zero can be reduced, and the conditional number X (k) of the transformation matrix B k can be improved. For example, according to a vortex-shaped microphone array, a flower-shaped microphone array, and a random-shaped microphone array, it is possible to eliminate the region where the value of the Bessel function becomes zero as shown in FIG.
[0135]
In FIG. 9, the horizontal axis represents the wave number k, and the vertical axis represents the order n of the spherical harmonic region. Further, the shading in FIG. 9 indicates the value of the Bessel function, and in particular, the black portion indicates the region where the value of the Bessel function is 0 (zero). More specifically, the value of the Bessel function shown in FIG. 9 is the maximum value of the values of the Bessel function for each subarray constituting the microphone array.
[0136]
In FIG. 9, the portion indicated by the arrow Q61 shows the value of the Bessel function in each region corresponding to the wave number k and the order n when the vortex-shaped microphone array is used.
[0137]
The part indicated by the arrow Q62 shows the value of the Bessel function in each region corresponding to the wave number k and the order n when the flower-shaped microphone array is used. Further, the part indicated by the arrow Q63 shows the value of the Bessel function in each region corresponding to the wave number k and the order n when a randomly shaped microphone array is used.
[0138]
In the examples shown by these arrows Q61 to Q63, in the frequency band of 0 kHz to 8 kHz, there is no region where the value of the Bessel function existing in the example shown in FIG. 1 becomes zero at a certain order n or less. You can see that. When there is no region where the value of the Bessel function becomes zero in this way, the conditional number X (k) of the transformation matrix B k can be reduced, which improves error tolerance in a wide frequency band at low cost. Can be done.
[0139]
Further, when spatial resolution control is applied to a microphone array to which the present technology is applied , the condition of the transformation matrix B k becomes better as the distance between the microphone units projected on the ring becomes shorter, for example, as shown in FIG .
[0140]
In FIG. 10, the horizontal axis represents the frequency, and the vertical axis represents the conditional number X (k) of the transformation matrix B k . Further, in the example of FIG. 10, the conditional number X (k) is obtained when the spatial resolution control described later is performed.
[0141]
In this example, curves L11 to L14 are the vortex-shaped microphone array MA11 shown in FIG. 3, the flower-shaped microphone array MA21 shown in FIG. 5, the random-shaped microphone array MA31 shown in FIG. 6, and FIG. 7. The conditional number X (k) for the random-shaped microphone array MA41 shown in is shown.
[0142]
Here, since the flower-shaped microphone array MA21 has the shortest distance between the microphone units projected on the ring, the conditional number X (k) of the flower-shaped microphone array MA21 is the smallest in the entire frequency band. You can see that.
[0143]
On the other hand, in the vortex-shaped microphone array MA11, the distance between the microphone units becomes relatively long for every eight microphone units. That is, the microphone unit forming the most central sub-array constituting the microphone array MA11 and the microphone unit forming the sub-array on the farthest side from the center are arranged apart from each other.
[0144]
Therefore, the distance between the microphone units projected on the ring is longer than that in the microphone array MA21, and the conditional number X (k) of the vortex-shaped microphone array MA11 is the conditional number X (k) of the flower-shaped microphone array MA21. It is slightly larger than k).
[0145]
Further, in the randomly shaped microphone array MA31 and the microphone array MA41, the distance between the microphone units projected on the ring becomes relatively long. Therefore, the conditional number X (k) of the microphone array MA31 and the microphone array MA41 is larger than the conditional number X (k) of the vortex-shaped microphone array MA11.
[0146]
By the way, as described above, in the present technology, the arrangement of each microphone unit in the microphone array can be determined parametrically.
[0147]
Here, a parameter indicating the arrangement of each microphone unit of the microphone array is referred to as an arrangement parameter, and a set of a plurality of arrangement parameters is referred to as an arrangement parameter set. That is, the arrangement of each microphone unit constituting the microphone array is determined by the arrangement parameter set.
[0148]
Specifically, for example, as placement parameters, the number of sub-arrays S, the radius r s of each sub-array (where s = 0,1, ..., S-1), and the rotation angle of each sub-array φ s (where s = 0). , 1,…, S-1), etc. can be considered.
[0149]
Here, the number of sub-arrays S is the number of sub-arrays constituting the microphone array, and the radius r s of the sub-array is the distance from the center position of the microphone array to the microphone unit constituting the sub-array. In the following, the vector consisting of the radii r s of each of the S subarrays will also be referred to as the radius vector r sub .
[0150]
The rotation angle φ s of the sub-array is the tilt angle of the sub-array with reference to a predetermined direction viewed from the center position of the microphone array. In other words, the rotation angle φ s of the sub-array is an angle in the rotation direction indicating the position of the sub-array in the rotation direction centered on the center position of the microphone array.
[0151]
Specifically, for example, the center position of the microphone array is set as the center O, and the reference direction is a predetermined reference direction viewed from the center O. At this time, for example, the angle formed by the straight line connecting the microphone unit and the center O, which form the sub-array, and the reference direction is defined as the rotation angle φ s .
[0152]
For example, the reference direction can be the direction of the microphone unit that is the reference of the sub-array on the most central O side. In this case, the rotation angle φ s of the sub-array indicates how much the sub-array is rotated on the sub-array closest to the center O side.
[0153]
In the following, the vector consisting of the rotation angles φ s of each of the S sub-arrays will also be referred to as the rotation angle vector φ sub .
[0154]
In the following, the number of subarrays S, the radius vector r sub , and the rotation angle vector φ sub as the placement parameters are also described as the placement parameter set P Q opt = {S, r sub , φ sub }.
[0155]
For example, the optimum placement parameters depend on the total number of microphone units Q, the operating frequency band [f min , f max ], the diameter D m of the microphone units , and the upper limit X max of the conditional number X (k) .
[0156]
Here, the total number of microphone units Q is the number of microphone units constituting the microphone array. The total number of microphone units Q determines the number of sub-arrays constituting the microphone array, that is, the number of sub-arrays S.
[0157]
Specifically, for example, when the total number of microphone units Q is 24, the value of the number of subarrays S can be 1, 2, 3, 4, 6, 12, or 24.
[0158]
The operating frequency band [f min , f max ] is a frequency band from the minimum value f min to the maximum value f max of the frequency of the target sound .
[0159]
In determining the placement parameter set P Q opt, the number of conditions is taken into consideration in this operating frequency band [f min , f max ], and each placement parameter is optimized.
[0160]
The diameter D m of the microphone unit is the diameter of the microphone unit constituting the microphone array, and this D m is the lower limit of the absolute value of the tolerance of the generalized arithmetic progression that determines the radius vector r sub .
[0161]
For example, let the radii r s of any two subarrays be the radii r i and the radii r j (where i ≠ j). In this case, the radius r i and the radius r j must satisfy the following equation (7). This is because if the condition of Eq. (7) is not satisfied, two microphone units having a diameter of D m cannot be arranged side by side even if the radius r s and the rotation angle φ s of the sub-array are taken into consideration. is there.
[0162]
[Number 7]
[0163]
Further, the upper limit value X max is the value of the worst (larger) conditional number X (k) acceptable in the operating frequency band [f min , f max ].
[0164]
Although it depends on the application, empirically, when the conditional number X (k) exceeds 100, it becomes a bad condition and the inverse matrix becomes unstable. However, in many cases, multicollinearity is not desired, so in reality, the upper limit X max should be around 30.
[0165]
Based on the total number of microphone units Q, the operating frequency band [f min , f max ], the diameter D m of the microphone unit , and the upper limit value X max of the conditional number X (k) as described above, the features F1 to F3 are set. By determining the optimal placement parameter set P Q opt to have, a microphone array with the proper microphone unit placement can be obtained.
[0166]
Specifically, for example, the optimal placement parameter set P Q opt is between the operating frequency bands [f min , f max ] under the constraints of the total number of microphone units Q, diameter D m , and upper limit X max . It is obtained by minimizing the average conditional number of the transformation matrix B k .
[0167]
The search for the placement parameter set P Q opt is guided by a 100% search for the possible placement parameters. Empirically, metaheuristic optimization methods such as differential evolution can provide approximately optimal results.
[0168]
For optimization methods such as differential evolution, see, for example, "R. Storn and K. Price," Different Evolution --A Simple and Efficient Heuristic for global Optimization over Continuous Spaces ", Journal of Global Optimization, 1997." , Also referred to as Reference 3) and the like.
[0169]
Next, spatial resolution control in the microphone array will be described.
[0170]
For example, in Reference 1 and Reference 2, it is recommended to select an appropriate spatial resolution for each frequency band to obtain better robustness. This suggests that proper selection of spatial resolution leads to a well-conditioned transformation matrix.
[0171]
In fact, given any kr for the wavenumber k and the radius r of the microphone array, the degree n is always large when the degree n is greater than or equal to a large degree (hereinafter referred to as n 0 (kr)). As it becomes, the value of the corresponding mode function (Bessel function) approaches zero.
[0172]
For example, as shown in the following equation (8), let N arr be the degree n determined with respect to the total number of microphone units Q of the microphone array .
[0173]
[Number 8]
[0174]
In this case, the spherical harmonics up to N arr , whose degree n is greater than n 0 (kr) , that is , the elements of the transformation matrix B k , do not contain reliable information for wave plane reproduction. This is because for degrees n up to N arr greater than n 0 (kr) , the value of the Bessel function for those degrees n is zero or near zero.
[0175]
Therefore, in this technique, information loss is minimized and the condition of the transformation matrix B k is improved by excluding such numerically small spherical harmonics .
[0176]
In this case, as spatial resolution control, a process of limiting the rows (number of rows) of the transformation matrix B k used in the calculation for calculating the spherical harmonic coefficient a mn (k) including the mode compensation is performed.
[0177]
That is, for example, the radius r of each subarray s maximum value max (r of the s When), the transformation matrix B k transformation matrix B obtained by the spatial resolution control for n0 k is the transformation matrix B k n from first line of It is a matrix containing up to the 0 (k × max (r s )) line. In other words, the spatial resolution control, order n 0 (k × max (r s based on)), the transformation matrix B is used in the calculation k number of rows n 0 (k × max (r s )) on the line The transformation matrix B n0 k is obtained as a transformation matrix with a limited number of rows .
[0178]
Here, the n 0 (k × max (r s )) line of the transformation matrix B k is the line corresponding to the degree n 0 (k × max (r s )). Degree n 0 (k × max (r s )) is order n 0 (kr) for a subarray with a radius of max (r s ) . That is, the degree n 0 (k × max (r s )) is the degree n 0 (kr) when the radius r = max (r s ) .
[0179]
Any method can be used to determine the order n 0 (kr) for the radius r , for example, the threshold th value may be 1 or 1.1, and n 0 (kr) = th × r. It may be determined by the calculation of the following equation (9). For example , the method of setting n 0 (kr) = th × r is described in detail in Reference 1 and Reference 2 described above.
[0180]
[Number 9]
[0181]
The threshold value th in the equation (9) may be a real number between 0 and 1, and a value close to 1 is recommended. Specifically, for example, the threshold value th is 0.95. Further, in FIGS. 3, 5 to 8 and 10 described above and FIG. 12 described later, the order n 0 (kr s ) defined by the equation (9 ) is used.
[0182]
By performing such spatial resolution control, the conditions of the transformation matrix can be improved and the error tolerance can be improved. For example, when spatial resolution control is not performed, the conditional number X (k) of the transformation matrix B k is as shown in FIG. 11 in the microphone array of each microphone unit arrangement when kr = 6 . In FIG. 11, the horizontal axis represents the frequency and the vertical axis represents the conditional number X (k).
[0183]
In FIG. 11, curves L21 to L23 show the number of conditions for the annular microphone array, the vortex-shaped microphone array MA11 shown in FIG. 3, and the flower-shaped microphone array MA21 shown in FIG.
[0184]
In this example, it can be seen that the conditional number X (k) of the transformation matrix B k is large in the low frequency band for any microphone array .
[0185]
This phenomenon is caused by the linear dependence of the transformation matrix B k due to the redundant rows, and can be dealt with by spatial resolution control or appropriate matrix regularization in the microphone array to which this technique is applied.
[0186]
On the other hand, in the microphone array of each microphone unit arrangement when the spatial resolution control is performed, the conditional number X (k) of the transformation matrix B n0 k is as shown in FIG. In FIG. 12, the horizontal axis represents the frequency and the vertical axis represents the conditional number X (k).
[0187]
In FIG. 12, curves L31 to L33 show the number of conditions for the annular microphone array, the vortex-shaped microphone array MA11 shown in FIG. 3, and the flower-shaped microphone array MA21 shown in FIG.
[0188]
In this example, it can be seen that the conditional number X (k) of the transformation matrix B n0 k is suppressed to be smaller in the low frequency band of any microphone array than in the example of FIG .
[0189]
Further, for the annular microphone array, the conditional number X (k) is large depending on the frequency. Such deterioration of the conditions of the annular microphone array is an inherent property caused by the value of the Bessel function becoming zero, and cannot be solved by spatial resolution control or matrix regularization.
[0190]
On the other hand, for the vortex-shaped microphone array MA11 and the flower-shaped microphone array MA21, the conditional number X (k) is 30 or less at most frequencies. From this result, it can be seen that a better conditional number X (k) can be obtained and error tolerance is improved by performing spatial resolution control on a microphone array with an appropriate microphone unit arrangement.
[0191]
Next, a recording system that records the wave surface (sound field) of sound using the microphone array described above, and a spherical harmonic coefficient a mn (k) obtained in the recording system. A configuration example of a reproduction system that reproduces (reproduces) the wave surface of sound based on the above will be described.
[0192]
For example, such a recording system and a playback system are configured as shown in FIG.
[0193]
In FIG. 13, the recording system is composed of the microphone array 11 and the recording device 12, and the reproduction system is composed of the reproduction device 13 and the speaker array 14.
[0194]
The microphone array 11 may be a part of the recording device 12, and the speaker array 14 may be a part of the playback device 13.
[0195]
In the recording system, a sound wave surface is recorded by a microphone array 11 composed of a plurality of microphone units, and a multi-channel signal, which is a sound signal obtained as a result, is supplied to the recording device 12. That is, the microphone array 11 records the wave surface of the sound by collecting the sound by each microphone unit, and outputs a signal composed of the audio signal obtained by the sound collection by each microphone unit as a multi-channel signal.
[0196]
The microphone array 11 is for recording a sound field, that is, a wave surface of sound, and is composed of a plurality of sub-arrays. Further, each sub-array is composed of a plurality of microphone units. In particular, the microphone array 11 is a microphone array having the above-mentioned features F1 to F3, such as those shown in FIGS. 3 and 5 to 8, and the microphone unit constituting the microphone array 11 is an omnidirectional microphone. Will be done.
[0197]
The recording device 12 calculates the spherical harmonic coefficient amn (k) based on the multi-channel signal supplied from the microphone array 11 and supplies it to the reproducing device 13.
[0198]
In this example, the recording device 12 has an input unit 21, a time frequency analysis unit 22, a parameter holding unit 23, a spatial resolution control unit 24, and a spherical harmonic coefficient calculation unit 25.
[0199]
The input unit 21 converts an analog multi-channel signal into a digital signal by performing AD (Analog to Digital) conversion on the multi-channel signal supplied from the microphone array 11, and causes the time-frequency analysis unit 22 to perform AD (Analog to Digital) conversion. Supply.
[0200]
The time frequency analysis unit 22 performs a short-time Fourier transform (STFT) on the multi-channel signal supplied from the input unit 21, and calculates the spherical harmonic coefficient of the time frequency spectrum obtained as a result. It is supplied to the unit 25. The time-frequency spectrum obtained by the time-frequency analysis unit 22 corresponds to the sound pressure p k (r l , θ l , φ l ) shown in the equation (4) .
[0201]
The parameter holding unit 23 is based on, for example, the total number of microphone units Q given in advance, the operating frequency band [f min , f max ], the diameter D m of the microphone unit , and the upper limit value X max of the conditional number X (k). Holds the determined placement parameter set P Q opt .
[0202]
For example a microphone array 11 is thus determined by placement parameter set P and Q opt has a shape of the microphone array defined by, the parameter storage unit 23, disposed parameter set P of the microphone array 11 Q opt is held Has been done. In other words, the placement parameter set P Q opt is geometric information indicating the microphone unit placement of the microphone array 11.
[0203]
The parameter holding unit 23 supplies the held arrangement parameter set P Q opt to the spatial resolution control unit 24 and the spherical harmonic coefficient calculation unit 25.
[0204]
The spatial resolution control unit 24 performs spatial resolution control based on the arrangement parameter set P Q opt supplied from the parameter holding unit 23 .
[0205]
In other words, the spatial resolution control unit 24, disposed parameter set P Q opt radius max subarrays constituting the microphone array 11 determined with respect to (r s based on), each frequency, that is the example described above for each wavenumber k Formula ( 9) order n by calculating the like 0 (k × max (r s )) is calculated (determined) the. Then, the spatial resolution control unit 24, that way order n obtained 0 (k × max (r s supplies)) to spherical harmonic coefficient calculation unit 25, the transformation matrix B k instructing restriction of the number of rows To do.
[0206]
The spherical harmonic coefficient calculation unit 25 includes a time frequency spectrum supplied from the time frequency analysis unit 22, an arrangement parameter set P Q opt supplied from the parameter holding unit 23 , and an order n 0 ( order n 0) supplied from the spatial resolution control unit 24. Calculate the spherical harmonics a mn (k) based on k × max (r s )) .
[0207]
For example, the spherical harmonic coefficient calculation unit 25 generates a transformation matrix B n0 k in which the number of rows is limited according to the instruction of the spatial resolution control unit 24 . Specifically, the spherical harmonization coefficient calculation unit 25 is n 0 (k × max (rs s) from the first row of the arrangement parameter set P Q opt , that is, the transformation matrix B k determined for the microphone unit arrangement of the microphone array 11. )) A matrix consisting up to the third row is generated as the final transformation matrix B n0 k .
[0208]
This transformation matrix B n0 k is based on the arrangement parameter set P Q opt which is the geometric information of the microphone array 11 and the order n 0 (k × max (r s )) which is the output of the spatial resolution control unit 24. , Generated for each wavenumber k, that is, for each STFT bin.
[0209]
The spherical harmonics calculation unit 25 performs the same calculation as the above equation (3) based on the generalized inverse matrix obtained for the transformation matrix B n0 k and the time frequency spectrum, and the spherical harmonics a mn ( Calculate k). For example, the spherical harmonics calculation unit 25 uses the Moore-Penrose pseudo-inverse matrix as the generalized inverse of the transformation matrix B n0 k . In other words, the transformation matrix B n0 k Moore about - pseudo-inverse of Penrose, the transformation matrix B n0 k is calculated as a generalized inverse matrix of.
[0210]
The spherical harmonic coefficient calculation unit 25 performs the same calculation as in the above equation (3), but in this calculation, spherical harmonic transformation (SHT (Spherical Harmonic Transform)) and mode compensation are performed at the same time. The mode compensation referred to here is a process corresponding to the division of p k (r, θ q , φ q ) Y * m n (θ q , φ q ) by b n (kr) in the equation (1) , that is, This is a process of dividing the time-frequency spectrum converted by spherical harmony with a mode function (Bessel function).
[0211]
Here, an example in which spherical harmonic conversion and mode compensation are performed at the same time when obtaining the spherical harmonic coefficient a mn (k) will be described, but these spherical harmonic conversion and mode compensation are performed separately. You may.
[0212]
In such a case, the spherical harmonic coefficient calculation unit 25 is provided with a processing block for performing spherical harmonic conversion and a processing block for performing mode compensation. Then, in the processing block that performs spherical harmonic conversion, spherical harmonic conversion is performed on the time frequency spectrum, and in the processing block that performs mode compensation, division by the mode function (Bessel function) is performed on the time frequency spectrum that has been spherically harmonized. Will be. At this time, in spherical harmonic conversion and mode compensation, operations up to the term determined by the order n 0 (k × max (r s )) are performed.
[0213]
Further, the spherical harmonics calculation unit 25 outputs (transmits) the calculated spherical harmonics a mn (k) to the reproduction system.
[0214]
In the reproduction system, a drive signal for driving the speaker array 14 is generated based on the spherical harmonics a mn (k) output from the spherical harmonics calculation unit 25 , and the wave surface of the sound is reproduced. The generation of the drive signal can be realized by correcting the speaker characteristics of the speaker array 14 or other algorithms.
[0215]
For example, the reproduction device 13 of the reproduction system includes a speaker arrangement information holding unit 31, a drive signal generation unit 32, a time frequency synthesis unit 33, and an output unit 34.
[0216]
The speaker arrangement information holding unit 31 holds the speaker arrangement information indicating the arrangement of the speakers constituting the speaker array 14, and supplies the held speaker arrangement information to the drive signal generation unit 32.
[0217]
The drive signal generation unit 32 receives the spherical harmonics a mn (k) transmitted from the spherical harmonics calculation unit 25 , and supplies the received spherical harmonics a mn (k) and the speaker arrangement information holding unit 31. A drive signal is generated based on the generated speaker arrangement information and supplied to the time-frequency synthesis unit 33.
[0218]
For example, in the drive signal generation unit 32, the calculation of the above equation (2) is performed, and a signal indicating the sound pressure p k (r q , θ q , φ q ) is calculated as a drive signal in the time frequency region. In the calculation of Eq. (2), the value of the radius of the reproduction area, which is the area for reproducing the wave surface of the sound, is used as the radius r q .
[0219]
In the calculation of equation (2), the multiplication of the Vessel function on the spherical harmonic coefficient a mn (k), that is, the generation of the driving signal in the spherical harmonic region and the inverse Spherical Harmonic Transform (ISHT) for the driving signal )) And) are performed at the same time, but the spherical harmonic inverse conversion may be performed after the drive signal of the spherical harmonic region is generated. In such a case, the drive signal generation unit 32 is provided with a processing block for generating a drive signal in the spherical harmonic region and a processing block for performing inverse spherical harmonic conversion.
[0220]
The time-frequency synthesizing unit 33 performs an inverse short-time Fourier transform (ISTFT) on the drive signal supplied from the drive signal generation unit 32, and obtains the drive signal in the time domain obtained as a result. It is supplied to the output unit 34.
[0221]
The output unit 34 performs DA (Digital to Analog) conversion on the drive signal supplied from the time frequency synthesis unit 33, and supplies the analog drive signal obtained as a result to the speaker array 14. The speaker array 14 reproduces (reproduces) the wave surface of the sound recorded by the recording system by outputting the sound based on the drive signal supplied from the output unit 34.
[0222]
For example, the speaker array 14 is obtained by arranging linear speaker arrays obtained by arranging the speakers in a linear shape in a rectangular shape, and the region inside the speaker array 14 is used as a wave surface reproduction area. The shape of the speaker array 14, that is, the speaker arrangement may be any.
[0223]
Next, the operations of the recording system and the reproduction system shown in FIG. 13 will be described.
[0224]
First, the recording process by the recording system will be described with reference to the flowchart of FIG. Before the start of the recording process, the parameter set P Q opt is determined in advance by the parameter holding unit 23 or another processing block , and the resulting placement parameter set P Q opt is held by the parameter holding unit 23. To.
[0225]
In step S11, the spatial resolution control unit 24 performs spatial resolution control based on the arrangement parameter set P Q opt supplied from the parameter holding unit 23 .
[0226]
For example the spatial resolution control unit 24, order n by calculating the like above expression (9) 0 (k × max (r s is supplied to the spherical harmonic coefficient calculation unit 25 calculates the)), the transformation matrix B k Instruct to limit the number of lines in.
[0227]
In step S12, the microphone array 11 picks up ambient sound by the microphone unit and supplies the resulting multi-channel signal to the input unit 21. The input unit 21 AD-converts the multi-channel signal supplied from the microphone array 11 and supplies it to the time-frequency analysis unit 22.
[0228]
In step S13, the time frequency analysis unit 22 performs a short-time Fourier transform on the multi-channel signal supplied from the input unit 21, and supplies the time frequency spectrum obtained as a result to the spherical harmonic coefficient calculation unit 25.
[0229]
In step S14, the spherical harmonic coefficient calculation unit 25 has a time frequency spectrum from the time frequency analysis unit 22, an arrangement parameter set P Q opt from the parameter holding unit 23 , and an order n 0 (k ×) from the spatial resolution control unit 24. Calculate the spherical harmonics a mn (k) based on max (r s )) .
[0230]
That is, spherical harmonic coefficient calculation unit 25 according to an instruction of the spatial resolution control unit 24, the degree n 0 (k × max (r s transformation matrix B based on)) n0 k to generate a, the transformation matrix B n0 Calculate the generalized inverse matrix of k . Then, the spherical harmonics calculation unit 25 performs the same calculation as in Eq. (3) based on the obtained generalized inverse matrix and the time frequency spectrum, and calculates the spherical harmonics a mn (k).
[0231]
The spherical harmonics calculation unit 25 outputs the spherical harmonics a mn (k) calculated in this way, and the recording process ends.
[0232]
As described above, the recording system uses the transformation matrix obtained by recording the wave surface using the microphone array 11 having a shape (microphone unit arrangement) determined for the arrangement parameter set P Q opt and performing spatial resolution control. To calculate the spherical harmonics a mn (k). By doing so, it is possible to perform wideband sound field recording at low cost.
[0233]
Subsequently, the reproduction process performed by the reproduction system will be described with reference to the flowchart of FIG. This reproduction process is started when the drive signal generation unit 32 of the reproduction device 13 receives the spherical harmonics a mn (k) transmitted by the recording system .
[0234]
In step S41, the drive signal generation unit 32 generates a drive signal based on the received spherical harmonic coefficient a mn (k) and the speaker arrangement information supplied from the speaker arrangement information holding unit 31, and the time frequency synthesis unit 32. Supply to 33. For example, in step S41, the above-mentioned equation (2) is calculated , and a signal indicating the sound pressure p k (r q , θ q , φ q ) is calculated as a drive signal in the time frequency region.
[0235]
In step S42, the time-frequency synthesis unit 33 performs an inverse short-time Fourier transform on the drive signal supplied from the drive signal generation unit 32, and supplies the drive signal in the time domain obtained as a result to the output unit 34. .. Further, the output unit 34 performs DA conversion on the drive signal supplied from the time-frequency synthesis unit 33, and supplies the analog drive signal obtained as a result to the speaker array 14.
[0236]
In step S43, the speaker array 14 reproduces (reproduces) the wave surface of the sound recorded by the recording system by outputting the sound based on the drive signal supplied from the output unit 34, and the reproduction process ends.
[0237]
As described above, the reproduction system generates a drive signal from the received spherical harmonic coefficient a mn (k), and reproduces the wave surface of the sound based on the drive signal. In the reproduction system, a wide band wave surface reproduction can be realized by reproducing the wave surface based on the spherical harmonics a mn (k) received from the recording system .
[0238]
By the way, the series of processes described above can be executed by hardware or software. When a series of processes are executed by software, the programs that make up the software are installed on the computer. Here, the computer includes a computer embedded in dedicated hardware and, for example, a general-purpose personal computer capable of executing various functions by installing various programs.
[0239]
FIG. 16 is a block diagram showing an example of hardware configuration of a computer that executes the above-mentioned series of processes programmatically.
[0240]
In a computer, a CPU (Central Processing Unit) 501, a ROM (Read Only Memory) 502, and a RAM (Random Access Memory) 503 are connected to each other by a bus 504.
[0241]
An input / output interface 505 is further connected to the bus 504. An input unit 506, an output unit 507, a recording unit 508, a communication unit 509, and a drive 510 are connected to the input / output interface 505.
[0242]
The input unit 506 includes a keyboard, a mouse, a microphone array, an image sensor, and the like. The output unit 507 includes a display, a speaker array, and the like. The recording unit 508 includes a hard disk, a non-volatile memory, and the like. The communication unit 509 includes a network interface and the like. The drive 510 drives a removable recording medium 511 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.
[0243]
In the computer configured as described above, the CPU 501 loads the program recorded in the recording unit 508 into the RAM 503 via the input / output interface 505 and the bus 504 and executes the above-described series. Is processed.
[0244]
The program executed by the computer (CPU 501) can be recorded and provided on a removable recording medium 511 as a package medium or the like, for example. Programs can also be provided via wired or wireless transmission media such as local area networks, the Internet, and digital satellite broadcasting.
[0245]
In a computer, the program can be installed in the recording unit 508 via the input / output interface 505 by mounting the removable recording medium 511 in the drive 510. Further, the program can be received by the communication unit 509 and installed in the recording unit 508 via a wired or wireless transmission medium. In addition, the program can be pre-installed in the ROM 502 or the recording unit 508.
[0246]
The program executed by the computer may be a program in which processing is performed in chronological order in the order described in this specification, or in parallel or at a necessary timing such as when a call is made. It may be a program in which processing is performed.
[0247]
Further, the embodiment of the present technology is not limited to the above-described embodiment, and various changes can be made without departing from the gist of the present technology.
[0248]
For example, the present technology can have a cloud computing configuration in which one function is shared by a plurality of devices via a network and jointly processed.
[0249]
Further, each step described in the above-mentioned flowchart can be executed by one device or can be shared and executed by a plurality of devices.
[0250]
Further, when a plurality of processes are included in one step, the plurality of processes included in the one step can be executed by one device or shared by a plurality of devices.
[0251]
Further, the present technology can also have the following configurations.
[0252]
(1)
A microphone array for a plurality of sub-arrays sound composed of field recording,
the sub-array is in the form of a discrete rotational symmetry with a predetermined radius while being composed of a plurality of microphones,
the A
microphone array in which when the radii of a plurality of the sub-arrays are arranged to form a sequence, the sequence becomes a generalized arithmetic progression .
(2) The
microphone array according to
(1) , wherein each of the plurality of microphones constituting the sub-array is arranged at a position at a distance of the radius of the sub-array from the center position of the microphone array.
(3) The
sub-array, according to
(1) or (2), which matches the other sub-array when at least one of enlargement, reduction, rotation, and inversion is performed on the sub-array . Microphone array.
(4)
When all the microphones constituting the microphone array are projected radially on the ring centered on the center position of the microphone array, the projected microphones are arranged at equal intervals on the ring.
The microphone array according to any one of (1) to (3) , wherein the microphone is arranged therein .
(5)
The plurality of microphones constituting the microphone array are all omnidirectional microphones, or at least one of the plurality of microphones constituting the microphone array is not an omnidirectional microphone
(1) to (4). The microphone array according to any one of ().
(6)
The sub-array is provided with a spherical harmony coefficient calculation unit that calculates a spherical harmony coefficient based on a multi-channel signal obtained by collecting sound from a microphone array for sound field recording composed of a plurality of sub-
arrays. , A discrete rotationally symmetric shape composed of a plurality of microphones and having a predetermined radius, and when the radii of the
plurality of sub-arrays are arranged to form a sequence, the sequence becomes a generalized arithmetic progression. made
recording apparatus.
(7) The recording device according
to
(6), wherein the spherical harmonics calculation unit calculates the spherical harmonics by performing mode compensation .
(8) The recording apparatus according to (7),
further including a spatial resolution control unit that limits the number of rows of the transformation matrix for mode compensation based on a predetermined order of the spherical harmonic region
.
(9) The recording device according
to
(8), wherein the spatial resolution control unit determines the predetermined order based on the maximum value of the radius of the plurality of sub-arrays .
(10) The
spherical harmonics calculation unit calculates the spherical harmonics by mode compensation based on the generalized inverse matrix of the transformation matrix with a limited number of rows and the multichannel signal
(8) or. The recording device according to (9).
(11) The
recording device
calculates a spherical harmony coefficient based on a multi-channel signal obtained by collecting sound from a microphone array for recording a sound field composed of a plurality of sub-arrays , and the sub-
array has a plurality of sub-arrays. A recording method that is composed of microphones and has a discrete rotationally symmetric shape having a predetermined radius, and when the radii of the
plurality of subarrays are arranged into a sequence, the sequence becomes a generalized arithmetic progression.
..
(12)
a plurality of microphone arrays for subarray from configured sound field From calculates the spherical harmonic coefficients based on the multi-channel signal obtained by sound pickup
by executing a process including a step to the computer,
the The sub-array is a discrete rotationally symmetric shape composed of a plurality of microphones and having a predetermined radius, and when the radii of the
plurality of sub-arrays are arranged into a sequence, the sequence is a generalized arithmetic progression. A
program that is a sequence .
Description of the sign
[0253]
11 Microphone array, 12 Recording device, 22-hour frequency analysis unit, 23 Parameter holding unit, 24 Spatial resolution control unit, 25 Spherical harmonic coefficient calculation unit
The scope of the claims
[Claim 1]
A microphone array for a plurality of sub-arrays sound composed of field recording,
the sub-array is in the form of a discrete rotational symmetry with a predetermined radius while being composed of a plurality of microphones,
the plurality of the A
microphone array in which the radii of the sub-array are arranged into a sequence, and the sequence becomes a generalized arithmetic progression .
[Claim 2]
The microphone array according to
claim 1 , wherein each of the plurality of microphones constituting the sub-array is arranged at a position at a distance of the radius of the sub-array from the center position of the microphone array.
[Claim 3]
The microphone array according to claim 1, wherein the sub-array matches the other sub-array when at least one of expansion, contraction, rotation, and inversion is performed on the sub- array.
[Claim 4]
When all the microphones constituting the microphone array are projected radially on a ring centered on the center position of the microphone array, the projected microphones are arranged at equal intervals on the ring.
The microphone array according to claim 1, wherein the microphone is arranged .
[Claim 5]
The first aspect of
claim 1, wherein the plurality of microphones constituting the microphone array are all omnidirectional microphones, or at least one of the plurality of the microphones constituting the microphone array is not an omnidirectional microphone . Microphone array.
[Claim 6]
The microphone array for recording a sound field composed of a plurality of sub-arrays is provided with a spherical harmonic coefficient calculation unit that calculates a spherical harmonic coefficient based on a multi-channel signal obtained by collecting sound, and the
sub-array includes a plurality of sub-arrays. A recording device composed of microphones and having a discrete rotationally symmetric shape having a predetermined radius, and when the radii of the
plurality of subarrays are arranged into a sequence, the sequence becomes a generalized arithmetic progression.
..
[Claim 7]
The recording device according to claim 6, wherein the spherical harmonics calculation unit calculates the spherical harmonics by performing mode compensation .
[Claim 8]
The recording apparatus according to claim 7, further comprising a spatial resolution control unit that limits the number of rows of the transformation matrix for mode compensation based on a predetermined order of the spherical harmonic region .
[Claim 9]
The recording device according to claim 8, wherein the spatial resolution control unit determines the predetermined order based on the maximum value of the radii of the plurality of sub-arrays .
[Claim 10]
The recording according to claim 8, wherein the spherical harmonics calculation unit calculates the spherical harmonics by mode compensation based on the generalized inverse matrix of the transformation matrix having a limited number of rows and the multichannel signal. apparatus.
[Claim 11]
The recording device
calculates a spherical harmony coefficient based on a multi-channel signal obtained by collecting sound from a microphone array for sound field recording composed of a plurality of sub-arrays , and the sub-
array is composed of a plurality of microphones. A recording method having a discrete rotationally symmetric shape having a predetermined radius, and when the radii of the
plurality of subarrays are arranged to form a sequence, the sequence becomes a generalized arithmetic progression
.
[Claim 12]
A
computer is made to execute a process including a step of calculating a spherical harmony coefficient based on a multi-channel signal obtained by collecting sound from a microphone array for recording a sound field composed of a plurality of sub-
arrays. It is a discrete rotationally symmetric shape composed of a plurality of microphones and having a predetermined radius, and when the radii of the
plurality of subarrays are arranged to form a sequence, the sequence becomes a generalized arithmetic progression.
program.
| # | Name | Date |
|---|---|---|
| 1 | 202017036722-TRANSLATIOIN OF PRIOIRTY DOCUMENTS ETC. [26-08-2020(online)].pdf | 2020-08-26 |
| 2 | 202017036722-STATEMENT OF UNDERTAKING (FORM 3) [26-08-2020(online)].pdf | 2020-08-26 |
| 3 | 202017036722-PRIORITY DOCUMENTS [26-08-2020(online)].pdf | 2020-08-26 |
| 4 | 202017036722-POWER OF AUTHORITY [26-08-2020(online)].pdf | 2020-08-26 |
| 5 | 202017036722-FORM 1 [26-08-2020(online)].pdf | 2020-08-26 |
| 6 | 202017036722-DRAWINGS [26-08-2020(online)].pdf | 2020-08-26 |
| 7 | 202017036722-DECLARATION OF INVENTORSHIP (FORM 5) [26-08-2020(online)].pdf | 2020-08-26 |
| 8 | 202017036722-COMPLETE SPECIFICATION [26-08-2020(online)].pdf | 2020-08-26 |
| 9 | 202017036722-Proof of Right [26-11-2020(online)].pdf | 2020-11-26 |
| 10 | 202017036722.pdf | 2021-10-19 |
| 11 | 202017036722-FORM 18 [09-02-2022(online)].pdf | 2022-02-09 |
| 12 | 202017036722-FER.pdf | 2022-07-12 |
| 13 | 202017036722-AbandonedLetter.pdf | 2024-02-16 |
| 1 | 202017036722E_12-07-2022.pdf |