Abstract: The present technology relates to a signal processing device and method, and a program which enable the reproducibility of a sound image to be improved with a small calculation amount. The signal processing device is provided with: a rendering scheme selection unit which selects, from among a plurality of different schemes, one or more schemes of a rendering process for localizing a sound image of an audio signal in a listening space; and a rendering processing unit which performs a rendering process on the audio signal by means of the scheme selected by the rendering scheme selection unit. The present technology can be applied to the signal processing device.
Specification
Title of invention: Signal processing device and method, and program
Technical field
[0001]
The present technology relates to a signal processing device and method, and a program, and particularly relates to a signal processing device and method, and a program capable of improving reproducibility of a sound image with a small amount of calculation.
Background technology
[0002]
Conventionally, object audio technology has been used in movies and games, and a coding method capable of handling object audio has been developed. Specifically, for example, the MPEG (Moving Picture Experts Group)-H Part 3:3D audio standard, which is an international standard, is known (see Non-Patent Document 1, for example).
[0003]
In such an encoding method, in addition to the conventional 2-channel stereo method and 5.1-channel multi-channel stereo method, a moving sound source is treated as an independent audio object, and the position information of the object is stored together with the signal data of the audio object. It can be encoded as metadata.
[0004]
By doing so, reproduction can be performed in various viewing environments in which the number and arrangement of speakers are different. Further, it is possible to easily process the sound of the specific sound source at the time of reproduction, such as adjusting the volume of the sound of the specific sound source and adding an effect to the sound of the specific sound source, which were difficult with the conventional encoding method.
[0005]
For example, in the standard of Non-Patent Document 1, a method called three-dimensional VBAP (Vector Based Amplitude Panning) (hereinafter, simply referred to as VBAP) is used for rendering processing.
[0006]
This is one of the rendering methods generally called panning, and among the speakers existing on the surface of the sphere whose origin is the listening position, the gains are added to the three speakers closest to the audio object that are also on the surface of the sphere. Is a method of rendering by distributing.
[0007]
In addition to VBAP, rendering processing by a panning method called Speaker-anchored coordinates panner that distributes gains to the x-axis, the y-axis, and the z-axis is also known (for example, Non-Patent Document 2). reference).
[0008]
On the other hand, as a method of rendering an audio object other than the panning process, a method of using a filter of a head related transfer function has been proposed (for example, refer to Patent Document 1).
[0009]
Generally, when rendering a moving audio object using a head related transfer function, a head related transfer function filter is often obtained as follows.
[0010]
That is, for example, it is general to spatially sample the moving space range and prepare in advance a number of head-related transfer function filters corresponding to individual points in the space. Further, for example, the head related transfer function at each position in the space measured at a constant distance interval may be used to obtain the filter related to the head related transfer function at the desired position by distance correction by the three-dimensional synthesis method.
[0011]
In Patent Document 1 described above, a method of generating a head-related transfer function filter at an arbitrary distance by using parameters necessary for generating a head-related transfer function filter obtained by sampling a spherical surface at a constant distance. Is listed.
Prior art documents
Non-patent literature
[0012]
Non-patent document 1: INTERNATIONAL STANDARD ISO/IEC 23008-3 First edition 2015-10-15 Information technology High efficiency coding and media delivery in heterogeneous environments Part 3: 3D audio
Non-patent document 2: ETSI TS 103 448 v1.1.1(2016 -09)
Patent literature
[0013]
Patent Document 1: Japanese Patent No. 5754414
Summary of the invention
Problems to be Solved by the Invention
[0014]
However, with the technique described above, when the sound image of the sound of the audio object is localized by rendering, it is difficult to obtain high sound image localization reproducibility with a small amount of calculation. That is, it is difficult to realize the sound image localization that causes the sound image to be perceived as if it were originally intended with a small amount of calculation.
[0015]
For example, in the rendering of the audio object by the panning process, it is assumed that there is one listening position. In this case, for example, when the audio object is close to the listening position, the difference between the arrival times of the sound waves reaching the left ear of the listener and the sound waves reaching the right ear of the listener cannot be ignored.
[0016]
However, when VBAP is performed as the panning process, even if the audio object is located inside or outside the sphere surface where the speaker is arranged, the rendering is performed as if the audio object is on the sphere surface. Then, when the audio object approaches the listening position, the sound image of the audio object during reproduction becomes far from what is expected.
[0017]
On the other hand, in the rendering using the head-related transfer function, high reproducibility of sound image localization can be realized even when the audio object is located close to the listener. Further, as FIR (Finite Impulse Response) filter processing of the head related transfer function, there is high-speed arithmetic processing such as FFT (Fast Fourier Transform) and QMF (Quadrature Mirror Filter).
[0018]
However, the amount of FIR filter processing of these head related transfer functions is much larger than the amount of panning processing. Therefore, when there are many audio objects, it may not be appropriate to perform rendering using the head related transfer function for all audio objects.
[0019]
The present technology has been made in view of such circumstances, and is to improve the reproducibility of a sound image with a small amount of calculation.
Means for solving the problems
[0020]
A signal processing device according to one aspect of the present technology includes a rendering method selection unit that selects one or more rendering processing methods for localizing a sound image of an audio signal in a listening space from a plurality of different methods, and the rendering method selection. A rendering processing unit that performs the rendering process of the audio signal according to a method selected by the unit.
[0021]
A signal processing method or program according to an aspect of the present technology selects one or more rendering techniques for localizing a sound image of an audio signal in a listening space from among a plurality of different techniques, and selects the audio according to the selected technique. Performing the rendering process of the signal.
[0022]
In one aspect of the present technology, one or more rendering processing methods for locating a sound image of an audio signal in a listening space are selected from a plurality of mutually different methods, and the rendering processing of the audio signal is performed by the selected method. Is done.
Effect of the invention
[0023]
According to one aspect of the present technology, it is possible to improve the reproducibility of a sound image with a small amount of calculation.
[0024]
Note that the effects described here are not necessarily limited, and may be any effects described in the present disclosure.
Brief description of the drawings
[0025]
FIG. 1 is a diagram illustrating VBAP.
FIG. 2 is a diagram showing a configuration example of a signal processing device.
FIG. 3 is a diagram showing a configuration example of a rendering processing unit.
FIG. 4 is a diagram showing an example of metadata.
FIG. 5 is a diagram illustrating audio object position information.
FIG. 6 is a diagram illustrating selection of a rendering method.
FIG. 7 is a diagram illustrating head related transfer function processing.
FIG. 8 is a diagram illustrating selection of a rendering method.
FIG. 9 is a flowchart illustrating an audio output process.
FIG. 10 is a diagram showing an example of metadata.
FIG. 11 is a diagram showing an example of metadata.
FIG. 12 is a diagram showing a configuration example of a computer.
MODE FOR CARRYING OUT THE INVENTION
[0026]
Hereinafter, embodiments to which the present technology is applied will be described with reference to the drawings.
[0027]
In the
present technology, when rendering an audio object, a plurality of different rendering methods are used for each audio object, depending on the position of the audio object in the listening space. By selecting one or more of these methods, the reproducibility of a sound image can be improved with a small amount of calculation. That is, the present technology is capable of realizing sound image localization that makes a person perceive as if a sound image were originally intended even with a small amount of calculation.
[0028]
Particularly, in the present technology, as a rendering processing method for localizing a sound image of an audio signal in a listening space, that is, as a rendering method, one or more of a plurality of rendering methods having different calculation amounts (calculation loads) and sound image localization performance are selected. A rendering method is selected.
[0029]
Note that, here, an example will be described in which the audio signal for which the rendering method is selected is an audio signal of an audio object (audio object signal). However, the audio signal to be selected as the rendering method is not limited to this, and may be any audio signal as long as it is an audio signal that localizes a sound image in the listening space.
[0030]
As described above, in VBAP, the gain is distributed to the three speakers closest to the audio object that also exist on the spherical surface among the speakers existing on the spherical surface whose origin is the listening position in the listening space.
[0031]
For example, as shown in FIG. 1, it is assumed that a listener U11 is in a listening space which is a three-dimensional space, and three speakers SP1 to SP3 are arranged in front of the listener U11.
[0032]
It is also assumed that the position of the listener U11's head is the origin O and that the speakers SP1 to SP3 are located on the surface of a sphere centered on the origin O.
[0033]
Now, assume that an audio object exists in a region TR11 surrounded by the speakers SP1 to SP3 on the surface of the sphere, and the sound image is localized at the position VSP1 of the audio object.
[0034]
In such a case, in VBAP, the gain of the audio object is distributed to the speakers SP1 to SP3 around the position VSP1.
[0035]
Specifically, in a three-dimensional coordinate system with the origin O as a reference (origin), the position VSP1 is represented by a three-dimensional vector P having the origin O as a starting point and the position VSP1 as an ending point.
[0036]
Further, the origin O as a starting point, a three-dimensional vector whose position the end point of the speakers SP1 to speaker SP3 vector L 1 to the vector L 3 When, the vector P, as shown in the following equation (1), the vector L It can be represented by a linear sum of 1 to vector L 3 .
[0037]
[Number 1]
[0038]
Here, vector L in the formula (1) 1 to the vector L 3 coefficients g are multiplied by 1 to coefficient g 3 is calculated and the coefficients g 1 through coefficient g 3 a, from each of the speakers SP1 to speaker SP3 With the gain of the sound to be output, the sound image can be localized at the position VSP1.
[0039]
For example , a vector having the coefficients g 1 to g 3 as elements is g 123 =[g 1 ,g 2 ,g 3 ], and a vector having the vectors L 1 to L 3 as elements is L 123 =[L 1 ,L 2 , L 3 ], the following equation (2) can be obtained by modifying the above equation (1).
[0040]
[Number 2]
[0041]
By using the coefficients g 1 to g 3 obtained by calculating Equation (2) as gains and outputting the audio object signals, which are the sound signals of the audio objects, to the speakers SP1 to SP3, A sound image can be localized at the position VSP1.
[0042]
The positions of the speakers SP1 to SP3 are fixed and the information indicating the positions of the speakers is known . Therefore , the inverse matrix L 123 −1 can be obtained in advance. Therefore, VBAP can perform rendering with relatively easy calculation, that is, with a small amount of calculation.
[0043]
Therefore, when the audio object is at a position sufficiently distant from the listener U11, the sound image can be properly localized with a small amount of calculation by performing rendering by panning processing such as VBAP.
[0044]
However, when the audio object is in a position close to the listener U11, it is difficult to express the difference in arrival time of the sound waves reaching the left and right ears of the listener U11 by panning processing such as VBAP, and it is sufficient. It was not possible to obtain high sound image reproducibility.
[0045]
Therefore, in the present technology, one or more rendering methods are selected from panning processing and rendering processing using a head-related transfer function filter (hereinafter, also referred to as head-related transfer function processing) according to the position of the audio object, The rendering process was performed.
[0046]
For example, the rendering method is selected based on the relative positional relationship between the listening position, which is the position of the listener in the listening space, and the position of the audio object.
[0047]
Specifically, as an example, when the audio object is located on the surface of the sphere where the speaker is arranged or outside the surface of the sphere, a panning process such as VBAP is selected as the rendering method.
[0048]
On the other hand, when the audio object is located inside the surface of the sphere on which the speaker is arranged, head related transfer function processing is selected as the rendering method.
[0049]
By doing so, a sufficiently high sound image reproducibility can be obtained with a small amount of calculation. That is, the reproducibility of the sound image can be improved with a small amount of calculation.
[0050]
Now, the present technology will be described in more detail.
[0051]
FIG. 2 is a diagram showing a configuration example of an embodiment of a signal processing device to which the present technology is applied.
[0052]
The signal processing device 11 shown in FIG. 2 has a core decoding processing unit 21 and a rendering processing unit 22.
[0053]
The core decoding processing unit 21 receives and decodes (decodes) the transmitted input bit stream, and supplies the audio processing object position information and the audio object signal obtained as a result to the rendering processing unit 22. In other words, the core decoding processing unit 21 acquires audio object position information and audio object signal.
[0054]
Here, the audio object signal is an audio signal for reproducing the sound of the audio object.
[0055]
The audio object position information is the metadata of the audio object, that is, the audio object signal, which is necessary for the rendering performed by the rendering processing unit 22.
[0056]
Specifically, the audio object position information is information indicating the position of the audio object in the three-dimensional space, that is, the listening space.
[0057]
The rendering processing unit 22 generates an output audio signal based on the audio object position information and the audio object signal supplied from the core decoding processing unit 21, and supplies the output audio signal to a speaker, a recording unit, or the like in the subsequent stage.
[0058]
Specifically, the rendering processing unit 22 selects one of a panning process, a head related transfer function process, or a panning process and a head related transfer function process as the rendering method based on the audio object position information. .
[0059]
Then, the rendering processing unit 22 performs the selected rendering process to perform rendering on a reproducing device such as a speaker or a headphone that is an output destination of the output audio signal to generate an output audio signal.
[0060]
The rendering processing unit 22 may of course select one or more rendering methods from among three or more different rendering methods including panning processing and head-related transfer function processing.
[0061]
Next, a more detailed configuration example of the rendering processing unit 22 of the signal processing device 11 shown in FIG. 2 will be described.
[0062]
The rendering processing unit 22 is configured, for example, as shown in FIG.
[0063]
In the example shown in FIG. 3, the rendering processing unit 22 includes a rendering method selection unit 51, a panning processing unit 52, a head related transfer function processing unit 53, and a mixing processing unit 54.
[0064]
The rendering method selection unit 51 is supplied with the audio object position information and the audio object signal from the core decoding processing unit 21.
[0065]
The rendering method selection unit 51 selects a rendering processing method for an audio object, that is, a rendering method, for each audio object based on the audio object position information supplied from the core decoding processing unit 21.
[0066]
In addition, the rendering method selection unit 51 uses at least the panning processing unit 52 and the head related transfer function processing unit 53 for the audio object position information and the audio object signal supplied from the core decoding processing unit 21 according to the selection result of the rendering method. Supply to either one.
[0067]
The panning processing unit 52 performs a panning process based on the audio object position information and the audio object signal supplied from the rendering method selection unit 51, and supplies a panning process output signal obtained as a result to the mixing processing unit 54.
[0068]
Here, the panning processing output signal is an audio signal of each channel for reproducing the sound of the audio object so that the sound image of the sound of the audio object is localized at the position in the listening space indicated by the audio object position information. is there.
[0069]
For example, here, the channel configuration of the output audio signal output destination is predetermined, and the audio signal of each channel of the channel configuration is generated as the panning processing output signal.
[0070]
As an example, when the output destination of the output audio signal is a speaker system including the speakers SP1 to SP3 shown in FIG. 1, the audio signals of the channels corresponding to the speakers SP1 to SP3 are the panning processing output signals. Is generated.
[0071]
Specifically, for example, when VBAP is performed as the panning process, an audio signal obtained by multiplying the audio object signal supplied from the rendering method selection unit 51 by a coefficient g 1 that is a gain is The output signal is the panning output signal of the channel corresponding to the speaker SP1. Similarly, the audio signal obtained by multiplying the audio object signal by each of the coefficient g 2 and the coefficient g 3 becomes the panning processing output signal of the channel corresponding to each of the speaker SP2 and the speaker SP3.
[0072]
In the panning processing unit 52, what kind of processing is performed as the panning processing, for example, VBAP adopted in the MPEG-H Part 3:3D audio standard, or processing by a panning method called Speaker-anchored coordinate panner. You may do it. In other words, the rendering method selection unit 51 may select VBAP as the rendering method, or may select the Speaker-anchored coordinates panner.
[0073]
The head-related transfer function processing unit 53 performs head-related transfer function processing based on the audio object position information and the audio object signal supplied from the rendering method selection unit 51, and outputs the resulting head-related transfer function processing output signal. It is supplied to the mixing processing unit 54.
[0074]
Here, the head-related transfer function processing output signal is such that the sound image of the sound of the audio object is localized in the position in the listening space indicated by the audio object position information so as to reproduce the sound of the audio object. It is an audio signal.
[0075]
That is, the head-related transfer function processing output signal corresponds to the panning processing output signal, and the head-related transfer function processing output signal and the panning processing output signal are the head-related transfer function processing when the audio signal is generated. It is different whether it is processing or panning processing.
[0076]
The panning processing unit 52 and the head related transfer function processing unit 53 described above function as a rendering processing unit that performs a rendering process such as a panning process or a head related transfer function process by the rendering method selected by the rendering method selection unit 51.
[0077]
The mixing processing unit 54 outputs an output audio signal based on at least one of the panning processing output signal supplied from the panning processing unit 52 and the head related transfer function processing output signal supplied from the head related transfer function processing unit 53. Generate and output to the latter stage.
[0078]
For example, it is assumed that audio object position information and audio object signal of one audio object are stored in the input bitstream.
[0079]
In such a case, when the panning processing output signal and the head related transfer function processing output signal are supplied, the mixing processing unit 54 performs correction processing to generate an output audio signal. In the correction process, the panning process output signal and the head related transfer function process output signal are combined (blended) for each channel to be an output audio signal.
[0080]
On the other hand, when only one of the panning processing output signal and the head-related transfer function processing output signal is supplied, the mixing processing unit 54 uses the supplied signal as it is as the output audio signal. ..
[0081]
Further, for example, it is assumed that the input bitstream stores audio object position information of a plurality of audio objects and audio object signals.
[0082]
In such a case, the mixing processing unit 54 performs correction processing as necessary to generate an output audio signal for each audio object.
[0083]
Then, the mixing processing unit 54 performs a mixing process of adding (combining) the output audio signals of the audio objects thus obtained for each channel, and finally outputs the output audio signals of the respective channels. Output audio signal. That is, the output audio signals of the same channel, obtained for each audio object, are added to form the final output audio signal of that channel.
[0084]
As described above, the mixing processing unit 54 is an output audio signal generation unit that generates an output audio signal by performing a correction process and a mixing process that combine the panning process output signal and the head-related transfer function process output signal as necessary. Function.
[0085]
By the way, the above-mentioned audio object position information is encoded using the format shown in FIG. 4 at predetermined time intervals (every predetermined number of frames) and stored in the input bitstream. ..
[0086]
In the metadata shown in FIG. 4, “num_objects” indicates the number of audio objects included in the input bitstream.
[0087]
Further, "tcimsbf" is an abbreviation for "Two's complement integer, most significant(sign) bit first", and the sign bit indicates the leading two's complement. "Uimsbf" is an abbreviation for "Unsigned integer, most significant bit first", and the most significant bit indicates the leading unsigned integer.
[0088]
Furthermore, “position_azimuth[i]”, “position_elevation[i]”, and “position_radius[i]” respectively indicate audio object position information of the i-th audio object included in the input bitstream.
[0089]
Specifically, "position_azimuth[i]" indicates the azimuth angle of the position of the audio object in the spherical coordinate system, and "position_elevation[i]" indicates the elevation angle of the position of the audio object in the spherical coordinate system. Further, "position_radius[i]" indicates the distance to the position of the audio object in the spherical coordinate system, that is, the radius.
[0090]
Here, the relationship between the spherical coordinate system and the three-dimensional orthogonal coordinate system is the relationship shown in FIG.
[0091]
In FIG. 5, the X axis, the Y axis, and the Z axis that pass through the origin O and are perpendicular to each other are axes of the three-dimensional orthogonal coordinate system. For example, in the three-dimensional orthogonal coordinate system, the position of the audio object OB11 in the space is X1, which is the X coordinate indicating the position in the X axis direction, Y1, which is the Y coordinate indicating the position in the Y axis direction, and the position in the Z axis direction. Z1 which is the Z coordinate indicating is used and is expressed as (X1, Y1, Z1).
[0092]
On the other hand, in the spherical coordinate system, the azimuth position_azimuth, the elevation position_elevation, and the radius position_radius are used to represent the position of the audio object OB11 in space.
[0093]
Now, let the straight line connecting the origin O and the position of the audio object OB11 in the listening space be a straight line r, and let the straight line obtained by projecting this straight line r on the XY plane be a straight line L.
[0094]
At this time, the angle θ between the X axis and the straight line L is set as the azimuth angle position_azimuth indicating the position of the audio object OB11, and this angle θ corresponds to the azimuth angle position_azimuth[i] shown in FIG.
[0095]
An angle φ formed by the straight line r and the XY plane is an elevation angle position_elevation indicating the position of the audio object OB11, and the length of the straight line r is a radius position_radius indicating the position of the audio object OB11.
[0096]
That is, the angle φ corresponds to the elevation angle position_elevation[i] shown in FIG. 4, and the length of the straight line r corresponds to the radius position_radius[i] shown in FIG.
[0097]
For example, the position of the origin O is the position of the listener (user) who listens to the sound of the content including the sound of the audio object, and the positive direction in the X direction (X axis direction), that is, the front direction in FIG. The front direction viewed from the listener, and the positive direction in the Y direction (Y-axis direction), that is, the right direction in FIG. 5, is the left direction viewed from the listener.
[0098]
In this way, in the audio object position information, the position of the audio object is represented by spherical coordinates.
[0099]
The position of the audio object in the listening space indicated by the audio object position information is a physical quantity that changes every predetermined time period. When reproducing the content, the sound image localization position of the audio object can be moved according to the change of the audio object position information.
[0100]
Next, a specific example of the rendering method selection by the rendering method selection unit 51 will be described with reference to FIGS. 6 to 8.
[0101]
6 to 8, parts corresponding to each other are designated by the same reference numerals, and the description thereof will be appropriately omitted. Further, although the present technology assumes that the listening space is a three-dimensional space, the present technology is also applicable when the listening space is a two-dimensional plane. 6 to 8, the description will be given assuming that the listening space is a two-dimensional plane for simplification of description.
[0102]
For example, as shown in FIG. 6, there is a listener U21 who is a user who listens to the sound of the content at the position of the origin O, and is used to reproduce the sound of the content on the circumference of a circle centered on the origin O and having a radius R SP. It is assumed that five speakers SP11 to SP15 are arranged. That is, on the horizontal plane including the origin O, the distance from the origin O to each of the speakers SP11 to SP15 is the radius R SP .
[0103]
Two audio objects OBJ1 and OBJ2 exist in the listening space. The origin O, that is, the distance from the listener U21 to the audio object OBJ1 is R OBJ1, and the distance from the origin O to the audio object OBJ2 is R OBJ2 .
[0104]
In particular, since the audio object OBJ1 is located outside the circle where the speakers are arranged here, the distance R OBJ1 has a value larger than the radius R SP .
[0105]
On the other hand, since the audio object OBJ2 is located inside the circle in which the speakers are arranged, the distance R OBJ2 is smaller than the radius R SP .
[0106]
The distance R OBJ1 and the distance R OBJ2 are the radius position_radius[i] included in the audio object position information of the audio object OBJ1 and the audio object OBJ2, respectively.
[0107]
The rendering method selection unit 51 selects a rendering method to be performed on the audio object OBJ1 and the audio object OBJ2 by comparing the predetermined radius R SP with the distance R OBJ1 and the distance R OBJ2 .
[0108]
Specifically, for example, when the distance from the origin O to the audio object is the radius R SP or more, the panning process is selected as the rendering method.
[0109]
On the other hand, when the distance from the origin O to the audio object is less than the radius R SP , head-related transfer function processing is selected as the rendering method.
[0110]
Therefore, in this example, the panning process is selected for the audio object OBJ1 whose distance R OBJ1 is the radius R SP or more, and the audio object position information and the audio object signal of the audio object OBJ1 are supplied to the panning processing unit 52. Then, in the panning processing unit 52, processing such as VBAP described with reference to FIG. 1 is performed as the panning processing on the audio object OBJ1.
[0111]
On the other hand, the head related transfer function processing is selected for the audio object OBJ2 whose distance R OBJ2 is less than the radius R SP , and the audio object position information and the audio object signal of the audio object OBJ2 are supplied to the head related transfer function processing unit 53. To be done.
[0112]
Then, the head related transfer function processing unit 53 performs head related transfer function processing using the head related transfer function on the audio object OBJ2 as shown in FIG. A processed output signal is generated.
[0113]
In the example shown in FIG. 7, first, the head related transfer function processing unit 53, based on the audio object position information of the audio object OBJ2, the left and right ears prepared in advance for the position of the audio object OBJ2 in the listening space. The head-related transfer function, more specifically, the head-related transfer function filter is read out.
[0114]
Here, for example, some points in the area inside the circle (on the origin O side) where the speakers SP11 to SP15 are arranged are sampling points. Then, for each of these sampling points, a head related transfer function indicating the transfer characteristic of sound from the sampling point to the ear of the listener U21 at the origin O is prepared in advance for each of the left and right ears, and the head related transfer function processing unit 53 is provided. Be held in.
[0115]
The head-related transfer function processing unit 53 reads the head-related transfer function at the sampling point closest to the position of the audio object OBJ2 as the head-related transfer function at the position of the audio object OBJ2. The head related transfer function at the position of the audio object OBJ2 may be generated from the head related transfer function at some sampling points near the position of the audio object OBJ2 by an interpolation process such as linear interpolation.
[0116]
In addition, for example, a head related transfer function for the position of the audio object OBJ2 may be stored in the metadata of the input bitstream. In such a case, the rendering method selection unit 51 supplies the audio object position information and the head related transfer function supplied from the core decoding processing unit 21 to the head related transfer function processing unit 53 as metadata.
[0117]
In the following, the head related transfer function for the position of the audio object will also be referred to as an object position related head related transfer function.
[0118]
Next, based on the position of the audio object OBJ2 in the listening space, the head-related transfer function processing unit 53 outputs, for the left and right ears of the listener U21, a signal of a sound presented to those ears as an output audio signal ( A speaker (channel) supplied as a head-related transfer function processing output signal) is selected. Hereinafter, the speaker that is the output destination of the output audio signal of the sound presented to the left or right ear of the listener U21 will be particularly referred to as a selected speaker.
[0119]
Here, for example, the head-related transfer function processing unit 53 selects the speaker SP11, which is located on the left side of the audio object OBJ2 as seen from the listener U21 and is closest to the audio object OBJ2, as the selected speaker for the left ear. To do. Similarly, the head-related transfer function processing unit 53 selects the speaker SP13, which is located on the right side of the audio object OBJ2 as seen from the listener U21 and is closest to the audio object OBJ2, as the selected speaker for the right ear. ..
[0120]
When the selection speakers for the left and right ears are selected in this way, the head related transfer function processing unit 53 obtains a head related transfer function for the arrangement positions of the selected speakers, more specifically, a head related transfer function filter.
[0121]
Specifically, for example, the head-related transfer function processing unit 53 appropriately performs interpolation processing based on the head-related transfer function of each sampling point that is held in advance, and the head at each position of the speaker SP11 and the speaker SP13. Generate a transfer function.
[0122]
Alternatively, the head related transfer function for the placement position of each speaker may be held in advance in the head related transfer function processing unit 53, or the head related transfer function for the placement position of the selected speaker may be used as metadata. It may be stored in the input bitstream.
[0123]
In the following, the head related transfer function of the placement position of the selected speaker is also referred to as a speaker position related head transfer function.
[0124]
Further, the head-related transfer function processing unit 53 convolves the audio object signal of the audio object OBJ2 with the left-ear object position head-related transfer function, and the resulting signal and the left-ear speaker position head-related transfer. The function and are convoluted to generate a left ear audio signal.
[0125]
Similarly, the head-related transfer function processing unit 53 convolves the audio object signal of the audio object OBJ2 with the right-ear object position head-related transfer function, and the resulting signal and the right-ear speaker position head. The partial transfer function is convoluted with the partial transfer function to generate a right ear audio signal.
[0126]
The left-ear audio signal and the right-ear audio signal present the sound of the audio object OBJ2 so that the listener U21 perceives the sound as if it comes from the position of the audio object OBJ2. It is a signal to do. That is, it is an audio signal that realizes sound image localization to the position of the audio object OBJ2.
[0127]
For example, by outputting sound from the speaker SP11 based on the left ear audio signal, the reproduced sound O2 SP11 is presented to the left ear of the listener U21, and at the same time, the sound is output from the speaker SP13 based on the right ear audio signal. It is assumed that the reproduced sound O2 SP13 is presented to the right ear of the listener U21 by outputting . In this case, the listener U21 perceives the sound of the audio object OBJ2 from the position of the audio object OBJ2 as if it were heard.
[0128]
In FIG. 7, the reproduced sound O2 SP11 is represented by the arrow connecting the speaker SP11 and the left ear of the listener U21, and the reproduced sound O2 SP13 is represented by the arrow connecting the speaker SP13 and the right ear of the listener U21. There is.
[0129]
However, when the sound is actually output from the speaker SP11 based on the left ear audio signal, the sound reaches not only the left ear of the listener U21 but also the right ear.
[0130]
In FIG. 7, when the sound is output from the speaker SP11 based on the audio signal for the left ear, the reproduced sound O2 SP11-CT propagating from the speaker SP11 to the right ear of the listener U21 is output from the speaker SP11 and the listener U21. It is represented by the arrow that connects it to the right ear.
[0131]
The reproduced sound O2 SP11-CT is a crosstalk component of the reproduced sound O2 SP11 leaking to the right ear of the listener U21 . That is, the reproduced sound O2 SP11-CT is a crosstalk component of the reproduced sound O2 SP11 that reaches the ear (here, the right ear) different from the purpose of the listener U21 .
[0132]
Similarly, when a sound is output from the speaker SP13 based on the audio signal for the right ear, the sound reaches not only the right ear of the target listener U21 but also the left ear of the target listener U21. become.
[0133]
In FIG. 7, when the sound is output from the speaker SP13 based on the audio signal for the right ear, the reproduced sound O2 SP13-CT propagating from the speaker SP13 to the left ear of the listener U21 is output from the speaker SP13 and the listener U21. It is represented by the arrow that connects it to the left ear. This reproduced sound O2 SP13-CT is a crosstalk component of the reproduced sound O2 SP13 .
[0134]
The reproduced sound O2 SP11-CT and reproduced sound O2 SP13-CT , which are crosstalk components, are factors that significantly impede the sound image reproducibility, and thus spatial transfer function correction processing including crosstalk correction is generally performed. ..
[0135]
That is, the head related transfer function processing unit 53 generates a cancel signal for canceling the reproduced sound O2 SP11-CT , which is a crosstalk component, based on the audio signal for the left ear, and the audio signal for the left ear and the cancel signal. The final left ear audio signal is generated based on and. Then, the final left ear audio signal including the crosstalk cancellation component and the spatial transfer function correction component obtained in this way is set as the head related transfer function processing output signal of the channel corresponding to the speaker SP11. R.
[0136]
Similarly, the head-related transfer function processing unit 53 generates a cancel signal for canceling the reproduced sound O2 SP13-CT , which is a crosstalk component, on the basis of the right ear audio signal, and outputs the cancel signal as the right ear audio signal. A final audio signal for the right ear is generated based on the cancellation signal. Then, the final right-ear audio signal including the crosstalk cancellation component and the spatial transfer function correction component obtained in this way is used as the head related transfer function processing output signal of the channel corresponding to the speaker SP13. .
[0137]
The process of rendering to the speaker including the crosstalk correction process of generating the left ear audio signal and the right ear audio signal as described above is called a transaural process. Such transaural processing is described in detail, for example, in JP-A-2016-140039.
[0138]
Here, an example in which one speaker is selected for each of the left and right ears as the selected speaker has been described, but two or more speakers are selected for each of the left and right ears as the selected speaker, and each selected speaker is selected. Alternatively, the left ear audio signal and the right ear audio signal may be generated. For example, all speakers constituting the speaker system, such as the speakers SP11 to SP15, may be selected as the selection speaker.
[0139]
Furthermore, for example, when the output destination of the output audio signal is a reproduction device such as headphones of left and right channels, binaural processing may be performed as head-related transfer function processing. The binaural process is a rendering process for rendering an audio object (audio object signal) on an output unit such as headphones attached to the left and right ears, using a head related transfer function.
[0140]
In this case, for example, when the distance from the listening position to the audio object is equal to or larger than a predetermined distance, the panning process for distributing the gain to the left and right channels is selected as the rendering method. On the other hand, when the distance from the listening position to the audio object is less than the predetermined distance, the binaural process is selected as the rendering method.
[0141]
By the way, in the description of FIG. 6 , according to whether the distance from the origin O (listener U21) to the audio object is the radius R SP or more, the panning process or the head related transfer function is used as the rendering method of the audio object. It has been described that one of the processes is selected.
[0142]
However, for example, as shown in FIG. 8, the audio object may gradually approach the listener U21 with time from a position at a distance equal to or greater than the radius R SP .
[0143]
In FIG. 8, the audio object OBJ2, which is located at a position longer than the radius R SP when viewed from the listener U21 at a predetermined time, is depicted to approach the listener U21 with time.
[0144]
Here, a region inside a circle having a radius R SP centered on the origin O is a speaker radius region RG11, a region inside a circle having a radius R HRTF centered on the origin O is a HRTF region RG12, and a speaker radius region RG11 The region other than the HRTF region RG12 is defined as the transition region R TS .
[0145]
That is, the transition area R TS is an area in which the distance from the origin O (the listener U21) is the distance from the radius R HRTF to the radius R SP .
[0146]
Now, for example, the audio object OBJ2 gradually moves from the position outside the speaker radius region RG11 to the listener U21 side , reaches the position in the transition region R TS at a certain timing, and then further moves to the HRTF. It is assumed that the area RG12 is reached.
[0147]
In such a case, if the rendering method is selected depending on whether the distance to the audio object OBJ2 is the radius R SP or more, the rendering method is suddenly switched when the audio object OBJ2 reaches the inside of the transition region R TS. It will be. Then, a discontinuity may occur in the sound of the audio object OBJ2, which may cause discomfort.
[0148]
Therefore, when the audio object is located in the transition area R TS , both the panning process and the head-related transfer function process are selected as the rendering method so that a sense of discomfort does not occur at the timing of switching the rendering method. You can
[0149]
In this case, when the audio object is on the boundary of the speaker radius region RG11 or outside the speaker radius region RG11, the panning process is selected as the rendering method.
[0150]
Further, when the audio object is within the transition region R TS , that is, when the distance from the listening position to the audio object is equal to or greater than the radius R HRTF and less than the radius R SP , the panning process and the head related transfer function process are performed as rendering methods. Both are selected.
[0151]
Then, when the audio object is in the HRTF region RG12, head related transfer function processing is selected as the rendering method.
[0152]
In particular, when the audio object is in the transition region R TS , the mixing ratio (blend ratio) of the head related transfer function processing output signal and the panning processing output signal in the correction processing is changed according to the position of the audio object. It is possible to prevent the occurrence of sound discontinuity points of the audio object in the time direction.
[0153]
At this time, the correction process is performed such that the closer the audio object is to the boundary position of the speaker radius region RG11 in the transition region R TS , the closer the final output audio signal is to the panning process output signal.
[0154]
On the contrary, the closer the audio object is to the boundary position of the HRTF region RG12 in the transition region R TS , the more the correction processing is performed so that the final output audio signal becomes closer to the head-related transfer function processing output signal. Be seen.
[0155]
By doing so, it is possible to prevent the generation of sound discontinuity points of the audio object in the time direction, and to realize more natural and comfortable sound reproduction.
[0156]
Here, as a specific example of the correction process, in the case where the audio object OBJ2 is located in the transition area R TS at a position where the distance from the origin O is R 0 (where R HRTF ≦R 0
Next, a specific operation of the signal processing device 11 will be described. That is, the audio output processing by the signal processing device 11 will be described below with reference to the flowchart in FIG. 9. Note that, for simplification of description, it is assumed that only data of one audio object is stored in the input bitstream.
[0166]
In step S11, the core decoding processing unit 21 decodes (decodes) the received input bitstream, and supplies the audio object position information and the audio object signal obtained as a result to the rendering method selection unit 51.
[0167]
In step S12, the rendering method selection unit 51 determines, based on the audio object position information supplied from the core decoding processing unit 21, whether to perform the panning process as the rendering of the audio object.
[0168]
For example, in step S12, if the distance from the listener to the audio object indicated by the audio object position information is equal to or greater than the radius R HRTF described with reference to FIG. 8, it is determined that the panning process is performed. That is, at least panning processing is selected as the rendering method.
[0169]
In addition, in addition, when a user or the like operating the signal processing device 11 has an instruction input for instructing whether or not to perform the panning process, and execution of the panning process is designated (instructed) by the instruction input, step S12 is performed. It may be determined that the panning process is performed. In this case, the rendering method to be executed is selected by the instruction input by the user or the like.
[0170]
When it is determined that the panning process is not performed in step S12, the process of step S13 is not performed, and then the process proceeds to step S14.
[0171]
On the other hand, when it is determined in step S12 that the panning processing is performed, the rendering method selection unit 51 supplies the audio object position information and the audio object signal supplied from the core decoding processing unit 21 to the panning processing unit 52. Then, the process proceeds to step S13.
[0172]
In step S13, the panning processing unit 52 performs panning processing based on the audio object position information and the audio object signal supplied from the rendering method selection unit 51, and generates a panning processing output signal.
[0173]
For example, in step S13, the above-described VBAP or the like is performed as the panning process. The panning processing unit 52 supplies the panning processing output signal obtained by the panning processing to the mixing processing unit 54.
[0174]
When the process of step S13 is performed or when it is determined that the panning process is not performed in step S12, the process of step S14 is performed.
[0175]
In step S14, the rendering method selection unit 51 determines, based on the audio object position information supplied from the core decoding processing unit 21, whether or not to perform head related transfer function processing as the rendering of the audio object.
[0176]
For example, in step S14, when the distance from the listener to the audio object indicated by the audio object position information is less than the radius R SP described with reference to FIG. 8, it is determined that the head related transfer function processing is performed. That is, at least head related transfer function processing is selected as the rendering method.
[0177]
In addition, there is an instruction input for instructing whether or not to perform the head related transfer function processing by a user operating the signal processing device 11, and the execution of the head related transfer function processing is designated (instructed) by the instruction input. If so, it may be determined to perform the head related transfer function processing in step S14.
[0178]
When it is determined in step S14 that the head related transfer function process is not performed, the processes of steps S15 to S19 are not performed, and then the process proceeds to step S20.
[0179]
On the other hand, when it is determined that the head related transfer function processing is performed in step S14, the rendering method selection unit 51 sets the audio object position information and the audio object signal supplied from the core decoding processing unit 21 to the head related transfer function. It is supplied to the processing unit 53, and then the process proceeds to step S15.
[0180]
In step S15, the head related transfer function processing unit 53 acquires the object position head related transfer function of the position of the audio object based on the audio object position information supplied from the rendering method selection unit 51.
[0181]
For example, as the object position head related transfer function, one that is held in advance may be read out, it may be obtained by interpolation processing from a plurality of head related transfer functions that are held in advance, or from the input bit stream. It may be read.
[0182]
In step S16, the head related transfer function processing unit 53 selects the selected speaker based on the audio object position information supplied from the rendering method selection unit 51, and acquires the speaker position head related transfer function at the position of the selected speaker. ..
[0183]
For example, the speaker position head-related transfer function may be read out in advance, may be obtained by interpolation processing from a plurality of head-related transfer functions retained in advance, or may be calculated from the input bit stream. It may be read.
[0184]
In step S17, the head-related transfer function processing unit 53 convolves the audio object signal supplied from the rendering method selection unit 51 and the object-position head-related transfer function obtained in step S15 for each of the left and right ears.
[0185]
In step S18, the head-related transfer function processing unit 53 convolves the audio signal obtained in step S17 and the speaker-position head-related transfer function for each of the left and right ears. As a result, the left ear audio signal and the right ear audio signal are obtained.
[0186]
In step S19, the head related transfer function processing unit 53 generates a head related transfer function processing output signal based on the left ear audio signal and the right ear audio signal, and supplies the head related transfer function processing output signal to the mixing processing unit 54. For example, in step S19, the cancel signal is appropriately generated as described with reference to FIG. 7, and the final head-related transfer function processing output signal is generated.
[0187]
Through the processing of steps S15 to S19, the transaural processing described with reference to FIG. 8 is performed as the head related transfer function processing, and the head related transfer function processing output signal is generated. When the output destination of the output audio signal is not a speaker but a reproducing device such as headphones, binaural processing or the like is performed as head-related transfer function processing, and a head-related transfer function processing output signal is generated.
[0188]
If the process of step S19 is performed or if it is determined in step S14 that the head related transfer function process is not performed, then the process of step S20 is performed.
[0189]
In step S20, the mixing processing unit 54 synthesizes the panning processing output signal supplied from the panning processing unit 52 and the head related transfer function processing output signal supplied from the head related transfer function processing unit 53, and outputs an output audio signal. To generate.
[0190]
For example, in step S20, the calculation of the above equation (3) is performed as a correction process, and an output audio signal is generated.
[0191]
Note that, for example, when the processing of step S13 is performed and the processing of steps S15 to S19 is not performed, or when the processing of steps S15 to S19 is performed and the processing of step S13 is not performed, the correction is performed. No processing is done.
[0192]
That is, for example, when only the panning process is performed as the rendering process, the panning process output signal obtained as a result is directly used as the output audio signal. On the other hand, when only the head-related transfer function processing is performed as the rendering processing, the head-related transfer function processing output signal obtained as a result is directly used as the output audio signal.
[0193]
It should be noted that although an example in which the data of one audio object is included in the input bitstream has been described here, when the data of a plurality of audio objects is included, the mixing processing unit 54 performs mixing processing. Be seen. That is, the output audio signals obtained for each audio object are added (synthesized) for each channel to form one final output audio signal.
[0194]
When the output audio signal is obtained in this way, the mixing processing unit 54 outputs the obtained output audio signal to the subsequent stage, and the audio output processing ends.
[0195]
As described above, the signal processing device 11 selects one or more rendering methods from the plurality of rendering methods based on the audio object position information, that is, based on the distance from the listening position to the audio object. Then, the signal processing device 11 performs rendering by the selected rendering method to generate an output audio signal.
[0196]
By doing so, the reproducibility of the sound image can be improved with a small amount of calculation.
[0197]
That is, for example, when the audio object is located far from the listening position, the panning process is selected as the rendering method. In this case, since the audio object is located far away from the listening position, it is not necessary to consider the difference in the arrival time of the sound to the left and right ears of the listener, and the sound image is localized with sufficient reproducibility even with a small amount of calculation. be able to.
[0198]
On the other hand, for example, when the audio object is located near the listening position, head related transfer function processing is selected as the rendering method. In this case, the sound image can be localized with sufficient reproducibility, although the amount of calculation is somewhat increased.
[0199]
In this way, by appropriately selecting panning processing and head-related transfer function processing according to the distance from the listening position to the audio object, sound image localization with sufficient reproducibility can be achieved while keeping the calculation amount low overall. Can be realized. In other words, the sound image reproducibility can be improved with a small amount of calculation.
[0200]
In the above description, the panning process and the head related transfer function process are selected as the rendering method when the audio object is in the transition region R TS .
[0201]
However, if the distance to the audio object is greater than or equal to the radius R SP , panning is selected as the rendering method, and if the distance to the audio object is less than the radius R SP , head related transfer function processing is used as the rendering method. It may be selected.
[0202]
In this case, for example, when the head related transfer function processing is selected as the rendering method, the head related transfer function processing is performed by using the head related transfer function according to the distance from the listening position to the audio object. It is possible to prevent the occurrence of discontinuous points.
[0203]
Specifically, in the head-related transfer function processing unit 53, the head-related transfer functions of the left and right ears become more approximate as the distance to the audio object increases, that is, as the position of the audio object approaches the boundary position of the speaker radius region RG11. Just make sure they are the same.
[0204]
In other words, in the head-related transfer function processing unit 53, the closer the distance to the audio object is to the radius R SP , the higher the degree of similarity between the head-related transfer function for the left ear and the head-related transfer function for the right ear. First, the head related transfer functions of the left and right ears used for head related transfer function processing are selected.
[0205]
For example, increasing the degree of similarity of the head related transfer function may be a decrease in the difference between the head related transfer function for the left ear and the head related transfer function for the right ear. In this case, for example, when the distance to the audio object becomes approximately the radius R SP , the head-related transfer function common to the left and right ears is used.
[0206]
Conversely, in the head-related transfer function processing unit 53, the shorter the distance to the audio object, that is, the closer the audio object is to the listening position, the head-related transfer function of each of the left and right ears, and the actual position of the audio object. The one close to the head related transfer function obtained by the measurement is used.
[0207]
By doing so, it is possible to prevent the occurrence of discontinuity points and realize natural sound reproduction without a feeling of strangeness. This is because when a head related transfer function processing output signal is generated using the same head related transfer function for the left and right ears, the head related transfer function processing output signal becomes the same as the panning processing output signal. Is.
[0208]
Therefore, by using the head related transfer function of each of the left and right ears depending on the distance from the listening position to the audio object, it is possible to obtain the same effect as the correction process of the above-described formula (3).
[0209]
Furthermore, when selecting the rendering method, the availability of resources of the signal processing device 11 and the importance of the audio object may be taken into consideration.
[0210]
For example, the rendering method selection unit 51 can allocate many resources for rendering when the signal processing device 11 has sufficient resources, and therefore selects the head related transfer function processing as the rendering method. On the contrary, the rendering method selection unit 51 selects the panning process as the rendering method when the availability of the resources of the signal processing device 11 is small.
[0211]
Further, for example, the rendering method selection unit 51 selects the head related transfer function processing as the rendering method when the importance of the audio object to be processed is equal to or higher than the predetermined importance. On the other hand, when the importance of the audio object to be processed is less than the predetermined importance, the rendering method selection unit 51 selects the panning process as the rendering method.
[0212]
As a result, it is possible to reduce the processing amount by localizing the sound image with higher reproducibility for the audio object having high importance and by locating the sound image with lesser reproducibility for the audio object having lower importance. As a result, the sound image reproducibility can be improved with a small amount of calculation as a whole.
[0213]
When selecting a rendering method based on the importance of audio objects, the importance of each audio object may be included in the input bitstream as metadata of those audio objects. Further, the importance of the audio object may be designated by an external operation input or the like.
[0214]
Also, an example in which trans-oral processing is performed as head-related transfer function processing has been described above. That is, the example in which the rendering to the speaker is performed in the head-related transfer function processing has been described.
[0215]
However, in addition, as head-related transfer function processing, rendering for headphone reproduction may be performed using, for example, the concept of a virtual speaker.
[0216]
For example, when a large number of audio objects are rendered on a headphone or the like, the calculation cost for performing the head related transfer function processing becomes large, as in the case of rendering for a speaker.
[0217]
Even in headphone rendering according to the MPEG-H Part 3:3D audio standard, all audio objects are once panned (rendered) to a virtual speaker by VBAP, and then the head related transfer function from the virtual speaker is used. Rendered to.
[0218]
In this way, the output audio signal is output to a playback device such as headphones that plays back left and right channels, and after rendering to the virtual speaker once, the playback device using the head-related transfer function is further used. The present technology is applicable even when rendering is performed.
[0219]
In such a case, the rendering method selection unit 51 may regard each of the speakers SP11 to SP15 shown in FIG. 8 as a virtual speaker and select one or more of the rendering methods at the time of rendering from the plurality of rendering methods. ..
[0220]
For example, if the distance from the listening position to the audio object is R SP or more, that is, if the audio object is far from the listening position farther than the virtual speaker position, panning is selected as the rendering method. It should be done.
[0221]
In this case, rendering on the virtual speaker is performed by the panning process. Then, based on the audio signal obtained by the panning process and the head-related transfer function for each of the left and right ears from the virtual speaker to the listening position, the head-related transfer function process further renders on a playback device such as headphones. And an output audio signal is generated.
[0222]
On the other hand, when the distance to the audio object is less than the radius R SP , head-related transfer function processing may be selected as the rendering method. In this case, by the binaural process as the head related transfer function process, the rendering is directly performed on the reproducing device such as the headphone and the output audio signal is generated.
[0223]
By doing so, it is possible to realize sound image localization with high reproducibility while suppressing the rendering processing amount as a whole. That is, the reproducibility of the sound image can be improved with a small amount of calculation.
[0224]
In
addition, when selecting a rendering method, that is, when switching a rendering method, one of the parameters necessary for selecting the rendering method at each time such as a frame All or part may be stored in the input bitstream and transmitted.
[0225]
In such a case, the encoding format based on the present technology, that is, the metadata of the audio object is as illustrated in FIG. 10, for example.
[0226]
In the example shown in FIG. 10, in addition to the example shown in FIG. 4 described above, “radius_hrtf” and “radius_panning” are further stored in the metadata.
[0227]
Here, radius_hrtf is information (parameter) indicating the distance from the listening position (origin O), which is used to determine whether or not to select the head related transfer function processing as the rendering method. On the other hand, radius_panning is information (parameter) indicating the distance from the listening position (origin O), which is used for determining whether to select the panning process as the rendering method.
[0228]
Therefore, in the example illustrated in FIG. 10, the metadata stores audio object position information of each audio object, the distance radius_hrtf, and the distance radius_panning, and these pieces of information are read by the core decoding processing unit 21 as metadata. It is output and supplied to the rendering method selection unit 51.
[0229]
In this case, the rendering method selection unit 51 selects the head-related transfer function processing as the rendering method if the distance from the listener to the audio object is the distance radius_hrtf or less, regardless of the radius R SP indicating the distance to each speaker. To do. Further, the rendering method selection unit 51 does not select the head-related transfer function processing as the rendering method if the distance from the listener to the audio object is longer than the distance radius_hrtf.
[0230]
Similarly, the rendering method selection unit 51 selects the panning process as the rendering method if the distance from the listener to the audio object is equal to or more than the distance radius_panning. In addition, the rendering method selection unit 51 does not select the panning process as the rendering method when the distance from the listener to the audio object is shorter than the distance radius_panning.
[0231]
The distance radius_hrtf and the distance radius_panning may be the same distance or different distances from each other. In particular, if the distance radius_hrtf is greater than the distance radius_panning, and if the distance from the listener to the audio object is greater than or equal to the distance radius_panning and less than or equal to the distance radius_hrtf, both panning processing and head-related transfer function processing are selected as rendering methods. Will be.
[0232]
In this case, in the mixing processing unit 54, the above equation (3) is calculated based on the panning processing output signal and the head related transfer function processing output signal to generate an output audio signal. That is, by the correction process, the panning process output signal and the head-related transfer function process output signal are proportionally divided according to the distance from the listener to the audio object, and the output audio signal is generated.
[0233]
Furthermore, on the output side of the input bitstream, that is, the content creator side, a rendering method at each time such as a frame for each audio object is set. Alternatively, the selection instruction information indicating the selection result may be stored in the input bitstream as metadata.
[0234]
The selection instruction information is information indicating an instruction to select a rendering method for the audio object, and the rendering method selection unit 51 renders based on the selection instruction information supplied from the core decoding processing unit 21. Select a method. In other words, the rendering method selection unit 51 selects the rendering method designated by the selection instruction information for the audio object signal.
[0235]
When the selection instruction information is stored in the input bitstream in this way, the encoding format based on the present technology, that is, the metadata of the audio object is as illustrated in FIG. 11, for example.
[0236]
In the example shown in FIG. 11, in addition to the example shown in FIG. 4 described above, “flg_rendering_type” is further stored in the metadata.
[0237]
flg_rendering_type is selection instruction information indicating which rendering method is used. In particular, here, the selection instruction information flg_rendering_type is flag information (parameter) indicating whether the panning process or the head-related transfer function process is selected as the rendering method.
[0238]
Specifically, for example, the value “0” of the selection instruction information flg_rendering_type indicates that the panning process is selected as the rendering method. On the other hand, the value “1” of the selection instruction information flg_rendering_type indicates that the head related transfer function process is selected as the rendering method.
[0239]
For example, such selection instruction information flg_rendering_type is stored in the metadata for each audio object for each frame (each time).
[0240]
Therefore, in the example shown in FIG. 11, the metadata stores audio object position information and selection instruction information flg_rendering_type for each audio object, and these pieces of information are read by the core decoding processing unit 21 as metadata. And is supplied to the rendering method selection unit 51.
[0241]
In this case, the rendering method selection unit 51 selects the rendering method according to the value of the selection instruction information flg_rendering_type regardless of the distance from the listener to the audio object. That is, the rendering method selection unit 51 selects the panning process as the rendering method when the value of the selection instruction information flg_rendering_type is “0”, and the head transmission as the rendering method when the value of the selection instruction information flg_rendering_type is “1”. Select function processing.
[0242]
In addition, although the example in which the value of the selection instruction information flg_rendering_type is either “0” or “1” has been described here, the selection instruction information flg_rendering_type may be any one of a plurality of values of three or more types. Good. For example, when the value of the selection instruction information flg_rendering_type is “2”, the panning process and the head related transfer function process may be selected as the rendering method.
[0243]
As described above, according to the present technology, as described in the first modification of the first embodiment to the third embodiment, for example, even when a large number of audio objects are present, the operation amount is high while suppressing the calculation amount. Sound image expression with reproducibility can be realized.
[0244]
In particular, the present technology is applicable not only to speaker reproduction using an actual speaker but also to headphone reproduction by rendering using a virtual speaker.
[0245]
Furthermore, according to the present technology, the content creator can control the selection of the rendering method by storing the parameters required for the selection of the rendering method as metadata in the encoding standard, that is, in the input bitstream. Becomes
[0246]
By the way, the series of processes described above can be executed by hardware or software. When the series of processes is executed by software, a program forming the software is installed in the computer. Here, the computer includes a computer incorporated in dedicated hardware and, for example, a general-purpose personal computer capable of executing various functions by installing various programs.
[0247]
FIG. 12 is a block diagram showing a configuration example of hardware of a computer that executes the series of processes described above by a program.
[0248]
In a computer, a CPU (Central Processing Unit) 501, a ROM (Read Only Memory) 502, and a RAM (Random Access Memory) 503 are connected to each other by a bus 504.
[0249]
An input/output interface 505 is further connected to the bus 504. An input unit 506, an output unit 507, a recording unit 508, a communication unit 509, and a drive 510 are connected to the input/output interface 505.
[0250]
The input unit 506 includes a keyboard, a mouse, a microphone, an image sensor, and the like. The output unit 507 includes a display, a speaker and the like. The recording unit 508 includes a hard disk, a non-volatile memory, or the like. The communication unit 509 includes a network interface or the like. The drive 510 drives a removable recording medium 511 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.
[0251]
In the computer configured as described above, the CPU 501 loads the program recorded in the recording unit 508 into the RAM 503 via the input/output interface 505 and the bus 504 and executes the program, thereby performing the above-described series of operations. Is processed.
[0252]
The program executed by the computer (CPU 501) can be provided by being recorded in a removable recording medium 511 such as a package medium, for example. Further, the program can be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting.
[0253]
In the computer, the program can be installed in the recording unit 508 via the input/output interface 505 by mounting the removable recording medium 511 in the drive 510. Further, the program can be received by the communication unit 509 via a wired or wireless transmission medium and installed in the recording unit 508. In addition, the program can be installed in the ROM 502 or the recording unit 508 in advance.
[0254]
The program executed by the computer may be a program in which processing is performed in time series in the order described in this specification, or in parallel or at a necessary timing such as when a call is made. It may be a program in which processing is performed.
[0255]
The embodiments of the present technology are not limited to the above-described embodiments, and various modifications can be made without departing from the scope of the present technology.
[0256]
For example, the present technology may have a configuration of cloud computing in which a plurality of devices share one function via a network and jointly process the functions.
[0257]
In addition, each step described in the above-described flowcharts can be executed by one device or shared by a plurality of devices.
[0258]
Furthermore, when one step includes a plurality of processes, the plurality of processes included in the one step can be executed by one device or shared by a plurality of devices.
[0259]
Furthermore, the present technology may be configured as below.
[0260]
(1)
A rendering method selection unit that selects one or more rendering processing methods for locating a sound image of an audio signal in a listening space from a plurality of mutually different methods; and a
method selected by the rendering method selection unit. A
signal processing apparatus comprising: a rendering processing unit that performs the rendering processing of an audio signal .
(2) The signal processing device according to (1),
wherein the audio signal is an audio signal of an audio object
.
(3) The signal processing device according to (1) or (2),
wherein the plurality of methods include panning processing
.
(4) The signal processing device according to any one of (1) to (3),
wherein the plurality of methods include the rendering process using a head related transfer function
.
(5) The signal processing device according to (4),
wherein the rendering process using the head-related transfer function is a transaural process or a binaural process
.
(6) The signal processing device according
to
(2), wherein the rendering method selection unit selects the rendering processing method based on the position of the audio object in the listening space .
(7) The signal processing according
to
(6) , wherein when the distance from the listening position to the audio object is equal to or greater than a predetermined first distance, the rendering method selection unit selects panning processing as the rendering processing method. apparatus.
(8) The signal processing according
to
(7) , wherein, when the distance is less than the first distance, the rendering method selection unit selects the rendering processing using a head related transfer function as the rendering processing method. apparatus.
(9)
When the distance is less than the first distance, the rendering processing unit performs the rendering processing by using the head related transfer function according to the distance from the listening position to the audio object
( The signal processing device according to 8).
(10) The
rendering processing unit may reduce the difference between the head related transfer function for the left ear and the head related transfer function for the right ear as the distance becomes closer to the first distance.
The signal processing device according to (9), wherein the head related transfer function used for rendering processing is selected .
(11)
When the distance is less than a second distance different from the first distance, the rendering method selection unit selects the rendering processing using a head related transfer function as the rendering processing method
( The signal processing device according to 7).
(12)
When the distance is equal to or more than the first distance and less than the second distance, the rendering method selection unit performs the panning processing and the rendering processing using the head related transfer function as a method of the rendering processing.
The signal processing device according to (11), which is selected .
(13)
An output audio signal generation unit is further provided which synthesizes a signal obtained by the panning process and a signal obtained by the rendering process using the head related transfer function to generate an output audio signal
(12). ) Signal processing device described in.
(14) The signal processing device according to any one of (1) to (5),
wherein the rendering method selection unit selects a method designated for the audio signal as a method of the rendering process
.
(15) The
signal processing device
selects one or more rendering processing methods for locating the sound image of the audio signal in the listening space from a plurality of mutually different methods, and the rendering processing of the
audio signal is performed by the selected method.
Signal processing method for performing .
(16)
One or more rendering processing methods for locating the sound image of the audio signal in the listening space are selected from a plurality of different methods, and the rendering processing of the
audio signal is performed by the selected method.
A program that causes a computer to execute processing including steps.
Explanation of symbols
[0261]
11 signal processing device, 21 core decoding processing unit, 22 rendering processing unit, 51 rendering method selecting unit, 52 panning processing unit, 53 head related transfer function processing unit, 54 mixing processing unit
The scope of the claims
[Claim 1]
The method of rendering processing for localizing the sound image of the audio signal to the listening space, and rendering method selection unit for selecting one or more from among a plurality of different techniques,
of the audio signal by the method selected by the rendering method selection unit A
signal processing apparatus comprising: a rendering processing unit that performs the rendering processing .
[Claim 2]
The signal processing device according to claim 1, wherein the audio signal is an audio signal of an audio object .
[Claim 3]
The signal processing apparatus according to claim 1, wherein the plurality of methods include panning processing .
[Claim 4]
The signal processing device according to claim 1, wherein the plurality of methods include the rendering process using a head related transfer function .
[Claim 5]
The signal processing device according to claim 4, wherein the rendering process using the head-related transfer function is a transaural process or a binaural process .
[Claim 6]
The signal processing device according to claim 2, wherein the rendering method selection unit selects the rendering method based on a position of the audio object in the listening space .
[Claim 7]
The signal processing device according to claim 6, wherein the rendering method selection unit selects a panning process as a method of the rendering process when a distance from the listening position to the audio object is equal to or more than a predetermined first distance .
[Claim 8]
The signal processing device according to claim 7, wherein the rendering method selection unit selects the rendering processing using a head related transfer function as a method of the rendering processing when the distance is less than the first distance .
[Claim 9]
The rendering processing unit, when the distance is less than said first distance, performs the rendering process using the head-related transfer function according to the distance from the listening position to the audio objects
in claim 8 The signal processing device described.
[Claim 10]
The rendering processing unit performs the rendering processing such that the closer the distance is to the first distance, the smaller the difference between the head related transfer function for the left ear and the head related transfer function for the right ear becomes.
The signal processing device according to claim 9, wherein the head related transfer function to be used is selected .
[Claim 11]
The rendering method selection unit, when said distance is from said first distance is less than a different second distance, selects the rendering processing using the head-related transfer function as a technique of the rendering process
to claim 7 The signal processing device described.
[Claim 12]
When the distance is equal to or more than the first distance and less than the second distance, the rendering method selection unit performs the panning processing and the rendering processing using the head related transfer function as a method of the rendering processing.
The signal processing device according to claim 11, which is selected .
[Claim 13]
A signal obtained by the panning process, further comprising an output audio signal generator and a signal obtained by the rendering process combined to produce an output audio signal using the head-related transfer function
according to claim 12 Signal processing equipment.
[Claim 14]
The signal processing device according to claim 1, wherein the rendering method selection unit selects a method designated for the audio signal as a method of the rendering processing .
[Claim 15]
Signal processing apparatus,
a method of rendering processing for localizing the sound image of the audio signal to the listening space, and select one or more from among a plurality of different techniques,
performing the rendering processing of the audio signals by the selected technique
signal Processing method.
[Claim 16]
A computer
performs
a process including a step of selecting one or more rendering processing methods for locating a sound image of an audio signal in a listening space from among a plurality of different methods, and performing the rendering processing of the audio signal by the selected method. The program to run.
| # | Name | Date |
|---|---|---|
| 1 | 202017023646-FORM-26 [30-10-2024(online)].pdf | 2024-10-30 |
| 1 | 202017023646-TRANSLATIOIN OF PRIOIRTY DOCUMENTS ETC. [05-06-2020(online)].pdf | 2020-06-05 |
| 2 | 202017023646-Written submissions and relevant documents [06-09-2024(online)].pdf | 2024-09-06 |
| 2 | 202017023646-STATEMENT OF UNDERTAKING (FORM 3) [05-06-2020(online)].pdf | 2020-06-05 |
| 3 | 202017023646-PRIORITY DOCUMENTS [05-06-2020(online)].pdf | 2020-06-05 |
| 3 | 202017023646-Correspondence to notify the Controller [20-08-2024(online)].pdf | 2024-08-20 |
| 4 | 202017023646-US(14)-HearingNotice-(HearingDate-22-08-2024).pdf | 2024-07-25 |
| 4 | 202017023646-POWER OF AUTHORITY [05-06-2020(online)].pdf | 2020-06-05 |
| 5 | 202017023646-FORM 1 [05-06-2020(online)].pdf | 2020-06-05 |
| 5 | 202017023646-CLAIMS [22-09-2022(online)].pdf | 2022-09-22 |
| 6 | 202017023646-DRAWINGS [05-06-2020(online)].pdf | 2020-06-05 |
| 6 | 202017023646-CORRESPONDENCE [22-09-2022(online)].pdf | 2022-09-22 |
| 7 | 202017023646-DRAWING [22-09-2022(online)].pdf | 2022-09-22 |
| 7 | 202017023646-DECLARATION OF INVENTORSHIP (FORM 5) [05-06-2020(online)].pdf | 2020-06-05 |
| 8 | 202017023646-FER_SER_REPLY [22-09-2022(online)].pdf | 2022-09-22 |
| 8 | 202017023646-COMPLETE SPECIFICATION [05-06-2020(online)].pdf | 2020-06-05 |
| 9 | 202017023646-Proof of Right [21-07-2020(online)].pdf | 2020-07-21 |
| 9 | 202017023646-PETITION UNDER RULE 137 [22-09-2022(online)].pdf | 2022-09-22 |
| 10 | 202017023646-FER.pdf | 2022-03-23 |
| 10 | 202017023646.pdf | 2021-10-19 |
| 11 | 202017023646-FORM 18 [20-10-2021(online)].pdf | 2021-10-20 |
| 12 | 202017023646-FER.pdf | 2022-03-23 |
| 12 | 202017023646.pdf | 2021-10-19 |
| 13 | 202017023646-PETITION UNDER RULE 137 [22-09-2022(online)].pdf | 2022-09-22 |
| 13 | 202017023646-Proof of Right [21-07-2020(online)].pdf | 2020-07-21 |
| 14 | 202017023646-COMPLETE SPECIFICATION [05-06-2020(online)].pdf | 2020-06-05 |
| 14 | 202017023646-FER_SER_REPLY [22-09-2022(online)].pdf | 2022-09-22 |
| 15 | 202017023646-DECLARATION OF INVENTORSHIP (FORM 5) [05-06-2020(online)].pdf | 2020-06-05 |
| 15 | 202017023646-DRAWING [22-09-2022(online)].pdf | 2022-09-22 |
| 16 | 202017023646-CORRESPONDENCE [22-09-2022(online)].pdf | 2022-09-22 |
| 16 | 202017023646-DRAWINGS [05-06-2020(online)].pdf | 2020-06-05 |
| 17 | 202017023646-CLAIMS [22-09-2022(online)].pdf | 2022-09-22 |
| 17 | 202017023646-FORM 1 [05-06-2020(online)].pdf | 2020-06-05 |
| 18 | 202017023646-POWER OF AUTHORITY [05-06-2020(online)].pdf | 2020-06-05 |
| 18 | 202017023646-US(14)-HearingNotice-(HearingDate-22-08-2024).pdf | 2024-07-25 |
| 19 | 202017023646-Correspondence to notify the Controller [20-08-2024(online)].pdf | 2024-08-20 |
| 19 | 202017023646-PRIORITY DOCUMENTS [05-06-2020(online)].pdf | 2020-06-05 |
| 20 | 202017023646-Written submissions and relevant documents [06-09-2024(online)].pdf | 2024-09-06 |
| 20 | 202017023646-STATEMENT OF UNDERTAKING (FORM 3) [05-06-2020(online)].pdf | 2020-06-05 |
| 21 | 202017023646-TRANSLATIOIN OF PRIOIRTY DOCUMENTS ETC. [05-06-2020(online)].pdf | 2020-06-05 |
| 21 | 202017023646-FORM-26 [30-10-2024(online)].pdf | 2024-10-30 |
| 22 | 202017023646-PETITION UNDER RULE 137 [05-12-2024(online)].pdf | 2024-12-05 |
| 23 | 202017023646-PatentCertificate06-12-2024.pdf | 2024-12-06 |
| 24 | 202017023646-IntimationOfGrant06-12-2024.pdf | 2024-12-06 |
| 1 | 202017023646E_23-03-2022.pdf |