Sign In to Follow Application
View All Documents & Correspondence

Signal Processing Device And Method

Abstract: The present art relates to a signal processing device, method, and program capable of improving encoding efficiency. The signal processing device comprises: an acquisition unit for acquiring reverb information including at least one of space reverb information inherent in a space around an audio object and object reverb information inherent in the audio object, and an audio object signal of the audio object; and a reverb processing unit for producing a signal of a reverb component of the audio object on the basis of the reverb information and the audio object signal. The present art is applicable to the signal processing device.

Get Free WhatsApp Updates!
Notices, Deadlines & Correspondence

Patent Information

Application #
Filing Date
13 April 2020
Publication Number
35/2020
Publication Type
INA
Invention Field
ELECTRONICS
Status
Email
patents@remfry.com
Parent Application
Patent Number
Legal Status
Grant Date
2024-01-05
Renewal Date

Applicants

SONY CORPORATION
1-7-1, Konan, Minato-ku, Tokyo 1080075

Inventors

1. HONMA Hiroyuki
c/o SONY CORPORATION, 1-7-1, Konan, Minato-ku, Tokyo 1080075
2. TSUJI Minoru
c/o SONY CORPORATION, 1-7-1, Konan, Minato-ku, Tokyo 1080075
3. CHINEN Toru
c/o SONY CORPORATION, 1-7-1, Konan, Minato-ku, Tokyo 1080075

Specification

Title of invention: Signal processing device and method, and program
Technical field
[0001]
 The present technology relates to a signal processing device and method, and a program, and particularly relates to a signal processing device and method, and a program capable of improving coding efficiency.
Background technology
[0002]
 Conventionally, object audio technology has been used in movies and games, and a coding method capable of handling object audio has been developed. Specifically, for example, the MPEG (Moving Picture Experts Group)-H Part 3:3D audio standard, which is an international standard, is known (see Non-Patent Document 1, for example).
[0003]
 In such an encoding method, in addition to the conventional 2-channel stereo method and 5.1-channel multi-channel stereo method, a moving sound source is treated as an independent audio object, and the position information of the object is stored together with the signal data of the audio object. It can be encoded as metadata.
[0004]
 By doing so, reproduction can be performed in various viewing environments with different numbers of speakers. Further, it is possible to easily process the sound of the specific sound source at the time of reproduction, such as adjusting the volume of the sound of the specific sound source and adding an effect to the sound of the specific sound source, which were difficult with the conventional encoding method.
[0005]
 For example, in the standard of Non-Patent Document 1, a method called three-dimensional VBAP (Vector Based Amplitude Panning) (hereinafter, simply referred to as VBAP) is used for rendering processing.
[0006]
 This is one of the rendering methods generally called panning. Among the speakers existing on the surface of the sphere with the listening position as the origin, the gain for the three speakers closest to the audio object also existing on the surface of the sphere. Is a method of rendering by distributing.
[0007]
 The rendering of the audio object by such panning is premised on that all audio objects are on the surface of a sphere whose origin is the viewing position. Therefore, the sense of distance when the audio object is close to the viewing position or far from the viewing position is controlled only by the magnitude of the gain for the audio object.
[0008]
 However, in reality, the expression of the sense of distance is far from the actual experience unless the attenuation factor varies depending on the frequency component and the reflection of the space in which the audio object exists is not taken into consideration.
[0009]
 In order to reflect these effects in the listening experience, it is first considered that physical reflection and attenuation of the space are physically calculated and used as the final output audio signal. However, such a method is effective for moving image contents such as movies that can be produced over a very long calculation time, but is difficult in the case of rendering an audio object in real time.
[0010]
 In addition, the final output obtained by physically calculating the reflection and attenuation of the space is difficult to reflect the intention of the content creator. Especially for music works such as music clips, you can apply your favorite reverb processing to vocal tracks etc. , A format that easily reflects the intention of the content creator is required.
Prior art documents
Non-patent literature
[0011]
Non-Patent Document 1: INTERNATIONAL STANDARD ISO/IEC 23008-3 First edition 2015-10-15 Information technology-High efficiency coding and media delivery in heterogeneous environments-Part 3: 3D audio
Summary of the invention
Problems to be Solved by the Invention
[0012]
 Therefore, data such as coefficients necessary for reverb processing that takes spatial reflection and attenuation into account for each audio object is stored in a file or transmission stream together with the position information of the audio object, and these are used to output the final output audio signal. It is desirable to obtain it for real-time playback.
[0013]
 However, storing reverb processing data required for each audio object in each frame in a file or a transmission stream leads to an increase in transmission rate, and data transmission with high coding efficiency is required.
[0014]
 The present technology has been made in view of such a situation, and is capable of improving coding efficiency.
Means for solving the problem
[0015]
 A signal processing device according to an aspect of the present technology is a reverb information including at least one of spatial reverb information unique to a space around an audio object and object reverb information unique to the audio object, and the audio object. An acquisition unit that acquires an audio object signal and a reverb processing unit that generates a signal of a reverb component of the audio object based on the reverb information and the audio object signal are provided.
[0016]
 A signal processing method or program according to an aspect of the present technology is to provide reverb information including at least one of spatial reverb information unique to a space around an audio object and object reverb information unique to the audio object; Obtaining an audio object signal of an object and generating a signal of a reverb component of the audio object based on the reverb information and the audio object signal.
[0017]
 In one aspect of the present technology, reverb information including at least one of spatial reverb information unique to a space around an audio object and object reverb information unique to the audio object, and an audio object signal of the audio object. And a signal of a reverb component of the audio object is generated based on the reverb information and the audio object signal.
Effect of the invention
[0018]
 According to the 1 side of this art, coding efficiency can be raised.
[0019]
 Note that the effects described here are not necessarily limited and may be any effects described in the present disclosure.
Brief description of the drawings
[0020]
FIG. 1 is a diagram showing a configuration example of a signal processing device.
FIG. 2 is a diagram showing a configuration example of a rendering processing unit.
FIG. 3 is a diagram showing an example of syntax of audio object information.
FIG. 4 is a diagram showing a syntax example of object reverb information and spatial reverb information.
FIG. 5 is a diagram for explaining localization positions of reverb components.
FIG. 6 is a diagram explaining an impulse response.
FIG. 7 is a diagram illustrating a relationship between an audio object and a viewing position.
FIG. 8 is a diagram illustrating a direct sound component, an early reflection sound component, and a rear reverberation component.
FIG. 9 is a flowchart illustrating an audio output process.
FIG. 10 is a diagram showing a configuration example of an encoding device.
FIG. 11 is a flowchart illustrating an encoding process.
FIG. 12 is a diagram showing a configuration example of a computer.
MODE FOR CARRYING OUT THE INVENTION
[0021]
 Hereinafter, embodiments to which the present technology is applied will be described with reference to the drawings.
[0022]

The
 present technology adaptively selects a reverb parameter encoding method according to the relationship between an audio object and a viewing position, thereby achieving a reverb with high encoding efficiency. It enables the transmission of parameters.
[0023]
 FIG. 1 is a diagram illustrating a configuration example of an embodiment of a signal processing device to which the present technology is applied.
[0024]
 The signal processing device 11 shown in FIG. 1 has a core decoding processing unit 21 and a rendering processing unit 22.
[0025]
 The core decoding processing unit 21 receives the transmitted input bit stream, decodes (decodes) the input bit stream, and supplies the audio processing information and the audio object signal obtained as a result to the rendering processing unit 22. In other words, the core decoding processing unit 21 functions as an acquisition unit that acquires audio object information and audio object signal.
[0026]
 Here, the audio object signal is an audio signal for reproducing the sound of the audio object.
[0027]
 The audio object information is an audio object, that is, metadata of an audio object signal. The audio object information includes information about the audio object, which is necessary for the processing performed by the rendering processing unit 22.
[0028]
 Specifically, the audio object information includes object position information, direct sound gain, object reverb information, object reverb sound gain, spatial reverb information, and spatial reverb gain.
[0029]
 Here, the object position information is information indicating the position of the audio object in the three-dimensional space. For example, the object position information includes a horizontal angle indicating the horizontal position of the audio object viewed from the reference viewing position, a vertical angle indicating the vertical position of the audio object viewed from the viewing position, and the viewing position to the audio object. It consists of a radius indicating the distance of.
[0030]
 Further, the direct sound gain is a gain value used for gain adjustment when generating the direct sound component of the sound of the audio object.
[0031]
 For example, the rendering processing unit 22 generates a signal of a direct sound component from the audio object, a signal of an object-specific reverb sound, and a signal of a space-specific reverb sound when rendering an audio object, that is, an audio object signal.
[0032]
 In particular, the signal of the object-specific reverb sound or the space-specific reverb sound is a signal of a component such as a reflected sound or a reverberation sound of the sound from the audio object, that is, a reverb component obtained by performing reverb processing on the audio object signal. It is a signal.
[0033]
 The object-specific reverb sound is an initial reflected sound component of the sound of the audio object, and is a sound with a large contribution of the state of the audio object such as the position of the audio object in the three-dimensional space. That is, the object-specific reverb sound is a reverb sound that depends on the position of the audio object and that changes greatly depending on the relative positional relationship between the viewing position and the audio object.
[0034]
 On the other hand, the space-specific reverb sound is the rear reverberation component of the sound of the audio object, and the contribution rate of the state of the audio object is small, and the contribution rate of the environment of the audio object, that is, the state of the space around the audio object. Is a loud sound.
[0035]
 That is, the space-specific reverb sound greatly changes depending on the relative positional relationship between the viewing position and the wall in the space around the audio object, the material of the wall or the floor, etc., but the relative position between the viewing position and the audio object. It hardly changes depending on the relationship. Therefore, it can be said that the space-specific reverb sound is a sound that depends on the space around the audio object.
[0036]
 During the rendering processing in the rendering processing unit 22, such direct sound component, object-specific reverb sound component, and space-specific reverb sound component from the audio object are generated by the reverb processing on the audio object signal. The direct sound gain is used to generate a signal of such a direct sound component.
[0037]
 The object reverb information is information about object-specific reverb sounds. For example, the object reverb information includes object reverb position information indicating the localization position of the sound image of the object-specific reverb sound, and coefficient information used to generate the object-specific reverb sound component during the reverb processing.
[0038]
 Since the object-specific reverb sound is a component peculiar to the audio object, the object reverb information can be said to be reverb information peculiar to the audio object used for generating the object-specific reverb sound component during the reverb processing.
[0039]
 Note that, hereinafter, the localization position of the sound image of the object-specific reverb sound in the three-dimensional space indicated by the object reverb position information is also referred to as the object reverb component position. It can be said that the object reverb component position is the arrangement position of the real speaker or the virtual speaker that outputs the object-specific reverb sound in the three-dimensional space.
[0040]
 The object reverb sound gain included in the audio object information is a gain value used for gain adjustment of the object-specific reverb sound.
[0041]
 The spatial reverb information is information about the spatial reverb sound. For example, the spatial reverb information includes spatial reverb position information indicating the localization position of the sound image of the spatial peculiar reverb sound, and coefficient information used for generating the spatial peculiar reverb sound component during the reverb processing.
[0042]
 Since the space-specific reverb sound is a space-specific component with a low contribution rate of the audio object, the spatial reverb information is used to generate the space-specific reverb sound component during the reverb processing, and the reverb information unique to the space around the audio object. It can be said that
[0043]
 Note that, hereinafter, the localization position of the sound image of the space-specific reverb sound in the three-dimensional space, which is indicated by the spatial reverb position information, is also referred to as the spatial reverb component position. It can be said that the position of the spatial reverb component is the arrangement position of the real speaker or the virtual speaker that outputs the space-specific reverb sound in the three-dimensional space.
[0044]
 The spatial reverb gain is a gain value used for adjusting the gain of the object-specific reverb sound.
[0045]
 The audio object information output from the core decoding processing unit 21 includes at least object position information among object position information, direct sound gain, object reverb information, object reverb sound gain, spatial reverb information, and spatial reverb gain. ing.
[0046]
 The rendering processing unit 22 generates an output audio signal based on the audio object information and the audio object signal supplied from the core decoding processing unit 21, and supplies the output audio signal to a speaker, a recording unit, or the like in the subsequent stage.
[0047]
 That is, the rendering processing unit 22 performs reverb processing based on the audio object information to generate a direct sound signal, an object-specific reverb sound signal, and a space-specific reverb sound signal of each of one or more audio objects.
[0048]
 Then, the rendering processing unit 22 performs a rendering process by VBAP for each of the obtained direct sound, object-specific reverb sound, and space-specific reverb sound signals, depending on a reproduction device such as a speaker system or headphones as an output destination. Generate an output audio signal of the channel configuration. Further, the rendering processing unit 22 adds the signals of the same channel of the output audio signal generated for each signal to obtain one final output audio signal.
[0049]
 When a sound is reproduced based on the output audio signal obtained in this way, the sound image of the direct sound of the audio object is localized at the position indicated by the object position information, and the sound image of the object-specific reverb sound is localized at the object reverb component position. Then, the sound image of the space-specific reverb sound is localized at the position of the space reverb component. Thereby, more realistic audio reproduction in which the sense of distance of the audio object is appropriately controlled is realized.
[0050]

 Next, a more detailed configuration example of the rendering processing unit 22 of the signal processing device 11 shown in FIG. 1 will be described.
[0051]
 Here, as a specific example, a case where there are two audio objects will be described. The number of audio objects may be any number, and it is possible to handle as many audio objects as the computing resources allow.
[0052]
 In the following, when distinguishing two audio objects, one audio object is also referred to as audio object OBJ1 and the audio object signal of the audio object OBJ1 is also referred to as audio object signal OA1. The other audio object is also referred to as an audio object OBJ2, and the audio object signal of the audio object OBJ2 is also referred to as an audio object signal OA2.
[0053]
 Further, hereinafter, the object position information, the direct sound gain, the object reverb information, the object reverb sound gain, and the spatial reverb gain for the audio object OBJ1, especially the object position information OP1, the direct sound gain OG1, the object reverb information OR1, the object reverb, Also referred to as sound gain RG1 and spatial reverb gain SG1.
[0054]
 Similarly, hereinafter, the object position information about the audio object OBJ2, direct sound gain, object reverb information, object reverb sound gain, and spatial reverb gain, especially object position information OP2, direct sound gain OG2, object reverb information OR2, object Also referred to as reverb sound gain RG2 and spatial reverb gain SG2.
[0055]
 When there are two audio objects in this way, the rendering processing unit 22 is configured as shown in FIG. 2, for example.
[0056]
 In the example shown in FIG. 2, the rendering processing unit 22 includes an amplification unit 51-1, an amplification unit 51-2, an amplification unit 52-1, an amplification unit 52-2, an object-specific reverb processing unit 53-1, and an object-specific reverb processing. It has a unit 53-2, an amplification unit 54-1, an amplification unit 54-2, a space-specific reverb processing unit 55, and a rendering unit 56.
[0057]
 The amplifying unit 51-1 and the amplifying unit 51-2, for the audio object signal OA1 and the audio object signal OA2 supplied from the core decoding processing unit 21, directly output the direct sound gain OG1 and the direct sound gain OG1 supplied from the core decoding processing unit 21, respectively. The gain adjustment is performed by multiplying the sound gain OG2, and the resulting direct sound signal of the audio object is supplied to the rendering unit 56.
[0058]
 Note that, hereinafter, the amplifier 51-1 and the amplifier 51-2 are simply referred to as the amplifier 51 unless it is necessary to distinguish them.
[0059]
 The amplifying unit 52-1 and the amplifying unit 52-2, for the audio object signal OA1 and the audio object signal OA2 supplied from the core decoding processing unit 21, the object reverb sound gain RG1 and the object reverb sound gain RG1 supplied from the core decoding processing unit 21, respectively. Gain is adjusted by multiplying the object reverb sound gain RG2. By this gain adjustment, the loudness of the reverb sound unique to each object is adjusted.
[0060]
 The amplification unit 52-1 and the amplification unit 52-2 supply the gain-adjusted audio object signal OA1 and audio object signal OA2 to the object-specific reverb processing unit 53-1 and the object-specific reverb processing unit 53-2.
[0061]
 Note that, hereinafter, the amplifying unit 52-1 and the amplifying unit 52-2 will be simply referred to as the amplifying unit 52 unless it is necessary to distinguish them.
[0062]
 The object-specific reverb processing unit 53-1 performs reverb processing on the gain-adjusted audio object signal OA1 supplied from the amplification unit 52-1 based on the object reverb information OR1 supplied from the core decoding processing unit 21. To do.
[0063]
 By this reverb processing, one or a plurality of object-specific reverb sound signals for the audio object OBJ1 are generated.
[0064]
 Further, the object-specific reverb processing unit 53-1 uses the object position information OP1 supplied from the core decoding processing unit 21 and the object reverb position information included in the object reverb information OR1 to identify each object in the three-dimensional space. Position information indicating the absolute localization position of the sound image of the unique reverb sound is generated.
[0065]
 As described above, the object position information OP1 is information including a horizontal angle, a vertical angle, and a radius indicating the absolute position of the audio object OBJ1 with reference to the viewing position in the three-dimensional space.
[0066]
 On the other hand, the object reverb position information can be information indicating the position (localization position) of the absolute sound image of the object-specific reverb sound viewed from the viewing position in the three-dimensional space, or in the three-dimensional space. It is also possible to use information indicating the position (localization position) of the sound image of the object-specific reverb sound relative to the audio object OBJ1 in.
[0067]
 For example, when the object reverb position information is information indicating the position of the absolute sound image of the object-specific reverb sound viewed from the viewing position in the three-dimensional space, the object reverb position information is based on the viewing position in the three-dimensional space. The horizontal angle, the vertical angle, and the radius indicating the absolute localization position of the sound image of the object-specific reverb sound.
[0068]
 In this case, the object-specific reverb processing unit 53-1 uses the object reverb position information as it is as position information indicating the absolute position of the sound image of the object-specific reverb sound.
[0069]
 On the other hand, when the object reverb position information is information indicating the position of the sound image of the object-specific reverb sound relative to the audio object OBJ1, the object reverb position information is the object-specific reverb sound viewed from the viewing position in the three-dimensional space. The information is composed of a horizontal angle, a vertical angle, and a radius that indicate the relative position of the sound image of the sound object to the audio object OBJ1.
[0070]
 In this case, the object-specific reverb processing unit 53-1 determines the absolute localization position of the sound image of the object-specific reverb sound based on the viewing position in the three-dimensional space, based on the object position information OP1 and the object reverb position information. The information including the horizontal angle, the vertical angle, and the radius is generated as position information indicating the absolute position of the sound image of the object-specific reverb sound.
[0071]
 The object-specific reverb processing unit 53-1 renders the pair of the signal of the object-specific reverb sound and the position information of the object-specific reverb sound obtained for each of the one or more object-specific reverb sounds in this way. Supply to.
[0072]
 In this way, by generating the signal of the object-specific reverb sound and the position information by the reverb processing, the signal of each object-specific reverb sound can be treated as an independent audio object signal.
[0073]
 Similarly, the object-specific reverb processing section 53-2, based on the object reverb information OR2 supplied from the core decoding processing section 21, with respect to the gain-adjusted audio object signal OA2 supplied from the amplification section 52-2. Perform reverb processing.
[0074]
 By this reverb processing, one or a plurality of object-specific reverb sound signals for the audio object OBJ2 are generated.
[0075]
 In addition, the object-specific reverb processing unit 53-2, based on the object position information OP2 supplied from the core decoding processing unit 21 and the object reverb position information included in the object reverb information OR2, each object in the three-dimensional space. Position information indicating the absolute localization position of the sound image of the peculiar reverb sound is generated.
[0076]
 Then, the object-specific reverb processing unit 53-2 supplies the rendering unit 56 with the pair of the object-specific reverb sound signal thus obtained and the position information of the object-specific reverb sound.
[0077]
 Note that, hereinafter, the object-specific reverb processing unit 53-1 and the object-specific reverb processing unit 53-2 are simply referred to as the object-specific reverb processing unit 53 unless it is necessary to distinguish them.
[0078]
 The amplifying unit 54-1 and the amplifying unit 54-2, for the audio object signal OA1 and the audio object signal OA2 supplied from the core decoding processing unit 21, receive the spatial reverb gain SG1 and the space supplied from the core decoding processing unit 21. Adjust the gain by multiplying the reverb gain SG2. By this gain adjustment, the loudness of the reverb sound unique to each space is adjusted.
[0079]
 Further, the amplification unit 54-1 and the amplification unit 54-2 supply the gain-adjusted audio object signal OA1 and audio object signal OA2 to the space-specific reverb processing unit 55.
[0080]
 Note that, hereinafter, the amplifier 54-1 and the amplifier 54-2 will be simply referred to as the amplifier 54 unless it is necessary to distinguish them.
[0081]
 The space-specific reverb processing unit 55, based on the spatial reverb information supplied from the core decoding processing unit 21, receives the gain-adjusted audio object signal OA1 and the audio object supplied from the amplifying unit 54-1 and the amplifying unit 54-2. Reverb processing is performed on the signal OA2. Further, the space-specific reverb processing unit 55 adds the signals obtained by the reverb processing for the audio object OBJ1 and the audio object OBJ2 to generate a space-specific reverb sound signal. The space-specific reverb processing unit 55 generates one or a plurality of space-specific reverb sound signals.
[0082]
 Further, the space-specific reverb processing unit 55, similar to the case in the object-specific reverb processing unit 53, the space reverb position information included in the space reverb information supplied from the core decoding processing unit 21, and the object position information OP1. It is generated based on the object position information OP2 as position information indicating the absolute localization position of the sound image of the space-specific reverb sound.
[0083]
 This position information is, for example, information including a horizontal angle, a vertical angle, and a radius indicating an absolute localization position of a sound image of a space-specific reverb sound based on a viewing position in a three-dimensional space.
[0084]
 The space peculiar reverb processing unit 55 supplies the pair of the signal and position information of the space peculiar reverb sound regarding the one or more space peculiar reverb sounds thus obtained to the rendering unit 56. Note that these space-specific reverb sounds, like the object-specific reverb sounds, can be treated as independent audio object signals because they have position information.
[0085]
 The amplification unit 51 to the space-specific reverb processing unit 55 function as a processing block that is provided in the preceding stage of the rendering unit 56 and that constitutes a reverb processing unit that performs reverb processing based on audio object information and an audio object signal.
[0086]
 The rendering unit 56 performs a rendering process by VBAP based on the supplied signals of the respective sounds and the position information of the signals of the sounds, and generates an output audio signal composed of the signals of the respective channels having a predetermined channel configuration. ,Output.
[0087]
 That is, the rendering unit 56 performs rendering processing by VBAP based on the object position information supplied from the core decoding processing unit 21 and the direct sound signal supplied from the amplification unit 51, and the audio object OBJ1 and the audio object OBJ2. Generate an output audio signal for each channel for each of the.
[0088]
 Further, the rendering unit 56 performs the rendering process by VBAP for each pair based on the pair of the signal of the object-specific reverb sound and the position information supplied from the object-specific reverb processing unit 53, and for each channel for each object-specific reverb sound. To produce an output audio signal.
[0089]
 Further, the rendering unit 56 performs the rendering process by VBAP for each pair based on the pair of the space-specific reverb sound signal supplied from the space-specific reverb processing unit 55 and the position information, and for each channel for each space-specific reverb sound. To produce the output audio signal.
[0090]
 Then, the rendering unit 56 adds the signals of the same channel of the output audio signals obtained for the audio object OBJ1, the audio object OBJ2, the object-specific reverb sound, and the space-specific reverb sound, respectively, and outputs the final output audio. Signal.
[0091]

 Here, an example of the format of the input bitstream supplied to the signal processing device 11 will be described.
[0092]
 For example, the format (syntax) of the input bitstream is as shown in FIG. In the example shown in FIG. 3, the character "object_metadata()" is the metadata of the audio object, that is, the audio object information.
[0093]
 The audio object information portion includes object position information about the audio objects by the number of audio objects indicated by the character "num_objects". In this example, horizontal angle position_azimuth[i], vertical angle position_elevation[i], and radius position_radius[i] are stored as the object position information of the i-th audio object.
[0094]
 In addition, the audio object information includes a reverb information flag that indicates whether or not reverb information such as object reverb information and spatial reverb information indicated by the character “flag_obj_reverb” is included.
[0095]
 Here, when the value of the reverb information flag flag_obj_reverb is "1", it indicates that the reverb information is included in the audio object information.
[0096]
 In other words, when the value of the reverb information flag flag_obj_reverb is “1”, it can be said that the reverb information including at least one of the spatial reverb information and the object reverb information is stored in the audio object information.
[0097]
 In more detail, depending on the value of a reuse flag use_prev described later, the audio object information includes identification information for identifying past reverb information as reverb information, that is, a reverb ID described later, and the object reverb information and space Reverb information may not be included.
[0098]
 On the other hand, when the value of the reverb information flag flag_obj_reverb is “0”, it indicates that the reverb information is not included in the audio object information.
[0099]
 When the value of the reverb information flag flag_obj_reverb is “1”, the audio object information includes the direct sound gain indicated by the character “dry_gain[i]” and the object reverb sound indicated by the character “wet_gain[i]” as the reverb information. The gain and the spatial reverb gain indicated by the letters "room_gain[i]" are stored for each audio object.
[0100]
 These direct sound gains dry_gain[i], object reverb sound gains wet_gain[i], and spatial reverb gain room_gain[i] determine the mixing ratio of direct sound, object-specific reverb sound, and spatial-specific reverb sound in the output audio signal. Determined.
[0101]
 Furthermore, in the audio object information, a reuse flag indicated by the character “use_prev” is stored as reverb information.
[0102]
 The reuse flag use_prev is flag information indicating whether to reuse the past object reverb information specified by the reverb ID as the object reverb information of the i-th audio object.
[0103]
 Here, a reverb ID is assigned to each piece of object reverb information transmitted in the input bitstream as identification information for identifying (specifying) the object reverb information.
[0104]
 For example, when the value of the reuse flag use_prev is “1”, it indicates that the past object reverb information is reused. In this case, the audio object information is indicated by the character “reverb_data_id[i]”, A reverb ID indicating object reverb information to be reused is stored.
[0105]
 On the other hand, when the value of the reuse flag use_prev is “0”, it indicates that the object reverb information is not reused. In this case, the audio object information is indicated by the character “obj_reverb_data(i)”. Object reverb information is stored.
[0106]
 The audio object information also stores a spatial reverb information flag indicated by the character "flag_room_reverb" as reverb information.
[0107]
 This spatial reverb information flag flag_room_reverb is a flag indicating the presence or absence of spatial reverb information. For example, when the value of the spatial reverb information flag flag_room_reverb is “1”, it indicates that there is spatial reverb information, and the audio object information stores the spatial reverb information indicated by the characters “room_reverb_data(i)”. ..
[0108]
 On the other hand, when the value of the spatial reverb information flag flag_room_reverb is “0”, it indicates that there is no spatial reverb information, and in this case, the spatial reverb information is not stored in the audio object information. Similar to the case of the object reverb information, the spatial reverb information may be stored with a reuse flag so that the spatial reverb information is appropriately reused.
[0109]
 Further, the format (syntax) of the object reverb information obj_reverb_data(i) and the spatial reverb information room_reverb_data(i) in the audio object information of the input bitstream is as shown in FIG. 4, for example.
[0110]
 In the example shown in FIG. 4, the reverb ID indicated by the character “reverb_data_id” as the object reverb information, the number of object-specific reverb sound components to be generated indicated by the character “num_out”, and the tap length indicated by the character “len_ir”. And are included.
[0111]
 In this example, the coefficient of the impulse response is stored as coefficient information used for generating the object-specific reverb sound component, and the tap length len_ir is the tap length of the impulse response, that is, the number of coefficients of the impulse response. Is shown.
[0112]
 The object reverb information includes object reverb position information of the object-specific reverb sounds corresponding to the number num_out of object-specific reverb sound components to be generated.
[0113]
 That is, the horizontal angle position_azimuth[i], the vertical angle position_elevation[i], and the radius position_radius[i] are stored as the object reverb position information of the i-th object-specific reverb sound component.
[0114]
 Further, as the coefficient information of the i-th object-specific reverb sound component, the impulse response coefficients impulse_response[i][j] are stored by the number of tap lengths len_ir.
[0115]
 On the other hand, the spatial reverb information includes the number of space-specific reverb sound components to be generated, which is indicated by the character “num_out”, and the tap length indicated by the character “len_ir”. The tap length len_ir is the tap length of the impulse response as the coefficient information used to generate the spatially unique reverb sound component.
[0116]
 Also, as the spatial reverb information, the spatial reverb position information of the spatial peculiar reverb sounds is included by the number num_out of the number of spatial peculiar reverb sound components to be generated.
[0117]
 That is, the horizontal angle position_azimuth[i], the vertical angle position_elevation[i], and the radius position_radius[i] are stored as the spatial reverb position information of the i-th spatial peculiar reverb sound component.
[0118]
 Further, as the coefficient information of the i-th space-specific reverberation sound component, the impulse response coefficients impulse_response[i][j] are stored by the number of tap lengths len_ir.
[0119]
 In the examples shown in FIGS. 3 and 4, the example in which the impulse response is used as the coefficient information used to generate the object-specific reverberation sound component and the space-specific reverberation sound component has been described. That is, the example in which the reverb processing using the sampling reverb is performed has been described. However, the present invention is not limited to this, and reverb processing may be performed by using parametric reverb or the like. Further, the coefficient information may be compressed using a lossless coding technique such as Huffman coding.
[0120]
 As described above, in the input bitstream, the information necessary for the reverb processing includes the information about the direct sound (direct sound gain), the information about the object-specific reverb sound such as the object reverb information, and the space-specific reverb such as the spatial reverb information. It is divided into information about sound and transmitted.
[0121]
 Therefore, information can be mixed and output with an appropriate transmission frequency for each information such as the information about the direct sound, the information about the object-specific reverb sound, the information about the space-specific reverb sound, and the like. That is, in each frame of the audio object signal, it is possible to selectively transmit only necessary information of each information such as information about direct sound based on the relationship between the audio object and the viewing position. As a result, the bit rate of the input bit stream can be suppressed and more efficient information transmission can be realized. That is, the coding efficiency can be improved.
[0122]

 Next, the direct sound, the object-specific reverb sound, and the space-specific reverb sound of the audio object reproduced based on the output audio signal will be described.
[0123]
 The relationship between the position of the audio object and the position of the object reverb component is as shown in FIG. 5, for example.
[0124]
 Here, around the position OBJ11 of one audio object, there are object reverb component positions RVB11 to RVB14 of four object-specific reverb sounds for that audio object.
[0125]
 Here, in the drawing, the horizontal angle (azimuth) and the vertical angle (elevation) indicating the object reverb component position RVB11 to the object reverb component position RVB14 are shown on the upper side. In this example, it can be seen that four object-specific reverb sound components are arranged around the origin O, which is the viewing position.
[0126]
 The localization position of the reverb sound peculiar to the object and what kind of sound the reverb sound peculiar to the object is greatly different depending on the position of the audio object in the three-dimensional space. Therefore, it can be said that the object reverb information is reverb information that depends on the spatial position of the audio object.
[0127]
 Therefore, in the input bitstream, the object reverb information is not associated with the audio object, but is managed by the reverb ID.
[0128]
 When the object reverb information is read from the input bitstream, the core decoding processing unit 21 holds the read object reverb information for a certain period. That is, the core decoding processing unit 21 always holds the object reverb information for a predetermined period in the past.
[0129]
 For example, it is assumed that the value of the reuse flag use_prev is “1” at a predetermined time, and reuse of the object reverb information is instructed.
[0130]
 In this case, the core decoding processing unit 21 acquires the reverb ID for the predetermined audio object from the input bitstream. That is, the reverb ID is read.
[0131]
 Then, the core decoding processing unit 21 reads the object reverb information specified by the read reverb ID from the past object reverb information held by the core decoding processing unit 21, and reads the object reverb information for a predetermined audio object at a predetermined time. Reuse as object reverb information.
[0132]
 By managing the object reverb information by the reverb ID in this way, the object reverb information transmitted as for the audio object OBJ1 can be reused as for the audio object OBJ2, for example. Therefore, the number of pieces of object reverb information temporarily stored in the core decoding processing unit 21, that is, the amount of data can be further reduced.
[0133]
 By the way, generally, when an impulse is emitted into a space, for example, as shown in FIG. 6, in addition to direct sound, an initial reflected sound is generated due to reflection of a floor or a wall existing in the surrounding space, and the reflection is also generated. A rear reverberation component generated by the repetition is generated.
[0134]
 Here, the portion indicated by the arrow Q11 indicates the direct sound component, and this direct sound component corresponds to the signal of the direct sound obtained by the amplification unit 51.
[0135]
 The portion indicated by arrow Q12 indicates the initial reflected sound component, and this initial reflected sound component corresponds to the signal of the object-specific reverb sound obtained by the object-specific reverb processing unit 53. Further, a portion indicated by an arrow Q13 indicates a rear reverberation component, and this rear reverberation component corresponds to the signal of the space-specific reverb sound obtained by the space-specific reverb processing unit 55.
[0136]
 The relationship between the direct sound, the early reflection sound, and the rear reverberation component will be described on a two-dimensional plane, for example, as shown in FIGS. 7 and 8. 7 and 8, parts corresponding to each other are designated by the same reference numerals, and the description thereof will be appropriately omitted.
[0137]
 For example, as shown in FIG. 7, it is assumed that there are two audio objects OBJ21 and OBJ22 in an indoor space surrounded by a wall represented by a rectangular frame. Further, it is assumed that the viewer U11 is at the reference viewing position.
[0138]
 Here, it is assumed that the distance from the viewer U11 to the audio object OBJ21 is R OBJ21 , and the distance from the viewer U11 to the audio object OBJ22 is R OBJ22 .
[0139]
 In this case, in the drawing as shown in FIG. 8, drawn in dashed-line arrow, generated by an audio object OBJ21, sound coming towards direct to viewers U11 is direct sound D audio object OBJ21 OBJ21 and Become. Similarly, in the figure, the sound generated by the audio object OBJ22 drawn by the one-dot chain line arrow and coming directly to the viewer U11 is the direct sound D OBJ22 of the audio object OBJ22 .
[0140]
 In the figure, drawn with a dotted line arrow occurs in audio object OBJ21, sound from the once reflected by the interior wall or the like coming towards the viewer U11 initial reflected sound E audio objects OBJ21 OBJ21 and Become. Similarly, the sound generated by the audio object OBJ22, which is drawn by the dotted arrow in the figure, is reflected once by the wall in the room and then comes toward the viewer U11. The initial reflected sound E OBJ22 of the audio object OBJ22. Becomes
[0141]
 Furthermore, the sound S OBJ21 that is generated by the audio object OBJ21 and is repeatedly reflected by the wall in the room and reaches the viewer U11, and the sound S OBJ21 that is generated by the audio object OBJ22 and is repeatedly reflected by the wall in the room. The sound component composed of the sound S OBJ22 reaching the viewer U11 becomes the rear reverberation component. Here, the rear reverberation component is depicted by a solid arrow.
[0142]
 Here, the distance R OBJ22 is shorter than the distance R OBJ21 , and the audio object OBJ22 is closer to the viewer U11 than the audio object OBJ21.
[0143]
 Therefore, for the audio object OBJ22, the direct sound D OBJ22 is dominant as the sound heard by the viewer U11 rather than the initial reflected sound E OBJ22 . Therefore, for the reverb of the audio object OBJ22, the direct sound gain is set to a large value, the object reverb sound gain and the spatial reverb gain are set to small values, and those gains are stored in the input bitstream.
[0144]
 On the other hand, the audio object OBJ21 is farther from the viewer U11 than the audio object OBJ22.
[0145]
 Therefore, for the audio object OBJ21, the early reflection sound E OBJ21 and the sound S OBJ21 of the rear reverberation component are dominant as the sound heard by the viewer U11 rather than the direct sound D OBJ21 . Therefore, for the reverb of the audio object OBJ21, the direct sound gain is set to a small value, the object reverb sound gain and the spatial reverb gain are set to large values, and those gains are stored in the input bitstream.
[0146]
 Further, when the audio objects OBJ21 and OBJ22 move, the initial reflected sound component largely changes depending on the positional relationship between the positions of these audio objects and the wall or floor of the room which is the surrounding space.
[0147]
 Therefore, the object reverb information of the audio object OBJ21 and the audio object OBJ22 needs to be transmitted at the same frequency as the object position information. Such object reverb information is information that largely depends on the position of the audio object.
[0148]
 On the other hand, since the rear reverberation component largely depends on the material of the space such as the wall and floor, the spatial reverb information is transmitted at the lowest necessary frequency, and only the magnitude relation is controlled according to the position of the audio object. As a result, sufficiently subjective quality can be secured.
[0149]
 Therefore, for example, the spatial reverb information is transmitted to the signal processing device 11 less frequently than the object reverb information. In other words, the core decoding processing unit 21 acquires the spatial reverb information at a lower frequency than the acquisition frequency of the object reverb information.
[0150]
 This technology reduces the amount of information (data) required for reverb processing by dividing the information required for reverb processing into sound components such as direct sound, object-specific reverb sound, and space-specific reverb sound. be able to.
[0151]
 Generally, sampling reverb requires a long impulse response data of about 1 second, but by dividing the necessary information for each sound component as in this technology, the impulse response can be a fixed delay and a short impulse response data. Can be realized as a combination of, and the amount of data can be reduced. This can reduce the number of stages of the biquad filter not only in the sampling reverb but also in the parametric reverb.
[0152]
 Moreover, according to the present technology, by dividing the information necessary for the reverb processing for each sound component and transmitting the information, it is possible to transmit the necessary information at a necessary frequency and improve the coding efficiency.
[0153]
 As described above, according to the present technology, when transmitting reverb information for controlling a sense of distance with respect to a panning-based rendering method such as VBAP, high transmission efficiency can be achieved even when there are many audio objects. Can be realized.
[0154]

 Next, a specific operation of the signal processing device 11 will be described. That is, the audio output processing by the signal processing device 11 will be described below with reference to the flowchart of FIG.
[0155]
 In step S11, the core decoding processing unit 21 decodes (data) the received input bitstream.
[0156]
 The core decoding processing unit 21 supplies the audio object signal obtained by the decoding to the amplification unit 51, the amplification unit 52, and the amplification unit 54, and at the same time, the direct sound gain, the object reverb sound gain, and the spatial reverb obtained by the decoding. The gains are supplied to the amplification unit 51, the amplification unit 52, and the amplification unit 54, respectively.
[0157]
 Further, the core decoding processing unit 21 supplies the object reverb information and the spatial reverb information obtained by the decoding to the object-specific reverb processing unit 53 and the space-specific reverb processing unit 55. Further, the core decoding processing unit 21 supplies the object position information obtained by the decoding to the object-specific reverb processing unit 53, the space-specific reverb processing unit 55, and the rendering unit 56.
[0158]
 At this time, the core decoding processing unit 21 temporarily holds the object reverb information read from the input bitstream.
[0159]
 In more detail, when the value of the reuse flag use_prev is “1”, the core decoding processing unit 21 uses the reverb ID read from the input bitstream among the object reverb information held by itself. The specified one is supplied to the object-specific reverb processing unit 53 as the object reverb information of the audio object.
[0160]
 In step S12, the amplification unit 51 multiplies the audio object signal supplied from the core decoding processing unit 21 by the direct sound gain supplied from the core decoding processing unit 21 to perform gain adjustment to thereby obtain a direct sound signal. Is generated and supplied to the rendering unit 56.
[0161]
 In step S13, the object-specific reverb processing unit 53 generates a signal of the object-specific reverb sound.
[0162]
 That is, the amplification unit 52 multiplies the audio object signal supplied from the core decoding processing unit 21 by the object reverb sound gain supplied from the core decoding processing unit 21 to perform gain adjustment, and the object-specific reverb processing unit. Supply to 53.
[0163]
 Further, the object-specific reverb processing unit 53 performs reverb processing on the audio object signal supplied from the amplification unit 52, based on the impulse response coefficient included in the object reverb information supplied from the core decoding processing unit 21. .. That is, the convolution process of the impulse response coefficient and the audio object signal is performed, and the signal of the object-specific reverb sound is generated.
[0164]
 Further, the object-specific reverb processing unit 53 generates and obtains position information of the object-specific reverb sound based on the object position information supplied from the core decoding processing unit 21 and the object reverb position information included in the object reverb information. The position information thus obtained and the signal of the object-specific reverb sound are supplied to the rendering unit 56.
[0165]
 In step S14, the space peculiar reverb processing unit 55 generates a signal of the space peculiar reverb sound.
[0166]
 That is, the amplification unit 54 multiplies the audio object signal supplied from the core decoding processing unit 21 by the spatial reverb gain supplied from the core decoding processing unit 21 to adjust the gain, and the spatial peculiar reverb processing unit 55. Supply to.
[0167]
 Further, the space-specific reverb processing unit 55 performs reverb processing on the audio object signal supplied from the amplification unit 54, based on the impulse response coefficient included in the space reverb information supplied from the core decoding processing unit 21. That is, a convolution process of the impulse response coefficient and the audio object signal is performed, the signals obtained for each audio object by the convolution process are added, and a signal of a space-specific reverb sound is generated.
[0168]
 Further, the space-specific reverb processing unit 55 generates and obtains position information of the space-specific reverb sound based on the object position information supplied from the core decoding processing unit 21 and the space reverb position information included in the space reverb information. The obtained position information and the signal of the space-specific reverb sound are supplied to the rendering unit 56.
[0169]
 In step S15, the rendering unit 56 performs rendering processing and outputs the obtained output audio signal.
[0170]
 That is, the rendering unit 56 performs the rendering process based on the object position information supplied from the core decoding processing unit 21 and the direct sound signal supplied from the amplification unit 51. In addition, the rendering unit 56 performs the rendering process based on the signal of the object-specific reverb sound supplied from the object-specific reverb processing unit 53 and the position information, and the space-specific reverb sound supplied from the space-specific reverb processing unit 55. Rendering processing is performed based on the signal and the position information.
[0171]
 Then, the rendering unit 56 adds the signals obtained by the rendering process of each sound component for each channel to generate a final output audio signal. The rendering unit 56 outputs the output audio signal thus obtained to the subsequent stage, and the audio output process ends.
[0172]
 As described above, the signal processing device 11 performs the reverb processing and the rendering processing based on the audio object information including the information divided into the components of the direct sound, the object-specific reverb sound, and the space-specific reverb sound, and outputs the result. Generate an audio signal. By doing so, the coding efficiency of the input bitstream can be improved.
[0173]

 Next, an encoding device that generates and outputs the input bit stream described above as an output bit stream will be described.
[0174]
 Such an encoding device is configured as shown in FIG. 10, for example.
[0175]
 The coding apparatus 101 shown in FIG. 10 has an object signal coding unit 111, an audio object information coding unit 112, and a packing unit 113.
[0176]
 The object signal encoding unit 111 encodes the supplied audio object signal by a predetermined encoding method and supplies the encoded audio object signal to the packing unit 113.
[0177]
 The audio object information encoding unit 112 encodes the supplied audio object information and supplies it to the packing unit 113.
[0178]
 The packing unit 113 stores the encoded audio object signal supplied from the object signal encoding unit 111 and the encoded audio object information supplied from the audio object information encoding unit 112 in a bitstream. The output bitstream. The packing unit 113 transmits the obtained output bitstream to the signal processing device 11.
[0179]

 Next, the operation of the encoding device 101 will be described. That is, the encoding process by the encoding device 101 will be described below with reference to the flowchart of FIG. 11. For example, this encoding process is performed for each frame of the audio object signal.
[0180]
 In step S41, the object signal encoding unit 111 encodes the supplied audio object signal by a predetermined encoding method and supplies the audio object signal to the packing unit 113.
[0181]
 In step S 42, the audio object information encoding unit 112 encodes the supplied audio object information and supplies the encoded audio object information to the packing unit 113.
[0182]
 Here, for example, audio object information including object reverb information and spatial reverb information is supplied and encoded so that the spatial reverb information is transmitted to the signal processing device 11 at a frequency lower than that of the object reverb information.
[0183]
 In step S43, the packing unit 113 stores the encoded audio object signal supplied from the object signal encoding unit 111 in the bitstream.
[0184]
 In step S44, the packing unit 113 stores the object position information included in the encoded audio object information supplied from the audio object information encoding unit 112 in the bitstream.
[0185]
 In step S45, the packing unit 113 determines whether the encoded audio object information supplied from the audio object information encoding unit 112 has reverb information.
[0186]
 Here, if neither object reverb information nor spatial reverb information is included as reverb information, it is determined that there is no reverb information.
[0187]
 If it is determined in step S45 that there is no reverb information, then the process proceeds to step S46.
[0188]
 In step S46, the packing unit 113 sets the value of the reverb information flag flag_obj_reverb to “0” and stores the reverb information flag flag_obj_reverb in the bitstream. As a result, an output bitstream containing no reverb information is obtained. When the output bitstream is obtained, the process then proceeds to step S54.
[0189]
 On the other hand, if it is determined in step S45 that there is reverb information, then the process proceeds to step S47.
[0190]
 In step S47, the packing unit 113 sets the value of the reverb information flag flag_obj_reverb to “1” and includes the reverb information flag flag_obj_reverb and the encoded audio object information supplied from the audio object information encoding unit 112. The gain information and the gain information are stored in the bitstream. Here, as the gain information, the above-described direct sound gain dry_gain[i], object reverb sound gain wet_gain[i], and spatial reverb gain room_gain[i] are stored in the bitstream.
[0191]
 In step S48, the packing unit 113 determines whether to reuse the object reverb information.
[0192]
 For example, when the encoded audio object information supplied from the audio object information encoding unit 112 does not include the object reverb information but does include the reverb ID, it is determined to be reused.
[0193]
 If it is determined in step S48 that the reuse is to be performed, then the process proceeds to step S49.
[0194]
 In step S49, the packing unit 113 sets the value of the reuse flag use_prev to “1” and includes the reuse flag use_prev and the encoded audio object information supplied from the audio object information encoding unit 112. The reverb ID and the reverb ID are stored in the bitstream. When the reverb ID is stored, the process then proceeds to step S51.
[0195]
 On the other hand, if it is determined in step S48 that the reuse is not performed, then the process proceeds to step S50.
[0196]
 In step S50, the packing unit 113 sets the value of the reuse flag use_prev to “0” and includes the reuse flag use_prev and the encoded audio object information supplied from the audio object information encoding unit 112. The object reverb information and the object reverb information are stored in the bitstream. When the object reverb information is stored, the process then proceeds to step S51.
[0197]
 When the process of step S49 or step S50 is performed, the process of step S51 is performed thereafter.
[0198]
 That is, in step S51, the packing unit 113 determines whether the encoded audio object information supplied from the audio object information encoding unit 112 has spatial reverb information.
[0199]
 If it is determined in step S51 that there is spatial reverb information, then the process proceeds to step S52.
[0200]
 In step S52, the packing unit 113 sets the value of the spatial reverb information flag flag_room_reverb to “1”, and sets the spatial reverb information flag flag_room_reverb and the encoded audio object information supplied from the audio object information encoding unit 112. Store the included spatial reverb information in the bitstream.
[0201]
 As a result, the output bitstream including the spatial reverb information is obtained. When the output bitstream is obtained, the process then proceeds to step S54.
[0202]
 On the other hand, if it is determined in step S51 that there is no spatial reverb information, then the process proceeds to step S53.
[0203]
 In step S53, the packing unit 113 sets the value of the spatial reverb information flag flag_room_reverb to “0” and stores the spatial reverb information flag flag_room_reverb in the bitstream. As a result, an output bitstream that does not include spatial reverb information is obtained. When the output bitstream is obtained, the process then proceeds to step S54.
[0204]
 When the process of step S46, step S52, or step S53 is performed to obtain the output bitstream, the process of step S54 is performed thereafter. The output bitstream obtained by these processes is, for example, a bitstream in the format shown in FIGS. 3 and 4.
[0205]
 In step S54, the packing unit 113 outputs the obtained output bitstream, and the encoding process ends.
[0206]
 As described above, the encoding apparatus 101 stores the audio object information that appropriately includes information divided for each component of the direct sound, the object-specific reverb sound, and the space-specific reverb sound, and outputs the bit stream. By doing so, the coding efficiency of the output bit stream can be improved.
[0207]
 Note that, in the above, an example in which gain information such as direct sound gain, object reverb sound gain, and spatial reverb gain is given as audio object information has been described, but even if these gain information is generated on the decoding side. Good.
[0208]
 In such a case, for example, the signal processing device 11 generates the direct sound gain, the object reverb sound gain, and the spatial reverb gain based on the object position information, the object reverb position information, the spatial reverb position information, etc. included in the audio object information. To do.
[0209]

 By the way, the series of processes described above can be executed by hardware or software. When the series of processes is executed by software, a program forming the software is installed in the computer. Here, the computer includes a computer incorporated in dedicated hardware and, for example, a general-purpose personal computer capable of executing various functions by installing various programs.
[0210]
 FIG. 12 is a block diagram showing a configuration example of hardware of a computer that executes the series of processes described above by a program.
[0211]
 In a computer, a CPU (Central Processing Unit) 501, a ROM (Read Only Memory) 502, and a RAM (Random Access Memory) 503 are connected to each other by a bus 504.
[0212]
 An input/output interface 505 is further connected to the bus 504. An input unit 506, an output unit 507, a recording unit 508, a communication unit 509, and a drive 510 are connected to the input/output interface 505.
[0213]
 The input unit 506 includes a keyboard, a mouse, a microphone, an image sensor, and the like. The output unit 507 includes a display, a speaker and the like. The recording unit 508 includes a hard disk, a non-volatile memory, or the like. The communication unit 509 includes a network interface or the like. The drive 510 drives a removable recording medium 511 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.
[0214]
 In the computer configured as described above, the CPU 501 loads the program recorded in the recording unit 508 into the RAM 503 via the input/output interface 505 and the bus 504 and executes the program, thereby performing the above-described series of operations. Is processed.
[0215]
 The program executed by the computer (CPU 501) can be provided, for example, by recording it on a removable recording medium 511 such as a package medium. Further, the program can be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting.
[0216]
 In the computer, the program can be installed in the recording unit 508 via the input/output interface 505 by mounting the removable recording medium 511 in the drive 510. Further, the program can be received by the communication unit 509 via a wired or wireless transmission medium and installed in the recording unit 508. In addition, the program can be installed in the ROM 502 or the recording unit 508 in advance.
[0217]
 The program executed by the computer may be a program that is processed in time series in the order described in this specification, or in parallel or at a necessary timing such as when a call is made. It may be a program in which processing is performed.
[0218]
 Further, the embodiments of the present technology are not limited to the above-described embodiments, and various modifications can be made without departing from the scope of the present technology.
[0219]
 For example, the present technology may have a configuration of cloud computing in which a plurality of devices share one function via a network and jointly process the functions.
[0220]
 In addition, each step described in the above-described flowcharts can be executed by one device or shared by a plurality of devices.
[0221]
 Further, when one step includes a plurality of processes, the plurality of processes included in the one step can be executed by one device or shared by a plurality of devices.
[0222]
 Furthermore, the present technology may be configured as below.
[0223]
(1)
 Reverb information including at least one of spatial reverb information unique to a space around an audio object and object reverb information unique to the audio object, and an acquisition unit for acquiring an audio object signal of the audio object. And
 a reverb processing unit that generates a signal of a reverb component of the audio object based on the reverb information and the audio object signal
 .
(2)  The signal processing device
 according to
(1), wherein the spatial reverb information is acquired at a frequency lower than that of the object reverb information .
(3)
 When the identification information indicating the past reverb information is acquired by the acquisition unit, the reverb processing unit, based on the reverb information indicated by the identification information and the audio object signal, the reverb component. The signal
 processing device according to (1) or (2), which generates the signal of 1.
(4) The
 identification information is information indicating the object reverb information, and the
 reverb processing unit reverb based on the object reverb information, the spatial reverb information, and the audio object signal indicated by the identification information.
 The signal processing device according to (3), which generates a component signal .
(5)

 The signal processing device according to any one of (1) to (4),  wherein the object reverb information is information that depends on a position of the audio object .
(6) The
 reverb processing unit
  generates a signal of the reverb component unique to the space based on the spatial reverb information and the audio object signal, and the audio object based on the
  object reverb information and the audio object signal.
 The signal processing device according to any one of (1) to (5), which generates a signal of the reverb component unique to the .
(7) The
 signal processing device
 includes reverb information including at least one of spatial reverb information unique to a space around an audio object and object reverb information unique to the audio object, and an audio object signal of the audio object. And
 a
 signal processing method for generating a signal of a reverb component of the audio object based on the reverb information and the audio object signal .
(8)
 Obtaining reverb information including at least one of spatial reverb information unique to a space around an audio object and object reverb information unique to the audio object, and an audio object signal of the audio object,

 A program that causes a computer to execute a process including  a step of generating a signal of a reverb component of the audio object based on the reverb information and the audio object signal .
Explanation of symbols
[0224]
 11 signal processing device, 21 core decoding processing unit, 22 rendering processing unit, 51-1, 51-2, 51 amplification unit, 52-1, 52-2, 52 amplification unit, 53-1, 53-2, 53 objects Unique reverb processing unit, 54-1, 54-2, 54 amplifying unit, 55 Spatial unique reverb processing unit, 56 rendering unit, 101 coding device, 111 object signal coding unit, 112 audio object information coding unit, 113 packing Department
The scope of the claims
[Claim 1]
 A unique spatial reverberation information in the space around the audio objects, and the reverberation information including at least one of the specific objects reverb information audio object, and acquisition unit that acquires an audio object signals of the audio objects,
 wherein A
 signal processing device comprising: a reverb processing unit that generates a signal of a reverb component of the audio object based on reverb information and the audio object signal .
[Claim 2]

 The signal processing device  according to claim 1, wherein the spatial reverb information is acquired at a frequency lower than that of the object reverb information .
[Claim 3]
 The reverb processing unit, when identification information indicating the past reverb information is acquired by the acquisition unit, the reverb information indicated by the identification information, and a signal of the reverb component based on the audio object signal.
 The signal processing device according to claim 1, which generates the signal.
[Claim 4]
 The identification information is information indicating the object reverb information, the
 reverb processing unit, the reverb component signal based on the object reverb information, the spatial reverb information, and the audio object signal indicated by the identification information.
 The signal processing device according to claim 3, which generates
[Claim 5]

 The signal processing device according to claim 1,  wherein the object reverb information is information depending on a position of the audio object .
[Claim 6]
 The reverb processing unit
  generates a signal of the reverb component unique to the space based on the spatial
  reverb information and the audio object signal, and unique to the audio object based on the object reverb information and the audio object signal.
 The signal processing device according to claim 1, wherein the signal of the reverb component is generated .
[Claim 7]
 A signal processing device
 obtains reverb information including at least one of spatial reverb information unique to a space around an audio object and object reverb information unique to the audio object, and an audio object signal of the audio object. ,
 based on the reverberation information and the audio object signal, and generates a signal of a reverb of the audio object
 signal processing method.
[Claim 8]
 Spatial reverb information peculiar to the space around the audio object and reverb information including at least one of object reverb information peculiar to the audio object, and an audio object signal of the audio object are acquired, and the
 reverb information and
 A program that causes a computer to execute processing including a step of generating a signal of a reverb component of the audio object based on the audio object signal .

Documents

Application Documents

# Name Date
1 202017015894-TRANSLATIOIN OF PRIOIRTY DOCUMENTS ETC. [13-04-2020(online)].pdf 2020-04-13
2 202017015894-STATEMENT OF UNDERTAKING (FORM 3) [13-04-2020(online)].pdf 2020-04-13
3 202017015894-PROOF OF RIGHT [13-04-2020(online)].pdf 2020-04-13
4 202017015894-PRIORITY DOCUMENTS [13-04-2020(online)].pdf 2020-04-13
5 202017015894-POWER OF AUTHORITY [13-04-2020(online)].pdf 2020-04-13
6 202017015894-FORM 1 [13-04-2020(online)].pdf 2020-04-13
7 202017015894-DRAWINGS [13-04-2020(online)].pdf 2020-04-13
8 202017015894-DECLARATION OF INVENTORSHIP (FORM 5) [13-04-2020(online)].pdf 2020-04-13
9 202017015894-COMPLETE SPECIFICATION [13-04-2020(online)].pdf 2020-04-13
10 202017015894-Proof of Right [07-07-2020(online)].pdf 2020-07-07
11 202017015894-FORM 18 [01-09-2021(online)].pdf 2021-09-01
12 abstract.jpg 2021-10-19
13 202017015894.pdf 2021-10-19
14 202017015894-FER.pdf 2022-04-20
15 202017015894-OTHERS [20-10-2022(online)].pdf 2022-10-20
16 202017015894-FER_SER_REPLY [20-10-2022(online)].pdf 2022-10-20
17 202017015894-DRAWING [20-10-2022(online)].pdf 2022-10-20
18 202017015894-CORRESPONDENCE [20-10-2022(online)].pdf 2022-10-20
19 202017015894-CLAIMS [20-10-2022(online)].pdf 2022-10-20
20 202017015894-PatentCertificate05-01-2024.pdf 2024-01-05
21 202017015894-IntimationOfGrant05-01-2024.pdf 2024-01-05

Search Strategy

1 202017015894E_20-04-2022.pdf

ERegister / Renewals

3rd: 27 Mar 2024

From 05/10/2020 - To 05/10/2021

4th: 27 Mar 2024

From 05/10/2021 - To 05/10/2022

5th: 27 Mar 2024

From 05/10/2022 - To 05/10/2023

6th: 27 Mar 2024

From 05/10/2023 - To 05/10/2024

7th: 30 Sep 2024

From 05/10/2024 - To 05/10/2025

8th: 25 Sep 2025

From 05/10/2025 - To 05/10/2026