Sign In to Follow Application
View All Documents & Correspondence

Encoding Device And Method, Decoding Device And Method

Abstract: The present technique pertains to an encoding device and method, a decoding device and method, and a program, which are configured so as to make it possible to improve convenience. The decoding device comprises: a decoding unit that decodes audio data, including object audio, included in an encoded bit stream, and reads the metadata of the object audio from regions in which desired data of the encoded bit stream can be stored; and an output unit that outputs the decoded audio data on the basis of the metadata. The present technique can be applied to a decoding device.

Get Free WhatsApp Updates!
Notices, Deadlines & Correspondence

Patent Information

Application #
Filing Date
27 March 2020
Publication Number
33/2020
Publication Type
INA
Invention Field
ELECTRONICS
Status
Email
mahua.ray@remfry.com
Parent Application
Patent Number
Legal Status
Grant Date
2024-01-03
Renewal Date

Applicants

SONY CORPORATION
1-7-1, Konan, Minato-ku, Tokyo 1080075

Inventors

1. HATANAKA Mitsuyuki
c/o SONY CORPORATION, 1-7-1, Konan, Minato-ku, Tokyo 1080075
2. CHINEN Toru
c/o SONY CORPORATION, 1-7-1, Konan, Minato-ku, Tokyo 1080075

Specification

Specification
Title of invention: Encoding device and method, decoding device and method, and program
Technical field
[0001]
 The present technology relates to an encoding device and method, a decoding device and method, and a program, and particularly relates to an encoding device and method, a decoding device and method, and a program that can improve convenience.
Background technology
[0002]
 Conventionally, audio encoding techniques such as AAC (Advanced Audio Coding) and MP3 (Moving Picture Experts Group 2 Audio Layer 3) are known (see, for example, Non-Patent Document 1 and Non-Patent Document 2).
[0003]
 These audio encoding technologies such as AAC and MP3 are used in various applications including music distribution, and have a track record in sound quality and mounting.
Prior art documents
Non-patent literature
[0004]
Non-Patent Document 1: ISO/IEC 14496-3:2009 Information technology
--Coding of audio-visual objects --Part 3: Audio Non-Patent Document 2: ISO/IEC 11172-3:1993 Information technology --Coding of moving pictures and associated audio for digital storage media at up to about 1.5 Mbit/s-Part 3: Audio
Summary of the invention
Problems to be Solved by the Invention
[0005]
 However, these audio encoding techniques, that is, the conventional audio encoding standards are encoding techniques on the assumption that channel-based audio data is handled.
[0006]
 Therefore, it was not possible to directly encode the object audio data used for reproducing the three-dimensional audio (3D Audio) and the 3D Audio metadata including the position information of each object audio data.
[0007]
 Therefore, in order to use the above audio coding standard in 3D Audio, it is desired to extend the standard while maintaining compatibility with the conventional standard.
[0008]
 The present technology has been made in view of such circumstances, and is to improve convenience.
Means for solving the problems
[0009]
 A decoding device according to a first aspect of the present technology decodes audio data including object audio included in a coded bitstream, and outputs the object from an area in which any data of the coded bitstream can be stored. A decoding unit that reads audio metadata and an output unit that outputs the decoded audio data based on the metadata are provided.
[0010]
 The decoding method or program according to the first aspect of the present technology decodes audio data including object audio included in a coded bitstream, and from an area in which arbitrary data of the coded bitstream can be stored. Reading metadata of the object audio and outputting the decoded audio data based on the metadata.
[0011]
 In the first aspect of the present technology, audio data including object audio included in a coded bitstream is decoded, and the object audio is recorded from an area in which any data of the coded bitstream can be stored. Is read out, and the decoded audio data is output based on the metadata.
[0012]
 The encoding device according to the second aspect of the present technology encodes audio data including object audio, stores the metadata of the object audio in an area capable of storing arbitrary data, and stores the metadata, and An encoding unit is provided that generates an encoded bitstream including the encoded audio data.
[0013]
 The encoding method or program according to the second aspect of the present technology encodes audio data including object audio, stores metadata of the object audio in an area capable of storing arbitrary data, and stores the metadata. , And generating an encoded bitstream containing the encoded audio data.
[0014]
 In the second aspect of the present technology, audio data including object audio is encoded, and metadata of the object audio is stored in an area capable of storing arbitrary data. An encoded bitstream including the generated audio data is generated.
Effect of the invention
[0015]
 According to the first aspect and the second aspect of the present technology, convenience can be improved.
[0016]
 Note that the effects described here are not necessarily limited, and may be any effects described in the present disclosure.
Brief description of the drawings
[0017]
FIG. 1 is a diagram explaining the present technology.
FIG. 2 is a diagram illustrating an AAC bitstream.
FIG. 3 is a diagram showing a configuration example of an encoding device.
FIG. 4 is a diagram showing a configuration example of a decoding device.
FIG. 5 is a diagram showing a configuration example of a 3D AAC bitstream.
FIG. 6 is a diagram illustrating application of 3D Audio metadata.
FIG. 7 is a diagram illustrating the use of externally acquired 3D Audio metadata.
FIG. 8 is a diagram illustrating position information of an audio object.
FIG. 9 is a diagram illustrating position information of an audio object.
FIG. 10 is a flowchart illustrating an encoding process.
FIG. 11 is a flowchart illustrating a decoding process.
FIG. 12 is a diagram illustrating divisional transmission of 3D Audio metadata.
FIG. 13 is a diagram illustrating divisional transmission of 3D Audio metadata.
FIG. 14 is a diagram illustrating divisional transmission of 3D Audio metadata.
FIG. 15 is a diagram illustrating intra-frame DSE division information.
FIG. 16 is a diagram illustrating frame division information.
FIG. 17 is a diagram illustrating division information.
FIG. 18 is a diagram illustrating division of 3D Audio metadata in a plurality of frames.
FIG. 19 is a flowchart illustrating an encoding process.
FIG. 20 is a flowchart illustrating a decoding process.
FIG. 21 is a diagram showing a configuration example of a computer.
MODE FOR CARRYING OUT THE INVENTION
[0018]
 Hereinafter, embodiments to which the present technology is applied will be described with reference to the drawings.
[0019]

The
 present technology is based on 3D Audio, that is, even if an encoding method such as AAC or MP3 that is not supposed to handle object audio data is used. By encoding and transmitting 3D Audio metadata and 3D Audio metadata so that 3D Audio playback can be realized on the decoding side, convenience is improved.
[0020]
 The present technology has the following features in particular.
[0021]
 Features (1)
 storing the 3D Audio metadata required to render a 3D Audio in ancillary data region in codecs AAC etc. that do not support 3D Audio
 feature (2)
 3D Audio metadata is one ancillary data area If the 3D Audio metadata is stored in multiple ancillary data areas if it does not fit in
 (3) If
 3D Audio metadata is present intermittently, it has already been acquired for the frame to be played. The
 feature of applying the last obtained 3D Audio metadata (4) By
 describing the information about the structure of the 3D Audio audio data in the comment field of the codec layer, the conventional codec (encoding) standard transmitting the audio data while maintaining the
 feature (5)
 when the 3D audio metadata does not fit in ancillary data region of one frame, it transmits the 3D audio metadata over a plurality of frames. At that time, the encoding device performs thinning processing of 3D Audio metadata that should be stored in each original frame, and transmits representative values ​​of 3D Audio metadata of multiple frames
 (6)
 3D Audio channel mapping Assign to existing channel mapping as much as possible, store the information indicating the channel mapping in the bitstream, and transmit.
[0022]
 With the present technology having the above characteristics, it is possible to realize the reproduction of 3D Audio by using the existing encoding technology such as AAC or MP3 while maintaining the compatibility of the encoding technology.
[0023]
 Specifically, for example, in 3D Audio, a sound image of a sound source such as an audio object can be localized at an arbitrary position in a three-dimensional space.
[0024]
 In such 3D Audio, 3D Audio metadata is required in addition to audio data of audio objects, that is, audio elements such as object-based audio data and channel-based audio data.
[0025]
 Here, the 3D audio metadata is data including position information indicating the position of the audio object in the three-dimensional space, gain information used for gain adjustment of audio data of the audio object, and the like.
[0026]
 However, in the encoding method such as AAC and MP3, since it is not originally assumed that such 3D Audio metadata is handled, 3D Audio metadata cannot be encoded and transmitted as it is.
[0027]
 Therefore, in the present technology, as shown by an arrow Q11 in FIG. 1, encoded 3D Audio metadata is stored in the extension area provided in the encoded bitstream in which audio data is stored.
[0028]
 As a result, it becomes possible to transmit 3D Audio metadata and reproduce 3D Audio even with an encoded bitstream of an encoding method that does not support 3D Audio.
[0029]
 Here, the extended area where 3D Audio metadata is stored is an area defined by existing encoding technology that allows users to record free data according to their own rules, that is, stores arbitrary data. This is a possible area and is also called an ancillary data area.
[0030]
 Specifically, for example, when the present technique is applied to the AAC encoding technique, the AAC bitstream that is the encoded bitstream as shown by an arrow Q12 has an area corresponding to the extension area of ​​the example shown by an arrow Q11. , DSE (Data Stream Element). Therefore, when audio data is encoded by the AAC encoding method, the encoded 3D Audio metadata is stored in the DSE as ancillary data.
[0031]
 For example, if the decoding device that has received the AAC bitstream indicated by arrow Q12 is 3D Audio, that is, 3D Audio metadata, that decoding device reads 3D Audio metadata from the AAC bitstream and decodes It can be used to play 3D Audio.
[0032]
 On the other hand, a decoding device that does not support 3D Audio metadata cannot interpret 3D Audio metadata as ancillary data described in DSE.
[0033]
 However, the decoding device can obtain data size information defined by the AAC standard and indicating the data size of the ancillary data recorded in the DSE. Therefore, the decoding device can skip the portion of the 3D Audio metadata in the DSE by using the data size information, and can maintain compatibility with the AAC standard.
[0034]
 When the present technology is applied to MP3, the MP3 bitstream that is a coded bitstream as shown by an arrow Q13 has an ancillary data area as an area corresponding to the extension area of ​​the example shown by an arrow Q11. Is provided. Therefore, when audio data is encoded by the MP3 encoding method, the encoded 3D Audio metadata is stored in the ancillary data area as ancillary data.
[0035]
 Now, the present technology will be described in more detail below. In the following, in order to give a specific description, an example in which the present technology is applied to the AAC standard, that is, an example in which 3D Audio reproduction is performed using the AAC encoding technology while maintaining compatibility with the AAC standard will be described.
[0036]
 When the present technology is applied to the AAC standard, an AAC bitstream, that is, an encoded bitstream obtained by encoding audio data according to AAC has the configuration shown in FIG.
[0037]
 That is, a PCE (Program Config Element) and a DSE are provided in the AAC bitstream, and AAC audio element data composed of audio data used for reproduction is stored.
[0038]
 For example, AAC audio element data includes at least one of audio object audio data, channel-based audio data, and HOA (High Order Ambisonic)-based audio data as audio data for realizing 3D Audio playback. Has been.
[0039]
 Here, the audio data of the audio object, that is, the object-based audio data is so-called object audio, which is an audio signal for reproducing the sound of the audio object. The channel-based audio data is an audio signal of each channel having a predetermined channel configuration such as 5.1 channel or 22.2 channel, that is, an audio signal of sound reproduced by a speaker corresponding to each channel.
[0040]
 Further, the HOA-based audio data is an ambisonic format audio signal. That is, the HOA-based audio data is a ambisonic format representation of a channel-based audio signal. Note that, hereinafter, the HOA-based audio data will also be referred to as ambisonic-based audio data.
[0041]
 In the following, audio data of an audio object, channel-based audio data, and ambisonic-based audio data will be simply referred to as audio data unless it is necessary to distinguish them.
[0042]
 In the comment field that can store arbitrary data in the PCE of the AAC bitstream, 3D Audio constituent element information indicating the constituent elements of the AAC audio element data is stored. Here, the 3D Audio component information is the number of channel bases that is the number of channel-based audio data that configures the AAC audio element data, the number of object audios that is the number of audio data of audio objects, and the ambisonic-based audio data. It is information indicating the number of HOA audios, which is the number of.
[0043]
 Furthermore, 3D Audio metadata, which is metadata of audio data forming AAC audio element data, is stored as ancillary data in the DSE of the AAC bit stream.
[0044]
 Here, 3D audio metadata includes channel-based metadata, object-based metadata, and ambisonic-based metadata.
[0045]
 For example, the channel-based metadata is metadata of channel-based audio data, and includes corresponding channel information indicating which channel each channel-based audio data belongs to.
[0046]
 The object-based metadata is metadata of audio data of an audio object, and includes object number information indicating the number of audio objects, position information indicating the position of each audio object in a three-dimensional space, and audio of each audio object. It is composed of gain information indicating the gain by which the data is multiplied.
[0047]
 Further, the ambisonic-based metadata is metadata of ambisonic-based audio data, and includes information indicating the ambisonic order and mode.
[0048]
 In the following, an AAC bitstream obtained by applying the present technology, that is, an encoded bitstream based on AAC and in which 3D Audio metadata is stored will be particularly referred to as a 3D AAC bitstream.
[0049]

 Subsequently, an encoding device that generates the 3D AAC bitstream described above and a decoding device that receives and decodes the 3D AAC bitstream output from the encoding device will be described. ..
[0050]
 FIG. 3 is a diagram illustrating a configuration example of an encoding device to which the present technology is applied.
[0051]
 The encoding device 11 shown in FIG. 3 has a 3D Audio metadata encoding unit 21 and an AAC encoding unit 22.
[0052]
 In the encoding device 11, channel-based metadata, object-based metadata, and ambisonic-based metadata are externally supplied to the 3D Audio metadata encoding unit 21 as 3D Audio metadata.
[0053]
 Also, externally supplied to the AAC encoding unit 22 is audio data of an audio object (object audio), various audio data such as channel-based audio data, ambisonic-based audio data, and 3D Audio component information. To be done. Here, each audio data is a PCM (Pulse Code Modulation) signal for reproducing the content, and the AAC encoding unit 22 is supplied with one or more audio data including at least object audio from the outside. Shall be.
[0054]
 The 3D Audio metadata encoding unit 21 encodes 3D Audio metadata supplied from the outside, and supplies the encoded 3D Audio metadata to the AAC encoding unit 22. Hereinafter, the 3D Audio metadata coded by the 3D Audio metadata coding unit 21 will also be referred to as coded 3D Audio metadata.
[0055]
 The AAC encoding unit 22 encodes the encoded 3D Audio metadata supplied from the 3D Audio metadata encoding unit 21 and the externally supplied 3D Audio component information and audio data by the AAC encoding method. Generate and output a 3D AAC bitstream.
[0056]
 Here, since the audio data supplied to the AAC encoder 22 is a PCM signal, it is possible to encode these audio data as AAC audio element data of the AAC standard.
[0057]
 Note that the 3D Audio component information is not externally supplied to the AAC encoding unit 22, but the AAC encoding unit 22 receives the necessary information from the 3D Audio metadata encoding unit 21 and outputs the information. You may make it generate|occur|produce 3D Audio component element information based on this.
[0058]

 Next, a configuration of a decoding device that receives and decodes the 3D AAC bitstream output from the encoding device 11 illustrated in FIG. 3 will be described. Such a decoding device is configured, for example, as shown in FIG.
[0059]
 The decoding device 51 illustrated in FIG. 4 includes an AAC decoding unit 61, a 3D Audio metadata decoding unit 62, a 3D Audio metadata acquisition unit 63, and a 3D Audio renderer 64.
[0060]
 The AAC decoding unit 61 receives the 3D AAC bitstream transmitted from the encoding device 11 and decodes the received 3D AAC bitstream.
[0061]
 The AAC decoding unit 61 supplies the encoded 3D Audio metadata obtained by decoding the 3D AAC bitstream to the 3D Audio metadata decoding unit 62, and at the same time, outputs the audio data obtained by the decoding, that is, the PCM signal, to the 3D Audio renderer. Supply to 64.
[0062]
 Here, the audio data supplied from the AAC decoding unit 61 to the 3D Audio renderer 64 is audio data of audio objects (object audio data), channel-based audio data, ambisonic-based audio data, and the like.
[0063]
 Further, the AAC decoding unit 61 appropriately supplies the 3D Audio component element information obtained by the decoding to a higher-level control unit that controls the decoding device 51.
[0064]
 The 3D Audio metadata decoding unit 62 decodes the encoded 3D Audio metadata supplied from the AAC decoding unit 61, and obtains channel-based metadata, object-based metadata, and ambience as 3D Audio metadata obtained as a result. Supplies sonic-based metadata to the 3D Audio renderer 64.
[0065]
 The 3D Audio metadata acquisition unit 63 appropriately acquires necessary information from the outside to generate 3D Audio metadata, or receives 3D Audio metadata from an external server or the like to receive 3D Audio metadata. It is acquired and supplied to the 3D Audio renderer 64.
[0066]
 The 3D Audio renderer 64 performs rendering processing to generate playback data for 3D Audio playback, and outputs the obtained playback data, that is, the output unit that outputs the decoded audio data based on the 3D Audio metadata. Functions as an output unit for outputting.
[0067]
 Specifically, the 3D Audio renderer 64 is based on the 3D Audio metadata supplied from the 3D Audio metadata decoding unit 62 or the 3D Audio metadata acquisition unit 63 and the audio data supplied from the AAC decoding unit 61. Rendering is performed to generate audio data of a predetermined number of channels as reproduction data.
[0068]
 The reproduction data is an audio signal for realizing 3D Audio reproduction, which includes audio data of each channel having a predetermined channel configuration such as 2 channels or 5.1 channels. If the sound is reproduced based on this reproduction data, for example, the sound image of the sound of the audio object can be localized at the position in the three-dimensional space indicated by the position information included in the 3D Audio metadata of the audio object. it can.
[0069]
 The 3D Audio renderer 64 outputs the obtained reproduction data to the outside. For example, the 3D Audio renderer 64 supplies reproduction data to a speaker to reproduce sound, or supplies reproduction data to a recording unit (not shown) for recording.
[0070]

 Here, a specific structural example of the 3D AAC bitstream exchanged between the encoding device 11 and the decoding device 51 as described above will be described.
[0071]
 For example, assume that the AAC audio element data is composed of channel-based audio data of 7.1 channels, audio data of two audio objects, and primary ambisonic audio data of ambisonic base.
[0072]
 In such a case, the 3D AAC bitstream is configured, for example, as shown in FIG. In this example, PCE and DSE are arranged at the head of the 3D AAC bit stream for one frame of audio data. The PCE stores 3D Audio component information, and the DSE stores encoded 3D Audio metadata.
[0073]
 Further, following the DSE, one SCE (Single Channel Element), three CPEs (Channel Pair Element), and one LFE (Low Frequency Effects) are arranged. Then, 7.1-channel channel-based audio data is stored in these SCE, CPE, and LFE to form a channel-based audio element.
[0074]
 Further, two SCEs are arranged behind the channel-based audio element, and audio data of two audio objects are stored in these SCEs to form an object-based audio element.
[0075]
 Four SCEs are arranged subsequent to the object-based audio element, and the ambisonic-based primary ambisonic audio data is stored in these SCEs to form an ambisonic-based audio element.
[0076]
 The FIL is placed after the ambisonic bass audio element, that is, at the end of the 3D AAC bit stream for one frame.
[0077]
 Thus, in the example shown in FIG. 5, the AAC audio element data is composed of a channel-based audio element, an object-based audio element, and an ambisonic-based audio element.
[0078]
 The AAC audio element data may be composed of at least one of these channel-based audio element, object-based audio element, and ambisonic-based audio element.
[0079]

 By the way, 3D Audio basically transmits 3D Audio metadata for each frame. However, when 3D Audio metadata is not necessarily required for each of a plurality of frames, such as a frame section in which an audio object does not move, the 3D Audio metadata may not be transmitted for each frame.
[0080]
 In such a case, for example, in the frame in which the 3D Audio metadata is not transmitted as shown in FIG. 6, among the 3D Audio metadata of the frames before the frame, the 3D Audio metadata acquired last in time. The data is used as that of the current frame.
[0081]
 In the example shown in FIG. 6, each of the rectangles indicated by the arrows A11 to A15 represents the respective 3D AAC bitstreams of the frame N to the frame (N+4). In the following, the 3D AAC bit stream for one frame will also be referred to as AAC frame data.
[0082]
 For example, in the frame N, the 3D Audio metadata of the frame N is stored in the AAC frame data. Therefore, in the decoding device 51, the 3D Audio metadata of the frame N obtained by decoding (hereinafter, 3D Audio metadata ( N) will be used for rendering.
[0083]
 On the other hand, in the next frame (N+1), 3D Audio metadata is not stored in the AAC frame data. Therefore, the decoding device 51 uses the 3D Audio metadata (N) of the frame N that is acquired last, that is, that is obtained by the decoding at the end, and renders the frame (N+1).
[0084]
 As described above, when there is no 3D Audio metadata of the current frame, the most current 3D Audio metadata of the frames temporally preceding the current frame that has already been read from the DSE of the 3D AAC bitstream is used. The 3D Audio metadata of the frame temporally close to the frame is used as the 3D Audio metadata of the current frame.
[0085]
 In other words, the 3D Audio metadata read last in time among the 3D Audio metadata read from the DSE in the frame before the current frame is used as the 3D Audio metadata of the current frame. And the current frame is rendered. Furthermore, as a result, when there is no 3D Audio metadata of the current frame, the 3D Audio metadata used in the frame immediately before the current frame is also used as the current frame.
[0086]
 As with frame (N+1), subsequent frames (N+2) and (N+3) did not store 3D Audio metadata in those frames, so it was obtained last. Rendering is performed using 3D Audio metadata (N) of frame N.
[0087]
 On the other hand, in the frame (N+4) following the frame (N+3), the 3D Audio metadata (N+4) of the frame (N+4) is obtained by decoding, so the 3D Audio metadata (N+4) is used for rendering.
[0088]
 When it is determined that nothing is stored in the DSE in the current frame and 3D Audio metadata is not stored, the 3D Audio metadata of the frame temporally preceding the current frame is referred to. Can be
[0089]
 In this case, it can be said that the fact that the 3D Audio metadata is not stored in the DSE of the current frame indicates the instruction to refer to the 3D Audio metadata of the frame temporally preceding the current frame.
[0090]
 In addition to the 3D Audio metadata, the DSE may also store a previous frame reference flag indicating whether to use (reference) the 3D Audio metadata of a frame temporally preceding the current frame. ..
[0091]
 That is, the previous frame reference flag uses the 3D Audio metadata read last in time among the 3D Audio metadata read from the DSE in the frame before the current frame as the 3D Audio of the current frame. This is information that indicates whether to use it as metadata.
[0092]
 In other words, the previous frame reference flag is the 3D audio metadata of the current frame, and the 3D audio metadata stored in the DSE of the frame temporally previous to the current frame is the 3D of the frame closest to the current frame. This is information indicating whether to use Audio metadata.
[0093]
 In this case, for example, when referring to the 3D Audio metadata of the previous frame, the value of the previous frame reference flag is set to "1", and only the previous frame reference flag whose value is 1 is described in DSE.
[0094]
 On the other hand, when the 3D Audio metadata of the previous frame is not referred to, the value of the previous frame reference flag is set to “0”, and the DSE is followed by the previous frame reference flag whose value is 0, followed by the 3D Audio meta data. The data is described. Hereinafter, the description will be continued assuming that the previous frame reference flag is stored in the DSE.
[0095]
 The 3D Audio metadata is data stored in the DSE and is independent of the AAC coding standard.
[0096]
 Therefore, when the 3D AAC bitstream is played from the middle, that is, when the content is played from the middle, even if the AAC standard frame (audio data) can be decoded, the decoding of the 3D Audio metadata may not be completed. is there. This is because when the 3D Audio metadata is encoded by differential encoding or the like, the starting point of the difference is unknown until a reset frame in which differential encoding is not performed appears.
[0097]
 Also, when playing the 3D AAC bitstream from the middle, some frames may omit the 3D Audio metadata.
[0098]
 Therefore, in the present technology, when the 3D Audio metadata of the current frame cannot be obtained, the 3D Audio metadata of the temporally previous frame is used as it is, and the 3D Audio metadata stored in advance on the decoding device 51 side is also used. The default value can be used.
[0099]
 Note that, hereinafter, other 3D Audio metadata different from the 3D Audio metadata read from the 3D AAC bitstream, such as a default value stored in advance on the decoding device 51 side, is also referred to as externally acquired 3D Audio metadata. I will.
[0100]
 For example, as shown in FIG. 7, it is assumed that 3D Audio metadata cannot be obtained in some frames at the beginning of the content.
[0101]
 In FIG. 7, each of the rectangles indicated by the arrows A21 to A25 represents the respective AAC frame data of the frame 0 to the frame 4.
[0102]
 In this example, from the first frame 0 to the fourth frame 3, the audio data according to the AAC standard could be decoded, but the 3D Audio metadata could not be decoded, or the 3D Audio metadata was stored in the DSE. There is no 3D Audio metadata because it is not stored.
[0103]
 Therefore, the 3D Audio metadata acquisition unit 63 acquires a predetermined initial metadata value as externally acquired 3D Audio metadata and supplies it to the 3D Audio renderer 64.
[0104]
 For example, in the initial metadata value, the position indicated by the position information of the audio object is a predetermined position such as the position directly in front of the user in the three-dimensional space.
[0105]
 When acquiring the initial metadata value, for example, the 3D Audio metadata acquisition unit 63 acquires the initial metadata value as externally acquired 3D Audio metadata based on the 3D Audio component element information read from the PCE. .. Specifically, for example, based on the number of object audios indicated by the 3D Audio component information, position information and gain information corresponding to the number of audio objects are acquired.
[0106]
 When acquiring externally acquired 3D Audio metadata, information indicating the structure of the AAC audio element data, that is, the structure of the audio data such as the number of channel bases and the number of object audios is required. Therefore, for example, when the value of the previous frame reference flag is set to “1” in the encoding device 11 and the 3D Audio metadata is not stored in the DSE, the 3D Audio component element information is always described in the PCE. ..
[0107]
 When the 3D Audio renderer 64 receives the supply of the initial metadata value as the externally acquired 3D Audio metadata, the 3D Audio renderer 64 renders the frames 0 to 3 using the supplied initial metadata value. Then, when the 3D Audio metadata is read from the DSE in the frame 4, the rendering of the frame 4 is performed based on the read 3D Audio metadata.
[0108]
 In this way, if the 3D Audio metadata cannot be obtained from the DSE for a long time, such as at the beginning or in the middle of the content, the externally acquired 3D Audio metadata obtained based on the 3D Audio component information etc. If used, 3D Audio can be played for the time being.
[0109]
 For example, on the encoding device 11 side, when there is no change in the 3D Audio metadata over a plurality of frames, such as when there is no motion in the object, the value of the previous frame reference flag is set to “1”, and the 3D Audio metadata is maintained for a long period of time. Data may not be transmitted. In such a case, if the 3D Audio metadata to be referred to on the decoding device 51 side is not obtained for some reason, it is possible to reproduce the 3D Audio by using the externally acquired 3D Audio metadata. it can.
[0110]
 When using externally acquired 3D Audio metadata, when using 3D Audio metadata read from the DSE after that, 3D Audio metadata is suddenly read from the DSE from the externally acquired one. When it changes into a thing, discomfort may occur, such as the position of the audio object becoming discontinuous. Therefore, for example, during a few frames, the interpolation process is performed based on the externally acquired 3D Audio metadata used last and the 3D Audio metadata read from the DSE, and the 3D Audio metadata obtained by the interpolation process. May be used so that the position of the audio object or the like smoothly transitions.
[0111]
 The externally acquired 3D Audio metadata is generated by the 3D Audio metadata acquisition unit 63 based on the information acquired from the outside, is acquired from the server via the network, or is encoded separately from the 3D AAC bitstream. It may be acquired in any way, such as being acquired from the encryption device 11. Alternatively, externally acquired 3D Audio metadata of each reproduction time may be prepared in advance for each content, and appropriate externally acquired 3D Audio metadata may be acquired according to the reproduction time of the content.
[0112]
 Furthermore, the 3D audio metadata may be changed according to a user's instruction or the like. As an example, for a content that is interactively shared and played by a plurality of users, such as a network battle game, the 3D audio meta-data is considered in consideration of not only the user operation of the decoding device 51 but also the operation of another user. The data can be changed.
[0113]
 As described above, when externally acquired 3D Audio metadata is acquired, the AAC audio element data encoded by the AAC encoding method is channel-based, object-based, or ambisonic. Identification information is required to identify the type of audio data, such as whether it is based. In other words, the configuration information indicating the configuration of the AAC audio element data is needed. Such information must be available without decoding the 3D Audio metadata.
[0114]
 Therefore, in the present technology, as described with reference to FIG. 2, 3D Audio component information is stored in the comment field of the PCE as information for identifying the type of audio data, that is, information indicating the configuration of AAC audio element data. Has been done. The information indicated by the 3D Audio component element information is also included in the 3D Audio metadata.
[0115]
 By storing the 3D Audio component information in the comment field of the PCE, the decoding device 51 can obtain the 3D Audio component information during the decoding process of the AAC standard in the AAC decoding unit 61. Since PCE is also used independently in the MPEG-4 File Format layer, storing 3D Audio component information in PCE has the advantage that 3D Audio component information can also be obtained from the system layer. ..
[0116]
 Regarding the comment field of PCE, as in the case of DSE, a decoding device that does not support this technology can skip the number of bytes in the comment field recorded in the AAC standard. Can be maintained.
[0117]

 By the way, according to the AAC standard, speaker arrangement information of each audio element, that is, each audio data, is transmitted to the front position (FRONT), horizontal position (SIDE), or rear position (in the PCE). BACK) can be recorded as position information on a two-dimensional plane. That is, it is possible to store in the PCE speaker arrangement information indicating which of the front, rear, and side directions the speaker reproduces the audio data.
[0118]
 On the other hand, in the 3D AAC bitstream, as position information of each audio object, that is, object-based audio data, the angle indicating the position in the plane direction of the audio object and the angle indicating the position in the height direction are included in 3D Audio metadata. include.
[0119]
 In the present technology, although there is an upper limit (limit) on the number of elements, the speaker placement information corresponding to the position of the audio object may be described in the PCE according to the angle in the plane direction as much as possible.
[0120]
 Specifically, for example, it is assumed that there are four audio objects EL11 to EL14 as audio objects as shown in FIG.
[0121]
 In FIG. 8, the near side is the front direction of the user U11 who is the viewer, that is, the front direction, and the far side is the rear side of the user U11 in the drawing.
[0122]
 In this example, the audio object EL11 is located on the front upper side of the user U11, and the audio object EL12 is located on the front upper right of the user U11. The audio object EL13 is located on the upper rear side of the user U11, and the audio object EL14 is located on the upper left side of the user U11.
[0123]
 In such a case, in the encoding device 11, the audio object EL11 and the audio object EL12 are regarded as the front position audio element located in the front. Then, the information indicating the front position is described in the PCE as the speaker arrangement information of those audio objects.
[0124]
 That is, the positions of the audio objects EL11 and EL12 in the three-dimensional space are remapped to the front positions on the two-dimensional plane. Then, according to the mapping result, the PCE describes position information of these audio objects, that is, information indicating a front position as speaker arrangement information.
[0125]
 Further, the audio object EL13 is set as a rear position audio element positioned rearward, and information indicating the rear position which is the position of the audio object EL13 on the two-dimensional plane is described in the PCE as speaker arrangement information.
[0126]
 Similarly, the audio object EL14 is a lateral position audio element positioned laterally, and information indicating the lateral position which is the position of the audio object EL14 on the two-dimensional plane is described in the PCE as speaker layout information.
[0127]
 Such speaker arrangement information is information indicating the speaker to which the audio data of the audio object should be output, that is, the speaker arrangement corresponding to the audio object. In other words, the speaker arrangement information is information indicating which channel the audio data of each audio object belongs to, that is, channel mapping of the audio object.
[0128]
 Such speaker arrangement information can also be said to be position information indicating the position of the audio object on the two-dimensional plane. Hereinafter, such speaker arrangement information is also referred to as position information indicating the position of the audio object on the two-dimensional plane.
[0129]
 In this way, by describing the position information on the two-dimensional plane of each audio object in the PCE, the 3D Audio metadata is used in the decoding device 51 for some reason, such as the decoding device 51 not supporting 3D Audio. Even if it is not possible, the position of each audio object on the two-dimensional plane can be specified.
[0130]
 Therefore, for example, as shown in FIG. 9, the sound image of each audio object can be localized at the correct position on the two-dimensional plane based on the position information on the two-dimensional plane described in the PCE. Note that in FIG. 9, portions corresponding to those in FIG. 8 are denoted by the same reference numerals, and description thereof will be omitted as appropriate.
[0131]
 In this example, in the decoding device 51, the position information (speaker arrangement information) indicating the positions of the audio objects EL11 to EL14 on the two-dimensional plane is read from the PCE, so that each audio is placed at the position indicated by the position information. Objects can be placed.
[0132]
 That is, although information in the height direction cannot be obtained, 3D Audio channel mapping can be assigned to existing channel mapping as much as possible. As a result, the sound based on the object-based audio element data can be output from the speaker at the approximate position on the plane of the audio object.
[0133]
 In the example shown in FIG. 9, the height of the audio objects EL11 to EL14 is different from the original height, and is the height of the head of the user U11. However, the audio object EL11 and the audio object EL12 are arranged in front of the user U11, the audio object EL13 is arranged behind the user U11, and the audio object EL14 is arranged on the left side of the user U11.
[0134]
 As described above, by using the position information indicating the position on the two-dimensional plane, it is possible to localize the sound images of the sounds of the audio objects EL11 to EL14 to the positions in substantially the same direction as the original direction.
[0135]

 Next , operations of the encoding device 11 and the decoding device 51 will be described.
[0136]
 First, the encoding process by the encoding device 11 will be described with reference to the flowchart in FIG. 10. This encoding process is started when audio data to be encoded is supplied.
[0137]
 In step S11, the 3D Audio metadata encoding unit 21 determines whether or not there is 3D Audio metadata to be transmitted for the current frame to be processed.
[0138]
 For example, when 3D Audio metadata is supplied from the outside, it is determined that there is 3D Audio metadata to be transmitted. When referring to the 3D Audio metadata of the previous frame, for example, when the 3D Audio metadata of the current frame to be processed and the frame immediately before it does not change, the 3D Audio metadata of the current frame to be processed is Not supplied.
[0139]
 When it is determined in step S11 that there is 3D Audio metadata to be transmitted, the 3D Audio metadata encoding unit 21 encodes 3D Audio metadata supplied from the outside in step S12. Process. Then, the 3D Audio metadata encoding unit 21 supplies the encoded 3D Audio metadata obtained by the 3D Audio metadata encoding process to the AAC encoding unit 22.
[0140]
 In step S13, the AAC encoding unit 22 performs AAC encoding on the encoded 3D Audio metadata supplied from the 3D Audio metadata encoding unit 21, the previous frame reference flag, and the audio data supplied from the outside, and one frame worth of data. Generate AAC frame data of.
[0141]
 That is, the AAC encoding unit 22 stores the previous frame reference flag whose value is “0” in the DSE of the AAC frame data, and then stores the encoded 3D Audio metadata in the DSE following the previous frame reference flag. To do.
[0142]
 Further, the AAC encoding unit 22 encodes one frame of audio data supplied from the outside by the AAC encoding method, and stores it as AAC audio element data in SCE, CPE, LFE, etc. of the AAC frame data.
[0143]
 Furthermore, the AAC encoding unit 22 also stores 3D Audio constituent element information supplied from the outside and position information (speaker arrangement information) indicating the position of the audio object on the two-dimensional plane in the PCE of the AAC frame data. As a result, for example, the AAC frame data having the configuration shown in FIG. 5 is obtained as a 3D AAC bit stream for one frame.
[0144]
 When the AAC encoding unit 22 outputs (transmits) the thus obtained 1D 3D AAC bitstream for one frame, the process proceeds to step S17.
[0145]
 On the other hand, if it is determined in step S11 that there is no 3D Audio metadata to be transmitted, the AAC encoding unit 22 determines in step S14 whether or not the decoding side refers to the 3D Audio metadata of the previous frame. To do.
[0146]
 If it is determined in step S14 that the 3D Audio metadata of the previous frame should be referred to, then the process proceeds to step S15.
[0147]
 In step S15, the AAC encoding unit 22 sets the value of the previous frame reference flag to "1" and AAC encodes the audio data supplied from the outside.
[0148]
 That is, the AAC encoding unit 22 stores the previous frame reference flag whose value is “1” in the DSE of the AAC frame data. Further, the AAC encoding unit 22 encodes one frame of audio data supplied from the outside by the AAC encoding method, and stores it as AAC audio element data in SCE, CPE, LFE, etc. of the AAC frame data.
[0149]
 Furthermore, the AAC encoding unit 22 also stores 3D Audio component element information supplied from the outside and position information indicating the position of the audio object on the two-dimensional plane in the PCE of the AAC frame data.
[0150]
 When the AAC encoding unit 22 outputs the AAC frame data thus obtained as a 3D AAC bit stream for one frame, the process proceeds to step S17.
[0151]
 On the other hand, if it is determined in step S14 that the 3D Audio metadata of the previous frame is not referred to, then the process proceeds to step S16.
[0152]
 In step S16, the AAC encoding unit 22 sets the value of the previous frame reference flag to "0" and AAC encodes the audio data supplied from the outside.
[0153]
 That is, the AAC encoding unit 22 stores the previous frame reference flag whose value is “0” in the DSE of the AAC frame data, and encodes the externally supplied audio data for one frame by the AAC encoding method. And store it in the AAC frame data. In the following, the encoded audio data will also be referred to as encoded audio data.
[0154]
 Furthermore, the AAC encoding unit 22 also stores 3D Audio component element information supplied from the outside and position information indicating the position of the audio object on the two-dimensional plane in the PCE of the AAC frame data.
[0155]
 When the AAC encoding unit 22 outputs the AAC frame data thus obtained as a 3D AAC bit stream for one frame, the process proceeds to step S17.
[0156]
 When the process of step S13, the process of step S15, or the process of step S16 is performed, the process of step S17 is performed thereafter.
[0157]
 In step S17, the encoding device 11 determines whether or not there is audio data to be encoded.
[0158]
 If it is determined in step S17 that there is still audio data to be encoded, the process returns to step S11, and the above-described process is repeated.
[0159]
 On the other hand, when it is determined in step S17 that there is no audio data to be encoded, that is, all the audio data of the content to be transmitted has been encoded and output, the encoding process ends.
[0160]
 As described above, the encoding device 11 stores the 3D Audio metadata in the DSE and generates the 3D AAC bitstream. By doing so, the existing AAC standard can be extended, and 3D Audio playback can be performed using 3D Audio metadata on the decoding side. Thereby, convenience can be improved.
[0161]
 Also, by storing the previous frame reference flag, the decoding side can refer to the 3D Audio metadata of the previous frame in terms of time, thus reducing the code amount of the 3D AAC bitstream by the amount of 3D Audio metadata. can do.
[0162]
 Furthermore, by describing the position information indicating the position of each audio object on the 2D plane in the PCE of the 3D AAC bitstream, even if 3D Audio metadata cannot be obtained on the decoding side, the sound image of each audio object can be obtained. Can be localized at the correct position on the two-dimensional plane. That is, audio reproduction can be performed without impairing the sense of presence.
[0163]

 Next, the decoding process performed by the decoding device 51 will be described with reference to the flowchart in FIG. 11. This decoding process is started when the 3D AAC bitstream transmitted by the encoding device 11 is received.
[0164]
 In step S41, the AAC decoding unit 61 performs an AAC decoding process, which is a decoding process in a decoding system corresponding to the AAC encoding system, on the received 3D AAC bit stream for one frame, that is, AAC frame data.
[0165]
 That is, the AAC decoding unit 61 reads the 3D Audio component information from the PCE and supplies it to the upper control unit, or reads the position information on the two-dimensional plane of each audio object from the PCE, that is, the speaker arrangement information, and directly Or, it is supplied to the 3D Audio renderer 64 via the upper control unit.
[0166]
 Further, the AAC decoding unit 61 decodes the encoded audio data stored in the AAC frame data, and supplies the audio data, which is PCM data obtained as a result, to the 3D Audio renderer 64.
[0167]
 In step S42, the AAC decoding unit 61 determines whether or not there is 3D Audio metadata in the DSE of AAC frame data.
[0168]
 For example, the AAC decoding unit 61 determines that there is 3D Audio metadata when encoded 3D Audio metadata is described following the value of the previous frame reference flag described at the beginning of the DSE. In this case, the value of the previous frame reference flag is "0".
[0169]
 When it is determined in step S42 that there is 3D Audio metadata, the AAC decoding unit 61 reads the encoded 3D Audio metadata from the DSE and supplies it to the 3D Audio metadata decoding unit 62, and then the process proceeds to step S43. And proceed.
[0170]
 In step S43, the 3D Audio metadata decoding unit 62 decodes the encoded 3D Audio metadata supplied from the AAC decoding unit 61, and supplies the 3D Audio metadata obtained as a result to the 3D Audio renderer 64. When the 3D Audio metadata is obtained in this way, the process then proceeds to step S47.
[0171]
 When it is determined in step S42 that there is no 3D Audio metadata, the AAC decoding unit 61 determines in step S44 whether or not the value of the previous frame reference flag described at the beginning of the DSE is “1”. judge.
[0172]
 If the value of the previous frame reference flag is not “1” in step S44, that is, the value of the previous frame reference flag is “0” but 3D Audio metadata is not obtained from DSE, then the process proceeds to step S48. move on.
[0173]
 On the other hand, if it is determined in step S44 that the value of the previous frame reference flag is "1", then the process proceeds to step S45. In this case, the encoded 3D Audio metadata is not stored in the DSE.
[0174]
 In step S45, the AAC decoding unit 61 determines whether there is past 3D Audio metadata. In other words, if there is 3D Audio metadata read from the DSE for the same content, that is, for the frame that is temporally previous to the current frame to be processed in the 3D AAC bitstream to be decoded, the past 3D Audio metadata is It is determined that there is.
[0175]
 If it is determined in step S45 that there is past 3D Audio metadata, the AAC decoding unit 61 determines the frame closest in time to the current frame among the past frames with 3D Audio metadata read from the DSE. Of the 3D Audio metadata, that is, the output of the last used 3D Audio metadata is instructed to the 3D Audio metadata decoding unit 62, and then the process proceeds to step S47.
[0176]
 In this case, the 3D Audio metadata decoding unit 62 supplies the 3D Audio metadata used in the past frame to the 3D Audio renderer 64 according to the instruction of the AAC decoding unit 61.
[0177]
 Specifically, assuming that the frame (N+1) is the current frame in the example shown in FIG. 6, the 3D Audio metadata decoding unit 62 outputs the 3D Audio metadata (N) of the frame N to the 3D Audio renderer. Supply to 64.
[0178]
 On the other hand, if it is determined in step S45 that there is no past 3D Audio metadata, the 3D Audio metadata used in the current frame could not be obtained from the 3D AAC bitstream for some reason, and then the process proceeds to step S46. And proceed.
[0179]
 In step S46, the AAC decoding unit 61 determines whether there is 3D Audio metadata acquired from other than the 3D AAC bitstream. That is, it is determined whether or not there is externally acquired 3D Audio metadata acquired by the 3D Audio metadata acquisition unit 63.
[0180]
 If it is determined in step S46 that there is no acquired 3D Audio metadata, then the process proceeds to step S48.
[0181]
 If it is determined in step S46 that there is acquired 3D Audio metadata, the AAC decoding unit 61 instructs the 3D Audio metadata acquisition unit 63 to output the acquired 3D Audio metadata, and then the processing is performed. Advances to step S47.
[0182]
 In this case, the 3D Audio metadata acquisition unit 63 supplies to the 3D Audio renderer 64 the externally acquired 3D Audio metadata acquired by being externally generated or being generated by itself.
[0183]
 If the process of step S43 is performed, it is determined that there is past 3D Audio metadata in step S45, or it is determined that there is 3D Audio metadata acquired in step S46, the process of step S47 is performed thereafter. Is done.
[0184]
 In step S47, the 3D Audio renderer 64 renders based on the 3D Audio metadata supplied from the 3D Audio metadata decoding unit 62 or the 3D Audio metadata acquisition unit 63 and the audio data supplied from the AAC decoding unit 61. I do.
[0185]
 Then, the 3D Audio renderer 64 outputs the reproduction data of the current frame for realizing the 3D Audio reproduction obtained by the rendering, and then the process proceeds to step S49.
[0186]
 If it is determined in step S44 that the value of the previous frame reference flag is not "1" or there is no 3D Audio metadata acquired in step S46, the process of step S48 is performed.
[0187]
 That is, in step S48, the 3D Audio renderer 64 performs the rendering defined by the normal AAC standard based on the audio data supplied from the AAC decoding unit 61, and outputs the reproduction data obtained as a result.
[0188]
 In this case, the 3D Audio renderer 64, based on the position information (speaker placement information) on the two-dimensional plane of each audio object read from the PCE, the two-dimensional sound image of each audio object is indicated by the position information. Audio data of each channel is generated as reproduction data so as to be localized at a position on the plane. As a result, even if the 3D Audio metadata cannot be obtained and the sound image localization control in the height direction of the audio object (sound source) cannot be performed, the sound image of each audio object is localized at the correct position on the two-dimensional plane. be able to.
[0189]
 When the reproduction data of the current frame is output in this way, the process then proceeds to step S49.
[0190]
 When the process of step S47 or the process of step S48 is performed and the reproduction data for one frame is output, the AAC decoding unit 61 determines in step S49 whether there is a 3D AAC bitstream to be decoded.
[0191]
 For example, in step S49, if all the received 3D AAC bitstreams have not been decoded yet, it is determined that there is a 3D AAC bitstream to be decoded.
[0192]
 If it is determined in step S49 that there is a 3D AAC bitstream to be decoded, the process returns to step S41, and the above process is repeated.
[0193]
 On the other hand, if it is determined in step S49 that there is no 3D AAC bitstream to be decoded, the decoding process ends.
[0194]
 As described above, the decoding device 51 reads 3D Audio metadata from the DSE and performs rendering using the read 3D Audio metadata. By doing so, it is possible to extend the existing AAC standard and perform 3D audio playback, and improve convenience.
[0195]
 Also, since the position information indicating the position of each audio object on the two-dimensional plane is described in the PCE of the 3D AAC bitstream, even if 3D Audio metadata cannot be obtained, the sound image of each audio object can be obtained. It can be localized at the correct position on the two-dimensional plane. That is, audio reproduction can be performed without impairing the sense of presence.
[0196]

 By the way, the larger the number of audio objects, that is, the number of object audios, the larger the data amount (code amount) of 3D Audio metadata.
[0197]
 Therefore, when the data amount of 3D Audio metadata increases, the size (data amount) of 3D Audio metadata may exceed the upper limit size of DSE defined in the AAC standard. In addition, the number of DSEs required to store 3D Audio metadata in one frame may exceed the upper limit number of DSEs defined in the AAC standard.
[0198]
 Therefore, in the present technology, when the amount of 3D Audio metadata is large, instead of transmitting the 3D Audio metadata of each of a plurality of frames, the 3D Audio metadata is appropriately thinned out, and the 3D Audio metadata is reduced. The representative value of is stored in the DSE of those multiple frames and transmitted. That is, the representative value is transmitted over a plurality of frames.
[0199]
 For example, as shown in FIG. 12, it is assumed that three DSEs can be provided in one frame of AAC frame data. It should be noted that in FIG. 12, the quadrangle with each of the characters “DSE 1” to “DSE 3” represents each DSE provided in one frame of AAC frame data. Hereinafter, those DSEs are also referred to as DSE 1 to DSE 3.
[0200]
 In this example, the amount of 3D Audio metadata for one frame is large, and the 3D Audio metadata cannot be stored in the first DSE 1.
[0201]
 In such a case, the encoding device 11 divides the 3D Audio metadata into three pieces of data. Here, three pieces of data obtained by dividing the 3D Audio metadata will be referred to as divided 3D Audio metadata (1) to divided 3D Audio metadata (3). Further, hereinafter, when it is not particularly necessary to distinguish each data obtained by dividing the 3D Audio metadata, it is also referred to as divided 3D Audio metadata.
[0202]
 The encoding device 11 stores each of the divided 3D Audio metadata (1) to the divided 3D Audio metadata (3) obtained by the division in each of DSE 1 to DSE 3.
[0203]
 At this time, the encoding device 11 can specify that the 3D Audio metadata is divided and which divided 3D Audio metadata is the number of data from the beginning so that the intra-frame DSE division counter can be specified. And an intra-frame DSE division end flag is generated.
[0204]
 The intra-frame DSE division counter is information indicating the number of the DSE in which the divided 3D Audio metadata is stored in the AAC frame data for one frame.
[0205]
 In other words, the intra-frame DSE division counter is information indicating how many divided 3D Audio metadata pieces the divided 3D Audio metadata stored in the DSE is stored in the DSE of one AAC frame. It can be said that there is.
[0206]
 For example, the divided 3D audio metadata stored in the DSE having the intra-frame DSE division counter value of “i−1” (1≦i) is i from the beginning stored in the DSE of one AAC frame. This is the 3rd divided 3D Audio metadata.
[0207]
 The intra-frame DSE division end flag is flag information indicating whether or not it is the last DSE in which the divided 3D Audio metadata is stored among the DSEs arranged in order in the AAC frame data for one frame. is there.
[0208]
 Specifically, when the value of the intra-frame DSE division end flag is “0”, it indicates that the DSE is not the last DSE in which the division 3D Audio metadata is stored.
[0209]
 On the other hand, when the value of the intra-frame DSE division end flag is “1”, it indicates that the DSE is the last DSE in the AAC frame data in which the division 3D Audio metadata is stored.
[0210]
 The encoding device 11 also stores an intra-frame DSE division counter and an intra-frame DSE division end flag together with the divided 3D Audio metadata in each DSE.
[0211]
 That is, in this example, the first DSE 1 indicates the first divided 3D Audio metadata, the intra-frame DSE division counter whose value is “0”, and the value indicating that it is not the last DSE. An intra-frame DSE division end flag having "0" and divisional 3D Audio metadata (1) are stored.
[0212]
 Also, the second DSE 2 has the intra-frame DSE division counter with a value of “1” indicating that it is the second divided 3D Audio metadata and the value “0” indicating that it is not the last DSE. Intra-frame DSE division end flag which is “” and division 3D Audio metadata (2) are stored. Furthermore, the 3rd DSE 3 indicates that it is the 3rd divided 3D Audio metadata, the intra-frame DSE division counter whose value is "2", and the value that indicates the last DSE is "1". Intra-frame DSE division end flag which is “” and division 3D Audio metadata (3) are stored.
[0213]
 In this way, by storing the intra-frame DSE division counter and the intra-frame DSE division end flag, the decoding side reads the divided 3D Audio metadata and correctly combines the divided 3D Audio metadata to obtain the original 3D Audio. You can get metadata.
[0214]
 Further, when 3D Audio metadata cannot be stored in one frame of AAC frame data, the encoding device 11 sets the representative value as 3D Audio metadata to a plurality of frames of AAC frame data as shown in FIG. It is divided into and stored. It should be noted that here, one DSE is provided for one frame for the sake of simplicity.
[0215]
 In FIG. 13, the quadrangle shown by each of the arrows A51 to A55 represents the respective AAC frame data of the frame (N-1) to the frame (N+3). In addition, here, the 3D Audio metadata of the original frame (i) (where N−1≦i≦N+3) is also referred to as 3D Audio metadata (i).
[0216]
 In this example, originally, 3D Audio metadata (N-1) to 3D Audio metadata (N+3) were prepared for frames (N-1) to frames (N+3). However, since the amount of data of 3D Audio metadata (N) to 3D Audio metadata (N+3) is large, these 3D Audio metadata cannot be stored in the corresponding DSE for one frame. It is in a state.
[0217]
 Therefore, the encoding device 11 generates one new 3D Audio metadata for obtaining the 3D Audio metadata for the frame (N) to the frame (N+3). In other words, some 3D Audio metadata is thinned out and a typical 3D Audio metadata is selected.
[0218]
 Specifically, for example, the encoding device 11 sets one 3D Audio metadata representing the 3D Audio metadata (N) to 3D Audio metadata (N+2) as a representative value.
[0219]
 Here, the representative value may be any one of 3D Audio metadata (N) to 3D Audio metadata (N+2), or may be 3D Audio metadata (N). The average value of the 3D audio metadata (N+2) may be used as the representative value.
[0220]
 In this case, for example, when focusing on one audio object, the average value of the position information of the frame (N) to the frame (N+2) of the audio object as the 3D Audio metadata is the 3D Audio metadata of the audio object. It is assumed to be the representative value of.
[0221]
 In addition, there are, for example, a plurality of audio objects, one particular audio object of the plurality of audio objects is the most important, and the movement of the particular audio object in the frame (N+1) is the most important. Suppose
[0222]
 In such a case, for example, position information as 3D Audio metadata (N+1) of a specific audio object is set as a representative value of the specific audio object. In addition, regarding all other audio objects, the position information as 3D Audio metadata (N+1) of those audio objects becomes the respective representative values ​​of those audio objects.
[0223]
 Further, for example, if there are multiple audio objects, one particular audio object of the plurality of audio objects is the most important, and the movement of the particular audio object in the frame (N+1) is the most important. Suppose In such a case, for example, the position information as 3D Audio metadata (N+1) of the specific audio object may be set as the representative value of all audio objects.
[0224]
 As yet another example, the average value of the position information in each of the frames in which the most movement of each of the plurality of audio objects is important is obtained, and the obtained average value is used as a representative value common to all audio objects. You may
[0225]
 When the encoding device 11 determines the representative value of the 3D Audio metadata (N) to 3D Audio metadata (N+2), the encoding device 11 divides the period into a plurality of frames for transmitting one 3D Audio metadata (hereinafter, 3D Audio metadata of the end frame of the divided transmission period) is generated as the end value.
[0226]
 For example, in this example, since the period from the frame (N) to the frame (N+3) is the divided transmission period, the end frame of the divided transmission period is the frame (N+3). Therefore, the encoding device 11 determines the termination value by some method such as using the 3D Audio metadata (N+3) as the termination value as it is.
[0227]
 Further, the encoding device 11 sets the frame to which the representative value is applied among the frames within the divided transmission period as the representative value applied frame, and generates the applied frame index indicating the representative value applied frame. Here, the representative value application frame is a frame whose representative value is used as it is as 3D Audio metadata on the decoding side.
[0228]
 For example, in the example of FIG. 13, when the representative value application frame is the frame (N+1), the encoding device 11 generates information indicating the frame (N+1) as the application frame index.
[0229]
 Specifically, for example, when the representative value application frame is the i-th (1≦i) frame from the beginning in the divided transmission period, the value of the application frame index is “i−1”. Therefore, in the example shown in FIG. 13, since the second frame (N+1) from the beginning of the divided transmission period is the representative value applicable frame, the value of the applicable frame index is “1”.
[0230]
 When the representative value, the end value, and the applicable frame index of the divided transmission period are obtained as described above, the encoding device 11 sets the representative value and the end value as the 3D Audio metadata of the divided transmission period and the 3D Split Audio metadata.
[0231]
 That is, the encoding device 11 divides the 3D Audio metadata of the divided transmission period into divided 3D Audio metadata of the number of frames included in the divided transmission period.
[0232]
 In this example, since the number of frames forming the divided transmission period is 4, the 3D Audio metadata is represented by the characters “divided 3D Audio metadata (1)” to “divided 3D Audio metadata (4)”. Split into 3D Audio metadata.
[0233]
 In addition, hereinafter, the divided 3D Audio metadata indicated by the characters “divided 3D Audio metadata (1)” to “divided 3D Audio metadata (4)” is referred to as divided 3D Audio metadata (1) to divided 3D Audio metadata. It is also referred to as (4).
[0234]
 When the divided 3D Audio metadata is obtained, the encoding device 11 stores the divided 3D Audio metadata in the DSE of the AAC frame data of each frame and outputs it.
[0235]
 Here, the quadrangle shown by each of the arrows A61 to A65 represents the respective AAC frame data of the frame (N-1) to the frame (N+3) actually output from the encoding device 11.
[0236]
 In this example, in the AAC frame data of the frame (N-1), the undivided 3D Audio metadata (N-1) is stored in the DSE as it is.
[0237]
 On the other hand, in the frame (N) AAC frame data, the divided 3D Audio metadata (1) and the applicable frame index are stored in the DSE, and in the frame (N+1) AAC frame data, the divided 3D Audio metadata (2). Is stored in DSE. Also, in the AAC frame data of the frame (N+2), the divided 3D Audio metadata (3) is stored in the DSE, and in the AAC frame data of the frame (N+3), the divided 3D Audio metadata (4) is the DSE. Stored in.
[0238]
 The applicable frame index may be stored in the DSE of any frame as long as it is the DSE of the frame within the divided transmission period.
[0239]
 Further, the decoding device 51 that has received such a 3D AAC bitstream performs decoding as shown in FIG. 14, for example. In addition, in FIG. 14, description of portions corresponding to the case in FIG. 13 will be appropriately omitted.
[0240]
 In the example illustrated in FIG. 14, the decoding device 51 receives the 3D AAC bitstream including the AAC frame data indicated by the arrows A71 to A75.
[0241]
 Here, the AAC frame data of each frame shown by each of the arrows A71 to A75 corresponds to the AAC frame data of each frame shown by each of the arrows A61 to A65 in FIG.
[0242]
 For a frame (N-1) in which the 3D Audio metadata has not been divided, the decoding device 51 determines from the AAC frame data of the frame (N-1) the 3D Audio metadata (N-1) as indicated by an arrow A81. -1) Read out. Then, the decoding device 51 performs rendering based on the read 3D Audio metadata (N-1).
[0243]
 On the other hand, for frames (N) to frames (N+3) in which the 3D Audio metadata has been divided, based on the divided 3D Audio metadata (1) to the divided 3D Audio metadata (4), 3D Audio metadata for those frames is generated.
[0244]
 That is, the AAC decoding unit 61 of the decoding device 51 divides the 3D Audio metadata for each frame from each of the DSEs of the frame (N) to the frame (N+3) to obtain the divided 3D Audio metadata ( Each of 1) to the divided 3D audio metadata (4) is read. Then, the AAC decoding unit 61 combines the read divided 3D audio metadata (1) to divided 3D audio metadata (4) into one encoded 3D audio metadata.
[0245]
 Then, the decoding device 51 decodes the obtained encoded 3D Audio metadata to obtain the representative value and the end value, and reads the applicable frame index from the DSE. Then, the decoding device 51 generates 3D Audio metadata of the frame within the divided transmission period by performing interpolation processing based on the obtained representative value, terminal value, and applied frame index.
[0246]
 Specifically, the decoding device 51 sets the representative value to 3D Audio metadata of the frame (N+1) indicated by the applicable frame index, as indicated by arrow A83.
[0247]
 Also, the decoding device 51, based on the 3D Audio metadata (N-1) of the frame (N-1) and the representative value that is the 3D Audio metadata of the frame (N+1), as indicated by the arrow A82. , 3D Audio metadata (N) of the frame (N) is generated by interpolation processing.
[0248]
 Further, the decoding device 51 uses the end value as it is as the 3D Audio metadata of the frame (N+3) as shown by an arrow A85. The decoding device 51 performs the interpolation processing based on the representative value that is the 3D Audio metadata of the frame (N+1) and the end value that is the 3D Audio metadata of the frame (N+3) as indicated by arrow A84. Generates 3D Audio metadata (N+2) for the frame (N+2).
[0249]
 In the frame (N) to the frame (N+3), the decoding device 51 performs rendering using the 3D Audio metadata (N) to the 3D Audio metadata (N+3) thus obtained.
[0250]
 The decoding device 51 reads the divided 3D Audio metadata stored by dividing the plurality of frames as described above, and obtains the 3D Audio metadata of each frame in the divided transmission period. By doing so, it is possible to efficiently transmit 3D Audio metadata, improve convenience, and realize realistic 3D Audio playback.
[0251]
 In addition, although the example in which the termination value is included in the 3D Audio metadata of the divided transmission period has been described here, the termination value may not be included. In such a case, 3D Audio metadata of the frame immediately after the end of the divided transmission period may be used as the end value.
[0252]
As
 described above, in one frame of AAC frame data, when the 3D Audio metadata is divided into a plurality of divided 3D Audio metadata and stored in the plurality of DSEs, it is necessary for the decoding side to do so. You need to be able to recognize.
[0253]
 Therefore, on the encoding side, the intra-frame DSE division information that is information about the division of the 3D Audio metadata in the frame, that is, in the AAC frame data, including the intra-frame DSE division counter and the intra-frame DSE division end flag described above, Described in DSE.
[0254]
 The syntax of such intra-frame DSE division information is as shown in FIG. 15, for example. In the example shown in FIG. 15, the character “dse_div_info()” represents a function in which intra-frame DSE division information is stored, and this function exists in 3D Audio metadata of each frame. Here, the intra-frame DSE division information exists for each DSE.
[0255]
 The intra-frame DSE division information includes an intra-frame DSE division counter indicated by the character “dse_div_cnt” and an intra-frame DSE division end flag indicated by the character “dse_div_terminate”.
[0256]
 As described above, the intra-frame DSE division counter is counter information starting from 0, which indicates which of the DSEs in the corresponding frame in which the divided 3D Audio metadata in the same frame is stored. Here, since the intra-frame DSE division counter is 3-bit information, it is possible to divide 3D Audio metadata or divided 3D Audio metadata into a maximum of 8 pieces of divided 3D Audio metadata in one frame. Has become. That is, a maximum of 8 DSEs can be provided.
[0257]
 Further, when the value of the intra-frame DSE division end flag is “0”, it indicates that the corresponding DSE is not the last DSE in the frame where the divided 3D Audio metadata is stored. On the other hand, when the value of the intra-frame DSE division end flag is "1", it indicates that the corresponding DSE is the last DSE in the frame in which the divided 3D Audio metadata is stored. ..
[0258]
 Furthermore, when the divided 3D Audio metadata obtained by dividing one 3D Audio metadata is stored in the DSEs of a plurality of frames, it is necessary for the decoding side to be able to recognize this.
[0259]
 Therefore, on the encoding side, similar to the case of the intra-frame DSE division information, 3D Audio metadata division between frames, that is, frame division information that is information about division of 3D Audio metadata for multiple frames is described in the DSE. To be done.
[0260]
 The syntax of such frame division information is as shown in FIG. 16, for example.
[0261]
 In the example shown in FIG. 16, the character “frame_div_info()” represents a function in which frame division information is stored, and this function exists in the 3D Audio metadata of each frame. Here, there is one frame division information for each frame.
[0262]
 The frame division information includes frame division mode information indicated by the character “frm_div_mode”.
[0263]
 This frame division mode information indicates whether or not one 3D Audio metadata is transmitted over a plurality of frames, that is, whether or not the frame division mode is one in which one 3D Audio metadata is divided and stored in the DSE of a plurality of frames. Information.
[0264]
 Here, when the value of the frame division mode information is “0”, it indicates that the frame division mode is not set, that is, 3D Audio metadata for one frame is stored in the DSE for one frame.
[0265]
 On the other hand, when the value of the frame division mode information is “1”, it indicates that the frame division mode is set, that is, one 3D Audio metadata is divided and stored in the DSE of a plurality of frames. There is.
[0266]
 When the value of the frame division mode information is “1”, the frame division information further stores the frame division counter indicated by the character “frm_div_cnt” and the frame division end flag indicated by the character “frm_div_terminate”. Has been done.
[0267]
 The frame division counter is counter information starting from 0, which indicates from the beginning of the frame that the corresponding frame constitutes the divided transmission period. Here, since the frame division counter is 3-bit information, the maximum division transmission period can be set to 8 frames. That is, one 3D Audio metadata can be transmitted over 8 frames.
[0268]
 Further, when the value of the frame division end flag is “0”, it indicates that the corresponding frame is not the last frame in the division transmission period. On the other hand, when the value of the frame division termination flag is "1", it indicates that the corresponding frame is the last (termination) frame in the division transmission period.
[0269]
 Furthermore, when the value of the frame division counter is “0”, the applicable frame index indicated by the character “apply_frm_Index” is also stored in the frame division information. That is, the applicable frame index is stored in the DSE of the first (first) frame in the divided transmission period.
[0270]
 This applied frame index is information indicating a representative value applied frame, and here, the value of the frame division counter of the representative value applied frame is the value of the applied frame index.
[0271]
 When the number of divided frames, that is, the number of frames constituting the divided transmission period is F, and the value of the applicable frame index is 1 smaller than F (F-1), the 3D audio metadata is set. Only representative or end values ​​are included.
[0272]
 In other words, when the frame indicated by the applicable frame index is the end frame of the divided transmission period, the end value becomes the representative value, so only the representative value or the end value is transmitted as 3D Audio metadata. become.
[0273]
 One frame of AAC frame data stores the intra-frame DSE division information and the frame division information described above. In the following, information consisting of intra-frame DSE division information and frame division information will also be referred to as division information. However, the division information is stored in each DSE of the AAC frame data, and although the division information does not necessarily include the frame division information, the intra-frame DSE division information is always included.
[0274]
 In such a case, the syntax of the division information stored in the AAC frame data is as shown in FIG. 17, for example.
[0275]
 In the example shown in FIG. 17, the character “div_info()” represents a function in which division information is stored, and this function exists in the 3D Audio metadata of each frame.
[0276]
 In addition, here, as the division information, the intra-frame DSE division information indicated by the character "dse_div_info()" is included. When the value of the intra-frame DSE division counter (dse_div_cnt) included in the intra-frame DSE division information is “0”, the division information further includes the frame division information indicated by the characters “frame_div_info()”. Has been.
[0277]
 That is, the first (first) DSE in the AAC frame data includes intra-frame DSE division information and frame division information as division information, and the second and subsequent DSEs in the AAC frame data include intra-frame division information. Only DSE partition information is included.
[0278]
 When division information is stored in AAC frame data as shown in FIG. 17, when one 3D Audio metadata is transmitted over three frames, for example, it becomes as shown in FIG.
[0279]
 In the example shown in FIG. 18, one piece of 3D Audio metadata is stored in the AAC frame data (1) to AAC frame data (3) which are the AAC frame data indicated by the arrows A101 to A103.
[0280]
 That is, the first AAC frame data (1) includes DSE1 which is the first DSE and DSE2 which is the second DSE.
[0281]
 In DSE1 of this AAC frame data (1), in-frame DSE division information (dse_div_info()) and frame division information (frame_div_info()) as division information, and division 3D Audio meta data that is division 3D Audio metadata Data (1-1) and are stored.
[0282]
 In this case, in the intra-frame DSE division information in DSE1 of AAC frame data (1), the value of the intra-frame DSE division counter (dse_div_cnt) is set to “0”, and the value of the intra-frame DSE division end flag (dse_div_terminate) is It is set to "0".
[0283]
 Also, in the frame division information in the DSE1 of the AAC frame data (1), the value of the frame division mode information (frm_div_mode) is “1”, the value of the frame division counter (frm_div_cnt) is “0”, and the frame The value of the division end flag (frm_div_terminate) is set to "0", and the applicable frame index of an appropriate value is included.
[0284]
 Further, in the DSE2 of the AAC frame data (1), there are intra-frame DSE division information (dse_div_info()) as division information and divided 3D Audio metadata (1-2) which is divided 3D Audio metadata. It is stored.
[0285]
 In this case, in the intra-frame DSE division information in DSE2 of AAC frame data (1), the value of the intra-frame DSE division counter (dse_div_cnt) is set to “1”, and the value of the intra-frame DSE division end flag (dse_div_terminate) is It is set to "1".
[0286]
 Similarly, the second AAC frame data (2) is provided with the first DSE1 and the second DSE2.
[0287]
 In DSE1 of the AAC frame data (2), in-frame DSE division information (dse_div_info()) and frame division information (frame_div_info()) as division information, and division 3D Audio metadata which is division 3D Audio metadata (2-1) and are stored.
[0288]
 In this case, in the intra-frame DSE division information in DSE1 of AAC frame data (2), the value of the intra-frame DSE division counter (dse_div_cnt) is set to “0”, and the value of the intra-frame DSE division end flag (dse_div_terminate) is It is set to "0".
[0289]
 Also, in the frame division information in DSE1 of the AAC frame data (2), the value of the frame division mode information (frm_div_mode) is set to "1", the value of the frame division counter (frm_div_cnt) is set to "1", and The value of the division end flag (frm_div_terminate) is set to "0". The frame division information does not include the applicable frame index.
[0290]
 Further, in the DSE2 of the AAC frame data (2), there is intra-frame DSE division information (dse_div_info()) as division information and divided 3D Audio metadata (2-2) which is divided 3D Audio metadata. It is stored.
[0291]
 In this case, in the intra-frame DSE division information in DSE2 of AAC frame data (2), the value of the intra-frame DSE division counter (dse_div_cnt) is set to “1”, and the value of the intra-frame DSE division end flag (dse_div_terminate) is It is set to "1".
[0292]
 Further, the third AAC frame data (3) is provided with the first DSE1 and the second DSE2.
[0293]
 The DSE1 of the AAC frame data (3) includes in-frame DSE division information (dse_div_info()) and frame division information (frame_div_info()) as division information, and divided 3D Audio metadata that is divided 3D Audio metadata. (3-1) and are stored.
[0294]
 In this case, in the intra-frame DSE division information in DSE1 of AAC frame data (3), the value of the intra-frame DSE division counter (dse_div_cnt) is set to “0”, and the value of the intra-frame DSE division end flag (dse_div_terminate) is It is set to "0".
[0295]
 In the frame division information in DSE1 of the AAC frame data (3), the value of the frame division mode information (frm_div_mode) is “1”, the value of the frame division counter (frm_div_cnt) is “2”, and the frame division information is “2”. The value of the division end flag (frm_div_terminate) is set to "1". The frame division information does not include the applicable frame index.
[0296]
 Further, in the DSE2 of the AAC frame data (3), the intra-frame DSE division information (dse_div_info()) as division information and the divided 3D Audio metadata (3-2) which is the divided 3D Audio metadata are included. It is stored.
[0297]
 In this case, in the intra-frame DSE division information in DSE2 of AAC frame data (3), the value of the intra-frame DSE division counter (dse_div_cnt) is set to “1”, and the value of the intra-frame DSE division end flag (dse_div_terminate) is It is set to "1".
[0298]
 Therefore, in the decoding device 51, the divided 3D Audio metadata (1-1), the divided 3D Audio metadata (1-2), and the divided 3D Audio included in the AAC frame data (1) to AAC frame data (3) are included. The metadata (2-1), the divided 3D Audio metadata (2-2), the divided 3D Audio metadata (3-1), and the divided 3D Audio metadata (3-2) are read out and their divided 3D are read. Audio metadata is combined into one 3D Audio metadata. Then, the representative value and the end value described above are obtained by decoding one obtained 3D Audio metadata.
[0299]

 Next, an encoding process and a decoding process performed when the 3D Audio metadata is divided and stored will be described as appropriate. First, the encoding process by the encoding device 11 will be described with reference to the flowchart in FIG.
[0300]
 In step S81, the 3D audio metadata encoding unit 21 sets the value of the frame division mode information (frm_div_mode) to “0” for the current frame to be processed.
[0301]
 Then, in step S82, the 3D Audio metadata encoding unit 21 determines whether or not there is 3D Audio metadata to be transmitted in the current frame to be processed. In step S82, the same processing as that in step S11 of FIG. 10 is performed.
[0302]
 When it is determined in step S82 that there is 3D Audio metadata to be transmitted, the 3D Audio metadata encoding unit 21 performs 3D Audio metadata encoding processing in step S83, and 3D Audio supplied from the outside is supplied. Encode the metadata. In step S83, the same processing as step S12 in FIG. 10 is performed.
[0303]
 In step S84, the 3D Audio metadata encoding unit 21 determines that the code amount (data amount) of the encoded 3D Audio metadata of the current frame to be processed obtained in the process of step S83 is within the predetermined allowable range. Or not.
[0304]
 Here, if the amount of encoded 3D Audio metadata is larger (larger) than the threshold value that is the allowable range and the encoded 3D Audio metadata cannot be stored in the DSE of the AAC frame data of the current frame, the allowable range is set. It is determined that it is not within.
[0305]
 If it is determined in step S84 that it is within the allowable range, the 3D Audio metadata encoding unit 21 supplies the encoded 3D Audio metadata obtained in the process of step S83 to the AAC encoding unit 22 as it is, and thereafter. The process proceeds to step S90.
[0306]
 More specifically, at this time, when the encoded 3D Audio metadata cannot be stored in one DSE, the 3D Audio metadata encoding unit 21 converts the encoded 3D Audio metadata into some divided 3D Audio metadata. To divide. Then, the 3D Audio metadata encoding unit 21 causes each of the divided 3D Audio metadata obtained by the division to be stored in each of the plurality of DSEs in the AAC frame data of the current frame.
[0307]
 Further, the 3D Audio metadata encoding unit 21 also generates appropriate intra-frame DSE division information and frame division information, and also supplies the intra-frame DSE division information and frame division information to the AAC encoding unit 22. In particular, here, since the 3D Audio metadata of the current frame is stored in the AAC frame data of the current frame, the value of the frame division mode information (frm_div_mode) is set to “0”. That is, the frame division mode information set in step S81 is directly supplied to the AAC encoding unit 22.
[0308]
 The division of the encoded 3D Audio metadata and the generation of the intra-frame DSE division information and the frame division information may be performed by the AAC encoding unit 22 instead of the 3D Audio metadata encoding unit 21.
[0309]
 On the other hand, if it is determined in step S84 that it is not within the allowable range, then the process proceeds to step S85.
[0310]
 In step S85, the 3D Audio metadata encoding unit 21 changes the value of the frame division mode information (frm_div_mode) of the current frame set in step S81 to "1".
[0311]
 In addition, the 3D Audio metadata encoding unit 21 sets the frame that has been the current frame so far as the first frame of the divided transmission period.
[0312]
 The 3D Audio metadata encoding unit 21 sets the temporally next frame of the first frame as a new processing target frame, and also acquires the 3D Audio metadata of the new processing target frame. Further, the 3D Audio metadata encoding unit 21 deletes the position information and the like of the audio object of low importance and the audio object having no motion from the 3D Audio metadata of the frame of the divided transmission period as appropriate to perform 3D audio metadata. The data amount of Audio metadata may be reduced.
[0313]
 In step S86, the 3D Audio metadata encoding unit 21 calculates the representative value and the end value for the 3D Audio metadata of each frame from the first frame of the divided transmission period to the frame currently processed. To do.
[0314]
 That is, the 3D Audio metadata encoding unit 21 assumes that the frame currently being processed is the end frame of the divided transmission period.
[0315]
 Then, the 3D Audio metadata encoding unit 21 determines the divided transmission period based on the 3D Audio metadata of each frame forming the divided transmission period, more specifically, the 3D Audio metadata of each frame excluding the end frame. Determine the representative value of 3D Audio metadata of. For example, when the representative value of the 3D Audio metadata is determined, as described above, the calculation for obtaining the average value is performed as necessary, and the representative value application frame to which the representative value is applied is also included in the 3D Audio metadata encoding unit 21. Determined by
[0316]
 Further, the 3D Audio metadata encoding unit 21 determines the final termination value, which is the final 3D Audio metadata of the end frame, based on the 3D Audio metadata of the end frame of the divided transmission period.
[0317]
 The 3D Audio metadata encoding unit 21 sets the representative value and the termination value determined in this way as 3D Audio metadata for a plurality of frames forming the divided transmission period. In this case, the 3D Audio metadata encoding unit 21 functions as a metadata determination unit that determines the representative value and the end value of the divided transmission period.
[0318]
 In step S87, the 3D Audio metadata encoding unit 21 performs 3D Audio metadata encoding processing using the representative value and the termination value obtained in step S86 as 3D Audio metadata, and encodes the 3D Audio metadata.
[0319]
 In step S88, the 3D Audio metadata encoding unit 21 determines that the code amount (data amount) of the encoded 3D audio metadata obtained in the process of step S87 is an allowable range determined by the number of frames forming the divided transmission period. It is determined whether or not
[0320]
 Here, the data amount of encoded 3D Audio metadata is larger (larger) than the threshold value that is the allowable range, and the encoded 3D Audio metadata cannot be stored in the DSE of the AAC frame data of all the frames that configure the divided transmission period. If not, it is determined that it is not within the allowable range.
[0321]
 If it is determined in step S88 that the data amount is not within the allowable range, the data amount of 3D Audio metadata is still too large for the divided transmission period, so the process returns to step S86, and the above-described process is repeated. At this time, the frame temporally next to the end frame of the period which has been the divided transmission period so far is set as the end frame of the new divided transmission period.
[0322]
 On the other hand, if it is determined in step S88 that it is within the allowable range, the 3D Audio metadata encoding unit 21 performs frame division processing in step S89.
[0323]
 That is, the 3D Audio metadata encoding unit 21 divides the encoded 3D Audio metadata obtained in the process of step S87 into a plurality of divided 3D Audio metadata. For example, in the example shown in FIG. 18, the encoded 3D Audio metadata is divided into divided 3D Audio metadata (1-1) to divided 3D Audio metadata (3-2).
[0324]
 Also, the 3D Audio metadata encoding unit 21 determines the frame division information including the frame division mode information whose value determined in step S85 is “1” according to the division result of the 3D Audio metadata, the division transmission period, and the like. , Intra-frame DSE division information is also generated. Here, the frame division information is generated for each frame of the divided transmission period, and the intra-frame DSE division information is generated for each DSE of each frame of the divided transmission period.
[0325]
 Then, the 3D Audio metadata encoding unit 21 supplies the generated intra-frame DSE division information and frame division information and the divided 3D Audio metadata to the AAC encoding unit 22, and then the process proceeds to step S90. .. The frame division process of step S89 may be performed by the AAC encoding unit 22.
[0326]
 If it is determined in step S84 that it is within the allowable range, or if the process of step S89 is performed, then the process of step S90 is performed.
[0327]
 In step S90, the AAC encoding unit 22 extracts the encoded 3D Audio metadata, the intra-frame DSE division information, and the frame division information supplied from the 3D Audio metadata encoding unit 21 and the audio data supplied from the outside. AAC encode and generate a 3D AAC bitstream.
[0328]
 Specifically, the AAC encoding unit 22 stores the intra-frame DSE division information and the frame division information in the DSE of the AAC frame data according to the syntax shown in FIGS. 15 to 17, and also appropriately divides the code. 3D Audio metadata is also stored in DSE.
[0329]
 For example, when the 3D Audio metadata is transmitted over a plurality of frames, the AAC encoding unit 22 divides the 3D Audio metadata for the plurality of frames into DSEs of the plurality of frames in the divided transmission period. Stores each of the divided 3D Audio metadata obtained by. At that time, the AAC encoding unit 22 also appropriately stores the frame division information including the frame division mode information and the applicable frame index, which have appropriate values, and the intra-frame DSE division information in the DSE.
[0330]
 Further, the AAC encoding unit 22 encodes audio data supplied from the outside by the AAC encoding method and stores it as AAC audio element data in SCE, CPE, LFE, etc. of the AAC frame data. Furthermore, the AAC encoding unit 22 also stores 3D Audio component element information supplied from the outside and position information indicating the position of the audio object on the two-dimensional plane in the PCE of the AAC frame data.
[0331]
 When the AAC encoding unit 22 outputs (transmits) the 3D AAC bitstream including the AAC frame data for one frame or a plurality of frames obtained in this way, the process then proceeds to step S92.
[0332]
 If it is determined in step S82 that there is no 3D Audio metadata to be transmitted, then the process proceeds to step S91.
[0333]
 In step S91, the AAC encoding unit 22 AAC encodes the audio data supplied from the outside.
[0334]
 That is, the AAC encoding unit 22 encodes one frame of audio data supplied from the outside by the AAC encoding method and stores it in the AAC frame data. Further, the AAC encoding unit 22 also stores 3D Audio component element information supplied from the outside and position information indicating the position of the audio object on the two-dimensional plane in the PCE of the AAC frame data.
[0335]
 When the AAC encoding unit 22 outputs the 3D AAC bitstream including the AAC frame data for one frame obtained in this way, the process proceeds to step S92.
[0336]
 Although the example in which the previous frame reference flag is not used is described here, the previous frame reference flag may of course be used.
[0337]
 In such a case, the previous frame reference flag is stored in DSE in step S90. If it is determined in step S82 that there is no 3D Audio metadata to be transmitted, then the same processing as steps S14 to S16 in FIG. 10 is performed.
[0338]
 When the process of step S90 or the process of step S91 is performed, the process of step S92 is performed thereafter.
[0339]
 In step S92, the encoding device 11 determines whether or not there is audio data to be encoded.
[0340]
 If it is determined in step S92 that there is still audio data to be encoded, the process returns to step S81, and the above-described process is repeated.
[0341]
 On the other hand, if it is determined in step S92 that there is no audio data to be encoded, that is, all the audio data of the content to be transmitted has been encoded and output, the encoding process ends.
[0342]
 As described above, the encoding device 11 stores the 3D Audio metadata in the DSE and generates the 3D AAC bitstream. By doing so, the existing AAC standard can be extended, and 3D Audio playback can be performed using 3D Audio metadata on the decoding side. Thereby, convenience can be improved. In particular, when 3D Audio metadata cannot be stored in one frame of AAC frame data, the representative value and the end value are set as 3D Audio metadata and stored over multiple frames of AAC frame data, so that 3D audio data can be efficiently stored in 3D. Audio metadata can be transmitted.
[0343]

 Next, the decoding process performed by the decoding device 51 will be described with reference to the flowchart in FIG.
[0344]
 In step S131, the AAC decoding unit 61 performs AAC decoding processing on the received 3D AAC bitstream in a decoding method corresponding to the AAC encoding method.
[0345]
 In step S131, the same processing as step S41 in FIG. 11 is performed, the 3D Audio component element information obtained by decoding is supplied to the higher-level control unit, and the position information of each audio object on the two-dimensional plane is obtained. The audio data read out or obtained by decoding is supplied to the 3D Audio renderer 64.
[0346]
 In step S132, the AAC decoding unit 61 determines whether or not there is 3D Audio metadata in the DSE of the AAC frame data of the 3D AAC bitstream.
[0347]
 If it is determined in step S132 that there is 3D Audio metadata, then the process proceeds to step S133.
[0348]
 In step S133, the AAC decoding unit 61 determines whether the frame is divided based on the frame division mode information (frm_div_mode) stored in the DSE of the AAC frame data. For example, when the value of the frame division mode information is “1”, it is determined that the frame is divided, that is, one piece of 3D Audio metadata is recorded across AAC frame data of a plurality of frames.
[0349]
 When it is determined in step S133 that the frame is not divided, that is, when the value of the frame division mode information is “0”, the process proceeds to step S134.
[0350]
 In this case, the AAC decoding unit 61 reads the encoded 3D Audio metadata from the DSE of the AAC frame data for one frame and supplies it to the 3D Audio metadata decoding unit 62.
[0351]
 At this time, when 3D Audio metadata is divided and stored in a plurality of DSEs of AAC frame data for one frame, the AAC decoding unit 61 refers to the in-frame DSE division information of each DSE and Read the split 3D Audio metadata from the DSE. That is, each of the divided 3D Audio metadata obtained by dividing the 3D Audio metadata for one frame is read from each of the plurality of DSEs of the AAC frame data for one frame.
[0352]
 Then, the AAC decoding unit 61 combines the divided 3D Audio metadata read from each DSE into one encoded 3D Audio metadata, and supplies the encoded 3D Audio metadata to the 3D Audio metadata decoding unit 62.
[0353]
 In step S134, the 3D Audio metadata decoding unit 62 decodes the encoded 3D Audio metadata supplied from the AAC decoding unit 61, and the 3D Audio metadata for one frame obtained as a result is transferred to the 3D Audio renderer 64. Supply.
[0354]
 When the 3D Audio metadata is obtained in this way, the process then proceeds to step S140.
[0355]
 On the other hand, when it is determined that the frame is divided in step S133, the AAC decoding unit 61 acquires (reads) the divided 3D Audio metadata from the DSE of the AAC frame data for one frame in step S135.
[0356]
 Also in this case, if the divided 3D audio metadata is stored in a plurality of DSEs in one frame of AAC frame data, the intra-frame DSE division information is referred to and the 3D Audio divided from those DSEs is referred to. The metadata is read.
[0357]
 In step S136, the AAC decoding unit 61 combines the divided 3D Audio metadata obtained by the processing up to the first frame of the divided transmission period. In the combining process for combining the divided 3D audio metadata, the AAC decoding unit 61 appropriately performs the combining while referring to the frame division information read from the DSE.
[0358]
 In step S137, the AAC decoding unit 61 determines whether the frame processed in step S135 immediately before is the end frame of the divided transmission period. For example, when the value of the frame division termination flag (frm_div_terminate) of the frame processed in step S135 immediately before is “1”, the AAC decoding unit 61 determines that the frame is the termination frame.
[0359]
 If it is determined in step S137 that the frame is not the end frame, there is a frame in the divided transmission period in which the divided 3D Audio metadata has not been read yet, and therefore the process returns to step S135 and the above-described process is repeated. That is, the divided 3D Audio metadata is read for the next frame.
[0360]
 On the other hand, when it is determined in step S137 that the frame is the last frame, the coded 3D audio metadata for the divided transmission period is obtained by the process in step S136 performed immediately before, and therefore, the subsequent process is performed. Advances to step S138.
[0361]
 At this time, the AAC decoding unit 61 uses the encoded 3D Audio metadata regarding the divided transmission period obtained by combining, and the applicable frame read from the DSE division information within the DSE frame or the frame division information regarding the divided transmission period. Information such as an index (apply_frm_Index) is supplied to the 3D Audio metadata decoding unit 62. Accordingly, the 3D Audio metadata decoding unit 62 can specify how many frames the divided transmission period is composed of and which frame the representative value is applied to.
[0362]
 In step S138, the 3D Audio metadata decoding unit 62 decodes the encoded 3D Audio metadata of the divided transmission period supplied from the AAC decoding unit 61. As a result, the representative value and the termination value for the divided transmission period can be obtained.
[0363]
 In step S139, the 3D audio metadata decoding unit 62 determines the representative value and the termination value obtained in the process of step S138, the already obtained 3D audio metadata of the frame immediately before the divided transmission period, and the applicable frame index. Interpolation processing is performed based on and.
[0364]
 For example, in the interpolation process, as described with reference to FIG. 14, the 3D Audio metadata and the representative value of the frame immediately before the divided transmission period are used, and the representative value is applied from the first frame of the divided transmission period. 3D Audio metadata of each frame up to the frame immediately before the frame is calculated by interpolation.
[0365]
 Similarly, the representative value and the end value are used, and the 3D Audio metadata of each frame from the frame immediately after the frame to which the representative value is applied to the frame immediately before the end frame of the divided transmission period is It is calculated by interpolation.
[0366]
 In this case, the 3D Audio metadata decoding unit 62 functions as a metadata generation unit that performs interpolation processing and generates 3D Audio metadata of the frame in the divided transmission period.
[0367]
 By such processing, 3D Audio metadata of each frame in the divided transmission period is obtained. During the interpolation processing, information supplied from the AAC decoding unit 61 such as the applicable frame index is appropriately referred to.
[0368]
 When the 3D Audio metadata of each frame in the divided transmission period is obtained, the 3D Audio metadata decoding unit 62 supplies the 3D Audio metadata to the 3D Audio renderer 64, and then the process proceeds to step S140.
[0369]
 When the processing of step S134 or step S139 is performed and 3D Audio metadata of one or a plurality of frames is obtained, then the processing of step S140 is performed.
[0370]
 That is, in step S140, the 3D Audio renderer 64 performs rendering based on the 3D Audio metadata supplied from the 3D Audio metadata decoding unit 62 and the audio data supplied from the AAC decoding unit 61. In step S140, the same processing as in step S47 of FIG. 11 is performed.
[0371]
 When the reproduction data is obtained by the rendering, the 3D Audio renderer 64 outputs the obtained reproduction data, and then the process proceeds to step S142.
[0372]
 On the other hand, if it is determined in step S132 that there is no 3D Audio metadata, then the process of step S141 is performed.
[0373]
 That is, in step S141, the 3D audio renderer 64 performs rendering specified by the normal AAC standard on the basis of the audio data supplied from the AAC decoding unit 61, and outputs the reproduction data obtained as a result. In step S141, the same processing as in step S48 of FIG. 11 is performed.
[0374]
 When the reproduction data is output in this manner, the process then proceeds to step S142.
[0375]
 Although the example in which the previous frame reference flag is not used is described here, the previous frame reference flag may of course be used.
[0376]
 In such a case, if it is determined in step S132 that there is no 3D Audio metadata, then the same processing as steps S44 to S46 and step S48 of FIG. 11 is performed.
[0377]
 When the process of step S140 or step S141 is performed, in step S142, the AAC decoding unit 61 determines whether there is a 3D AAC bitstream to be decoded.
[0378]
 If it is determined in step S142 that there is a 3D AAC bitstream to be decoded, the process returns to step S131, and the above process is repeated.
[0379]
 On the other hand, if it is determined in step S142 that there is no 3D AAC bitstream to be decoded, the decoding process ends.
[0380]
 As described above, the decoding device 51 reads 3D Audio metadata from the DSE and performs rendering using the read 3D Audio metadata. By doing so, it is possible to extend the existing AAC standard and perform 3D audio playback, and improve convenience.
[0381]
 Especially when the 3D Audio metadata of the divided transmission period is stored across AAC frame data of multiple frames, interpolation processing is performed based on the representative value and the end value as the 3D Audio metadata, and the 3D Audio of each frame is By obtaining the metadata, 3D Audio can be played back efficiently even with a small code amount.
[0382]
 As described above, according to the present technology, 3D audio data can be transmitted within the standard of an existing audio codec having an ancillary data area without depending on the audio codec (encoding method).
[0383]
 When the size of the ancillary data area such as DSE is limited, 3D Audio metadata can be transmitted by dividing the 3D Audio metadata and storing it in multiple ancillary data areas. You can
[0384]
 Furthermore, by storing the position information indicating the position of the audio object on the two-dimensional plane, that is, the information about the channel mapping (speaker arrangement information) in the PCE or the like, and transmitting it, the existing decoding device that does not support 3D Audio can be used. Also, audio reproduction can be performed using position information as much as possible.
[0385]

 By the way, the series of processes described above can be executed by hardware or software. When the series of processes is executed by software, the programs forming the software are installed in the computer. Here, the computer includes a computer incorporated in dedicated hardware and, for example, a general-purpose personal computer capable of executing various functions by installing various programs.
[0386]
 FIG. 21 is a block diagram showing a configuration example of hardware of a computer that executes the series of processes described above by a program.
[0387]
 In a computer, a CPU (Central Processing Unit) 501, a ROM (Read Only Memory) 502, and a RAM (Random Access Memory) 503 are connected to each other by a bus 504.
[0388]
 An input/output interface 505 is further connected to the bus 504. An input unit 506, an output unit 507, a recording unit 508, a communication unit 509, and a drive 510 are connected to the input/output interface 505.
[0389]
 The input unit 506 includes a keyboard, a mouse, a microphone, an image sensor, and the like. The output unit 507 includes a display, a speaker and the like. The recording unit 508 includes a hard disk, a non-volatile memory, or the like. The communication unit 509 includes a network interface or the like. The drive 510 drives a removable recording medium 511 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.
[0390]
 In the computer configured as described above, the CPU 501 loads the program recorded in the recording unit 508 into the RAM 503 via the input/output interface 505 and the bus 504 and executes the program, thereby performing the above-described series of operations. Is processed.
[0391]
 The program executed by the computer (CPU 501) can be provided by being recorded in a removable recording medium 511 such as a package medium, for example. In addition, the program can be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting.
[0392]
 In the computer, the program can be installed in the recording unit 508 via the input/output interface 505 by mounting the removable recording medium 511 in the drive 510. Further, the program can be received by the communication unit 509 via a wired or wireless transmission medium and installed in the recording unit 508. In addition, the program can be installed in the ROM 502 or the recording unit 508 in advance.
[0393]
 It should be noted that the program executed by the computer may be a program in which processing is performed in time series in the order described in this specification, or in parallel or at a necessary timing such as when a call is made. It may be a program for processing.
[0394]
 The embodiments of the present technology are not limited to the above-described embodiments, and various modifications can be made without departing from the scope of the present technology.
[0395]
 For example, the present technology may have a configuration of cloud computing in which one device shares and jointly processes one function via a network.
[0396]
 In addition, each step described in the above-described flowcharts can be executed by one device or shared by a plurality of devices.
[0397]
 Further, when one step includes a plurality of processes, the plurality of processes included in the one step can be executed by one device or shared by a plurality of devices.
[0398]
 Furthermore, the present technology may have the following configurations.
[0399]
(1) A
 decoding unit that decodes audio data including object audio included in an encoded bitstream and reads out metadata of the object audio from an area that can store arbitrary data of the encoded bitstream. A  decoding device that outputs the decoded
 audio data based on the metadata
.
(2) The
 decoding unit reads each of a plurality of data obtained by dividing the metadata from each of the plurality of regions in the encoded bitstream for one frame, and outputs the plurality of data as the plurality of data.
 The decoding device according to (1), wherein the metadata is obtained based on the decoding device.
(3) The
 decoding unit reads each of a plurality of data obtained by dividing the metadata from each of the regions of a plurality of frames in the encoded bitstream, and based on the plurality of the data,
 The decoding device according to (1), wherein the metadata for the plurality of frames is obtained .
(4)
 the relative said plurality of frame metadata, the representative value of the metadata for the period consisting of the plurality of frames, and the termination value is metadata at the end frame of the period
 according to (3) Decryption device.
(5)
 The decoding unit further reads, from the area, applicable frame information indicating a frame that uses the representative value as the metadata
 , and performs interpolation processing based on the representative value, the termination value, and the applicable frame information.
 The decoding device according to (4), further including a metadata generation unit that generates the metadata of the frames within the period .
(6) The
 output unit decodes the predetermined frame based on the metadata read last in time among the metadata read from the area in a frame before the predetermined frame.
 The decoding device according to any one of (1) to (5), wherein the decoded audio data is output .
(7) The
 output unit is responsive to an instruction as to whether to use the last-read metadata, based on the last-read metadata, and outputs the decoded audio of the predetermined frame.
 The decoding device according to (6), which outputs data .
(8)
 When the metadata of the predetermined frame is not obtained, the output unit decodes the predetermined frame based on other metadata different from the metadata read from the area.
 The decoding device according to any one of (1) to (5), which outputs audio data .
(9) The
 encoded bitstream is an AAC bitstream,
 The said decoding part is
 a decoding apparatus as described in any one of (1) thru|or (8) which reads the said metadata from DSE as said area|region .
(10) The
 encoded bitstream is an AAC bitstream, and the
 decoding unit reads the configuration information of the audio data or the speaker arrangement information of the audio data from the PCE of the AAC bitstream
 (1) to ( 9. The decoding device according to any one of 9).
(11) The
 decoding device
 decodes the audio data including the object audio included in the encoded bitstream, and at the same time, extracts the metadata of the object audio from an area in which arbitrary data of the encoded bitstream can be stored. reading,
 on the basis of the metadata, and outputs the decoded the audio data is
 decoding method comprising the steps.
(12)
 Decode audio data including object audio included in the encoded bitstream, read metadata of the object audio from an area in the encoded bitstream that can store arbitrary data, and convert the
 metadata.
 A program that causes a computer to execute a process including a step of outputting the decoded audio data based on the data .
(13)
 Encoding bits that encode audio data including object audio, store metadata of the object audio in an area that can store arbitrary data, and include the metadata and the encoded audio data. An encoding
 device including an encoding unit that generates a stream .
(14)
 the encoding unit to each of the plurality of the regions in the coded bit stream of one frame, stores a plurality of data obtained by dividing the metadata
 described in (13) Encoding device.
(15) The
 encoding unit stores each of a plurality of data obtained by dividing the metadata for the plurality of frames in each of the regions of the plurality of frames in the encoded bitstream
 (13). ) Encoding device described in.
(16)  The code according to (15),
 further including a metadata determination unit that determines a representative value of the metadata for a period including the plurality of frames, and sets the representative value as the metadata for the plurality of frames.
Device.
(17)
 the metadata determination unit determines the termination value is metadata at the end frame of the period, the representative value and the end edge value and the metadata for the plurality of frames
 according to (16) Encoding device.
(18)

 The encoding device according  to (16) or (17), wherein the encoding unit stores applicable frame information indicating a frame using the representative value as the metadata in the area .
(19) The
 encoding unit may use, as the metadata of the predetermined frame, a frame closest to the predetermined frame among the metadata stored in the area of ​​the frame temporally preceding the predetermined frame.
 The encoding device according to any one of (13) to (18), wherein information that indicates whether to use the metadata is stored in the area .
(20) The
 encoded bitstream is an AAC bitstream, and the
 encoding unit stores the metadata in a DSE as the area
 (13) to (19). apparatus.
(21) The
 encoded bitstream is an AAC bitstream, and the
 encoding unit stores the configuration information of the audio data or the speaker arrangement information of the audio data in the PCE of the AAC bitstream
 (13). The encoding device according to any one of (1) to (20).
(22) The
 encoding device is
 Encoding bits that encode audio data including object audio, store metadata of the object audio in an area that can store arbitrary data, and include the metadata and the encoded audio data. An
 encoding method including the step of generating a stream .
(23)
 Audio data including object audio is encoded, and metadata of the object audio is stored in an area in which arbitrary data can be stored, and the metadata and the encoded audio data are included.
 A program that causes a computer to execute a process including a step of generating an encoded bitstream .
Explanation of symbols
[0400]
 11 encoding device, 21 3D Audio metadata encoding unit, 22 AAC encoding unit, 51 decoding device, 61 AAC decoding unit, 62 3D Audio metadata decoding unit, 63 3D Audio metadata acquisition unit, 64 3D Audio renderer
The scope of the claims
[Claim 1]
 Included in the encoded bit stream, with decoding the audio data including an object audio, a decoding unit for reading metadata of the object audio any data from the available space storing the encoded bit stream,
 said meta An output unit that outputs the decoded audio data based on data
 .
[Claim 2]
 The decoding unit reads each of a plurality of data obtained by dividing the metadata from each of the plurality of regions in the encoded bitstream for one frame, and based on the plurality of the data,
 The decoding device according to claim 1, wherein the decoding device obtains metadata .
[Claim 3]
 The decoding unit reads each of a plurality of data obtained by dividing the metadata from each of the regions of a plurality of frames in the encoded bitstream, and based on the plurality of the data, the plurality of data. The
 decoding device according to claim 1, wherein the metadata for the frame is obtained .
[Claim 4]

 The decoding device according to claim 3,  wherein the metadata for the plurality of frames is a representative value of the metadata for a period including the plurality of frames and a termination value that is the metadata in a termination frame of the period. ..
[Claim 5]
 The decoding unit further reads, from the area, applicable frame information indicating a frame that uses the representative value as the metadata
 , and performs interpolation processing based on the representative value, the termination value, and the applicable frame information.
 The decoding device according to claim 4, further comprising a metadata generation unit that generates the metadata of the frames within the period .
[Claim 6]
 The output unit decodes the predetermined frame based on the metadata read last in time among the metadata read from the area in a frame before a predetermined frame.
 The decoding device according to claim 1, which outputs audio data .
[Claim 7]
 The output unit outputs the decoded audio data of the predetermined frame based on the last-read metadata, in response to an instruction as to whether to use the last-read metadata.
 The decoding device according to claim 6.
[Claim 8]
 When the metadata of the predetermined frame is not obtained, the output unit outputs the decoded audio data of the predetermined frame based on other metadata different from the metadata read from the area.
 The decoding device according to claim 1, which outputs .
[Claim 9]
 The decoding device according to claim 1,  wherein the encoded bitstream is an AAC bitstream, and the
 decoding unit reads the metadata from a DSE as the area
.
[Claim 10]
 The decoding device according to claim 1,  wherein the encoded bitstream is an AAC bitstream, and the
 decoding unit reads configuration information of the audio data or speaker arrangement information of the audio data from a PCE of the AAC bitstream.
..
[Claim 11]
 Decoding apparatus,
 included in the encoded bit stream, with decoding the audio data including an object audio reads metadata of the object audio from any possible data storing area of the encoded bit stream,
 said
 A decoding method, comprising the step of outputting the decoded audio data based on metadata .
[Claim 12]
 Decoding audio data including object audio included in the coded bitstream, reading metadata of the object audio from an area in which any data of the coded bitstream can be stored,
 and based on the metadata And
 a program for causing a computer to execute processing including a step of outputting the decoded audio data .
[Claim 13]
 Encoding bits that encode audio data including object audio, store metadata of the object audio in an area that can store arbitrary data, and include the metadata and the encoded audio data. An encoding
 device including an encoding unit that generates a stream .
[Claim 14]
 The encoding according to
 claim 13, wherein the encoding unit stores each of the plurality of data obtained by dividing the metadata in each of the plurality of areas in the encoded bitstream for one frame . apparatus.
[Claim 15]
 The encoding unit stores each in each of the areas of the plurality of frames in the coded bit stream, the plurality of data obtained by dividing the metadata for the plurality of frames
 according to claim 13 Encoding device.
[Claim 16]

 The encoding device according to claim 15,  further comprising a metadata determination unit that determines a representative value of the metadata for a period including the plurality of frames, and sets the representative value as the metadata for the plurality of frames .
[Claim 17]

 The encoding according  to claim 16, wherein the metadata determination unit determines a termination value that is the metadata in the termination frame of the period, and sets the representative value and the termination value as the metadata for the plurality of frames . apparatus.
[Claim 18]

 The encoding device according  to claim 16, wherein the encoding unit stores applicable frame information indicating a frame using the representative value as the metadata in the area .
[Claim 19]
 The encoding unit, as the metadata of the predetermined frame, the metadata of the frame closest to the predetermined frame among the metadata stored in the area of ​​the frame temporally preceding the predetermined frame.
 The encoding device according to claim 13, wherein information indicating whether to use is stored in the area .
[Claim 20]
 The encoding device according
 to
claim 13  , wherein the encoded bitstream is an AAC bitstream, and the encoding unit stores the metadata in a DSE as the area .
[Claim 21]
 14. The encoded bitstream is an AAC bitstream, and the
 encoding unit stores the configuration information of the audio data or speaker arrangement information of the audio data in a PCE of the AAC bitstream
 . Encoding device.
[Claim 22]
 The encoding device encodes
 audio data including object audio, stores metadata of the object audio in an area in which arbitrary data can be stored, and the metadata and the encoded audio data are A coding
 method comprising the step of generating a coded bitstream included .
[Claim 23]
 Encoding bits that encode audio data including object audio, store metadata of the object audio in an area that can store arbitrary data, and include the metadata and the encoded audio data.
 A program that causes a computer to execute a process including a step of generating

Documents

Application Documents

# Name Date
1 202017013466-TRANSLATIOIN OF PRIOIRTY DOCUMENTS ETC. [27-03-2020(online)].pdf 2020-03-27
2 202017013466-STATEMENT OF UNDERTAKING (FORM 3) [27-03-2020(online)].pdf 2020-03-27
3 202017013466-PROOF OF RIGHT [27-03-2020(online)].pdf 2020-03-27
4 202017013466-PRIORITY DOCUMENTS [27-03-2020(online)].pdf 2020-03-27
5 202017013466-POWER OF AUTHORITY [27-03-2020(online)].pdf 2020-03-27
6 202017013466-FORM 1 [27-03-2020(online)].pdf 2020-03-27
7 202017013466-DRAWINGS [27-03-2020(online)].pdf 2020-03-27
8 202017013466-DECLARATION OF INVENTORSHIP (FORM 5) [27-03-2020(online)].pdf 2020-03-27
9 202017013466-COMPLETE SPECIFICATION [27-03-2020(online)].pdf 2020-03-27
10 202017013466-FORM 3 [25-06-2020(online)].pdf 2020-06-25
11 202017013466-FORM 18 [17-08-2021(online)].pdf 2021-08-17
12 202017013466.pdf 2021-10-19
13 202017013466-FER.pdf 2022-03-11
14 202017013466-OTHERS [12-09-2022(online)].pdf 2022-09-12
15 202017013466-FER_SER_REPLY [12-09-2022(online)].pdf 2022-09-12
16 202017013466-CORRESPONDENCE [12-09-2022(online)].pdf 2022-09-12
17 202017013466-CLAIMS [12-09-2022(online)].pdf 2022-09-12
18 202017013466-PatentCertificate03-01-2024.pdf 2024-01-03
19 202017013466-IntimationOfGrant03-01-2024.pdf 2024-01-03

Search Strategy

1 Search_Strategy_202017013466E_11-03-2022.pdf

ERegister / Renewals

3rd: 22 Mar 2024

From 21/09/2020 - To 21/09/2021

4th: 22 Mar 2024

From 21/09/2021 - To 21/09/2022

5th: 22 Mar 2024

From 21/09/2022 - To 21/09/2023

6th: 22 Mar 2024

From 21/09/2023 - To 21/09/2024

7th: 18 Sep 2024

From 21/09/2024 - To 21/09/2025

8th: 16 Sep 2025

From 21/09/2025 - To 21/09/2026