Sign In to Follow Application
View All Documents & Correspondence

Information Processing Device, Information Processing Method And Program

Abstract: [Problem] To enable a priority to be set for each piece of acquired object audio data. [Solution] Provided is an information processing device provided with a processing unit which sets a priority for each piece of acquired object audio data, determines, from one or two or more pieces of the object audio data, the object audio data to be included in a generated segment file, on the basis of the priority, and generates, as priority information, a new priority to be set for the generated segment file, on the basis of the priority.

Get Free WhatsApp Updates!
Notices, Deadlines & Correspondence

Patent Information

Application #
Filing Date
19 June 2020
Publication Number
40/2020
Publication Type
INA
Invention Field
MECHANICAL ENGINEERING
Status
Email
mahua.ray@remfry.com
Parent Application

Applicants

SONY CORPORATION
1-7-1, Konan, Minato-ku, Tokyo 1080075

Inventors

1. KATSUMATA, Mitsuru
c/o SONY CORPORATION, 1-7-1, Konan, Minato-ku, Tokyo 1080075
2. HIRABAYASHI, Mitsuhiro
c/o SONY CORPORATION, 1-7-1, Konan, Minato-ku, Tokyo 1080075
3. HAMADA, Toshiya
c/o SONY CORPORATION, 1-7-1, Konan, Minato-ku, Tokyo 1080075

Specification

Specification
Title of invention: Information processing device, information processing method and program
Technical field
[0001]
 The present disclosure relates to information processing devices, information processing methods and programs.
Background technology
[0002]
 In recent years, the mainstream of streaming services on the Internet has become OTT-V (Over The Top Video). MPEG-DASH (Moving Picture Experts Group phase
-Dynamic Adaptive Streaming over HTTP) is beginning to spread as this basic technology (see, for example, Non-Patent Document 1).
[0003]
 In the distribution of audio content using MPEG-DASH, the distribution server prepares audio data for each object (the data is called "object audio data"), and the client responds to the conditions of the transmission line, etc. Adaptive streaming distribution is realized by requesting the optimum object audio data group.
Prior art literature
Non-patent literature
[0004]
Non-Patent Document 1: MPEG-DASH (Dynamic Adaptive Streaming over HTTP) (URL: http://standards.iso.org/ittf/PubliclyAvailableStandards/index.html)
Outline of the invention
Problems to be solved by the invention
[0005]
 However, in the MPEG-DASH standard shown in Non-Patent Document 1, it is not possible to set the priority for each object audio data. Therefore, the client must acquire the audio file once, in other words, acquire the entire file including the data that should not be needed, and then parse the object audio metadata of the audio file, which makes the bandwidth effective. There was a problem that it could not be utilized and that the client had a processing overhead.
[0006]
 Therefore, this disclosure has been made in view of the above problems, and the purpose of this disclosure is new and improved information that can set a priority for each acquired object audio data. To provide processing equipment, information processing methods and programs.
Means to solve problems
[0007]
 According to the present disclosure, a priority is set for each acquired object audio data, and based on the priority, the object audio data to be included in a segment file generated from one or more of the object audio data is included. An information processing apparatus is provided that includes a processing unit that determines and generates new priority as priority information, which is determined and set for the generated segment file based on the priority.
[0008]
 Further, the present disclosure provides an information processing method executed by a computer, which has information processing of stream data in which priority is set for each acquired object audio data.
[0009]
 Further, according to the present disclosure, a program is provided for causing a computer to process stream data in which a priority is set for each acquired object audio data.
Effect of the invention
[0010]
 As described above, according to the present disclosure, it is possible to set a priority for each acquired object audio data.
[0011]
 It should be noted that the above effects are not necessarily limited, and either in combination with or in place of the above effects, any of the effects shown herein, or any other effect that can be grasped from this specification. May be played.
A brief description of the drawing
[0012]
[Fig. 1] Fig. 1 is a diagram illustrating the background of the present disclosure.
FIG. 2 is a diagram illustrating the background of the present disclosure.
FIG. 3 is a diagram illustrating the background of the present disclosure.
FIG. 4 is a diagram illustrating the background of the present disclosure.
FIG. 5 is a diagram illustrating the background of the present disclosure.
FIG. 6 is a diagram showing a system configuration example of an information processing system according to the present embodiment.
FIG. 7 is a block diagram showing a functional configuration example of the server 100 according to the present embodiment.
FIG. 8 is a block diagram showing a functional configuration example of the client 200 according to the present embodiment.
[Fig. 9] Fig. 9 is a diagram illustrating an example of generating an audio file based on priority.
[Fig. 10] Fig. 10 is a diagram illustrating an example of generating an audio file based on priority.
FIG. 11 is a diagram illustrating an example of generating an audio file based on priority.
[Fig. 12] Fig. 12 is a diagram illustrating an example of generating an audio file based on priority.
[Fig. 13] Fig. 13 is a diagram illustrating an example of signaling priority information when the priority does not change with the passage of time.
[Fig. 14] Fig. 14 is a diagram illustrating a file structure when the priority changes with the passage of time.
[Fig. 15] Fig. 15 is a diagram illustrating ISO BMFF of an audio file in MPEG-H 3D Audio.
[Fig. 16] Fig. 16 is a diagram for explaining the Box structure of ISOBMFF.
FIG. 17 is a diagram illustrating ISO BMFF (Example 1, RAW method) of a metadata file in MPEG-H 3D Audio.
FIG. 18 is a diagram illustrating ISO BMFF (Example 1, MHAS method) of a metadata file in MPEG-H 3D Audio.
FIG. 19 is a diagram illustrating ISO BMFF (Example 2, RAW method) of a metadata file in MPEG-H 3D Audio.
FIG. 20 is a diagram illustrating ISOBMFF (Example 2, MHAS method) of a metadata file in MPEG-H 3D Audio.
[Fig. 21] Fig. 21 is a diagram illustrating ISO BMFF of an audio file in AAC 3D Audio.
FIG. 22 is a diagram illustrating ISOBMFF (Example 3) of a metadata file in AAC 3D Audio.
FIG. 23 is a diagram illustrating ISO BMFF (Example 4) of a metadata file in AAC 3D Audio.
[Fig. 24] Fig. 24 is a diagram for explaining an example of associating an audio file and a metadata file.
FIG. 25 is a diagram illustrating an example of associating an audio file and a metadata file (Example 1).
FIG. 26 is a diagram illustrating an example of associating an audio file and a metadata file (Example 2).
FIG. 27 is a diagram illustrating an example of associating an audio file with a metadata file (Example 3).
FIG. 28 is a diagram illustrating an example of associating an audio file and a metadata file (Example 4).
FIG. 29 is a diagram illustrating an example of associating an audio file with a metadata file (Example 5).
FIG. 30 is a diagram illustrating a signaling example (Example 1) in the case where the bit rate does not change with the passage of time.
FIG. 31 is a diagram illustrating a signaling example (Example 2) in the case where the bit rate does not change with the passage of time.
FIG. 32 is a diagram illustrating a signaling example (Example 3) in the case where the bit rate does not change with the passage of time.
FIG. 33 is a diagram illustrating an example of signaling (Example 4) in the case where the bit rate does not change with the passage of time.
[Fig. 34] Fig. 34 is a diagram for explaining an example of signaling when the bit rate changes with the passage of time.
FIG. 35 is a diagram illustrating an example of signaling (Example 5) when the bit rate changes with the passage of time.
FIG. 36 is a diagram illustrating a signaling example (Example 6) in which the bit rate changes with the passage of time.
FIG. 37 is a diagram illustrating an example of signaling (Example 7) when the bit rate changes with the passage of time.
FIG. 38 is a diagram illustrating an example of signaling (Example 8) when the bit rate changes with the passage of time.
[Fig. 39] Fig. 39 is a diagram illustrating an example of signaling of description information.
FIG. 40 is a flowchart showing a processing example until the client 200 acquires an audio file used for reproducing audio content when the priority does not change with the passage of time.
FIG. 41 is a diagram for explaining a processing example until the client 200 acquires an audio file used for playing back audio content when the priority does not change with the passage of time.
FIG. 42 is a flowchart showing an example of processing until the client 200 acquires an audio file used for reproducing audio content when the priority changes with the passage of time.
FIG. 43 is a diagram for explaining a processing example until the client 200 acquires an audio file used for playing back audio content when the priority changes with the passage of time.
FIG. 44 is a block diagram showing a hardware configuration example of an information processing device 900 embodying a server 100 or a client 200.
[Fig. 45] Fig. 45 is a diagram showing the structure of 3da_meta_data ().
[Fig. 46] Fig. 46 is a diagram showing the structure of 3da_meta_data () stored in DSE.
[Fig. 47] Fig. 47 is a diagram showing the structure of DSE.
[Fig. 48] Fig. 48 is a diagram showing the structure of 3da_ancillary_data stored in data_stream_byte in DSE.
Mode for carrying out the invention
[0013]
 Preferred embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. In the present specification and the drawings, components having substantially the same functional configuration are designated by the same reference numerals, so that duplicate description will be omitted.
[0014]
 The explanations will be given in the following order.
 1. 1. Background
 2. Configuration example
 3. File generation example based on priority
 4. Signaling example of priority information
 5. Bit rate information signaling example
 6. Example of signaling of description information
 7. Processing example of client 200
 8. Hardware configuration example
[0015]
  <1. Background>
 First, the background of the present disclosure will be described.
[0016]
 MPEG-H 3D Audio and AAC 3D Audio are standards that can handle object audio data, which is audio data generated for each object. The audio content is composed of a plurality of object audio data which are waveform data of a sound source, and object metadata including information on the position of an object, the spread of sound, various effects, and the like.
[0017]
 For example, as shown in FIG. 1, object metadata and a plurality of object audio data (in FIG. 1, object audio data 1 to object audio data n are shown) are provided to a client by a server or the like. When the client functioning as an object renderer receives object metadata and object audio data, it renders based on the playback environment information (for example, the position or number of speakers), and waveform data for the playback environment such as speakers. By providing, the playback of audio content is realized.
[0018]
 Here, in MPEG-H 3D Audio and AAC 3D Audio, not all object audio data need to be rendered. This can be done, for example, by the server not providing the object audio data itself that is not the target of rendering to the client, or excluding the object audio data that is not the target of rendering from the object metadata.
[0019]
 Further, in these standards, a plurality of object audio data may be reproduced at different bit rates. For example, as shown in 2A of FIG. 2, it is assumed that the audio content has object metadata and object audio data 1 to object audio data 3 for which high bit rates and low bit rates are prepared, respectively. In this case, the combination of bit rates of the reproduced object audio data is free. For example, as shown in 2B, the object audio data 1 and the object audio data 2 may be reproduced at a high bit rate, and the object audio data 3 may be reproduced at a low bit rate.
[0020]
 When audio content is provided, the object audio data is stored in an audio file and transmitted to the client. Here, a specific example will be described with reference to FIG. As shown in 3A of FIG. 3, when object audio data 1 to object audio data 3 and object metadata corresponding to these data exist, the mode of storing in the audio file is shown in 3B-1 to 3B-3. Multiple combinations are possible.
[0021]
 More specifically, as shown in 3B-1, all the data may be stored in one audio file, or as shown in 3B-3, one object audio data and the corresponding object metadata. May be stored in one audio file. Further, as shown in 3B-2, one or more object audio data and the corresponding object metadata may be stored in the audio file, respectively.
[0022]
 By the way, when the audio content is provided by MPEG-DASH, audio files having different bit rates are generated, and the client can select a desired audio file from these audio files.
[0023]
 For example, as shown in FIG. 4, it is assumed that object audio data 1 to object audio data 3 having bit rates of 64 [kbps] and 32 [kbps] are generated, respectively. In this case, the combination of a client can get audio files 2 3 exist ways. For example, a combination of file 1-1, file 2-1 and file 3-2 (total 160 [kbps]) or a combination of file 1-1, file 2-2 and file 3-1 (total 160 [kbps]). ]), There are combinations in which the total bit rate is the same.
[0024]
 However, the client cannot determine which combination is more appropriate based on the information about the total bit rate alone. Therefore, it may be considered to set the priority as information indicating which object audio data should be reproduced with higher sound quality (in other words, higher bit rate) so that the client can determine a more appropriate combination. ..
[0025]
 As one of the object metadata that can be used as priority information, there is "Priority" in MPEG-H 3D Audio or the like. However, since Priority is data stored in the object metadata, when Priority is applied to MPEG-DASH, the client once obtains the audio file and then obtains the Priority from the object metadata to obtain the audio file. Will determine the bit rate of. In other words, the client will have to retrieve unnecessary data.
[0026]
 Therefore, in view of the above circumstances, the disclosing party of this case has come to create the technology related to this disclosure. The present disclosure sets a priority for each acquired object audio data, enables efficient acquisition, and can realize appropriate distribution and reproduction of audio content based on the priority.
[0027]
 Further, in MPEG-DASH, the client basically determines the audio file to be acquired based on the bit rate of the audio file. Therefore, in the absence of the priority information, for example, when four types of audio files are generated by combining the bit rates of the object audio data 1 and the object audio data 2 as shown in FIG. 5, the file 2 and the file 2 are generated. Since the bit rates of files 3 are both 96 [kbps], the client cannot determine which audio file is more appropriate based on the bit rate information alone, and the content author intended. It is not possible to provide the client with a combination of object audio data and its bit rate.
[0028]
 On the other hand, the present disclosure can provide a client with bit rate information for each object audio data stored in an audio file. As a result, the client can determine which of the file 2 and file 3 audio files is more appropriate by also considering the above priority information.
[0029]
 In addition, the present disclosure can provide the description information of the object audio data to the client. As a result, the user who operates the client can reproduce the desired object audio data at a high bit rate.
[0030]
 Hereinafter, one embodiment of the present disclosure will be described in more detail.
[0031]
  <2. Configuration Example>
 The background of the present disclosure has been described above. Subsequently, a configuration example of the information processing system according to the embodiment of the present disclosure will be described with reference to FIGS. 6 to 8.
[0032]
 (2-1. System Configuration Example)
 First, a system configuration example of the information processing system according to the present embodiment will be described with reference to FIG.
[0033]
 As shown in FIG. 6, the information processing system according to the present embodiment includes a server 100 and a client 200. The server 100 and the client 200 are connected to each other by the Internet 300.
[0034]
 The server 100 is an information processing device (transmission device) that distributes (streams) object audio data used for audio content to the client 200 based on MPEG-DASH. More specifically, the server 100 acquires audio data used for audio content for each object, and generates stream data by encoding the data for each object. Then, the server 100 generates an audio file by converting the stream data into a file every time unit of several seconds to 10 seconds, which is called a segment, or for all the contents.
[0035]
 The object is a sound source, and the audio data of each object is acquired by a microphone or the like attached to the object. The object may be an object such as a fixed microphone stand, or a moving object such as a person.
[0036]
 Further, the server 100 encodes object metadata including position information of each object. The server 100 generates a metadata file by filing the encoded data of the object metadata in segment units.
[0037]
 Further, the server 100 generates an MPD (Media Presentation Description) file (control information) that manages the audio file.
[0038]
 Then, the server 100 transmits the above audio file, metadata file, MPD file, or the like to the client 200 in response to a request from the client 200.
[0039]
 The client 200 is an information processing device (receiver) that reproduces audio content. More specifically, the client 200 acquires an MPD file from the server 100, and acquires a metadata file and an audio file from the server 100 based on the MPD file. Then, the client 200 decodes the audio file acquired from the server 100, synthesizes the audio file, and outputs the file to reproduce the audio content.
[0040]
 The configuration example of the information processing system according to the present embodiment has been described above. The above configuration described with reference to FIG. 6 is merely an example, and the configuration of the information processing system according to the present embodiment is not limited to such an example. For example, some of the functions of the server 100 may be provided in the client 200 or other external device. For example, software that provides a part of the functions of the server 100 (for example, a WEB application using a predetermined API (Application Programming Interface)) may be executed on the client 200. On the contrary, a part of the functions of the client 200 may be provided in the server 100 or other external device. The configuration of the information processing system according to the present embodiment can be flexibly modified according to specifications and operations.
[0041]
 (2-2. Functional Configuration Example of Server 100) In the
 above, a system configuration example of the information processing system according to the present embodiment has been described. Subsequently, a functional configuration example of the server 100 will be described with reference to FIG. 7.
[0042]
 As shown in FIG. 7, the server 100 includes a processing unit 110, a control unit 120, a communication unit 130, and a storage unit 140.
[0043]
 The processing unit 110 has a functional configuration for performing processing related to the provision of audio content. As shown in FIG. 7, the processing unit 110 includes a data acquisition unit 111, a coding processing unit 112, a segment file generation unit 113, and an MPD file generation unit 114.
[0044]
 The data acquisition unit 111 has a functional configuration for acquiring audio data used for audio content for each object (in other words, acquiring object audio data). The data acquisition unit 111 may acquire the object audio data from the server 100, or may acquire the object audio data from an external device connected to the server 100. Further, the data acquisition unit 111 may set a priority for each income object audio data, or may acquire object audio data for which the priority has been set in advance. The data acquisition unit 111 provides the acquired object audio data to the coding processing unit 112.
[0045]
 The coding processing unit 112 has a functional configuration for generating stream data by encoding the object audio data provided by the data acquisition unit 111 for each object. Further, the coding processing unit 112 encodes the object metadata including the object position information of each object input from the outside. The coding processing unit 112 provides the stream data of each object and the coded data of the object metadata to the segment file generation unit 113.
[0046]
 The segment file generation unit 113 has a functional configuration for generating a segment file which is unit data that can be distributed as audio content. More specifically, the segment file generation unit 113 generates an audio file by converting the stream data of each object provided by the coding processing unit 112 into a file in segment units. There are various aspects of generating audio files. For example, the segment file generation unit 113 generates an audio file by storing one or more object audio data in one audio file based on the priority of each object audio data. Further, the segment file generation unit 113 can also set priority information for each file. Details of priority-based audio file generation will be described later.
[0047]
 Further, the segment file generation unit 113 generates a metadata file that stores only the object metadata by converting the coded data of the object metadata provided by the coding processing unit 112 into a file for each segment. You can also. The contents of the metadata file and the case where the metadata file is generated will be described later.
[0048]
 The MPD file generation unit 114 has a functional configuration for generating an MPD file. In the present embodiment, the MPD file generation unit 114 generates an MPD file including priority information, bit rate information, or description information. The details of the contents of the MPD file will be described later.
[0049]
 The control unit 120 has a functional configuration that comprehensively controls all the processing performed by the server 100. For example, the control unit 120 can control the start and stop of each configuration based on the request information from the client 200 received via the communication unit 130. The control content of the control unit 120 is not particularly limited. For example, the control unit 120 may control processing generally performed in a general-purpose computer, a PC, a tablet PC, or the like.
[0050]
 The communication unit 130 has a functional configuration for performing various communications with the client 200 (also functions as a transmission unit). For example, the communication unit 130 receives the request information from the client 200, or transmits an MPD file, a metadata file, an audio file, or the like to the client 200 as a response to the request information. The communication content of the communication unit 130 is not limited to these.
[0051]
 The storage unit 140 has a functional configuration for storing various types of information. For example, the storage unit 140 stores object audio data, audio files, object metadata, metadata files, MPD files, and the like, and stores programs or parameters used by each functional configuration of the server 100. The information stored in the storage unit 140 is not limited to these.
[0052]
 The functional configuration example of the server 100 has been described above. The above-mentioned functional configuration described with reference to FIG. 7 is merely an example, and the functional configuration of the server 100 is not limited to such an example. For example, the server 100 does not necessarily have all of the functional configurations shown in FIG. 7. Further, the functional configuration of the server 100 can be flexibly modified according to the specifications and operation.
[0053]
 (2-3. Example of Functional Configuration of Client 200) In the
 above, an example of functional configuration of the server 100 has been described. Subsequently, a functional configuration example of the client 200 will be described with reference to FIG.
[0054]
 As shown in FIG. 8, the client 200 includes a processing unit 210, a control unit 220, a communication unit 230, and a storage unit 240.
[0055]
 The processing unit 210 has a functional configuration for performing processing related to reproduction of audio content. As shown in FIG. 8, the processing unit 210 includes an MPD file acquisition unit 211, an MPD file processing unit 212, a segment file selection unit 213, a segment file acquisition unit 214, a decoding processing unit 215, and a synthesis processing unit 216. And.
[0056]
 The MPD file acquisition unit 211 has a functional configuration for acquiring an MPD file from the server 100 prior to reproducing the audio content. More specifically, the MPD file acquisition unit 211 generates the MPD file request information based on the user operation or the like, and provides the request information to the server 100 via the communication unit 230 to provide the MPD file to the server. Get from 100. The MPD file acquisition unit 211 provides the acquired MPD file to the MPD file processing unit 212.
[0057]
 The MPD file processing unit 212 has a functional configuration for performing processing related to the MPD file provided by the MPD file acquisition unit 211. More specifically, the MPD file processing unit 212 recognizes the information (for example, URL, etc.) necessary for acquiring the audio file or the corresponding metadata file by analyzing the MPD file, and each object audio. Recognize the priority or bit rate of data (or audio file). The MPD file processing unit 212 provides such information to the segment file selection unit 213.
[0058]
 The segment file selection unit 213 has a functional configuration for selecting a segment file to be acquired. More specifically, the segment file selection unit 213 selects an audio file or a metadata file to be acquired based on the above-mentioned various information provided by the MPD file processing unit 212. For example, the segment file selection unit 213 selects the audio file to be acquired so that the bit rate of the audio file having a higher priority is higher. At that time, the segment file selection unit 213 also considers the available transmission band, and selects, for example, the audio file having the highest bit rate in the available transmission band.
[0059]
 The method of selecting the file to be acquired is not limited to the above. For example, the segment file selection unit 213 may select an audio file desired by the user as an acquisition target based on the input from the user. The segment file selection unit 213 provides the segment file acquisition unit 214 with information about the file to be acquired.
[0060]
 The segment file acquisition unit 214 has a functional configuration for acquiring a segment file. More specifically, the segment file acquisition unit 214 generates request information of an audio file or a metadata file based on various information provided from the segment file selection unit 213, and the request information is transmitted via the communication unit 230. By providing to the server 100, these files are acquired from the server 100. The segment file acquisition unit 214 provides these acquired files to the decoding processing unit 215.
[0061]
 The decoding processing unit 215 has a functional configuration for decoding data included in an audio file or a metadata file provided by the segment file acquisition unit 214. The decoding processing unit 215 provides the object audio data or the like obtained by the decoding processing to the synthesis processing unit 216.
[0062]
 The compositing processing unit 216 is a functional configuration that synthesizes and outputs a plurality of object audio data provided by the decoding processing unit 215. The synthesis processing unit 216 provides the combined data to the control unit 220.
[0063]
 The control unit 220 has a functional configuration that comprehensively controls the overall processing performed by the client 200. For example, the control unit 220 controls an output unit (not shown) such as a display or a speaker, and outputs audio content to the user by outputting the combined data provided by the composition processing unit 216. Further, the control unit 220 controls various processes based on the input performed by the user using an input unit (not shown) such as a mouse and a keyboard. The control content of the control unit 220 is not particularly limited. For example, the control unit 220 may control processing generally performed in a general-purpose computer, a PC, a tablet PC, or the like.
[0064]
 The communication unit 230 has a functional configuration for performing various communications with the server 100 (also functions as a receiving unit). For example, the communication unit 230 transmits request information to the server 100 based on user input or the like, or receives an MPD file, a metadata file, an audio file, or the like from the server 100 as a response to the request information. The communication content of the communication unit 230 is not limited to these.
[0065]
 The storage unit 240 has a functional configuration for storing various types of information. For example, the storage unit 240 stores object audio data, audio files, object metadata, metadata files, MPD files, etc. provided by the server 100, programs or parameters used by each functional configuration of the client 200, and the like. To remember. The information stored in the storage unit 240 is not limited to these.
[0066]
 The functional configuration example of the client 200 has been described above. The above-mentioned functional configuration described with reference to FIG. 8 is merely an example, and the functional configuration of the client 200 is not limited to such an example. For example, the client 200 does not necessarily have all of the functional configurations shown in FIG. Further, the functional configuration of the client 200 can be flexibly modified according to the specifications and operation.
[0067]
  <3. Audio file generation example based on priority> In the
 above, a functional configuration example of the client 200 has been described. Next, an example of generating an audio file based on the priority will be described.
[0068]
 As described above, the segment file generation unit 113 of the server 100 generates an audio file by combining one or more object audio data and storing them in one audio file by using the priority information for each object audio data. To do. Further, the segment file generation unit 113 can also set priority information for each file.
[0069]
 Here, an example of generating an audio file based on the priority will be described with reference to FIGS. 9 to 12. For example, as shown in 9A of FIG. 9, it is assumed that object audio data 1 to object audio data 4 exist and their priorities are set in advance. More specifically, it is assumed that the priority of the object audio data 1 is 3, the priority of the object audio data 2 and the object audio data 3 is 2, and the priority of the object audio data 4 is 1. , The higher the value, the higher the priority).
[0070]
 In this case, the segment file generation unit 113 may combine object audio data whose priorities are closer to each other (the difference in priority is equal to or less than a predetermined value) and store them in one audio file. For example, as shown in 9B, the segment file generation unit 113 is closer to the object audio data 1 which is the first data in the series of data and the priority 3 of the data (for example, the difference in priority is 1 or less). The object audio data 2 and the object audio data 3 having the priority 2 may be combined and stored in one audio file. Then, the segment file generation unit 113 may store the remaining object audio data 4 in another audio file.
[0071]
 Then, the segment file generation unit 113 may set the highest priority among the priorities of the object audio data stored in one audio file as the priority of the audio file. For example, as shown in 9B, the segment file generation unit 113 sets the highest priority 3 among the priorities of the object audio data 1 to the object audio data 3 as the priority of the audio file in which these data are stored. It may be set. The method of setting the priority of the audio file is not limited to this. For example, the segment file generation unit 113 sets the priority of the object audio data stored in one audio file with the largest number of data having the same priority as the priority of the audio file. You may. Further, the segment file generation unit 113 may set the average value of the priorities of the object audio data stored in one audio file as the priority of the audio file.
[0072]
 Further, as shown in 10A of FIG. 10, when the priority is not set for the object audio data 1 to the object audio data 4 or all the priorities are the same, a segment file is generated as shown in 10B. Unit 113 may store all object audio data in the same audio file. Then, the segment file generation unit 113 may not set the priority of the audio file, or may set the same priority set for each object audio data as the priority of the audio file.
[0073]
 Further, as shown in 11A of FIG. 11, when the priorities of the object audio data 1 to the object audio data 4 are different, the segment file generation unit 113 is different in each object audio data as shown in 11B. It may be stored in an audio file. Then, the segment file generation unit 113 may set the same value as the priority of each object audio data as the priority of each audio file.
[0074]
 Further, as shown in 12A of FIG. 12, it is assumed that the object audio data 1 to the object audio data 3 have a priority of 2 and the object audio data 4 has a priority of 1. In this case, as shown in 12B, the segment file generation unit 113 stores the object audio data 1 to the object audio data 3 having the same priority in one audio file, and stores the object audio data 4 having different priorities in one audio file. It may be stored in another audio file. Also in 12B, the same value as the priority of each object audio data is set as the priority of each audio file.
[0075]
 Here, in MPEG-DASH, the client 200 performs acquisition control for each audio file. Therefore, as shown in 9B of FIG. 9, 10B of FIG. 10, and 12B of FIG. 12, the segment file generation unit 113 stores a plurality of object audio data in one audio file, so that the content creator can obtain the content creator. It is possible to control the combination of object audio data and its bit rate. In other words, the content creator can provide the audio content with the intended combination of object audio data and bit rate. On the other hand, the degree of freedom for the user to acquire the object audio data is low.
[0076]
 On the other hand, as shown in 11B of FIG. 11, when one object audio data is stored in one audio file, the user can acquire only the desired object audio data, so that the object audio The degree of freedom in acquiring data is increased. On the other hand, it is difficult for the content creator to prevent the audio content from being played back with an unintended combination of the object audio data and the bit rate. As described above, the content creator can adjust how much freedom the user has in acquiring the object audio data according to the mode of storing the object audio data in the audio file.
[0077]
 The method of generating the audio file based on the priority or the method of setting the priority of the audio file is not limited to the above, and may be changed as appropriate.
[0078]
  <4. Signaling example of priority information> In the
 above, an example of generating an audio file based on priority has been described. Next, an example of signaling priority information will be described.
[0079]
 (4-1. Signaling example when the
 priority does not change with the passage of time ) First, a signaling example when the priority does not change with the passage of time will be described.
[0080]
 Audio files with different bit rates that store the same object audio data have the same priority. Therefore, signaling by the Adapter Set of the MPD file is appropriate. More specifically, the present disclosure newly defines schemaIdUri indicating that it is priority information by using SupplementalProperty, and the MPD file generation unit 114 of the server 100 stores the priority information in value.
[0081]
 Here, an example of signaling priority information will be described with reference to FIG. As shown in 13A of FIG. 13, consider the case where there is a difference in the bit rate of the audio file in which the object audio data 1 and the object audio data 2 are stored and a difference in the bit rate of the audio file in which the object audio data 3 is stored. (In the figure, the object audio data is referred to as "obj" for convenience). Then, it is assumed that the priority of the audio file in which the object audio data 1 and the object audio data 2 are stored is 2, and the priority of the audio file in which the object audio data 3 is stored is 1.
[0082]
 In this case, the MPD file generation unit 114 generates an MPD file as shown in 13B. More specifically, in the preselection Components of Preselection, the Adaptation Set containing the object audio data to be simultaneously played in the audio content is listed. The first AdaptationSet is an AdaptationSet of an audio file in which the object audio data 1 and the object audio data 2 are stored (the id of the AdaptationSet is o1). The MPD file generation unit 114 sets a Supplemental Property indicating that the priority is 2 in the Adaptation Set.
[0083]
 The second AdaptationSet is an AdaptationSet of an audio file in which the object audio data 3 is stored (the id of the AdaptationSet is o2). The MPD file generation unit 114 sets a Supplemental Property indicating that the priority is 1 in the Adaptation Set. As a result, the MPD file processing unit 212 of the client 200
can grasp the priority of each audio file based on the MPD file.
[0084]
 The above is just an example, and the signaling method is not limited to the above. For example, objectAcquisitionPriority may be added as an Attribute of AdaptationSet instead of using SupplementalProperty. More specifically, the MPD file generation unit 114 of the server 100 does not use SupplementalProperty in the MPD file, but "" and "".
[0085]
 (4-2. Signaling example when the
 priority changes with the passage of time ) In the above, an example of signaling when the priority does not change with the passage of time has been described. Next, an example of signaling when the priority changes with the passage of time will be described.
[0086]
 If the priority changes over time, the object metadata changes over time. Therefore, if it is described in the MovieBox area of ​​the MPD or ISOBMFF file, it is not appropriate from the viewpoint of increasing the processing load of the server and the processing overhead of the client. Therefore, the segment file generation unit 113 of the server 100 generates a metadata file that stores only the object metadata corresponding to each object audio data.
[0087]
 For example, as shown in FIG. 14, an audio file 1 in which object audio data 1, object audio data 2 and object metadata 1 + 2 corresponding to each are stored, object audio data 3 and object metadata 3 corresponding thereto are stored. Consider the case where the stored audio file 2 exists. In this case, the segment file generation unit 113 generates a metadata file by creating a file of object metadata 1 + 2 + 3 that integrates object metadata 1 + 2 and object metadata 3. It is assumed that the metadata file is a file that stores only object metadata, but the metadata file is not limited to this, and other data may be stored as appropriate.
[0088]
 Here, since the file size of the metadata file is smaller than the file size of the audio file in which the object audio data is stored, the client 200 obtains the metadata file before acquiring the audio file. The acquisition control of the audio file can be appropriately performed based on the priority information of the object metadata 1 + 2 + 3 stored in the file.
[0089]
 The method of storing the metadata file in ISOBMFF and the handling in MPEG-DASH are not specified, and are newly specified in this disclosure. In the following, the ISOBMFF (ISO Base Media File Format) of the metadata file in MPEG-H 3D Audio and AAC 3D Audio, and the method of associating the object metadata stored in the metadata file with the object audio data. Will be explained.
[0090]
 (4-2-1. ISOBMFF of
 MPEG-H 3D Audio ) Before explaining the ISOBMFF of the metadata file in MPEG-H 3D Audio, first, the existing file format will be described. MPEG-H 3D Audio defines a method for storing object audio data and object metadata in an audio file. For example, as shown in 15A of FIG. 15, an audio file 1 in which object audio data 1, object audio data 2, and object metadata 1 + 2 are stored, and an audio file in which object audio data 3 and object metadata 3 are stored. Consider the case where 2 exists.
[0091]
 In this case, the ISO BMFF of each of the audio file 1 and the audio file 2 is 15B. More specifically, in each audio file, the configuration of the object included in the sample is described in the MHADecoderConfigurationRecord of the MHAConfigurationBox ('mhaC') included in the SampleEntry of the MovieBox ('moov'). In addition, the MHAMultiStreamBox ('maeM') included in the SampleEntry describes a streamID for uniquely identifying each audio file when the audio content is provided by a plurality of audio files. Each sample data included in MediaDataBox ('mdat') includes object audio data (ES (Elementary Stream)) and object metadata for each time.
[0092]
 The Box structure of ISOBMFF described in FIG. 15 and the following is omitted as appropriate. More specifically, the Movie Box shown in 15B of FIG. 15 includes various components such as a Track Box as shown in FIG. Then, the Sample Entry is included in the Sample Description Box in the Box structure.
[0093]
 (4-2-1-1. ISOBMFF of metadata file
 in MPEG-H 3D Audio (Example 1)) In the above, ISOBMFF of an audio file in MPEG-H 3D Audio has been described. Subsequently, ISOBMFF (Example 1) of the metadata file in MPEG-H 3D Audio will be described with reference to FIG. More specifically, the ISO BMFF of the metadata file in MPEG-H 3D Audio in which the object metadata 1 + 2 + 3 corresponding to the object audio data 1 to the object audio data 3 shown in 15A of FIG. 15 is stored will be described. .. Further, in MPEG-H 3D Audio, there are two types of storage methods, a RAW method and an MHAS method. In the following, the RAW method will be described first.
[0094]
 Here, the present disclosure newly defines MHAMetadataSampleEntry ('mham'). As shown in 17A and 17B of FIG. 17, num_reference_streamID and reference_streamID are described in SampleEntry as information about the audio file corresponding to the metadata file. More specifically, num_reference_streamID indicates the number of audio files corresponding to the metadata file, and reference_streamID indicates the streamID of the audio file corresponding to the metadata file.
[0095]
 In addition, the SampleEntry contains an MHAMetadataConfigurationBox ('mhmC') to indicate the contents of each object metadata. The MHAMetadataConfigurationBox ('mhmC') contains the same MHADecoderConfigurationRecord as the audio file shown in 15B of FIG. However, in the metadata file MHADecoderConfigurationRecord,
it is possible to show that only object metadata is included by excluding the Configuration related to the Elementary stream.
[0096]
 The sample of MediaDataBox ('mdat') stores the object metadata corresponding to the audio file indicated by each streamID. In other words, the segment file generation unit 113 of the server 100 stores the object metadata stored in each audio file in the sample in the metadata file. At that time, size indicating the data size of each object metadata is added to the beginning of each object metadata.
[0097]
 In this embodiment, num_reference_streamID, reference_streamID, etc. are shown by SampleEntry in the Box structure of ISOBMFF, but the present invention is not limited to this. For example, num_reference_streamID, reference_streamID, etc. may be indicated in Sample Description Box, Sample Group, Sub-Sample Information Box, and the like.
[0098]
 Subsequently, the MHAS method ISOBMFF will be described with reference to FIG. As shown in FIG. 18, in the MHAS method, MHAMetadataSampleEntry ('mhmm') is used. Further, in the MHAS method, since the MHAMetadataConfigurationBox can be stored as the MHAConfigurationBox in the sample, it does not have to be stored in the SampleEntry (in the figure, an example in which the MHAMetadataConfigurationBox is stored in the SampleEntry is shown). Other points are the same as the RAW method described above, so the description thereof will be omitted.
[0099]
 (4-2-1-2. ISOBMFF of the metadata file
 in MPEG-H 3D Audio (Example 2)) Subsequently, with reference to FIG. 19, ISOBMFF of the metadata file in MPEG-H 3D Audio (ISOBMFF). The RAW method of the second embodiment will be described.
[0100]
 In the embodiment, the object metadata 1 + 2 stored in the audio file 1 shown in 15A of FIG. 15 and the object metadata 1 + 2 + 3 in which the object metadata 3 stored in the audio file 2 are integrated are stored in the sample. The object. In other words, when the coding processing unit 112 of the server 100 generates the object metadata 1 + 2 + 3 by integrating the object metadata 1 + 2 and the object metadata 3, and the segment file generation unit 113 generates the metadata file. The object metadata 1 + 2 + 3 is stored in the sample.
[0101]
 In this embodiment, the number of objects in the object metadata stored in the sample is shown for each streamID. More specifically, as shown in 19A of FIG. 19, it is shown that the number of objects (object_num) of the object metadata corresponding to the audio file 1 having the streamID (reference_streamID) of 1 is 2. In other words, it is shown that the object audio data 1 and the object audio data 2 are stored in the audio file 1 having the streamID 1. Further, it is shown that the number of objects (object_num) of the object metadata corresponding to the audio file 2 having the streamID (reference_streamID) of 2 is 1. In other words, it is shown that the audio file 2 having the streamID 2 stores the object audio data 3 next to the object audio data 2.
[0102]
 Other matters such as the MHAMetadataConfigurationBox shown in 19A and 19B of FIG. 19 are the same as those described with reference to FIG. 17, and thus the description thereof will be omitted. In this embodiment, reference_streamID and the like are shown by SampleEntry in the Box structure of ISOBMFF, but the present invention is not limited to this. For example, reference_streamID and the like may be indicated in Sample Description Box, Sample Group and Sub-Sample Information Box.
[0103]
 FIG. 20 is a diagram showing ISO BMFF of the MHAS method in this embodiment. Except for the point that the number of objects (object_num) in the object metadata is shown, the description is the same as that described with reference to FIG.
[0104]
 (4-2-2. ISOBMFF of AAC 3D Audio) In the
 above, ISOBMFF of MPEG-H 3D Audio has been described. Next, ISO BMFF of AAC 3D Audio will be described.
[0105]

Before discussing ISOBMFF for metadata files in  AAC 3D Audio, let's first talk about existing file formats. Since AAC 3D Audio does not specify a method for signaling multiple audio files, it is newly specified in this disclosure.
[0106]
 More specifically, as shown in 21A and 21B of FIG. 21, for SampleEntry, AAC3DAudioSampleEntry ('a3a2') indicating that a plurality of audio files are used is used (when one file is used). SampleEntry ('a3a1') is used). Further, as information for signaling a plurality of audio files, MHA MultiStreamBox ('maeM') is used as in the case of MPEG-H 3D Audio. This makes
it possible to show the relationship between audio files using streamID as in MPEG-H 3D Audio.
[0107]
 Also, the relationship between Tracks is indicated by the track reference. More specifically, the connection from the main track (file1 containing object audio data 1 and object audio data 2 in FIG. 21) to the auxiliary track (file2 containing object audio data 3 in FIG. 21) is main. Indicated by the track reference ('maux') of the track. The connection from the auxiliary track to the main track is indicated by the track reference ('mbas') of the auxiliary track.
[0108]
 The Box structure of ISOBMFF described with reference to FIG. 21 is also omitted as appropriate. More specifically, the Movie Box shown in 21A and 21B of FIG. 21 includes various components such as a Track Box as shown in FIG. Then, the Sample Entry is included in the Sample Description Box in the Box structure (however, the Sample Entry is ('a3a2') instead of ('mham') shown in FIG. 16).
[0109]
 (4-2-2-1. ISOBMFF of metadata file
 in AAC 3D Audio (Example 3)) In the above, ISOBMFF of an audio file in AAC 3D Audio has been described. Subsequently, ISOBMFF (Example 3) of the metadata file in AAC 3D Audio will be described with reference to FIG. In other words, ISOBMFF of the metadata file in AAC 3D Audio in which the object metadata 1 + 2 + 3 corresponding to the object audio data 1 to the object audio data 3 shown in 15A of FIG. 15 is stored will be described.
[0110]
 In this disclosure, A3AMetadataSampleEntry ('a3am') indicating that it is a metadata file in AAC 3D Audio is newly specified. As shown in 22A and 22B of FIG. 22, num_reference_streamID and reference_streamID are described in SampleEntry as information about the audio file corresponding to the metadata file. More specifically, num_reference_streamID indicates the number of audio files corresponding to the metadata file, and reference_streamID indicates the streamID of the audio file corresponding to the metadata file. In AAC 3D Audio, the Object Metadata Configuration is stored in the Sample of MediaDataBox ('mdat'), so it is not signaled by the Sample Entry.
[0111]
 The sample of MediaDataBox ('mdat') stores the object metadata corresponding to the audio file indicated by each streamID. In other words, the segment file generation unit 113 of the server 100 stores the object metadata stored in each audio file in the sample in the metadata file. At that time, size indicating the data size of each object metadata is added to the beginning of each object metadata.
[0112]
 Here, for the structure of the sample, 3da_meta_data () of AAC 3D Audio may be used, or 3da_meta_data () stored in DSE, which is used in the Elementary stream of AAC 3D Audio, may be used. The structure of 3da_meta_data () is shown in FIG. 45, the structure of 3da_meta_data () stored in the DSE is shown in FIG. 46, and the structure of the DSE is shown in FIG. 47. The structure of 3da_ancillary_data stored in data_stream_byte in DSE shown in FIG. 47 is shown in FIG. 48. However, if the size of 3da_meta_data () is larger than the maximum size of data_stream_byte of DSE, 3da_meta_data () is divided and stored in multiple DSEs.
[0113]
 The Box structure of ISOBMFF described in FIG. 22 and the following is also omitted as appropriate. More specifically, the Movie Box shown in 22A of FIG. 22 includes various components such as a Track Box as shown in FIG. Then, the Sample Entry is included in the Sample Description Box in the Box structure (however, the Sample Entry is ('a3am') instead of ('mham') shown in FIG. 16).
[0114]
 Further, in this embodiment, num_reference_streamID, reference_streamID, etc. are shown by SampleEntry in the Box structure of ISOBMFF, but the present invention is not limited to this. For example, num_reference_streamID, reference_streamID, etc. may be indicated in Sample Description Box, Sample Group, Sub-Sample Information Box, and the like.

The scope of the claims
[Claim 1]
 A priority is set for each of the acquired object audio data, and
based on the priority , the object audio data to be included in the generated segment file is determined from one or more of the object audio data, and the
priority is determined. An information
 processing apparatus including a processing unit that generates new priorities to be set for the generated segment file as priority information based on the above .
[Claim 2]
 The processing unit further generates a metadata file having encoded object metadata corresponding to the object audio data, and the object metadata
includes the priority information, according to
 claim 1. Information processing device.
[Claim 3]
 The first aspect of the present invention, wherein the processing unit further generates an MPD file for the segment file and includes the priority information in the MPD file when the priority information does not change with the passage of time
 . Information processing device.
[Claim 4]

 The information processing apparatus according to claim 3,  wherein the processing unit includes the priority information in the adaptation set of the MPD file .
[Claim 5]

 The information processing apparatus according to claim 2,  wherein when the priority information changes with the passage of time, the processing unit further generates an MPD file for the segment file and the metadata file .
[Claim 6]

 The information processing apparatus according to claim 4 or 5,  wherein the processing unit includes information regarding a bit rate of the object audio data in the MPD file for each object audio data .
[Claim 7]

 The information processing device according to claim 2, wherein the processing unit  stores information used for identifying the segment file in the Sample Description Box of the Movie Box of the metadata file .
[Claim 8]

 The information processing device according  to claim 7, wherein the processing unit stores information used for identifying the segment file in the Sample Entry in the Sample Description Box .
[Claim 9]

 The information processing apparatus according  to claim 8 , wherein the information used for identifying the segment file includes a stream ID for uniquely identifying the segment file .
[Claim 10]

 An information processing method executed by a computer  , which has information processing of stream data in which a priority is set for each acquired object audio data .
[Claim 11]

 A program  that allows a computer to process stream data for which priority is set for each acquired object audio data .

Documents

Application Documents

# Name Date
1 202017025926-TRANSLATIOIN OF PRIOIRTY DOCUMENTS ETC. [19-06-2020(online)].pdf 2020-06-19
2 202017025926-STATEMENT OF UNDERTAKING (FORM 3) [19-06-2020(online)].pdf 2020-06-19
3 202017025926-PRIORITY DOCUMENTS [19-06-2020(online)].pdf 2020-06-19
4 202017025926-POWER OF AUTHORITY [19-06-2020(online)].pdf 2020-06-19
5 202017025926-FORM 1 [19-06-2020(online)].pdf 2020-06-19
6 202017025926-DRAWINGS [19-06-2020(online)].pdf 2020-06-19
7 202017025926-DECLARATION OF INVENTORSHIP (FORM 5) [19-06-2020(online)].pdf 2020-06-19
8 202017025926-COMPLETE SPECIFICATION [19-06-2020(online)].pdf 2020-06-19
9 202017025926-Proof of Right [17-07-2020(online)].pdf 2020-07-17
10 202017025926.pdf 2021-10-19