Sign In to Follow Application
View All Documents & Correspondence

Information Processor And Information Processing Method

Abstract: The present invention relates to an information processor and an information processing method in which a prescribed type of audio data can easily be played back from among multiple types of audio data. A file creation device creates audio files in which an audio stream of multiple groups is arranged as being divided into tracks for each group or for each set of more than one group and information pertaining to multiple groups is arranged. The present invention can be applied to an information processing system or the like configured from e.g. a file creation device for creating files a Web server for recording files created by the file creation device and a video playback terminal for playing back the files.

Get Free WhatsApp Updates!
Notices, Deadlines & Correspondence

Patent Information

Application #
Filing Date
22 December 2016
Publication Number
14/2017
Publication Type
INA
Invention Field
ELECTRONICS
Status
Email
remfry-sagar@remfry.com
Parent Application

Applicants

SONY CORPORATION
1 7 1 Konan Minato ku Tokyo 1080075

Inventors

1. HIRABAYASHI Mitsuhiro
c/o SONY CORPORATION 1 7 1 Konan Minato ku Tokyo 1080075
2. YAMAMOTO Yuki
c/o SONY CORPORATION 1 7 1 Konan Minato ku Tokyo 1080075
3. CHINEN Toru
c/o SONY CORPORATION 1 7 1 Konan Minato ku Tokyo 1080075
4. SHI Runyu
c/o SONY CORPORATION 1 7 1 Konan Minato ku Tokyo 1080075

Specification

DESCRIPTION INFORMATION PROCESSINGAPPARATUS AND INFORMATION PROCESSING METHOD 5 TECHNICAL FIELD [00011 The present disclosure relates to an information processing apparatus and an information processing method, andespeciallyrelatestoaninformationprocessingapparatus 10 and an formation processing method that enable easy reproduction of audio data of a predetermined kind, of audio data of a plurality of kinds. BACKGROUND ART 15 [0002] In recent years, the mainstream of streaming services on the Internet has been over the top video (OTT-V). A technology growing popular as a basic technology is moving picture experts group phase-dynamic adaptive streaming over 20 HTTP (MPEG-DASH) (for example, see Non-Patent Document I). [00031 In MPEG-DASH, a distribution server prepares moving image data groups with different screen sizes and encoding speeds, for one piece of moving image content, and a 25 reproduction terminal requires the moving image data group I with an optimum screen size and an optimum encoding speed according to a state of atransmission path, sothat adaptive streaming distribution is realized. 30 CITATION LIST NON-PATENT DOCUMENT [0004] Non-Patent Document 1: Dynamic Adaptive Streaming over HTTP (MPEG-DASH) (UP,~:http://mpeg.chiar~glione.org/standards/mpeg-dash/med 5 ia-presentation-description-and-segment-formats/text-isoi ec-23009-12012-dam-1) SUMMARY OF INVENTION PROBLEMS TO BE SOLVED BY THE INVENTION 10 [0005] However, easy reproduction of audio data of a predetermined group, of audio data of a plurality of groups, has not been considered. [0006] 15 The present disclosure has been made in view of the foregoing, and enables easy reproduction of audio data of a desired group, of audio data of a plurality of groups. SOLUTION TO PROBLEMS 20 [0007] An inforination processing apparatus of a firsL aspect of the present disclosure is an information processing apparatus including a file generation unit that generates a file in which audio data of a plurality of kinds is divided 25 into tracks for each one or more of the kinds and arranged, andinformationrelatedtothepluralityof kindsis arranged. [00081 An information processing method of the first aspect of the present disclosure corresponds to the information 30 processing apparatus of the first aspect of the present disclosure. [0009] In the f i r s t aspect of the present disclosure, the f i l e in which audio data of a p l u r a l i t y of kinds is divided i n t o tracks for each one or more of the kinds and arranged, and 5 information r e l a t e d to the p l u r a l i t y of kinds is arranged i s generated. [ O O l O ] An information processing apparatus of a second aspect of the present disclosure is an information processing 10 apparatusincludingareproductionunitthatreproduces, from a f i l e in which audio data of a p l u r a l i t y of kinds is divided i n t o t r a c k s for each one or more of the kinds and arranged, andinformationrelatedtothepluralityofkindsis arranged, the audio data in a predetermined t r a c k . 15 [ O O l l ] An information processing method of the second aspect of the present disclosure corresponds t o the information processing apparatus of the second aspect of the present disclosure. 20 [0012] In the second aspect of the present disclosure, the audio data of a predetermined track is reproduced from the f i l e i n which audiodata o f a p l u r a l i t y o f k i n d s i s d i v i d e d i n t o t r a c k s foreachoneormore o f t h e kinds andarranged, andinformation 25 r e l a t e d t o the p l u r a l i t y of kinds is arranged. [00131 Notethat t h e i n f o r m a t i o n p r o c e s s i n g a p p a r a t u s e s o f t h e f i r s t andsecondaspects canbe r e a l i z e d b y c a u s i n g a computer t o execute a program. 30 [0014] Further, t o r e a l i z e the information processing apparatuses of the f i r s t and second aspects, the program executed by the computer can be transmitted through a transmission medium, or can be recorded on a recording medium and provided. 5 EFFECTS OF THE INVENTION [0015] Accordingtothe f i r s t aspect o f t h e p r e s e n t disclosure, a f i l e canbe generated. Further, according t o the f i r s t aspect 10 of the present disclosure, a f i l e t h a t enables easy reproduction of audio data of a predetermined kind, of audio data of a p l u r a l i t y of kinds, can be generated. [0016] According t o the secondaspect of the present disclosure, 15 audiodatacanbereproduced. Further, accordingtothesecond a s p e c t o f t h e p r e s e n t d i s c l o s u r e , audiodataofapredetermined kind, of audio data of a p l u r a l i t y of kinds, can be e a s i l y reproduced. 20 BRIEF DESCRIPTION OF DRAWINGS [0017] Fig. 1 is a diagram i l l u s t r a t i n g a s t r u c t u r e of an MPD f i l e . Fig. 2 i s a diagram i l l u s t r a t i n g r e l a t i o n s h i p among 25 "Period", "Representation", and "Segment". Fig. 3 is adiagramillustratingahierarchical s t r u c t u r e of the MPD f i l e . Fig. 4 i s a diagram i l l u s t r a t i n g r e l a t i o n s h i p between a s t r u c t u r e and a time axis of the MPD f i l e . 3 0 Fig. 5 i s a diagram for describing an outline of a t r a c k of a 30 audio r i l e format of MP4. Fig. 6 is a diagram illustrating a structure of a moov box. Fig. 7 is a diagramillustrating a hierarchical structure of 3D audio. 5 Fig. 8,is a diagram for describing an outline of an information processing system in a first embodiment to which the present disclosure is applied. Fig. 9is a diagram for describing an outline of a first example of atrackinthe first embodiment towhichthepresent 10 disclosure is applied. Fig. 10 is a diagram illustrating an example of syntax of sample entry of a base track. Fig. 11 is a diagram illustrating an example of syntax of sample entry ofatrack of a group that forms switch Group. 15 Fig. 12 is a diagram illustrating a first example of a segment structure. Fig. 13 is a diagram illustrating a second example of the segment structure. Fig. 14 is adiagramillustratingadescriptionexample 20 of a level assignment box. Fig. 15 is a diagram illustrating a first description example of the MPD file in the first embodiment to which the present disclosure is applied. Fig. 16is ablockdiagramillustratinga configuration 25 example of a file generation device of Fig. 8. Fig. 17 is a flowchart for describing file generation processing of the file generation device of Fig. 16. Fig. 18 is ablockdiagramillustrating a configuration example of a streaming reproduction unit realized with a moving 30 image reproduction terminal of Fig. 8. Fig. 19 is a flowchart for describing reproduction p r o c e s s i n g of t h e streaming reproduction u n i t of Fig. 18. Fig. 2 0 i s a d i a g r a m f o r d e s c r i b i n g a n o u t l i n e o f a s e c o n d example o€ t h e t r a c k i n t h e f irst embodiment t o which t h e p r e s e n t d i s c l o s u r e is a p p l i e d . 5 Fig. 21 is a diagram i l l u s t r a t i n g an example of syntax of sample group e n t r y of a t r a c k of a group t h a t forms switch Group. Fig. 22 i s a diagram i l l u s t r a t i n g an example of syntax of sample e n t r y of a t r a c k of each of groups. 10 F i g . 23 is adiagramfordescribinganoutlineoliathird example of t h e t r a c k of an audio f i l e . Fig. 24 is a diagram i l l u s t r a t i n g a second d e s c r i p t i o n example of t h e MPD f i l e . Fig. 25 is a diagram i l l u s t r a t i n g another example of 15 t h e second d e s c r i p t i o n example of t h e MPD f i l e . Fig. 2 6 i s a d i a g r a m f o r d e s c r i b i n g a n o u t l i n e o f a f o u r t h example of t h e t r a c k of t h e audio f i l e . Fig. 27 is a diagram i l l u s t r a t i n g a t h i r d d e s c r i p t i o n example of t h e MPD f i l e . 20 Fig. 2 8 i s a d i a g r a m f o r d e s c r i b i n g a n o u t l i n e o f a f i f t h example of t h e t r a c k of t h e audio f i l e . Fig. 29 is a diagram i l l u s t r a t i n g an example of syntax of sample e n t r y i n which 4cc is "mha3". Fig. 30 is a diagram i l l u s t r a t i n g a n o t h e r example of 25 t h e syntax of t h e sample e n t r y where 4cc is "mha3". Fig. 31 is a diagram i l l u s t r a t i n g a f o u r t h d e s c r i p t i o n example of a MPD f i l e . Fig. 3 2 i s a d i a g r a m f o r d e s c r i b i n g a n o u t l i n e o f a n o t h e r example of t h e t h i r d example of t h e t r a c k of t h e audio f i l e . 30 Fig. 3 3 i s a d i a g r a m f o r d e s c r i b i n g a n o u t l i n e o f a n o t h e r example of t h e f o u r t h example of t h e t r a c k of t h e audio f i l e . Fig. 3 4 i s a d i a g r a m f o r d e s c r i b i 1 1 g a n o u t l i n e o f a n o t h e r example of t h e f i f t h example of t h e t r a c k of t h e audio f i l e . Fig. 35 is a d i a g r a m f o r d e s c r i b i n g a n o u t l i n e o f a s i x t h example of t h e t r a c k of t h e audio f i l e . 5 Fig. 36 is a diagram i l l u s t r a t i n g an example of syntax of sample e n t r y of a base t r a c k and a group t r a c k o f F i g . 3 5 . Fig. 3 7 i s a d i a g r a m i l l u s t r a t i n g s t i l l another example of t h e syntax of t h e sample e n t r y where 4cc is "mha3". Fig. 3 8 i s adiagramfordescribinganoutlineofatrack 10 i n a second embodiment t o which t h e p r e s e n t d i s c l o s u r e is a p p l i e d . Fig. 39 is a diagram i l l u s t r a t i n g a f i r s t d e s c r i p t i o n example o f an MPD f i l e i n t h e second embodiment t o which t h e p r e s e n t d i s c l o s u r e is a p p l i e d . 15 Fig. 40 is a diagram f o r d e s c r i b i n g an o u t l i n e of an informationprocessingsysteminthesecondembodimenttowhich t h e p r e s e n t d i s c l o s u r e is a p p l i e d . Fig. 4 1 i s ablockdiagramillustratinga c o n f i g u r a t i o n example o f a f i l e g e n e r a t i o n d e v i c e of Fig. 40. 2 0 Fig. 42 is a flowchart f o r d e s c r i b i n g f i l e g e n e r a t i o n p r o c e s s i n g of the f i l e g e n e r a t i o n d e v i c e o f Fig. 4 1 . Fig. 4 3 i s ablockdiagramillustratinga c o n f i g u r a t i o n example of a streaming reproduction u n i t r e a l i z e d with amoving image reproduction t e r m i n a l o f F i g . 40. 2 5 F i g . 44 is a flowchart f o r d e s c r i b i n g an example of reproduction p r o c e s s i n g of t h e streaming reproduction u n i t of Fig. 43. Fig. 45 is a diagram i l l u s t r a t i n g a second d e s c r i p t i o n example of t h e MPD f i l e i n t h e second embodiment t o which t h e 30 p r e s e n t d i s c l o s u r e is a p p l i e d . Fig. 46 is a diagram i l l u s t . r a t i n g a t h i r d d e s c r i p t i o n example of the MPD file in the second embodiment to which the present disclosure is applied. Fig. 47 is a diagram illustrating a fourth description example of the MPD file in the second embodiment to which the 5 present disclosure is applied. Fig. 48 is a diagram illustrating a fifth description example of the MPD file in the second embodiment to which the present disclosure is applied. Fig. 49 is a diagram illustrating a sixth description 10 example of the MPD file in the second embodiment to which the present disclosure is applied. Fig. 50is a diagramillustratinga seventhdescription example of the MPD file in the second embodiment to which the present disclosure is applied. 15 Fig. 51 is a diagram illustrating an example of a track structure of an audio file including aplurality of base tracks. Fig. 52 is a diagram illustrating another example of the track structure ofthe audio file includingtheplurality of base tracks. 2 0 Fig. 53is ablockdiagramillustrating a corifiquration example of hardware of a computer. MODE FOR CARRYING OUT THE INVENTION [0018] 2 5 Hereinafter, presuppositionsofthepresent disclosure and embodiments for implementing the present disclosure (hereinafter, referred to as embodiments) will be described. Note that the description will be given as follows: 0. Presuppositions of Present Disclosure (Figs. 1 to 7) 30 1. First Embodiment (Figs. 8 to 37) 2. Second Embodiment (Figs. 38 to 50) 3. Another Example of Base Track (Figs. 51 and 52) 4. Third Embodiment (Fig. 53) [0019] (Description of 5 Structure of MPD File) Flg. 1 is a diagramillustrating a structure of amedia presentation description (MPD) file of MPEG-DASH. [0020] In an analysis (parsing) of an MPD file, an optimum one 10 is selected from "Representation" attributes included in "Periods" of the MPD file (Media Presentation of Fig. 1). [0021] Then, a file is acquired and processed by reference to a uniform resource locator (URL) and the like of 15 "Initialization Segment" in a head of the selected "Representation". Following that, a file is acquired and reproduced by reference to a URL and the like of subsequent "Media Segment". [0022] 20 Note that relationship among "Period", "Representatic~n",a nd "Segment" in the MPD file is illustrated in Fig. 2. That is, one piece of moving image content can bemanagedinunits of a longertime than the segmentby "Period", and can be managed in units of a segment by "Segment" in each 25 of "Periods". Further, in each of "Periods", the moving image content can be managed in units of an attribute of a stream by "Representation". [0023] Therefore, the MPD file has a hierarchical structure 30 illustrated in Fig. 3 in and under "Period". Further, arrangement of the structure of the MPD file on a time axis is illustrated in the example of Fig. 4. As is clear from Fig. 4, aplurality of "Representations" exists with respect to the same segment. By adaptively selecting any of these "Representations", a stream of a desired attribute of a user 5 can be acquired and reproduced. [0024 j (Outline of 3D Audio File Format) Fig. 5is a diagram for describing an outline of atrack of a 3D audio file format of MP4. 10 [0025j In an MP4 file, codec information of the moving image content, and position information indicating a position in a filecanbemanagedforeachtrack. Ina 3Daudio file format of MP4, all of audio streams (elementary streams (ESs) ) of 15 30 audio (Channel audio/Object audio/SAOC Object audio/HOA audio/metadata) are recordedas one trackinunits ofa sample (frame). Further, the codec information (Pro file/level/audio configuration) of the 3D audio is stores as sample entry. 20 [00261 The Channel audio that configure sthe 3Daudio is audio data in units of a channel, and the Object audio is audio data in units of an object. Note that an object is a sound source, and the audio data in units of an object is acquired with a 25 microphone or the like attached to the object. The object maybe a substance such as a fixedmicrophone standor amoving body such as a person. [0027] Further, the SAOC Object audio is audio data of spatial 30 audio object coding (SAOC), the HOA audio is audio data of higher order ambisonics (HOA) , and the metadata is metadata of the Channel audio, the Object audio, the SAOC Object audio, and the HOA audio. [0028] (Structure of moov Box) 5 Fig. 6 is a diagram illustrating a structure of a moov box of the MP4 file. [0029] As illustrated in Fig. 6, in the MP4 file, image data and audio data are recorded as different tracks. In Fig. 6, 10 although details are not described, the track of the audio data is similar to the track of the image data. The sample entry is included in sample description arranged in a stsd box in the moov box. [O03O] 15 By theway, inbroadcasting or local storage reproduction of theMP4 file, typically, a server side sends the audio streams of all of the 3D audio. Then, a client side decodes and outputs only the audio streams of necessary 3D audio while parsing the audio streams of all of the 3D audio. However, in a case 20 where a bit rate is high or there is a restriction on a reading rate of the local storage, it is desirable to reduce a load ofthedecodingprocessingbyacquiringonlytheaudio streams of the necessary 3D audio. [0031] 25 Further, in stream reproduction of the MP4 file conformable to MPEG-DASH, the server side prepares the audio streams at a plurality of encoding speeds. Therefore, the client side can select and acquire the audio streams at an encoding speed optimum for a reproduction environment by 30 acquiring only the audio streams of necessary 3D audio. 100321 As described above, in the present disclosure, by dividing the audio streams of the 3Daudio into tracks according to kinds, and arranging the audio streams in an audio file, onlythe audio streams ofapredetermined kindofthe 3Daudio 5 can be efficiently acquired. Accordingly, in the broadcasting or the local storage reproduction, the load of the decoding processing can be reduced. Further, in stream reproduction, the audio streams with highest quality, ofthe audio streams of the necessary 30 audio, can be reproduced 10 according to a band. 100331 (Description of Hierarchical Structure of 3D Audio) Fig. 7 is a diagram illustrating ahierarchical structure of the 3D audio. 15 100341 As illustrated in Fig. 7, the audio data of the 3D audio is an audio element (Element) that is different in each audio data. Types of the audio elements include a single channel element (SCE) and a channel pair element (CPE) . The type of 20 the audio element of the audio data of one channel is the SCE and the type oL the audio element corresponding to the audio data of two channels is the CPE. [0035] The audio elements of the same audio kind 25 (Channel/Object/SAOC Object/HOA) form a group. Therefore, examples of a group type (GroupType) include Channels, Objects, SAOC Objects, and HOA. Two or more groups can form switch Group or group Preset as needed. [0036] 30 The switch Group is a group (exclusive reproduction group) in which an audio stream ofthe groupincludedtherein is exclusively reproduced. That is, as i l l u s t r a t e d i n Fig. 7, in a case where there are a group of the Object audio for English (EN) and a group of the object audio for French (FR) , only one of the groups should be reproduced. Therefore, the 5 switch Group i s formed of the group of the Object audio for Englishwitha group I D of 2, andthe group o f t h e Object audio f o r F r e n c h w i t h a g r o u p I D o f 3 . Accordingly, t h e o b j e c t audio for English or the Object audio for French i s exclusively reproduced. 10 [a0371 Meanwhile, the group Preset defines a combination of the groups intended by a content c r e a t o r . [a0381 Further, t h e m e t a d a t a o f t h e 3DaudioisExtelement (Ext 15 Element) t h a t i s d i f f e r e n t i n each metadata. Types of the Extelement include Object Metadata, SAOC 3D Metadata, HOA Metadata, DRC Metadata, SpatialFrame, SaocFrame, and the l i k e . The Extelement of the Object Metadata is metadata of a l l of the Object audio, and the Extelement of the SAOC 3D Metadata 20 ismetadata of a l l o f t h e SAOCaudio. Further, theExtelement of the HOA Metadata is metadata of a l l of the HOA audio, and Extelementofdynamicrangecontrol (DRC) Metadataismetadata of a l l of the Object audio, the SAOC audio, and the HOA audio. [a0391 2 5 As described above, division u n i t s of the audio data, of the 3D audio, include the audio element, the group type, thegroup, theswitchGroup, a n d t h e g r o u p p r e s e t . Therefore, the audio streams of the audio data, of the 3D audio, can be divided i n t o d i f f e r e n t tracks i n each kind, where the kind 30 is the audio element, the group type, the group, the switch Group, or the group Preset. [00401 Further, divisionunits of themetadata, of the 3Daudi0, include a type of the Extelement and the audio element corresponding tothemetadata. Therefore, the audio streams 5 of the metadata of the 30 audio can be divided into different tracks in each kind, where the kind is the Extelement or the audio element corresponding to the metadata. [0041] In the embodiment below, the audio streams of the audio 10 data are divided into the tracks in each one or more groups, and the audio streams of the metadata are divided into the tracks in each type of the Extelement. [0042] 15 (Outline of Information Processing System) Fig. 8 is a diagram for describing an outline of an information processing system in a first embodiment to which the present disclosure is applied. [00431 2 0 An information processing system 140 of Fig. 8 is configured such that a web server 142 connected wil:h a file generationdevice141andamovingimagereproductionterminal 144 are connected through the Internet 13. [00441 2 5 Intheinformationprocessingsystem140,thewebserver 142 distributes the audio streams of the tracks in the group to be reproduced to the moving image reproduction terminal 144 by a method conforming to MPEG-DASH. [00451 3 0 To be specific, the file generation device 141 encodes the audio data and the mPtadata of the 313 audio of the moving image content a t p l u r a l i t y of encoding speeds t o generate the audio streams. The f i l e generation device 1 4 1 makes f i l e s of a l l o f t h e audio streams a t the encoding speeds and in each timeunitfromseveralseconds t o t e n seconds, which i s called 5 segment, t o generate the audio f i l e . A t t h i s time, the f i l e generationdevice141dividestheaudiostreams foreachgroup a n d e a c h t y p e o f t h e E x t e l e m e n t , andarrangestheaudiostreams i n t h e audio f i l e as the audio s t r e a m s i n t h e d i f f e r e n t t r a c k s . The f i l e generation device 1 4 1 uploads the generated audio 10 f i l e onto the web server 1 4 2 . [00461 Further, the f i l e generation device 1 4 1 generates the MPD f i l e (management f i l e ) t h a t manages the audio f i l e and the l i k e . The f i l e g e n e r a t i o n d e v i c e 1 4 l u p l o a d s t h e M P D f i l e 15 onto the web server 1 4 2 . [ 0 0 4 7 ] Thewebserver142 s t o r e s t h e a u d i o f i l e o f e a c h e n c o d i n g speed and segment, and the MPD f i l e uploaded by the f i l e generation device 1 4 1 . The web server 142 transmits the stored 20 audio f i l e , the MPD f i l e , and the l i k e , t o the moving image reproduction terminal 1 4 4 , i n response t o a request from the moving image reproduction terminal 1 4 4 . [0048] The moving image reproduction terminal 1 4 4 executes 25 control software of streaming data ( h e r e i n a f t e r , r e f e r r e d to as control software) 161, moving image reproduction software 162, c l i e n t software for hypertext t r a n s f e r protocol (HTTP) access ( h e r e i n a f t e r , r e f e r r e d t o a s access software) 163, and the l i k e . 30 [0049] The c o n t r o l s o f t w a r e 1 6 1 i s software t h a t c o n t r o l s d a t a streamedfromtheweb server142. Tobe specific, the control software 161 causes the moving image reproduction terminal 144 to acquire the MPD file from the web server 142. [0050] 5 Further, the control software 161 commands the access software 163 to send a transmission request of the group to be reproduced specified by the moving image reproduction software 162, and the audio streams of the tracks of the type of Extelement corresponding to the group, on the basis of the 10 MPD file. (00511 Themoving image reproduction software 162 is software thatreproducestheaudiostreamsacquiredfromthewebserver 142. To be specific, the moving image reproduction software 15 162 specifies the group to be reproduced and the type of the E x t e l e m e n t c o r r e s p o n d i n g t o t h e g r o u p , t o t h e c o n t r o l s o f t w a r e 161. Further, the moving image reproduction software 162 decodes the audio streams received from the moving image reproduction terminal 144 when receiving notification of 20 receptionstartfromtheaccesssoftware163. Themovingimage reproduction software 162 synthesizes and outputs the audio data obtained as a result of the decoding, as needed. [00521 The access software 163 is software that controls 25 communicationbetween themoving image reproduction terminal 144 and the web server 142 through the Internet 13 using the HTTP. Tobe specific, the access software 163 causes themoving image reproduction terminal 144 to transmit a transmission request of the audio stream of the track to be reproduced 30 included in the audio file in response to the command of the controlsoftwarel61. Further,theaccesssoftware163causes themoving image reproduction t e r m i n a l 1 4 4 t o s t a r t r e c e p t i o n of t h e a u d i o s t r e a m s t r a n s m i t t e d from t h e web s e r v e r 142 i n response t o t h e t r a n s m i s s i o n r e g u e s t , and s u p p l i e s n o t i f i c a t i o n of t h e r e c e p t i o n s t a r t t o t h e moving image 5 reproduction software 162. [0053] Note t h a t , i n t h e p r e s e n t s p e c i f i c a t i o n , o n l y t h e audio f i l e o f t h e m o v i n g image c o n t e n t w i l l b e d e s c r i b e d . However, i n r e a l i t y , a corresponding image f i l e is generated and 10 reproduced t o g e t h e r with t h e audio f i l e . [a0541 ( O u t l i n e of F i r s t Example of Track o f Audio F i l e ) Fig. 9 i s a d i a g r a m f o r d e s c r i b i n g an o u t l i n e of a f i r s t example of t h e t r a c k of t h e audio f i l e . 15 [a0551 Note t h a t , i n Fig. 9, only t h e t r a c k of t h e audio d a t a , o f t h e 3 D a u d i o , i s i l l u s t r a t e d f o r c o n v e n i e n c e o f d e s c r i p t i o n . The same a p p l i e s t o Figs. 20, 23, 26, 28, 30, 32 t o 35, and 38. 20 [00561 As i l l u s t r a t e d i n Fig. 9, t h e audio streams of a l l of t h e 3D audio a r e s t o r e d i n one audio f i l e (3dauio.mp4). In t h e audio f i l e (3dauio.mp4), t h e a u d i o s t r e a m s of t h e groups of t h e 30 audio a r e r e s p e c t i v e l y divided i n t o t h e d i f f e r e n t 25 t r a c k s and arranged. F u r t h e r , i n f o r m a t i o n r e l a t e d t o t h e e n t i r e 3D audio is arranged a s t h e base t r a c k (Base T r a c k ) . [00571 Track Reference is arranged i n a t r a c k box of each of t h e t r a c k s . The Track Reference i n d i c a t e s r e f e r e n c e 30 relationshipbetweenacorrespondingtrackandanothertrack. Tobe s p e c i f i c , t h e T r a c k R e f e r e n c e i n d i c a t e s an I D o f a n o t h e r track in the reference r e l a t i o n s h i p , unique t o the track ( h e r e i n a f t e r , r e f e r r e d t o as track I D ) . [0058] Intheexampleof Fig. 9 , t h e t r a c k I D s o f t h e b a s e t r a c k , 5 the track in a group #1 with a group I D of 1, t h e t r a c k i n a group #2 with a group I D of 2, t h e t r a c k i n a group #3 with a group I D of 3, t h e t r a c k in a group #4 with a group I D of 4, are 1, 2, 3, 4 , and 5. Further, the Track Reference of the base track i s 2, 3, 4, and 5, and the Track Reference of 10 the tracks in the groups #1 t o #4 is 1 t h a t is the track I D of the base track. Therefore, the base track, and the tracks i n t h e g r o u p s # 1 t o # 4 a r e i n t h e r e f e r e n c e r e l a t i o n s h i p . That is, the base track i s referenced a t the time of reproduction of the tracks in the groups #1 t o #4. 15 [0059] Further, 4cc (character code) of the sample entry of the base track is "mha2", and i n the sample entry of the base track, an mhaC box including config information of a l l of the groups of the 3D audio or config information necessary for 20 decoding only the base track, and an mhas box including information r e l a t e d t o a l l o f t h e groups and the switch Group o f t h e 3D audio are arranged. The information r e l a t e d t o t h e groups is configured from the IDS of the groups, information indicating content of data o f t h e elementclassifiedintothe 25 groups, and the l i k e . The information r e l a t e d t o the switch Group is configured from an I D of the switch Group, the I D S of the groups t h a t form the switch Group, and the l i k e . [0060] The 4cc of the sample entry of t h e t r a c k of each of the 30 groups is "mhgl", and i n the sample entry o f t h e t r a c k o f each of the groups, an mhgC box including information r e l a t e d to the group may be arranged. In a case where a group forms the switch Group, an mhsC box including information related to the switch Group is arranged in the sample entry of the track in the group. 5 [00611 In a sample of the base track, reference information to samples of the tracks in the groups or config information necessaryfordecodingthereferenceinformationisarranged. By arranging the samples of the groups referenced by the 10 referenceinformationinorderofarrangementofthereference information, the audio streams of the 3D audio before being divided into the tracks can be generated. The reference information is configured from positions and sizes of the samples of the tracks the groups, the group types, and the 15 like. [00621 (Example of Syntax of Sample Entry of Base Traclc) Fig. 10 is a diagram illustrating an example of syntax of the sample entry of the base track. 20 [0063] As illustrated in Fig. 10, in the sample entry of the base track, the mhaC box (MHAConfigration Box), the mhas box (MHAAudioSceneInfo Box), and the like are arranged. In the mhaC box, the config information of all of the groups of the 25 3 D a u d i o o r t h e c o n f i g i n f o r m a t i o n n e c e s s a r y f o r d e c o d ~ n g o n l y the base track is described. Further, in the mhas box, Audioscene information including the information related to all of the groups and the switch Group of the 3D audio is described. The Audioscene information describes the 30 hierarchical structure of Fig. 7 [00641 (Examp1.e of Syntaxof Sample Entry of Trackof each of Groups) Fig. 11 is a diagram illustrating an example of syntax of sample entry of the track of each of the groups. [0065] 5 As illustrated in Fig. 11, in the sample entry of the track of each of the groups, the mhaC box (MHAConfigration Box), the mhgC box (MHAGroupDefinition Box), the mhsC box (MHASwitchGropuDefinition Box), and the like are arranged. [0066] 10 In the mhaC box, Config information necessary for decoding the corresponding track is described. Further, in the mhgC box, AudioScene information related to the corresponding group is describedas GroupDefinition. In the mhsC box, AudioScene information related to the switch Group 15 is described in SwitchGroupDefinition in a case where the corresponding group forms the switch Group. [0067] (First Example of Segment Structure of Audio File) Fig. 12 is a diagram illustrating a first example of 20 a segment structure of the audio file. [0068] Inthe segment structure of Fig. 12, an Initial segment is configured from an ftyp box and a moov box. In the moov box, a trak box is arranged for each track included in the 25 audio file. Further, in the moov box, an mvex box including information indicating corresponding relationship between the track ID of each of the tracks and a level used in an ssix box in a media segment, and the like are arranged [0069] 3 0 Further, the media segment is configured from an sidx box, an ssix box, and one or more subsegments. In the sidx box, position information indj-cating positions of the subsegments in the audio file is arranged. In the ssix box, position information of the audio streams of the levels arranged in an mdat box is arranged. Note that the level 5 correspondsto the track. Further, the position information of the first track is the position information of data made of an moof box and the audio stream of the first track. [0070] The subsegment is provided for eacharbitrary time length, 10 and the subsegment is provided with a pair of the moof box and the mdat box, which is common to. all of the tracks. In the mdat box, the audio streams of all of the tracks are collectivelyarrangedby an arbitrarytime length, and in the moof box, management information of the audio streams is 15 arranged. The audio streams of the tracks arranged in the mdat box are successive in each track. [0071] In the example of Fig. 12, Track1 with the track ID of 1 is the base track, and Track2 to TrackN with the tracks ID 20 of 2 to N are tracks in the groups with the group ID of 1 to N - 1. The same applies to Fig. 13 described below. [0072] (Second Example of Segment Structure of Audio File) Fig. 13 is a diagram illustrating a second example of 25 the segment structure of the audio file. [0073] The segment structure of Fig. 13 is different from the segment structure of Fig. 12 in that the moof box and the mdat box are provided for each track. 30 [0074] That is, the Initial segment of Fig. 13 is similar to the Initial segment of Fig. 12. Further, the media segment of Fig. 13 is configured from the sidx box, the ssix box, and one or more subsegments, similarly to the media segment of Fig. 12. In the sidx box, the positlon information of the 5 subsegments is arranged, similarly to the sidx box of Fig. 12. In the ssix box, position information of data of the levels made of the moof box and the mdat box is included. [0075] The subsegment isprovidedfor eacharbitrary time length, 10 and the subseqment is provided with a pair of the moof box and the mdat box for each track. That is, in the mdat box of each of the tracks, the audio streams of the tracks are collectively arranged (interleave storage) by an arbitrary time length, and in the moof box, management information of 15 the audio streams is arranged. [0076] As illustrated in Figs. 12 and-13, the audio streams of the tracks are collectively arranged by an arbitrary time length. Therefore, audio stream acquisition efficiency 20 t h r o u g h t h e H T T P o r t h e l i k e i s i m p r o v e d , compared with a case where the audio streams are collectively arranged jn units of a sample. [0077] (Description Example of mvex Box) 2 5 Fig. 14 is adiagramillustratinga descriptionexample of a level assignment box arranged in the mvex box of Figs. 12 and 13. [0078] The level assignment box is a box that associates the 30 track ID of each of the tracks and the level used in the ssix box. In the example of Fig. 14, the base trackwith thc track IDof ].is associatedwith alevel0, anda channel audio track withthe track TDof 2 is associatedwith a levell. Further, an HOA audio track with the track ID of 3 is associated with a level 2, and an object metadata track with the track ID of 5 4 i s a s s o c i a t e d w i t h a l e v e l 3 . Further, an object audiotrack with the track ID of 5 is associated with a level 4. 100791 (First Description Example of MPD File) Fig. 15 is a diagram illustrating a first description 10 example of the MPD file. [0080] As illustrated in Fig. 15, in the MPD file, "Representation" that manages the segment of the audio file (3daudio.mp4) of the 3D audio, "SubRepresentation" that 15 manages the track included in the segment, and the like are described. [OOSl] The "Representation" and the "SubRepresentation" include "codecs" that indicates the kind (profile or level) 20 of codec of the corresponding segment as a whole or the track in a 30 audio file format. [0082] The "SubRepresentation" includes a "level" that is a value set in the level assignment box as a value that indicates 25 the level of the corresponding track. "SubRepresentation" includes "dependencylevel" that is avalue indicating the level corresponding to another track (hereinafter, referred to as reference track) having the reference relationship (having dependency) . 30 [0003] Further, the "SubRepresentation" includes is not described in "SubRepresentationn corresponding to the group. The same, applies to Figs. 24, 25, 31, 39, 45, 47, 48, and 50, described I below. [00861 30 (Configuration Example of File Generation Device) Fig. 16is a b l o c k d i a g r a m i l l u s t r a t i n g a c o ~ ~ f i g u r a t i o n example of t h e f i l e g e n e r a t i o n d e v i c e 1 4 1 of Fig. 8. 100871 The f i l e g e n e r a t i o n d e v i c e 1 4 1 of Fig. 1 6 i s configured from an audio encoding processing u n i t 171, an audio f i l e 5 g e n e r a t i o n u n i t 1 7 2 , anMPD g e n e r a t i o n u n i t 173, a n d a s e r v e r upload processing u n i t 174. [a0881 The audio encoding processing u n i t 171 of t h e f i l e g e n e r a t i o n device 1 4 1 e n c o d e s t h e audio d a t a a n d t h e m e t a d a t a 10 of t h e 3D audio of t h e moving image content a t a p l u r a l i t y of encoding speeds t o g e n e r a t e t h e audio streams. The audio e n c o d i n g p r o c e s s i n g u n i t 1 7 l s u p p l i e s t h e a u d i o s t r e a m o f e a c h encoding speed t o t h e audio f i l e g e n e r a t i o n u n i t 172. [0089] 15 The audio f i l e g e n e r a t i o n u n i t 172 a l l o c a t e s t h e t r a c k t o t h e audio stream supplied from t h e audio encoding processing u n i t 171 f o r each group and each type of t h e Extelement. The audio f i l e generation u n i t 172 g e n e r a t e s t h e audio f i l e i n t h e segment s t r u c t u r e of Fig. 12 o r 13, i n which t h e audio 20 streams o f t h e t r a c k s a r e arrange d i n u r i i t s o f t h e subseyment, f o r eachencodirig speed and segment. The audio f i l e g c r ~ e r a t i o n u n i t 172 s u p p l i e s t h e g e n e r a t e d a u d i o f i l e t o t h e MPD g e n e r a t i o n u n i t 173. [0090] 2 5 The MPD g e n e r a t i o n u n i t 173 determines t h e URL of t h e web s e r v e r 1 4 2 inwhich t h e audio f i l e s u p p l i e d fromthe audio f i l e g e n e r a t i o n u n i t 1 7 2 i s t o b e s t o r e d , a n d t h e l i k e . Then, t h e MPD generation u n i t 173 g e n e r a t e s t h e MPD f i l e i n which t h e URL of t h e audio f i l e and t h e l i k e a r e arranged i n t h e 30 "Segment" of t h e "Representation" f o r t h e audio f i l e . The MPD g e n e r a t i o n u n i t 173 s u p p l i e s t h e generated MPD f i l ~ ean d t h e audio f i l e t o t h e s e r v e r upload p r o c e s s i n g u n i t 174. [0091] The s e r v e r u p l o a d p r o c e s s i n g u n i t 174 u p l o a d s t h e audio f i l e and t h e MPD f i l e supplled from t h e MPD g e n e r a t i o n u n i t 5 173 onto t h e web s e r v e r 142. [0092] ( D e s c r i p t i o n of Processing of F i l e Generation Device) Fig. 1 7 i s a flowchart f o r d e s c r i b i n g f i l e g e n e r a t i o n p r o c e s s i n g of t h e f i l e g e n e r a t i o n device 1 4 1 of Fig. 16. 10 [0093] In s t e p 5191 of Fig. 17, t h e audio encoding p r o c e s s i n g u n i t 171 encodes t h e audio d a t a and t h e metadata of t h e 3D audio of t h e moving image c o n t e n t a t a p l u r a l i t y of encoding speeds t o g e n e r a t e t h e audio streams. The audio encoding 15 p r o c e s s i n g u n i t 1 7 1 s u p p l i e s t h e a u d i o s t r e a m o f e a c h e n c o d i n g speed t o t h e audio f i l e g e n e r a t i o n u n i t 172. [0094] I n s t e p 5192, t h e audio f i l e g e n e r a t i o n u n i t 172 a l l o c a t e s t h e t r a c k t o t h e a u d i o streamsuppliedfromtheaudio 20 encoding p r o c e s s i n g u n i t 171 f o r each group and each t y p e of t h e Extelement. [0095] In s t e p S193, t h e audio f i l e g e n e r a t i o n u n i t 172 g e n e r a t e s t h e audio f i l e i n t h e segment s t r u c t u r e of Fig. 12 25 o r 13, i n which t h e audio streams of t h e t r a c k s a r e arranged i n u n i t s of t h e subsegment, f o r each encoding speed and segment. The audio f i l e g e n e r a t i o n u n i t 172 s u p p l i e s t h e generated audio f i l e t o t h e MPD g e n e r a t i o n u n i t 173. [0096] 3 0 In s t e p S194, theMPDgeneration u n i t 1 7 3 g e n e r a t e s t h e MPD f i l e i n c l u d i n g t h e URL of t h e audio f i l e and t h e l i k e . The MPD generation unit 173 supplies the generated MPD file and the audio file to the server upload processing unit 174. [00971 In step 5195, the server upload processing unit 174 5 uploads the audio file and the MPD file supplied from the MPD generation unit 173 onto the web server 142. Then, the processing is terminated. [00981 (Functional Configuration Example of Moving Image 10 Reproduction Terminal) Fig. 18 is ablockdiagramillustratingaconfiguration example of a streaming reproduction unit realized such that themoving image reproduction terminal144 of Fig. 8 executes the control software 161, the moving image reproduction 15 software 162, and the access software 163. [0099] A streaming reproduction unit -190 of Fig. 18 is configured fromanMPD acquisitionunit 91, anMPDprocessing unit191, a n a u d i o f i l e a c q u i s i t i o n u n i t 1 9 2 , anaudiodecoding 20 processing unit 194, and an audio synthesis processing unit 195. [ OlOO] TheMPDacquisitionunit 91 of the streaming reproduction unit 190 acquires the MPD file from the web server 142, and 25 supplies the MPD file to the MPD processing unit 191. [01011 The MPD processing unit 191 extracts the information of the URL of the audio file of the segment to be reproduced described in the "Segment" for the audio file, and the like, 30 from the MPD file supplied from the MPD acquisition unit 91, and supplies the information to the audio file acquisition u n i t 192. [01021 The audio f i l e a c q u i s i t i o n u n i t 192 r e q u e s t s t h e web s e r v e r 142 and a c q u i r e s t h e audio stream of t h e t r a c k t o be 5 r e p r o d u c e d i n t h e audio f i l e i d e n t i f i e d w i t h t h e URL s u p p l i e d fromtheMPDprocessingunit 191. The audio f i l e a c q u i s i t i o n u n i t 192 s u p p l i e s t h e a c q u i r e d a u d i o stream t o t h e audio decoding processing u n i t 194. [01031 10 Theaudiodecodingprocessingunit194 d e c o d e s t h e a u d i o streamsuppliedfromtheaudio f i l e a c q u i s i t i o n u n i t 1 9 2 . The audio decoding p r o c e s s i n g u n i t 194 s u p p l i e s t h e audio d a t a obtained a s a r e s u l t of t h e decoding t o t h e audio s y n t h e s i s p r o c e s s i n g u n i t 1 9 5 . T h e a u d i o s y n t h e s i s p r o c e s s i n g u n i t 1 9 5 15 s y n t h e s i z e s t h e audio data s u p p l i e d f r o m t h e audio decoding p r o c e s s i n g u n i t 194, a s needed, and o u t p u t s t h e audio d a t a . [0104] theaudiodecodingprocessingunit194, a r i d t h e a u d i o s y n t h e s i s 20 p r o c e s s i n g u n i t 195 f u n c t i o n a s a r e p r o d u c t i o n u n i t , and a c q u i r e and reproduce t h e audio stream of t h e t r a c k t o be reproduced from t h e audio f i l e s t o r e d i n t h e web s e r v e r 142. [01051 ( D e s c r i p t i o n of Processing of Moving Image Reproduction 25 Terminal) Fig. 19 is a flowchart f o r d e s c r i b i n g reproduction p r o c e s s i n g o f t h e streaming reproduction u n i t 190 of Fig. 18. [0106] In s t e p S211 of Fig. 1 9 , t h e MPD a c q u i s i t i o n u n i t 91 30 of t h e streaming reproduction u n i t 190 a c q u i r e s t h e MPD f i l e from t h e web s e r v e r 142, and s u p p l i e s t h e MPD f i l e t o t h e MPD processing unit 191. [01071 In step S212, the MPD processing unit 191 extracts the information of the URL of the audio file of the segment to 5 be reproduced described in the "Segment" for the audio file, andthe like, fromtheMPDfile suppliedfromtheMPDacquisition unit 91, and supplies the information to the audio fiie acquisition unit 192. [0108] 10 In step 5213, the audio file acquisition unit 192 requests the web server 142 and acquires the audio stream of the track to be reproduced in the audio file identified by the URL on the basis of the URL supplied from the MPD processing unit 191. The audio file acquisition unit 192 supplies the 15 acquired audio stream to the audio decoding processing unit 194. [0109] In step S214, the audio decoding processing unit 194 decodes the audio stream supplied from the audio file 20 acquisitionunit192. The audiodecodingprocessingunit194 supplies the audio data obtained as a result of the decoding to the audio synthesis processing unit 195. In step 5215, the audio synthesis processingunit195 synthesizesthe audio data supplied from the audio decoding processing unit 194, 25 as needed, and outputs the audio data. [OllO] (Outline of Second Example of Track of Audio File) Notethat,intheabovedescription,theGroupDefinition 30 However, as illustrated in Fig. 20, the GroupDefinition and the SwitchGroupDefinitionmaybe arrangedin sample groupe~try t h a t is the sample e n t r y of each group of a subsample i n t h e t r a c k . [ O l l l ] I n t h i s c a s e , a s i l l u s t r a t e d i n Flg. 21, t h e sample group 5 e n t r y oE t h e t r a c k of t h e group t h a t forms t h e switch Group i n c l u d e s t h e GroupDefinition and t h e SwitchGroupDefj.nition. Although i l l u s t r a t i o n is omitted, t h e sample group e n t r y of t h e t r a c k of t h e group t h a t does not form t h e switch Group i n c l u d e s only t h e GroupDefinition. 10 [O1121 F u r t h e r , t h e sample e n t r y of t h e t r a c k of each of t h e groups becomes one i l l u s t r a t e d i n Fig. 22. That is, a s i l l u s t r a t e d i n Fig. 22, i n t h e sample e n t r y of t h e t r a c k of each o f t h e groups, MHAGroupAudioConfigrationBox i n which 15 Config information such as a p r o f i l e (MPEGHAudioProfile) of t h e audio stream of t h e corresponding t r a c k , a l e v e l (MPEGHAudioLevel), and t h e l i k e a r e d e s c r i b e d . [01131 (Outline of Third Example of Track of Audio F i l e ) 2 0 Pig. 2 3 i s a d i a g r a m f o r d e s c r i b i n g a n o u t l i n e o f a t h i r d example of t h e t r a c k of t h e audio f i l e . [0114] The c o n f i g u r a t i o n o f t h e t r a c k o f t h e a u d i o d a t a o f Fig. 23 is d i f f e r e n t from t h e c o n f i g u r a t i o n o f F i g . 9 i n t h a t t h e 25 audiostreamsofoneormoregroups o f t h e 3 D a u d i o a r e i n c l u d e d i n t h e base t r a c k , and t h e number o f groups corresponding t o t h e audio streams divided i n t o t h e t r a c k s ( h e r e i n a f t e r , r e f e r r e d t o a s group t r a c k s ) t h a t donot i n c l u d e t h e information r e l a t e d t o t h e 3D audio a s a whole is 1 or more. 30 [01151 That i s , t h e sample e n t r y of t h e base t r a c k of Fig. 23 i s the sample entry with the 4cc of "mha2", which includes the syntax for base track of when the audio streams of the audio data, of the 30 audio, are divided i n t o a p l u r a l i t y of tracks and arranged, s i m i l a r l y t o Fig. 9 (Fig. 1 0 ) . 5 [0116] Further, the sample entry of the group track is the sample entry with the 4cc of "mhgl", which includes the syntax for group track of when the audio streams of the audio data, of the 3Daudi0, are dividedinto a p l u r a l i t y of tracks andarranged, 10 s i m i l a r l y t o Fig. 9 (Fig. 11). Therefore, the base track and the group track are i d e n t i f i e d with the 4cc of the sample entry, and dependency between the tracks can be recognized. [0117] Further, s i m i l a r l y t o Fig. 9, the Track Reference is 15 arranged in t h e t r a c k box of each of the t r a c k s . Therefore, even i n a case where which of "mha2" and "mhgl" is the 4cc of the sample entry of the base track or the group track is unknown, the dependencybetweenthetracks canbe recognized with the Track Reference. 20 [01181 Note that themhgC box andthemhsc boxmay not be described i n the sample entry of the group t r a c k . Further, in a case where the mhaC box including the config information of a l l o f t h e groups o f t h e 3D audio is d e s c r i b e d i n t h e sample entry 25 of the base t r a c k , the mhaC box may not be described in the sample entry of the group t r a c k . However, in a case where the mhaC box including the config information t h a t can independently reproduce the base track is described in the sample entry of the base t r a c k , t h e mhaC box including the 30 configinformationthatcanindependentlyreproduccthegroup track i s described in the sample entry of the group t r a c k . Whether it is in the former state or in the latter state can be recognized according to existence/non-existence of the config information in the sample entry. However, the recognition can be made by describing a flag in the sample 5 entry or by changingthetype ofthe sample entry. Note that, although illustration is omitted, in a case of making the former state and the latter state recognizable by changing the type of the sample entry, the 4cc of the sample entry of the base track is "mha2" in the case ofthe former state, and is "mha4" 10 in the case of the latter state. [Oil91 (Second Description Example of MPD File) Fig. 24 is adiagramillustratinga descriptionexample o f t h e M P D f i l e i n a c a s e w h e r e t h e configurationofthe tracks 15 of the audio file is the configuration of Fig. 23. [0120] The MPD file of Fig. 24 is different from the MPD file of Fig. 15 in that the "SubRepresentation" of the base track is described. 20 [0121] In the "SubRepresentation" of the base track, the "codecs", the "level", the "dependencyLeve1" of the base track, and 25 are described, similarly to the "SubRepresentation" of the group track. [0122] Intheexampleof Fig. 24, t h e " c o d e c s " o f t h e b a s e t r a c k is "mha2.2.1r', andthel'level" is "0" as avaluethatindicates 30 the level of the base track. The "dependencyLevel" is "1" and "2" as valuesthat indicate the levels ofthe group track. Further, the "dataTypeU is "3" as a number that indicates Lhe AudioScene information as a kind described in the mhas box of the sample entry of the base track, and the "definition" is binary data of the AudioScene information encoded by the 5 base64 method. [0123] Note that, as illustrated in Fig. 25, in the "SubRepresentation" of the base track, the AudioScene information may be divided and described. 10 [0124] In the example of Fig. 25, "1" is set as a number that indicates "Atmo" as a kind, "Atmo" indicating content of the group with the group ID "I", of the AudioScene information (Fig. 7) described in the mhas box of the sample entry of the 15 base track. [0125] Further, "2" to "7" are set as numbers that respectively indicate, as kinds, "Dialog EN" that indicates the content ofthe groupwith the group ID "2", "Dialog FRWthatindicates 20 the content of the group with the group ID "3", "Voiceover GE" that indicates the content of the group with the group ID1'4", "Effects" that indicate sthe content ofthe groupwith the group ID "5", "Effect" that indicates the content of the group with the group ID "6", and "Effect" that indicates the 25 content of the group with the group ID "7". [01261 Therefore, inthe "SubRepresentation" ofthebasetrack of Fig. 25, is not described in the "SubRepresentation" of the group track. [0135] Note that, although illustration is omitted, the Audioscene information may be divided and described in the "SubRepresentation" of the base track, similarly to the case 15 of Fig. 25. [0136] (Outline of Fifth Example of Track of Audio File) Fig. 28 isadiagramfordescribinganoutlineofafifth example of the track of the audio file. 20 [01371 The configuration of the tracks of the audio data of Fig. 28 is different from the configuration of Fig. 23 in that the sample entries of the base track and the group track are the sample entry including syntax suitable for both of the 25 base track and the group track of a case where the audio streams of the audiodata, of the 3Daudi0, are dividedinto theplurality of tracks. I [01381 In the case of Fig. 28, both of the 4ccs of the sample 30 entries of the base track and the group track are "mha3" that I I is the 4cc of the sample entry including the syntax suitable for both of the base track and the group track. [0139] Therefore, similarly to the case of Fig. 26, the dependency between the tracks is recognized with the Track 5 Reference arranged in the track box in each of the tracks. Further, because the 4ccs of the sample entries are "mha3", the corresponding track being the track of when the audio streams of the audio data, of the 3D audio, are divided into the plurality of tracks and arranged can be recognized. 10 [0140] (Example of Syntax of Sample Entry Where 4cc is "mha3") Fig. 29 is adiagramillustratinganexampleof the syntax of the sample entry with the 4cc of "mha3". [0141] 15 AsillustratedinFig. 2 9 , t h e s y n t a x o f t h e s a m p l e e n t r y with the 4cc of "mha3" is syntax obtained by synthesizing the syntax of Fig. 10 and the syntax of -Fig. 11. LO1421 That is, in the sample entry with the 4cc of "mha3", 20 the mhaC box (MHAConfigration Box), the mhas box (MHAAudioSceneInfo Box), the mhgC box (MHAGroupDefinitionBox), the mhsC box (MHASwitchGropuDefinition Box), and the like are arranged. lo1431 25 In the mhaC box of the sample entry of the base track, the config information of all of the groups of the 3D audio or the config information that can independently reproduce the base track is described. Further, in the mhas box, the Audioscene information including the information related to 30 all of the groups and the switch Group of the 30 audio is described, andthemhgcbox andthemhscbox are notarranged. [01441 In a case where the mhaC box including the config information of a l l of the groups of the 31) audio i s described i n the sample entry of the base t r a c k , the mhaC box may not 5 be d e s c r i b e d i n t h e sample entry o f t h e grouptrack. However, in a case wherethemhacbox includingthe config information t h a t can independently reproduce the base track is described i n the sample entry of the base track, the mhaC box including the config information t h a t can independently reproduce the 10 group track is described in the sample entry of the group t r a c k . Whether it is in the former s t a t e or i n the l a t t e r s t a t e can be recognized according t o existence/non-existence of the config information in the sample entry. However, the former s t a t e and the l a t t e r s t a t e can be recognized by describing 15 a f l a g i n the sample entry, or by changing the type of the sample entry. Note t h a t , although i l l u s t r a t i o n is omitted, i n a case of making the former s t a t e and the l a t t e r s t a t e recognizable by changing the type of the sample entry, the 4ccs of the sample e n t r i e s of the base track and the group 20 t r a c k a r e , forexample, "mha3"inthecase o f t h e former s t a t e , and are "mha5" in the case of the l a t t e r s t a t e . FurLher, the mhasboxisnotarrangedinthe s a m p l e e n t r y o f t h e g r o u p t r a c k . The mhgC box and the mhsC box may be or may not be arranged. [01451 2 5 Notethat, a s i l l u s t r a t e d i n F i g . 30, i n t h e sampleentry of the base t r a c k , the mhas box, the mhgC box, and the mhsC box are arranged, and both of the mhaC box i n which config information t h a t can independently reproduce only the base track is described and the mhaC box including the config 30 information of a l l of the groups of the 3Daudiomay be arranged. In t h i s case, the mhaC box in which the config information of a l l of the groups of the 3 D audio are described, and the mhaC box in which config information t h a t can independently reproduce onlythebasetrackisdescribedare recognizedwith f l a g s included in these mhaC boxes. Further, i n t h i s case, 5 the mhaC box may not be described i n the sample entry of the group track. WhetherthemhaC box is d e s c r i b e d i n t h e sample entry of the group track can be recognized according t o existence/non-existence of the mhaC box i n the sample entry of the group t r a c k . However, whether themhaC box is described 10 i n the sample entry of the group track can be recognized by describing a f l a g in the sample entry, o r b y c h a n g i n g t h e t y p e of the sample entry. Note t h a t , although i l l u s t r a t i o n is omitted, i n a case o f m a k i n g w h e t h e r t h e m h a C b o x i s described inthesampleentryofthegrouptrackrecognizablebychanging 15 the type of the sample entry, the 4ccs of the sample e n t r i e s o f t h e b a s e t r a c k a n d t h e group t r a c k a r e , for example, "mha3" i n a case where the mhaC box is described i n the sample entry of the group t r a c k , and are "mha5" i n a case where the mhaC box is not described i n the sample entry of the group t r a c k . 20 Note t h a t , i n Fig. 30, the mhgC box and the mhsC box may not be described in the sample entry of the base t r a c k . [0146] (Fourth Description Example of MPD f i l e ) Fig. 3 1 i s a d i a g r a m i l l u s t r a t i n g a description example 25 o f t h e M P D f i l e i n a c a s e w h e r e t h e c o n f i g u r a t i o n o f t h e tracks of the audio f i l e is the configuration of Fig. 28 or 30. [0147] The MPD f i l e of Fig. 31 is d i f f e r e n t from the MPD f i l e of Fig. 24 i n t h a t the "codecs" of the "Representation" i s 30 "mha3.3.lU, and the "codecs" of the "SubRepresentation" i s "mha3.2.1". [ 0 1 4 8 ] Note t h a t , although i l l u s t r a t i o n is omitted, the AudioScene information may be divided and described in the "SubRepresentation" o f t h e base t r a c k , s i m l l a r l y t o t h e case 5 of Fig. 25. [0149] Further, in the above-description, the TrackReference is arranged i n the track box in each of the t r a c k s . However, the Track Reference may not be arranged. For example, Figs. 10 32to34 arediagramsrespectivelyillustratingcasesinwhich the Track Reference is not arranged in the track boxes of the tracks of the audio f i l e s of Figs. 23, 26, and 28. In the case of Fig. 32, the Track Reference is not arranged, but the 4ccs of the sample e n t r i e s of the base track and the group 15 traclc a r e d i f f e r e n t , and thus the dependency between t h e t r a c k s can be recognized. In the cases of Figs. 33 and 34, because the mhas box is arranged, whether the track is the base track can be recognized. [0150] 2 0 The MPD f i l e s of the cases where the configurations of the tracks of the audio f i l e are the configurations of Figs. 32 t o 34 are respectively the same as the MPD f i l e s of Figs. 24, 27, and 31. Note t h a t , i n t h i s case, the AudioScene information may be divided and described i n the 25 "SubRepresentation" of the base t r a c k , s i m i l a r l y t o the case of Fig. 25. [0151] (Outline of Sixth Example of Track of Audio F i l e ) Fig. 3 5 i s a d i a g r a m f o r d e s c r i b i n g a n o u t l i n e o f a sixth 30 example of the track of the audio f i l e . [0152] The configuration of t h e t r a c k s of the audio data of Fig. 35 i s d i f f e r e n t fromthe configuration of Fig. 33 i n t h a t the reference information to the' samples of the tracks of the groups and the config information necessary for decodingthe 5 reference information are not arranged i n the sample of the base t r a c k , theaudiostreamsof 0 ormore groupsareincluded, and the reference information t o the samples of the tracks of the groups is described i n the sample entry of the base t r a c k . 10 LO1531 Tobe s p e c i f i c , anmhmtboxthatdescribes which tracks thegroupsdescribedintheAudioSceneinformationaredivided i n t o is newly arranged in the sample entry with the 4cc of "mha2", which includes the syntax for base track of when the 15 audio streams of the audio data, of the 30 audio, are divided i n t o a p l u r a l i t y of t r a c k s . 101541 (Another Exampleof Syntaxof Sample Entrywith 4ccof "mha2") Fig. 36 is a diagram i l l u s t r a t i n g an example of syntax 20 of the sample e n t r i e s of the base track and the group track of Fig. 35 where the 4cc is "mha2". LO1551 The configuration of the sample entry with the 4cc of "mha2" of Fig. 36 is d i f f e r e n t from the configuration of Fig. 25 10 in t h a t an MHAMultiTrackDescription box (mhmt box) is arranged. (01561 In the mhmt box, as the reference information, corresponding r e l a t i o n s h i p between the group I D (group - I D ) 30 and the track I D (track - I D ) is described. Note t h a t , i n the mhmt box, the audio element and the track I D may be described in association with each other. [0157] Inacasewherethereferen'ceinformationis not changed in each sample, the reference information canbe efficiently 5 described by arranging the mhmt box in the sample entry. [0158] Note that, although illustration is omitted, in the cases of Figs. 9, 20, 23, 26, 28, 30, 32, and 34, the mhmt box can be similarly arranged in the sample entry of the back track, 10 insteadofdescribingthereferenceinformationtothesamples of the tracks of the groups, to the sample of the base track. [0159] In this case, the syntax of the sample entry with the 4cc of "mha3" becomes one illustrated in Fig. 37. That is, 15 the configuration of the sample entry with the 4cc of "mha3" of Fig. 37 is different from the configuration of Fig. 29 in thattheMHAMultiTrackDescriptionbox-(mhmtbox)i sarranged. [0160] Further, in Figs. 23, 26, 28, 30, 32 to 34, and 35, the 20 audio streams of one or more groups of the 3D audio may not be includedinthe base track, similarlyto Fig. 9. Further, the number of the groups corresponding to the audio streams divided into the group tracks may be 1. [0161] 2 5 Further, in Figs. 23, 26, 28, 30, 32 to 34, and 35, the GroupDefinitionandtheSwitchGroupDefinitionmaybearranged in the sample group entry, similarly to the case of Fig. 20. i01621 30 (Outline of Track) Fig. 38 is a diagramfordescribingan outline oftracks in a second embodiment to which the present disclosure is applied. LO1631 As illustrated in Fig. 38, the second embodiment is 5 different from the first embodiment in that the tracks are recorded as different files (3da - base.rnp4/3da_groupl.mp4/3da - group2.mp4/3da - group3.m~ 4/3da - group 4.mp4). In this case, by acquiring a file of a desiredtrackthrough an HTTP, only data ofthe desiredtrack 10 can be acquired. Therefore, the data of the desired track through the HTTP can be efficiently acquired. [01641 (Description Examples of MPD file) Fig. 39is a diagramillustrating description examples 15 of the MPD file in the second embodiment to which the present disclosure is applied. [0165] As illustrated in Fig. 39, in an MPD file, "Representation" that manages segments of the audio files 20 (3da - base.mp4/3da - groupl.mp4/3da - group2.mp4/3da - group3.m~ 4/3da - group 4 .mp4) of 3D audio, and the like are described. [0166] The "Representation" includes "codecs", "id", "associationId", and "assciationType". The "id" is an ID of 25 the"Representation"inc1udingthesame. The"association1d" is information indicating reference relationship between corresponding track and anothet track, and is "id" of a reference track. The "assciationTypet' is a code indicating meaning of reference relationship (dependency) with the 30 reference track, and for example, a value that is the same as a value of track reference of MP4 is used. [0167] Further, the "Representation" of the tracks of groups include . 5 In the example of Fig. 39, the "Representations" that manage the segments of the audio files are provided under one "AdaptationSet". However, the "AdaptationSet" may be provided for each of the segments of the audio files, and the "Representation" that manages the segment may be provided 10 thereunder. In this case, in the "AdaptationSet", the "associationId" and 5 is included in the "Representation" of the base track. To be specific, the "codecs" of the "Representation" of the base track of the MPD of Fig. 45 is "mha2.2.lW, and the "associationIdW is "gl" and "92" that are the "ids" of the group tracks. 10 [0200] Further, the MPD of Fig. 46 is different from the MPD of Fig. 45 in the "codecs" of the group track, and in that 15 is not included in the "Representation" of the group track. To be specific, the "codecs" of the group track of the MPD of Fig. 46 is "mha2.2.1". [0201] Further, the MPD of Fig. 47 is differenL from the MPD 20 of Fig. 45 in the "codecs" of the base track and the group track. To be specific, the "codecs" of the group Lrack of the MPD of Fig. 47 is "mha3.2.1". [02021 Note that, intheMPDof Figs. 45to 47, "Adaptationset" 25 can be divided for each "Representation", as illustrated in Figs. 48 to 50. [0203] In the above description, only one base track is provided. 30 However, a plurality of the base tracks may be provided. In this case, the base track is provided for each viewpoint of the 3D audio ( d e t a i l s w i l l be given below), for example, and i n the base t r a c k s , mhaC boxes including config information of a l l of the groups of the 3Daudio of the viewpoints are arranged. Notethat, i n t h e b a s e t r a c k s , mhasboxesincluding 5 the Audioscene informationofthe viewpointsmaybe arranged. [0204] The viewpoint of the 3D audio is a position where the 3D audio can be heard, such as a viewpoint of an image reproduced a t the same time with the 3Daudio or apredeterminedposition 10 s e t i n advance. [02051 As described above, i n a case where the base track is dividedforeachviewpoint, audiodifferentforeachviewpoint can be reproduced from the audio stream of the same 3D audio 15 on the basis of the position of an object on a screen and the l i k e included in the config information of each of the viewpoints. As a r e s u l t , a data amount of the audio streams of the 3D audio can be reduced. [0206] 2 0 That is, i n a case where the viewpoints of the 3D audio areapluralityofviewpoints of images of abaseball stadium, which can be reproduced a t the same time with the 30 audio, animagehavingaviewpointinacenterbackscreenisprepared as a main image t h a t is an image of a b a s i c viewpoint. Further, 25 images having viewpoints in a s e a t behind t h e p l a t e , a f i r s t - b a s e i n f i e l d bleacher s e a t , a third-base i n f i e l d bleacher s e a t , aleftoutfieldbleacherseat, a r i g h t o u t f i e l d bleacher s e a t , andthe l i k e are preparedas multi-imagesthat are images of the viewpoints other than the basic viewpoint. 30 102071 In t h i s case, i f tne 3D audio of a l l of the viewpoints is prepared, the data amount of the 3D audio becomes l a r g e . Therefore, by descrj-bing, t o the base t r a c k s , the posj.tions of the object on the screen and the l i k e i n the viewpoints, the audio streams such as Object audio and SAOC Object audio, 5 which are changed according t o the positions of the object on the screen, can be shared by the viewpoints. As a r e s u l t , the data arnount of the audio streams of the 3D audio can be reduced. [02081 10 A t t h e t i m e o f reproductionofthe 3Daudi0, for example, d i f f e r e n t audio i s reproducedaccording t o the viewpoint, using the audio streams such as the Object audio and the SAOC Object audio of the basic viewpoint, and the base track corresponding t o the viewpoint of the main image or the 15 multi-imagereproducedatthesametimewiththeaudiostream. [0209] Similarly, for example, i n a case where the viewpoints of the 3 D audio are positions of a p l u r a l i t y of s e a t s of a s t a d i u m s e t i n a d v a n c e , t h e d a t a a m o u n t o f t h e 3 D a u d i o b e c o m e s 20 large i f the 3D audio of a l l of the viewpoints is prepared. Therefore, by describing, t o the base t r a c k s , the positions o f t h e object onthe screen, intheviewpoints, the audio streams suchastheObjectaudioandtheSAOCObjectaudiocanbe shared by the viewpoints. As a r e s u l t , d i f f e r e n t audio can be 25 reproduced according t o t h e s e a t selected by the user using a s e a t i n g c h a r t , using the Object audio and the SAOC Object audio of one viewpoint, and the data amount of the audio streams of the 3D audio can be reduced. [02101 30 In a case where the base track is provided for each viewpoint of the 3D audio i.n t h e t r a c k s t r u c t u r e of Fig. 28, the track structure becomes one as illustrated in Fig. 51. In the example of Fig. 51, the number of viewpoints of the 3Daudioisthree. Further, intheexampleof Fig. 51, Channel audio is generated for each viewpoint of the 30 audio, and 5 other audiodata are sharedbythe viewpoints of the 3D audio. The same applies to the example of Fig. 52 described below. [02111 In this case, three base tracks are provided for each viewpoint of the 30 audio, as illustrated in Fig. 51. Track 10 Reference is arranged in the track box of each of the base tracks. Further, syntax of sample entry of each of the base tracks is the same as the syntax of the sample entry with 4cc of "mha3". The 4cc is "mhcf" that indicates that the base track is provided for each viewpoint of the 3D audio. 15 LO2121 The mhaC box including config information of all of groups of the 30 audio of each of the viewpoints is arranged in the sample entry of each ofthe base tracks. As the config information of all of the groups of the 3D audio of each of 20 the viewpoillts is the position of the object on the screen, in the viewpoiri l:, for example. Further, the mhas box including the Audioscene information of each of the viewpoints is arranged in each of the base tracks. I02131 2 5 The audio streams of the groups of the Channel audio i ofthe viewpoints are arrangedin samples ofthe base tracks. LO2141 I Note that, in a case where Object Metadata that describes I the position of the object on the screen, in each of the 30 viewpoints, inunits of a sample, exists, the Object Metadata is also arranged in the sample of each of the base tracks. [0215] That is, in a case where the object is a moving body (for example, a sport athlete),' the position of the object on the screenineachofthe viewpoints istemporallychanged. 5 Therefore, the position is described as Object Metadata in units of the sample. In this case, the Object Metadata in units of the sample is arranged, for each viewpoint, in the sample of the base track corresponding to the viewpoint. [0216] 10 The configurations of the group tracks of Fig. 51 are the same as the configuration of Fig. 28 except that the audio stream of the group of the Channel audio is not arranged, and thus description is omitted. [0217] 15 Note that, in the track structure of Fig. 51, the audio streams of the groups of the Channel audio of the viewpoints may not be arranged in the base track, and may be arranged in the different group tracks. In this case, the track structure becomes one illustrated in Fig. 52. 20 [0218] In the example of Fig. 52, the audio stream of i:he group of the Channel audio of the viewpoint corresponding to the base track with the track ID of ''1" is arranged in the group track with the track ID of "4". Further, the audio stream 25 of the group of the Channel audio of the viewpoint corresponding to the base track with the track ID of "2" is arranged in the group track with the track ID of "5". [0219] Further, the audio stream of the group of the Channel 30 audio of the viewpoint corresponding to the base track with the track ID of "3" is arranged in the group track with the tracli ID of "6". [02201 Note that, in the examples of Figs. 51 and 52, the 4cc of the sample entry of the base track is "mhcf". However, 5 the 4cc may be "mha3" that is the same as the case of Fig. 28. [02211 Further, although illustration is omitted, a case in which the base track is provided for each viewpoint of the 10 3Daudioin all ofthe above-describedtrack structures other than the track structure of Fig. 28 is similar to the case of Figs. 51 and 52. [02221 15 (Description of Computer to Which Present Disclosure Is Applied) The series of processing of the web server 142 (212) can be executed by hardware or can be executed by software. In a case of executing the series of processing by software, 20 a program that configures the software is installed to the computer. Here, the computer includes a computer incorporated in special hardware, and a general-purpose personal computerthat canexecute various types of functions by installing various types of programs, and the like. 25 [02231 Fig. 53is ablockdiagramillustrating a configuration example of hardware of the computer that executes the series of processing of the web server 142 (212) with a program. [02241 30 In the computer, a central processing unit (CPU) 601, a read only memory (ROM) 602, and a random access memory (RAM) 603 are mutually connected by a bus 604. 102251 An input/output interface605 is further connected to the bus 604. An input unit 606, an output unit 607, a storage 5 unit 608, acommunicationunit 609, andadrive 610 are connected to the input/output interface 605. [0226] The input unit 606 is made of a keyboard, a mouse, a microphone, and the like. The output unit 607 is made of a 10 display, a speaker, and the like. The storage unit. 608 is made of a hard dislc, a non-volatile memory, and the like. The communication unit 609 is made of a network interface, and the like. The drive 610 drives a removable medium 611 such as amagnetic disk, an optical disk, or amagneto-optical disk, 15 or a semiconductor memory. 102271 In the computer configured as described above, the CPU 601 loads the program stored in the storage unit 608 onto the RAM 603 through the input/output interface 605 and the bus 20 604, andexecutestheprogram, sothattheseriesofprocessing is performed. LO2281 The program executed by the computer (CPU 601) can be provided by being recorded in the removable medium 611 as a 25 package medium, for example. Further, the program can be providedthroughawiredorwirelesstransmissionmedium, such as a local area network, the Internet, or digital satellite broadcasting. [0229] 30 In the computer, the program can be installed to the storage unit 608 through the input/output interface 605 by a t t a c h i n g t h e r e m o v a b l e m e d i u m 6 1 1 t o t h e d r i v e 610. Further, the program can be received by the communicatiorl unit 609 throughawiredorwirelesstransmissionmedium, andinstalled to the storage unit 608. In addition, the program can be 5 installed to the ROM 602 or the storage unit 608 in advance. [0230] Note that the program executed by the computer may be a program processed in time series according to the order described in the present specification, or may be a program 10 processed in parallel or at necessary timing such as when called. [0231] Further, thehardware configurationofthemoving image reproduction terminal 144 (214) can have a similar 15 configuration to the computer of Fig. 53. In this case, for example, the CPU 601 executes the control software 161 (221), the moving image reproduction software 162, and the access software 163 (223). The processing of the moving image reproduction terminal 144 (214) car1 be executed by hardware. 20 [02321 In the present specification, a systemmeans a co:llective of a plurality of configuration elements (devices, modules (components), and the like), and all of the configuration elements may or may not be in the same casing. Therefore, 25 bothofapluralityofdevicesaccommodatedinseparatecasings and connected via a network, and a single device in which a plurality of modules are accommodated in a single casing are the systems. [0233] 3 0 Note that embodiments ofthepresent disclosure arenot limited to the above-described embodiments, and various changes canbemadewithoutdepartingfromthe spirit andscope of the present disclosure. [0234] Further, the present disclosure can be applied to an 5 information processing system that performs broadcasting or local storage reproduction, instead of streaming reproduction. [0235] In the embodiments of the MPD, the information is 10 describedby EssentialProperty having descriptor definition that can be ignored when the content described by the schema cannot be understood. However, the information may be described by SupplementalProperty having descriptor definition that can be reproduced even if the content described 15 by the schema cannot be understood. This descriptionmethod is selectedby the side that creates the content with intention. [02361 Further, the present disclosure can employ the configurations like below. 20 [0237] (1) An information processing apparatus including: a file generation unit configured to generate a file in which audio data of a plurality of kinds is divided into 25 tracks for each one or more of the kinds and arranged, and information related to the plurality of kinds is arranged. (2) The informationprocessing apparatus according to (I), wherein 30 the information related to the plurality of kinds is arranged in sample entry of a predetermined track. ( 3 ) The informationprocessing apparatus according to (21, wherein the predetermined track is one of the tracks in which 5 theaudiodataofapluralityof kindsisdividedandarranged. (4) The information processing apparatus according to any one of (1) to ( 3 ) , wherein, for each of the tracks, information related tothe kind 10 corresponding to the track is arranged in the file. (5) The information processing apparatus according to (4), wherein, for each of the tracks, information related to an 15 exclusive reproduction kind made of the kind corresponding to the track, and the kind corresponding to the audio data exclusively reproduced from the audio data of the kind corresponding to the track is arranged in the file. (6) 2 0 The informationprocessing apparatus according to (5), wherein information related to the kind corresponding to the trackandtheinformationrelatedtoanexclusivereproduction kind are arrangedin sample entry ofthe corresponding track. 25 (7) The information processing apparatus according to (5) or (6), wherein the file generation unit generates a management file thatmanagesthefileincludinginformationindicatingwhether 30 the information related to an exclusive reproduction kind exists for each of the tracks. (8) The information processing apparatus according to any one of (1) to ( 7 1 , wherein reference information to the tracks corresponding to 5 the plurality of kinds is arranged in the file. (9) The informationprocessingapparatus according to (8), wherein the reference information is arranged in a sample of 10 the predetermined track. (10) The informationprocessing apparatus according to (9), wherein the predetermined track is one of the tracks in which 15 theaudiodataofapluralityof kindsisdividedandarranged. (11) The information processing apparatus according to any one of (1) to (lo), wherein information indicating reference relationship among 20 the tracks is arranged in the file. (12) The information processing apparatus according to any one of (1) to (11), wherein the file generation unit generates a management file 25 that manages the file including information indicating reference relationship among the tracks. (13) The information processing apparatus according to any one of (1) to (12), wherein 3 0 the file is one file. (14) The information processing apparatus according t o any one of (1) to ( 1 2 ) , wherein the f i l e is a f i l e of each of the tracks. (15) 5 Aninformationprocessingmethodincludingthe s t e p o f : by an information processing apparatus, generating a f i l e i n which audio data of a p l u r a l i t y of kinds is divided i n t o tracks for each one or more of the kinds and arranged, andinformationrelatedtothepluralityof k i n d s i s a r r a n g e d . 10 ( 1 6 ) An information processing apparatus including: areproductionunitconfiguredtoreproduce, froma f i l e i n which audio data of a p l u r a l i t y of kinds is divided i n t o tracks for each one or more of the kinds and arranged, and 15 information r e l a t e d t o the p l u r a l i t y of kinds is arranged, the audio data of a predetermined t r a c k . ( 1 7 ) Aninformationprocessingmethodincludingthe s t e p o f : by an information processing apparatus, reproducing, 20 fromfileinwl~ichaudiodataofapluralityofkindsisdivided i n t o tracks for each one or more of the kinds and arranged, andinformationrelatedtothepluralityofkindsis arranged, the audio data of a predetermined t r a c k . 25 REFERENCE SIGNS LIST [0238] 11 File generation device 192 Audio f i l e a c q u i s i t i o n unit 194 Audio decoding processing unit 30 195 Audio synthesis processing unit 211 File generation device 264 Audio file acquisition unit CLAIMS 1. (Amended) An information processing apparatus comprising: 5 a file generation unit configured to allocate a track to an audio stream made of one track including a plurality of groups, the groupbeingindicatedwith a group IDandbeing configured from one or more audio elements, for each of the groups, to generate a file made of a plurality of the tracks. 10 2. (Amended) The information processing apparatus according to claim 1, wherein the file includes information indicating correspondence between the plurality of groups and the 15 plurality of tracks. 3. (Amended) The information processing apparatus according to claim 2, wherein the information indicating correspondence between the 20 plurality of groups and the plurality of tracks includes the group IDS of Llie plurality of groups. 4. (Amended) The information processing apparatus according to claim 2, wherein 2 5 the information indicating correspondence between the plurality of groups and the plurality of tracks includes the group IDS of the plurality of groups and track IDS corresponding to the plurality of tracks. 30 5. (Amended) The information processing apparatus according to claim 2, wherein the information indicating correspondence between the plurality of groups and the plurality of tracks is included in a base track. 5 6. (Amended) The information processing apparatus according to claim 2, wherein t h e f i l e g e n e r a t i o n u n i t s e t s t h e i n f o r m a t i o n i n d i c a t i n g correspondence between the plurality of groups and the plurality of tracks to a box different from audio scene 10 information related to the plurality of groups and config information of the plurality of groups. 7. (Amended) The information processing apparatus according to claim 1, wherein 15 information related to the plurality of groups is arrangedin sample entry of apredeterminedtrackinthe file. 8. (Amended) The information processing apparatus according to claim 1, wherein, 2 0 foreachofthetracks, informationrelatedtothe group corresponding to the track is arranged in the file. 9. (Amended) The information processing apparatus according to claim 1, wherein, 2 5 for each of the tracks, information related to an exclusive reproduction groupmade ofthe group corresponding tothe track, andthegroupcorrespondingtothe audioelement exclusively reproduced from the audio element of the group corresponding to the track is arrangedin the file. 30 10. (Amended) The ififormation processing apparatus according to claim 9, wherein information related to the group corresponding to the trackandtheinformationrelatedtoanexclusivereproduction g r o u p a r e a r r a n g e d i n s a m p l e e n t r y o f t h e c o r r e s p o n d i n g t r a c k . 5 11. (Amended) The information processing apparatus according to claim 9, wherein the file generation unit generates a management file thatmanagesthefileincludinginformationindicatingwhether 10 the information related to an exclusive reproduction group exists for each of the tracks. 12. (Amended) The information processing apparatus according to claim 1, wherein 15 reference information to the plurality of tracks is arrangedin the file. 13. (Amended) The information processing apparatus according to claim 12, wherein 2 0 the reference information is arranged in a sample of a predetermined track. 14. (Amended) The information processing apparatus according to claim 1, wherein 2 5 information indicating reference relationship among the tracks is arranged in the file. 15. (Amended) The information processing apparatus according to claim 1, wherein 3 0 the file generation unit generates a management file that manages the file including information indicating reference relati~onship among the tracks 16. (Amended) The information processing apparatus according to claim 1, wherein 5 the fi.le is one file. 17. (Amended) The information processing apparatus according to claim 1, wherein the file is a file of each of the tracks. 10 18. (Added) An information processing method comprising the step of: allocating atrack to an audio streammade of one track including a plurality of groups, the group being indicated 15 with a group ID and being configured from one or more audio elements, for each of the groups, to generate a file made of a plurality of the tracks. 19. (Added) An information processing apparatus comprising: 2 0 a reproduction unit configured to reproduce a predeterminedtrackfroma filemade ofapluralityoitracks, the file being generated by allocating, for each of groups, atrack to an audio streammade of one trackincludingaplurality of the groups, the group being indicated with a group ID and 25 being configured from one or more audio elements. 20. (Added) An information processing method comprising the step of: reproducing a predetermined track from a file made of 30 apluralityoftracks, the filebeinggeneratedbyallocating, for each of groups, a track to an audio stream made of one track including plurality of the groups, the group being indicated with a group ID and being configured from one or more audio elements.

Documents

Application Documents

# Name Date
1 Priority Document [22-12-2016(online)].pdf 2016-12-22
2 Power of Attorney [22-12-2016(online)].pdf 2016-12-22
3 Form 5 [22-12-2016(online)].pdf 2016-12-22
4 Form 3 [22-12-2016(online)].pdf 2016-12-22
5 Form 1 [22-12-2016(online)].pdf 2016-12-22
6 Drawing [22-12-2016(online)].pdf 2016-12-22
7 Description(Complete) [22-12-2016(online)].pdf_275.pdf 2016-12-22
8 Description(Complete) [22-12-2016(online)].pdf 2016-12-22
9 201617043896.pdf 2016-12-26
10 abstract.jpg 2017-01-24
11 Other Patent Document [28-01-2017(online)].pdf 2017-01-28
12 201617043896-OTHERS-300117.pdf 2017-02-01
13 201617043896-Correspondence-300117.pdf 2017-02-01
14 Form 3 [03-05-2017(online)].pdf 2017-05-03
15 201617043896-FORM 3 [25-07-2017(online)].pdf 2017-07-25
16 201617043896-FORM 18 [07-05-2018(online)].pdf 2018-05-07
17 201617043896-FER.pdf 2021-10-17

Search Strategy

1 searchstrategyE_28-04-2020.pdf