Sign In to Follow Application
View All Documents & Correspondence

Information Processing Device And Information Processing Method

Abstract: The present invention pertains to an information processing device and an information processing method configured to enable an improvement in the acquisition efficiency of prescribed types of audio data among a plurality of types of audio data. Audio data for prescribed tracks among files arranged by separating a plurality of types of audio data into a plurality of tracks according to type is obtained in the present invention. The present invention can be applied to, for example, a file generation device for generating files, a web server for recording a file generated by the file generation device, or an information processing system composed of a video playback terminal for playing back a file.

Get Free WhatsApp Updates!
Notices, Deadlines & Correspondence

Patent Information

Application #
Filing Date
23 November 2016
Publication Number
34/2017
Publication Type
INA
Invention Field
ELECTRONICS
Status
Email
remfry-sagar@remfry.com
Parent Application

Applicants

SONY CORPORATION
1 7 1 Konan Minato ku Tokyo 1080075

Inventors

1. HIRABAYASHI Mitsuhiro
c/o SONY CORPORATION 1 7 1 Konan Minato ku Tokyo 1080075
2. CHINEN Toru
c/o SONY CORPORATION 1 7 1 Konan Minato ku Tokyo 1080075
3. YAMAMOTO Yuki
c/o SONY CORPORATION 1 7 1 Konan Minato ku Tokyo 1080075
4. SHI Runyu
c/o SONY CORPORATION 1 7 1 Konan Minato ku Tokyo 1080075

Specification

Technical field [0001]  The present disclosure relates to an information processing apparatus and an information processing method, and more particularly to an information processing apparatus and an information processing method which make it possible to improve the efficiency of obtaining a predetermined type of audio data of a plurality of types of audio data . Background technique [0002]  In recent years, the mainstream of the streaming service on the Internet has become a OTT-V (Over The Top Video). It is MPEG-DASH're beginning to spread as the basic technology - a (Moving Picture Experts Group phase Dynamic Adaptive Streaming over HTTP) (e.g., see Non-Patent Document 1). [0003]  In MPEG-DASH, distribution server prepared video data groups screen sizes and coding rates are different for the single video content, the optimum screen size and coding rate of the video playback device in response to the channel conditions by requesting data group, adaptive streaming can be realized. CITATION Non-patent literature [0004] 非特許文献1 : MPEG-DASH(Dynamic Adaptive Streaming over HTTP)(URL:http://mpeg.chiariglione.org/standards/mpeg-dash/media-presentation-description-and-segment-formats/text-isoiec-23009-12012-dam-1) Summary of the Invention Problems that the Invention is to Solve [0005]  However, it is not considered to improve the efficiency of obtaining a predetermined type of audio data of a plurality of types of audio data of the moving image content. [0006]  The present disclosure has been made in view of such circumstances, it is desirable to make it possible to improve the efficiency of obtaining a predetermined type of audio data of a plurality of types of audio data. Means for Solving the Problems [0007]  Acquiring information processing apparatus of the first aspect of the present disclosure, to obtain a predetermined audio data of the track of the files in which a plurality of types of audio data are arranged is divided into a plurality of tracks according to the type parts is an information processing apparatus including a. [0008]  An information processing method of the first aspect of the present disclosure correspond to the information processing apparatus of the first aspect of the present disclosure. [0009]  In the first aspect of the present disclosure, a plurality of types of audio data predetermined audio data of the track of the files that are arranged is divided into a plurality of tracks in accordance with the type is acquired. [0010]  The information processing apparatus of the second aspect of the present disclosure is an information processing apparatus including a generation unit which generates a file in which a plurality of types of audio data are arranged is divided into a plurality of tracks according to the type. [0011]  The information processing method of the second aspect of the present disclosure correspond to the information processing apparatus of the second aspect of the present disclosure. [0012]  In the second aspect of the present disclosure, the file in which a plurality of types of audio data are arranged is divided into a plurality of tracks according to the type is generated. [0013]  The information processing apparatus of the first and second aspects can be realized by causing a computer to execute a program. [0014]  Further, in order to realize an information processing apparatus of the first and second side surfaces, a program to be executed by a computer, by transmitting via a transmission medium or by being recorded on a recording medium, it is possible to provide . Effect of the Invention [0015]  According to a first aspect of the present disclosure, it is possible to obtain the audio data. Further, according to the first aspect of the present disclosure, a predetermined type of audio data of a plurality of types of audio data can be efficiently acquired. [0016]  According to a second aspect of the present disclosure, it can generate a file. Further, according to the second aspect of the present disclosure, it is possible to generate a file for improving the efficiency of obtaining a predetermined type of audio data of a plurality of types of audio data. BRIEF DESCRIPTION OF THE DRAWINGS [0017] FIG. 1 is a diagram illustrating an outline of a first example of the applied information processing system of the present disclosure. It is a diagram illustrating an example of FIG. 2 tiles. Is a diagram illustrating a FIG. 3 object. 4 is a diagram for explaining the object position information. 5 is a diagram for explaining the frame size information. Is a diagram illustrating the structure of FIG. 6] MPD file. [7] "Period" is a diagram showing the relationship of "Representation", and "Segment". Is a diagram showing the hierarchical structure of FIG. 8] MPD file. 9 is a diagram showing the relationship between the structure and the time axis of the MPD file. It is a diagram illustrating a description example of FIG. 10] MPD file. It is a block diagram showing a configuration example of FIG. 11 file generation device. It is a flowchart illustrating a file generating process in FIG. 12 file generation device. 13 is a block diagram showing a configuration example of a streaming playback unit. 14 is a flowchart illustrating a streaming playback process of streaming reproducing unit. It is a diagram illustrating another description example of FIG. 15] MPD file. 16 is a diagram explaining still another example of the description of MPD file. Is a diagram showing an example of the arrangement of FIG. 17 audio streams. Is a diagram illustrating a description example of FIG. 18] gsix. [Figure 19] Sample group Is a diagram illustrating an example of information indicating a correspondence relationship entry and the object ID. It is a diagram illustrating a description example of FIG. 20] AudioObjectSampleGroupEntry. Is a diagram illustrating a description example of FIG. 21] type assignment box. 22 is a diagram for describing the outline of a second example of the information processing system of the present disclosure. FIG. 23 is a block diagram showing a configuration example of streaming playback of the information processing system of the present disclosure. Is a diagram for explaining a method for determining the position of FIG. 24 objects. Is a diagram for explaining a method for determining the position of FIG. 25 objects. FIG. 26 is a diagram for explaining a method for determining the position of an object. [27] Horizontal angle theta Ai and horizontal angle theta Ai is a diagram showing the relationship '. [FIG. 28] is a flowchart illustrating a streaming playback process of streaming reproducing portion of FIG. 23. FIG. 29 is a flowchart illustrating the details of the position determination process in FIG. 28. [Figure 30] horizontal angle θ in FIG. 29 Ai is a flowchart illustrating the details of the 'estimation process. [FIG. 31] is a diagram for describing the outline of the track of the MP4 of 3D audio file formats. [FIG. 32] is a diagram showing a structure of a moov box. [33] is a diagram for describing the outline of the track in the first embodiment according to the present disclosure. [34] is a diagram showing an example of a syntax of sample entries in the base track of Figure 33. Is a diagram illustrating an example of FIG. 35 syntax of the sample entry channel audio track of Figure 33. [FIG. 36] is a diagram illustrating an example of syntax of sample entries of object audio tracks of FIG. 33. Is a diagram illustrating an example of syntax of the sample entry HOA audio tracks [37] Figure 33. [FIG. 38] is a diagram illustrating an example of syntax of sample entries of object metadata track of Figure 33. It is a diagram showing a first example of FIG. 39 segment structure. It is a diagram showing a second example of FIG. 40 segment structure. It is a diagram illustrating a description example of FIG. 41] level assignment box. Is a diagram illustrating a description example of MPD file in the first embodiment according to the FIG. 42 the present disclosure. [FIG 43 is a diagram showing the definition of essential Property. Is a diagram illustrating an outline of an information processing system in a first embodiment according to the FIG. 44 the present disclosure. It is a block diagram showing a configuration example of a file generation device of FIG 45 FIG 44. [FIG. 46] is a flowchart illustrating a file generation processing of the file generation apparatus of FIG. 45. [FIG. 47] is a block diagram showing a configuration example of a streaming playback unit realized by video playback device of FIG. 44. FIG 48 is a flowchart illustrating a channel audio reproduction process of the streaming reproduction of Figure 47. [49] is a flow chart for explaining the object identification processing of streaming playback of Figure 47. It is a flowchart illustrating a specific object audio reproduction process of streaming reproducing portion of FIG. 50 FIG. 47. [FIG. 51] is a diagram for describing the outline of the track in the second embodiment according to the present disclosure. [FIG. 52] is a diagram showing an example of a syntax of sample entries in the base track of Figure 51. [FIG. 53] is a diagram showing a structure of a base sample. Is a diagram illustrating an example of FIG. 54] based sample syntax. Is a diagram illustrating an example of data in FIG. 55] extractor. [FIG. 56] is a diagram for describing the outline of the track in the third embodiment according to the present disclosure. Is a diagram for describing the outline of the track in the fourth embodiment according to the FIG. 57 the present disclosure. It is a diagram illustrating a description example of MPD file in the fourth embodiment according to the FIG. 58] The present disclosure. Is a diagram illustrating an outline of an information processing system in the fourth embodiment according to the FIG. 59] The present disclosure. It is a block diagram showing a configuration example of a file generation device of FIG 60] FIG 59. [FIG. 61] is a flowchart illustrating a file generation processing of the file generation device of FIG. 60. Is [62] a block diagram showing a configuration example of a streaming playback unit realized by video playback device of FIG. 59. Is a flowchart illustrating an example of FIG. 63] channel audio reproduction process of streaming reproducing portion of FIG. 62. [FIG. 64] is a flowchart illustrating a first example of an object audio reproduction process of the streaming reproduction of Figure 62. It is a flowchart illustrating a second example of an object audio reproduction process of the streaming reproduction of FIG 65] FIG 62. [FIG. 66] is a flowchart illustrating a third example of an object audio reproduction process of the streaming reproduction of Figure 62. [FIG. 67] is a diagram showing an example of the selected object based on priorities. [FIG. 68] is a diagram illustrating an outline of a track in the fifth embodiment according to the present disclosure. [FIG. 69] is a diagram for describing the outline of the track in a sixth embodiment according to the present disclosure. Diagrams [Figure 70] shows the hierarchical structure of the 3D audio. [FIG. 71] is a diagram for explaining a first example of the processing of the Web server. [FIG. 72] is a flowchart illustrating a track division processing of the Web server. [FIG. 73] is a diagram for explaining a first example of the processing of the speech decoding processor. [FIG. 74] is a flowchart illustrating the details of a first example of decoding processing of the audio decoding process unit. [FIG. 75] is a diagram for explaining a second example of the processing of the audio decoding process unit. [FIG. 76] is a flowchart illustrating the details of a second example of decoding process of the audio decoding process unit. [FIG. 77] is a diagram for explaining a second example of the process of the Web server. [FIG. 78] is a diagram for explaining a third example of the processing of the speech decoding processor. [79] is a flowchart illustrating the details of a third example of the decoding process of the audio decoding processing section. It is a diagram showing a second example of FIG. 80] in the Config information arranged in the base sample syntax. [FIG. 81] is a diagram showing an example of a syntax of Config information for Ext element of Figure 80. [FIG. 82] is a diagram showing an example of a syntax of Config information for Extractor in Figure 81. Is a diagram showing a second example of FIG. 83] in the data of the frame units are disposed on the base sample syntax. [FIG. 84] is a diagram showing an example of a syntax of data Extractor in Figure 83. Is a diagram showing a third example of syntax of Config information arranged in FIG. 85] based samples. Is a diagram showing a third example of FIG. 86] in the data of the frame units are disposed on the base sample syntax. It is a diagram illustrating a configuration example of an audio stream in the seventh embodiment of an information processing system according to the FIG. 87] The present disclosure. [FIG. 88] is a diagram for describing the outline of the track in the seventh embodiment. [FIG. 89] is a flowchart illustrating a file generation processing according to the seventh embodiment. FIG 90 is a flowchart illustrating an audio reproduction process in the seventh embodiment. [FIG. 91] is a diagram for describing the outline of the track in the eighth embodiment is applied to an information processing system of the present disclosure. Is a diagram illustrating a configuration example of FIG. 92] the audio file. [FIG. 93] is a diagram showing another configuration example of the audio file. is a diagram illustrating yet another configuration example of FIG. 94] the audio file. [FIG. 95] is a block diagram showing a configuration example of hardware of a computer. DESCRIPTION OF THE INVENTION [0018]  Hereinafter, embodiments of the premise and the disclosure of the present disclosure (hereinafter, referred to as embodiments) will be described. The description will be made in the following order.  0. Premise of the present disclosure (FIGS. 1 to  30) 1. First Embodiment (FIGS. 31 to  50) 2. Second Embodiment (FIG. 51 through FIG.  55) 3. Third Embodiment (FIG.  56) 4. Fourth Embodiment (FIG. 57 through FIG.  67) 5. Fifth Embodiment (FIG.  68) 6. Sixth Embodiment (FIG.  69) Description of 7.3D audio hierarchy (Fig.  70) 8. Description of a first example of processing of the Web server (FIG. 71 and FIG. 72)  9. Description of a first example of processing of the audio decoding process section (FIGS. 73 and  74) 10. Description of a second example of processing of the audio decoding process section (FIG. 75 and FIG.  76) 11. Description of a second example of processing of the Web server (FIG.  77) 12. Description of the third example of the processing of the audio decoding processing section (FIGS. 78 and  79) 13. A second example of the base sample syntax (FIG. 80 through FIG.  84) 14. A third example of the base sample syntax (FIG. 85 and FIG.  86) 15. Seventh Embodiment (FIG. 87 through FIG.  90) 16. Eighth Embodiment (FIG. 91 through FIG.  94) 17. Ninth Embodiment (FIG. 95) [0019]    (Summary of first embodiment of the information processing system)  Fig. 1 is a diagram for describing the outline of a first example of the information processing system of the present disclosure. [0020]  The information processing system 10 of FIG. 1, Web server 12 and video reproducing terminal 14 to be connected to the file generating unit 11 is configured by being connected through the Internet 13. [0021]  In the information processing system 10, in a manner equivalent to MPEG-DASH, Web server 12 distributes the image data of the video content video reproducing terminal 14 with a tile unit (tiled streaming) to. [0022]  Specifically, the file generation unit 11 obtains the image data of the moving image content, generates a video stream encoded in tile. File generating apparatus 11, the video stream of each tile is filed for each time unit of 10 seconds from a few seconds to called segments. File generating unit 11 uploads the image file for each tile resulting in Web server 12. [0023]  The file generating unit 11 acquires the audio data of the moving image content for each object (details will be described later), and generates an audio stream encoded in object units. File generating unit 11, a file the audio stream of each object in units of segments, to upload audio file for each object obtained as a result to the Web server 12. [0024]  Note that the object is a sound source, audio data of each object is acquired by a microphone attached to the object. Object may be an object such as a fixed microphone stand, or may be a moving object such as a person. [0025]  File generating unit 11, the object position information indicating the (acquisition position of the audio data) position of each object (sound position information), encoding an audio metadata including object ID or the like which is a unique ID to the object. File generating unit 11, a file encoded data of audio metadata in segments, to upload audio metafile resulting in Web server 12. [0026]  The file generation unit 11, frame size information that represents the frame size of the image of the video content, including tile position information or the like indicating a position on the image of each tile, MPD to manage image files and audio files (Media to generate a Presentation Description) file (control information). File generation device 11 uploads the MPD file to the Web server 12. [0027]  Web server 12 stores uploaded image file from the file generator 11, sound files, audio metafile, and the MPD file. [0028]  In the example of FIG. 1, the Web server 12, a segment group consisting of an image file of a plurality of segments of the tiles the tile ID "1", segment group consisting of an image file of a plurality of segments of the tiles the tile ID "2" door is stored. Also, the Web server 12, stores the segment group consisting of audio files of a plurality of segments of the object of the object ID "1", and the segment group consisting of audio files of a plurality of segments of the object of the object ID "2" ing. Although not shown, it is stored similarly segment groups of audio metafile. [0029]  In the following, the tile tile ID is i called a tile #i, object ID is the object of i that object #i. [0030]  Web server 12 functions as a transmission unit, in response to a request from the video reproduction device 14, and transmits the image file storing voice files, audio metafile, the MPD file like the video playback device 14. [0031]  Video playback terminal 14, software (hereinafter, the control of software) for control of the streaming data 21, video playback software 22, HTTP (HyperText Transfer Protocol) client software for access (hereinafter, software for the access that) 23 to run, and the like. [0032]  Control software 21 is software that controls the data to be streamed from the Web server 12. Specifically, the control software 21, to acquire the MPD file to the video reproducing terminal 14 from the Web server 12. [0033]  Further, the control software 21, based a display area is an area for displaying the image of the moving image content to be instructed from the video playback software 22, to the tile position information included in the MPD file, the display area to identify the tile. Then, the control software 21 commands the transmission request for image file of the tile access software 23. [0034]  Further, the control software 21 commands the transmission request for the audio meta file to access the software 23. Then, the control software 21, the display region, on the basis of the object position information included frame size information included in the MPD file, and the audio meta file, to identify the object corresponding to the image in the display area. Then, the control software 21 commands the transmission request for the audio file of the object to access the software 23. [0035]  Video playback software 22 is software for reproducing image and audio files acquired from the Web server 12. Specifically, video playback software 22, when the display area is designated by the user, and instructs the display area to the control software 21. Also, video playback software 22 decodes the image file and the audio file acquired from a Web server 12 in response to the instruction is synthesized and output. [0036]  Software for accessing software 23 is software that controls the communication with the Web server 12 via the Internet 13 using HTTP. Specifically, access software 23, in accordance with the instruction of the control software 21, to send the image file, audio file, and a request for transmission of audio metafile video reproducing terminal 14. Also, access software 23, in response to the transmission request, the image file, to receive audio files, and the audio metafile video reproducing terminal 14 transmitted from the Web server 12. [0037]  (Example tile)  FIG. 2 is a diagram illustrating an example of a tile. [0038]  As shown in FIG. 2, the video content image is divided into a plurality of tiles, the tile ID is given from 1 to each tile sequentially. In the example of FIG. 2, the image of the video content is divided into four tiles # 1 to tile # 4. [0039]  (Description of the object)  FIG. 3 is a diagram for explaining the object. [0040]  In the example of FIG. 3, as the audio video content, audio eight objects in the image have been acquired, the object ID in order from 1 is given to each object. Object # 1 to object # 5 are each a moving object, object # 6 to object # 8 is an immovable object. Further, in the example of FIG. 3, the image of the video content is divided into tiles 5 (vertical) × 7 (horizontal). [0041]  In this case, as shown in FIG. 3, the display area 31 composed of the user by 2 (vertical) × 3 (lateral) tiles is designated, the display area 31, object # 1, object # 2, and only object # 6 is included. Thus, video playback terminal 14, for example, to retrieve the object # 1, object # 2, and only the sound files of the object # 6 from the Web server 12 and reproduces. [0042]  Objects in the display area 31, as described below, can be identified based on the frame size information and the object position information. [0043]  (Description of object position information)  FIG. 4 is a diagram for explaining the object position information. [0044]  As shown in FIG. 4, the object position information, the horizontal angle theta of objects 40 A (-180 ° ≦ theta A ≦ 180 °), the vertical angle gamma A (-90 ° ≦ gamma A ≦ 90 °), and the distance r a (0  (Outline of 3D audio file formats)  Before describing a first embodiment according to the present disclosure, first with reference to FIG. 31, outline of the tracks of MP4 of 3D audio file formats It will be described. [0204]  The MP4 file, for each track, it is possible to manage the position information indicating the position of the codec information and file of the video content. The MP4 of 3D audio file formats, all of the audio streams of the 3D audio (Channel audio / Object audio / HOA audio / metadata) (ES (Elementary Stream)) is recorded as one track in a sample (frame) unit. Also, 3D audio codec information (Profile / level / audio configuration) is stored as the sample entries (sample entry). [0205]  Channel audio constituting a 3D audio is the audio data of the channel units, Object audio is audio data of an object units, HOA audio are spherical audio data, metadata is, Channel audio / Object audio / HOA it is a meta data of audio. Here, Object audio is assumed to be audio data of the object units may be audio data of the SAOC. [0206]  (Structure of moov box)  FIG. 32 is a diagram showing a structure of a moov box of the MP4 file. [0207]  As shown in FIG. 32, the MP4 file, image data and audio data are recorded as different tracks. In Figure 32, it is not described details of the track of the audio data, the same track of the image data. sample entry is included in sample description disposed stsd box inside moov box. [0208]  Meanwhile, in the broadcasting and local storage playback, Web server side, and sends all audio streams, video reproducing terminal (client) while parsing all audio streams (, stream), an audio stream of 3D audio required decoded (decode-), outputs (rendering). And when the bit rate (Bitrate) is high, when there are restrictions on the rate of load of a local storage, by obtaining only the audio stream of the 3D audio required, it is desirable to reduce the load of the decode processing. [0209]  Further, the stream playback, video playback terminal (client), by obtaining only 3D audio encoded data required, be desirable to permit obtaining an audio stream of the optimal coding rate reproduction environment ing. [0210]  Accordingly, in the present disclosure, by arranging the audio file is divided into tracks according to the type of encoded data of a 3D audio and efficiently be obtained only a predetermined type of the encoded data. Accordingly, in the broadcasting and local storage playback, it is possible to reduce the load on the system. Further, the stream reproduction, depending on the band, can be reproduced having the highest quality of the encoded data of the 3D audio required. Further, since the need be recorded position information of the audio stream of 3D audio in the audio file in track units of sub-segments, as compared with the case where the encoded data of the object units are arranged in the sub-sample, the amount of position information it can be reduced. [0211]  (Overview of the track)  FIG. 33 is a diagram for describing the outline of the track in the first embodiment according to the present disclosure. [0212]  As shown in FIG. 33, in the first embodiment, Channel audio / Object audio / HOA audio / metadata constituting the 3D audio, respectively, different tracks (Channel audio track / Object audio track (s) / HOA audio track / Object is an audio stream metadata track). The object metadata track, is disposed an audio stream of audio metadata. [0213]  Further, as a track for arranging the information about the entire 3D audio, base track (Base Track) are provided. The base track of FIG. 33, but information about the entire 3D audio is placed on the sample entry, not arranged anything as a sample. Further, Base track, Channel audio track, Object audio track (s), HOA audio track, and Object metadata track is recorded as the same audio file (3dauio.mp4). [0214]  Track Reference is disposed, for example, track box, representing the reference relationship with other tracks in the corresponding track. Specifically, Track Reference is a unique ID to track other tracks in the reference relationship (hereinafter, referred to as a track ID) representing the. In the example of FIG. 33, Base track, Channel audio track, HOA audio track, Object metadata track, Object Track ID of audio track (s) has been 1,2,3,4,10 ... and. Further, Track Reference of Base track is 2,3,4,10 ..., Track Reference of Channel audio track / HOA audio track / Object metadata track / Object audio track (s) is a track ID of the Base track it is a certain one. [0215]  従って、Base trackと、Channel audio track/HOA audio track/Object metadata track/Object audio track(s)とは、参照関係にある。即ち、Base trackは、Channel audio track/HOA audio track/Object metadata track/Object audio track(s)の再生時に参照される。 [0216]  (Examples of the syntax of sample entries in the base track)  FIG. 34 is a diagram showing an example of a syntax of sample entries in the base track of Figure 33. [0217]  ConfigurationVersion in Figure 34, MPEGHAudioProfile, and MPEGHAudioLevel as information about the entire 3D audio, respectively, config information, profile information of the entire audio stream 3D audio (audio streams of the usual 3D audio), representing the level information. Further, width and height in FIG. 34, the information about the entire 3D audio, respectively, represent the number of pixels in the horizontal direction of the video content, the number of pixels in the vertical direction. theta1, theta2, gamma1, and gamma2 as information about the entire 3D audio, respectively, horizontal angle theta of the leftmost image frame of the frame size information for video content v1 , the right end of the horizontal angle theta v2 , the upper end of the picture frame vertical angle gamma v1 , vertical angle gamma of lower v2 represents a. [0218]  (Channel audio exemplary syntax of sample entries in a track)  FIG. 35 is a diagram showing an example of a syntax of the sample entry channel audio track (Channel audio track) in FIG. 33. [0219]  configurationVersion in Figure 35, MPEGHAudioProfile, and MPEGHAudioLevel, respectively, config information Channel Audio, profile information, representative of the level information. [0220]  (Syntax example of sample entries of object audio tracks)  FIG. 36 is a diagram showing an example of a syntax of sample entries of object audio tracks of FIG. 33 (Object audio track). [0221]  configurationVersion in Figure 36, MPEGHAudioProfile, and MPEGHAudioLevel, respectively, config information of one or more Object audio included in the object audio track, profile information, representative of the level information. object_is_fixed represents whether the object of one or more Object audio included in the object audio track is fixed. If object_is_fixed is 1, indicates that the object is fixed, if it is 0, indicating that the object moves. mpegh3daConfig represents the config identification information of an object of one or more Object audio included in the object audio track. [0222]  Further, objectTheta1 / objectTheta2 / objectGamma1 / objectGamma2 / objectRength represents object information of one or more Object audio included in the object audio track. This object information is valid information when Object_is_fixed = 1. [0223]  maxobjectTheta1, maxobjectTheta2, maxobjectGamma1, maxobjectGamma2 /, and maxobjectRength represents the maximum value of the object information when the object of one or more Object audio included in the object audio track is moved. [0224]  (HOA sample entries of exemplary syntax of an audio track)  FIG. 37 is a diagram showing an example of a syntax of a sample entry HOA audio track of FIG. 33 (HOA audio track). [0225]  configurationVersion in Figure 37, MPEGHAudioProfile, and MPEGHAudioLevel represent config information HOA audio, profile information, the level information. [0226]  (Object metadata exemplary syntax of sample entries in a track)  FIG. 38 is a diagram showing an example of a syntax of sample entries of object metadata tracks of FIG. 33 (Object metadata track). [0227]  configurationVersion in Figure 38 represents the config information metadata. [0228]  (First Example of segment structure of the audio files of 3D audio)  FIG. 39 is a diagram showing a first example of a segment structure of the audio files of 3D audio in the first embodiment according to the present disclosure. [0229]  The segment structure of FIG. 39, Initial segment is composed of ftyp box and moov box. The moov box, trak box is placed in each track included in the audio file. Further, the moov box, a track ID of the track, is arranged mvex box including information representing the correspondence between the level used in ssix box in media segment. [0230]  Also, media segment is, SIDX box, Ssix box, and constituted by one or more Subsegment. The sidx box, position information indicating a position in the audio file of each subsegment is located. The ssix box includes position information of each level of the audio stream that is arranged in mdat box. Note that level, which corresponds to the track. Also, the position information of the first track, the position information of the data consisting of audio streams moof box and the first track. [0231]  Subsegment is provided for each arbitrary time length, the Subsegment, pairs common set of moof box and mdat boxes to all tracks are provided. The mdat box, all track audio stream is arranged arbitrary time length only collectively, the moof box, management information of the audio stream is arranged. Audio stream for each track arranged in mdat box is continuous for each track. [0232]  In the example of FIG. 39, a track ID is 1 Track1 is base Track, Track2 to trackN track ID is 2 to N, respectively, Channel Audio Track, Object audio track (s), HOA audio track, object metadata is a track. This is the same in FIG. 40 to be described later. [0233]  (Second example of a segment structure of the audio files of 3D audio)  FIG 40 is a diagram showing a second example of a segment structure of the audio files of 3D audio in the first embodiment according to the present disclosure. [0234]  Segment structure of FIG. 40, that the moof box and mdat boxes is provided for each track is different from the segment structure of FIG. 39. [0235]  That, Initial segment of Figure 40 is similar to the Initial segment of Figure 39. Also, media segment of Figure 40 is similar to the media segment in FIG. 39, SIDX box, Ssix box, and constituted by one or more Subsegment. The sidx box, like the sidx box in Figure 39, the position information of each subsegment is located. The ssix box includes position information of each level of data consisting moof box and mdat boxes. [0236]  Subsegment is provided for each arbitrary time length, the Subsegment, moof box and mdat boxes pair is provided for each track. That is, the mdat box of each track, the audio stream of the track, any time length only collectively arranged (interleaved storage), the moof box, management information of the audio stream is arranged. [0237]  As shown in FIGS. 39 and 40, the audio streams of each track, for placement any time length collectively, as compared to when placed together in sample units, the audio stream over HTTP, etc. acquisition efficiency is improved. [0238]  (Description example of mvex box)  FIG. 41 is a diagram showing a description example of a level assignment box disposed mvex box in Figure 39 and Figure 40. [0239]  level assignment box is a box for associating a level (level) to be used in the track ID and ssix box of each track. In the example of FIG. 41, the base track track ID is 1 is associated with level 0, channel audio tracks is associated with a level one track ID is 2. Further, HOA audio track track ID is 3 is associated with level 2, the object metadata track is associated with the level 3 track ID is 4. Furthermore, object audio track is associated with the level 4 track ID is 10. [0240]  (Description example of MPD file)  FIG. 42 is a diagram showing a description example of MPD file in the first embodiment according to the present disclosure. [0241]  As shown in FIG. 42, the MPD file manages a segment of the audio file (3daudio.mp4) of 3D audio "Representation", and "SubRepresentation" for managing the tracks included in the segment are described. [0242]  The "Representation" to "SubRepresentation" is a type of codec for the corresponding segment or track includes "codecs" represented by code defined in the 3D audio file format. In addition, in the "Representation" includes "id", "associationId", and "assciationType". [0243]  "Id" is the ID of "Representation" that contains it. "AssociationId" is information that references the corresponding track and the other track is "id" in the reference track. "AssciationType", see the relationship between the reference track is a code that represents the meaning of (dependencies), the same as is used with the value of example MP4 of track reference. [0244]  The "SubRepresentation" includes a value set by the level assignment box as a value representing a level corresponding to the corresponding track "level". The "SubRepresentation" has a reference relationship (dependent) another track (hereinafter, referred to as reference track) is a value representing a level corresponding to include "dependencyLevel". [0245]  In addition, the "SubRepresentation" is, as the information necessary for selection of 3D audio, are included. [0246]  In addition, the "SubRepresentation" of Object audio track is, are included. The theta, gamma, r, when the object corresponding to the "SubRepresentation 'are fixed, respectively, horizontal angle, vertical angle of the object position information is the distance. On the other hand, if the object moves, the theta, gamma, r, respectively, the maximum value of the horizontal angle of the maximum value of the object position information, the maximum value of the vertical angle, the maximum value of the distance. [0247]  Figure 43 is a diagram showing the definition of the Essential Property of Figure 42. [0248]  Upper left side of FIG. 43, indicates the definition of AudioType of. AudioType represents the type of the 3D audio corresponding track. [0249]  In the example of FIG. 43, when AudioType is 1, the audio data of the corresponding track, indicates that the Channel audio of the 3D audio, if AudioType is 2, audio data of the corresponding track, HOA indicating that it is the audio. Also, if AudioType is 3, the audio data of the corresponding track, indicates that the Object audio, if AudioType is 4, the audio data of the corresponding track, indicating that it is an metadata. [0250]  Also, the right side of FIG. 43, indicates the definition of Contentkind of. contentkind represents the contents of the corresponding voice. In the example of FIG. 43, for example, if contentkind is 3, the corresponding voice is music (music). [0251]  As shown in the lower left of FIG. 43, Priority are those defined in 23008-3, represents the processing priority of the corresponding Object. The Priority, is described a value representing a processing priority of the Object only if not changed in the middle of an audio stream, 0 is described in vary. [0252]  (Overview of Information Processing System)  FIG. 44 is a diagram for explaining an outline of an information processing system in a first embodiment according to the present disclosure. [0253]  In the configuration shown in FIG. 44 are denoted by the same reference numerals are the same as the structural elements of FIG. Redundant description will be omitted as appropriate. [0254]  The information processing system 140 in FIG. 44, Web server 142 and the video playback device 144 connected to the file generation unit 141 is configured by being connected through the Internet 13. [0255]  In the information processing system 140, in a manner equivalent to MPEG-DASH, Web server 142 delivers the video stream of video content to video playback device 144 in units of tiles (tiled streaming) to. Further, in the information processing system 140, Web server 142, Object audio corresponding to reproduction object tile, it delivers an audio stream of Channel audio or HOA audio, the video playback device 144. [0256]  File generation device 141 of the information processing system 140, the audio file generating unit 56 generates the audio file in the first embodiment, except for the MPD generator 57 is such that it generates an MPD file in the first embodiment is the same as the file generating unit 11 in FIG. 11. [0257]  Specifically, the file generation unit 141 acquires the image data of the moving image content, generates a video stream encoded in tile. File generating unit 141, a video stream of each tile is filed for each segment. File generating apparatus 141 uploads the image file for each tile resulting in the Web server 142. [0258]  Further, the file generation unit 141 obtains the 3D audio video content, generates an audio stream by encoding each type of 3D audio (Channel audio / Object audio / HOA audio / metadata). File generating unit 141, the audio stream, assigning a track for each type of 3D audio. File generating unit 141 generates the audio file segment structure of FIG. 39 or FIG. 40 the audio stream of each track are arranged in the sub-segment unit, is uploaded to the Web server 142. [0259]  File generating unit 141 generates MPD file including frame size information, tile position information, the object position information and the like. File generating apparatus 141 uploads the MPD file to the Web server 142. [0260]  Web server 142 stores uploaded image file from the file generation unit 141, audio files, and the MPD file. [0261]  In the example of FIG. 44, the Web server 142, the tile # and segment group consisting of first image file of a plurality of segments, the tile # and a segment group consisting of second image file of the plurality of segments are stored. Also, the Web server 142, segment group is stored consisting of audio files 3D audio. [0262]  Web server 142, in response to a request from the video reproduction terminal 144, and transmits the image file storing, audio files, the MPD file like the video playback device 144. [0263]  Video playback terminal 144, control software 161, video playback software 162 executes such access software 163. [0264]  Control software 161 is software for controlling the data to be streamed from the Web server 142. Specifically, the control software 161, to acquire the MPD file from the Web server 142 to the video playback device 144. [0265]  Further, the control software 161 includes a display area to be instructed from the video playback software 162, based on the tile position information included in the MPD file, it identifies a tile of the display area. Then, the control software 161 commands the transmission request for image file of the tile access software 163. [0266]  制御用ソフトウエア161は、Object audioを再生対象とする場合、音声ファイル内の画枠サイズ情報の送信要求をアクセス用ソフトウエア163に指令する。また、制御用ソフトウエア161は、metadataのオーディオストリームの送信要求をアクセス用ソフトウエア163に指令する。制御用ソフトウエア161は、その指令に応じてWebサーバ142から送信されてくる画枠サイズ情報およびmetadataのオーディオストリームに含まれるオブジェクト位置情報、並びに、表示領域に基づいて、表示領域内の画像に対応するオブジェクトを特定する。そして、制御用ソフトウエア161は、そのオブジェクトのオーディオストリームの送信要求をアクセス用ソフトウエア163に指令する。 [0267]  また、制御用ソフトウエア161は、Channel audioまたはHOA audioを再生対象とする場合、Channel audioまたはHOA audioのオーディオストリームの送信要求をアクセス用ソフトウエア163に指令する。 [0268]  動画再生ソフトウエア162は、Webサーバ142から取得された画像ファイルと音声ファイルを再生するソフトウエアである。具体的には、動画再生ソフトウエア162は、ユーザにより表示領域が指定されると、その表示領域を制御用ソフトウエア161に指示する。また、動画再生ソフトウエア162は、その指示に応じてWebサーバ142から取得された画像ファイルと音声ファイルを復号する。動画再生ソフトウエア162は、復号の結果得られるタイル単位の画像データを合成して出力する。また、動画再生ソフトウエア162は、復号の結果得られるObject audio、Channel audio、またはHOA audioを必要に応じて合成し、出力する。 [0269]  アクセス用ソフトウエア163は、HTTPを用いたインターネット13を介したWebサーバ142との通信を制御するソフトウエアである。具体的には、アクセス用ソフトウエア163は、制御用ソフトウエア161の指令に応じて、画像ファイル、音声ファイル内の画枠サイズ情報や所定のオーディオストリームの送信要求を、動画再生端末144に送信させる。また、アクセス用ソフトウエア163は、その送信要求に応じて、Webサーバ142から送信されてくる画像ファイル、音声ファイル内の画枠サイズ情報や所定のオーディオストリームを動画再生端末144に受信させる。 [0270]  (ファイル生成装置の構成例)  図45は、図44のファイル生成装置141の構成例を示すブロック図である。 [0271]  図45に示す構成のうち、図11の構成と同じ構成には同じ符号を付してある。重複する説明については適宜省略する。 [0272]  図45のファイル生成装置141の構成は、音声符号化処理部55、音声ファイル生成部56、MPD生成部57、サーバアップロード処理部58の代わりに、音声符号化処理部171、音声ファイル生成部172、MPD生成部173、サーバアップロード処理部174が設けられる点が、図11のファイル生成装置11の構成と異なる。 [0273]  具体的には、ファイル生成装置141の音声符号化処理部171は、外部から入力される動画コンテンツの3Dオーディオを種類(Channel audio/Object audio/HOA audio/metadata)ごとに符号化し、オーディオストリームを生成する。音声符号化処理部171は、3Dオーディオの種類ごとのオーディオストリームを音声ファイル生成部172に供給する。 [0274]  音声ファイル生成部172は、音声符号化処理部171から供給されるオーディオストリームに対して、3Dオーディオの種類ごとにトラックを割り当てる。音声ファイル生成部172は、各トラックのオーディオストリームがサブセグメント単位で配置される図39または図40のセグメント構造の音声ファイルを生成する。このとき、音声ファイル生成部172は、外部から入力される画枠サイズ情報をサンプルエントリに格納する。音声ファイル生成部172は、生成された音声ファイルをMPD生成部173に供給する。 [0275]  MPD生成部173は、画像ファイル生成部53から供給される各タイルの画像ファイルを格納するWebサーバ142のURL等を決定する。また、MPD生成部173は、音声ファイル生成部172から供給される音声ファイルを格納するWebサーバ142のURL等を決定する。 [0276]  MPD生成部173は、画像情報生成部54から供給される画像情報をMPDファイルの画像用の「AdaptationSet」に配置する。また、MPD生成部173は、各タイルの画像ファイルのURL等を、そのタイルの画像ファイル用の「Representation」の「Segment」に配置する。 [0277]  MPD生成部173は、音声ファイルのURL等を、その音声ファイル用の「Representation」の「Segment」に配置する。また、MPD生成部173は、外部から入力される各オブジェクトのオブジェクト位置情報等を、そのオブジェクトのObject metadata track用の「Sub Representation」に配置する。MPD生成部173は、以上のようにして各種の情報が配置されたMPDファイル、並びに、画像ファイルおよび音声ファイルをサーバアップロード処理部174に供給する。 [0278]  サーバアップロード処理部174は、MPD生成部173から供給される各タイルの画像ファイル、音声ファイル、およびMPDファイルを、Webサーバ142にアップロードする。 [0279]  (ファイル生成装置の処理の説明)  図46は、図45のファイル生成装置141のファイル生成処理を説明するフローチャートである。 [0280]  図46のステップS191乃至S195の処理は、図12のステップS11乃至S15の処理と同様であるので、説明は省略する。 [0281]  ステップS196において、音声符号化処理部171は、外部から入力される動画コンテンツの3Dオーディオを種類(Channel audio/Object audio/HOA audio/metadata)ごとに符号化し、オーディオストリームを生成する。音声符号化処理部171は、3Dオーディオの種類ごとのオーディオストリームを音声ファイル生成部172に供給する。 [0282]  ステップS197において、音声ファイル生成部172は、音声符号化処理部171から供給されるオーディオストリームに対して、3Dオーディオの種類ごとにトラックを割り当てる。 [0283]  ステップS198において、音声ファイル生成部172は、各トラックのオーディオストリームがサブセグメント単位で配置される図39または図40のセグメント構造の音声ファイルを生成する。このとき、音声ファイル生成部172は、外部から入力される画枠サイズ情報をサンプルエントリに格納する。音声ファイル生成部172は、生成された音声ファイルをMPD生成部173に供給する。 [0284]  ステップS199において、MPD生成部173は、画像情報生成部54から供給される画像情報、各ファイルのURL、オブジェクト位置情報等を含むMPDファイルを生成する。MPD生成部173は、画像ファイル、音声ファイル、およびMPDファイルをサーバアップロード処理部174に供給する。 [0285]  ステップS200において、サーバアップロード処理部174は、MPD生成部173から供給される画像ファイル、音声ファイル、およびMPDファイルを、Webサーバ142にアップロードする。そして、処理は終了する。 [0286]  (動画再生端末の機能的構成例)  図47は、図44の動画再生端末144が制御用ソフトウエア161、動画再生ソフトウエア162、およびアクセス用ソフトウエア163を実行することにより実現されるストリーミング再生部の構成例を示すブロック図である。 [0287]  図47に示す構成のうち、図13の構成と同じ構成には同じ符号を付してある。重複する説明については適宜省略する。 [0288]  図47のストリーミング再生部190の構成は、MPD処理部92、音声選択部94、音声ファイル取得部95、音声復号処理部96、音声合成処理部97の代わりに、MPD処理部191、音声選択部193、音声ファイル取得部192、音声復号処理部194、音声合成処理部195が設けられる点、および、メタファイル取得部93が設けられない点が、図13のストリーミング再生部90の構成と異なる。 [0289]  ストリーミング再生部190は、選択されたオブジェクトの再生対象の音声データを取得する方法等が異なる点等を除いて、図13のストリーミング再生部90と同様である。 [0290]  具体的には、ストリーミング再生部190のMPD処理部191は、MPD取得部91から供給されるMPDファイルから音声ファイル用の「Segment」に記述される再生対象のセグメントの音声ファイルのURL等の情報を抽出し、音声ファイル取得部192に供給する。 [0291]  MPD処理部191は、MPDファイルから画像用の「AdaptationSet」に記述されるタイル位置情報を抽出し、画像選択部98に供給する。MPD処理部191は、MPDファイルから、画像選択部98から要求されるタイルの画像ファイル用の「Segment」に記述されるURL等の情報を抽出し、画像選択部98に供給する。 [0292]  音声ファイル取得部192は、Object audioを再生対象とする場合、MPD処理部191から供給されるURL等の情報に基づいて、そのURLで特定される音声ファイル内のBase trackのInitial SegmentをWebサーバ142に要求し、取得する。 [0293]  また、音声ファイル取得部192は、音声ファイルのURL等の情報に基づいて、そのURLで特定される音声ファイル内のobject metadata trackのオーディオストリームをWebサーバ142に要求し、取得する。音声ファイル取得部192は、object metadata trackのオーディオストリームに含まれるオブジェクト位置情報、Base trackのInitial Segmentに含まれる画枠サイズ情報、および音声ファイルのURL等の情報を音声選択部193に供給する。 [0294]  また、Channel audioを再生対象とする場合、音声ファイル取得部192は、音声ファイルのURL等の情報に基づいて、そのURLで特定される音声ファイル内のChannel audio trackのオーディオストリームをWebサーバ142に要求し、取得する。音声ファイル取得部192は、取得されたChannel audio trackのオーディオストリームを音声復号処理部194に供給する。 [0295]  HOA audioを再生対象とする場合、音声ファイル取得部192は、Channel audioを再生対象とする場合と同様の処理を行う。その結果、HOA audio trackのオーディオストリームが音声復号処理部194に供給される。 [0296]  なお、Object audio,Channel audio、およびHOA audioのいずれを再生対象とするかは、例えば、ユーザからの指令により決定される。 [0297]  音声選択部193は、音声ファイル取得部192から供給される画枠サイズ情報とオブジェクト位置情報とに基づいて、各オブジェクトの画像上の位置を算出する。音声選択部193は、各オブジェクトの画像上の位置に基づいて、ユーザにより指定される表示領域内のオブジェクトを選択する。音声選択部193は、音声ファイル取得部192から供給される音声ファイルのURL等の情報に基づいて、そのURLで特定される音声ファイル内の、選択されたオブジェクトのObject audio trackのオーディオストリームをWebサーバ142に要求し、取得する。音声選択部193は、取得されたObject audio trackのオーディオストリームを音声復号処理部194に供給する。 [0298]  音声復号処理部194は、音声ファイル取得部192から供給されるChannel audio trackまたはHOA audio trackのオーディオストリーム、もしくは、音声選択部193からObject audio trackのオーディオストリームを復号する。音声復号処理部194は、復号の結果得られるChannel audio,HOA audio、またはObject audioを音声合成処理部195に供給する。 [0299]  音声合成処理部195は、音声復号処理部194から供給されるObject audio,Channel audio、またはHOA audioを必要に応じて合成し、出力する。 [0300]  (動画再生端末の処理の説明)  図48は、図47のストリーミング再生部190のチャンネルオーディオ再生処理を説明するフローチャートである。このチャンネルオーディオ再生処理は、例えば、ユーザによりChannel audioが再生対象として選択される場合に行われる。 [0301]  図48のステップS221において、MPD処理部191は、MPD取得部91から供給されるMPDファイルを解析し、「SubRepresentation」に記述されるessential property とcodecに基づいて、再生対象のセグメントのChannel audioの「SubRepresentation」を特定する。また、MPD処理部191は、MPDファイルから、再生対象のセグメントの音声ファイル用の「Segment」に記述されるURL等の情報を抽出し、音声ファイル取得部192に供給する。 [0302]  In step S222, MPD processor 191, based on the dependencyLevel the identified at step S221 "SubRepresentation" to identify the level of a reference track Base track, and supplies to the audio file obtaining unit 192. [0303]  In step S223, the audio file acquisition section 192, based on the URL information or the like supplied from the MPD processing unit 191 requests the Initial Segment to be reproduced segment to the Web server 142, and acquires. [0304]  In step S224, the audio file acquisition section 192, the Level assignment box in Initial Segment, acquires the track ID corresponding to the level of the Base track is a reference track and channel audio track. [0305]  In step S225, the audio file obtaining unit 192 on the basis of the track ID of the Base track is a reference track and channel audio track, the Initial Segment, to obtain the sample entries in the trak box corresponding to the track ID. Audio file acquisition section 192 supplies the codec information included in the sample entries acquired in the speech decoding unit 194. [0306]  In step S226, the audio file acquisition section 192, based on the URL information or the like supplied from the MPD processing unit 191, requests the Web server 142, the sidx box and ssix box from the head of the audio file to be played back segment get. [0307]  In step S227, the audio file acquisition section 192, the sidx box and ssix boxes obtained in step S223, acquires the position information of the reference track and channel audio track of subsegments reproduced. Here, since the a reference track Base track contains no audio stream, the position information of the reference track are not. [0308]  In step S228, the audio file acquisition section 192, channel audio position information of track and on the basis of information such as the URL of the audio file to be played back segment, channel audio track Web server 142 an audio stream that is arranged in mdat box the request to, to get. Audio file obtaining unit 192 supplies the audio stream of the acquired channel audio track to the audio decoding processor 194. [0309]  In step S229, the speech decoding unit 194, based on the codec information supplied from the audio file obtaining unit 192, performs decoding on the audio stream of the channel audio track. Audio file obtaining unit 192 supplies the resulting channel audio to the speech synthesis processing unit 195. [0310]  In step S230, the speech synthesis processing unit 195 outputs channel audio, the process ends. [0311]  Although not shown, HOA audio reproduction process of reproducing the HOA audio by streaming playback unit 190 is carried out in the same manner as channel audio reproduction process of FIG. 48. [0312]  Figure 49 is a flow chart for explaining the object identification processing of streaming playback unit 190 of FIG. 47. The object identification process may be, for example, Object audio has been selected for playback by the user is performed when the playback area is changed. [0313]  In step S251 of FIG. 49, the audio selecting unit 193 by a user's operation, acquires the display region specified by the user. [0314]  In step S252, MPD processor 191 analyzes the MPD file supplied from the MPD acquisition unit 91, based on essential property and codec described in "SubRepresentation", the "SubRepresentation" of metadata to be played segment Identify. Also, MPD processor 191 extracts information such as the URL of the audio file to be played back segment described in the "Segment" for audio from MPD file, and supplies to the audio file obtaining unit 192. [0315]  In step S253, MPD processor 191, based on the dependencyLevel the identified at step S252 "SubRepresentation" to identify the level of a reference track Base track, and supplies to the audio file obtaining unit 192. [0316]  In step S254, the audio file acquisition section 192, based on the URL information or the like supplied from the MPD processing unit 191 requests the Initial Segment to be reproduced segment to the Web server 142, and acquires. [0317]  In step S255, the audio file acquisition section 192, the Level assignment box in Initial Segment, acquires the track ID corresponding to the level of the Base track is a reference track and object metadata track. [0318]  In step S256, the audio file obtaining unit 192 on the basis of the track ID of the Base track is a reference track and object metadata track, the Initial Segment, to obtain the sample entries in the trak box corresponding to the track ID. Audio file acquisition section 192 supplies the frame size information included in the sample entries of a reference track Base track to the audio selection unit 193. The audio file obtaining unit 192 supplies the Initial Segment to the audio selection unit 193. [0319]  In step S257, the audio file acquisition section 192, based on the URL information or the like supplied from the MPD processing unit 191, requests the Web server 142, the sidx box and ssix box from the head of the audio file to be played back segment get. [0320]  In step S258, the audio file acquisition section 192, the sidx box and ssix boxes obtained in step S257, acquires the position information of the reference track and object metadata track of subsegments reproduced. Here, since the a reference track Base track contains no audio stream, the position information of the reference track are not. Audio file acquisition section 192 supplies the sidx box and ssix box to the audio selection unit 193. [0321]  In step S259, the audio file acquisition section 192, object metadata position information of track and on the basis of information such as the URL of the audio file to be played back segment, object metadata track Web server 142 an audio stream that is arranged in mdat box the request to, to get. [0322]  In step S260, the audio file acquisition section 192, on the basis of the codec information included in the sample entry acquired in step S256, decodes the audio stream of the acquired object metadata track in step S259. Audio file obtaining unit 192 supplies the object location information included in the metadata obtained as a result of the decoding to the audio selection unit 193. The audio file obtaining unit 192 supplies information such as the URL of the audio file supplied from MPD processor 191 to the audio selection unit 193. [0323]  In step S261, the audio selection unit 193, frame size information and the object position information supplied from the audio file obtaining unit 192, and, based on the display area specified by the user, selects an object in the display area. Then, the process is terminated. [0324]  Figure 50 is a flowchart for explaining a specific object audio reproduction process performed by a streaming reproduction unit 190 after the object identification process of FIG. 49. [0325]  In step S281 of FIG. 50, MPD processor 191 analyzes the MPD file supplied from the MPD acquisition unit 91, based on essential property and codec described in "SubRepresentation", the object audio of the selected object to identify the "SubRepresentation". [0326]  In step S282, MPD processor 191, based on the dependencyLevel the identified at step S281 "SubRepresentation" to identify the level of a reference track Base track, and supplies to the audio file obtaining unit 192. [0327]  In step S283, the audio file acquisition section 192, the Level assignment box in Initial Segment, acquires the track ID corresponding to the level of the Base track is a reference track and object audio track, and supplies to the audio selection unit 193. [0328]  In step S284, the audio selection unit 193, based on the track ID of the Base track is a reference track and object audio track, the Initial Segment, to obtain the sample entries in the trak box corresponding to the track ID. The Initial Segment are those supplied from the audio file obtaining unit 192 in step S256 of FIG. 49. Audio selection unit 193 supplies the codec information included in the sample entries acquired in the speech decoding unit 194. [0329]  In step S285, the audio selection unit 193, from sidx box and ssix box supplied from the audio file obtaining unit 192 in step S258, the position information of the object audio track and the reference track of the object selected sub-segments of the reproduced get. Here, since the a reference track Base track contains no audio stream, the position information of the reference track are not. [0330]  In step S286, the audio selection unit 193, based on the positional information and the information such as the URL of the audio file to be played back segment of the object audio track, it is arranged in mdat box, audio object audio track of the selected object It requests a stream to the Web server 142, and acquires. Audio selection unit 193 supplies the audio stream of the acquired object audio track to the audio decoding processor 194. [0331]  In step S287, the speech decoding unit 194, based on the codec information supplied from the audio selection unit 193, decodes the audio stream of object audio track. Audio selection unit 193 supplies the object audio obtained as a result of decoding the speech synthesis processing unit 195. [0332]  In step S288, the speech synthesis processing unit 195, and outputs the synthesized object audio supplied from the audio decoding unit 194. Then, the process is terminated. [0333]  As described above, in the information processing system 140, the file generation unit 141 generates an audio file 3D audio are arranged is divided into a plurality of tracks according to the type of 3D audio. The video playback device 144 acquires the audio stream of the track of a given type of 3D audio out of the audio file. Therefore, video reproduction terminal 144 can efficiently retrieve an audio stream of a predetermined type of 3D audio. Thus, the file generation unit 141 can be said to generate a sound file to improve the efficiency of obtaining the predetermined kind of 3D audio of the audio stream. [0334]    (Summary of the track)  FIG. 51 is a diagram for describing the outline of the track in the second embodiment according to the present disclosure. [0335]  As shown in FIG. 51, in the second embodiment, as a sample of the Base track, that the base sample is recorded, different from the first embodiment. Base sample composed of Channel audio / Object audio / HOA audio / reference information sample f of metadata. Samples of Channel audio / Object audio / HOA audio / metadata referenced by reference information included in the base sample, by arranging the order of arrangement of the reference information, and generates an audio stream before the 3D audio are divided into tracks be able to. [0336]  (Examples of the syntax of sample entries in the base track)  FIG. 52 is a diagram showing an example of a syntax of sample entries in the base track of Figure 51. [0337]  Syntax of FIG. 52, sample entry, instead of "mha1" indicating that the sample entries of the Base track of Figure 33, "mha2" is described which indicates that a sample entry Base track of Figure 51 except for the point, it is identical to the syntax of FIG. 34. [0338]  (Structure Example of the base sample)  FIG. 53 is a diagram showing a structural example of a base sample. [0339]  As shown in FIG. 53, the base sample is configured to Channel audio / Object audio / HOA audio / metadata of extractor sample units as a sub-sample. Channel audio / Object audio / HOA audio / metadata of extractor is the type of extractor, and are constituted by offset and size of the corresponding Channel audio track / Object audio track (s) / HOA audio track / Object metadata track subsample that. This offset is a position in the file of the sub-samples of the base sample, which is a difference between the position in the file of Channel audio track / Object audio track (s) / HOA audio track / Object metadata track of the sample. That, offset is information indicating the position in the file of a sample of another track corresponding to the sub-samples of the base sample containing it. [0340]  Figure 54 is a diagram showing an example of a base sample syntax. [0341]  As shown in FIG. 54, the base sample, SCE element object audio is stored in a sample of Object audio track is, are replaced on the EXT element to store the extractor. [0342]  Figure 55 is a diagram showing an example of a data extractor. [0343]  As shown in FIG. 55, the extractor, the type of extractor, as well as the corresponding Channel audio track / Object audio track (s) / HOA audio track / Object metadata offset and size of the sub-sample of track is written. [0344]  Incidentally, are defined in AVC (Advanced Video Coding) / HEVC (High Efficiency Video Coding), by utilizing the structure of NAL (Network Abstraction Layer), it may be extended to store audio elementary and config information . [0345]  Processing by the information processing system and an information processing system in the second embodiment is the same as the first embodiment, description thereof will be omitted. [0346]    (Summary of the track)  FIG. 56 is a diagram for describing the outline of the track in the third embodiment according to the present disclosure. [0347]  As shown in FIG. 56, in the third embodiment, as a sample of the Base track, based samples and metadata sample is recorded, that the Object metadata track is not provided, it is different from the first embodiment. [0348]  Processing by the information processing system and an information processing system in the third embodiment, in order to obtain the object position information, except that the audio stream of the Base track is acquired instead of Object metadata track, the first embodiment are the same as the form, explanation is omitted. The scope of the claims [Claim 1]  Acquisition unit for acquiring speech data of a predetermined said track of the files in which a plurality of types of audio data are arranged is divided into a plurality of tracks according to the type  information processing apparatus including a. [Claim 2]  The types, elements of the audio data, the type of the element or the element is a group that is classified,  configured as  an information processing apparatus according to claim 1. [Claim 3]  Decoding unit for decoding the audio data of the predetermined track acquired by the acquisition unit  further comprises a  processing apparatus according to claim 1. [Claim 4]  Said decoding unit, when the number of the predetermined track is plural, synthesizes the audio data of the predetermined track acquired by the acquisition unit, decodes the audio data after synthesis  according to claim 3 the information processing apparatus. [Claim 5]  The file, along with the audio data of a plurality of objects units are arranged is divided into the track that is different for each of the object, the metadata of the audio data of all of the object units are summarized in different tracks and the track are arranged,  the acquisition unit, the audio data of the track to be reproduced object, acquires as the audio data of the predetermined track, acquires the metadata,  the decoding unit by the acquisition unit from the acquired metadata, said extract metadata of playback target object, combines the meta data and the audio data acquired by the acquisition unit  configured as  described in claim 4 the information processing apparatus. [Claim 6]  The file, along with the audio data of a plurality of objects units are arranged is divided into the track that is different for each of the object, the metadata of the audio data of all of the object units are summarized in different tracks and the track are arranged,  the acquisition unit, the audio data of the track to be reproduced object, acquires as the audio data of the predetermined track, acquires the metadata,  the decoding unit is not a reproduction object zero data the decoding result as the audio data of the track is zero, to synthesize said audio data and the metadata acquired by the acquiring unit  is configured as  an information processing apparatus according to claim 4 . [Claim 7]  Wherein the file, along with the audio data of a plurality of objects units are arranged is divided into the track that is different for each of the object, the metadata of the audio data of the object units are arranged in different track for each of the objects,  the acquisition unit, the audio data of the track to be reproduced object, acquires as the audio data of the predetermined track, to retrieve metadata of the object to be reproduced,  the decoding unit, the acquisition unit the and the acquired speech data and the metadata is synthesized by  configured as  an information processing apparatus according to claim 4. [8.]  Wherein the audio data of a plurality of tracks are arranged in one of said file  configured as  an information processing apparatus according to claim 1. [Claim 9]  Wherein the audio data of a plurality of tracks are arranged in said different file for each of the tracks  are configured as  an information processing apparatus according to claim 1. [Claim 10]  Wherein the file information relating to the audio data of the plurality of types are arranged as different track from said plurality of tracks  are configured as  an information processing apparatus according to claim 1. [Claim 11]  It said information on the audio data of the plurality of types includes a frame size information that indicates the frame size of the image data corresponding to the audio data  have been configured as  an information processing apparatus according to claim 10. [Claim 12]  Wherein the file, as the audio data of different tracks and said plurality of tracks, information is arranged to indicate the position of the speech data of the other of the track corresponding to the voice data  configured to  claim 1 the information processing apparatus according to. [Claim 13]  The file as the data of different tracks and said plurality of tracks, and information indicating the position of the speech data of the other of the track corresponding to the data, the metadata of the audio data of the other of the track is disposed  is configured as  an information processing apparatus according to claim 1. [Claim 14]  Wherein the metadata of the audio data includes information indicating the acquisition position of the voice data  configured as  an information processing apparatus according to claim 13. [Claim 15]  The file includes information that references the other of the track of the track  is constructed as  an information processing apparatus according to claim 1. [Claim 16]  The file includes a codec information of the audio data of each track  is constructed as  an information processing apparatus according to claim 1. [Claim 17]  Given the type of audio data is information indicating the acquisition position of the other of said types of voice data  configured as  an information processing apparatus according to claim 1. [Claim 18]  The information processing apparatus,  the predetermined acquisition step of acquiring speech data of the track of the files in which a plurality of types of audio data are arranged is divided into a plurality of tracks in accordance with the type  information processing method comprising. [Claim 19]  Generator in which a plurality of types of audio data to produce a file that is located is divided into a plurality of tracks according to the type  information processing apparatus including a. [Claim 20]  The information processing apparatus,  generating step of generating a file in which a plurality of types of audio data are arranged is divided into a plurality of tracks in accordance with the type  information processing method comprising

Documents

Application Documents

# Name Date
1 Priority Document [23-11-2016(online)].pdf 2016-11-23
2 Power of Attorney [23-11-2016(online)].pdf 2016-11-23
3 Form 1 [23-11-2016(online)].pdf 2016-11-23
4 201617040061-STATEMENT OF UNDERTAKING (FORM 3) [23-11-2016(online)].pdf 2016-11-23
5 201617040061-PRIORITY DOCUMENTS [23-11-2016(online)].pdf 2016-11-23
6 201617040061-POWER OF AUTHORITY [23-11-2016(online)].pdf 2016-11-23
7 201617040061-FORM 1 [23-11-2016(online)].pdf 2016-11-23
8 201617040061-DRAWINGS [23-11-2016(online)].pdf 2016-11-23
9 201617040061-DECLARATION OF INVENTORSHIP (FORM 5) [23-11-2016(online)].pdf 2016-11-23
10 201617040061-COMPLETE SPECIFICATION [23-11-2016(online)].pdf 2016-11-23
11 201617040061.pdf 2016-11-24
12 Other Patent Document [15-12-2016(online)].pdf 2016-12-15
13 201617040061-OTHERS-191216.pdf 2016-12-20
14 201617040061-Correspondence-191216.pdf 2016-12-20
15 abstract.jpg 2017-01-16
16 Form 3 [18-04-2017(online)].pdf 2017-04-18
17 201617040061-FORM 18 [16-04-2018(online)].pdf 2018-04-16
18 201617040061-FER.pdf 2021-10-17

Search Strategy

1 SearchStrategyMatrixE_15-05-2020.pdf