Sign In to Follow Application
View All Documents & Correspondence

Signal Processing Device Method And Program

Abstract: The present technology relates to a signal processing device method and program for making it possible to more easily obtain appropriate spread information. This signal processing device is provided with a spread information generation unit for generating spread information about a frame to be processed for an audio object on the basis of object position information about the frame to be processed for the audio object and one from among spread information about another frame different from the frame to be processed for the audio object zoom information and information about the size of the audio object. The present technology can be applied to an encoding device and a decoding device.

Get Free WhatsApp Updates!
Notices, Deadlines & Correspondence

Patent Information

Application #
Filing Date
18 October 2019
Publication Number
50/2019
Publication Type
INA
Invention Field
ELECTRONICS
Status
Email
mahua.ray@remfry.com
Parent Application
Patent Number
Legal Status
Grant Date
2024-05-16
Renewal Date

Applicants

SONY CORPORATION
1-7-1, Konan, Minato-ku, Tokyo 1080075

Inventors

1. YAMAMOTO Yuki
c/o SONY CORPORATION, 1-7-1, Konan, Minato-ku, Tokyo 1080075
2. CHINEN Toru
c/o SONY CORPORATION, 1-7-1, Konan, Minato-ku, Tokyo 1080075
3. TSUJI Minoru
c/o SONY CORPORATION, 1-7-1, Konan, Minato-ku, Tokyo 1080075

Specification

0001]This technique, the signal processing apparatus and method, and a program, the signal processing apparatus and method capable of obtaining a more easily appropriate spread information, and a program.
BACKGROUND
[0002]Conventionally, as a coding technique that can handle object audio, such as an international standard MPEG (Moving Picture Experts Group) -H Part 3: such 3D audio standards are known (e.g., see Non-Patent Document 1).
[0003]
 In such a coding system, it is possible to perform spread processing to expand the sound image of the sound of the audio objects.
[0004]
 Specifically, the encoding apparatus side, together with the object position information indicating a position of the audio object in the bit stream, spread information indicating the spreading degree of a sound image of the sound of the audio object is included as metadata for the audio objects.
[0005]
 For example, when the sound image of the sound of audio objects is spread symmetrically in the horizontal and vertical directions, spread information, any angle in the range of from 0 ° indicating the spread degree of the sound image in the horizontal direction and the vertical direction to 180 degrees ( hereinafter, consisting of a single value indicating the spread angle also referred to).
[0006]
 In contrast, if the sound image of the sound of the audio object is widened asymmetrically in the horizontal and vertical directions, spread information consists of two values ​​indicating the spread angle of the horizontal spread angle and a vertical direction. Incidentally, spread angle and vertical spread angle in the horizontal direction is the angle of both the range of from 0 degrees to 180 degrees.
[0007]
 Such spread information may be a different value for each frame of the audio signal of the audio objects.
[0008]
 Further, in the decoding apparatus, on the basis of the spread information extracted from the bit stream, rendering processing to expand the sound image of the audio object, i.e. the spread process is performed. By performing such a spread process, such as greater, widen the area of ​​the sound image of an object near small, it is possible to narrow the area of ​​the sound image of an object in the distance. As a result, it is possible to realize a content reproduction high realistic sensation.
CITATION
Non-patent literature
[0009]
非特許文献1 : INTERNATIONAL STANDARD ISO/IEC 23008-3 First edition 2015-10-15 Information technology - High efficiency coding and media delivery in heterogeneous environments - Part 3: 3D audio
Summary of the Invention
Problems that the Invention is to Solve
[0010]
 However, it is not easy to obtain a proper spread information.
[0011]
 For example spread information, high cost to impart manually in every time (frame) for each and audio objects. In some cases, the first place may not have spread information is given.
[0012]
 Furthermore, for example, the size of an object changes in the zoom process is performed in the screen by the user, such as the decoding side, widening the spread information obtained from the bit stream region of the sound image be used as it is in the appropriate size it can not be. Then, the size of the object in the screen, not the size corresponding region of a sound image of the audio object corresponding to the object, realism is impaired.
[0013]
 This technology has been made in view of such circumstances, it is desirable to make it possible to obtain a more easily correct spread information.
Means for Solving the Problems
[0014]
 Signal processing apparatus according to an embodiment of the present technology, the spread information different from the frame and the processing target frame of object position information and the audio object frame to be processed in the audio object, and a zoom information, the audio object based on one of the size information includes a spread information generation unit for generating spread information of the processing target frame of the audio object.
[0015]
 Wherein the spread information generation unit, the other of said spread information and the object position information and the frame of the processing target based on the processing object frame and the other frames of the audio objects in the frame of the audio objects thereby generating the spread information.
[0016]
 The object position information, may be the distance from the reference position to the audio object.
[0017]
 Wherein the spread information generation unit, and the ratio of the distance in the distance and the other frames in the frame of the processing target, the other of said frame of said processed the arctangent of the product of the tangent of the spread information of the frame it can be generated as the spread information.
[0018]
 Wherein the spread information generation unit, said zoom information, based on said spread information of the processing target frame of the audio object, thereby generating the spread information of the processing target frame of the audio object after zooming be able to.
[0019]
 Wherein the spread information generating unit based on the relationship between the position and size of the virtual screen before and after the zoom determined by the zoom information, thereby generating the spread information of the processing target frame of the audio object after the zoom can.
[0020]
 Wherein the spread information generation unit, and the size information, wherein the processing target frame of the audio object based on the object position information, the spread information of the frame of the processing target can be produced.
[0021]
 The size information, may be the distance to the outer shell from the center of the audio object.
[0022]
 The object position information, may be the distance from the reference position to the audio object.
[0023]
 Wherein the spread information generation unit, can be generated and the size information, the inverse tangent of the ratio between the object position information of the frame of the processing target as the spread information of the frame of the processing target.
[0024]
 Signal processing method or a program according to an embodiment of the present technology, the spread information different from the frame and the processing target frame of object position information and the audio object frame to be processed in the audio object, and a zoom information, the based on one of the size information of the audio object, comprising the step of generating a spread information of the processing target frame of the audio object.
[0025]
 In one aspect of the present technology, the spread information different from the frame and the processing target frame object position information and the audio object frame to be processed in the audio object, and zoom information, the audio object size based on one of the information, spread information of the processing target frame of the audio objects is generated.
The invention's effect
[0026]
 According to the embodiments of the present technology, it is possible to obtain a more easily correct spread information.
[0027]
 Here, the advantages described in the present invention is not necessarily limited, it may be any of the effects described in the present disclosure.
BRIEF DESCRIPTION OF THE DRAWINGS
[0028]
FIG. 1 is a diagram illustrating the generation of the spread information.
Is a diagram illustrating a configuration example of FIG. 2 encoder.
Is a flow chart illustrating the FIG. 3] encoding process.
Is a diagram illustrating a configuration example of FIG. 4 decoding apparatus.
Is a flowchart illustrating a FIG. 5 decoding process.
6 is a diagram illustrating the generation of the spread information.
7 is a flowchart illustrating an encoding process.
Is a flowchart illustrating a FIG. 8 decoding process.
9 is a diagram for explaining zoom.
Is a diagram illustrating a modification of the object position information corresponding to FIG. 10 Zoom.
Is a diagram illustrating a configuration example of FIG. 11 decoder.
Is a flowchart illustrating the FIG. 12 decoding process.
13 is a diagram showing a configuration example of a computer.
DESCRIPTION OF THE INVENTION
[0029]
 Hereinafter, with reference to the accompanying drawings, a description will be given of an embodiment according to the present technology.
[0030]

 This technique, spread information of audio objects, object position information of the audio object, zoom information, size information, etc. spread information at different times, the spread to be obtained information by generated based on different other information and is intended to be able to obtain a more easily correct spread information.
[0031]
 In the following, an audio signal and the audio signal of the audio objects of the multi-channel will be described as being encoded according to such predetermined standards. In the following, it will be referred to as simply both object audio objects.
[0032]
 For example, an audio signal of each channel and each object is transmitted is encoded for each frame.
[0033]
 That, and the audio signal encoded, stored information needed to decode such audio signal into a plurality of elements (bit stream element), those composed of elements bitstream is transmitted to the decoding side from the encoding side .
[0034]
 More specifically, for example, 1 to frame of bitstream beginning plurality of elements are arranged in order from, is disposed identifier indicating that the end is the end position related information of the frame.
[0035]
 Then, the elements arranged on the top are the ancillary data region called DSE (Data Stream Element), etc. Information and identification information on the downmix audio signal is DSE, information is described regarding each of a plurality of channels .
[0036]
 In addition, each element following the DSE, encoded audio signals are stored. In particular, the elements that audio signal of a single channel is stored called SCE is called (Single Channel Element), the elements that audio signals of two channels in the pair is stored CPE (Coupling Channel Element) It has been. Audio signals of each object are stored in SCE.
[0037]
 In this technique, spread information of the audio signals of each object is generated and stored in the DSE. More specifically, the metadata of the object that contains the generated spread information is stored in the DSE.
[0038]
 Here, the meta data includes at least spread information and object position information.
[0039]
 Spread information is spread angle showing the angle information indicating the range of the sizes of the sound of the object, i.e. the spread degree of the sound of the sound object. In other words, spread information may also be referred to as information indicating the size of the area of ​​the object.
[0040]
 Furthermore, the object position information is information indicating the position of the object in three-dimensional space, is for example, a horizontal direction angle indicating the position of the object as viewed from the reference position (origin), the vertical angles, and a radial coordinate information that.
[0041]
 Horizontal angle, horizontal angle (azimuth angle) in the horizontal direction position of the object as viewed from the reference position is a position where the user exists, that is the direction of the object as viewed from the direction and the reference position serving as a reference in the horizontal direction is the angle formed between.
[0042]
 Here, when the horizontal angle is 0 degrees, the object is located in front of the user, when the horizontal angle is 90 degrees or -90 degrees, objects that are located just beside the user Become. Further, when the horizontal angle is 180 degrees or -180 degrees, the object will be positioned directly behind the user.
[0043]
 Similarly the vertical angle is an angle in the vertical direction indicates the vertical direction of the position of the object as viewed from the reference position (elevation angle), that is the angle between the direction of the object as viewed from a reference direction and the reference position in the vertical direction is there.
[0044]
 Moreover, the radius is the distance from the reference position in the space to the position of the object.
[0045]
 However, the high cost to impart to each frame spread information manually for each object. For example, to handle for a long time many of the objects in the movie content, it can be said that the cost is particularly high by hand.
[0046]
 Therefore, in the applied coding apparatus to which the present technology, an object indicating the spread information of a frame to be processed of an object, the spread information of the processing object frame and different other frames in time are the spatial position of the object and to generate on the basis of the position information.
[0047]
 Specifically, for example, the first frame of the audio signal of the object, the particular one frame (hereinafter, also a reference frame referred to) spread information only is applied manually. And, for frames other than the reference frame, based on the spread information of the reference frame and the object position information, the spread information is generated by the encoding device.
[0048]
 Incidentally, the one frame of the audio signal may be a reference frame constituting the content, may be a plurality of frames are used as reference frames. In particular, when providing a plurality of reference frames, to equal intervals frame may be the reference frame, any frame may be a reference frame arranged at unequal intervals. Further, originally not a reference frame frame after generating spread information of the frame, may be used as a reference frame.
[0049]
 Moreover, here it will be described an example in which the spread information in a frame unit is generated, and the number frames, spread information may be generated at a predetermined time unit.
[0050]
 Hereinafter, for simplicity of explanation, the description will be continued assuming that one frame is a reference frame.
[0051]
 For example, it is assumed that the sound image of the sound of the object is widened symmetrically in the horizontal and vertical directions, spread information frame index i indicating the frame, that is, to mark the spread angle and s (i). Further, the frame index i indicating the current frame is the processing object marked IFRm, a frame index i of the reference frame and that referred to Rfrm.
[0052]
 Further, it is assumed that referred in the frame the frame index of the object is i, the radius of the object included in the object position information r (i).
[0053]
 In such a case, for example, the size of the object, as shown in Figure 1 is utilized that does not change, the current frame IFRm spread information s (IFRm) is generated.
[0054]
 In the example shown in FIG. 1, which is the position of the user's head is the position where the position of the origin O is the reference, the object was in position of the reference frame in rfrm point OB11 is, the point OB21 the current frame ifrm It has moved to the position.
[0055]
 At this time, the distance of the length of the straight line L11 connecting the origin O and the point OB11 serving as a reference, that is from the origin O to the point OB11 becomes the radius r (rfrm) included in the object position information of the reference frame Rfrm. Similarly, the distance of the length of the straight line L21 connecting the origin O and the point OB21, i.e. from the origin O to the point OB21 becomes the radius r (IFRm) included in the object position information of the current frame IFRm.
[0056]
 Further, the spread angle indicated by the spread information of the reference frame Rfrm, i.e. spread information s (rfrm) are known. Passing the origin O, and the angle between the straight line L11 is a straight line L12 linearly is s (rfrm).
[0057]
 Spread information s (Rfrm), the size of the range of the sound of the object, that is because it is angle information indicating a range of a region of an object, the end portion (outer shell portion) of the area of ​​the object is to be located on the straight line L12 become.
[0058]
 That is, through the point OB11, and a straight line L11 perpendicular straight line a straight line L13, when the intersection of the straight line L12 and the line L13 as a point OBE11, end portions (outer shell portion) of the area of ​​the position of this point OBE11 objects the position. Accordingly, the length from the point OB11 to point OBE11 (distance) consists of the center of the object in the reference frame rfrm and length to the outer shell.
[0059]
 Such point length from OB11 to point OBE11, i.e. the length of the straight line L13, using the spread information s (rfrm) and radius r (rfrm), and r (rfrm) × tan (s (rfrm)) it can be expressed.
[0060]
 Further, for the current frame IFRm spread information s to be obtained from now (IFRm), passing the origin O, and the angle between the straight line L21 is a straight line L22 linearly is s (IFRm). Further, as the point OB21, and a straight line L23 and the straight line L21 perpendicular straight line, when the intersection of the straight line L22 and the line L23 as a point OBE21, the end part of the object area the position of the point OBE21 is in the current frame IFRm (outer the position of the shell portion).
[0061]
 Accordingly, the length from the point OB21 to point OBE21, the length of the words linear L23, becomes the length from the center of the object to the outer shell in the current frame IFRm, using spread information s (IFRm) and radius r (IFRm) it can be expressed as r (ifrm) × tan (s (ifrm)) Te.
[0062]
 Here, the size of the object, i.e. the size of the sound image area of ​​an object sound is a is constant regardless of the frame (no change), and the length from the point OB11 to point OBE11, from point OB21 to point OBE21 It is equal to the length of. That r (rfrm) × tan (s (rfrm)) = r (ifrm) × tan (s (ifrm)) is established.
[0063]
 Therefore, it is possible from the relationship obtain appropriate spread information s (IFRm) of the current frame IFRm by the following equation (1).
[0064]
[Number 1]

[0065]
 In equation (1), the ratio of the radius r (Rfrm) at radius r (IFRm) and the reference frame Rfrm in the current frame IFRm object r (rfrm) / r (ifrm), the reference frame Rfrm objects spread information s ( the product of the tangent of rfrm) tan (s (rfrm)) × r (rfrm) / r (ifrm) is required. Then, generated as a product tan (s (rfrm)) × r (rfrm) / r of the current frame IFRm of arctangent of (IFRm) object spreads information s (ifrm).
[0066]
 Thus the radius r of the current frame IFRm (IFRm), by calculating the spread information s of the reference frame rfrm (rfrm) and radius r (Rfrm) and formula based on (1) the current frame by a simple operation IFRm it is possible to generate a proper spread information s (IFRm) of.
[0067]
 In particular, it is sufficient to impart spread information manually only several frames of the audio signal of the object in this example, it is possible to significantly reduce the applied cost spread information manual.
[0068]
 Incidentally, the case has been described where the sound image of the sound of the object is widened symmetrically in the horizontal direction and the vertical direction in the above, spread by calculation similar also when the sound image of the sound of the object is widened asymmetrically in the horizontal and vertical directions it is possible to generate the information.
[0069]
 For example, the frame index is i-frame horizontal spread information, i.e. the spread angle in the horizontal direction s width marked (i), the vertical spread information frame frame index is i, i.e. the spread angle in the vertical direction the S Height to be referred to as (i).
[0070]
 In this case, each of the horizontal and vertical directions may be performed the same calculation as equation (1) described above.
[0071]
 That is, the horizontal spread information s of the current frame IFRm by calculating the following equation (2) width can be obtained (IFRm), vertical current frame IFRm by calculating the following equation (3) spread information S Height it is possible to obtain a (ifrm).
[0072]
[Number 2]

[0073]
[Number 3]

[0074]
 As described above, the object position information of the current frame, by using the object position information and the spread information of the reference frame, it is possible to more easily obtain a proper spread information of the current frame.
[0075]

 Next, a specific embodiment of the applied coding apparatus to which the present technology.
[0076]
 Figure 2 is a diagram illustrating a configuration example of a coding apparatus to which the present technology is applied.
[0077]
 Encoding device 11 shown in FIG. 2 includes a channel audio encoder 21, an object audio encoding unit 22, the metadata input unit 23, spread information generation unit 24 and packing unit 25.
[0078]
 The channel audio coding unit 21, the audio signals of each channel of the multi-channel number of channels is M is supplied. For example an audio signal of each channel is supplied from the microphone corresponding to those channels. In Figure 2, the character "# 0" to "# M-1" represents the channel number of each channel.
[0079]
 Channel audio encoder 21 encodes the audio signal of each channel supplied, and supplies to the packing unit 25 the coded data obtained by coding.
[0080]
 The object audio encoding unit 22, the audio signal of each of the N object is supplied. For example audio signals of each object is supplied from the microphone attached to those objects. In Figure 2, the character "# 0" to "# N-1" represents the object number of each object.
[0081]
 Object audio encoding unit 22 encodes the audio signals of each object supplied, and supplies to the packing unit 25 the obtained encoded data.
[0082]
 The metadata input section 23 supplies the metadata for each object to spread information generating unit 24.
[0083]
 For example, the meta-data of the reference frame of the object includes the object position information and the spread information. In contrast, the meta-data of the frame is not a reference frame of the object, have been included in the object position information is not included in the spread information.
[0084]
 Spread information generating unit 24, based on the metadata supplied from the metadata input unit 23, generates a spread information not a reference frame frame, and stores the metadata. Thus, the spread information generation unit 24, all the frames for meta containing the object position information and spreads information data is obtained.
[0085]
 Spread information generating unit 24 supplies the metadata for each object in the object position information and the spread information to the packing unit 25.
[0086]
 Packing unit 25, bit packed-channel audio encoder 21 supplied encoded data from the object audio encoding unit 22 supplied encoded data from, and the metadata supplied from the spread information generation unit 24 It generates a stream and outputs.
[0087]
 Thus a bit stream obtained by the each channel of coded data for each frame includes the object coded data, and each object in the metadata.
[0088]
 Here, one frame of M audio signal of each channel to be stored in the bit stream, and the N audio signals of each object is an audio signal of the same frame to be reproduced simultaneously.
[0089]

 Next, description will be given of a processing performed by the encoding device 11.
[0090]
 Encoder 11, an audio signal and a plurality of audio signals of each object of the plurality of respective channels are supplied by one frame min, performs encoding processing, the bit stream including encoded audio signals Output.
[0091]
 Hereinafter, with reference to the flowchart of FIG. 3 will be described encoding process by the encoding device 11.
[0092]
 In step S11, spread information generation unit 24, a frame to be processed is determined whether a reference frame.
[0093]
 For example or metadata input unit 23 receives an input operation of the user, or by communicating with the outside, by or performs reading from outside of the recording area, it acquires the metadata for each object in the processing-frame , and supplies the spread information generating unit 24.
[0094]
 Spread information generation unit 24 determines that if it contains spread information in the meta data of the frame to be processed for each object supplied from the metadata input unit 23, a frame to be processed is a reference frame.
[0095]
 Here, for simplicity of description, a frame is a reference frame for all the objects will be described as being the same, if the frame is different to be a reference frame for each object, step S11 and step the process in S12 may be performed for each object.
[0096]
 In step S11, if it is determined that the reference frame, spread information generation unit 24 supplies the meta data of a frame to be processed for each object supplied from the metadata input section 23 directly to the packing unit 25, the processing It proceeds to step S13.
[0097]
 The other hand, if it is determined not to be the reference frame in step S11, the process proceeds to step S12.
[0098]
 In step S12, the spread information generation unit 24, for each object, and generates a spread information of the frame to be processed based on the metadata supplied from the metadata input section 23.
[0099]
 For example spread information generation unit 24, the object position information included in the metadata of the processing target frame, to the the object position information and spreads information included in the meta data of the reference frame supplied from the metadata input unit 23 past based on equation (1) described above, or by calculating the above equations (2) and (3), and generates a spread information processing target frame.
[0100]
 Spread information generation unit 24, as for each object to produce a spread information, supplies the generated spread information, stored in their metadata for each object, the meta data including the object position information and the spread information to the packing unit 25 to.
[0101]
 For each object in this manner, the metadata including the object position information and the spread information is obtained, then the process proceeds to step S13.
[0102]
 Or reference is judged to be the frame, or when the spread information in step S12 is generated in step S11, processing in step S13 is performed.
[0103]
 That is, the packing unit 25 in step S13 stores the metadata for each object supplied from the spread information generation unit 24 to the DSE bitstream.
[0104]
 In step S14, channel audio encoder 21 encodes the audio signal of each channel supplied, supplies the encoded data of each channel obtained result to the packing unit 25.
[0105]
 In step S15, the packing unit 25 the coded data of the audio signals of each channel supplied from channel audio encoder 21, and stores the SCE or CPE bitstream. That is, the encoded data is stored in each element being arranged subsequent to the DSE in the bitstream.
[0106]
 In step S16, object audio encoding unit 22 encodes the audio signals of each object supplied, supplies the encoded data of each object obtained result to the packing unit 25.
[0107]
 In step S17, the packing unit 25 the coded data of the audio signals of each object supplied from the object audio encoding unit 22, and stores the SCE bitstream. That is, the encoded data is stored in a number of elements arranged later than DSE in the bitstream.
[0108]
 With the above processing, the frame to be processed, encoded data of the audio signals of all channels, and the audio signals of all the object metadata and bitstream encoded data is stored is obtained.
[0109]
 In step S18, the packing unit 25 outputs one frame of the bit stream obtained.
[0110]
 In step S19, the encoding apparatus 11 determines whether to end the process. For example, it is determined that the encoding of all the frame when finished, the process ends in step S19.
[0111]
 If it is determined not to end the process in step S19, the process returns to step S11, processing for the next frame is described above is a frame processed is repeatedly performed.
[0112]
 In contrast, if it is determined that the process ends in step S19, each unit of the encoding device 11 stops the process performed, the coding process ends.
[0113]
 Encoding device 11 as described above may be stored in the bit stream to generate a spread information of the frame is not a reference frame for each object, and outputs. Thus, it is possible to reduce the cost of the spread information giving manual, obtained more easily correct spread information.
[0114]

 In the above example has been described that generates a spread information in the encoding device 11, to generate a spread information in the decoding device it may be. In such a case, an input bit stream outputted from the encoding device, decoding device for decoding encoded data included in the bit stream is configured as shown in FIG. 4, for example.
[0115]
 Decoding device 101 shown in FIG. 4 includes unpacking / decrypting section 111, spread information generation unit 112, the rendering unit 113 and the mixing portion 114,.
[0116]
 Unpacking / decrypting unit 111 obtains the bit stream outputted from the encoding device performs unpacking and decoding the bit stream.
[0117]
 Unpacking / decrypting section 111 supplies the audio signals of each object obtained by the unpacking and decoding to the rendering unit 113, the metadata for each object obtained by the unpacking and decoding the spread information generating unit 112 supplies. Here, only the meta data of the reference frame includes the spread information, the meta data is not a reference frame frame is assumed to include the spread information.
[0118]
 Further, unpacking / decrypting section 111 supplies the audio signal of each channel obtained by the unpacking and decoding the mixing unit 114.
[0119]
 Spread information generating unit 112, based on the object metadata supplied from the unpacking / decrypting unit 111, and generates a spread information not a reference frame frame, and stores the metadata. Thus, the spread information generation unit 112, all the frames for meta containing the object position information and spreads information data is obtained.
[0120]
 Spread information generating unit 112 supplies the metadata for each object that contains the object position information and the spread information to the rendering unit 113.
[0121]
 Rendering unit 113 generates an audio signal of M channels based on each object in the metadata supplied audio signals of each object supplied from the unpacking / decrypting unit 111, and the spread information generation unit 112, a mixing unit and supplies to 114. Rendering unit 113 at this time, the sound image of each object, localized at the position indicated by the object position information of the objects, audio of the M respective channels as the sound image spreads in breadth degree indicated by the spread information of the object to generate a signal.
[0122]
 Mixing unit 114, an audio signal of each channel supplied from the unpacking / decrypting unit 111, an audio signal of each channel supplied from the rendering unit 113 and weighted addition for each channel, the final audio of each channel to generate a signal. Mixing unit 114, a final audio signal of each channel obtained in this manner is supplied to a speaker corresponding to each channel of the external, to reproduce the sound.
[0123]

 Next, the operation of the decoding device 101.
[0124]
 Decoding device 101, the bit stream is supplied from the encoding device to generate an audio signal by performing decoding processing, and outputs it to the speaker. Hereinafter, with reference to the flowchart of FIG. 5 will be described decoding processing performed by the decoding device 101.
[0125]
 In step S51, the unpacking / decrypting unit 111 obtains the bit stream transmitted from the encoding device by one frame. That is, the bit stream of a frame to be processed is received.
[0126]
 In step S52, unpacking / decrypting unit 111 performs frame decoding process.
[0127]
 That is, the unpacking / decrypting unit 111, a bit stream acquired processing target frame unpacking at step S51, acquires the encoded data of each channel, each object coded data, and metadata for each object .
[0128]
 The unpacking / decrypting unit 111 decodes the encoded data of each channel, and supplies the audio signal of each channel obtained result to the mixing unit 114. Further, unpacking / decrypting unit 111 decodes the encoded data of each object, and supplies the audio signals of each object obtained as a result to the rendering unit 113.
[0129]
 Further unpacking / decrypting section 111 supplies the metadata for each object acquired in spread information generation unit 112.
[0130]
 In step S53, spread information generation unit 112, a frame to be processed is determined whether a reference frame.
[0131]
 For example spread information generation unit 112 determines, if it contains spread information in the meta data of the frame to be processed for each object supplied from the unpacking / decrypting unit 111, a frame to be processed is a reference frame .
[0132]
 In step S53, if it is determined that the reference frame, spread information generating unit 112 supplies the metadata of the frame to be processed for each object supplied from the unpacking / decrypting section 111 to the rendering unit 113 as it is, the process proceeds to step S55.
[0133]
 The other hand, if it is determined not to be the reference frame in step S53, the process proceeds to step S54.
[0134]
 In step S54, the spread information generating unit 112, for each object, and generates a spread information of the frame to be processed based on the supplied metadata from the unpacking / decrypting section 111.
[0135]
 For example spread information generating unit 112, based on the object position information included in the metadata of the processing target frame (current frame), to the object position information and spreads information included in the meta data of the reference frame supplied in the past above equation (1), or by calculating the above equations (2) and (3), and generates a spread information processing target frame.
[0136]
 Spread information generating unit 112, as for each object to produce a spread information, supplies the generated spread information, stored in their metadata for each object, the meta data including the object position information and the spread information to the rendering unit 113 to.
[0137]
 For each object in this manner, the metadata including the object position information and the spread information is obtained, then the process proceeds to step S55.
[0138]
 Or reference is judged to be the frame, or when the spread information in step S54 is generated in step S53, step S55 is performed.
[0139]
 In step S55, the rendering unit 113 performs rendering of the audio signal of the object based audio signal of an object supplied from the unpacking / decrypting unit 111, and the metadata supplied from the spread information generation unit 112.
[0140]
 For example rendering unit 113, for each object, based on the spread information and object position information, and generates an audio signal of each channel by VBAP (Vector Base Amplitude Pannning), supplied to the mixing section 114. In step S55, the sound image of the object is localized at the position indicated by the object position information, the audio signal as a sound image spreads in breadth degree indicated by the spread information is generated.
[0141]
 In step S56, the mixing unit 114, weighted addition and the audio signal of each channel supplied from the unpacking / decrypting unit 111, an audio signal of each channel supplied from the rendering unit 113 for each channel, an external speaker supplied to. Thus, each speaker, since the audio signal of the channel corresponding to those of the speaker is supplied, each speaker reproduces sound based on the supplied audio signal.
[0142]
 In step S57, the decoding apparatus 101 determines whether to end the process. For example, it is determined that the output to the speaker of the audio signal for every frame if ended, the process ends in step S57.
[0143]
 If it is determined not to end the process at step S57, the process returns to step S51, processing for the next frame is described above is a frame processed is repeatedly performed.
[0144]
 In contrast, if it is determined that the process ends in step S57, the respective units of the decoding apparatus 101 stops the process performed, the decoding process ends.
[0145]
 As described above, the decoding apparatus 101 performs rendering to generate spread information of the frame is not a reference frame of each object.
[0146]
 Thus, it is possible to reduce the cost of the spread information giving manual, obtained more easily correct spread information. Further, if to generate spread data frame other than the reference frame in the decoding device 101 side, it is not necessary to store the spread information of a frame other than the reference frame in the bit stream. Therefore, it is possible to reduce the bit rate of the bit stream to be transmitted to the decoding device 101.
[0147]

 In the way first embodiment has described the case where application of the spread information is manually for some frames of the audio signal of the object.
[0148]
 However, even many exist content that spreads information has not been completely granted.
[0149]
 For example, the encoding scheme that can handle object audio, some are configured to be capable of switching whether to include a spread information to the metadata in the flag of the header portion. That is, the presence of a bit stream spread information is not applied is allowed.
[0150]
 Further, in the first place coding method object audio spreads information is not included in the bitstream is also present.
[0151]
 Against this background, many present content that is not properly applied in spreads information all the time (frame), as a result, for their content could not be performed with high realism playback.
[0152]
 Therefore, by generating the spread information based and object position information of the object, on the size information indicating the size of the object, without the application of spread information manual, it is possible to obtain an appropriate spread information it may be so.
[0153]
 For example, in the encoding method of some object audio, there is one that contains size information of each object in the bit stream. If there is a size information of the object, without the application of spread information manual, it can generate a spread information using the object position information and size information.
[0154]
 Specifically, for example, that the sound image of the object sound is spread symmetrically in the horizontal and vertical directions, information indicating the distance from the center of the object to the outer shell (end) is the size information of the object and things. Hereinafter, it is assumed that marks the size information of the frame the frame index is i and d (i).
[0155]
 In this case, for example, the center position of the object and the point OB21 in the current frame ifrm as shown in FIG. 6, when the position of the end part of the object (shell portion) and the point OBE21, distance from the point OB21 to the point OBE21 size is information d (ifrm), that is, the distance indicated by the size information. Note that portions corresponding to the case in FIG. 1 in FIG. 6 are denoted with the same reference numerals, and description thereof will be omitted as appropriate.
[0156]
 In this example, the position of the point OB21 is determined by the radius r (IFRm) included in the object position information. Also, the position of the point OB21, the position of the color point OBE21 and size information d (IFRm) is determined. Here, the point OBE21 is assumed to be located on a straight line orthogonal to the straight line L21 connecting the origin O and the point OB21. That is, the straight line L23 connecting the point OB21 and the point OBE21 is assumed to be a straight line perpendicular to the straight line L21.
[0157]
 Thus, when the object position information and size information of the object is given, and r (IFRm) is the length of the straight line L21, and d (IFRm) is the length of the straight line L23 are known.
[0158]
 When a straight line connecting the origin O and the point OBE21 a straight line L22, spread information s of the current frame ifrm (ifrm), i.e. the spread angle is the angle between the straight line L22 and the line L21. Therefore, it is possible to determine the spread information s (IFRm) by the following equation using the radius r (IFRm) and size information d (ifrm) (4).
[0159]
[Formula 4]

[0160]
 In Equation (4), the size information d of the current frame IFRm object (IFRm), the ratio of the radius r (IFRm) of the current frame IFRm object d (IFRm) / r (IFRm) is obtained, the ratio d (ifrm) / arctangent of r (IFRm) is generated as a spread information s (IFRm) of the current frame IFRm object.
[0161]
 Thus calculate the equation (4) based on the radius r (IFRm) and size information d of the current frame IFRm object (IFRm), suitable spread information s of the current frame IFRm object by a simple operation ( ifrm) can be generated.
[0162]
 In particular, it is not necessary to provide a spread information manually in this example, it is possible to significantly reduce the applied cost spread information. That is, even when the frame spread information is added do not exist, it is possible to easily generate the appropriate spread information to all frames. Thus, it is possible to realize a high content reproducing realistic using spread information.
[0163]
 Further, it is possible to generate a spread information by calculation similar to that of equation (4) also when the sound image of the sound of the object is widened asymmetrically in the horizontal and vertical directions.
[0164]
 That is, for example shell from the center of the object in the horizontal direction, i.e. the distance to the end portion of the horizontal direction and the horizontal distance, the outer shell from the center of the object in the vertical direction, i.e. the distance to the end portion in the vertical direction and vertical distance.
[0165]
 Further, the object of the current frame IFRm, horizontal distance d width (IFRm) and the vertical distance d height and (IFRm) are included as size information.
[0166]
 In this case, similarly to the equation (4) above, the horizontal spread information s of the current frame IFRm by calculating the following equation (5) width can be obtained (IFRm), the following equation (6) vertical spread information s of the current frame IFRm by calculating the height can be obtained (IFRm).
[0167]
[Formula 5]

[0168]
[Number 6]

[0169]
 In equation (5), the radius r (IFRm) of the current frame IFRm, horizontal distance d as size information width (IFRm) and horizontal spread information s based on the width (IFRm) is calculated. Similarly, in Formula (6), the radius r (IFRm) of the current frame IFRm, vertical distance as size information d height (IFRm) and vertical spread information s based on the height (IFRm) is calculated that.
[0170]

 If the spread information is generated based on object position information and size information as described above, in the encoding device 11 shown in FIG. 2, spread information generation unit from the metadata input unit 23 the 24, metadata and size information is supplied. At this time, the meta data includes the object position information, in a state that is not included spread information.
[0171]
 Spread information generation unit 24 generates a spread information based on the object position information and size information supplied from the metadata input section 23.
[0172]
 If spread information in this manner is produced, the encoding process shown in FIG. 7, the encoding apparatus 11 is performed. Hereinafter, with reference to the flowchart of FIG. 7 will be described encoding process by the encoding device 11.
[0173]
 In step S81, the spread information generation unit 24 generates a spread information.
[0174]
 That is, the meta data input unit 23 or receives the user input operation, or communicate with the outside by or performs reading from outside of the recording area, the metadata and the size of each object in the processing object frame It is to obtain information, and supplies the spread information generating unit 24.
[0175]
 Then, spread information generating unit 24, based on the metadata and size information supplied from the metadata input unit 23, and generates a spread information for each object.
[0176]
 Specifically, spread information generation unit 24, the object position information included in the metadata of the processing target frame, described above on the basis of the size information of the processing target frame formula (4), or the above-described formula (5) and by calculating the equation (6), and generates a spread information processing target frame.
[0177]
 The spread information generating unit 24 supplies the generated spread information, stored in their metadata for each object, and meta data including object position information and the spread information, and size information to the packing unit 25 .
[0178]
 When the spread information is generated, then the processing of step S82 to step S88 is performed the encoding process is terminated, since these processes are the same as the processing in steps S13 through step S19 in FIG. 3, the description thereof is omitted. However, in step S82, the size information of the objects with metadata also stored in the DSE of the bitstream.
[0179]
 Encoding device 11 as described above may be stored in the bit stream to generate a spread information based on the object position information and size information, and outputs. Thus, it is possible to reduce the cost of the spread information addition, to obtain a more easily correct spread information.
[0180]

 In addition, if the spread information is generated based on object position information and size information, the spread information in the decoding device 101 shown in FIG. 4 generation can be.
[0181]
 If the spread information is generated in the decoding apparatus 101, decoding process is performed as shown in FIG. 8 the decoding device 101. Hereinafter, decoding processing by the decoding apparatus 101 with reference to the flowchart of FIG. 8 will be described.
[0182]
 The processing of steps S111 and step S112 is the same as the processing in steps S51 and step S52 in FIG. 5, a description thereof will be omitted.
[0183]
 However, in step S112, the metadata and the size information of each object is acquired in the unpacking / decrypting unit 111 by unpacking the bitstream, these metadata and size information is supplied to the spread information generating unit 112 that. In this case, it has been in the metadata acquired by the unpacking. The object position information is included, does not contain the spread information.
[0184]
 In step S113, the spread information generation unit 112 generates a spread information of the frame to be processed based on the unpacking / supplied metadata and magnitude information from the decoding unit 111 for each object.
[0185]
 For example spread information generation unit 112, the object position information included in the metadata of the processing target frame, described above on the basis of the size information of the processing target frame formula (4), or the above-described formula (5) and by calculating the equation (6), and generates a spread information processing target frame.
[0186]
 Spread information generating unit 112, when generating the spread information for each object, supplies the generated spread information, stored in their metadata for each object, the meta data including the object position information and the spread information to the rendering unit 113 to.
[0187]
 When the process of step S113 is performed, then, although the decoding processing process is performed in step S114 to step S116 ends, since these processes are the same as those in step S55 to step S57 of FIG. 5, the description thereof is omitted.
[0188]
 As described above, to generate a spread information based decoding device 101 to object position information and size information of each object, perform rendering.
[0189]
 Thus, it is possible to reduce the cost of the spread information addition, to obtain a more easily correct spread information. Since it is not necessary to store the spread information bit stream, it is possible to reduce the bit rate of the bit stream to be transmitted to the decoding device 101.
[0190]

 Incidentally, in the encoding scheme that can handle a number of object audio is to modify the object position information based on the zoom information the user specifies, corresponding to the zoom It is realized the movement of the object.

claims

And spread information different from the frame and the processing target frame object position information and the audio object to be processed in the frames of audio objects, and zoom information, to any one of the size information of the audio objects based on, and a spread information generation unit for generating spread information of the processing target frame of the audio object
 signal processor.
[Requested item 2]
 The spread information generating unit, and the spread information of the other frames of the audio objects, wherein the processed audio object frame and the other of said object position information and the frame of the processing target based on the frame generating the spread information
 signal processing apparatus according to claim 1.
[Requested item 3]
 The object position information, from the reference position is the distance to the audio object
 signal processing apparatus according to claim 2.
[Requested item 4]
 The spread information generation unit, and the ratio of the distance in the distance and the other frames in the frame of the processing target, the other said of the processing target arctangent of the product of the tangent of the spread information in Frame It is generated as the spread information
 signal processing apparatus according to claim 3.
[Requested item 5]
 The spread information generating unit is configured and zoom information, based on said spread information of the processing target frame of the audio objects, generating the spread information of the processing target frame of the audio objects in the after zooming
 claims the signal processing apparatus according to claim 1.
[Requested item 6]
 The spread information generating unit based on the relationship between the position and size of the virtual screen before and after the zoom determined by the zoom information, and generates the spread information of the processing target frame of the audio object after the zoom
 claim the signal processing apparatus according to 5.
[Requested item 7]
 The spread information generating unit, and the size information, on the basis of the said object position information of the processing target frame of audio objects, generating the spread information of the frame of the processing target
 of claim 1 signal processor.
[Requested item 8]
 The size information is the distance to the outer shell from the center of the audio object
 signal processing apparatus according to claim 7.
[Requested item 9]
 The object position information, from the reference position is the distance to the audio object
 signal processing apparatus according to claim 8.
[Requested item 10]
 The spread information generating unit generates said size information, the inverse tangent of the ratio between the object position information of the frame of the processing target as the spread information of the frame of the processing target
 signal processing according to claim 9 apparatus.
[Requested item 11]
 And spread information different from the frame and the processing target frame object position information and the audio object to be processed in the frames of audio objects, and zoom information, to any one of the size information of the audio objects based on, it generates a spread information of the processing target frame of the audio object
 signal processing method comprising the steps.
[Requested item 12]
 And spread information different from the frame and the processing target frame object position information and the audio object to be processed in the frames of audio objects, and zoom information, to any one of the size information of the audio objects based on, it generates a spread information of the processing target frame of the audio object
 program for executing the processing including a step in the computer.

Documents

Application Documents

# Name Date
1 201917042351.pdf 2019-10-18
2 201917042351-TRANSLATIOIN OF PRIOIRTY DOCUMENTS ETC. [18-10-2019(online)].pdf 2019-10-18
3 201917042351-STATEMENT OF UNDERTAKING (FORM 3) [18-10-2019(online)].pdf 2019-10-18
4 201917042351-PROOF OF RIGHT [18-10-2019(online)].pdf 2019-10-18
5 201917042351-PRIORITY DOCUMENTS [18-10-2019(online)].pdf 2019-10-18
6 201917042351-POWER OF AUTHORITY [18-10-2019(online)].pdf 2019-10-18
7 201917042351-FORM 1 [18-10-2019(online)].pdf 2019-10-18
8 201917042351-DRAWINGS [18-10-2019(online)].pdf 2019-10-18
9 201917042351-DECLARATION OF INVENTORSHIP (FORM 5) [18-10-2019(online)].pdf 2019-10-18
10 201917042351-COMPLETE SPECIFICATION [18-10-2019(online)].pdf 2019-10-18
11 abstract.jpg 2019-10-19
12 201917042351-OTHERS-221019.pdf 2019-10-25
13 201917042351-Correspondence-221019.pdf 2019-10-25
14 201917042351-FORM 3 [27-03-2020(online)].pdf 2020-03-27
15 201917042351-FORM 18 [10-03-2021(online)].pdf 2021-03-10
16 201917042351-FER.pdf 2022-02-01
17 201917042351-OTHERS [01-08-2022(online)].pdf 2022-08-01
18 201917042351-FER_SER_REPLY [01-08-2022(online)].pdf 2022-08-01
19 201917042351-DRAWING [01-08-2022(online)].pdf 2022-08-01
20 201917042351-CORRESPONDENCE [01-08-2022(online)].pdf 2022-08-01
21 201917042351-COMPLETE SPECIFICATION [01-08-2022(online)].pdf 2022-08-01
22 201917042351-CLAIMS [01-08-2022(online)].pdf 2022-08-01
23 201917042351-PatentCertificate16-05-2024.pdf 2024-05-16
24 201917042351-IntimationOfGrant16-05-2024.pdf 2024-05-16

Search Strategy

1 SearchHistory(6)E_07-09-2021.pdf

ERegister / Renewals

3rd: 23 Jul 2024

From 11/04/2020 - To 11/04/2021

4th: 23 Jul 2024

From 11/04/2021 - To 11/04/2022

5th: 23 Jul 2024

From 11/04/2022 - To 11/04/2023

6th: 23 Jul 2024

From 11/04/2023 - To 11/04/2024

7th: 23 Jul 2024

From 11/04/2024 - To 11/04/2025

8th: 05 Apr 2025

From 11/04/2025 - To 11/04/2026