Sign In to Follow Application
View All Documents & Correspondence

Signal Processing Device Method And Program

Abstract: The present art pertains to a signal processing device method and program that enable the calculation amount for decoding to be reduced at low cost. The signal processing device is provided with a priority information generation unit for generating priority information for audio objects on the basis of a plurality of elements representing characteristics of the audio objects. The present art is applicable to a coding device and a decoding device.

Get Free WhatsApp Updates!
Notices, Deadlines & Correspondence

Patent Information

Application #
Filing Date
18 October 2019
Publication Number
50/2019
Publication Type
INA
Invention Field
ELECTRONICS
Status
Email
patents@remfry.com
Parent Application
Patent Number
Legal Status
Grant Date
2025-02-19
Renewal Date

Applicants

SONY CORPORATION
1-7-1, Konan, Minato-ku, Tokyo 1080075

Inventors

1. YAMAMOTO Yuki
c/o SONY CORPORATION, 1-7-1, Konan, Minato-ku, Tokyo 1080075
2. CHINEN Toru
c/o SONY CORPORATION, 1-7-1, Konan, Minato-ku, Tokyo 1080075
3. TSUJI Minoru
c/o SONY CORPORATION, 1-7-1, Konan, Minato-ku, Tokyo 1080075

Specification

0001]This technique, the signal processing apparatus and method, and a program, the signal processing apparatus and method capable of reducing the computational complexity of the decoding at a low cost, and a program.
BACKGROUND
[0002]Conventionally, as a coding technique that can handle object audio, such as an international standard MPEG (Moving Picture Experts Group) -H Part 3: such 3D audio standards are known (e.g., see Non-Patent Document 1).
[0003]
 In such coding schemes, by transmitting the priority information indicating the priority of each audio object decoding apparatus, reduction of the calculation amount at the time of decoding is realized.
[0004]
 For example, if a large number of audio objects, when to perform decoding only high audio object priority based on the priority information, it is possible to reproduce the content with sufficient quality with a small calculation amount.
CITATION
Non-patent literature
[0005]
非特許文献1 : INTERNATIONAL STANDARD ISO/IEC 23008-3 First edition 2015-10-15 Information technology - High efficiency coding and media delivery in heterogeneous environments - Part 3: 3D audio
Summary of the Invention
Problems that the Invention is to Solve
[0006]
 However, the high cost to impart manually every each and audio object priority information time. For example, in order to deal with for a long time a lot of audio objects in the movie content, it can be said that the cost is particularly high by hand.
[0007]
 Also, many existing content priority information is not added. For example, the above-mentioned MPEG-H Part 3: The 3D audio standard, whether to include priority information in the encoded data can be switched by the header portion of the flag. That is, even the presence of encoded data priority information is not assigned is permitted. Furthermore, the first place priority information also exists encoding method object audio that is not included in the encoded data.
[0008]
 Against this background, the priority information There are many coded data that has not been granted, so that, for those of the encoded data can not be reduced calculation amount of the decoding.
[0009]
 This technology has been made in view of such circumstances, it is desirable to make it possible to reduce the computational complexity of the decoding at a low cost.
Means for Solving the Problems
[0010]
 Signal processing apparatus according to an embodiment of the present technology, based on the plurality of elements representing the characteristics of audio objects, comprises a priority information generator for generating priority information of the audio object.
[0011]
 The element may be metadata of the audio object.
[0012]
 The element may be a position of the audio object in the space.
[0013]
 It can be a distance of said elements from a reference position on the space to the audio object.
[0014]
 It can be a horizontal angle of said elements indicating the horizontal position of the audio object on the space.
[0015]
 Wherein the priority information generating unit, the priority information corresponding to the moving speed of the audio object based on the metadata can be generated.
[0016]
 It can be a gain information that is multiplied with the elements in the audio signal of the audio objects.
[0017]
 The priority information generating unit, and the gain information in units of the processing target time, based on the difference between the average value of the gain information of the plurality of unit time, the priority information of the unit of processing target time it can be generated.
[0018]
 Wherein the priority information generating unit, the priority information can be generated based on the sound pressure of the audio signal in which the gain information is multiplied.
[0019]
 The element can be spread information.
[0020]
 Wherein the priority information generating unit, on the basis of the spread information, the priority information corresponding to the area of ​​the region of the audio object can be generated.
[0021]
 It may be information indicating an attribute of the element tone of the audio object.
[0022]
 The element may be an audio signal of the audio objects.
[0023]
 Wherein the priority information generator may generate the priority information based on the result of the speech section detection processing on the audio signal.
[0024]
 Wherein the priority information generating unit, to perform the time direction smoothing on the generated the priority information may be a final said priority information.
[0025]
 Signal processing method or a program according to an embodiment of the present technology, based on the plurality of elements representing a feature of the audio object, comprising the step of generating the priority information of the audio object.
[0026]
 In one aspect of the present technology, based on the plurality of elements representing a feature of the audio object, the priority information of the audio objects is generated.
The invention's effect
[0027]
 According to the embodiments of the present technology, it is possible to reduce the computational complexity of the decoding at a low cost.
[0028]
 Here, the advantages described in the present invention is not necessarily limited, it may be any of the effects described in the present disclosure.
BRIEF DESCRIPTION OF THE DRAWINGS
[0029]
Is a diagram illustrating a configuration example of FIG. 1 encoder.
It is a diagram illustrating a configuration example of FIG. 2 object audio encoding unit.
Is a flow chart illustrating the FIG. 3] encoding process.
Is a diagram illustrating a configuration example of FIG. 4 decoding apparatus.
Is a diagram illustrating a configuration example of FIG. 5 unpacking / decrypting unit.
Is a flowchart illustrating a FIG. 6 decoding process.
7 is a flowchart illustrating selection decoding process.
Is a diagram illustrating a configuration example of FIG. 8 computer.
DESCRIPTION OF THE INVENTION
[0030]
 Hereinafter, with reference to the accompanying drawings, a description will be given of an embodiment according to the present technology.
[0031]

 This technology, and meta data of the audio objects, content information, based on the distinctive element of the audio objects, such as audio signals of audio objects, an audio object by generating the priority information, it is to be able to reduce the calculation amount of the decoding at a low cost.
[0032]
 In the following, an audio signal and the audio signal of the audio objects of the multi-channel will be described as being encoded according to such predetermined standards. In the following it will be referred to as simply both object audio objects.
[0033]
 For example, an audio signal of each channel and each object is transmitted is encoded for each frame.
[0034]
 That, and the audio signal encoded, stored information needed to decode such audio signal into a plurality of elements (bit stream element), those composed of elements bitstream is transmitted to the decoding side from the encoding side .
[0035]
 More specifically, for example, 1 to frame of bitstream beginning plurality of elements are arranged in order from, is disposed identifier indicating that the end is the end position related information of the frame.
[0036]
 Then, the elements arranged on the top are the ancillary data region called DSE (Data Stream Element), etc. Information and identification information on the downmix audio signal is DSE, information is described regarding each of a plurality of channels .
[0037]
 In addition, each element following the DSE, encoded audio signals are stored. In particular, the elements that audio signal of a single channel is stored called SCE is called (Single Channel Element), the elements that audio signals of two channels in the pair is stored CPE (Coupling Channel Element) It has been. Audio signals of each object are stored in SCE.
[0038]
 In this technique, priority information of the audio signals of each object is generated and stored in the DSE.
[0039]
 Here, the priority information is information indicating the priority of the object, in particular the value of the priority indicated by the priority information, i.e. the greater the numerical value indicating the priority level, the priority of the object is high, an important object it is shown that there is.
[0040]
 In the applied coding apparatus to which the present technique is based on the meta data of the object, the priority information of each object is generated. Accordingly, even when the priority information is not assigned to the content, it is possible to reduce the computational complexity of the decoding. In other words, without the application of priority information manual, it is possible to reduce the computational complexity of the decoding at a low cost.
[0041]
 Next, detailed embodiments of the applied coding apparatus to which the present technology.
[0042]
 Figure 1 is a diagram illustrating a configuration example of a coding apparatus to which the present technology is applied.
[0043]
 Encoding device 11 shown in FIG. 1 has a channel audio encoder 21, an object audio encoding unit 22, the metadata input unit 23 and packing unit 24.
[0044]
 The channel audio coding unit 21, the audio signals of each channel of the multi-channel number of channels is M is supplied. For example an audio signal of each channel is supplied from the microphone corresponding to those channels. In Figure 1, the character "# 0" to "# M-1" represents the channel number of each channel.
[0045]
 Channel audio encoder 21 encodes the audio signal of each channel supplied, supplies the encoded data obtained by encoding to the packing unit 24.
[0046]
 The object audio encoding unit 22, the audio signal of each of the N object is supplied. For example audio signals of each object is supplied from the microphone attached to those objects. In Figure 1, the character "# 0" to "# N-1" represents the object number of each object.
[0047]
 Object audio encoding unit 22 encodes the audio signals of each object supplied. Also, object audio encoding unit 22, the supplied audio signal, metadata generates priority information based from the input unit 23 on the supplied metadata and content information such as, encoded data obtained by encoding If, supplies the priority information to the packing unit 24.
[0048]
 The metadata input unit 23 supplies the meta data and content information of each object to the object audio encoding unit 22 and packing unit 24.
[0049]
 For example, object meta data includes such as gain information indicating a gain of the spread information, an audio signal of the object that indicates the size range of the sound image of the object position information, object indicating the position of the object in space. The content information includes information about the attributes of the sound of each object in the content.
[0050]
 Packing unit 24, the encoded data supplied from the channel audio encoder 21, and the supplied encoded data from an object audio encoding unit 22 priority information, and the metadata supplied from the metadata input unit 23 packing the content information to generate a bit stream, and outputs.
[0051]
 Thus a bit stream obtained by are included each channel of coded data for each frame, each object of the encoded data, priority information of each object, and meta data and content information for each object.
[0052]
 Here, one frame of M audio signal of each channel to be stored in the bit stream, and the N audio signals of each object is an audio signal of the same frame to be reproduced simultaneously.
[0053]
 Here, as the priority information of the audio signals of each object will be described an example in which the priority information is generated for each audio signal for each frame as a unit any given time, for example, several it may be one of the priority information to the frame of audio signal is generated.
[0054]

 In addition, object audio encoding unit 22 of FIG. 1, is configured as shown in FIG. 2 for example and more.
[0055]
 Object audio encoding unit 22 shown in FIG. 2 includes a coding unit 51 and the priority information generating unit 52.
[0056]
 Coding section 51 is provided with MDCT (Modified Discrete Cosine Transform) unit 61, encoding unit 51 encodes the audio signals of each object supplied from the outside.
[0057]
 That, MDCT unit 61 performs MDCT (Modified Discrete Cosine Transform) to the audio signals of each object supplied from the outside. Coding section 51, the MDCT coefficients for each object obtained by the MDCT encoding, resulting each object coded data, i.e. supplying the encoded audio signal to the packing unit 24.
[0058]
 Further, priority information generating unit 52, the audio signals of each object supplied from the outside, the metadata supplied from the metadata input unit 23, and from the metadata input unit 23 to at least one of the supplied content information based generates priority information of the audio signals of each object, and supplies the packing unit 24.
[0059]
 In other words, the priority information generator 52, an audio signal and metadata, such as content information, based on one or more distinctive element of the object, and generates the priority information of the object. For example an audio signal is a component representing the features of sounds of an object, the metadata is an element representing a feature such as spreading degree, the gain of the position and the sound image of the object, element content information representing the characteristics about the attributes of a sound object it is.
[0060]

 will now be described priority information of objects to be generated in the priority information generator 52.
[0061]
 For example, it is conceivable to generate the priority information based on only the sound pressure of the audio signal of the object.
[0062]
 However, the object meta data is stored gain information, since the audio signal the gain information is multiplied will be used as an audio signal of the final object, the audio signals before and after the multiplication of the gain information is the sound pressure changes.
[0063]
 Accordingly, even when generating the priority information based on only the sound pressure of the audio signal, not necessarily appropriate priority information is obtained. Therefore, the priority information generator 52, the priority information is generated sound pressure or outside the information of at least the audio signal is used. Thus, it is possible to obtain an appropriate priority information.
[0064]
 Specifically, the priority information is generated by at least one of the following methods of (1) to (4).
[0065]
(1) generates the priority information based on the object metadata
(2) generates the priority information based on information other than the metadata
(3) combining the resulting priority information of a plurality of methods generating one of the priority information Te
generates a final one of the priority information (4) to smooth the priority information in the time direction
[0066]
 First, discussion will be made on generation of the priority information based on the object metadata.
[0067]
 As described above, it includes an object position information, the spread information, and gain information on the metadata of the object. Therefore, and these object position information, the spread information, it is considered that by using the gain information to generate the priority information.
[0068]
 (1-1) for generating priority information based on the object position information
 First, an example of generating the priority information based on object position information.
[0069]
 Object position information is information indicating the position of the object in three-dimensional space, for example the horizontal angle a indicating the position of the object as viewed from the reference position (origin), the coordinate information comprising a vertical angle e, and the radius r It is.
[0070]
 Horizontal angle a, horizontal angle (azimuth angle) in the horizontal direction position of the object as viewed from the reference position is a position where the user exists, that is the object as seen from the direction and the reference position serving as a reference in the horizontal direction it is the angle between the direction.
[0071]
 Here, when the horizontal angle a is 0 degrees, the object is located in front of the user, when the horizontal angle a is 90 degrees or -90 degrees, the object is located just beside the user It will be. Further, when the horizontal angle a is 180 degrees or -180 degrees, the object will be positioned directly behind the user.
[0072]
 Similarly the vertical angle e is the angle in the vertical direction indicated by the vertical position of the object as viewed from the reference position (elevation angle), i.e. the angle between the direction of the object as viewed from a reference direction and the reference position in the vertical direction it is.
[0073]
 Further, the radius r is the distance from the reference position to the position of the object.
[0074]
 For example a short distance object from a location of the user origin (reference position), i.e. the radius r is small, the object at the position close to the origin, is considered to be more important than the object from the origin to the distant position. Therefore, it is possible to make the priority indicated by the priority information as the radius r is smaller becomes higher.
[0075]
 In this case, for example, the priority information generator 52, based on the radius r of the object by calculating the following equation (1), and generates the priority information of the object. In the following it is assumed that also referred to as a priority the priority information.
[0076]
[Number 1]

[0077]
 In the example shown in equation (1), the value of the priority information priority higher radius r is small becomes large, the priority becomes higher.
[0078]
 Also, human hearing is known to have a high sensitivity to the front than the rear. Therefore, the object behind the user, the effect on the hearing of the user even if the different decoding process shall be made originally to lower the priority is considered to be small.
[0079]
 Therefore, the more objects behind the user, i.e. it is possible priority indicated by the object as priority information in a position close to the right behind the user to be lower. In this case, for example, the priority information generator 52, based on the horizontal angle a of an object by calculating the following equation (2), and generates the priority information of the object. However, when the horizontal angle a is less than 1 °, the value of the priority information priority objects are 1.
[0080]
[Number 2]

[0081]
 Incidentally, abs (a) in formula (2) represents the absolute value of the horizontal angle a. Therefore, the horizontal angle a in this example is small, the value of the more priority information priority closer to the direction of the position directly in front position of the object is viewed from the user is increased.
[0082]
 Furthermore, objects moving time variation is large objects in the object position information, i.e., at a high rate is considered likely an important object in the content. Therefore, it is possible to make larger the time variation of the object position information, i.e. the priority the moving speed of the object is indicated by faster the priority information is high.
[0083]
 In this case, for example, the priority information generator 52, the horizontal angle a included in the object position information of the object, by vertical angles e, and based on the radius r to calculate the following formula (3), the object generating the priority information corresponding to the moving speed.
[0084]
[Number 3]

[0085]
 Incidentally, a (i) in equation (3), e (i), and r (i) is in the current frame to be respectively processed, it indicates the horizontal angle a, the vertical angle e, and the radius r of the object ing. Further, a (i-1), e (i-1), and r (i-1) is, in the temporally previous frame of the current frame to be respectively processed, the horizontal angle a of the object, shows a vertical angle e, and the radius r.
[0086]
 Thus, for example, (a (i) -a (i-1)) shows the horizontal speed of the object, the right side of the equation (3) corresponds to the speed of the entire object. That is, the value of the priority information priority represented by the formula (3) is increased as the speed of the object it is fast.
[0087]
 (1-2) for generating priority information based on the gain information
 Next, an example of generating the priority information is described based on the gain information.
[0088]
 For example, the metadata of the object, the coefficient values ​​to be multiplied to the audio signal of the object at the time of decoding is included as the gain information.
[0089]
 The value of the gain information, that is, as the coefficient value of the gain information is large, the sound pressure of the final object of the audio signal after the coefficient value multiplied is increased, thereby believed sound objects is likely to be perceived by humans. Further, the object to increase the sound pressure by applying a large gain information is considered to be an important object in the content.
[0090]
 Therefore, as the value of the gain information is large, it is possible to make the priority indicated by priority information of the object becomes higher.
[0091]
 In such a case, for example, the priority information generator 52, by calculating the following equation (4) based on the gain information of the object, that is, the coefficient value g is a gain indicated by the gain information, the priority of the object to generate the information.
[0092]
[Formula 4]

[0093]
 In the example shown in Equation (4), the coefficient value g itself is the priority information priority is the gain information.
[0094]
 Further, the time average value of the gain information of a plurality of frames of one object (coefficient value g) the time average value g ave and be referred to. For example, the time average value g ave is a like time average value of the gain information of a plurality of consecutive frames in the past than the frame to be processed.
[0095]
 For example the gain information and the time average value g ave frame difference is large and, more coefficient value g is the time average value g more ave In the frame much greater than the coefficient value g and the time average value g ave is the difference between the importance of the object as compared to the small frame is considered to be high. In other words, the rapidly frame coefficient value g is increased, the importance of the object is considered to be high.
[0096]
 Therefore, the gain information and the time average value g ave as frame difference is large and can be made to the priority indicated by priority information of the object becomes higher.
[0097]
 In such a case, for example, the priority information generator 52, the gain information of the object, namely a coefficient value g, the time average value g ave by calculating the following equation (5) based on the priority of the object to generate the information. In other words, the coefficient value g of the current frame, the time average value g ave priority information is generated based on the difference between.
[0098]
[Formula 5]

[0099]
 G (i) denotes the coefficient values g of the current frame in equation (5). Thus, in this example, the coefficient value g of the current frame (i) is the time average value g ave larger than, the value of the priority information priority increases. That is, in the example shown in Equation (5), the frame gain information becomes abruptly larger is the high importance of an object, the higher the priority indicated by priority information.
[0100]
 The time average value g ave is past index average value or based on the gain information of a plurality of frames (coefficient value g) of an object, or an average value of the gain information of the object across the content.
[0101]
 (1-3) for generating priority information based on the spread information
 subsequently, an example of generating the priority information is described on the basis of the spread information.
[0102]
 Spread information, angle information indicating the size range of the sound of the object, that is, the angle information indicating the spreading degree of a sound image of a sound object. In other words, spread information may also be referred to as information indicating the size of the area of ​​the object. Hereinafter, as indicated by the spread information, and that the angle indicating the range of the sizes of the sound of the object is referred to as a spread angle.
[0103]
 Object spread angle is large, an object that is largely reflected in the screen. Thus, the object is spread angle large, considered likely an important object in the content as compared with object spread angle is small. Therefore, it is possible to make the priority becomes higher as indicated by the more objects spread angle is large priority information indicated by the spread information.
[0104]
 In such a case, for example, the priority information generator 52, on the basis of the spread information of the object by calculating the following equation (6), generates the priority information of the object.
[0105]
[Number 6]

[0106]
 Incidentally, it s in Formula (6) shows the spread angle indicated by the spread information. Area of ​​the region of the object in this example, i.e. to reflect the breadth of the range of the sound image to the value of the priority information priority, the square value of the spread angle s is the value of the priority information priority. Therefore, the calculation of equation (6), the area of ​​the region of the object, that will be the priority information corresponding to the area of ​​the sound image of a region of the object sound is generated.
[0107]
 Further, it may different directions as spread information, that is, perpendicular horizontal and vertical spread angles given.
[0108]
 For example, as a spread information, horizontal spread angle s width spread angle s of the vertical height and are included with. In this case, the horizontal and vertical directions sizes depending spread information is different, i.e. it is possible to spread condition to represent different objects.
[0109]
 Spread angle s as spread information such width and spread angle s height when it is included, the priority information generator 52, by calculating the following equation (7) based on the spread information of the object, the object to generate a priority information.
[0110]
[Number 7]

[0111]
 In equation (7), spread angle s width and spread angle s height product of is the priority information priority. By generating the priority information by Equation (7), as in equation (6), as the object is spread angle larger, that is, as the area of the object is large, the priority is increased as indicated by the priority information it can be so.
[0112]
 Further, in the above, it has been described object position information, the spread information, and that gain information about an example of generating the priority information based on the object metadata. However, it is also possible to generate the priority information based on information other than the metadata.
[0113]
 (2-1) for generating priority information based on the content information
 first, as a generation example of the priority information based on information other than the metadata, an example of generating the priority information is described with reference to content information.
[0114]
 For example, in the encoding method of some object audio, there is one that contains content information as the information on each object. For example the attribute of the sound object is specified by the content information. That includes information indicating the attributes of a sound object in the content information.
[0115]
 Specifically, for example, whether the sound of the object by the content information is dependent on the language, the language of the sound type of the object, whether the sound of the object is a voice, and the sound of the object is in the environmental sound whether it is possible to identify.
[0116]
 For example, when the sound of the object is a voice, the object is considered in comparison with objects such as other environmental sounds, is more important. This is because, in the content such as movies or news, the amount of information by voice is larger than the amount of information by other sounds, also, the human hearing is because it is more sensitive to sound.
[0117]
 Therefore, it is possible priority object is speech, to be higher than the priority of the objects of other attributes.
[0118]
 In this case, for example, the priority information generator 52, the computation of the following equation based on the content information of the object (8), and generates the priority information of the object.
[0119]
[Number 8]

[0120]
 Incidentally, object_class in equation (8) indicates the attribute of the sound of the object indicated by the content information. In equation (8), if the attribute of the sound object indicated by the content information is a voice (speech), the value of the priority information is a 10, if the attribute of the sound object indicated by the content information is not the voice , i.e. for example in the case which is the environmental sound value of the priority information is a 1.
[0121]
 (2-2) for generating priority information based on the audio signal
 also, whether the object is a voice can be identified by using a VAD (Voice Activity Detection) technique.
[0122]
 Therefore, VAD relative to e.g. objects of the audio signal, i.e. performs voice section detection processing, may be generated priority information of the object based on the detection result (processing result).
[0123]
 Similar to the case of using the content information also in this case, as a result of the speech section detection processing, when the sound of the object is obtained detection result indicating the voice, than when other detection result is obtained also, the priority indicated by the priority information is to be higher.
[0124]
 Specifically, for example, the priority information generator 52 performs voice section detection processing on the audio signal of the object, and generates the priority information of the object by calculation of the following equation (9) based on the detection result .
[0125]
[Number 9]

[0126]
 Incidentally, Object_class_vad shows the attributes of a sound of an object obtained as a result of the speech section detection processing in Equation (9). In equation (9), when the attribute of the sound object is a voice, that is, when the sound of the object as a detection result voice (speech) a is that of the detection result obtained by the speech section detection processing, the priority information the value is set to 10. Further, in Formula (9), when the attribute of the sound object is not a voice, that is, when the speech section detection processing detection result as detection indicating the sound object is the voice result of is not obtained, the value of the priority information It is set to 1.
[0127]
 Further, when the value of ness speech section is obtained as a result of the speech section detection processing, the priority information may be generated based on the value of the speech segment ness. In such a case, the current frame of the object is so that priority as likely speech section becomes higher.
[0128]
 (2-3) Audio signal and for generating the priority information based on the gain information
 further example, as described above, it is also conceivable to generate the priority information based on only the sound pressure of the audio signal of the object. However, the decoding side, the gain information included in the object meta data is multiplied by the audio signal, the sound pressure of the audio signal before and after multiplication of the gain information changes.
[0129]
 Therefore, even on the basis of the sound pressure of the gain information multiplication previous audio signal to generate the priority information may not appropriate priority information is obtained. Therefore, based on the sound pressure of the signal obtained by multiplying the gain information to the audio signal of the object, may generate the priority information. That may generate the priority information based on the gain information and the audio signal.
[0130]
 In this case, for example, the priority information generator 52, a gain information by multiplying the audio signal of the object, determining the sound pressure of the audio signal after the multiplication gain information. The priority information generator 52 generates the priority information based on the obtained sound pressure. In this case, for example, the sound as the pressure increases, the priority information is generated as the priority becomes high.
[0131]
 In the above, it explained metadata and content information of an object, such as an audio signal, for example of generating the priority information based on the distinctive element of the object. However, not limited to the example described above, for example, a value obtained by calculation such as equations (1), with respect to the calculated priority information, further or multiplied by a predetermined coefficient, and or by adding a predetermined constant things may be the final priority information.
[0132]
 (3-1) for generating priority information based on the object position information and the spread information
 also respectively coupled by such a linear combination or a nonlinear combination of the priority information obtained by different methods from each other (synthetic), final it may be one of the priority information. In other words, it may generate the priority information based on the plurality of elements representing a feature of an object.
[0133]
 By coupling a plurality of priority information, namely by combining a plurality of priority information, it is possible to obtain a more appropriate priority information.
[0134]
 Here, the priority information calculated on the basis of the object position information First, the priority information calculated on the basis of the spread information by linearly combining for an example of the final one of the priority information is described.
[0135]
 For example, the object even when it is in hard user backward to be perceived by the user, when the magnitude of the sound image of the object is considered that the object is an important object. On the contrary, even if the object is in front of the user, when the small size of the sound image of an object, the object is not considered to be important objects.
[0136]
 Therefore, for example, the priority information determined based on the object position information, by the linear sum of the priority information obtained based on spread information, it may be obtained a final priority information.
[0137]
 In this case, the priority information generator 52, for example, a plurality of priority information by calculating the following equation (10) linearly combined to produce a final one of the priority information about the object.
[0138]
[Formula 10]

[0139]
 In the equation (10), priority (position) indicates the priority information determined based on the object position information, priority (spread) shows the priority information determined on the basis of the spread information .
[0140]
 Specifically, priority (position), for example the formula (1) or Equation (2) shows a priority information obtained by such formula (3). priority (spread) illustrates, for example a priority information obtained by the equation (6) or expression (7).
[0141]
 Also, A and B in formula (10) represents the coefficient of linear sum. In other words A and B, it can be said that indicates the weighting coefficients used to generate the priority information.
[0142]
 For example, a method of setting the weighting factor of these A and B, can be considered the following two setting methods.
[0143]
 That is, as the first setting method, a method of setting the weight is equal in accordance with the value range by generation equation of the priority information being a linear combination (hereinafter, referred to as setting method 1) can be considered. Further, as the second setting method, a method of changing the weighting factor reported in the case (hereinafter, referred to as setting method 2) are considered.
[0144]
 Here it will be specifically described example of setting the weighting factor A and the weighting coefficient B by setting method 1.
[0145]
 For example, priority information determined by Equation (2) mentioned above is a priority (position), and priority information determined by the above-mentioned equation (6) is a priority (spread).
[0146]
 In this case, priority information priority (position) of the range is next 1 from 1 / [pi, range priority information priority (spread) from 0 [pi 2 becomes.
[0147]
 Therefore, becomes the value prevailing of formula (10) in the priority information priority (spread), and finally the value of the priority information priority obtained is hardly depends on the value of the priority information priority (position) it becomes a thing.
[0148]
 Therefore, in consideration of the range of both the priority information priority (position) and the priority information priority (spread), for example, the ratio of the weighting factors A and the weighting factor B [pi: if 1, eventually a more equal weight it can generate the Do priority information priority.
[0149]
 In this case, the weighting factor A is π / (π + 1), and the weighting coefficient B becomes 1 / (π + 1).
[0150]
 (3-2) for generating priority information based on the content information and other information
 further to a non-linear combination of the respective priority information obtained by different methods from each other, and the final one of the priority information examples will be described.
[0151]
 Here, for example, a calculated priority information based on the content information, for example the priority information calculated based on information other than the content information nonlinear coupling to a final one of the priority information is described.
[0152]
 For example, referring to the content information, it is possible to sound the object to identify whether the speech. If the sound object is voice, be any information other information other than the content information used for generating the priority information, the value of the priority information finally obtained is preferably large. This object generally speech it is considered that many information than the other objects, the more important objects.
[0153]
 Therefore, the calculated priority information based on the content information, if the final priority information by combining the calculated priority information based on information other than the content information, for example, the priority information generator 52 , using the weight coefficient determined by the setting method 2 described above to calculate the following equation (11), to generate the final one of the priority information.
[0154]
[Number 11]

[0155]
 In the equation (11), priority (object_class) shows the priority information determined based on the content information, for example, the priority information determined by the above-mentioned equation (8). Also, priority (others) shows information other than the content information, for example, object position information and gain information, spread information, the priority information determined based on the audio signal or the like of an object.
[0156]
 Further, although A and B in formula (11) is a power of the value of the non-linear sum, these A and B, it can be said that indicates the weighting coefficients used to generate the priority information.
[0157]
 For example, by setting method 2, if such weighting factor A = 2.0 and the weighting coefficient B = 1.0, when the sound of the object is a voice, the final value of the priority information priority becomes sufficiently large, non-speech object never priority information is smaller than. On the other hand, the magnitude relation of the priority information of the two objects is a sound, the second term is a priority (others) of formula (11) B will be determined by the value of.
[0158]
 As described above, it is possible to obtain different determined by several methods, by coupling a plurality of priority information by linear combination or non-linear combination, a more appropriate priority information with each other. The invention is not limited thereto, it may be to produce a final one of the priority information according to the conditional expression of a plurality of priority information.
[0159]
(4) smoothing in the time direction of the priority information
 In the above, etc. and generate a priority information from the metadata and content information of the object, the final one of the priority by combining a plurality of priority information It has been described for generating the time information. However, it is not desirable to magnitude of the priority information of a plurality of object changes many times during a short period.
[0160]
 In example decoding side, in case of switching the presence or absence of decoding processing for each object on the basis of the priority information, no longer hear or hear object sound for each short time by a change in the magnitude of the priority information of a plurality of objects It will be or. When this occurs, degradation of the auditory occurs.
[0161]
 Change in magnitude relation of such priority information (switching) is larger the number of objects, also can occur as a method of generating the priority information is if more complex becomes higher.
[0162]
 Therefore, in the priority information generator 52, for example, if the following equation smoothing the priority information in the time direction by calculating an exponential average performed as shown in (12), switched magnitude relation of the priority information of the object in a short time it is possible to suppress that.
[0163]
[Number 12]

[0164]
 In the equation (12) i denotes an index indicating the current frame, i-1 denotes an index indicating a temporally previous frame of the current frame.
[0165]
 priority (i) indicates the priority information before smoothing obtained for the current frame, priority (i), for example due to any of the formulas of the above-mentioned formula (1) through (11) is the obtained priority information.
[0166]
 Further, priority_smooth (i) is the priority information after the smoothing of the current frame, that is, shows the final priority information, priority_smooth (i-1) is after smoothing of the previous frame of the current frame shows the priority information. Furthermore shows the smoothing coefficient α is exponential averaging in equation (12), the smoothing coefficient α is a value between 0 and 1.
[0167]
 From smoothing factor alpha is multiplied priority information priority (i), a value obtained by subtracting (1-alpha) is multiplied by the priority information priority_smooth (i-1), the final priority information smoothing of the priority information by the priority_smooth (i) is performed.
[0168]
 In other words, by performing the time direction smoothing relative priority information priority of the current frame generated (i), the final priority information priority_smooth the current frame (i) is generated.
[0169]
 In this example, the smaller the value of smoothing factor alpha, the weight of the value of the priority information prior to smoothing of the current frame priority (i) becomes small, as a result, prioritized smoother is performed switching of the magnitude relationship between degree information is suppressed.
[0170]
 As an example of the smoothing of the priority information, has been described smoothing by exponential average, not limited to this, simple moving average or weighted moving average, such as smoothing using a low pass filter, what other it may be smoothing the priority information by the smoothing technique.
[0171]
 According to the technology described in the above, because it generates the priority information of the object based on the metadata or the like, it is possible to reduce the applied cost of the priority information of the object by hand. Further, even in the encoded data that have not been properly applied for all the time priority information of an object (frame), it is possible to impart proper priority information, as a result, reduce the calculation amount of the decoding it can be.
[0172]

 Next, description will be given of a processing performed by the encoding device 11.
[0173]
 Encoder 11 is reproduced simultaneously, the audio signal and the plurality of audio signals of each object of the plurality of respective channels are supplied by one frame, performs coding processing, the encoded audio signal outputs a bit stream contained.
[0174]
 Hereinafter, with reference to the flowchart of FIG. 3 will be described encoding process by the encoding device 11. Incidentally, this coding process is performed for each frame of the audio signal.
[0175]
 In step S11, the priority information generator 52 of object audio encoding unit 22 generates a priority information of the audio signals of each object supplied, and supplies to the packing unit 24.
[0176]
 For example meta-data input unit 23 or receives the user input operation, or communicate with the outside by or performs reading from outside of the recording area, acquires the metadata and content information of each object, priority degree information and supplies to the generating unit 52 and packing unit 24.
[0177]
 Priority information generating unit 52, for each object, based on the supplied audio signal, to at least one of the metadata the metadata supplied from the input unit 23, and the content information supplied from the metadata input unit 23 to generate a priority information of the object Te.
[0178]
 The method Specifically, the priority information generator 52, to generate the one or the priority information based on the audio signal and gain information of the object of the above-mentioned formula (1) through (9), the formula ( 10) and equation (11), and generates the priority information of each object due equation (12).
[0179]
 In step S12, the packing unit 24 stores the priority information of the audio signals of each object supplied from the priority information generator 52 to the DSE bitstream.
[0180]
 In step S13, the packing unit 24 stores the metadata and content information of each object supplied from the metadata input unit 23 to the DSE bitstream. By the above processing, the DSE of the bit stream, so that the priority information of the audio signals of all the objects, and metadata and content information of all objects are stored.
[0181]
 In step S14, channel audio encoder 21 encodes the audio signal of each channel supplied.
[0182]
 More specifically, the channel audio encoder 21 performs MDCT on the audio signals of the channels, the MDCT coefficients of each channel obtained by the MDCT encoding, the encoding of each channel obtained as a result data to be supplied to the packing unit 24.
[0183]
 In step S15, the packing unit 24 the coded data of the audio signals of each channel supplied from channel audio encoder 21, and stores the SCE or CPE bitstream. That is, the encoded data is stored in each element being arranged subsequent to the DSE in the bitstream.
[0184]
 In step S16, the encoding unit 51 of the object audio encoding unit 22 encodes the audio signals of each object supplied.
[0185]
 More specifically, the MDCT section 61 performs MDCT on the audio signals of each object, the encoding unit 51, the MDCT coefficients for each object obtained by the MDCT encoding, resulting in each object supplying the encoded data to the packing unit 24.
[0186]
 In step S17, the packing unit 24 the coded data of the audio signals of each object supplied from the encoding unit 51, and stores the SCE bitstream. That is, the encoded data is stored in a number of elements arranged later than DSE in the bitstream.
[0187]
 With the above processing, the frame to be processed, encoded data of the audio signals of all channels, the priority information and the encoded data of the audio signals of all the objects, and all objects metadata and content information is stored bitstream is obtained.
[0188]
 In step S18, the packing unit 24 outputs the obtained bit stream, the encoding process ends.
[0189]
 Encoding device 11 as described above may be stored in the bit stream to generate a priority information of the audio signals of each object, and outputs. Accordingly, the decoding side, which audio signal is to be able to easily understand whether those higher priority level.
[0190]
 Accordingly, the decoding side, in accordance with the priority information, the decoding of the encoded audio signal can be selectively performed. As a result, while minimizing the degradation of sound quality of the sound reproduced by the audio signal, it is possible to reduce the computational complexity of the decoding.
[0191]
 In particular, the priority information of the audio signals of each object by storing the bit stream, the decoding side can not only reduce the calculation amount of the decoding, making it possible to also reduce the amount of calculation processing such as subsequent rendering it can.
[0192]
 Also, in the encoding apparatus 11, and the metadata of the object and generating the priority information of the object based, such as the content information, the audio signal of the object, it is possible to obtain a more appropriate priority information at a low cost .
[0193]

 In the above, an example has been described in which the bit stream outputted from the encoding device 11 contains priority information, the encoding device may also be a bit stream does not contain priority information.
[0194]
 Therefore, it is also possible to generate the priority information in the decoding device. In such a case, an input bit stream outputted from the encoding device, decoding device for decoding encoded data included in the bit stream is configured as shown in FIG. 4, for example.
[0195]
 Decoding device 101 shown in FIG. 4 includes unpacking / decrypting unit 111, the rendering unit 112 and the mixing portion 113,.
[0196]
 Unpacking / decrypting unit 111 obtains the bit stream outputted from the encoding device performs unpacking and decoding the bit stream.
[0197]
 Unpacking / decrypting section 111 supplies the audio signals of each object obtained by the unpacking and decoding, and metadata of each object to the rendering unit 112. At this time, the unpacking / decrypting unit 111 generates a priority information for each object based on the metadata and content information of the object, performs the decoding of the encoded data of each object according to the obtained priority information .
[0198]
 Further, unpacking / decrypting section 111 supplies the audio signal of each channel obtained by the unpacking and decoding the mixing unit 113.
[0199]
 The rendering unit 112 generates an audio signal of M channels based on the object position information included audio signals of each object supplied from the unpacking / decrypting unit 111, and the metadata for each object, supplied to the mixing unit 113 to. Rendering unit 112 at this time, the sound image of each object, and generates an audio signal of each of the M channels to be localized at the position indicated by the object position information of the objects.
[0200]
 Mixing unit 113, an audio signal of each channel supplied from the unpacking / decrypting unit 111, an audio signal of each channel supplied from the rendering unit 112 and weighted addition for each channel, the final audio of each channel to generate a signal. Mixing unit 113, a final audio signal of each channel obtained in this manner is supplied to a speaker corresponding to each channel of the external, to reproduce the sound.
[0201]

 The unpacking / decrypting unit 111 of the decoding device 101 shown in FIG. 4, is configured as shown in FIG. 5, for example more.
[0202]
 Unpacking / decrypting unit 111 shown in FIG. 5, channel audio signal acquiring unit 141, channel audio signal decoding unit 142, IMDCT (Inverse Modified Discrete Cosine Transform) unit 143, object audio signal acquiring unit 144, object audio signal decoding unit 145 has a priority information generator 146, the output selection unit 147,0 value output unit 148 and the IMDCT unit 149,.
[0203]
 Channel audio signal acquiring unit 141, from the supplied bit storms obtains encoded data of each channel, and supplies the channel audio signal decoding unit 142.
[0204]
 Channel audio signal decoding unit 142 decodes the encoded data of each channel supplied from channel audio signal acquiring unit 141, and supplies the resulting MDCT coefficients IMDCT unit 143.
[0205]
 IMDCT unit 143 generates an audio signal by performing IMDCT on the basis of the MDCT coefficient supplied from the channel audio signal decoding unit 142, and supplies to the mixing unit 113.
[0206]
 The IMDCT unit 143, IMDCT (inverse modified discrete cosine transform) is performed on the MDCT coefficients, the audio signal is generated.
[0207]
 Object audio signal acquiring unit 144 acquires the coded data of each object from the supplied bit stream, and supplies the object audio signal decoding unit 145. Also, the object audio signal acquiring unit 144, from the supplied bit stream to obtain the meta-data and content information of each object, and supplies the metadata and content information to the priority information generating unit 146, renders the metadata supplied to the part 112.
[0208]
 Object audio signal decoding unit 145 decodes the encoded data of each object supplied from the object audio signal acquiring unit 144, and supplies the resulting MDCT coefficients to the output selection unit 147 and the priority information generating unit 146 .
[0209]
 Priority information generating unit 146, metadata supplied from the object audio signal acquiring unit 144, object audio signals supplied content information from the acquisition unit 144, and one at least of the supplied MDCT coefficient from the object audio signal decoding unit 145 generates priority information of each object based on whether, to the output selection unit 147.
[0210]
 Output selection unit 147, based on the priority information of each object supplied from the priority information generator 146, selectively switches the output destination of the MDCT coefficients for each object supplied from the object audio signal decoding unit 145.
[0211]
 That is, the output selection unit 147, when the priority information for a given object is less than a predetermined threshold Q, and supplies the MDCT coefficient of the object as 0 to 0 value output section 148. Further, the output selection unit 147, when the priority information for a given object is equal to or greater than a predetermined threshold Q, supplied from the object audio signal decoding unit 145 supplies the MDCT coefficient of the object to the IMDCT unit 149.
[0212]
 The value of the threshold Q is suitably determined according to the example of the decoding device 101 computing power and the like. By determining the threshold value Q appropriately, the calculation amount of the decoding of the audio signal, the decoding apparatus 101 can be reduced to calculate the amount in the range that can be decoded in real time.
[0213]
 0 value output unit 148 generates an audio signal based on the MDCT coefficient supplied from the output selection unit 147, and supplies the rendering unit 112. In this case, since the MDCT coefficient is 0, silent audio signal is generated.
[0214]
 IMDCT unit 149 generates an audio signal by performing IMDCT on the basis of the MDCT coefficient supplied from the output selection unit 147, and supplies the rendering unit 112.
[0215]

 Next, the operation of the decoding device 101.
[0216]
 Decoding device 101, one frame of the bit stream from the encoder is supplied to generate an audio signal by performing decoding processing, and outputs it to the speaker. Hereinafter, with reference to the flowchart of FIG. 6, described decoding processing performed by the decoding device 101.
[0217]
 In step S51, the unpacking / decrypting unit 111 obtains the bit stream transmitted from the encoding apparatus. That is, the bit stream is received.
[0218]
 In step S52, unpacking / decrypting unit 111 performs selection decoding process.
[0219]
 Although details of the selected decoding process will be described later, together with the coded data for each channel is decoded by the selective decoding, the priority information for each object is generated, based on the coded data priority information of the object It is selectively decoded.
[0220]
 Then, the audio signals of each channel are supplied to the mixing unit 113, the audio signals of each object are fed to the rendering unit 112. Furthermore, each object in the metadata acquired from the bit stream is supplied to the rendering unit 112.
[0221]
 In step S53, the rendering unit 112 performs rendering of the audio signal of the object based on object position information included audio signals supplied from the unpacking / decrypting section 111 objects, and the object metadata.
[0222]
 For example rendering unit 112, the VBAP (Vector Base Amplitude Pannning) based on the object position information, and generates an audio signal of each channel to be localized at a position where the sound image of the object is indicated by the object position information, the mixing unit 113 supplies. Incidentally, if it contains spread information in the metadata is performed even spread process on the basis of the spread information during rendering, the sound image of the object is widened.
[0223]
 In step S54, the mixing unit 113, weighted addition and the audio signal of each channel supplied from the unpacking / decrypting unit 111, an audio signal of each channel supplied from the rendering unit 112 for each channel, an external speaker supplied to. Thus, each speaker, since the audio signal of the channel corresponding to those of the speaker is supplied, each speaker reproduces sound based on the supplied audio signal.
[0224]
 When the audio signal of each channel is supplied to the speaker, the decoding process ends.
[0225]
 As described above, the decoding apparatus 101 generates priority information, decodes the encoded data of each object in accordance with the priority information.
[0226]

 Next, with reference to the flowchart of FIG. 7 will be described selection decoding process corresponding to the process in step S52 in FIG. 6.
[0227]
 In step S81, the channel audio signal acquiring unit 141 sets 0 to the channel number of the channel to be processed, held.
[0228]
 In step S82, the channel audio signal acquiring unit 141, the held channel number is equal to or less than the number of channels M.
[0229]
 In step S82, the case where the channel number is determined to be less than M, in step S83, the channel audio signal decoding unit 142 decodes the encoded data of the audio signal of the channel to be processed.
[0230]
 That is, the channel audio signal acquiring unit 141 supplies the supplied bit stream, a channel audio signal decoding unit 142 obtains encoded data of the channel to be processed. Then, channel audio signal decoding unit 142 decodes the encoded data supplied from the channel audio signal acquiring unit 141, and supplies the resulting MDCT coefficients IMDCT unit 143.
[0231]
 In step S84, IMDCT unit 143 performs IMDCT on the basis of the MDCT coefficient supplied from the channel audio signal decoding unit 142, generates an audio signal of a channel to be processed is supplied to the mixing unit 113.
[0232]
 In step S85, the channel audio signal acquiring unit 141 adds 1 to the channel number held, to update the channel number of the channel to be processed.
[0233]
 When the channel number is updated, then the process returns to step S82, the process described above is repeated. That is, the audio signal of a new processing target channel is generated.
[0234]
 Further, in step S82, the case where the channel number of the channel to be processed is determined not to be less than M, since the audio signal is obtained for all the channels, the processing proceeds to step S86.
[0235]
 In step S86, the object audio signal acquiring unit 144 sets 0 to the object number of the object to be processed, held.
[0236]
 In step S87, the object audio signal acquiring unit 144, object number holding is equal to or less than the number of objects N.
[0237]
 In step S87, if the object ID is determined to be less than N, in step S88, object audio signal decoding unit 145 decodes the encoded data of the audio signal to be processed object.
[0238]
 That is, the object audio signal acquiring unit 144 supplies the supplied bit stream, the object audio signal decoding unit 145 obtains encoded data of the processing target object. Then, object audio signal decoding unit 145 decodes the encoded data supplied from the object audio signal acquiring unit 144, and supplies the resulting MDCT coefficients in the priority information generator 146 and the output selection unit 147.
[0239]
 Also, the object audio signal acquiring unit 144 acquires the metadata and content information of an object to be processed from the supplied bit stream, and supplies the metadata and content information to the priority information generating unit 146, the metadata and supplies to the rendering unit 112.
[0240]
 In step S89, the priority information generating unit 146 generates the priority information of the audio signal of the processing target object, and supplies the output selection unit 147.
[0241]
 That is, the priority information generating unit 146, metadata supplied from the object audio signal acquiring unit 144, object audio signals supplied content information from the acquisition unit 144, and the object audio signal decoding unit 145 of the MDCT coefficients supplied from at least one of out to generate the priority information based on.
[0242]
 In step S89, the priority information is generated similarly to step S11 in FIG. 3 is performed. The method Specifically, the priority information generator 146, which generates the priority information based on either or sound pressure and gain information of an audio signal of an object of the above-mentioned formula (1) through (9) , formula (10) or equation (11), such as by equation (12) generates the priority information of the object. For example the generation of the priority information, if the sound pressure of the audio signal is used, priority information generating unit 146, the square sum of the MDCT coefficient supplied from the object audio signal decoding unit 145 as a sound pressure of the audio signal used.
[0243]
 In step S90, the output selection unit 147, the priority information of the priority information generating unit 146 object to be processed, which is supplied from and whether a threshold Q or designated by the controller or the like of the upper (not shown) the judges. In this case, the threshold value Q is for example determined according to the computing power or the like of the decoding device 101.
[0244]
 In step S90, if the priority information is determined to be the threshold value Q or more, the output selection unit 147 supplies supplied from object audio signal decoding unit 145, the MDCT coefficients of the object to be processed to the IMDCT unit 149 , the process proceeds to step S91. In this case, the decoding of the object to be processed, IMDCT is performed more.
[0245]
 In step S91, IMDCT unit 149 performs IMDCT on the basis of the MDCT coefficient supplied from the output selection unit 147, and generates an audio signal of the processing target object, and supplies the rendering unit 112. When the audio signal is generated, then the process proceeds to step S92.
[0246]
 In contrast, supply in step S90, if the priority information is determined to be less than the threshold value Q, the output selection unit 147, a 0 value output section 148 MDCT coefficients as zero.
[0247]
 0 value output unit 148, the MDCT coefficients is zero, which is supplied from the output selection unit 147, and generates an audio signal of the processing target object, and supplies the rendering unit 112. Therefore, the 0 value output section 148, in effect treatment for generating an audio signal such as the IMDCT nothing is done. In other words, IMDCT decoding the coded data, and more particularly for the MDCT coefficient is not substantially performed.
[0248]
 Incidentally, the audio signal generated by the 0 value output unit 148 is a silent signal. When the audio signal is generated, then the process proceeds to step S92.
[0249]
 Or priority information in step S90 is determined to be less than the threshold value Q, or an audio signal in step S91 is generated, in step S92, object audio signal acquiring unit 144, a 1 to the object number held in addition, to update the object number of the object to be processed.
[0250]
 If the object number is updated, then the process returns to step S87, the process described above is repeated. That is, the audio signal of a new object to be processed is generated.
[0251]
 Further, in step S87, if the object number of the processing target object is determined not to be less than N, selected decoding because the audio signal is obtained for all the channels and the necessary objects is terminated, then the process diagram proceed to the 6 step S53 of.
[0252]
 As described above, the decoding apparatus 101 generates the priority information for each object, while determining whether to perform decoding of the priority information and the threshold value and comparing the encoded audio signal, coding decoding the audio signal.
[0253]
 Thus, only the high audio signal priority degree in accordance with the reproduction environment can be selectively decoded, while minimizing the degradation of sound quality of the sound reproduced by the audio signal, thereby reducing the computational complexity of the decoding be able to.
[0254]
 Moreover, based on the priority information of the audio signals of each object, by performing the decoding of encoded audio signals, not only the calculation of the decoded audio signal, such as processing in the rendering unit 112 or the like, subsequent processing the amount of calculations can be reduced.
[0255]
 Also, and the object metadata, content information, by generating priority information of the object on the basis of such MDCT coefficients of an object, appropriate priority at a low cost even if it does not contain priority information in the bit stream it is possible to obtain the information. In particular, when generating the priority information at the decoding device 101, it is not necessary to store the priority information in the bit stream, thereby also reducing the bit rate of the bit stream.
[0256]

 The series of processes described above can be executed by hardware or can be executed by software. When executing the series of processing by software, a program constituting the software is installed into a computer. Here, the computer includes a computer incorporated in dedicated hardware, by installing various programs, which can execute various functions include, for example, such as a general-purpose personal computer.
[0257]
 Figure 8 is a block diagram showing a configuration example of hardware of a computer that executes the series of processes described above.
[0258]
 In the computer, CPU (Central Processing Unit) 501, ROM (Read Only Memory) 502, RAM (Random Access Memory) 503 are connected to each other via a bus 504.
[0259]
 The bus 504 is further output interface 505 is connected. Output interface 505, an input unit 506, output unit 507, recording unit 508, a communication unit 509, and a drive 510 are connected.
[0260]
 Input unit 506 includes a keyboard, a mouse, a microphone, made of an imaging device. The output unit 507 includes a display and a speaker. Recording unit 508, a hard disk and a nonvolatile memory. Communication unit 509 including a network interface. Drive 510 drives a magnetic disk, an optical disk, a magneto-optical disk, or a removable recording medium 511 such as a semiconductor memory.
[0261]
 Series In the computer configured as described above, CPU 501 is, for example, a program recorded in the recording unit 508 via the input-output interface 505 and the bus 504 and executes the loaded into RAM 503, the above-mentioned processing of is performed.
[0262]
 Program computer (CPU 501) is executed, for example, can be provided by being recorded in a removable recording medium 511 as a package medium or the like. Further, the program may be provided via a local area network, the Internet, or digital satellite broadcasting, a wired or wireless transmission medium.
[0263]
 In the computer, by mounting the removable recording medium 511 into the drive 510, it can be through the input-output interface 505, installed in the recording unit 508. The program via a wired or wireless transmission medium and received by the communication unit 509, can be installed in the recording unit 508. Alternatively, the program may be in the ROM502 and the recording unit 508 installed in advance.
[0264]
 The program which the computer executes may be a program in which processes are performed in time series in the order described herein, at a necessary timing such as when the parallel or call was made processing may be a program to be carried out.
[0265]
 Further, embodiments of the present technology is not limited to the embodiments described above, but various modifications are possible without departing from the scope of the present disclosure.
[0266]
 For example, the present technology, sharing one function by a plurality of devices via a network, it is possible to adopt a configuration of cloud computing which processes jointly.
[0267]
 Further, each step described in the above flowcharts may be executed by one device, it can be performed by allocating a plurality of apparatuses.
[0268]
 Further, when a plurality of processes are included in one step, the plurality of processes included in the one step may be executed by one device, it can be performed by allocating a plurality of apparatuses.
[0269]
 Additionally, the present technology may also be configured as follows.
[0270]
(1)
 based on a plurality of elements representing the characteristics of audio objects, comprises a priority information generator for generating priority information of the audio object
 signal processor.
(2)
 said element is a metadata of the audio object
 signal processing apparatus according to (1).
(3)
 said element is a position of the audio object in the space
 signal processing apparatus according to (1) or (2).
(4)
 the element is the distance from the reference position on the space to the audio object
 signal processing apparatus according to (3).
(5)
 the element is a horizontal angle indicating the horizontal position of the audio object on the spatial
 signal processing apparatus according to (3).
(6)
 the priority information generating unit generates the priority information corresponding to the moving speed of the audio object based on the metadata
 signal processing according to any one of (2) to (5) apparatus.
(7)
 the element is the gain information to be multiplied to the audio signals of the audio objects
 (1) to the signal processing apparatus according to any one of (6).
(8)
 the priority information generating unit, and the gain information in units of the processing target time, based on the difference between the average value of the gain information of the plurality of unit time, the priority of the unit of processing target time It generates information
 signal processing apparatus according to (7).
(9)
 the priority information generating unit generates the priority information based on the sound pressure of the audio signal in which the gain information is multiplied
 signal processing apparatus according to (7).
(10)
 said elements are spread information
 signal processing apparatus according to any one of (1) to (9).
(11)
 the priority information generating unit, on the basis of the spread information, and generates the priority information corresponding to the area of the region of the audio object
 signal processing apparatus according to (10).
(12)
 the element is information indicating an attribute of the sound of the audio object
 signal processing apparatus according to any one of (1) to (11).
(13)
 the element is an audio signal of the audio object
 signal processing apparatus according to any one of (1) to (12).
(14)
 The priority information generating unit generates the priority information based on the result of the speech section detection processing for the audio signal
 signal processing apparatus according to (13).
(15)
 the priority information generating unit performs temporal smoothing to the generated the priority information, a final said priority information
 (1) to to any one of (14) the signal processing apparatus according.
(16)
 based on a plurality of elements representing a feature of the audio objects, and generates the priority information of the audio object
 signal processing method comprising the steps.
(17)
 based on a plurality of elements representing a feature of the audio objects, and generates the priority information of the audio object
 program for executing the processing including a step in the computer.
DESCRIPTION OF SYMBOLS
[0271]
 11 encoder, 22-object audio encoding unit, 23 the metadata input section, 51 encoding section, 52 priority information generating unit, 101 decoding unit, 111 unpacking / decrypting unit, 144 object audio signal acquiring unit, 145 object audio signal decoding section, 146 priority information generating unit, 147 output selection unit

WE CLAIM

Based on the plurality of elements representing the characteristics of audio objects, it comprises a priority information generator for generating priority information of the audio object
 signal processor.
[Requested item 2]
 It said element is a metadata of the audio object
 signal processing apparatus according to claim 1.
[Requested item 3]
 It said element is the location of the audio object in the space
 signal processing apparatus according to claim 1.
[Requested item 4]
 It said element is the distance from the reference position on the space to the audio object
 signal processing apparatus according to claim 3.
[Requested item 5]
 It said element is a horizontal angle indicating the horizontal position of the audio object on the spatial
 signal processing apparatus according to claim 3.
[Requested item 6]
 The priority information generating unit generates the priority information corresponding to the moving speed of the audio object based on the metadata
 signal processing apparatus according to claim 2.
[Requested item 7]
 It said element is a gain information to be multiplied to the audio signals of the audio object
 signal processing apparatus according to claim 1.
[Requested item 8]
 The priority information generating unit generates said gain information of the unit of the processing target time, based on the difference between the average value of the gain information of the plurality of unit time, the priority information of the unit of processing target time to
 signal processing apparatus according to claim 7.
[Requested item 9]
 The priority information generating unit, the generating the priority information based on the sound pressure of the audio signal in which the gain information is multiplied
 signal processing apparatus according to claim 7.
[Requested item 10]
 It said element is a spread information
 signal processing apparatus according to claim 1.
[Requested item 11]
 The priority information generating unit on the basis of the spread information, and generates the priority information corresponding to the area of the region of the audio object
 signal processing apparatus according to claim 10.
[Requested item 12]
 It said element is information indicating an attribute of the sound of the audio object
 signal processing apparatus according to claim 1.
[Requested item 13]
 It said element is an audio signal of the audio object
 signal processing apparatus according to claim 1.
[Requested item 14]
 The priority information generating unit generates the priority information based on the result of the speech section detection processing for the audio signal
 signal processing apparatus according to claim 13.
[Requested item 15]
 The priority information generator performs time direction smoothing on the generated the priority information, a final said priority information
 signal processing apparatus according to claim 1.
[Requested item 16]
 Based on the plurality of elements representing a feature of the audio objects, and generates the priority information of the audio object
 signal processing method comprising the steps.
[Requested item 17]
 Based on the plurality of elements representing a feature of the audio objects, and generates the priority information of the audio object
 program for executing the processing including a step in the computer.

Documents

Application Documents

# Name Date
1 201917042320.pdf 2019-10-18
2 201917042320-TRANSLATIOIN OF PRIOIRTY DOCUMENTS ETC. [18-10-2019(online)].pdf 2019-10-18
3 201917042320-STATEMENT OF UNDERTAKING (FORM 3) [18-10-2019(online)].pdf 2019-10-18
4 201917042320-PROOF OF RIGHT [18-10-2019(online)].pdf 2019-10-18
5 201917042320-PRIORITY DOCUMENTS [18-10-2019(online)].pdf 2019-10-18
6 201917042320-POWER OF AUTHORITY [18-10-2019(online)].pdf 2019-10-18
7 201917042320-FORM 1 [18-10-2019(online)].pdf 2019-10-18
8 201917042320-DRAWINGS [18-10-2019(online)].pdf 2019-10-18
9 201917042320-DECLARATION OF INVENTORSHIP (FORM 5) [18-10-2019(online)].pdf 2019-10-18
10 201917042320-COMPLETE SPECIFICATION [18-10-2019(online)].pdf 2019-10-18
11 abstract.jpg 2019-10-19
12 201917042320-OTHERS-221019.pdf 2019-10-25
13 201917042320-Correspondence-221019.pdf 2019-10-25
14 201917042320-FORM 3 [27-03-2020(online)].pdf 2020-03-27
15 201917042320-FORM 18 [15-03-2021(online)].pdf 2021-03-15
16 201917042320-FER.pdf 2022-02-01
17 201917042320-OTHERS [01-08-2022(online)].pdf 2022-08-01
18 201917042320-FER_SER_REPLY [01-08-2022(online)].pdf 2022-08-01
19 201917042320-DRAWING [01-08-2022(online)].pdf 2022-08-01
20 201917042320-CORRESPONDENCE [01-08-2022(online)].pdf 2022-08-01
21 201917042320-COMPLETE SPECIFICATION [01-08-2022(online)].pdf 2022-08-01
22 201917042320-CLAIMS [01-08-2022(online)].pdf 2022-08-01
23 201917042320-ABSTRACT [01-08-2022(online)].pdf 2022-08-01
24 201917042320-US(14)-HearingNotice-(HearingDate-06-05-2024).pdf 2024-04-24
25 201917042320-Correspondence to notify the Controller [03-05-2024(online)].pdf 2024-05-03
26 201917042320-Written submissions and relevant documents [21-05-2024(online)].pdf 2024-05-21
27 201917042320-Annexure [21-05-2024(online)].pdf 2024-05-21
28 201917042320-PatentCertificate19-02-2025.pdf 2025-02-19
29 201917042320-IntimationOfGrant19-02-2025.pdf 2025-02-19

Search Strategy

1 201917042320E_19-01-2022.pdf

ERegister / Renewals