Sign In to Follow Application
View All Documents & Correspondence

Audio Signal Output Device And Method Encoding Device And Method Decoding Device And Method

Abstract: The present invention relates to an audio signal output device and method capable of more realistic audio reproduction and also to an encoding device and method a decoding device and method and a program. Given input of an audio signal that is generated to be outputted as sound by a virtual speaker i.e. a hypothetical speaker arranged in the ideal position the distance between the position of the virtual speaker and the position of the real reproduction speaker is calculated. Then the gain of the audio signal is adjusted on the basis of the gain corresponding to the calculated distance and the gain adjusted audio signal is reproduced by the reproduction speaker. By this means even if there is positional deviation between the virtual speaker and the reproduction speaker more realistic audio reproduction can be achieved. The present invention can be applied to a reproduction device.

Get Free WhatsApp Updates!
Notices, Deadlines & Correspondence

Patent Information

Application #
Filing Date
20 November 2015
Publication Number
09/2016
Publication Type
INA
Invention Field
COMMUNICATION
Status
Email
remfry-sagar@remfry.com
Parent Application
Patent Number
Legal Status
Grant Date
2023-08-28
Renewal Date

Applicants

SONY CORPORATION
1 7 1 Konan Minato ku Tokyo 1080075

Inventors

1. SHI Runyu
c/o SONY CORPORATION 1 7 1 Konan Minato ku Tokyo 1080075
2. CHINEN Toru
c/o SONY CORPORATION 1 7 1 Konan Minato ku Tokyo 1080075
3. YAMAMOTO Yuki
c/o SONY CORPORATION 1 7 1 Konan Minato ku Tokyo 1080075
4. HATANAKA Mitsuyuki
c/o SONY CORPORATION 1 7 1 Konan Minato ku Tokyo 1080075

Specification

AUDIO SIGNAL OUTPUT DEVICE AND METHOD, ENCODING DEVICE
AND METHOD, DECODING DEVICE AND METHOD, AND PROGRAM
TECHNICAL FIELD [0001]
The present technology relates to an audio signal output device and a method, an encoding device and a method, a decoding device and a method, and a program, and more particularly, to an audio signal output device and a method, an encoding device and a method, a decoding device and a method, and a program that are designed to be capable of audio reproduction with a more realistic feeling.
BACKGROUND ART
[0002]
In multichannel audio reproduction, the positions of the speakers on the reproducing side preferably correspond to the positions of the sound sources. In reality, however, the positions of the speakers on the reproducing side often differ from the positions of the sound sources.
[0003]
Where the positions of the speakers on the reproducing side differ from the positions of the sound sources, there is occurrence of a sound source that is not located in the speaker's position, therefore how to reproduce the sound of such sound sources is a critical issue.
[0004]

A technique called VBAP (Vector Base Amplitude Panning) has been suggested as a method of reproducing the sound of a sound source located in a desired position through a speaker located in a desired position (see Non-Patent Document 1, for example). [0005]
By VBAP, a target normal position of a sound image is expressed by a linear sum of vectors extending toward two or three speakers located around the normal position. The coefficients by which the respective vectors are multiplied in the linear sum are used as the gains of the audio signals to be output from the respective speakers, and gain adjustment is performed so that a sound image is fixed in the target position.
CITATION LIST NON-PATENT DOCUMENT [0006]
Non-Patent Document 1: Ville Pulkki, "Virtual Sound Source Positioning Using Vector Base Amplitude Panning", Journal of AES, vol.45, no.6, pp.456-466, 1997
SUMMARY OF THE INVENTION
PROBLEMS TO BE SOLVED BY THE INVENTION
[0007]
Meanwhile, a sound reproduction method has been suggested for a conventional situation where the number of channels and the speaker arrangement on the sound source side, and the number of channels of speakers and the speaker arrangement on the reproducing side are determined in advance, like 7.1 channel arrangement and

5.1 channel arrangement, 5.1 channel arrangement and 2.1 channel arrangement, or 22.2 channel arrangement and 5.1 channel arrangement, as recommended in several international standardization conferences. In such a case, sounds are output from the respective speakers with appropriate gains by virtue of a down-mixing process, and audio reproduction with a realistic feeling can be realized. [0008]
In the other cases such as a case where the sound sources or the speakers are arranged in positions that differ from predetermined positions, however, sound might not be reproduced by the suggested reproduction method, or the sound quality and the sound image definition might be severely degraded though reproduction can be performed by the suggested reproduction method. [0009]
In a case where channel-based sound sources are reproduced by the above-escribed VBAP, most sound images of the channel-based sound sources differ in position from the ideal speakers reproducing the sound sources. As a result, the sound image definition is severely degraded. [0010]
By the above-escribed technology, it is difficult to realize audio reproduction with a realistic feeling. [0011]
The present technology has been developed in view of those circumstances, and aims at realizing audio reproduction with a more realistic feeling.

SOLUTIONS TO PROBLEMS [0012]
An audio signal output device of a first aspect of the present technology includes: a distance calculating unit that calculates the distance between the position of an ideal speaker that reproduces an audio signal and the position of a real speaker that reproduces the audio signal; a gain calculating unit that calculates a reproduction gain of the audio signal based on the distance; and a gain adjusting unit that performs gain adjustment on the audio signal based on the reproduction gain. [0013]
The gain calculating unit can calculate the reproduction gain based on curve information for obtaining the reproduction gain corresponding to the distance. [0014]
The curve information can be information indicating a polyline curve or a function curve. [0015]
When the ideal speaker is not located on a unit circle having a predetermined reference point as the its center point, the gain adjusting unit can further perform gain adjustment on the audio signal with a gain determined based on the distance from the reference point to the ideal speaker and the radius of the unit circle. [0016]
The gain adjusting unit can delay the audio signal based on a delay time determined based on the distance from the reference point to the ideal speaker and the

radius of the unit circle. [0017]
When the real speaker is not located on a unit circle having a predetermined reference point as the its center point, the gain adjusting unit can further perform gain adjustment on the audio signal with a gain determined based on the distance from the reference point to the real speaker and the radius of the unit circle. [0018]
The gain adjusting unit can delay the audio signal based on a delay time determined based on the distance from the reference point to the real speaker and the radius of the unit circle. [0019]
The audio signal output device may further include a gain correcting unit that corrects the reproduction gain based on the distance between the position of an ideal center speaker and the position of the real speaker. [0020]
The audio signal output device may further include a lower limit correcting unit that corrects the reproduction gain when the reproduction gain is smaller than a predetermined lower limit. [0021]
The audio signal output device may further include a total gain correcting unit that calculates a ratio between the total power of an output sound based on the audio signal subjected to the gain adjustment with the reproduction gain and the total power of an input sound, and corrects the reproduction gain based on the ratio, the ratio being calculated based on the reproduction gain

and an expected value of the sound pressure of the input
sound based on the audio signal input.
[0022]
An audio signal output method or a program of the first aspect of the present technology includes the steps of: calculating the distance between the position of an ideal speaker that reproduces an audio signal and the position of a real speaker that reproduces the audio signal; calculating a reproduction gain of the audio signal based on the distance; and performing gain adjustment on the audio signal based on the reproduction gain. [0023]
In the first aspect of the present technology, the distance between the position of an ideal speaker that reproduces an audio signal and the position of a real speaker that reproduces the audio signal is calculated, a reproduction gain of the audio signal is calculated based on the distance, and gain adjustment is performed on the audio signal based on the reproduction gain. [0024]
An encoding device of a second aspect of the present technology includes: a correction information generating unit that generates correction information for correcting a gain of an audio signal in accordance with the distance between the position of an ideal speaker that reproduces the audio signal and the position of a real speaker that reproduces the audio signal; an encoding unit that encodes the audio signal; and an output unit that outputs a bit stream including the correction information and the encoded audio signal.

[0025]
An encoding method of the second aspect of the present technology includes the steps of: generating correction information for correcting a gain of an audio signal in accordance with the distance between the position of an ideal speaker that reproduces the audio signal and the position of a real speaker that reproduces the audio signal; encoding the audio signal; and outputting a bit stream including the correction information and the encoded audio signal. [0026]
In the second aspect of the present technology, correction information for correcting a gain of an audio signal in accordance with the distance between the position of an ideal speaker that reproduces the audio signal and the position of a real speaker that reproduces the audio signal is generated, the audio signal is generated, and a bit stream including the correction information and the encoded audio signal is output. [0027]
A decoding device of a third aspect of the present technology includes: an extracting unit that extracts, from a bit stream, correction information for correcting a gain of an audio signal in accordance with the distance between the position of an ideal speaker that reproduces the audio signal and the position of a real speaker that reproduces the audio signal, and the encoded audio signal; a decoding unit that decodes the encoded audio signal; and an output unit that outputs the decoded audio signal and the correction information. [0028]

The correction information can be the location information about the ideal speaker. [0029]
The correction information can be curve information for obtaining the gain corresponding to the distance. [0030]
The curve information can be information indicating a polyline curve or a function curve. [0031]
A decoding method of the third aspect of the present technology includes the steps of: extracting, from a bit stream, correction information for correcting a gain of an audio signal in accordance with the distance between the position of an ideal speaker that reproduces the audio signal and the position of a real speaker that reproduces the audio signal, and the encoded audio signal; decoding the encoded audio signal; and outputting the decoded audio signal and the correction information. [0032]
In the third aspect of the present technology, correction information for correcting a gain of an audio signal in accordance with the distance between the position of an ideal speaker that reproduces the audio signal and the position of a real speaker that reproduces the audio signal, and the encoded audio signal are extracted from a bit stream, the encoded audio signal is decoded, and the decoded audio signal and the correction information are output.
EFFECTS OF THE INVENTION [0033]

According to the first through third aspects of the present technology, audio reproduction with a more realistic feeling can be performed.
BRIEF DESCRIPTION OF DRAWINGS [0034]
Fig. 1 is a diagram for explaining the outline of the present technology.
Fig. 2 is a diagram for explaining a polyline curve. Fig. 3 is a diagram for explaining a function curve.
Fig. 4 is a diagram for explaining reproduction gains.
Fig. 5 is a diagram showing an example structure of a reproduction device.
Fig. 6 is a flowchart for explaining a down-mixing process.
Fig. 7 is a diagram showing an example configuration of an audio system.
Fig. 8 is a diagram for explaining metadata.
Fig. 9 is a flowchart for explaining an encoding process.
Fig. 10 is a flowchart for explaining a decoding process.
Fig. 11 is a diagram showing an example configuration of a computer.
MODES FOR CARRYING OUT THE INVENTION [0035]
The following is a description of embodiments to which the present technology is applied, with reference to the drawings.

[0036]

<0utline of the Present Technology>
The present technology relates to a reproduction method of reproducing the sound source of a channel with a desired number of speakers, and techniques for encoding and decoding the necessary information (metadata) for realizing the reproduction method. [0037]
First, the outline of the present technology is described. [0038]
Audio signals of channels and the metadata of these audio signals are supplied to a reproduction device, and the reproduction device controls sound reproduction based on the metadata and the audio signals, for example. [0039]
The audio signals of the respective channels are signals generated to be reproduced through speakers placed at ideal positions indicated by the metadata. In the description below, the virtual speakers that are placed at positions indicated by the metadata and reproduce the audio signals of the respective channels will be referred to as the ideal speakers. Also, the real speakers that output sounds based on audio signals output from the reproduction device will be referred to as the reproduction speakers. [0040]
In the present technology, audio signals of all the channels are classified into audio signals for LFE (Low Frequency Effect) and audio signals not for LFE. That is,

all the ideal speakers are classified into speakers for LF~E and speakers not for LFE. Likewise, the reproduction speakers are classified into speakers for LFE and speakers not for LFE. [0041]
First, reproduction of audio signals of channels not for LFE is described. [0042]
In reproducing audio signals of channels not for LFE, audio signal gain adjustment i£ performed based on the distances between an ideal speaker and reproduction speakers, as shown in Fig. 1f for example. [0043]
In Fig. 1, an ideal speaker VSPl and reproduction speakers RSP11-1 through RSP11-3 are disposed on the surface of a sphere PHll that has a radius ru and has its center at the position of a user Ull who is the viewer. The ideal speaker VSPl and the reproduction speakers RSPH-1 through RSPll-3 are speakers not for LFE. [0044]
Hereinafter, the reproduction speakers RSP11-1 through RSPll-3 will be also referred to simply as the reproduction speakers RSP11, if there is no particular need to distinguish them from one another. Although only one ideal speaker and three reproduction speakers are shown in this example, other ideal speakers and reproduction speakers exist in reality. [0045]
For example, a sound based on an audio signal of the channel corresponding to the ideal speaker VSPl ideally fixes a sound image at the position of the ideal

speaker VSP1. [0046]
Therefore, in the present technology, the reproduction gains of the respective reproduction speakers RSPll are determined in accordance with the distances between the ideal speaker VSPl and the reproduction speakers RSPll, and a sound based on an audio signal is output from each of the reproduction speakers RSPll with the determined reproduction gains, so that a sound image is fixed at the position of the ideal speaker VSPl. [0047]
Specifically, the distance between the ideal speaker VSPl and a reproduction speaker RSPll is the angle between a vector in the direction from the user Ull toward the ideal speaker VSPl and a vector in the direction from the user Ull toward the reproduction speaker RSPll. [0048]
In other words, the distance between the ideal speaker VSPl and a reproduction speaker RSPll on the surface of the sphere PHll, or the length of the arc connecting the two speakers, is the distance between the ideal speaker VSPl and the reproduction speaker RSPll. [0049]
In the example shown in Fig. 1, the angle between an arrow All and an arrow A12 is the distance DistMl between the ideal speaker VSPl and the reproduction speaker RSP11-1. Likewise, the angle between the arrow All and an arrow A13 is the distance DistM2 between the ideal speaker VSPl and the reproduction speaker RSPll-2,

and the angle between the arrow All and an arrow A14 is the distance DistM3 between the ideal speaker VSPl and the reproduction speaker RSP11-3. [0050]
An audio signal of the channel of the ideal speaker VSPl is subjected to gain adjustment based on the distance DistMl, and is reproduced by the reproduction speaker RSP11-1. The audio signal of the channel of the ideal speaker VSPl is also subjected to gain adjustment based on the distance DistM2 and the distance DistM3, and is reproduced by the reproduction speaker RSPll-2 and the reproduction speaker RSPll-3. [0051]
Accordingly, even in a case where there are differences in position between the ideal speaker VSPl and the reproduction speakers RSP11, differences caused in the sound image by the differences in position can be reduced, and audio reproduction with a more realistic feeling can be realized. [0052]
Next, reproduction of audio signals of channels not for LFE is described in greater detail. [0053]
Specifically, in the example described below, audio signals of M ideal speakers not for LFE, or of M channels, are down-mixed to generate audio signals of N channels, and the audio signals of the N channels are reproduced by N reproduction speakers not for LFE. [0054]
In the down-mixing process, the six processes STEl through STE6 shown below are mainly performed in

sequential order. [0055]
Process STE1: The distances between the ideal speakers and the reproduction speakers are determined.
Process STE2: The reproduction gains of the respective reproduction speakers are determined for each ideal speaker based on the determined distances and a predetermined attenuation curve.
Process STE3: The reproduction gains are corrected in accordance with the position of a reproduction speaker.
Process STE4: The reproduction gains are corrected based on a lower limit.
Process STE5: The reproduction gains are corrected so that the energy of the total output sound approximates the energy of the total input sound.
Process STE6: The reproduction gains are applied to audio signals, and gain adjustment is performed. [0056]
These processes STEl through STE6 are further described below. [0057]

First, in the process STEl, the distances between speakers are determined. The position of each speaker is represented by a horizontal angle G (-180° < G ^ + 180°), a vertical angle y {-90° < y < +90°), and a distance from the user to the speaker r (0 ^ r ^ +■*>) . [0058]
For example, Fig. 1 shows a three-dimensional coordinate system formed with the x-axis, the y-axis, and the z-axis, with the position of the user Ull being the

origin. [0059]
Where the x-y plane is the plane including a straight line extending in the depth direction of the drawing and a straight line extending in the transverse direction of the drawing, the angle between a straight line extending in the reference direction in the x-y plane, or the y~axis, and the vector in the direction from the user Ull toward the speaker is the horizontal angle 9, for example. That is, the horizontal angle G is an angle in the horizontal direction in Fig. 1. [0060]
Also, the angle between the vector in the direction from the user Ull toward the speaker and the x-y plane is the vertical angle y, and the length of the straight line connecting the user Ull and the speaker is the distance r. [0061]
The horizontal angles 9, the vertical angles y, and the distances r, which indicate the positions of the respective ideal speakers, are supplied as the metadata of audio signals to the reproduction device. The horizontal angles 8, the vertical angles y, and the distances r, which indicate the positions of the respective reproduction speakers, are also supplied to the reproduction device. [0062]
In the description below, the horizontal angle 9, the vertical angle y, and the distance r of the mth ideal speaker among the M ideal speakers will be represented by Uim/ Yimr and rim, respectively. Likewise, the horizontal angle G, the vertical angle y, and the distance r of the

nth reproduction speaker among the N reproduction speakers will be represented by 9Qn/ Yon/ arid ron, respectively. [0063]
The reproduction device calculates the distances between each of M ideal speakers and the N reproduction speakers. [0064]
For example, the distance Dist(m, n) between the mth ideal speaker and the nth reproduction speaker is calculated according to the equation (1) shown below. [0065]
[Mathematical Formula 1]
Dist(m, n) = arccos[ COSBim X COSU0n X COS (Yim - Yon) + sinGim x sin8on] ... (1) [0066]
The reproduction device performs calculation according to the equation (1) for each of the combinations of the M ideal speakers .and the N reproduction speakers, and calculates a total of (M x N) distances Dist{m, n). [0067]
If the respective ideal speakers and the respective reproduction speakers are located on a unit circle having the radius ru or on the sphere PH11 shown in Fig. 1, sounds output from the respective speakers reach the user Ull at the same time. If one of the speakers is not located on the sphere PHll, however, the sound from the speaker reaches the user Ull earlier or later than the sounds from the other speakers, and furthermore, a change is caused in the sound pressure of the sound to be heard

17

SP353112WO00

by the user. [0068]
Therefore, the reproduction device performs sound pressure correction using a correction value 5 SoundPressureCorrectionim on the audio signal of the
ideal speaker having a distance rim not equal to ru, and performs a delay process using a delay time Delayim. [0069]
In this manner, the ideal speaker can be regarded 10 as being located on the sphere PH11. [0070]
Specifically, calculation according to the equation (2) shown below is performed based on the distance rim and the radius ru, so that the correction value 15 SoundPressureCorrectionim is obtained. [0071] [Mathematical Formula 2]
n
III!
(dB)
SoundPressureCorrectionjm=-10x logio
■ ■ ■ (2)
[0072]
20 The correction value SoundPressureCorrectionira
determined according to the equation (2) is used in the correction to be performed on the audio signal of the ideal speaker side or on the audio signal of the channel m that is input to the reproduction device. In the
25 description below, an audio signal that is input to the reproduction device will be also referred to as an input audio signal, and an audio signal that is output from the reproduction device will be also referred to as an output audio signal.

18
SP353112WO00

[0073]
The delay time Delays for the delay process to be performed on the input audio signal of the ideal speaker is calculated according to the equation (3) shown below 5 based on the distance rim and the radius ru. If rim > ru, the delay time Delays has a negative value, and, in the delay process, the audio signal is delayed in the negative direction, or the audio signal is shifted backward in terms of time.
10 [0074]
[Mathematical Formula 3]
Delayim = (ru - rim) x sound speed (s) ... (3)
[0075]
The correction value SoundPressureCorrectionim and
15 the delay time Delays are calculated for each ideal
speaker having a distance rira not equal to ru. Likewise, the correction value SoundPressureCorrectionon arid the delay time Delayon are also calculated for each reproduction speaker having a distance ron not equal to ru.
20 [0076]
Specifically, the correction value SoundPressureCorrectionon is calculated according to the equation (4) shown below, and the delay time Delayon is calculated according to the equation (5) shown below.
25 [0077]
[Mathematical Formula 4]
on
(dB)
$oundPressureCorrectionon=-~10x logjo
... (4)
[0078]
[Mathematical Formula 5]

19
SP353112WO00
Delayon = (ru - ron) * sound speed (s) . . . (5) [0079]
The correction value SoundPressureCorrectionon and the delay time Delayon calculated in the above manner are 5 the sound pressure correction value and the delay time for the reproduction speaker side or an output audio signal. Therefore, the reproduction device performs sound pressure correction using the correction value SoundPressureCorrectionon on the audio signal supplied to 10 a reproduction speaker having a distance ron not equal to ru, and performs a delay process using the delay time
Delayon-[0080]

15 In the process STE2, the reproduction gains of the respective reproduction speakers are calculated with respect to each ideal speaker. [0081]
First, for each of the M ideal speakers, a check is
20 made to determine whether there is a reproduction speaker at a distance Dist(m, n) of "0" from the ideal speaker. The respective ideal speakers are then classified into speakers located in reproduction speaker positions and speakers not located in reproduction speaker positions.
25 [0082]
For the mth ideal speaker determined to be a speaker located in a reproduction speaker position, the reproduction gain MixGain(m, n) of the nth reproduction speaker with respect to the audio signal of the channel m
30 corresponding to the mth ideal speaker is calculated according to the equation (6) shown below.

20
SP353112WO00
[0083]
[Mathematical Formula 6]
... n . , v [OdB, Dist(m.n)=0 MixGain(m,n) = ^ ,D n. ., . .
L-oodB, Dist(m, n)>0 ■ ■ ■ (6)
[0084] 5 According to the equation (6) , the reproduction gain MixGain(m, n) of a reproduction speaker at a distance Dist(m, n) of "0" or a reproduction speaker located in the same position as the mth ideal speaker is 0 dB. Also, the reproduction gain MixGain(m, n) of a
10 reproduction speaker at a dxstance Dist{m, n) that is not "0" or a reproduction speaker located in a different position from that of the mth ideal speaker is -°° dB. [0085]
Accordingly, the audio signal of the channel m
15 corresponding to the mth ideal speaker is reproduced by the reproduction speaker located in the same position as the ideal speaker. That is, any sound component of the channel m is not output from the other reproduction speakers.
20 [0086]
For the mth ideal speaker determined to be a speaker not located in a reproduction speaker position, on the other hand, the reproduction gain MixGain{m, n) of each reproduction speaker with respect to the ideal
25 speaker is calculated with the use of an attenuation curve that is a polyline curve or a function curve. [0087]
Specifically, the metadata to be supplied to the reproduction device includes curve information indicating
30 which one of a polyline curve and a function curve is to

21
SP353112WO00
be used in calculating a reproduction gain, and the reproduction device calculates a reproduction gain using the curve of the type indicated by the curve information included in the metadata. 5 [0088]
The metadata also includes a curve index specifically indicating which one of the curves indicated in the curve information is to be used. The curve index might be information indicating a new curve that is not
10 recorded in the reproduction device. [0089]
In a case where the curve index is information indicating a predetermined curve, the reproduction device calculates a reproduction gain, using information that is
15 recorded in advance and is designed for obtaining a curve such as coefficients. In a case where the curve index is information indicating a new curve, on the other hand, the reproduction device reads information for obtaining a new curve from the metadata, and calculates a
20 reproduction gain, using the curve obtained from the information. [0090]
For example, the polyline curve to be used in calculating a reproduction gain is expressed as a
25 numerical sequence formed with the values of the
reproduction gains corresponding to the respective
distances Dist(m, n).
[0091]
Specifically, as the numerical sequence formed with
30 the values of reproduction gains, [0, -1.5, -4.5, -6, -9, -10.5, -12, -13.5, -15, -15, -16.5, -16.5, -18, -18, -18,

22
SP353112WO00
-19.5, -19.5, -21, -21, -21, ~~, -~, —, -~, —, -~] (dB) is the information for obtaining a reproduction gain. [0092]
In such a case, the value at the start of the 5 numerical sequence is the reproduction gain at the time when the distance Dist(m, n) is 0 degrees, and the value at the end of the numerical sequence is the reproduction gain at the time when the distance Dist(m, n) is 180 degrees. Also, the value at the kth point in the 10 numerical sequence is the reproduction gain at the time when the distance Dist(m, n) is as expressed by the equation (7) shown below. [0093] [Mathematical Formula 7]
180°
15 Dist{m,n)-(k-l)x ... (7)
Length of numerical sequence - 1
[0094]
Between adjacent points in the numerical sequence,
the reproduction gain linearly varies depending on the
distance Dist(m, n). The polyline curve obtained with 20 such a numerical sequence is the curve representing the
mapping of the reproduction gain MixGain(m, n) and the
distance Dist(m, n).
[0095]
For example, the polyline curve shown in Fig. 2 is 25 obtained from the above-escribed numerical sequence.
[0096]
In Fig. 2, the ordinate axis indicates the value of
the reproduction gain, and the abscissa axis indicates
the distance between an ideal speaker and a reproduction 30 speaker. Also, a polyline CVll represents the polyline

ill

23
SP353112WO00
curve, and each square on the polyline curve represents a numerical value of the numerical sequence formed with the values of the reproduction gain. [0097] 5 In this example, when the distance Dist(m, n)
between the nth reproduction speaker and the mth ideal speaker is DistMl, the reproduction gain MixGain(m, n) of the nth reproduction speaker is -3.5 dB, which is the value of the gain at DistMl on the polyline curve.
10 [0098]
Also, the reproduction gain MixGain(m, n) of the reproduction speaker at a distance Dist(m, n) of DistM2 is -8 dB, which is the value of the gain at DistM2 on the polyline curve, and the reproduction gain MixGain(m, n)
15 of the reproduction speaker at a distance Dist(m, n) of DistM3 is -16.5 dB, which is the value of the gain at DistM3 on the polyline curve. [0099]
Meanwhile, the function curve to be used in
20 calculating a reproduction gain is expressed with three coefficients coef1, coef2, and coef3, and a gain value MinGain, which is a predetermined lower limit. [0100]
In this case, the reproduction device performs
25 calculation according to the equation (9) shown below, using the function f(Dist{m, n)) shown in the equation (8) expressed with the coefficients coefl through coef3, the gain value MinGain, and the distance Dist(m, n). By doing so, the reproduction device calculates the
30 reproduction gain MixGain(m, n) of each reproduction speaker with respect to the mth ideal speaker.

24

SP353112WO00


[0101]
[Mathematical Formula 8]
f (Dist0n.n))=MinGainx(Coef1 x

Pi st (m. n) 180°

+Coef2x|»l»tMY+Coef3X^irt0n.n)
180c
180'

(8)

f(Dist(m,n)), -oodB,
[0102]
[Mathematical Formula 9]
OdB, M i xGa i n (m, n)

f(Dist(m, n))>0dB
otherwise
Dist(m,n)>Cut_thre

(9)
[0103]
In the equation (9), Cutthre represents the smallest value that satisfies the equation (10) shov/n
10 below. [0104]
[Mathematical Formula 10]
f(Cut_thre) = MinGain = -21dB, f ". 5 The dashed line in the drawing represents the values of the original function f(Dist(m, n)) at the respective distances Dist(m, n). [0107]
In this example, when the distance Dist(m, n)
10 between the nth reproduction speaker and the mth ideal
speaker is DistMl, the reproduction gain MixGain(m, n) of the nth reproduction speaker is -6 dB, which is the value of the gain at DistMl on the function curve. [0108]
15 Also, the reproduction gain MixGain(m, n) of the
reproduction speaker at the distance Dist(m, n) of DistM2 is -12 dB, which is the value of the gain at DistM2 on the function curve, and the reproduction gain MixGain(m, n) of the reproduction speaker at the distance Dist(m, n)
20 of DistM3 is -18 dB, which is the value of the gain at DistM3 on the function curve. [0109]
In a case where the reproduction gain MixGain(m, n) is calculated from the function curve, the combination
25 [coefl, coef2, coef3] of the coefficients coefl through coef3 may be [8, -12, 6], [1, -3, 3], or [2, -5.3, 4.2], for example. [0110]
Through the above process, the reproduction gains
30 MixGain(m, n) of the N reproduction speakers are obtained for each of the M ideal speakers. The values of the

26
SP353112WO00
reproduction gains of these reproduction speakers are greater where the distance Dist(m, n) to the ideal speaker is shorter. The same applies to the volumes of sounds from these reproduction speakers. Where M > N, 5 the reproduction gains MixGain{m, n) are mix gains. [0111]

Further, in the process STE3, the (M x N} reproduction gains MixGain(m, n) obtained in the process
10 STE2 are corrected in accordance with the position of the nth reproduction speaker. [0112]
For example, if a sound from a sound source located in front of a user comes from behind the user, the user
15 will find it strange. If a sound from a sound source
located behind the user comes from ahead of the user, the
user will not find it very strange.
[0113]
Therefore, the reproduction gains of the respective
20 reproduction speakers are corrected in accordance v/ith the positions of the N reproduction speakers located in front of or behind the user, so that the output sounds will not cause a feeling of strangeness depending on the positions of the reproduction speakers. That is, in a
25 case where an audio signal of an ideal speaker is
reproduced by two reproduction speakers that are at the same distance Dist(m, n) from the ideal speaker and are located in front of the user and behind the user,-correction is performed so that the reproduction gain of
30 the reproduction speaker behind the user becomes smaller than the reproduction gain of the reproduction speaker in

27
SP353112WO00
front of the user. [0114]
Specifically, the reproduction device first obtains.
information indicating whether it is necessary to correct 5 reproduction gains in accordance with the positions of
reproduction speakers from the metadata. If the obtained
information indicates that there is no need to correct
reproduction gains, the process STE3 is not carried out.
That is, after the process STE2, the process STE3 is 10 skipped, and the process STE4 is carried out.
[0115]
If the information obtained from the metadata
indicates that it is necessary to correct reproduction
gains, on the other hand, the reproduction device 15 performs the same calculation as the equation (1), and
determines the distances Dist(n, C) between a spatial
origin C and the N reproduction speakers.
[0116]
Here, the spatial origin C is the reference 20 positxon in the space in which the reproduction speakers
are placed, and the position of the spatial origin C is
expressed with a horizontal angle 9 of 0, a vertical
angle y of 0, and a distance r equal to ru, for example.
In this case, the spatial origin C is located on the unit 25 circle or on the sphere PH11 shown in Fig. 1, and is
located in front of the user Ull. The position of such a
spatial origin C is the position of an ideal center
speaker.
[0117] 30 After the distances Dist(n, C) from the spatial
origin C to the N reproduction speakers are determined,

28

SP353112WO00

the correction coefficient
spkr_pos_correction_coeffcient(n) of each of the N reproduction speakers is determined through calculation according to the equation (11) shown below. 5 [0118]
[Mathematical Formula 11]
spkr„pos„cor rect i on_coeffc i ent(n)
,., . ^ Dist(n,c) =Max_spkr_pos„cor rect i on_coef f c i ent x ——-—
loU
■ " " (11)
[0119]
In the equation (11) ,
10 Max_spkr_pos_correction^coeffcient represents the
correction coefficient at the time when the distance
Dist(n, C) is maximized (180 degrees).
[0120]
Further, the reproduction gain MixGain(m, n) of the
15 nth reproduction speaker with respect to the mth ideal speaker is multiplied by the obtained correction coefficient spkr_pos_correction_coeffcient(n) , so that a corrected reproduction gain MixGain_pos_corr(m, n) is obtained. That is, calculation is performed according to
20 the equation (12) shown below. [0121] [Mathematical Formula 12]

29
SP353112WO00
MixGain_pos_corr (m, n)
=MixGain(m, n) x spkr_pos_correction_coeffcient(n)
-MaxM ixGa in (n) x (Maxjspkr j3os_cor rect i orLCoef f c i ent-1)
xDist(n'c) ■■ - -(12)
180°
[0122]
In the equation (12), MaxMixGain(n) represents the largest value of M reproduction gains of the nth 5 reproduction speaker or the reproduction gains MixGain(m, n) having the same value as n. In the equation (12), the term including MaxMixGain (n) is the terra of reverse correction for preventing excess correction from being performed with spkr pos_correction_coeffcient(n). 10 [0123]
Through the above process, (M * N) reproduction gains MixGain_pos_corr(m, n) , which have been appropriately corrected in accordance with the positions of the reproduction speakers, are obtained. 15 [0124]
In a case where reproduction gain correction in accordance with the positions of the reproduction speakers is not performed, the reproduction gains MixGain(m, n) are used as the reproduction gains 20 MixGain_pos_corr(m, n). [0125]

In the process STE4 to be carried out after the process STE3, the reproduction gains are corrected so 25 that audio signals are reproduced by at least one

30
SP353112WO00
reproduction speaker with a predetermined lower limit of reproduction gain. Here , the audio signals are of an ideal speaker with which all the reproduction speakers have small reproduction gain values. 5 [0126]
Specifically, the largest value MaxMixGaini (m) of the reproduction gains of each ideal speaker obtained in the process STE3 or the N reproduction gains MixGain_pos_corr(m, n) having the same value as m is
10 determined, and the largest value MaxMixGaini (m) is compared with a lower limit MixGainHinThre> [0127]
If the largest value MaxMixGaini (m) with respect to the predetermined mth ideal speaker is smaller than the
15 lower limit MixGainr.iinThre, a correction value
MinGaincorrectioni (m) is added to the N reproduction gains MixGain_pos_corr(m, n) with respect to the mth ideal speaker. Here, the correction value MinGaincorrectioni (m) is the difference between the largest value MaxMixGaini (m)
20 and the lower limit MixGainMinThre/ as shown in the equation (13) shown below. [0128]
[Mathematical Formula 13] MinGaincorrectioni (m) = MaxMixGaini (m) - MixGainMinThre ...(13)
25 [0129]
Through this correction, the audio signal of the channel m is reproduced by at least one reproduction speaker with the predetermined smallest reproduction gain, and the sound from a certain channel can be prevented
30 from becoming inaudible. [0130]

31
SP353112WO00

In the process STE5, the reproduction gains MixGain_pos_corr(m, n) are corrected so that the energy of the total output sound approximates the energy of the 5 total input sound. [0131]
First, the reproduction device reads expected values SPR_i(m) of the relative sound pressures between the respective channels of the ideal speakers from the 10 metadata, and assumes the absolute sound pressure of the ideal speaker having the highest sound pressure to be 0 dBFS. The reproduction device then calculates the sound pressures of the sounds of the audio signals of the respective channels from the expected values SPR_i(m) of 15 the respective ideal speakers, and determines the power
value pow_i of the total sound of the input audio signals. [0132]
Here, the power value powi is the power of the total sound that is output from the ideal speakers as a 20 result of reproduction of the audio signals of the M
channels (the total sound output from the ideal speakers will be hereinafter also referred to as the input sound). Also, the sound that is output from the reproduction speakers as a result of reproduction of the audio signals 25 of the N channels will be hereinafter also referred to as the output sound. [0133]
The reproduction device then multiplies the reproduction gains MixGain_pos_corr(m, n) obtained in the 30 process STE4 by the expected values SPRi(m), to
determine the expected values SPR o(n) of the sound

32
SP353112WO00
pressures of the output sounds from the respective reproduction speakers. The reproduction device then determines the power value pow_o of the total output sound from the expected values SPR_o(n) 5 [0134]
The reproduction device then multiplies all the reproduction gains MixGain_pos__corr {m, n) obtained in the process STE4 by the power value ratio between the input sound and the output sound (pow_o/pow_i), to correct the
10 sound pressure of the total output sound. The
reproduction gains obtained in this manner are the ultimate reproduction gains of the reproduction speakers with respect to each ideal speaker. [0135]
15 In this example, the absolute sound pressure of the ideal speaker having the highest sound pressure is assumed to be 0 dB, and the power value ratio between the input sound and the output sound (pow o/pow i) is then determined. The determined power value ratio is the same
20 as the power value ratio between the input sound and the output sound (pow_o/pow_i) determined with the use of the actual absolute sound pressure. Even in a case where the absolute sound pressure of the actual input sound is unknown, if the absolute sound pressure of the input
25 sound is assumed in the above manner, the power value ratio between the input sound and the output sound (pow__o/pow_i) can be determined. The assumed sound pressure value may not be 0 dB but may be some other value, to obtain the same power value ratio as above.
30 [0136]


33
SP353112WO00
Reproduction of audio signals of channels for LFE is described. [0137]
For example, the number of ideal speakers for LFE 5 is zero, one, or two. Likewise, the number of
reproduction speakers for LFE is zero, one, or two. [0138]
In a case where the number of ideal speakers for LFE or the number of reproduction speakers for LFE is 10 zero, the audio signal of any channel for LFE cannot be reproduced, and the gain of the audio signal is -°°. [0139]
In a case where the number of ideal speakers for LFE and the number of reproduction speakers for LFE are 15 one or two, on the other hand, the reproduction device generates the audio signal of each channel for LFE with the reproduction gains shown in Fig. 4, for example. [0140]
That is, in a case where both the number of ideal 20 speakers for LFE and the number of reproduction speakers for LFE are one or two, the audio signal(s) of the ideal speaker(s) for LFE are reproduced as the audio signal(s) of the reproduction speaker(s) for LFE. [0141] 25 In a case where there are one ideal speaker for LFE and two reproduction speakers for LFE, or where there are two ideal speakers for LFE and one reproduction speaker for LFE, the audio signals of the respective channels are evenly distributed. 30 [0142]
That is, in a case where two reproduction speakers

34
SP353112WO00
for LFE are provided for one ideal speaker for LFE, the audio signal of the ideal speaker is subjected to gain adjustment with the same reproduction gain, and is reproduced by the two reproduction speakers. In a case 5 where one reproduction speaker for LFE is provided for two ideal speakers for LFE, the audio signals of the ideal speakers are combined into one audio signal with the same reproduction gain, and the audio signal is reproduced by the reproduction speaker. 10 [0143]

Next, a specific embodiment of the reproduction device described above is described. [0144] 15 The reproduction device has the structure shown in Fig. 5, for example. [0145]
The reproduction device 11 shown in Fig. 5 receives metadata and an audio signal from a decoder or the like 20 (not shown), performs gain adjustment on the audio signal based on the metadata, and supplies the resultant audio signal to speakers 12-1 through 12-N. [0146]
Fig. 5 shows only the functional blocks of the 25 reproduction device 11 for reproducing audio signals of channels not for LFE, and does not show the functional blocks for reproducing audio signals of channels for LFE. [0147]
In Fig. 5, audio signals of M channels are supplied 30 to the corresponding M ideal speakers not for LFE. The audio signals of the M channels are converted into audio

35
SP353112WO00
signals of N channels, and are then output- Further, the speakers 12-1 through 12-N correspond to the above described reproduction speakers not for LFE. [0148] 5 Hereinafter, when there is no particular need to distinguish the speakers 12-1 through 12-N from one another, the speakers 12-1 through 12-N will be also referred to simply as the speakers 12. The respective speakers 12 are also the speakers corresponding to the
10 above-escribed reproduction speakers RSPll, and therefore, the speakers 12 will be also referred to as the reproduction speakers 12. [0149]
The reproduction device 11 shown in Fig. 5 includes
15 a distance calculating unit 21, a reproduction gain
calculating unit 22, a correcting unit 23, a lower limit correcting unit 24, a total gain correcting unit 25, and a gain adjusting unit 26. The gain adjusting unit 26 includes an amplifier 31, an amplifier 32, and an
20 amplifier 33. [0150]
The location information about the respective ideal speakers not for LFE and the location information about the respective reproduction speakers 12, which are
25 included in the metadata, are supplied to the distance calculating unit 21. The distance calculating unit 21 calculates distances Dist(m, n) based on the location information about the ideal speaker and the location information about the reproduction speakers 12, and
30 supplies the distances Dist(m, n) to the reproduction gain calculating unit 22.


36

SP353112WO00

[0151]
Here, the location information about each speaker is information formed with a horizontal angle 0, a vertical angle y, and a distance r. 5 [0152]
The distance calculating unit 21 calculates correction values SoundPressureCorrectionim and delay times Delayim of the ideal speaker side, and supplies the correction values and the delay times to the amplifier 31,
10 as necessary. The distance calculating unit 21 also
calculates correction values SoundPressureCorrectionon and delay times Delayon of the side of the reproduction speakers 12, and supplies the correction values and the delay times to the amplifier 33. That is, the process
15 STEl is performed in the distance calculating unit 21. [0153]
The curve information and the curve index included in the metadata are supplied to the reproduction gain calculating unit 22. The reproduction gain calculating
20 unit 22 calculates reproduction gains MixGain{m, n) using the curve information and the curve index as well as the distances supplied from the distance calculating unit 21, and supplies the reproduction gains MixGainfm, n) to the correcting unit 23. That is, the process STE2 is
25 performed in the reproduction gain calculating unit 22. [0154]
The location information about the reproduction speakers 12, the information that is included in the metadata and indicates whether it is necessary to correct
30 the reproduction gains in accordance with the positions of the reproduction speakers 12, and the correction

37
SP353112WO00
coefficient Max_spkr_pos__correction_coef fcient are
supplied to the correcting unit 23.
[0155]
Based on the supplied information, the correcting 5 unit 23 corrects the reproduction gains supplied from the reproduction gain calculating unit 22 in accordance with the positions of the reproduction speakers 12, and supplies the resultant reproduction gains MixGain_pos_corr (m, n) to the lower limit correcting unit
10 24. That is, the process STE3 is performed in the correcting unit 23. [0156]
The reproduction gain lower limit MixGainMinThre included in the metadata is supplied to the lower limit
15 correcting unit 24. Based on the lower limit
MixGainMinThI:e/ the lower limit correcting unit 24 corrects the reproduction gains supplied from the correcting unit 23, and supplies the corrected reproduction gains to the total gain correcting unit 25. That is, the process STE4
20 is performed in the lower limit correcting unit 24. [0157]
The expected values SPR_i(m) that are included in the metadata and are of the relative sound pressures between the respective channels of the ideal speakers are
25 supplied to the total gain correcting unit 25. Based on the expected values SPR_i(m), the total gain correcting unit 25 corrects the reproduction gains supplied from the lower limit correcting unit 24, and supplies the resultant ultimate reproduction gains to the amplifier 32.
30 The process STE5 is performed in the total gain correcting unit 25.

38
SP353112WO00
[0158]
The gain adjusting unit 26 generates the audio signals of the N channels by performing gain adjustment on the audio signals of the M ideal speakers supplied 5 from the decoder (not shown), and supplies the audio signals of the respective channels to the reproduction speakers 12 for reproduction. The process STE6 is performed in the gain adjusting unit 26. [0159]
10 That is, based on the correction values and the
delay times supplied from the distance calculating unit 21, the amplifier 31 performs gain correction and a delay process on the supplied audio signals of the M channels as appropriate, and supplies the resultant audio signals
15 to the amplifier 32. [0160]
The amplifier 32 multiplies the audio signals of the M channels supplied from the amplifier 31 by the reproduction gains supplied from the total gain
20 correcting unit 25. The amplifier 32 also generates the audio signals of the N channels by adding the audio signals of the respective ideal speakers multiplied by the reproduction gains, and supplies the generated audio signals to the amplifier 33.
25 [0161]
Based on the correction values and the delay times supplied from the distance calculating unit 21, the amplifier 33 performs gain correction and a delay process on the audio signals of the N channels supplied from the
30 amplifier 32 as appropriate, and supplies the resultant audio signals to the reproduction speakers 12.

39
SP353112WOOO
[0162]
Next, the operation of the reproduction device 11 is described. 5 [0163]
When the audio signals and the metadata of the respective ideal speakers are supplied to the reproduction device 11, the reproduction device 11 generates the audio signals to be supplied to the 10 reproduction speakers with respect to audio signals for LFE and audio signals not for LFE, and then outputs the generated audio signals. [0164]
Referring to the flowchart in Fig. 6, the down-15 mixing process to be performed by the reproduction device 11 on the audio signals not for LFE is described below. [0165]
In step Sll, the distance calculating unit 21 determines the distances Dist(m, n) between the ideal 20 speakers and the reproduction speakers 12 based on the
location information about the ideal speakers not for LFE and the location information about the reproduction speakers 12 not for LFE, which are included in the metadata, and supplies the distances Dist(m, n) to the 25 reproduction gain calculating unit 22. Specifically, the calculation according to the equation (1) is performed for each of the combinations of the ideal speakers and the reproduction speakers 12, to determine (M x N) distances Dist (in, n) . 30 [0166]
In step S12, the distance calculating unit 21

40
SP353112WO00
determines the correction values and the delay times of the ideal speaker side and the side of the reproduction speakers 12, as necessary. [0167] 5 Specifically, for the ideal speakers each having a distance rim not equal to ru, the distance calculating unit 21 calculates the correction values SoundPressureCorrectionim and the delay times Delays by performing the calculation according to the equation (2)
10 and the equation (3) based on the distances rim serving
as the location information about the ideal speakers, and supplies the correction values and the delay times to the amplifier 31. [0168]
15 For the reproduction speakers each having a
distance ron not equal to ru, the distance calculating unit 21 also calculates the correction values SoundPressureCorrectionon and the delay times Delayon by performing the calculation according to the equation (4)
20 and the equation (5) based on the distances ron serving as the location information about the reproduction speakers 12, and supplies the correction values and the delay times to the amplifier 33. [0169]
25 In step S13, the reproduction gain calculating unit 22 determines the reproduction gains of the respective reproduction speakers 12 for each ideal speaker based on the distances Dist(m, n) supplied from the distance calculating unit 21.
30 [0170]
For example, for an ideal speaker having a

41
SP353112WO00
reproduction speaker 12 at a distance Dist(m, n) of "0" between the ideal speaker and the reproduction speaker 12, the reproduction gain calculating unit 22 performs the calculation according to the equation (6), to calculate 5 the reproduction gains MixGain(m, n) of the respective reproduction speakers 12 with respect to the ideal speaker. [0171]
For an ideal speaker having no reproduction
10 speakers 12 at the distance Dist(m, n) of "0", the
reproduction gain calculating unit 22 obtains the curve indicated by the curve information included in the metadata, which is a polyline curve or a function curve. In doing so, the reproduction gain calculating unit 22
15 refers to the curve index, and reads the polyline curve or the function curve from the metadata, as necessary. [0172]
Having obtained the polyline curve or the function curve, the reproduction gain calculating unit 22
20 determines the gain values corresponding to the distances Dist(m, n) based on the obtained curve, and sets the determined gain values as the reproduction gains MixGain(m, n) of the reproduction speaker 12 with respect to the ideal speaker. At this point, the calculation
25 according to the equation (7) and the equation (9) is performed, as necessary. [0173]
Having obtained the reproduction gains MixGain(m, n) of the respective reproduction speakers 12 for each
30 ideal speaker, the reproduction gain calculating unit 22 supplies the reproduction gains MixGain{m, n) to the

42
SP353112WO00
correcting unit 23. [0174]
In step S14, based on the information that is included in the metadata and indicates whether it is 5 necessary to correct the reproduction gains, the correcting unit 23 corrects the reproduction gains supplied from the reproduction gain calculating unit 22 in accordance with the positions of the reproduction speakers 12, as necessary, and supplies the corrected
10 reproduction gains to the lower limit correcting unit 24. [0175]
Specifically, the correcting unit 23 calculates the reproduction gains MixGain_pos_corr{m, n) by performing the calculation according to the equation (11) and the
15 equation (12) using the location information about the respective reproduction speakers 12 and the correction coefficient Max_spkr_pos_correction_coeffcient included in the metadata. [0176]
20 In step S15, based on the lower limit MixGainMinThre included in the metadata, the lower limit correcting unit 24 corrects the reproduction gains supplied from the correcting unit 23, as necessary, and supplies the corrected reproduction gains to the total gain correcting
25 unit 25. Specifically, the calculation according to the equation (13) is performed as necessary, and the correction value MinGaincorrectioni (m) is added to the reproduction gains MixGain_pos_corr(m, n). [0177]
30 In step S16, the total gain correcting unit 25
performs sound pressure correction on the total output

43
SP353112WO00
sound. [0178]
That is, the total gain correcting unit 25 calculates the power value ratio between the input sound 5 and the output sound (pow__o/pow_i) based on the expected values SPR_i(m) included in the metadata and the reproduction gains MixGain__pos__corr (m, n) supplied from the lower limit correcting unit 24. The total gain correcting unit 25 then multiplies the reproduction gains
10 MixGain_pos_corr(m, n) by the power value ratio
(pow_o/pow_i) to obtain the ultimate reproduction gains, and supplies the ultimate reproduction gains to the amplifier 32. [0179]
15 In step S17, the amplifier 31 performs audio signal gain adjustment based on the correction values and delay values of the ideal speaker side supplied from the distance calculating unit 21. [0180]
20 Specifically, as for the audio signal of a channel m for which a correction value and a delay value have been supplied, the amplifier 31 multiplies the audio signal by the correction value SoundPressureCorrectionim, delays the resultant audio signal by the delay time
25 Delays in the temporal direction, and supplies the delayed audio signal to the amplifier 32. [0181]
In step SI8, the amplifier 32 generates the audio signals of the respective reproduction speakers 12 based
30 on the reproduction gains supplied from the total gain correcting unit 25 and the audio signals supplied from

44
SP353112WO00
the amplifier 31, and supplies the generated audio
signals to the amplifier 33.
[0182]
Specifically, with one of the N channels 5 corresponding to the reproduction speakers 12 being an attention channel nc, the amplifier 32 multiplies the reproduction gains of the respective ideal speakers with respect to the attention channel nc by the audio signals of the respective ideal speakers. The amplifier 32 then
10 sets the one audio signal obtained by combining the audio signals of the respective ideal speakers multiplied by the reproduction gains, or the M audio signals, as the audio signal of the attention channel nc. The same process as above is performed on each of the N channels
15 as the attention channel, so that the audio signals of the M respective ideal speakers are converted into the audio signals of the N reproduction speakers 12. [0183]
In step SI9, the amplifier 33 performs gain
20 adjustment on the audio signals supplied from the
amplifier 32 based on the correction values and delay values of the side of the reproduction speakers 12 supplied from the distance calculating unit 21. [0184]
25 Specifically, as for the audio signal of a channel n for which a correction value and a delay value have been supplied, the amplifier 33 multiplies the audio signal by the correction value SoundPressureCorrectionon, delays the resultant audio signal by the delay time
30 Delayon in the temporal direction, and supplies the
delayed audio signal to the reproduction speakers 12.

45

SP353112WO00

[0185]
After the audio signals of the respective channels are output to the reproduction speakers 12, the down-mixing process comes to an end. Also, the reproduction 5 speakers 12 reproduce sounds based on the audio signals supplied from the reproduction device 11. [0186]
In the above-escribed manner, the reproduction device 11 performs gain adjustment (gain correction) on
10 audio signals in accordance with the distances between
the positions of the ideal speakers and the positions of the real reproduction speakers 12. Accordingly, even in a case where there are differences in position between the ideal speaker and the reproduction speakers 12,
15 degradation of the sound quality of output sounds and
degradation of the sound image definition can be reduced, and audio reproduction with a more realistic feeling can be realized. [0187]
20 Through the above-escribed process, the input audio signal(s) of one or more channels can be reproduced by one or more reproduction speakers placed in one or more desired position. Even in a case where the input audio signals of the respective channels are audio signals from
25 respective objects serving as sound sources, audio
reproduction in the correct sound image position can be performed through the same down-mixing process as above. [0188]

30 Next, the encoder that encodes the metadata to be supplied to the reproduction device 11, and the decoder

46
SP353112WO00
that decodes the encoded metadata are described.
[0189]
As shown in Fig. 7, for example, in an audio system
to which the present technology is applied, metadata is 5 supplied from an encoder 61 to a decoder 62, and the
metadata is further supplied from the decoder 62 to the
reproduction device 11.
[0190]
The encoder 61 obtains the necessary information 10 for obtaining the metadata from the outside and the audio
signals of the M ideal speakers, and generates a bit
stream formed with the metadata and the audio signals
that have been encoded.
[0191] 15 The encoder 61 includes a metadata generating unit
71, an audio signal encoding unit 72, and an output unit
73.
[0192]
The metadata generating unit 71 obtains the 20 necessary information from the outside, and generates
encoded metadata by encoding the obtained information as
necessary.
[0193]
The metadata includes the location information 25 about the respective ideal speakers, the number of ideal
speakers for LFE (the number of channels) among the ideal
speakers, the curve information, and the curve index, for
example. The metadata also includes the information
indicating whether it is necessary to correct 30 reproduction gains in accordance with the positions of
the reproduction speakers 12, the correction coefficient

47
SP353112WO00
Max_spkr_pos_correction_coeffcient depending on the positions of the reproduction speakers 12, the gain lower limit MixGainHinThre/ and the expected values SPR_i (m) of the relative sound pressures between the channels. 5 [0194] -. ■
The audio signal encoding unit 72 encodes audio signals supplied from the outside. The output unit 73 generates a bit stream containing the encoded metadata - and the encoded audio signals, and outputs the bit stream
10 to the decoder 62. [0195]
The decoder 62 includes an extracting unit 81, an audio signal decoding unit 82, and an output unit 83. The decoder 62 receives the bit stream transmitted from
15 the encoder 61, and the extracting unit 81 extracts the metadata and the audio signals from the received bit stream. At this point, the extracting unit 81 decodes the metadata, as necessary. [0196]
20 The audio signal decoding unit 82 decodes the audio signals extracted by the extracting unit 81. The output unit 83 supplies the metadata extracted by the extracting unit 81 and the audio signals decoded by the audio signal decoding unit 82 to the reproduction device 11.
25 [0197]
Part of the metadata written in a bit stream to be output from the encoder 61 to the decoder 62 is as shown in Fig. 8, for example. That is, Fig. 8 shows the syntax of part of the metadata.
30 [0198]
In the example shown in Fig. 8, at the start of the

48
SP353112WO00
header, "down mix coef exist flag" is placed as the information indicating whether the necessary information for down-mixing is included in the metadata. [0199] 5 Also, in the metadata, "down mix coef mode" is placed as the curve information, and, under the curve information, "polyline curve idx" or "function curve idx" is placed as the curve index. [0200]
10 The "polyline curve idx" indicates a polyline curve, and, if the value thereof is a binary number "111", the polyline curve is a new polyline curve. In this case, "polyline curve coeffcient[j]" is written as the information for obtaining a new polyline curve.
15 [0201]
The information for obtaining a new polyline curve is the information for identifying the respective squares on the polyline CV11 shown in Fig. 2 (these squares will be hereinafter referred to as description points), for
20 example, or for identifying the respective values constituting a numerical sequence. [0202]
Specifically, the reproduction gain axis (the ordinate axis) is divided into sixteen, so that sixteen
25 divided lines are defined. The respective description points are sequentially placed on the respective divisional lines along the ordinate axis. [0203]
In the metadata, the description points are
30 represented by "0"s, and the information indicating on
which divided lines the respective description points are

49
SP353112WO00
placed is represented by "l"s. [0204]
In Fig. 2, the description points are sequentially
written from left.' First, the information indicating on 5 which divided line counted from the bottom the first
description point from left is located is written with
the number "1", and thereafter, "0"s representing
description points are written. Here, the first
description point from left is located on the uppermost 10 divided line, only a "0" representing a description point
is written.
[0205]
Thereafter, the information indicating that the
description point is located Q divided lines below the 15 divided line on which the last description line is
located is written with Q "l"s, followed by a "0"
representing a description point.
[0206]
For example, the third description point from left 20 is located two divided lines below the second description
point. Therefore, two "l"s are written, followed by one
"0". Also, the tenth description point from left is
located on the same divided line as the ninth description
line, or is located zero divided lines below the ninth 25 description line. Therefore, no "l"s are written, and
only one "0" is written.
[0207]
The description is conducted by the above method.
If all the description points have been written, one ul" 30 is written to indicate that the information about the
polyline curve has been written.. If the number of

50
SP353112WO00
description points is large, and the description points cannot be written even with 64 "l"s and "0"s in total, the description is conducted until the number of "l"s and "0"s reaches 64, and the description is then ended. 5 [0208]
Therefore, in a case where the information for obtaining a polyline curve is read from the metadata, the information for sequentially obtaining the respective description points is read until 16 "l"s or 64 "l"s and
10 "0"s in total (the sum of the number of "l"s and the number of "0"s being 64) have been read out. In this manner, a polyline curve is generated. [0209]
The "function curve idx" indicates a function curve,
15 and, if the value thereof is a binary number "111", the function curve is a new function curve. In this case, "functioncurveco-effcient[i]" is written as the coefficient of a new function curve. [0210]
20 Meanwhile, "minimun_gain_threshold__idx" written in the metadata is the index indicating the gain low limit MixGainMinThre- Further, "gain_correction__coeff cient" written in the metadata is the correction coefficient Max__spkr_pos__correction_coeffcient required in correcting
25 reproduction gains in accordance with the positions of the reproduction speakers 12. If the value of Max_spkr_pos_correction_coeffcient is "1", there is no need to correct reproduction gains in accordance with the positions of the reproduction speakers 12.
30 [0211]
Further, in the metadata, "sound_level_exist_flag"

51
SP353112WO00
is written as the information indicating whether the expected values SPR_i(m) of the relative sound pressures between channels are written in the metadata, and "channel sound level[i]" is v/ritten in accordance with 5 the value of "sound_level_exist_flag". Here, "channel sound level[i]" represents the expected values SPR_i(m). [0212]
Explanation of the Encoding Process>
The operations of the encoder 61 and the decoder 62 10 are further described. [0213]
Referring first to the flowchart in Fig. 9, the encoding process to be performed by the encoder 61 is described. 15 [0214]
In step S41, the metadata generating unit 71 obtains the necessary information from the outside, and generates encoded metadata by encoding the obtained information. For example, the metadata generating unit 20 71 generates the metadata corresponding to the syntax shown in Fig. 8. [0215]
In step S42, the audio signal encoding unit 72 encodes audio signals supplied from the outside. 25 [0216]
In step S43, the output unit 73 generates a bit stream containing the encoded metadata and the encoded audio signals, and outputs the bit stream to the decoder 62. After the bit stream is output, the encoding process 30 comes to an end. [0217]

52
SP353112WO00
In the above manner, the encoder 61 generates and outputs the metadata including the location information about the ideal speakers, the curve information, and the like. As the information formed with the location 5 information about the ideal speakers, the curve
information, and the like is generated as the metadata, the reproduction device 11 can perform appropriate gain correction, such as gain correction in accordance with the distances between the positions of the ideal speakers
10 and the positions of the real reproduction speakers 12. As a result, audio reproduction with a more realistic feeling can be performed. [0218]

15 Referring now to the flowchart in Fig. 10, the decoding process to be performed by the decoder 62 is described. [0219]
In step S71, the decoder 62 receives a bit stream
20 transmitted from the encoder 61, and the extracting unit 81 extracts metadata and audio signals from the received bit stream. The extracting unit 81 also decodes the metadata. [0220]
25 In step S72, the audio signal decoding unit 82 decodes the audio signals extracted by the extracting unit 81. [0221]
In step S73, the output unit 83 outputs the decoded
30 metadata and the decoded audio signals to the
reproduction device 11, and the decoding process then

53
SP353112WO00
comes to an end. [0222]
In the above manner, the decoder 62 decodes the metadata and the audio signals, and outputs the metadata 5 including the location information about the ideal
speakers, the curve information, and the like, and the audio signals to the reproduction device 11. As the information formed with the location information about the ideal speakers, the curve information, and the like
10 is output as the metadata, the reproduction device 11 can perform appropriate gain correction, such as gain correction in accordance with the distances between the positions of the ideal speakers and the positions of the real reproduction speakers 12. As a result, audio
15 reproduction with a more realistic feeling can be performed. [0223]
The above-escribed series of processes may be performed by hardware or may be performed by software.
20 Where the series of processes are to be performed by software, the program that forms the software is installed into a computer. Here, the computer may be a computer incorporated into special-purpose hardware, or may be a general-purpose computer that can execute
25 various kinds of functions as various kinds of programs are installed thereinto. [0224]
Fig. 11 is a block diagram showing an example structure of the hardware of a computer that performs the
30 above-described series of processes in accordance with a program.

54
SP353112WO00'
[0225]
In the computer, a CPU 501, a ROM 502, and a RAM 503 are connected to one another by a bus 504. [0226] 5 An input/output interface 505 is further connected to the bus 504. An input unit 506, an output unit 507, a recording unit 508, a communication unit 509, and a drive 510 are connected to the input/output interface 505. [0227]
10 The input unit 506 is formed with a keyboard, a mouse, a microphone, an imaging device, and the like. The output unit 507 is formed with a display, a speaker, and the like. The recording unit 508 is formed with a hard disk, a nonvolatile memory, or the like. The
15 communication unit 509 is formed with a network interface or the like. The drive 510 drives a removable medium 511 such as a magnetic disk, an optical disk, a magnetooptical disk, or a semiconductor memory. [0228]
20 In the computer having the above-escribed structure, the CPU 501 loads a program recorded in the recording unit 508 into the RAM 503 via the input/output interface 505 and the bus 504, for example, and executes the program, so that the above-escribed series of processes
25 are performed. [0229]
The program to be executed by the computer (the CPU 501) may be recorded on the removable medium 511 as a packaged medium to be provided, for example.
30 Alternatively, the program can be provided via a v/ired or wireless transmission medium such as a local area network,

55

SP353112WO00

the Internet, or digital satellite broadcasting. [0230]
In the computer, the program can be installed into the recording unit 508 via the input/output interface 505 5 when the removable medium 511 is mounted on the drive 510. The program can also be received by the communication unit 509 via a wired or wireless transmission medium, and be installed into the recording unit 508. Alternatively, the program may be installed beforehand into the ROM 502
10 or the recording unit 508. [0231]
The program to be executed by the computer may be a program for performing processes in chronological order in accordance with the seguence described in this
15 specification, or may be a program for performing processes in parallel or performing a process when necessary, such as when there is a call. [0232]
It should be noted that embodiments of the present
20 technology are not limited to the above-escribed
embodiments, and various modifications may be made to them without departing from the scope of the present technology. [0233]
25 For example, the present technology can be embodied in a cloud computing structure in which one function is shared among devices via a network, and processing is performed by the devices cooperating with one another. [0234]
30 The respective steps described with reference to the above-escribed flowcharts can be carried out by one

56
SP353112WO00
device or can be shared among devices. [0235]
In a case where more than one process is included in one step, the processes included in the step can be 5 performed by one device or can be shared among devices. [0236]
Further, the present technology may take the following forms. [0237] 10 [1]
An audio signal output device including:
a distance calculating unit that calculates the distance between the position of an ideal speaker that reproduces an audio signal and the position of a real 15 speaker that reproduces the audio signal;
a gain calculating unit that calculates a reproduction gain of the audio signal based on the distance; and
a gain adjusting unit that performs gain adjustment 20 on the audio signal based on the reproduction gain. [2]
The audio signal output device of [1], wherein the gain calculating unit calculates the reproduction gain based on curve information for obtaining the reproduction 25 gain corresponding to the distance. [3]
The audio signal output device according to [2], wherein the curve information is information indicating a polyline curve or a function curve. 30 [4]
The audio signal output device according to [1] or

57
SP353112WO00
[2], wherein, when the ideal speaker is not located on a unit circle having a predetermined reference point as its center point, the gain adjusting unit further performs gain adj ustment on the audio signal with a gain 5 determined based on the distance from the reference point to the ideal speaker and the radius of the unit circle. [5]
The audio signal output device according to [ 4 j , wherein the gain adjusting unit delays the audio signal
10 based on a delay time determined based on the distance from the reference point to the ideal speaker and the radius of the unit circle. [6]
The audio signal output device according to [1] or
15 [2], wherein, when the real speaker is not located on a
unit circle having a predetermined reference point as its center point, the gain adjusting unit further performs gain adjustment on the audio signal with a gain determined based on the distance from the reference point
20 to the real speaker and the radius of the unit circle. [7]
The audio signal output device according to [6], wherein the gain adjusting unit delays the audio signal based on a delay time determined based on the distance
25 from the reference point to the real speaker and the radius of the unit circle. [8]
The audio signal output device according to any one of [1] through [7], further including
30 a gain correcting unit that corrects the
reproduction gain based on the distance between the

58
SP353112WO00
position of an ideal center speaker and the position of the real speaker.
[9]
The audio signal output device according to any one 5 of [1] through [8], further including
a lower limit correcting unit that corrects the reproduction gain when the reproduction gain is smaller than a predetermined lower limit. [10] 10 The audio signal output device according to any one of [1] through [9], further including
a total gain correcting unit that calculates a ratio between the total power of an output sound based on the audio signal subjected to the gain adjustment with 15 the reproduction gain and the total power of an input sound, and corrects the reproduction gain based on the ratio, the ratio being calculated based on the reproduction gain and an expected value of the sound pressure of the input sound based on the audio signal 20 input.
[11]
An audio signal output method including the steps of:
calculating the distance between the position of an 25 ideal speaker that reproduces an audio signal and the position of a real speaker that reproduces the audio signal;
calculating a reproduction gain of the audio signal based on the distance; and 30 performing gain adjustment on the audio signal based on the reproduction gain.

[12]
A program for causing a computer to perform a process including the steps of:
calculating the distance between the position of an ideal speaker that reproduces an audio signal and the position of a real speaker that reproduces the audio signal;
calculating a reproduction gain of the audio signal based on the distance; and
performing gain adjustment on the audio signal based on the reproduction gain. [13]
An encoding device including:
a correction information generating unit that generates correction information for correcting a gain of an audio signal in accordance with the distance between the position of an ideal speaker that reproduces the audio signal and the position of a real speaker that reproduces the audio signal;
an encoding unit that encodes the audio signal; and
an output unit that outputs a bit stream including the correction information and the encoded audio signal.
[14]
An encoding method including the steps of:
generating correction information for correcting a gain of an audio signal in accordance with the distance between the position of an ideal speaker that reproduces the audio signal and the position of a real speaker that reproduces the audio signal;
encoding the audio signal; and
outputting a bit stream including the correction

information and the encoded audio signal. [15]
A decoding device including: an extracting unit that extracts, from a bit stream, correction information for correcting a gain of an audio signal in accordance with the distance between the position of an ideal speaker that reproduces the audio signal and the position of a real speaker that reproduces the audio signal, and the encoded audio signal;
a decoding unit that decodes the encoded audio signal; and
an output unit that outputs the decoded audio signal and the correction information. [16]
The decoding device according to [15], wherein the correction information is the location information about the ideal speaker. [17]
The decoding device according to [15] or [16], wherein the correction information is curve information for obtaining a gain corresponding to the distance.
[18]
The decoding device according to [17], wherein the
curve information is information indicating a polyline
curve or a function curve.
[19]
A decoding method including the steps of: extracting, from a bit stream, correction
information for correcting a gain of an audio signal in
accordance with the distance between the position of an
ideal speaker that reproduces the audio signal and the

position of a real speaker that reproduces the audio
signal, and the encoded audio signal;
decoding the encoded audio signal; and outputting the decoded audio signal and the
correction information.
REFERENCE SIGNS LIST
[0238]
11 Reproduction device
21 Distance calculating unit
22 Reproduction gain calculating unit
23 Correcting unit
24 Lower limit correcting unit
25 Total gain correcting unit
26 Gain adjusting unit

61 Encoder
62 Decoder

71 Metadata generating unit
72 Audio signal encoding unit
73 Output unit

81 Extracting unit
82 Audio signal decoding unit
83 Output unit


CLAIMS
1. An audio signal output device comprising:
a distance calculating unit configured to calculate a distance between a position of an ideal speaker reproducing an audio signal and a position of a real speaker reproducing the audio signal;
a gain calculating unit configured to calculate a reproduction gain of the audio signal based on the distance; and
a gain adjusting unit configured to perform gain adjustment on the audio signal based on the reproduction gain.
2. The audio signal output device according to claim 1,
wherein the gain calculating unit calculates the
reproduction gain based on curve information for
obtaining the reproduction gain corresponding to the
distance.
3. The audio signal output device according to claim 2, wherein the curve information is information indicating one of a polyline curve and a function curve.
4. The audio signal output device according to claim 1, wherein, when the ideal speaker is not located on a unit circle having a predetermined reference point as a center point thereof, the gain adjusting unit further performs gain adjustment on the audio signal with a gain determined based on a distance from the reference point to the ideal speaker and a radius of the unit circle.

5. The audio signal output device according to claim 4, wherein the gain adjusting unit delays the audio signal based on a delay time determined based on the distance from the reference point to the ideal speaker and the radius of the unit circle.
6. The audio signal output device according to claim 1, wherein, when the real speaker is not located on a unit circle having a predetermined reference point as a center point thereof, the gain adjusting unit further performs gain adjustment on the audio signal with a gain determined based on a distance from the reference point to the real speaker and a radius of the unit circle.
7. The audio signal output device according to claim 6, wherein the gain adjusting unit delays the audio signal based on a delay time determined based on the distance from the reference point to the real speaker and the radius of the unit circle.
8. The audio signal output device according to claim 1, further comprising
a gain correcting unit configured to correct the reproduction gain based on a distance betv/een a position of an ideal center speaker and the position of the real speaker.
9. The audio signal output device according to claim 1,
further comprising
a lower limit correcting unit configured to correct

the reproduction gain when the reproduction gain is smaller than a predetermined lower limit.
10. The audio signal output device according to claim 1,
further comprising
a total gain correcting unit configured to calculate a ratio between total power of an output sound based on the audio signal subjected to the gain adjustment with the reproduction gain and total power of an input sound, and corrects the reproduction gain based on the ratio, the ratio being calculated based on the reproduction gain and an expected value of sound pressure of the input sound based on the audio signal input.
11. An audio signal output method comprising the steps
of:
calculating a distance between a position of an ideal speaker reproducing an audio signal and a position of a real speaker reproducing the audio signal;
calculating a reproduction gain of the audio signal based on the distance; and
performing gain adjustment on the audio signal based on the reproduction gain.
12. A program for causing a computer to perform a
process including the steps of:
calculating a distance between a position of an ideal speaker reproducing an audio signal and a position of a real speaker reproducing the audio signal;
calculating a reproduction gain of the audio signal based on the distance; and

performing gain adjustment on the audio signal based on the reproduction gain.
13. An encoding device comprising:
a correction information generating unit configured to generate correction information for correcting a gain of an audio signal in accordance with a distance between a position of an ideal speaker reproducing the audio signal and a position of a real speaker reproducing the audio signal;
an encoding unit configured to encode the audio signal; and
an output unit configured to output a bit stream including the correction information and the encoded audio signal.
14. An encoding method comprising the steps of:
generating correction information for correcting a
gain of an audio signal in accordance with a distance between a position of an ideal speaker reproducing the audio signal and a position of a real speaker reproducing the audio signal;
encoding the audio signal; and
outputting a bit stream including the correction information and the encoded audio signal.
15- A decoding device comprising:
an extracting unit configured to extract, from a bit stream, correction information for correcting a gain of an audio signal in accordance with a distance between a position of an ideal speaker reproducing the audio

signal and a position of a real speaker reproducing the audio signal, and the encoded audio signal;
a decoding unit configured to decode the encoded audio signal; and
an output unit configured to output the decoded audio signal and the correction information.
16. The decoding device according to claim 15, wherein the correction information is location information about the ideal speaker.
17. The decoding device according to claim 15, wherein the correction information is curve information for obtaining a gain corresponding to the distance.
18. . The decoding device according to claim 17, wherein-
the curve information is information indicating one of a
polyline curve and a function curve.
19. A decoding method comprising the steps of:
extracting, from a bit stream, correction
information for correcting a gain of an audio signal in accordance with a distance between a position of an ideal speaker reproducing the audio signal and a position of a real speaKer reproducing the audio signal, and the encoded audio signal;
decoding the encoded audio signal; and outputting the decoded audio signal and the correction information.

Documents

Application Documents

# Name Date
1 Priority Document [20-11-2015(online)].pdf 2015-11-20
2 Power of Attorney [20-11-2015(online)].pdf 2015-11-20
3 Form 5 [20-11-2015(online)].pdf 2015-11-20
4 Form 3 [20-11-2015(online)].pdf 2015-11-20
5 Form 1 [20-11-2015(online)].pdf 2015-11-20
6 Drawing [20-11-2015(online)].pdf 2015-11-20
7 Description(Complete) [20-11-2015(online)].pdf 2015-11-20
8 10666-DELNP-2015.pdf 2015-11-23
9 10666-delnp-2015-Form-1-(29-12-2015).pdf 2015-12-29
10 10666-delnp-2015-Correspondence Others-(29-12-2015).pdf 2015-12-29
11 10666-delnp-2015-Form-1-(13-01-2016).pdf 2016-01-13
12 10666-delnp-2015-Correspondence Others-(13-01-2016).pdf 2016-01-13
13 10666-delnp-2015-Form-3-(29-02-2016).pdf 2016-02-29
14 10666-delnp-2015-Correspondence Others-(29-02-2016).pdf 2016-02-29
15 Form 18 [19-05-2017(online)].pdf 2017-05-19
16 10666-DELNP-2015-FER.pdf 2019-07-26
17 10666-DELNP-2015-OTHERS [07-01-2020(online)].pdf 2020-01-07
18 10666-DELNP-2015-FER_SER_REPLY [07-01-2020(online)].pdf 2020-01-07
19 10666-DELNP-2015-CORRESPONDENCE [07-01-2020(online)].pdf 2020-01-07
20 10666-DELNP-2015-COMPLETE SPECIFICATION [07-01-2020(online)].pdf 2020-01-07
21 10666-DELNP-2015-CLAIMS [07-01-2020(online)].pdf 2020-01-07
22 10666-DELNP-2015-ABSTRACT [07-01-2020(online)].pdf 2020-01-07
23 10666-DELNP-2015-Power of Attorney-080120.pdf 2020-01-13
24 10666-DELNP-2015-Correspondence-080120.pdf 2020-01-13
25 10666-DELNP-2015-PatentCertificate28-08-2023.pdf 2023-08-28
26 10666-DELNP-2015-IntimationOfGrant28-08-2023.pdf 2023-08-28

Search Strategy

1 search_25-07-2019.pdf

ERegister / Renewals