7-35 KITASHINAGAWA 6-CHOME,
SHINAGAWA-KU, TOKYO 141-001,
JAPAN
Inventors
1. MINORU TSUJI SHIRO SUZUKI
7-35, KITASHINAGAWA 6-CHOME,
SHINAGAWA-KU , TOKYO,
JAPAN
2. KEISUKE TOYAMA
7-35 Kitashinagawa 6-chome
Shinagawa-ku,
Tokyo
Specification
FORM 2
THE PATENTS ACT 1970
[39 OF 1970}
&
THE PATENTS RULES, 2003
as amended by
THE PATENTS (AMENDMENT) RULES, 2006
COMPLETE SPECIFICATION
[See Section 10; rule 13]
AN ACOUSTIC SIGNAL ENCODING METHOD AND APPARATUS
THEREOF
SONY CORPORATION, a Japanese company, of 7-35 Kitashinagawa chome, Shinagawa-ku, Tokyo 141-001, Japan,
The following specification particularly describes the invention and manner in which it is to be performed:
The present invention relates an acoustic encoding method and apparatus thereof.
DESCRIPTION Acoustic Signal Encoding Method and Apparatus, Acoustic Signal Decoding Method and Apparatus and Recording Medium
Technical Field
The present invention relates to an acoustic signal encoding method and apparatus, and an acoustic signal decoding method and apparatus, in which acoustic signals are encoded and transmitted or recorded on a recording medium or the encoded * acoustic signals are received or reproduced and decoded on a decoding side. This invention also relates to an acoustic signal encoding- program, an acoustic signal decoding program and to a recording medium having recorded thereon a code string encoded by the acoustic signal encoding apparatus.
Background Art
A variety of techniques exist for high efficiency encoding of digital audio signals or speech signals. Examples of these techniques include a sub-band coding (SBC) of splitting e.g., time-domain audio signals into plural frequency bands, and encoding the signals from one frequency band to another, without blocking the time-domain-signals as a.non-blocking frequency band sphtting system, and a blocking, frequency band splitting system, or transform encoding, of converting time-domain signals by an orthogonal transform into frequency-domain signals, which frequency-
domain signals are encoded from one frequency band to another. There is also a technique of high efficiency encoding consisting in the combination of the sub-band coding and transform coding. In this case, the time-domain signals are divided into plural frequency bands by sub-band coding, and the resulting band-based signals are orthogonal-transformed into signals in the frequency domain, which signals are then encoded from one frequency band to another.
There are known techniques for orthogonal transform including the technique of dividing the digital input audio signals into blocks of a predetermined time duration, by way of blocking, and processing the resulting blocks using a Discrete Fourier Transform (DFT), discrete cosine transform (DCT) or modified DCT (MDCT) to convert the signals from the time axis to the frequency axis. Discussions of a MDCT may be found in J.P. Princen and A.B. Bradley, Subband/Transform Coding Using Filter Bank Designs Based on Time Domain Aliasing Cancellation, ICASSP, 1987, Univ. of Surrey Royal Melbourne Inst, of Tech.
By quantizing the signals, divided from band to band, using a filter or orthogonal transform, it is possible to control the band susceptible to quantization noise and, by exploiting such properties as masking effect, it is possible to achieve psychoacoustically more efficient encoding. If, prior to quantization, the signal components of the respective bands are normalized using the maximum absolute value of the signal components of each band, the encoding efficiency may be improved further.
In quantizing the frequency components, resulting from the division of the frequency spectrum, it is known to divide the frequency spectrum into widths which take characteristics of the human acoustic system into account. That is, audio signals are divided into plural bands, such as 32 bands, in accordance with band widths increasing with increasing frequency. In encoding the band-based data, bits are allocated fixedly or adaptively from band to band. When applying adaptive bit allocation to coefficient data resulting from MDCT, the MDCT coefficient data are encoded with an adaptively allocated number of bits from one frequency band resulting from the block-based MDCT to another.
It should be noted that, in orthogonal transform encoding and decoding of time-domain acoustic signals, the noise contained in tonal acoustic signals, the energy of which is concentrated in a specified frequency, is extremely harsh to the ear and hence may prove to be psychoacoustically highly objectionable. For this reason, a sufficient number of bits need to be used for encoding the tonal components. However, if the quantization step is determined fixedly from one band to another, as described above, the encoding efficiency is lowered because the bits are allocated uniformly to the totality of spectral components in an encoding unit containing the tonal components.
For coping with this deficiency, there is proposed in for example the International Patent Publication W094/28633 or Japanese Laying-Open Patent Publication 7-168593 a technique in which the spectral components are divided into
tonal and non-tonal components and finer quantization steps are used only for the tonal components.
In this technique, the spectral components with a locally high energy level, that is tonal components T, are removed from the spectrum on the frequency axis as shown in Fig. 1 A. The spectrum of noisy components, freed of tonal components, is shown in Fig.IB. The tonal and noisy components are quantized using sufficient optimum quantization steps.
However, in orthogonal transform techniques, such as MDCT, it is presupposed that the waveform in a domain being analyzed is repeated periodically outside the domain being analyzed. Consequently, the frequency components which really do not exist are observed. For example, if a sine wave of a certain frequency is input, and orthogonal-transformed by MDCT, the resulting spectrum covers not only the inherent frequency but also the ambient frequency, as shown in Fig. 1 A. Thus, if the sine wave is to be represented to high accuracy, not only the inherent sole frequency but also plural spectral components neighboring to the inherent frequency on the frequency axis need to be quantized with sufficient quantization steps, even though it is only being attempted by the above technique to quantize only the tonal components with high accuracy as shown in Fig. 1 A. As a result, more bits are needed, thus lowering the encoding efficiency.
Disclosure of the Invention
In view of the above depicted status of the art, it is an object of the present invention to provide an acoustic signal encoding method and apparatus, an acoustic signal decoding method and apparatus, an acoustic signal encoding program, an acoustic signal decoding program and a recording medium having recorded thereon a code string encoded by the acoustic signal encoding apparatus, whereby it is possible to prevent the encoding efficiency from being lowered due to a tonal component existing at a localized frequency.
An acoustic signal encoding method for encoding acoustic time-domain signals according to the present invention includes a tonal component encoding step of extracting tonal component signals from the acoustic time-domain signals and encoding the so extracted tonal component signals, and a residual component encoding step of encoding residual time-domain signals obtained on extracting the tonal component signals from the acoustic time-domain signals by the tonal component encoding step.
With this acoustic signal encoding method, tonal component signals are extracted from the acoustic time-domain signals and the tonal component signals as well as residual time-domain signals freed of the tonal component signals on extraction for the acoustic time-domain signals are encoded.
An acoustic signal decoding method for decoding acoustic signals in which tonal component signals are extracted from acoustic time-domain signals and encoded, and in which a code string obtained on encoding residual time-domain signals
corresponding to the acoustic time-domain signals freed on extraction of the tonal component signals is input and decoded, according to the present invention, includes a code string resolving step of resolving the code string, a tonal component decoding step of decoding the tonal component time-domain signals in accordance with the tonal component information obtained by the code string resolving step, a residual component decoding step of decoding residual component time-domain signals in accordance with the residual component information obtained by the code string resolving step, and a summation step of summing the tonal component time-domain signals obtained by the tonal component decoding step to the residual component time-domain signals obtained by the residual component decoding step to restore the acoustic time-domain signals.
With this acoustic signal decoding method, a code string obtained on extraction of tonal component signals from the acoustic time-domain signals and on encoding the tonal component signals as well as residual time-domain signals freed of the tonal component signals on extraction from the acoustic time-domain signals is decoded to restore acoustic time-domain signals.
An acoustic signal encoding method for encoding acoustic time-domain signals according to the present invention includes a frequency band splitting step of splitting the acoustic time-domain signals into a plurality of frequency bands, a tonal component encoding step of extracting tonal component signals from the acoustic time-domain signals of at least one frequency band and encoding the so extracted tonal
component signals, and a residual component encoding step of encoding residual time-domain signals freed on extraction of the tonal component by the tonal component encoding step from the acoustic time-domain signals of at least one frequency range.
With this acoustic signal encoding method, tonal component signals are extracted from the acoustic time-domain signals for at least one of plural frequency bands into which the frequency spectrum of the acoustic time-domain signals is split, and the residual time-domain signals, obtained on extracting the tonal component signals from the acoustic time-domain signals, are encoded.
An acoustic signal decoding method in which acoustic time-domain signals are split into a plurality of frequency bands, tonal component signals are extracted from the acoustic time-domain signals in at least one frequency band and encoded, a code string, obtained on encoding residual time-domain signals, obtained in turn on extracting the tonal component signals from the acoustic time-domain signals of at least one frequency band, is input, and in which the code string is decoded, according to the present invention, includes a code string resolving step of resolving the code string, a tonal component decoding step of synthesizing, for the at least one frequency band, tonal component time-domain signals in accordance with the residual component information obtained by the code string resolving step, a residual component decoding step of generating, for the at least one frequency band, residual component time-domain signals in accordance with the residual component information obtained by the
code string resolving step, a summation step of summing the tonal component time-domain signals obtained by the tonal component decoding step to the residual component time-domain signals obtained by the residual component decoding step, and a band synthesizing step of band-synthesizing decoded signals for each band to restore the acoustic time-domain signals.
With this acoustic signal decoding method, tonal component signals are extracted from the acoustic time-domain signals for at least one frequency band of the acoustic time-domain signals split into plural frequency bands, and the residual time-domain signals, obtained on extracting tonal component signals from the acoustic time-domain signals, are encoded to form a code string, which is then decoded to restore acoustic time-domain signals.
An acoustic signal encoding method for encoding acoustic signals according to the present invention includes a first acoustic signal encoding step of encoding the acoustic time-domain signals by a first encoding method including a tonal component encoding step of extracting tonal component signals from the acoustic time-domain signals and encoding the tonal component signals, a residual component encoding step of encoding residual signals obtained on extracting the tonal component signals from the acoustic time-domain signals by the tonal component encoding step, and a code string generating step of generating a code string from the information obtained by the tonal component encoding step and the information obtained from the residual component encoding step, a second acoustic signal encoding step of encoding the
t acoustic time-domain signals by a second encoding method, and an encoding efficiency decision step of comparing the encoding efficiency of the first acoustic signal encoding step to that of the second acoustic signal encoding step to select a code string with a better encoding efficiency.
With this acoustic signal encoding method, a code string obtained by the first acoustic signal encoding process of encoding the acoustic time-domain signals by a first encoding method of extracting tonal component signals from the acoustic time-domain signals, and encoding the residual time-domain signals, obtained on extracting tonal component signals from the acoustic time-domain signals, or a code string obtained by a second encoding process of encoding the acoustic time-domain signals by a second encoding method, whichever has a higher encoding efficiency, is selected.
An acoustic signal decoding method for decoding a code string which is selectively input in such a manner that a code string encoded by a first acoustic signal encoding step or a code string encoded by a second acoustic signal encoding step, whichever is higher in encoding efficiency, is selectively input and decoded, the first acoustic signal encoding step being such a step in which the acoustic signals are encoded by a first encoding method comprising generating a code string from the information obtained on extracting tonal component signals from acoustic time-domain signals and on encoding the tonal component signals and from the information obtained on encoding residual signals obtained on extracting the tonal component signals from the acoustic time-domain signals, the second acoustic signal encoding
step being such a step in which the acoustic signals are encoded by a second encoding method, according to the present invention, is such a method wherein, if the code string resulting from encoding in the first acoustic signal encoding step is input, the acoustic time-domain signals are restored by a first acoustic signal decoding step including a code string resolving sub-step of resolving the code string into the tonal component information and the residual component information, a tonal component decoding step of generating the tonal component time-domain signals in accordance with the tonal component information obtained in the code string resolving sub-step, a residual component decoding step of generating residual component time-domain signals in accordance with the residual component information obtained in the code string resolving sub-step and a summation sub-step of summing the tonal component time-domain signals to the residual component time-domain signals, and wherein, if the code string obtained on encoding in the second acoustic signal encoding step is input, the acoustic time-domain signals are restored by a second acoustic signal decoding sub-step corresponding to the second acoustic signal encoding step.
With this acoustic signal decoding apparatus, a code string obtained by a first acoustic signal encoding method of encoding the acoustic time-domain signals by a first encoding method of extracting tonal component signals from the acoustic time-domain signals, and encoding the residual time-domain signals, obtained on extracting tonal component signals from the acoustic time-domain signals, or a code string obtained by a second encoding process of encoding the acoustic time-domain signals
by a second encoding method, whichever has a higher encoding efficiency, is input and decoded by an operation which is the counterpart of the operation performed on the side encoder.
An acoustic signal encoding apparatus for encoding acoustic time-domain signals, according to the present invention, includes tonal component encoding means for extracting tonal component signals from the time-domain signals and encoding the so extracted signals, and residual component encoding means for encoding residual time-domain signals, freed on extraction of the tonal component information from the acoustic time-domain signals by the tonal component encoding means.
With this acoustic signal encoding apparatus, the tonal component signals are extracted from the acoustic time-domain signals and the tonal component signals as well as the residual time-domain signals freed of the tonal component signals on extraction by the tonal component encoding means from the acoustic time-domain signals are encoded.
An acoustic signal decoding apparatus in which a code string resulting from extracting tonal component signals from acoustic time-domain signals, encoding the tonal component signals and from encoding residual time-domain signals corresponding to the acoustic time-domain signals freed on extraction of the tonal component signals, is input and decoded, according to the present invention, includes code string resolving means for resolving the code string, tonal component decoding means for decoding the tonal component time-domain signals in accordance with the
tonal component information obtained by the code string resolving means, residual component decoding means for decoding the residual time-domain signals in accordance with the residual component information obtained by the code string resolving means, and summation means for summing the tonal component time-domain signals obtained from the tonal component decoding means and the residual component time-domain signals obtained from the residual component decoding means to restore the acoustic time-domain signals.
With this acoustic signal decoding apparatus, a code string obtained on extracting the tonal component signals from the acoustic time-domain signals and on encoding the tonal component signals as well as the residual time-domain signals freed of the tonal component signals on extraction by the tonal component encoding means from the acoustic time-domain signals is decoded to restore the acoustic time-domain signals.
A computer-controllable recording medium, having recorded thereon an acoustic signal encoding program configured for encoding acoustic time-domain signals, according to the present invention, is such a recording medium in which the acoustic signal encoding program includes a tonal component encoding step of extracting tonal component signals from the time-domain signals and encoding the so extracted signals, and a residual component encoding step of encoding residual time-domain signals, freed on extraction of the tonal component signals from the acoustic time-domain signals by the tonal component encoding step.
On this recording medium, there is recorded an acoustic signal encoding program of extracting the tonal component signals from the acoustic time-domain signals and on encoding the tonal component signals as well as the residual time-domain signals freed of the tonal component signals on extraction by the tonal component encoding means from the acoustic time-domain signals.
A computer-controllable recording medium, having recorded thereon an acoustic signal decoding program of decoding acoustic time-domain signals, according to the present invention, is such a recording medium in which the acoustic signal decoding program includes a code string resolving step of resolving the code string, a tonal component decoding step of decoding the tonal component time-domain signals in accordance with the tonal component information obtained by the code string resolving step, a residual component decoding step of decoding the residual time-domain signals in accordance with the residual component information obtained by the code string resolving step, and a summation step of summing the tonal component time-domain signals obtained from the tonal component decoding step and the residual component time-domain signals obtained from the residual component decoding step to restore the acoustic time-domain signals.
On this recording medium, there is recorded an acoustic signal decoding program of decoding a code string obtained on extracting the tonal component signals from the acoustic time-domain signals and on encoding the tonal component signals as well as the residual time-domain signals freed of the tonal component signals on
extraction by the tonal component encoding means from the acoustic time-domain signals to restore the acoustic time-domain signals.
A recording medium according to the present invention has recorded thereon a code string obtained on extracting tonal component signals from acoustic time-domain signals, encoding the tonal component signals and on encoding residual time-domain signals corresponding to the acoustic time-domain signals freed on extraction of the tonal component signals from the acoustic time-domain signals.
Other objects, features and advantages of the present invention will become more apparent from reading the embodiments of the present invention as shown in the drawings.
Brief Description of the Drawings
Figs. 1A and IB illustrate a conventional technique of extracting a tonal component, Fig. 1A illustrating the spectrum prior to removal of the tonal component and Fig. IB illustrating the spectrum of noisy components subsequent to removal of the tonal component.
Fig. 2 illustrates a structure of an encoding apparatus for acoustic signals embodying the present invention.
Figs.3A to 3C illustrate a method for smoothly linking extracted time domain signals to a directly previous frame and to the next frame, Fig.3A showing a frame in MDCT, Fig.3B showing a domain from which to extract the tonal component and
Fig.3C showing a window function for synthesis of the directly previous frame and the next frame.
Fig.4 illustrates a structure of a tonal component encoding unit of the encoding apparatus for acoustic signals.
Fig.5 illustrates a first structure of the tonal component encoding unit in which the quantization error is contained in residual time-domain signals.
Fig. 6 illustrates a first structure of the tonal component encoding unit in which the quantization error is contained in residual time-domain signals.
Fig.7 illustrates an instance of determining normalization coefficients using the maximum amplitude values of extracted plural sine waves as reference.
Fig.8 is a flowchart for illustrating a sequence of operations of an acoustic signal encoding apparatus having the tonal component encoding unit of Fig. 6.
Figs.9A and 9B illustrate parameters of a waveform of a pure sound, Fig.9A showing an example of using the frequency and the amplitudes of sine and cosine waves and Fig.9B showing an example of using the frequency, amplitudes and the phase.
Fig. 10 is a flowchart showing a sequence of operations of an acoustic signal encoding apparatus having the tonal component encoding unit of Fig.5.
Fig. 11 illustrates a structure of an acoustic signal decoding apparatus embodying the present invention.
Fig. 12 illustrates a structure of a tonal component decoding unit of the acoustic
signal decoding apparatus.
Fig. 13 is a flowchart showing a sequence of operations of the acoustic signal decoding apparatus.
Fig. 14 illustrates another structure of the a residual component encoding unit of the acoustic signal decoding apparatus.
Fig. 15 shows an illustrative structure of a residual signal decoding unit as a counterpart of the residual component encoding unit shown in Fig. 14.
Fig. 16 illustrates a second illustrative structure of the acoustic signal encoding apparatus and the acoustic signal decoding apparatus.
Fig. 17 shows a third illustrative structure of the acoustic signal encoding apparatus and the acoustic signal decoding apparatus.
Best Mode for Carrying out the Invention
Referring to the drawings, certain preferred embodiments of the present invention will be explained in detail.
An illustrative structure of the acoustic signal encoding apparatus embodying the present invention is shown in Fig.2, in which an acoustic signal encoding apparatus 100 is shown to include a tonal noise verification unit 110, a tonal component encoding unit 120, a residual component encoding unit 130, a code string generating unit 140 and a time domain signal holding unit 150.
The tonal noise verification unit 110 verifies whether the input acoustic time-
domain signals S are a tonal signal or a noise signal to output a tone/noise verification code T/N depending on the verified results to switch the downstream side processing.
The tonal component encoding unit 120 extracts a tonal component from an input signal to encode the tonal component signal, and includes a tonal component extraction unit 121 for extracting a tonal component parameter N-TP from an input signal determined to be tonal by the tonal noise verification unit 110, and a normalization/quantization unit 122 for normalizing and quantizing the tonal component parameter N-TP obtained in the tonal component extraction unit 121 to output a quantized tonal component parameter N-QTP.
The residual component encoding unit 130 encodes residual time-domain signals RS, resulting from extraction by the tonal component extraction unit 121 of the tonal component from the input signal determined to be tonal by the tonal noise verification unit 110, or the input signal determined to be noisy by the tonal noise verification unit 110. The residual component encoding unit 130 includes an orthogonal transform unit 131 for transforming these time-domain signals into the spectral information NS by for example modified discrete cosine transformation (MDCT), and a normalization/quantization unit 132 for normalizing and quantizing the spectral information NS, obtained by the orthogonal transform unit 131, to output the quantized spectral information QNS.
The code string generating unit 140 generates and outputs a code string C, based on the information from the tonal component encoding unit 120 and the residual
component encoding unit 130.
The time domain signal holding unit 150 holds the time domain signals input to the residual component encoding unit 130. The processing in the time domain signal holding unit 150 will be explained subsequently.
Thus, the acoustic signal encoding apparatus 100 of the present embodiment switches the downstream side encoding processing techniques, from one frame to the next, depending on whether the input acoustic time domain signals are tonal or noisy. That is, the acoustic signal encoding apparatus extracts the tonal component signals of the tonal signal to encode parameters thereof, using the generalized harmonic analysis (GHA), as later explained, while encoding the residual signals, obtained on extracting the tonal signal component from the tonal signal, and the noisy signal, by orthogonal transform with for example MDCT, and subsequently encoding the transformed signals.
Meanwhile, in MDCT used in general in orthogonal transform, a frame for analysis (encoding unit) needs one-half frame overlap with each of directly forward and directly backward frames, as shown in Fig. 3 A. Moreover, the frame for analysis in the generalized harmonic technique analysis in tonal component encoding processing may be endowed with one-half frame overlap with the directly forward and directly backward frames, such that the extracted time domain signals can be smoothly linked to the extracted time domain signals of the directly forward and directly backward frames.
However, since there is the one-half frame overlap in the analysis frame of MDCT, as described above, the time domain signals of a domain A during analysis of the first frame must not differ from the time domain signals of the domain A during analysis of the second frame. Thus, in the residual component encoding processing, extraction of the tonal component during the domain A needs to be completed at a time point the first frame has been orthogonal transformed. Consequently, the following processing is desirably performed.
First, in encoding the tonal components, pure sound analysis is carried out by generalized harmonic analysis in a domain of the second frame shown in Fig.3B. Subsequently, waveform extraction is carried out on the basis of the produced parameters. The domain of extraction is to be overlapped with the first frame. The analysis of pure tone by generalized harmonic analysis in a domain of the first frame has already been finished, such that waveform extraction in this domain is carried out based on the parameters obtained in each of the first and second frames. If the first frame has been determined to be noisy, waveform extraction is carried out based only on the parameters obtained in the second frame.
Next, the time-domain signals, extracted in each frame, are synthesized as follows: That is, the time domain signals by parameters analyzed in each frame is multiplied with a window function which on summation gives unity, such as Harming function shown in the following equation (1):
-(1) where 0 < t < L, to synthesize time-domain signals in which transition from the first frame to the second frame is smooth, as shown in Fig.3C. In the equation (1), L stands for the frame length, that is the length of one encoding unit.
The synthesized time domain signals are extracted from the input signal. Thus, residual time domain signals in the overlap domain of the first and second frames are found. These residual time domain signals serve as residual time-domain signals of the latter one-half of the first frame. The encoding of the residual components of the first frame is by forming residual time-domain signals of the first frame by the residual time-domain signals of the latter one-half of the first frame and by the residual time-domain signals of the former one-half of the first frame already held, orthogonal-transforming the residual time-domain signals of the first frame and by normalizing and quantizing the so produced spectral information. By generating the code string by the tonal component information of the first frame and the residual component information of the first frame, it is possible to synthesize the tonal components and the residual components in one frame at the time of decoding.
Meanwhile, if the first frame is the noisy signal, there lack tonal component parameters of the first frame. Consequently, the above-mentioned window function is
multiplied only with the time-domain signals extracted in the second frame. The so produced time-domain signals are extracted from the input signal, with the residual time-domain signals similarly serving as residual time-domain signals of the latter one-half of the first frame.
The above enables extraction of smooth tonal component time-domain signals having no discontinuous points. Moreover, it is possible to prevent frame-to-frame non-matching in MDCT in encoding the residual components.
For carrying out the above processing, the acoustic signal encoding apparatus 100 includes the time domain signal holding unit 150 ahead of the residual component encoding unit 130, as shown in Fig.2. This time domain signal holding unit 150 holds residual time-domain signals every one-half frame. The tonal component encoding unit 120 includes parameter holding portions 2115,2217 and 2319, as later explained, and outputs waveform parameters and the extracted waveform information of the previous frame.
The tonal component encoding unit 120, shown in Fig.2, may specifically be configured as shown in Fig.4. For frequency analysis in tonal component extraction, tonal component synthesis and tonal component extraction, the generalized harmonic analysis, as proposed by Wiener, is applied. This technique is such an analysis technique in which the sine wave which gives the smallest residual energy in an analysis block is extracted from the original time-domain signals, with this processing being repeated for the resulting residual signals. With this technique, frequency
components can be extracted one by one in the time domain without being influenced by the analysis window. Moreover, the frequency resolution can be freely set, such that frequency analysis can be achieved more precisely than is possible with Fast Fourier transform (FFT) or MDCT.
A tonal component encoding unit 2100, shown in Fig.4, includes a tonal component extraction unit 2110 and a normalization/quantization unit 2120. The tonal component extraction unit 2110 and the normalization/quantization unit 2120 are similar to the component extraction unit 121 and the normalization/quantization unit 122 shown in Fig. 2.
In the tonal component encoding unit 2100, a pure sound analysis unit 2111 analyzes a pure sound component, which minimizes the energy of the residual signals, from the input acoustic time-domain signals S. The pure sound analysis unit then sends the pure sound waveform parameter TP to a pure sound synthesis unit 2112 and to a parameter holding unit 2115.
The pure sound synthesis unit 2112 synthesizes a pure sound waveform time-domain signals TS of the pure sound component, analyzed by the pure sound analysis unit 2111. A subtractor 2113 extracts the pure sound waveform time-domain signals TS, synthesized by the pure sound synthesis unit 2112, from the input acoustic time-domain signals S.
An end condition decision unit 2114 checks whether or not the residual signals obtained by pure sound extraction in the subtractor 2113 meet the end condition for
tonal component extraction, and effects switching for repeating pure sound extraction, with the residual signal as the next input signal for the pure sound analysis unit 2111, until the end condition is met. This end condition will be explained subsequently.
The parameter holding unit 2115 holds the pure sound waveform parameter TP of the current frame and a pure sound waveform parameter of the previous frame PrevTP to route the pure sound waveform parameter of the previous frame PrevTP to a normalization/quantization unit 2120, while routing the pure sound waveform parameter TP of the current frame and the pure sound waveform parameter of the previous frame PrevTP to an extracted waveform synthesis unit 2116.
The extracted waveform synthesis unit 2116 synthesizes the time-domain signals by the pure sound waveform parameter TP in the current frame to the time-domain signals by the pure sound waveform parameter of the previous frame PrevTP, using the aforementioned Hanning function, to generate tonal component time-domain signals N-TS for an overlap domain. A subtractor 2117 extracts the tonal component time-domain signals N-TS from the input acoustic time-domain signals S to output residual time-domain signals RS for the overlap domain. These residual time-domain signals RS are sent to and held by the time domain signal holding unit 150 shown in Fig.2.
The normalization/quantization unit 2120 normalizes and quantizes the pure sound waveform parameter of the previous frame PrevTP, supplied from the parameter holding unit 2115, to output a quantized tonal component parameter of the previous
frame PrevN-QTP.
It should be noted that the configuration shown in Fig.4 is susceptible to quantization error in encoding the tonal component. In order to combat this, such a configuration may be used, in which the quantization error is contained in the residual time-domain signals, as shown in Figs.5 and 6.
As a first configuration for having the quantization error included in the residual time-domain signals, a tonal component encoding unit 2200, shown in Fig. 5, includes a normalization/quantization unit 2212 in the tonal component extraction unit 2210, for normalizing and quantizing the tonal signal information.
In the tonal component encoding unit 2200, a pure sound analysis unit 2211 analyzes a pure sound component, which minimizes the residual signals, from the input acoustic time-domain signals S, to route the pure sound waveform parameter TP to the normalization/quantization unit 2212.
The normalization/quantization unit 2212 normalizes and quantizes the pure sound waveform parameter TP, supplied from the pure sound analysis unit 2211, to send the quantized pure sound waveform parameter QTP to an inverse quantization inverse normalization unit 2213 and to a parameter holding unit 2217.
The inverse quantization inverse normalization unit 2213 inverse quantizes and inverse normalizes the quantized pure sound waveform parameter QTP to route inverse quantized pure sound waveform parameter TP' to a pure sound synthesis unit 2214 and to the parameter holding unit 2217.
The pure sound synthesis unit 2214 synthesizes the pure sound waveform time-domain signals Ts of the pure sound component, based on the inverse quantized pure sound waveform parameter TP', to extract at subtractor 2215 the pure sound waveform time-domain signals TS, synthesized by the pure sound synthesis unit 2214, from the input acoustic time-domain signals S.
An end condition decision unit 2216 checks whether or not the residual signals obtained on pure sound extraction by the subtractor 2215 meets the end condition of tonal component extraction and effects switching for repeating pure sound extraction, with the residual signal as the next input signal for the pure sound analysis unit 2211, until the end condition is met. This end condition will be explained subsequently.
The parameter holding unit 2217 holds the quantized pure sound waveform parameter QTP and an inverse quantized pure sound waveform parameter TP' to output the quantized tonal component parameter of the previous frame PrevN-QTP, while routing the inverse quantized pure sound waveform parameter TP' and the inverse quantized pure sound waveform parameter of the previous frame PrevTP' to an extracted waveform synthesis unit 2218.
The extracted waveform synthesis unit 2218 synthesizes time-domain signals by the inverse quantized pure sound waveform parameter TP' in the current frame to the time-domain signals by the inverse quantized pure sound waveform parameter of the previous frame PrevTP', using the aforementioned Hanning function, to generate tonal component time-domain signals N-TS for an overlap domain. A subtractor 2219
extracts the tonal component time-domain signals N-TS from the input acoustic time-domain signals S to output residual time-domain signals RS for the overlap domain. These residual time-domain signals RS are sent to and held by the time domain signal holding unit 150 shown in Fig.2.
As a second configuration of having the quantization error included in the residual time-domain signals, a tonal component encoding unit 2300, shown in Fig.6, also includes a normalization/quantization unit 2315, adapted for normalizing and quantizing the information of the tonal signals, in a tonal component extraction unit 2310.
In the tonal component encoding unit 2300, a pure sound analysis unit 2311 analyzes the pure sound component, which minimizes the energy of the residual signals, from the input acoustic time-domain signals S. The pure sound analysis unit routes the pure sound waveform parameter TP to a pure sound synthesis unit 2312 and to a normalization/quantization unit 2315.
The pure sound synthesis unit 2312 synthesizes the pure sound waveform time-domain signals TS, analyzed by the pure sound analysis unit 2311, and a subtractor 2313 extracts the pure sound waveform time-domain signals TS, synthesized by the pure sound synthesis unit 2312, from the input acoustic time-domain signals S.
An end condition decision unit 2314 checks whether or not the residual signals obtained by pure sound extraction by the subtractor 2313 meets the end condition for tonal component extraction, and effects switching for repeating pure sound extraction,
with the residual signal as the next input signal for the pure sound analysis unit 2311, until the end condition is met.
The normalization/quantization unit 2315 normalizes and quantizes the pure sound waveform parameter TP, supplied from the pure sound analysis unit 2311, and routes the quantized pure sound waveform parameter N-QTP to an inverse quantization inverse normalization unit 2316 and to a parameter holding unit 2319.
The inverse quantization inverse normalization unit 2316 inverse quantizes and inverse normalizes the quantized pure sound waveform parameter N-QTP to route the inverse quantized pure sound waveform parameter N-TP' to the parameter holding unit 2319.
The parameter holding unit 2319 holds the quantized pure sound waveform parameter N-QTP and the inverse quantized pure sound waveform parameter N-TP' to output the quantized tonal component parameter of the previous frame PrevN-QTP. The parameter holding unit also routes the inverse quantized pure sound waveform parameter for the current frame N-TP' and the inverse quantized pure sound waveform parameter of the previous frame PrevN-TP' to the extracted waveform synthesis unit 2317.
The extracted waveform synthesis unit 2317 synthesizes time-domain signals by the inverse quantized pure sound waveform parameter of the current frame N-TP' to the inverse quantized pure sound waveform parameter of the previous frame PrevN-TP', using for example the aforementioned Harming function, to generate the tonal
component time-domain signals N-TS for the overlap domain. A subtractor 2318 extracts the tonal component time-domain signals N-TS from the input acoustic time-domain signals S to output the residual time-domain signals RS for the overlap domain. These residual time-domain signals RS are sent to and held in the time domain signal holding unit 150 of Fig.2.
Meanwhile, in the illustrative structure of Fig. 5, the normalization coefficient for the amplitude is fixed for a value not less than the maximum value that can be assumed. For example, if the input signal is the acoustic time-domain signals, recorded on a music Compact Disc (CD), quantization is carried out using 96 dB as the normalization coefficient. Meanwhile, the normalization coefficient is of a fixed value and hence need not be included in the code string.
Conversely, with the illustrative structures shown in Figs.4 and 6, it is possible to determine the normalization coefficient with the maximum amplitude value of the extracted plural sine waves as a reference, as shown for example in Fig.7. That is, an optimum normalization coefficient is selected from among the plural normalization coefficients, provided at the outset, and the amplitude values of the totality of the sine waves are quantized using this normalization coefficient. In this case, the information indicating the normalization coefficient used in the quantization is included in the code string. In the case of the illustrative structures, shown in Figs.4 and 6, as compared to the illustrative structure of Fig. 5, quantization may be achieved to a higher accuracy, even though the quantity of bits is increased by a value corresponding to the
information indicating the normalization coefficient.
The processing by the acoustic signal encoding apparatus 100 in case the tonal component encoding unit 120 of Fig.2 is configured as shown in Fig.6 is now explained in detail with reference to the flowchart of Fig. 8.
First, at step S1, the acoustic time-domain signals are input for a certain preset analysis domain (number of samples).
At the next step S2, it is checked whether or not the input time-domain signals
are tonal. While a variety of methods for decision may be envisaged, it may be
contemplated to process e.g., the input time-domain signal x(t) with spectral analysis,
such as by FFT, and to give a decision that the input signal is tonal when the average
value AVE (X(k)) and the maximum value Max (X(k)) of the resulting spectrum X(k)
meet the following equation (2):
Max(X(k)) ^ AVE(X(k))
•••(2) that is when the ratio thereof is larger than a preset threshold Th,one.
If it is determined at step S2 that the input signal is tonal, processing transfers to step S3. If it is determined that the input signal is noisy, processing transfers to stepS 10.
At step S3, such frequency component which give the smallest residual energy
is found from the input time-domain signals. The residual components, when the pure sound waveform with a frequency f is extracted from the input time-domain signals x()(t), are depicted by the following equation (3):
RSj(t) = Xo(t) - S,sin(27ift) - Cfcos(2nft)
•••(3) where L denotes the length of the analysis domain (number of samples).
In the above equation (3), St and C, may be depicted by the following equations (4) and (5):
"•(5). In this case, the residual energy Ef is given by the following equation (6):
•••(6). The above analysis is carried out for the totality of frequencies f to find the frequency f, which will give the smallest residual energy Ef.
At the next step S4, the pure sound waveform of the frequency f1, obtained at step S3, is extracted from the input time-domain signals x0(t) in accordance with the following equation (7):
-(7).
At step S5, it is checked whether or not the end condition for extraction has been met. The end condition for extraction may be exemplified by the residual time-domain signals not being tonal signals, the energy of the residual time-domain signals having fallen by not less than a preset value from the energy of the input time-domain signals, the decreasing amount of the residual time-domain signals resulting from the pure sound extraction being not higher than a threshold value, and so forth.
If, at step S5, the end condition for extraction is not met, program reverts to step S3 where the residual time-domain signals obtained in the equation (7) are set as the next input time-domain signals x,(t). The processing as from step S3 to step S5 is repeated N times until the end condition for extraction is met. If, at step S5, the end condition for extraction is met, processing transfers to step S6.
At step S6, the N pure sound information obtained, that is the tonal component information N-TP, is normalized and quantized. The pure sound information may, for example, be the frequency fn, amplitude Sfn or amplitude Cfn of the extracted pure sound waveform, shown in Fig. 9A, or the frequency fn, amplitude Afn or phase Pfn, shown in Fig.9B where 0 < n < N. The frequency fn, amplitude Sfn, amplitude Cft
amplitude A& and the phase F^ are correlated with one another in accordance with the following equations (8) to (10):
•••(8)
•"(9)
-(10).
At the next step S7, the quantized pure sound waveform parameter N-QTP is inverse quantized and inverse normalized to obtain the inverse quantized pure sound waveform parameter N-TP'. By first normalizing and quantizing the tonal component information and subsequently inverse quantizing and inverse normalizing the component information, time-domain signals, which may be completely identified with the tonal component time-domain signals, extracted here, may be summed during the process of decoding the acoustic time-domain signals.
At the next step S8, the tonal component time-domain signals N-TS is generated in accordance with the following equation (11):
(0
Documents
Application Documents
#
Name
Date
1
167-mumnp-2003-power of authority(27-12-2002).pdf
2002-12-27
2
167-mumnp-2003-form 5(03-02-2003).pdf
2003-02-03
3
167-mumnp-2003-form 3(03-02-2003).pdf
2003-02-03
4
167-mumnp-2003-form 1a(04-02-2003).pdf
2003-02-04
5
167-mumnp-2003-form 18(18-05-2006).pdf
2006-05-18
6
167-mumnp-2003-power of authority(16-01-2008).pdf
2008-01-16
7
167-mumnp-2003-form-pct-isa-210(16-01-2008).pdf
2008-01-16
8
167-mumnp-2003-form 2(granted)-(16-01-2008).pdf
2008-01-16
9
167-mumnp-2003-form 1a(16-01-2008).pdf
2008-01-16
10
167-mumnp-2003-form 13(16-01-2008).pdf
2008-01-16
11
167-mumnp-2003-drawing(16-01-2008).pdf
2008-01-16
12
167-mumnp-2003-claim(granted)-(16-01-2008).pdf
2008-01-16
13
167-mumnp-2003-cancelled page(16-01-2008).pdf
2008-01-16
14
167-mumnp-2003-abstract(16-01-2008).pdf
2008-01-16
15
167-mumnp-2003-petition under rule 138(01-02-2008).pdf
2008-02-01
16
167-mumnp-2003-petition under rule 137(01-02-2008).pdf