Sign In to Follow Application
View All Documents & Correspondence

Audio Processing Device Audio Processing Method And Program

Abstract: ABSTRACT The present invention relates to a speech processing apparatus, a speech processing method and a program which, when multichannel audio signals are downmixed and coded, prevevent delay and an increase in the computation amount upon decoding of the audio signals.An inverse multiplexing unit (101) acquires coded data on which a BC parameter is multiplexed. An uncorrelated frequency-time transform unit (102) performs IMDCT transform" and IMDST transform of frequency spectrum coefficients of a monaural signal (XM) obtained from this coded data to generate the monaural signal XM) which is a time domain signal and a signal (XD" ) which is substantially uncorrelated with this monaural signal (XM) . The stereo synthesis unit (103) generates a stereo signal by synthesizing the monaural signal (XM) and the signal (XD") using the BC parameter. The" present invention is applicable to, for example, a speech processing apparatus which decodes a downmixed and coded stereo signal.

Get Free WhatsApp Updates!
Notices, Deadlines & Correspondence

Patent Information

Application #
Filing Date
10 September 2012
Publication Number
02/2014
Publication Type
INA
Invention Field
ELECTRONICS
Status
Email
remfry-sagar@remfry.com
Parent Application

Applicants

SONY CORPORATION
1 7 1 Konan Minato ku Tokyo 1080075

Inventors

1. TOGURI Yasuhiro
c/o SONY CORPORATION 1 7 1 Konan Minato ku Tokyo 1080075
2. SUZUKI Shiro
c/o SONY CORPORATION 1 7 1 Konan Minato ku Tokyo 1080075
3. MATSUMOTO Jun
c/o SONY CORPORATION 1 7 1 Konan Minato ku Tokyo 1080075
4. MAEDA Yuuji
c/o SONY CORPORATION 1 7 1 Konan Minato ku Tokyo 1080075
5. MATSUMURA Yuuki
c/o SONY CORPORATION 1 7 1 Konan Minato ku Tokyo 1080075

Specification

DESCRIPTION SPEECH PROCESSING APPARATUS, SPEECH PROCESSING METHOD AND PROGRAM TECHNICAL FIELD [0001] The present invention relates to a speech processing apparatus, a speech processing method and a program and, more particularly, relates to a speech processing apparatus, a speechprocessingmethodandaprogramwhich, whenmultichannel audio signals are downmixed and coded, prevent delay and an increase in the computation amount upon decoding of the audio signals'. BACKGROUND ART [0002] A coding apparatus which codes multichannel audio signals can perform highly efficient coding by utilizing a relationship between channels.. This coding includes, for. example, intensity coding, M/S stereo coding and spatial coding. A coding apparatus which performs spatial coding downmixes an n channel audio signal into a m (m < n) channel audio signal and codes the signal, finds spatial parameters representing the inter-channel relationship upon downmixing and transmits the spatial parameters together with the coded data. A decoding apparatus which receives the spatial parameters and the coded data decodes the coded data, and restores the original n channel audio signal from the m channel audio signal obtained as a result of decoding using the spatial parameter. [0003] This spatial coding is known as "binaural cue coding". For the spatial parameters (hereinafter, referred to as "BC parameters"), for example, ILD (Inter-channel Level Difference), IPD (Inter-channel Phase Difference) and ICC (Inter-channel Correlation) are used. The ILD refers to a parameter indicating, the ratio of the magnitude of an inter-channel signal. The IPD refers to a parameter indicating an inter-channel phase difference, and the ICC refers to a parameter indicating an inter-channel correlation. [0004] Fig. 1 is a block diagram illustrating a configuration example of a coding apparatus which performs spatial coding. [0005] In addition, n = 2 and m = 1 for ease of description. That is, a coding target audio signal is a stereo audio signal (hereinafter, referred to as "stereo signal") , and coded data obtained as a result of coding is coded data of a monaural audio signal (hereinafter, referred to as "monaural signal"). [0006] A coding apparatus 10 in Fig. 1 includes a channel donwmix unit 11, a spatial parameter detection unit 12, an audio signal coding unit 13 and amultiplexing unit 14 . The coding apparatus 10 receives an input of a stereo signal including a left audio signal XL and a right audio signal XR as a coding target, and outputs coded data of a monaural signal. [0007] More specifically, the channel downmix unit 11 of the coding apparatus 10 downmixes the stereo signal input as the coding target, to the monaural signal XM. Further, the channel downmix unit 11 supplies the monaural signal to the spatial parameter detection unit 12 and the audio signal coding unit 13. [0008] The spatial parameter detection unit 12 detects the BC parameters based on the monaural signal XM supplied from the channel downmix unit 11 and the stereo signal input as the coding target, and supplies the BC parameters to the multiplexing unit 14. [0009] The audio signal coding unit 13 codes the monaural signal supplied from the channel downmix unit 11, and supplies resulting coded data to the multiplexing unit 14. [0010] The multiplexing unit 14 multiplexes and outputs the coded data supplied from the audio signal coding unit 13 and the BC parameter supplied from the spatial parameter detection unit 12. [0011] Fig. 2 is ,a block diagram illustrating a configuration example of the audio signal coding unit 13 in Fig. 1. [0012] In addition, the audio signal coding unit 13 in Fig. 2 employs a configuration where the audio signal coding unit 13 performs coding according to, for example, MPEG-2 AAC LC (Moving Picture Experts Group phase" 2 Advanced Audio Coding Low Complexity) profile. Meanwhile, the configuration is simplified and illustrated in Fig. 2 for ease of description. [0013] The audio signal coding unit 13 in Fig. 2 includes a MDCT (Modified Discrete Cosine Transform) unit 21, a spectrum quantization unit 22, an entropy coding unit 23 and a multiplexing unit 24. [0014] The MDCT unit 21 performs MDCT of the monaural signal supplied from the channel downmix unit 11, and transforms a monaural signal which is a time domain signal, into a MDCT coefficient which is a frequency domain coefficient. The MDCT unit 21 supplies the. MDCT coefficient obtained as a result. of transform, to the spectrum quantization unit 22 as a frequency spectrum coefficient. [0015] The spectrum quantization unit 22 quantizes the frequency spectrum coefficient supplied from the MDCT unit 21, and supplies the frequency spectrum coefficient to the entropy coding unit 23. Further, the spectrum quantization unit 2 2 supplies quantization information which is information related to this quantization, to the multiplexing unit 24. The quantization information includes, for example, a scale factor and quantization bit information. [0016] The entropy coding unit 23 performs entropy coding such as Huffman coding or arithmetic coding of the quantized frequency spectrum coefficient supplied from the spectrum quantization unit 22, and losslessly compresses the frequency spectrum coefficient. The entropy coding unit 23 supplies data obtained as a result of entropy coding, to the multiplexing unit 24. [0017] The multiplexing unit 24 multiplexes the data supplied from the entropy coding unit 23 and the quantization information supplied from the spectrum quantization unit 22, and supplies resulting data to the multiplexing unit 14 (Fig. 1) as coded data. [0018] Fig. 3 is a block diagram illustrating another configuration example of the audio signal coding unit 13 in Fig. 1. [0019] In addition, the audio signal coding unit 13 in Fig. 3 employs a configuration of performing coding according to, for example, a MPEG-2 AAC SSR (Scalable Sample Rate) profile or MP3 (MPEG Audio Layer-3). Meanwhile, the configuration is simplified and illustrated in Fig. 3 for ease of description. [0020] The audio signal coding unit 13 in Fig. 3 includes an analysis filter bank 31, MDCT units 32-1 to 32-N (N is an arbitrary integer) , a spectrum quantization unit 33, an entropy. coding unit 34 and a multiplexing unit 35. [0021] The analysis filter bank 31 includes, for example, a QMF (Quadrature Mirror Filterbank) bank or a PQF (Poly-phase Quadrature Filter) bank. The analysis filter bank 31 divides the monaural signal supplied from the channel downmix unit 11, into N groups according to a frequency. The analysis filter bank 31 supplies N subband signals obtained as a result of division, to the MDCT units 32-1 to 32-N. [0022] The MDCT units 32-1 to 32-N each perform MDCT of the subband signal supplied from the analysis filter bank 31, and transforms the subband signal which is a time domain signal, into a MDCT coefficient which is a frequency domain coefficient. Further, the MDCT units 32-1 to 32-N-each supply the MDCT coefficient of each subband signal to the spectrumquantization unit 33 as the frequency spectrum coefficient. [0023] The spectrum quantization unit 33 quantizes each of the N frequency spectrum coefficients supplied from the MDCT units 32-1 to 32-N, and supplies the N frequency spectrum coefficients to the entropy coding unit 34. Further, the Spectrum quantization unit- 33 supplies quantization information about this quantization, to the multiplexing unit 35. [0024] The entropycoding unit 34 performs entropy coding such as Huffman coding or arithmetic coding of each of the quantized N frequency spectrum coefficients supplied from the spectrum quantization unit 33, and losslessly compresses the N frequency spectrum coefficients. The entropy coding unit 34 supplies N items of data obtained as a result of entropy coding, to the multiplexing unit 35. [0025] The multiplexing unit 35 multiplexes the N items of data supplied from the entropy coding unit 34 and the quantization information supplied from the spectrum quantization unit 33, and supplies resulting data to the multiplexing unit 14 (Fig. 1) as coded data. [0026] Fig. 4 is a block diagram illustrating a configuration examgle of a decoding apparatus which decodes coded data which is spatially coded by the coding apparatus 10 in Fig. 1. [0027] A decoding apparatus 40 in Fig. 4 includes an inverse multiplexing unit 41, an audio sigha decoding unit 42, a generation parameter calculation unit 43 and a stereo signal generation unit 44. The decoding apparatus 40 decodes the coded data supplied from the coding apparatus in Fig. 1, and generates a stereo signal. [0028] More specifically, the inverse multiplexing unit 41 of the decoding apparatus 40 inversely multiplexes the multiplexed coded data supplied from the coding apparatus 10 in Fig. 1, and obtains the coded-data and the BC parameter. The inverse, multiplexing unit 41 supplies the coded data to the audio signal decoding unit 42, and supplies the BC parameter to the generation parameter calculation unit 43. [0029] The audio signal decoding unit 42 decodes the coded data supplied from the inverse multiplexing unit 41, and supplies the resulting monaural signal XM which is a time domain signal, to the stereo signal generation unit 44. [0030] The generation parameter calculation unit 43 calculates generation parameters which are parameters for generating a stered) signal from a monaural signal which is a decoding result of the multiplexed coded data, using the BC parameter supplied from the" inverse multiplexing unit 41. The generation parameter calculation unit 43 supplies these generation parameters to the stereo signal generation unit' 44. [0031] The stereo signal generation unit 44 generates the left audio signal XL and the right audio, signal XR from the monaural signal XM supplied from the audio signal decoding unit 42 using the generation parameters supplied from the generation parameter calculation unit 43. The stereo signal generation unit 44 outputs the left audio signal XL and the right audio signal XR as stereo signals. [0032] Fig. 5 is a block diagram illustrating a configuration example of the audio signal decoding unit 42 in Fig. 4. [0033] In addition,' the audio signal decoding unit 42 in Fig. employs a configuration where coded data coded according to, for example, the MPEG-2 AAC LC profile is input to the decoding apparatus 40. That is, the audio signal decoding unit 42 in Fig. 5 decodes the coded data coded by the audio signal coding unit 13 in Fig. 2. [0034] The audio signal decoding unit 42 in Fig. 5 includes an inverse multiplexing unit 51, an entropy decoding unit 52, a spectrum inverse quantization unit 53 and an IMDCT unit 54. [0035] The inverse multiplexing unit 51 inversely multiplexes the coded data supplied from the inverse multiplexing unit 41 in Fig. 4, and obtains the quantized and entropy-coded frequency spectrum coefficient and the quantization information. The inverse multiplexing-unit 51 supplies the quantized and entropy-coded frequency spectrum coefficient to the entropy decoding unit 52, and supplies the quantization information to the spectrum inverse quantization unit 53. [0036] The entropy decoding unit 52 performs entropy decoding such as Huffman decoding or arithmetic decoding of the frequency spectrum coefficient supplied "from the inverse multiplexing unit 51, and restores the quantized frequency spectrum coefficient. The entropy decoding unit 52 supplies this frequency spectrum coefficient to the spectrum inverse quantization unit 53. [0037] The spectrum inverse quantization unit 53 inversely quantizes the quantized frequency spectrum coefficient supplied from the entropy decoding unit 52 based on the quantization information supplied from the inverse . multiplexing unit 51, and restores the frequency spectrum coefficient. Further, the spectrum inverse quantization unit 53 supplies the frequency spectrum coefficient to the IMDCT (Inverse MDGT) (Inverse Modified Discrete Cosine Transform) unit 54. [0038] The IMDCT unit 54 performs IMDCT of the frequency spectrum coefficient supplied from the spectrum inverse quantization unit 53, and transforms the frequency spectrum- coefficient into the monaural signal XM which is a time domain signal. The IMDCT unit 54 supplies this monaural signal XM to the stereo signal generation unit 44 (Fig. 4). [0039] Fig. 6 is a block diagram illustrating another configuration example of the audio signal decoding unit 42 in Fig. 4. [0040] In addition, the audio signal decoding unit 42 in Fig. 6 employs a contiguration where coded data coded according to,for example, the MPEG-2 AAC SSR profile or a method such as MP3 is input to the decoding apparatus 40. That is, the audio-signal decoding unit 42 in Fig. 6 decodes the coded data coded by the audio signal coding unit 13 in Fig. 3. [0 041] The audio signal decoding unit 42 in Fig. 6 includes an inverse multiplexing unit 61, an entropy decoding unit 62, a spectrum inverse quantization unit 63, IMDCT units 64-1 to 64-N and a synthesis filter bank 65. [0042] The inverse multiplexing unit 61 inversely multiplexes the coded data supplied from the inverse multiplexing unit 41 in Fig.-4, and obtains the quantized and- entropy-coded frequency spectrum coefficients of the N subband signals and the quantization information. The inverse multiplexing unit 61 supplies the quantized and entropy-coded frequency spectrum coefficients, of the N subband signals to the entropy decoding unit 62, and supplies the quantization information to the spectrum inverse quantization unit 63. [0043] The entropy decoding unit 62 performs entropy decoding such Huffman decoding or arithmetic decoding of the frequency spectrum coefficients of the N subband signals supplied from the inverse multiplexing unit 61, and supplies the frequency spectrum coefficients to the spectrum inverse quantization unit 63. [0044] The Spectrum inverse quantization unit 63 inversely quantizes each of the frequency spectrum coefficients of the N subband signals which are supplied from the entropy decoding unit 62 and which are obtained as a result of entropy decoding, based on the quantization information supplied from the inverse multiplexing unit 61. By this means, the frequency spectrum coefficients of the N subband signals are restored. The spectrum inverse quantization unit 63 supplies the restored frequency spectrum coefficients of the N subband signals to the IMDCT units 64-1 to 64-N one by one. [0045] The IMDCT units 64-1 to 64-N each perform IMDCT of the frequency spectrum coefficient supplied from the spectrum inverse quantization unit 63, and transform the frequency spectrum coefficient into a subband signal which is a time domain signal.The IMDCT units 64-1 to 64-N each supply the subband signal -obtained as a result of transform, to the synthesis filter bank 65. [0046] The synthesis filter bank 65 includes, for example, an inverse PQF and an inverse QMF. The synthesis bank 65 synthesizes the N subband signals supplied from the IMDCT units 64-1 to 64-N, and supplies the resulting signal to the stereo signal generation unit 44 (Fig. 4) as the monaural signal XM. [0047] Fig. 7 is a block diagram illustrating a configuration example of the stereo signal generation unit 44 in Fig. 4. [0048] The stereasignal generation unit 44 in Fig. 7 includes a reverb signal generation unit 71 and. a stereo synthesis unit 72. [0049] The reverb signal generation unit 71 generates a signal XD which is uncorrelated with this monaural signal XM using the monaural signal XM supplied from'the audio signal decoding unit 4.2 in Fig. 4. For the reverb signal generation unit 71, a comb filter or an all pass filter is generally used. In this case, the reverb signal generation unit 71 generates a reverb signal of the monaural signal XM as the signal XD. [0050] In addition, for the reverb signal generation unit 71, a feedback delay network (FDN) is used in some cases (see, for example, Patent Document 1). [0051] The reverb signal generation unit 71 supplies the generated signal XD to the stereo synthesis unit 72. [0052] The stereo synthesis unit 72 synthesizes the monaural signal XM supplied from the audio signal decoding unit 42 in Fig. 4 and the signal XD supplied from the reverb signal generation Unit 71 using the generation parameters supplied from the generation parameter calculation unit 43 in Fig. 4. Further, the stereo synthesis unit 72 outputs the left audio signal XL and the right audio signal XR obtained as a result of synthesis as stereo signals. [0053] Fig. 8 is a block diagram illustrating another configuration example of the stereo signal generation unit 44 in Fig. 4. [0054] The stereo signal generation unit 44 in Fig. 8 includes an analysis filter bank 81, subband stereo signal generation , units 82-1 to 82-P (,P is an arbitrary number) and a synthesis filter bank 83. [0055] In addition, when the stereo signal generation unit 44 in Fig. 4 employs the configuration illustrated in Fig. 8, the spatial parameter detection unit 12 of the coding-apparatus 10 in Fig. 1 detects the BC parameter per subband signal. [0056] More specifically, for example the spatial parameter detection unit 12 has two analysis filter banks. Further, in the spatial parameter detection unit 12, one analysis filter bank divides the stereo signal according to a frequency, and the other analysis filter bank divides the monaural signal from the channel downmix unit 11 according to a frequency. The spatial parameter detection unit 12 detects the BC parameter per subband signal based on the subband signal of the stereo signal and the subband signal of the monaural signal obtained as a result of division-. Further, the generation parameter calculation unit 43 in Fig. 4 receives a supply of the BC parameter of each subband signal from the inverse multiplexing unit 41, and generates generation parameters per subband signal. [0057] The analysis filter bank 81 includes, for example, a QMF (Quadrature Mirror Filter) bank. The analysis filter bank 81 divides the monaural signal XM supplied from the audio signal decoding unit 42 in Fig. 4 into P groups according to a frequency The analysis filter bank 81 supplies P subband signals obtained as a result of division, to the subband stereo signal generation units 82-1 to 82-P. [0058] The subband stereo signal generation units 82-1 to. 82-P each include a reverb signal generation unit and a stereo synthesis unit. The configuration of each of the subband stereo signal generation units 82-1-to 82-P is the same, and therefore only the subband stereo signal generation unit 82-B will be described. [0059] The subband stereo signal generation unit 82-B includes a reverb signal generation unit 91 and a stereo synthesis unit 92. The reverb signal generation unit 91 generates a signal XDB which is irrelevant to this subband signal XmB using the subband signal XmB of the monaural signal supplied from the analysis filter bank 81, and supplies the signal XDB to the stereo synthesis unit 92. [0060] The stereo synthesis unit 92 synthesizes the subband signal XmB supplied from the analysis-filter bank 81 and the signal XDB supplied from the reverb signal generation unit 91 using the generation parameters of the subband signal XmB supplied from the generation parameter calculation unit 43 in Fig. 4. Further, the stereo synthesis unit 92 supplies the left audio signal XLB and the right audio signal XRB obtained as a result of synthesis, to the synthesis filter bank 83 as subband signals of the stereo signals. [0061] The synthesis filter bank 83 synthesizes left and right stereo signals of each subband signal supplied from the subband stereo signal generation units 82-1 to 82-P at a time. The synthesis filter bank 83 outputs the resulting left audio signal XL and right audio signal XR as stereo signals.. [0062] In addition,.. the configuration of the stereo signal generation unit 44 in Fig. 8 is disclosed, in for example, Patent Document 2. [0063] Further, a coding apparatus which performs intensity codingmixes the frequency spectrum coefficient of each channel at a frequency equal to or more than a predetermined frequency band of the input stereo signal, and generates the frequency spectrum coefficient of the monaural signal. Further, the coding apparatus outputs a level ratio of the frequency spectrum coefficient of this monaural signal and an inter-channel frequency spectrum coefficient as a coding result. [0064] More specifically, the coding apparatus which performs intensity coding performs MDCT with respect to the stereo signal, andmixes and shares the frequency spectrumcoef ficient of each channel at a frequency equal to or more than a predetermined frequency band among resulting frequency spectrum cqefficients of channels. Further, the coding apparatus which performs intensity coding quantizes and entropy-codes the shared frequency spectrum coefficient, and multiplexes resulting data and quantization Information as coded data. Furthermore, the coding apparatus which performs intensity coding finds the level ratio of the inter-channel, frequency spectrum coefficients, and multiplexes and outputs the level ratio and the coded data. [0065] Still further, a decoding apparatus which performs intensity decoding inversely multiplexes the coded data on which the level ratio of the inter-channel frequency spectrum coefficients is multiplexed, entropy-decodes resulting coded data and inversely quantizes the coded data based on the quantization information. Moreover, the decoding apparatus which performs intensity decoding-restores the frequency spectrum coefficient of each channel based on the level ratio of the frequency spectrum coefficient obtained as' a result of inverse quantization and the inter-channel frequency spectrum coefficients multiplexed on the coded data. Moreover, the decoding apparatus whi.ph performs intensity decoding performs IMDCT of the restored frequency spectrum coefficient of each channel, and obtains a stereo signal at a frequency equal to or more than a predetermined frequency band. [0066] Although such intensity coding ratio is usually used to improve a coding efficiency, a high band frequency spectrum Coefficient of a stereo signal-is monaural coded and, represented only by an inter-channel level difference, and therefore the original stereophonic effect is slightly lost. CITATION LIST PATENT DOCUMENTS [0067] Patent Document 1: Japanese Patent Application Laid-Open No. 2006-325162 Patent Document 2: Japanese Patent Application Laid-Open No. 2006-524832 SUMMARY OF THE INVENTION PROBLEMS TO BE SOLVED BY THE INVENTION [0068] As described above, the decoding apparatus 4.0 which decodes conventional spatially coded data generates the signal Xp and signals XD1 to XDP which are irrelevant to the monaural . signal XM used upon generation of a-stereo -signal, using the monaural signal XM which is a time domain signal. [0069] Therefore, the reverb signal generation unit 71 which generates the signal XD, and the analysis filter bank 81 and the reverb signal generation units 91 of the subband stereo signal generation units 82-1 to 82-P which generate the signals XD1 to XDP cause delay, and increases algorithm delay of the decoding apparatus 40 . This causes aproblem when, for example, the decoding apparatus 40 is requested to provide immediate response performance or the decoding apparatus 40 is used in real-time communication, that is, when low delay property is important. [0070] Further, filter computation in the reverb signal generation unit 71, and the analysis filter bank 81 and the reverb.signal generation units 91 of the subband stereo signal generation units 82-1 to 82-P increases the computation amount, and also increases the required buffer capacity. [0071] In light of such a situation, the present invention can preventdelay and an increase in the computation amount upon, decoding of audio signals when multichannel audio signals are donwmixed and coded. SOLUTIONS TO PROBLEMS [0072] A speech processing apparatus according to an aspect of the present invention includes: an acquisition unit which acquires frequency domain coefficients of speech signals of channels which are generated from speech signals which are speech time domain, signals of a plurality of channels, and the number of which is less than a plurality of channels, and a parameter representing a relationship between the plurality of channels'; a first transform unit which transforms the frequency domain coefficients acquiredby the acquisition unit, into first time domain signals; a second transform unit which transforms the frequency domain coefficients acquired by the acquisition unit, into second time domain signals; and a synthesis unit which generates the speech signals of the plurality of channels by synthesizing the first time domain signals and the second time domain signals using the parameter, wherein a base of transform performed by the first transform unit and a base of transform performed by the second transform unit are orthogonal. [0073] A speech processing method and a program according to an aspect of the present invention support a speech processing apparatus according to an aspect of the present invention. [0074] According to an aspect of the present invention, frequency domain coefficients of speech signals of channels which are generated from speech signals which are speech time domain signals of a plurality of channels, and the number of which is less than a plurality of channels, and a parameter representing a relationship between the plurality of channels are acquired, the acquired frequency domain coefficients are transformed into first time domain signals, the acquired frequency domain coefficients are transformed into second time domain signals, and the speech signals of the plurality of channels are generated by synthesizing the first time domain signals and the second time domain signals using the parameter. In addition,a case of transform into the first time domain s~ig£als and a base of transform into the second time domain signals are orthogonal. [0075] The speech processing apparatus according to an aspect of the present invention may be an independent apparatus or may be an internal block which forms one apparatus. EFFECTS OF THE INVENTION [0076] According to an aspect of the present invention, it is possible to prevent delay and an increase in the computation amount upon decoding of audio signals when multichannel audio signals are downmixed and-coded BRIEF DESCRIPTION OF DRAWINGS [0077] Fig. 1 is a block diagram illustrating a configuration example of a coding apparatus which performs spatial coding. Fig. 2 is a block diagram illustrating a configuration example of can audio signal coding unit in Fig. 1. Fig.3 is a block diagram illustrating another configuration example of the audio signal coding unit in Fig. 1. Fig. 4 is a block diagram illustrating a configuration example of a decoding apparatus, which decodes spatially coded data.; Fig. 5 is a block diagram illustrating a configuration example ofan audio signal decoding unit in Fig. 4. Fig. 6 is a block diagram illustrating another; configuration example of the audio signal decoding unit in Fig. 4. Fig. 7 is a block diagram illustrating a configuration example of a stereo signal generation unit in Fig. 4. Fig. 8 is a block diagram illustrating another configuration example of the stereo signal generation unit in Fig. 4. Fig. 9 is a block diagram illustrating a configuration example of a speech processing apparatus to which the present invention is applied according to a first embodiment. Fig. 10 is a block diagram illustrating a detailed configuration example of an uncorrelated frequency-time transform unitin Fig. 9. Fig. 11 is a block diagram illustrating another detailed configuration example of the; uncorrslated frequency-time transform unit in Fig. 9. Fig. 12 is a block diagram illustrating a detailed configuration example of a stereo synthesis unit in Fig. 9. Fig. 13 illustrates a view illustrates a vector of each signal. Fig. 14 is a flowchart for describing decoding processing of the speech processing apparatus in Fig. 9. Fig. 15 is a block diagram illustrating a configuration. example of a speech processing apparatus to which the present invention is applied according to a second embodiment. Fig. 16 is a flowchart for describing decodingprocessing of the speech processing apparatus in Fig. 15. Fig. 17 is a block diagram illustrating a configuration example of a speech processing apparatus to which the present invention is applied according to a third embodiment.. Fig. 18 is a flowchart for describing decoding processing of the speech processing apparatus in Fig. 17. Fig. 19 is a block diagram illustrating a configuration examgle of a speech processing apparatus to which the present invention is applied according to a fourth embodiment. Fig. 20 is a flowchart for describing decoding processing of the speech processing apparatus in Fig. 19. Fig. 21 is a view illustrating configuration example of a computer according to an embodiment. MODE FOR CARRYING OUT THE INVENTION [0078] [Configuration Example of Speech Processing Apparatus according to Second Embodiment] Fig. 15 is a block diagram illustrating a configuration, example of a speech processing apparatus to which the present invention is applied according to' a second embodiment. [0133] The same configuration illustrated in Fig. 15 as the configuration in Fig. 9 will be assigned the same reference numerals. Overlapping description will be adequately skipped. [0134] The configuration of a speech processing apparatus 200 in Fig. 15 differs from the configuration in Fig. 9 mainly in that a band division unit 201, an IMDCT unit 202, an adder 203 and an adder 204 are additionally provided. [0135] The speech processing apparatus-200 decodes, for example, coded data for which the same spatial coding as in a coding apparatus 10 in Fig. 1 which has an audio signal coding unit 13 in Fig. 2 is performed, and on which the BC parameter of a high band is multiplexed, and stereo-codes only the monaural signal XM in a high band. [0136] More specifically, the band division unit 2 01 (division unit) of the speech processing apparatus 200 divides the frequency spectrum coefficient obtained by a spectrum inverse quantization unit 53, into two groups of high band frequency spectrum coefficients and low band frequency spectrum coefficients according to frequencies. Further, the band division unit 201 supplies; the low band-frequency spectrum coefficients to the IMDCT unit 202', and supplies the high band frequency spectrum coefficients to an uncorrelated frequency-time transform unit 102. [0137] The IMDCT unit 202 (third transform unit) performs IMDCT of the low band frequency spectrum coefficients supplied from the band division unit 201, and obtains a monaural signal XM (third time domain signal) which is a low band time domain signal. The IMDCT unit 202 supplies the low band monaural signal XMlow to the adder 203 as a low band left audio signal, and to the adder 204 as the low band right audio signal. [0138] The adder 203 receives an input of a high band left audio signal XLHigh obtained as a result of processing the high band frequency slpectrum coefficient output from the band division unit 201 in the uncorrelated frequency-time trans form unit 102 and the stereo synthesis unit 103. The adder 203 adds the highland left audio signal XLHigh and the low band monaural signal XMlow supplied from the IMDCT unit 2 02 as the low band left audio signal, and generates an entire frequency band left audio signal XL. [0139] The adder 2 04 receives an input of a high band right audio signal XRHigh obtained as a result of processing the high band frequency spectrum coefficient output from the band division unit 201 in the uncorrelated frequency-time transform unit 102 and the stereo synthesis unit 103. The adder 204 adds the high band right audio signal XRHigh and the low band monaural signal XMl0W supplied from the IMDCT unit 2 02 as the low band right audio signal, and generates an entire frequency band right audiosignal XR. [0140] [Description of Processing of Speech Processing Apparatus]; Fig. 16 is a flowchart for describing decodingprocessing of the speech processing apparatus 200 in Fig. 15. This decoding processing is started when coded data for which the same spatial coding as in the coding apparatus 10 in Fig. 1 which has the audio signal coding unit 13 in Fig. 2 is performed and on which a BC parameter of a high band is multiplexed is input to the speech processing apparatus 200. [0141] Steps S31 to S33 in Fig 16 are the same as processing in steps Sll to S13 in Fig. 14, and will not be repeatedly described. [0142] In step S34, the band division unit 201 divides frequency spectrum coefficients obtained by the spectrum/inverse quantization unit-it 53, into two groups of high band frequency spectrum coefficients and low band frequency spectrum coefficients according to frequencies. Further, the band division unit 2 01 supplies the low band frequency spectrum coefficients to the IMDCT unit 202, and supplies the high band frequency spectrum coefficients to the uncorrelated frequency-time transform unit 102. [0143] In step S35, the IMDCT unit 202 performs IMDCT of the low band frequency spectrum coefficients supplied from the band division unit 201, and obtains the monaural signal XMlow which is a low band time domain signal. The IMDCT unit 2 02 supplies the low band monaural signal XMl0W to the adder 203 3S the low -band left,audio- signal, and to the -adder 204 as the low band right audio signal; [0144] In step S36, stereo signal generation processing is performed for high band frequency spectrum coefficients supplied from the band division unit 201 by the uncorrelated frequency-time transform unit 102, the stereo synthesis unit 103, and the generation parameter calculation unit 104. More specifically, the uncorrelated frequency-time transform unit 102, the stereo synthesis unit 103 and the generation parameter calculation unit 104 perform processing in steps S14 to S18 in Fig. 14. The resulting high band left audio signal XLHigh and high band right audio signal XRHigh are input to the adder 203 and the adder 204, respectively. [0145] In step S37, the adder 203 adds the low band monaural signal XMlow supplied from the IMDCT unit 20'2 as a low band left audio signal and the high band left audio signal XLHigh supplieo?from the uncorrelated frequency-time transform unit 102 and generates an entire frequency band left audio signal XT. Further, the adder 203 outputs the entire frequency band left audio signal XL. [0146] In step S38, the adder 2 04 adds the low band monaural signal XMlow supplied from the IMDCT unit 202 as the low band right audio signal and the high band right audio signal XRHigh supplied from the uncorrelated frequency-time transform unit 102, and generates the entire frequency band right audio signal XR. Further, the adder 204 outputs this entire frequency band right audio signal XR. [0147] As described above, the speech processing apparatus 200 decodes coded data of the entire frequency band monaural signal XM, and stereo-codes only the high band. Consequently, it is possible to prevent sound from being unnatural due to stereo coding of the low band monaural signal XM. [0148] In addition, although, with the speech processing apparatus 200, the band division unit 201 divides frequency spectrum coefficients into high band frequency spectrum coefficients and low band frequency spectrum coefficients, the band division band unit 201 may divide frequency spectrum coefficients into predetermined frequency band frequency spectrum coefficients and other -frequency band frequency spectrum coefficients . That is, whether'or not stereo coding is performed may be selected depending oh whether a frequency band is a predetermined frequency band or other frequency bands instead of whether a frequency band is a low band or a high band. [0149] [Configuration Example of Speech Processing Apparatus according to Third Embodiment] Fig. 17 is a block diagram illustrating a configuration example of a speech processing apparatus to which the present invention is applied according to a third embodiment. [0150] The same configuration illustrated in Fig. 17 as the configurations in Figs. 4, 6 and 9 will be assigned the same reference numerals. Overlapping description will be adequately skipped. [0151] A configuration of a speech processing apparatus 300 in Fig. 17 differs from a configuration of a decoding apparatus 40 in Fig. 4 which has an audio signal decoding unit 42 in Fig. 6 and a stereo signal generation unit 44 in Fig. 7 mainly in that an inverse multiplexing unit 301 is provided instead of an inverse multiplexing unit 41 and an inverse multiplexing unit 61, IMDCT units 304-1 to 304-(N-l) are provided instead of IMDCT unit 64-1 to IMDCT unit 64-(N-l), a stereo coding unit 305 is provided instead of an IMDCT unit 64-N and a stereo signal generation unit 44 and a generation parameter calculation unit 104 and a synthesis filter bank 306 are provided instead of a generation parameter calculation unit 43 and a synthesis filter bank 65. [0152] 11 The.'speech processing apparatus 300 in Fig. 17 decodes, for example, coded data for which the same spatial coding as in a coding apparatus 10 in Fig. 1 which has an audio signal coding unit 13 in Fig. 3 is performed, and on which a BC parameter of a predetermined subband signal -is multiplexed. [0163] More specifically, the inverse multiplexing unit 301 of the speech processing apparatus 30 0 corresponds to the inverse multiplexing unit 41 in Fig. 4 and the inverse multiplexing unit 61 in Fig. 6. That is, the inverse multiplexing unit 301 receives an input of coded data for which the same spatial coding as in the coding apparatus 10 in Fig. 1 which has the audio signal coding unit 13 in Fig. 3 is performed, and in which a BC parameter of a predetermined subband signal is multiplexed. The inverse multiplexing unit 301 inversely multiplexes the input coded data, and obtains the coded data and the BC parameter of the predetermined subband signal. Further, the inverse, multiplexing unit -301 supplies the BC parameter of the predetermined subband signal to the generation parameter calculation unit 104. [0154] Furthermore, the inverse multiplexing unit 301 inversely multiplexes the coded data, and obtains quantized and entropy-coded frequency spectrum coefficients of N subband signals and quantization information. The inverse multiplexing unit 301 supplies the quantized and entropy-coded-frequency spectrum coefficients of the N subband signals to the entropy decoding unit 62, and supplies the quantization information to the spectrum inverse quantization unit 63. [0155] The IMDCTunits 304-1 to 304- (N-l) (third transformunit) and the stereo coding unit 3 05 receive an input of the frequency spectrum coefficients of the N subband signals restored by the spectrum inverse quantization unit 63 one by one. [0156] The IMDCT units 304-1 to 304--(N-l) each perform IMDCT of the input frequency spectrum coefficient, and transform the frequency spectrum coefficient into a subband.signal XMI (i = 1, 2, ... and N-l) of the monaural signal XM which is a time domain signal. The IMDCT units 304-1 to 304-(N-l) each supply the subband signal XM1 to the synthesis filter bank 306 as a left audio signal XL1 and a right audio signal XR1. [0157] The stereo coding unit 305 includes an uncorrelated frequency-time transform unit 102 and a stereo synthesis unit 103 in Fig. 9. The stereo coding unit 305 generates a subband signal XLA of a left audio signal and a subband signal XRA of a right audio signal which are time domain signal, from frequency spectrum coefficients of the, predetermined subband signal input from the spectrum inverse quantization unit 63, using the generation parameters generated by the generation parameter calculation unit 104. Further, the stereo coding unit 305 supplies the left subband signal XLA and the right subband signal XRA to the synthesis filter bank 306. [0158] The synthesis filter bank 306 (addition unit) includes a left synthesis filter bank for synthesizing a subband signal of a left audio signal, and a right synthesis filter bank for synthesizing a subband signal of a right audio signal. The left synthesis filter bank of the synthesis filter bank 306 synthesizes left subband signals XL1 to XLN_1 from the IMDCT units|304-l to 304-(N-l) , and the left subband signal XLA from the stereo'coding unit 305. Further, theleft synthesis filter bank outputs the entire frequency band left audio signal XL obtained as a result of synthesis. [0159] Furthermore, the right synthesis filter bank of the synthesis filter bank 30 6 synthesizes right subband signals. XR1 to XRN-1 from the IMDCT units 304-1 to 304-(N-l), and the right subbah'd signal XRA from the stereo coding unit 305. Still further, the right synthesis filter bank outputs the entire frequency band right audio signal XR obtained as a result of synthesis. [0160] In addition, although the speech processing apparatus 300 in Fig. 17 stereo-codes one subband signal alone, the speech processing apparatus 300 can stereo-codes a plurality of subband signals. Further, a subband signal which is stereo-coded may be dynamically set on a coding side instead of being set in advance. In this case,for example, information for specifying a subband signal which is a stereo coding target is included in a BC parameter. [0161] [Description of Processing of Speech Processing Apparatus] Fig. 18 is a flowchart for describing decodingprocessing of the speech processing apparatus 300 in Fig. 17. This decoding processing is started when, for example, coded data for which the same spatial coding as in the coding apparatus 10 in Fig. 1 which has the audio signal coding unit 13 in Fig. 3 is performed, and on which a BC parameter of a predetermined subband signal ismultiplexed is input to the speechprocessing apparatus 300. [0162] In step S51 in Fig. 18, the inverse multiplexing unit 301 inversely multiplexes the input multiplexed coded data, and obtains the coded data and the BC parameter of the predetermined subband signal. Further, the inverse multiplexing unit 301 supplies the BC parameter of the predetermined subband signal to the generation parameter calculation unit 104 . Furthermore, the inverse multiplexing unit 301 inversely multiplexes the coded data, and obtains quantized and entropy-coded frequency-spectrum coefficients of N subband signals and quantization information. The inverse multiplexing unit 301 supplies the quantized and entropy-coded frequency spectrum coefficients of the N subband signals to the entropy decoding unit 62, and supplies the quantization information to the spectrum inverse quantization unit 63. [0163] Instep S52,the..entropy decoding unit 62 entropy-decodes the frequency spectrum coefficients of the N subband signals supplied from the inverse multiplexing unit 101, and supplies the frequency spectrum coefficients to the spectrum inverse quantization unit 63. [0164] In step S53, the spectrum inverse quantization unit 63 inversely quantizes the frequency spectrum coefficients of the N subband signals supplied from the entropy decoding unit 62 and obtained as a result of entropy decoding, based on the quantization information supplied from the inverse multiplexing unit 301. Further, the spectrum inverse quantization unit 63 supplies the resulting restored frequency spectrum, coefficients of the N subband signals, to the IMDCT unites 304-1 to 304-(N-l) and the stereb coding unit 305 one by one. [0165] In step S54, the IMDCT units 304-1 to 304(N-l) each perform IMDCT of the frequency spectrum coefficient supplied fronuthe spectrum inverse quantization unit 63. Further, the IMDCT units 304-1 to 304-(N-l) each supply the resulting subband signal XKi (i = 1, 2, ... and N-l) of a monaural signal to the synthesis filter bank 306 as the subband signal XLi of the left audio signal and the suibband signal XLI of the right audio signal. [0166] In step S55, the stereo coding unit 305 performs stereo signal generation processing of the frequency spectrum coefficient of a predetermined subband signal supplied from the spectrum inverse quantization unit 63, using the generation parameters supplied from the generation parameter calculation unit 104. Further,the stereo coding unit 305-supplies the resulting subband signal XLA of the left audio signal and subband signal XRA of the right audio signal which are time domain signals, to the synthesis filter bank 306. [0167] In step S56, the left synthesis filter bank of the synthesis filter bank 306 synthesizes all subband signals of left audio signals supplied from the IMDCT units 304-1 to 304-(N-l) and the stereo coding unit 305, and generates the entire frequency band left audio signal XL. Further, the left synthesis filter bank outputs this entire frequency band left audio signal XL. [0168] In step S57, the right synthesis filter bank,of the synthesis filter bank 306 synthesizes all subband signals of right audip signals supplied from the IMDCT units 304-1 to 304-(N-l) and the stereo coding unit 305, and generates the entire frequency band right audio signal XR. Further, the right synthesiss filter bank outputs this entire frequency band right audio signal XR. [0169] [Configuration Example of Speech Processing Apparatus according to Fourth Embodiment] Fig. 19 is a block diagram illustrating a configuration example of a speech processing apparatus to which the present invention is applied according to a fourth embodiment. [0170] The same configuration illustrated in Fig. 19 as the configuration in Fig. 15 will be assigned the same reference numerals. Overlapping description will be adequately skipped. [0171] The configuration of a speech processing apparatus 400 in Fig. 19 differs from the configuration in Fig. 15 mainly in that a spectrum separation unit 401 is provided instead of a band division unit 201, IMDCTs 402 and 403 are provided instead of an IMDCT unit 202, and an adder 404 and an adder 405 are provided instead of an adder 203 and an adder 204. [0172] The speech processing apparatus 400 decodes coded data for which intensity coding is performed, and on which a BC parameter at a frequency equal to or more than an intensity start frequency Fis is multiplexed instead of a conventional levellratio of inter-channel frequency spectrum coefficients. [0173] That is, the coded data decoded by the speech processing apparatus 400 is generated by a coding apparatus which detects the BC parameter by, for example, downmixing a coding target stereo signal to a monaural signal XM and extracting the resulting monaural signal XM and a component at a frequency equal to or more than the intensity start frequency Fis of the coding target stereo signal by means of, for example, a bypass filter. [017 4] The spectrum separation unit 401 (separation unit) of the speechprocessing apparatus 400 obtains frequency spectrum coefficients restored by a spectrum inverse quantization unit 53 . The spectrum separation unit 401 separates this frequency spectrum coefficient into a frequency spectrum coefficient of a stereo signal at a frequency lower than the intensity start frequency Fis and a frequency spectrum coefficient of a monaural-signal XMhigh at a frequency equal to or more than the intensity start frequency Fis. The spectrum separation unit 401 supplies the frequency spectrum coefficient of the left audio ;signal XLl0W of the stereo signal at a frequency lower than the intensity start frequency Fis, to the IMDCT. unit 402, and supplies the frequency spectrum coefficient of the right audio signal XRlow to the IMDCT unit 403. Further, the spectrum separation unit 401 supplies the frequency spectrum coefficient of the monaural signal XMhigh to an uncorrelated frequency-time transform unit 102. [0175] The IMDCT unit 402 (third trans form unit) performs IMDCT of the frequency spectrum coefficient of the left audio signal XLlow Supplied from the spectrum separation unit 401, and supplies the resulting left audio signal XLl0W to the adder 404. [0176] The IMDCTunit 403 (third trans form unit) per/forms IMDCT of the frequency spectrum coefficient of the right audio signal XRlow supplied from the spectrum separation unit 401, and supplies the resulting right audio signal XRlow to the adder 405. [0177] The adder 404 (addition unit) adds the left audio signal XLhigh which is generated by the stereo synthesis unit 103 and which is a time domain signal at a frequency equal to or more than an intensity start frequency Fis, and the left audio signal XLlow supplied from the IMDCT unit 402. The adder 404 outputs the resulting audio signal as the entire frequency band left audio signal XL. [0178] The adder 405.(addition unit:) adds the righ taudio signal XRhigH which is generated by thestereo synthesis unit 103 and which is a time domain signal at a frequency equal to or more than the intensity start frequency Fis, and the right audio signal XRlow supplied from the IMDCT unit 402. The adder 405 outputs the resulting audio signal as the entire frequency band right audio signal XR. [0179] As described above, the speech processing apparatus 400 stereo-codes a component of the frequency equal to or more than the intensity start frequency Fis monaural-coded by intensity coding, using the BC parameter multiplexed on intensity-coded data. Consequently, it is possible to restore a stereophonic effect of the component of the frequency . equal to or more than the intensity start frequency Fis compared-to an intensity decoding apparatus which performs stereo-coding using a conventional level ratio of inter-channel frequency spectrum coefficients. [0180] [Description of Processing of Speech Processing Apparatus] Fig. 20 is a flowchart for describing decodingprocessing of the speech processing apparatus 400 in Fig. 19. This decoding processing is started when, for example, coded data which is intensity coded and on which the BC parameter of the frequency equal to or more than the intensity start frequency Fis is multiplexed is input. [0181] Processing in steps S71 to S73 in Fig. 20 are the same as the processing in steps S31 to S33 in Fig. 16, and therefore will not be described. [0182] In step S74, the spectrum separation unit 401 separates the frequency spectrum coefficients restored by the spectrum inverse quantization unit 53 into frequency spectrum coefficients of stereo signals at a frequency lower than the intensity start frequency Fis and the frequency spectrum coefficient of the monaural signal XMhigh at a frequency equal to or more than the intensity start frequency Fis . The spectrum separation unit 401 supplies the frequency spectrum coefficient of the left audio signal XLl0W of the stereo signal at a frequency lower than the intensity start frequency Fis, to the IMDCT unit 402, and the frequency spectrum coefficient of the right audip signal XRlow to the IMDCT unit 4 03. Further, the spectrum separation unit 4 01 supplies the frequency spectrum coefficient of the monaural signal XMhlgh to the uncorrelated frequency-time transform unit 102. [0183] In step S75, the IMDCT unit 402 performs IMDCT of the , frequency spectrum coefficient of the left audio signal XLlow supplied from the spectrum separation unit 401. Further, the IMDCT unit 402 supplies the resulting left audio signal XLlow to the adder 404. [0184] In step S76, the IMDCT unit 403 performs IMDCT of the frequency spectrum coefficient of the right audio signal XRl0W supplied from the spectrum separation unit 401. Further, the IMDCT unit 403 supplies the resulting right audio signal XRlow to the adder 405. [018.5] In step S77, the uncorrelated frequency-time transform unit 102, the stereo synthesis unit 103 and the generation parameter calculation unit 104 perform stereo signal generation processing of the frequency spectrum coefficient of the monaural signal XMhagh from the spectrum separation unit 401. The resulting left audio signal XLhigh which is a time domain signal is supplied to the adder 404, and the right audio signal XRhigh is supplied to the adder 405. [0186] In s'tfep S78, the adder 404 adds the left audio signal XLl0W at a frequency lower than the intensity start frequency. Fis from the IMDCT unit 402 and the left audio signal XLhigh at a frequency equal to or more than the intensity start frequency Fis from the stereo synthesis unit 103, and generates the entire frequency band left!:audio signal XL. Further, the adderH04 outputs this left audio signal XL. [0187] In step S79, the adder 405 adds the right audio signal XRlow at a frequency lower than the intensity start frequency Fis from the IMDCT unit 403 and the right audio signal XRhigh at a frequency equal to or more than the intensity start frequency Fis from the stereo synthesis unit 103, and generates the entire frequency band right audio signal XR. Further, the adder 4'05 outputs this right audio signal XR. [0188] In addition, although, with the, above description, a speech processing apparatus 100 (200, 300 and 400) decodes coded data which is time-frequency transformed by MDCT, and therefore IMDCT is performed upon frequency-time transform, IMDST is performed upon frequency-time transform when coded data, which is time-frequency transformed by MDST is decoded. [0189] Further, although, with the above description, the uncorrelated time-frequency transform unit 102 uses IMDCT transform and IMDST transform where bases are orthogonal to each other, other lapped orthogonal transform such as sine transform or cosine transform may be used. [0190] [Description of Computer to which Present Invention is applied] Next; a series or the above processing can be executed by hardware or by software. When a series of the processing are executed by software, a program configuring this software is installed to, for example, a general-purpose computer. [0191] Fig. 21 illustrates a configuration example of a computer in which a program for executing a series of the above processing are installed according to an embodiment. [0192] The program can be recorded in advance in a memory unit 508 or a ROM (Read Only Memory) 502 which is a recording medium , built in the computer. [0193] Alternatively, the program can be stored (recorded) in a removable media 511. This removable media 511 can beprovided as so-called package software. Meanwhile, the removable media 511 includes, for example, afiexible disc, a CD-ROM (Compact Disc Read Only Memory) , a MO (Magneto Optical) disc, a DVD (Digital Versatile Disc), a magnetic disc and a semiconductor memory. [0194] In addition, the program can be installed to a computer from the above removable media 511 through a drive 510, and, in addition, may be downloaded to a computer through a communication network or a broadcasting network-or installed in the built-in memory unit 508. That is, the program can be wirelessly transferred, for example, from a download site to a computer through a digital satellite broadcasting satellite, or can be transferred to a computer by way of a wire through a network such as LAN (Local Area Network) or Internet. [0195] The computer has abuilt-in CPU (Central Processing Unit) 501, and the CPU 501 is connected with an input/output interface 505 through a bus 504. [0196] The CPU 50.1 executes the program stored in the ROM 502 according to a command when receiving, an input of the command according, to, for example, a user' s operation of an input unit 506 through the input/output interface 505. Alternatively, the CPU 501 loads the program stored in the memory unit 508 to a RAM ( Random Access Memory) 503 and executes the program. [0197] Thus, the CPU 501 executes processing according to the above flowchart or processing executed by the configuration in the above block diagram. Further, the CPU 501 outputs this processing result from an output unit 507 through the input/output interface 505, transmits the processing result from a communication unit 509 or records the processing result in the memory unit 508. [0198] In addition, the input unit 506 includes a keyboard, a mouse or a microphone . Further, the output unit 507 includes a LCD (Liquid Crystal Display) or speakers. [0199] Meanwhile, in this descriptions-processing executed by the computer according to the program does not necessarily need to be executed in a chronological order disclosed as a flowchart. That is, the processing executed by the computer according to the program include processing (such as parallel processing or processing by an object) executed in parallel or individually. [0200] Further, the program may be processed by one computer (processor) or processed in a distributedmanner by aplurality of computers. Furthermore, the program may be transferred to a distant computer and executed. [0201] (The present invention is applicable to a pseudo stereo coding technique for audio signals. [0202] The embodiments of the present invention are by no means limited to the above embodiments, and can be variously modified , within a scope which does not deviate from the spirit of the present invention. REFERENCE SIGNS LIST [0203] 54 IMDCT unit 100 Speech processing apparatus 101 Inverse multiplexing unit 103 Stereo synthesis unit 111 IMDST unit 121 Spectrum inversion unit 122 IMDCT unit 123 Sign inversion unit 200 Speech processing apparatus 201 Band division unit 202 . IMDCT unit 203, 204 Adder 300 Speech processing apparatus 301 Inverse multiplexing unit 304-1 to 304-N IMDCT unit 305 Stereo coding unit 306 Synthesis filter bank 400 Speech processing apparatus 401 Spectrum separation unit 402, 403 IMDCT unit 404, 405 Addar CLAIMS 1. A speech processing apparatus comprising: an acquisition unit which acquires frequency domain coefficients of speech, signals of channels which are generated from speech signals which are speech time domain signals of a plurality of channels, and the number of which is less than a plurality of channels, and a parameter representing a relationship between the plurality of channels; a first transform unit which transforms the frequency domain coefficients acquired by the acquisition unit, into first time domain signals; a second transform unit which transforms the frequency domain coefficients acquired by the acquisition unit, into second time domain signals; and a synthesis unit which generates the speech signals of the plurality of channels by synthesizing the first time domain signals and the second time domain signals using the parameter, wherein a base of transform performed by the first transform ilnit and a, base of transform performed by the second transform unit are orthogonal. 2. The speech processing apparatus according to claim 1, further comprising: a division unit which divides the frequency domain coefficients acquired by the acquisition unit, into a plurality of groups according to a frequency; a third transform unit which transforms the frequency domain coefficients divided into a first group among the plurality of groups, into third time domain signals; and an addition unit which adds the third time domain signals which are speech signals of respective channels in a frequency-band of the first group and the speech signals of the plurality of channels generated by the synthesis unit per channel, and generates the speech signals of the plurality of channels in an entire frequency band, wherein. the acquisition unit acquires the frequency domain coefficients and the parameter in a frequency band of a second group which" is a group other than the first group, the first transform unit transforms the frequency domain coefficients divided into the second group, into the first time domain- signals, the .second transform unit transforms the frequency domain coefficients divided into the second group, into the second time domain signals, and the synthesis unit generates the speech signals of the plurality of channels in the frequency band of the second group by synthesizing the first time domain signals and the second time .domain signals using the parameter. 3. A speech processing apparatus according to claim 1, further comprising: a third tr ans form unit which transforms frequency domain coef ficifents of/ a first group among the frequency domain coefficients acquired by the acquisition unit and divided into a plurality of groups according to a frequency, into third time domain signals; and an addition unit which adds the third time domain signals which are speech signals of respective channels in the frequency band of the first group and the speech signals of the plurality of channels generated by the synthesis unit per channel, and generates the speech signals of the plurality of channels in an entire frequency band, wherein the acquisition unit acquires the frequency domain coefficients of each group and the parameter of a frequency band of a second group which is a group other than the first group among the plurality of groups the first trans form unit transforms the frequency domain coefficients divided into the second group, into the first time domain;-signals, the second transform unit transforms the frequency domain coefficients divided into the second group, into the second time domain signals, and the synthesis unit generates the speech signals of the plurality of channels in a frequency band of the second group by synthesizing the first time domain signals and the second time domain signals using the parameter. 4. The speech processing apparatus according to claim 1, wherein the frequency domain coefficients are generated from frequency domain coefficients of the speech signals of the plurality of channels. 5. A speech processing apparatus according to claim 4, further comprising: a separation unit which separates the frequency domain coefficients in a predetermined frequency band acquired by the acquisition unit, and the frequency domain coefficients of the speech signals of a plurality of channels in a frequency band other than the predetermined frequency band; a third transform unit which transforms the frequency domain coefficients of the speech signals of the plurality of channels separated by the separation unit, into third time domain signals of the plurality of channels; and an addition unit which adds the third time domain signals of the plurality of channels which are the speech signals of the plurality of channels in the frequency band other than the predetermined frequency band and the speech signals" of the plurality of. channels generated by the synthesis unit, and generates the speech signals of the plurality of channels in an entir,e frequency band, wherein the acquisition unit acquires the frequency domain coefficients in the predetermined frequency band, the frequency domain coefficients of the speech signals of the plurality of channels in the frequency band other than the predetermined frequency band, and the parameter in the predetermined frequency band, the first transform unit transforms the frequency domain coefficients in the predetermined frequency band separated by the separatioh unit, into the first time domain signals, the second transform unit, transforms the frequency domain coefficients in the predetermined frequency band separated by the separation unit, into the second time domain signals, and the synthesis unit generates the speech signals of the plurality of channels in the predetermined frequency band by synthesizing the first time domain signals and the second time domain signals using the parameter. 6. The speech processing apparatus according to any one of claims 1 to 5, wherein the frequency domain coefficients are MDCT (Modified Discrete Cosine Transform) coefficients, trans form performed by the first transform unit is IMDCT (Inverse Modified Discrete Cosine Transform), and trans formperformed by the second transform unit is IMDST (Inverse Modified Discrete Sine Transform). 7. The speech processing apparatus according to any one of claims 1 to 5, wherein the second transform unit comprises: a spectrum inversion unit which inverts the frequency domain coefficients such that frequencies are in an inverse order; an IMDCT unit which obtains time domain signals by perf orming lMDCT (Inverse Modified Discrete Cosine Transform) of the frequency domain coefficients obtained as a result of inversion by the spectrum inversion unit; and a sign inversion unit which inverts a sign of each sample of the time domain signals obtained by the IMDCT unit every other sign, and; the frequency domain coefficients are MDCT (Modified Discrete Cosine Transform) coefficients, and transform performed by the first transform unit is IMDCT. 8. A speech signal processing method to be performed by a speech processing apparatus, the method comprising: an acquisition step of acquiring frequency domain coefficients of speech signals of channels which are generated from speech, signals which are speech time domain signals of a plurality of channels, and the number of which is less than a plurality of channels, and a parameter representing a relationship between the plurality of channels; a first transform step of transforming the frequency domain coefficients acquired by processing in the acquisition step, into first time domain signals; a second transform step of transforming the frequency-domain coefficients acquired by processing in the acquisition step, into second time domain signals; and a synthesis step of generating the speech signals of the plurality of channels by synthesizing the first time domain signals and the second time domain signals using the parameter, wherein a base of transform in processing in the first transform step and a base of transform in processing in the second transform step are orthogonal. 9. A program for causing a computer to execute: an acquisition step of acquiring frequency domain coefficients of speech signals of channels which are generated from speech signals which are speech time domain signals of a plurality of channels, and the number of which is less than a plurality of channels, and a parameter representing a relationship between the plurality of channels; a first transform step of transforming the frequency domain coefficients acquired by processing in the acquisition step, into first time domain signals; a second transform step of transforming the frequency domain coefficients acquired by processing in the acquisition step/' into second time domain signals; and a synthesis step of generating the speech signals of the plurality of channels by synthesizing the first time domain signals and the second time domain signals using the parameter, wherein a base of transform in processing in the first transform step and a base of transform in processing in the second transform step are orthogonal.

Documents