JP3387084B2

JP3387084B2 - Recording medium, audio decoding device

Info

Publication number: JP3387084B2
Application number: JP32595799A
Authority: JP
Inventors: 美昭田中; 昭治植野; 徳彦渕上
Original assignee: Victor Company of Japan Ltd
Current assignee: Victor Company of Japan Ltd
Priority date: 1998-11-16
Filing date: 1999-11-16
Publication date: 2003-03-17
Anticipated expiration: 2019-11-16
Also published as: JP2000214894A

Description

Detailed Description of the Invention

【０００１】[0001]

【発明の属する技術分野】本発明は、マルチチャネルの
音声信号を圧縮して記録した記録媒体及び音声復号装置
に関する。BACKGROUND OF THE INVENTION 1. Field of the Invention The present invention relates to a recording medium in which a multi-channel audio signal is compressed and recorded, and an audio decoding device .

【０００２】[0002]

【従来の技術】音声信号を圧縮する方法として、本発明
者は先の出願（特願平９−２８９１５９号）において１
チャネルの原デジタル音声信号に対して、特性が異なる
複数の予測器により時間領域における過去の信号から現
在の信号の複数の線形予測値を算出し、原デジタル音声
信号と、この複数の線形予測値から予測器毎の予測残差
を算出し、予測残差の最小値を選択する予測符号化方法
を提案している。2. Description of the Related Art As a method for compressing an audio signal, the present inventor has described in 1
For the original digital audio signal of the channel, multiple predictors with different characteristics are used to calculate multiple linear prediction values of the current signal from past signals in the time domain, and the original digital audio signal and the multiple linear prediction values. We propose a predictive coding method that calculates the prediction residual for each predictor and selects the minimum value of the prediction residual.

【０００３】なお、上記方法では原デジタル音声信号が
サンプリング周波数＝９６ｋＨｚ、量子化ビット数＝２
０ビット程度の場合にある程度の圧縮効果を得ることが
できるが、近年のＤＶＤオーディオディスクではこの２
倍のサンプリング周波数（＝１９２ｋＨｚ）が使用さ
れ、また、量子化ビット数も２４ビットが使用される傾
向がある。また、マルチチャネルにおけるサンプリング
周波数と量子化ビット数はチャネル毎に異なることもあ
る。In the above method, the original digital audio signal has a sampling frequency = 96 kHz and a quantization bit number = 2.
Although it is possible to obtain a certain degree of compression effect when the number of bits is about 0, in recent DVD audio discs, this 2
The double sampling frequency (= 192 kHz) is used, and the number of quantization bits tends to be 24 bits. Further, the sampling frequency and the number of quantization bits in the multi-channel may be different for each channel.

【０００４】[0004]

【発明が解決しようとする課題】ところで、予測符号化
方式のような圧縮方式は圧縮率が可変（ＶＢＲ：バリア
ブル・ビット・レート）であるので、マルチチャネルの
音声信号を予測符号化するとチャネル毎のデータ量が時
間的に大きく変化する。また、このようなデータを伝送
する場合には、チャネル毎にパラレルではなくデータス
トリームとして伝送される。したがって、再生側（デコ
ード側）においてこのような可変長のデータストリーム
をチャネル毎に同期して再生（プレゼンテーション）可
能にする必要がある。By the way, since a compression method such as a predictive coding method has a variable compression rate (VBR: variable bit rate), when predictive coding a multi-channel audio signal, The data amount of changes greatly with time. Further, when transmitting such data, it is transmitted not as parallel for each channel but as a data stream. Therefore, it is necessary for the reproducing side (decoding side) to be capable of reproducing (presenting) such a variable-length data stream in synchronization with each channel.

【０００５】そこで本発明は、マルチチャネルの音声信
号を圧縮符号化する場合に再生側の復号効率を改善する
ことができるデータを記録した記録媒体及び音声復号装
置を提供することを目的とする。Therefore, the present invention is directed to a recording medium and a voice decoding device which record data capable of improving the decoding efficiency on the reproducing side when compressing and coding a multi-channel voice signal.
The purpose is to provide storage .

【０００６】本発明は上記目的を達成するために、以下
の１）〜４）の手段からなるものである。すなわち、In order to achieve the above object, the present invention comprises the following means 1 ) to 4) . That is,

【０００７】１）あるサンプリング周波数と量子化ビッ
ト数のマルチチャネルの音声信号を、そのままのチャネ
ル又は互いに相関をとったチャネル毎に、入力される音
声信号に応答して先頭サンプル値を得ると共に、特性が
異なる複数の線形予測方法により時間領域の過去から現
在の信号の線形予測値がそれぞれ予測され、その予測さ
れる線形予測値と前記音声信号とから得られる予測残差
が最小となるような線形予測方法を選択して予測符号化
するステップと、前記ステップにより選択されたチャネ
ル毎の線形予測方法と予測残差と所定の先頭サンプル値
を含む予測符号化データを格納するサブパケットと、再
生側において元のアナログ音声信号に復元される際に用
いられるサンプリング周波数及び量子化ビット数を含む
同期情報部と、を有するデータ構造にフォーマット化す
るステップとにより、前記データ構造にフォーマット化
されたデータが記録され、その記録されたデータうち前
記予測符号化データは元の音声信号を復元するために用
いられる予測値を算出するためのデータとして記録され
ていることを特徴とする記録媒体。２）請求項１記載のサブパケット及び同期情報部と、更
にＳＣＲ情報を含むパックヘッダと、を含んで一つのパ
ックとしてフォーマット化されて記録されると共に、前
記ＳＣＲ情報は、前記パックを再生する際の時間管理情
報として用いられることを特徴とする記録媒体。３）請求項１記載の記録媒体に記録されたデータから元
のマルチチャネルの音声信号を復号する音声復号装置で
あって、前記データ構造をサブパケットと同期情報部に
分離する手段と、前記サブパケット内の圧縮データをチ
ャネル毎に伸長する伸長手段と、前記チャネル毎に伸長
された音声データから元のマルチチャネルの伸長された
音声信号に変換する手段と、前記変換されたマルチチャ
ネルの音声データを前記同期情報部内のサンプリング周
波数及び量子化ビット数に基づいてアナログ音声信号に
変換する手段とを、有する音声復号装置。４）請求項２記載の記録媒体に記録されたデータから元
のマルチチャネルの音声信号を復号する音声復号装置で
あって、前記ヘッダに含まれるＳＣＲ情報を分離する第
１の分離手段と、前記分離されたＳＣＲ情報に基づいて
前記サブパケット及び同期情報部を保持するためのバッ
ファと、前記バッファに保持された前記サブパケットと
同期情報部とを分離する第２の分離手段と、前記同期情
報部内の識別子に基づいて前記サブパケット内の圧縮デ
ータをチャネル毎に伸長する伸長手段と、前記チャネル
毎に伸長された音声データから元のマルチチャネルの伸
長された音声信号に変換する手段と、前記変換されたマ
ルチチャネルの音声データを前記同期情報部内のサンプ
リング周波数及び量子化ビット数に基づいてアナログ音
声信号に変換する手段とを、有する音声復号装置。1) A multi-channel audio signal having a certain sampling frequency and a quantized bit number is obtained in response to an input audio signal for a channel as it is or for each channel correlated with each other, and a head sample value is obtained. Linear prediction values of the current signal are predicted from the past in the time domain by a plurality of linear prediction methods having different characteristics, and the prediction residual obtained from the predicted linear prediction value and the speech signal is minimized. A step of selecting a linear prediction method for predictive coding, a linear prediction method for each channel selected in the step, a subpacket for storing prediction coded data including a prediction residual and a predetermined head sample value, and reproduction A synchronization information part including a sampling frequency and a quantization bit number used when the original analog audio signal is restored on the side, Formatting into a data structure to record the formatted data in the data structure, the predictive coded data of the recorded data is a prediction value used to restore the original audio signal. Recorded as data for calculation
Recording medium characterized Tei Rukoto. 2) The sub-packet and the synchronization information part according to claim 1 and a pack header including SCR information are further formatted and recorded as one pack, and the SCR information reproduces the pack. A recording medium characterized by being used as time management information at the time. 3) A voice decoding device for decoding an original multi-channel voice signal from the data recorded on the recording medium according to claim 1, said means for separating said data structure into a sub-packet and a synchronization information part; Decompression means for decompressing the compressed data in the packet for each channel, means for converting the decompressed audio data for each channel into the original decompressed multi-channel audio signal, and the converted multi-channel audio data To an analog audio signal based on the sampling frequency and the number of quantization bits in the synchronization information section. 4) An audio decoding device for decoding an original multi-channel audio signal from the data recorded on the recording medium according to claim 2, wherein the first separating means separates the SCR information contained in the header; A buffer for holding the sub-packet and the synchronization information part based on the separated SCR information; a second separating means for separating the sub-packet and the synchronization information part held in the buffer; and the synchronization information. Decompressing means for decompressing the compressed data in the subpacket for each channel based on the identifier in the section; means for converting the decompressed audio data for each channel into an original multichannel decompressed audio signal; The converted multi-channel audio data is converted into an analog audio signal based on the sampling frequency and the number of quantization bits in the synchronization information section. A stage, having speech decoding apparatus.

【０００８】[0008]

【発明の実施の形態】以下、図面を参照して本発明の実
施の形態を説明する。図１は本発明が適用される音声符
号化装置及び音声復号装置の第１の実施形態を示すブロ
ック図、図２は図１の符号化部を詳しく示すブロック
図、図３は図１、図２の符号化部により符号化されたビ
ットストリームを示す説明図、図４はＤＶＤのパックの
フォーマットを示す説明図、図５はＤＶＤのオーディオ
パックのフォーマットを示す説明図、図６は図５のオー
ディオデータエリアのフォーマットを詳しく示す説明
図、図７は図１の復号化部を詳しく示すブロック図、図
８は図７の入力バッファの書き込み／読み出しタイミン
グを示すタイミングチャート、図９はアクセスユニット
毎の圧縮データ量を示す説明図、図１０はアクセスユニ
ットとプレゼンテーションユニットを示す説明図であ
る。BEST MODE FOR CARRYING OUT THE INVENTION Embodiments of the present invention will be described below with reference to the drawings. 1 is a block diagram showing a first embodiment of a speech coding apparatus and speech decoding apparatus to which the present invention is applied , FIG. 2 is a block diagram showing the coding unit of FIG. 1 in detail, and FIG. 3 is FIG. 2 is an explanatory diagram showing a bitstream encoded by the encoding unit 2; FIG. 4 is an explanatory diagram showing a DVD pack format; FIG. 5 is an explanatory diagram showing a DVD audio pack format; and FIG. FIG. 7 is an explanatory diagram showing the format of the audio data area in detail, FIG. 7 is a block diagram showing the decoding section of FIG. 1 in detail, FIG. 8 is a timing chart showing the write / read timing of the input buffer of FIG. 7, and FIG. 9 is for each access unit. FIG. 10 is an explanatory diagram showing the amount of compressed data, and FIG. 10 is an explanatory diagram showing an access unit and a presentation unit.

【０００９】ここで、マルチチャネル方式としては、例
えば次の４つの方式が知られている。（１）４チャネル方式ドルビーサラウンド方式の
ように、前方Ｌ、Ｃ、Ｒの３チャネル＋後方Ｓの１チャ
ネルの合計４チャネル（２）５チャネル方式ドルビーＡＣ−３方式のＳ
Ｗチャネルなしのように、前方Ｌ、Ｃ、Ｒの３チャネル
＋後方ＳＬ、ＳＲの２チャネルの合計５チャネル（３）６チャネル方式ＤＴＳ（Digital Theater
System）方式や、ドルビーＡＣ−３方式のように６チャ
ネル（Ｌ、Ｃ、Ｒ、ＳＷ（Ｌｆｅ）、ＳＬ、ＳＲ）（４）８チャネル方式ＳＤＤＳ（Sony Dynamic D
igital Sound）方式のように、前方Ｌ、ＬＣ、Ｃ、Ｒ
Ｃ、Ｒ、ＳＷの６チャネル＋後方ＳＬ、ＳＲの２チャネ
ルの合計８チャネルHere, as the multi-channel method, for example, the following four methods are known. (1) 4-channel method Like the Dolby Surround method, a total of 4 channels of 3 channels of front L, C, and R + 1 channel of rear S (2) 5-channel method S of Dolby AC-3 method
A total of 5 channels (3 channels of front L, C, and R + 2 channels of rear SL and SR) (3) 6-channel system DTS (Digital Theater)
System) and 6 channels (L, C, R, SW (Lfe), SL, SR) like the Dolby AC-3 system (4) 8 channel system SDDS (Sony Dynamic D
forward L, LC, C, R, like the igital Sound) method
6 channels of C, R and SW + 2 channels of rear SL and SR, total 8 channels

【００１０】図１に示す符号化側の６チャネル（ch）ミ
クス＆マトリクス回路１’は、マルチチャネル信号の一
例としてフロントレフト（Ｌｆ）、センタ（Ｃ）、フロ
ントライト（Ｒｆ）、サラウンドレフト（Ｌｓ）、サラ
ウンドライト（Ｒｓ）及びＬｆｅ（Low Frequency Effe
ct）の６chのＰＣＭデータを次式（１）により前方グル
ープに関する２ch「１」、「２」と他のグループに関す
る４ch「３」〜「６」に分類して変換し、２ch「１」、
「２」を第１符号化部２’−１に、また、４ch「３」〜
「６」を第２符号化部２’−２に出力する。「１」＝Ｌｆ＋Ｒｆ「２」＝Ｌｆ−Ｒｆ「３」＝Ｃ−（Ｌｓ＋Ｒｓ）／２「４」＝Ｌｓ＋Ｒｓ「５」＝Ｌｓ−Ｒｓ「６」＝Ｌｆｅ−ａ×Ｃただし、０≦ａ≦１ …（１）The 6-channel (ch) mix & matrix circuit 1'on the encoding side shown in FIG. 1 is a front left (Lf), center (C), front right (Rf), surround left ( Ls), surround light (Rs) and Lfe (Low Frequency Effe)
ct) 6ch PCM data are classified into 2ch “1” and “2” related to the front group and 4ch “3” to “6” related to other groups by the following formula (1), and converted to 2ch “1”,
"2" to the first encoding unit 2'-1 and 4ch "3" to
"6" is output to the second encoding unit 2'-2. “1” = Lf + Rf “2” = Lf−Rf “3” = C− (Ls + Rs) / 2 “4” = Ls + Rs “5” = Ls−Rs “6” = Lfe−a × C where 0 ≦ a ≦ 1 (1)

【００１１】符号化部２’を構成する第１及び第２符号
化部２’−１、２’−２はそれぞれ、図２に詳しく示す
ように２ch「１」、「２」と４ch「３」〜「６」のＰＣ
Ｍデータをチャネル毎に予測符号化し、予測符号化デー
タを図３に示すようなビットストリームで記録媒体５や
衛星回線や電話回線等の通信媒体６を介して復号側に伝
送する。復号側では復号化部３’を構成する第１及び第
２復号化部３’−１、３’−２により、図７に詳しく示
すようにそれぞれ前方グループに関する２ch「１」、
「２」と他のグループに関する４ch「３」〜「６」の予
測符号化データをチャネル毎にＰＣＭデータに復号す
る。As shown in detail in FIG. 2, the first and second encoding units 2'-1, 2'-2 constituting the encoding unit 2'respectively have 2ch "1", "2" and 4ch "3". ] ~ "6" PC
The M data is predictively coded for each channel, and the predictively coded data is transmitted to the decoding side via a recording medium 5 and a communication medium 6 such as a satellite line or a telephone line as a bit stream as shown in FIG. On the decoding side, the first and second decoding units 3′-1 and 3′-2, which constitute the decoding unit 3 ′, respectively, as shown in detail in FIG.
Predicted coded data of 4 channels “3” to “6” related to “2” and other groups is decoded into PCM data for each channel.

【００１２】次いでミクス＆マトリクス回路４’により
式（１）に基づいて元の６ch（Ｌｆ、Ｃ、Ｒｆ、Ｌｓ、
Ｒｓ、Ｌｆｅ）を復元するとともに、この元の６chと係
数ｍij（ｉ＝１，２，ｊ＝１，２〜６）により次式
（２）のようにステレオ２chデータ（Ｌ、Ｒ）を生成す
る。Ｌ＝ｍ11・Ｌｆ＋ｍ12・Ｒｆ＋ｍ13・Ｃ＋ｍ14・Ｌｓ＋ｍ15・Ｒｓ＋ｍ16・ＬｆｅＲ＝ｍ21・Ｌｆ＋ｍ22・Ｒｆ＋ｍ23・Ｃ＋ｍ24・Ｌｓ＋ｍ25・Ｒｓ＋ｍ26・Ｌｆｅ …（２）Then, based on the equation (1), the original 6ch (Lf, C, Rf, Ls,
Rs, Lfe) is restored, and stereo 2ch data (L, R) is generated by the original 6ch and coefficients mij (i = 1, 2, j = 1, 2 to 6) as shown in the following equation (2). To do. L = m11 ・ Lf + m12 ・ Rf + m13 ・ C + m14 ・ Ls + m15 ・ Rs + m16 ・ Lfe R = m21 ・ Lf + m22 ・ Rf + m23 ・ C + m24 ・ Ls + m25 ・ Rs + m26 ・ Lfe (2)

【００１３】図２を参照して符号化部２’−１、２’−
２について詳しく説明する。各ch「１」〜「６」のＰＣ
Ｍデータは１フレーム毎に１フレームバッファ１０に格
納される。そして、１フレームの各ch「１」〜「６」の
サンプルデータがそれぞれ予測回路１３Ｄ１、１３Ｄ
２、１５Ｄ１〜１５Ｄ４に印加されるとともに、各ch
「１」〜「６」の各フレームの先頭サンプルデータ（後
述のリスタートヘッダ内に格納される）がアンパッキン
グ回路８及びフォーマット化回路１９に印加される。ま
た、ＰＣＭデータがＡ／Ｄ変換されたときのサンプリン
グ周波数（ｆｓ）と量子化ビット数（Ｑｂ）がパッキン
グ回路１８及びフォーマット化回路１９に印加される。
予測回路１３Ｄ１、１３Ｄ２、１５Ｄ１〜１５Ｄ４はそ
れぞれ、各ch「１」〜「６」のＰＣＭデータに対して、
特性が異なる複数の予測器（不図示）により時間領域に
おける過去の信号から現在の信号の複数の線形予測値を
算出し、次いで原ＰＣＭデータと、この複数の線形予測
値から予測器毎の予測残差を算出する。続くバッファ・
選択器１４Ｄ１、１４Ｄ２、１６Ｄ１〜１６Ｄ４はそれ
ぞれ、予測回路１３Ｄ１、１３Ｄ２、１５Ｄ１〜１５Ｄ
４により算出された各予測残差を一時記憶して、選択信
号／ＤＴＳ（デコーディング・タイム・スタンプ）生成
器１７により指定されたサブフレーム毎に予測残差の最
小値を選択する。Referring to FIG. 2, coding units 2'-1, 2'-
2 will be described in detail. PC of each channel "1" to "6"
The M data is stored in the frame buffer 10 for each frame. Then, the sample data of each channel “1” to “6” of one frame is input to the prediction circuits 13D1 and 13D, respectively.
2, 15D1 to 15D4 applied to each channel
The leading sample data (stored in a restart header described later) of each frame of “1” to “6” is applied to the unpacking circuit 8 and the formatting circuit 19. The sampling frequency (fs) and the number of quantization bits (Qb) when the PCM data is A / D converted are applied to the packing circuit 18 and the formatting circuit 19.
The prediction circuits 13D1, 13D2, and 15D1 to 15D4 respectively respond to the PCM data of each channel “1” to “6”,
A plurality of predictors (not shown) having different characteristics are used to calculate a plurality of linear prediction values of a current signal from a past signal in the time domain, and then original PCM data and prediction of each predictor from the plurality of linear prediction values. Calculate the residuals. Continued buffer
The selectors 14D1, 14D2, 16D1 to 16D4 are prediction circuits 13D1, 13D2, 15D1 to 15D, respectively.
The prediction residuals calculated in step 4 are temporarily stored, and the minimum value of the prediction residuals is selected for each subframe designated by the selection signal / DTS (decoding time stamp) generator 17.

【００１４】選択信号／ＤＴＳ生成器１７は予測残差の
ビット数フラグをパッキング回路１８とフォーマット化
回路１９に対して印加し、また、予測残差が最小の予測
器を示す予測器選択フラグと、式（１）における相関係
数ａと、復号化側が入力バッファ２２ａ（図７）からス
トリームデータを取り出す時間を示すＤＴＳをフォーマ
ット化回路１９に対して印加する。パッキング回路１８
はバッファ・選択器１４Ｄ１、１４Ｄ２、１６Ｄ１〜１
６Ｄ４により選択された６ch分の予測残差を、選択信号
／ＤＴＳ生成器１７により指定されたビット数フラグに
基づいて指定ビット数でパッキングする。またＰＴＳ生
成器１７ｃは、復号化側が出力バッファ１１０（図７）
からＰＣＭデータを取り出す時間を示すＰＴＳ（プレゼ
ンテーション・タイム・スタンプ）を生成してフォーマ
ット化回路１９に出力する。The selection signal / DTS generator 17 applies a bit number flag of the prediction residual to the packing circuit 18 and the formatting circuit 19, and a predictor selection flag indicating a predictor with the minimum prediction residual. , The correlation coefficient a in Expression (1) and the DTS indicating the time at which the decoding side takes out the stream data from the input buffer 22a (FIG. 7) are applied to the formatting circuit 19. Packing circuit 18
Is the buffer / selector 14D1, 14D2, 16D1-1
The prediction residuals for 6 channels selected by 6D4 are packed with the specified number of bits based on the bit number flag specified by the selection signal / DTS generator 17. The decoding side of the PTS generator 17c is the output buffer 110 (FIG. 7).
PTS (Presentation Time Stamp) indicating the time to take out the PCM data is generated and output to the formatting circuit 19.

【００１５】続くフォーマット化回路１９は図３〜図６
に示すようなユーザデータにフォーマット化する。図３
に示すユーザデータ（サブパケット）は、前方グループ
に関する２ch「１」、「２」の予測符号化データを含む
可変レートビットストリーム（サブストリーム）ＢＳ０
と、他のグループに関する４ch「３」〜「６」の予測符
号化データを含む可変レートビットストリーム（サブス
トリーム）ＢＳ１と、サブストリームＢＳ０、ＢＳ１の
前に設けられたビットストリームヘッダ（リスタートヘ
ッダ）により構成されている。また、サブストリームＢ
Ｓ０、ＢＳ１の１フレーム分は・フレームヘッダと、・各ch「１」〜「６」の１フレームの先頭サンプルデー
タと、・各ch「１」〜「６」のサブフレーム毎の予測器選択フ
ラグと、・各ch「１」〜「６」のサブフレーム毎のビット数フラ
グと、・各ch「１」〜「６」の予測残差データ列（可変ビット
数）と、・ch「６」の係数ａとが、多重化されている。このような予測符号化によれば、原
信号が例えばサンプリング周波数（ｆｓ）＝９６ｋＨ
ｚ、量子化ビット数（Ｑｂ）＝２４ビット、６チャネル
の場合、７１％の圧縮率を実現することができる。The following formatting circuit 19 is shown in FIGS.
Format the user data as shown in. Figure 3
The user data (sub-packet) shown in (1) is a variable rate bit stream (sub-stream) BS0 including predictive encoded data of 2ch “1” and “2” related to the front group.
And a variable rate bitstream (substream) BS1 including predictive coded data of 4ch "3" to "6" relating to other groups, and a bitstream header (restart header) provided before the substreams BS0 and BS1. ). Also, substream B
For one frame of S0 and BS1, a frame header, head sample data of one frame of each channel "1" to "6", and a predictor selection for each subframe of each channel "1" to "6" A flag, a bit number flag for each subframe of each ch “1” to “6”, a prediction residual data string (variable number of bits) of each ch “1” to “6”, and a ch “6” And the coefficient a of “. According to such predictive coding, the original signal has a sampling frequency (fs) = 96 kHz, for example.
In the case of z, the number of quantization bits (Qb) = 24 bits, and 6 channels, a compression rate of 71% can be realized.

【００１６】図２に示す符号化部２’−１、２’−２に
より予測符号化された可変レートビットストリームデー
タを、記録媒体の一例としてＤＶＤオーディオディスク
に記録する場合には、図４に示すオーディオ（Ａ）パッ
クにパッキングされる。このパックは２０３４バイトの
ユーザデータ（Ａパケット、Ｖパケット）に対して４バ
イトのパックスタート情報と、６バイトのＳＣＲ（Syst
em Clock Reference：システム時刻基準参照値）情報
と、３バイトのMux レート（rate）情報と１バイトのス
タッフィングの合計１４バイトのパックヘッダが付加さ
れて構成されている（１パック＝合計２０４８バイ
ト）。この場合、タイムスタンプであるＳＣＲ情報を、
先頭パックでは「１」として同一タイトル内で連続とす
ることにより同一タイトル内のＡパックの時間を管理す
ることができる。When the variable rate bit stream data predictively coded by the coding units 2'-1 and 2'-2 shown in FIG. 2 is recorded on a DVD audio disc as an example of a recording medium, FIG. The audio (A) pack shown is packed. This pack includes pack start information of 4 bytes for user data (A packet and V packet) of 2034 bytes, and SCR (Syst of 6 bytes).
em Clock Reference: System time reference reference value information, 3-byte Mux rate information, and 1-byte stuffing, a total of 14-byte pack header is added (1 pack = 2048 bytes in total). . In this case, the SCR information, which is the time stamp,
In the first pack, "1" is set to be consecutive within the same title, so that the time of the A pack in the same title can be managed.

【００１７】圧縮ＰＣＭのＡパケットは図５に詳しく示
すように、９〜２２バイトのパケットヘッダと、圧縮Ｐ
ＣＭのプライベートヘッダと、図３に示すフォーマット
の１ないし２０１５バイトのオーディオデータ（圧縮Ｐ
ＣＭ）により構成されている。そして、ＤＴＳとＰＴＳ
は図５のパケットヘッダ内に（具体的にはパケットヘッ
ダの１０〜１４バイト目にＰＴＳが、１５〜１９バイト
目にＤＴＳが）セットされる。圧縮ＰＣＭのプライベー
トヘッダは、・１バイトのサブストリームＩＤと、・２バイトのＵＰＣ／ＥＡＮ−ＩＳＲＣ（Universal Pr
oduct Code/European Article Number-International S
tandard Recording Code）番号、及びＵＰＣ／ＥＡＮ−
ＩＳＲＣデータと、・１バイトのプライベートヘッダ長と、・２バイトの第１アクセスユニットポインタと、・４バイトのオーディオデータ情報（ＡＤＩ）と、・０〜７バイトのスタッフィングバイトとに、より構成
されている。As shown in detail in FIG. 5, the compressed PCM A packet has a packet header of 9 to 22 bytes and a compressed P packet.
CM private header and 1 to 2015 bytes of audio data in the format shown in FIG. 3 (compressed P
CM). And DTS and PTS
Is set in the packet header of FIG. 5 (specifically, PTS is at the 10th to 14th bytes of the packet header and DTS is at the 15th to 19th bytes). The private header of the compressed PCM includes: 1-byte substream ID, 2-byte UPC / EAN-ISRC (Universal Pr
oduct Code / European Article Number-International S
tandard Recording Code) number and UPC / EAN-
ISRC data, 1-byte private header length, 2-byte first access unit pointer, 4-byte audio data information (ADI), and 0-7 bytes of stuffing bytes. ing.

【００１８】そして、ＡＤＩ内に１秒後のアクセスユニ
ットをサーチするための前方アクセスユニット・サーチ
ポインタと、１秒前のアクセスユニットをサーチするた
めの後方アクセスユニット・サーチポインタがともに１
バイトでセットされる。具体的には、ＡＤＩの１バイト
目に前方アクセスユニット・サーチポインタが、８バイ
ト目に後方アクセスユニット・サーチポインタがセット
される。このようにＡＤＩは、圧縮ＰＣＭでは４バイト
に減少させるためオーディオデータを２０１５バイトま
で収納できる。The forward access unit search pointer for searching the access unit one second later and the rear access unit search pointer for searching the access unit one second before are both 1 in the ADI.
Set in bytes. Specifically, the forward access unit search pointer is set at the 1st byte of the ADI and the backward access unit search pointer is set at the 8th byte. In this way, the ADI can store up to 2015 bytes of audio data because the compressed PCM is reduced to 4 bytes.

【００１９】図５に示す圧縮ＰＣＭ（ＰＰＣＭ）のオー
ディオパケットにおけるオーディオデータエリアは、図
６に示すように複数のＰＰＣＭアクセスユニットにより
構成され、ＰＰＣＭアクセスユニットはＰＰＣＭシンク
情報とサブパケットにより構成されている。最初のＰＰ
ＣＭアクセスユニット内のサブパケットは、ディレクト
リと、サブストリーム「ＢＳ０」と、ＣＲＣ（１バイト
又は２バイト）と、サブストリーム「ＢＳ１」と、ＣＲ
Ｃとエクストラ情報により構成され、サブストリーム
「ＢＳ０」、「ＢＳ１」はＰＰＣＭブロックのみにより
構成されている。２番目以降のＰＰＣＭアクセスユニッ
ト内のサブパケットも、ディレクトリと、サブストリー
ム「ＢＳ０」と、ＣＲＣと、サブストリーム「ＢＳ１」
と、ＣＲＣとエクストラ情報により構成され、サブスト
リーム「ＢＳ０」、「ＢＳ１」はリスタートヘッダとＰ
ＰＣＭブロックにより構成されている。そして、エクス
トラ情報は、少なくとも、サイズ調整機能を有してい
る。すなわち、入来データが固定レート（ＣＢＲ）の場
合には、上述したようにサンプリング周波数ｆｓによっ
て１パケット当たりのサンプリング数が４０，８０，１
６０のいずれかに定められており、そのため、決定され
たサンプリング数によっては１パケット当たりのデータ
長とサブパケットのサイズとが合わない場合があり、そ
れをサブパケットのサイズに合わせるために、例えば、
０，０…等を付加してサイズ調整を行う。また、このサ
イズ調整用のデータはテキストデータ等を利用すること
も可能である。The audio data area in the compressed PCM (PPCM) audio packet shown in FIG. 5 is composed of a plurality of PPCM access units as shown in FIG. 6, and the PPCM access unit is composed of PPCM sync information and subpackets. There is. First PP
The subpacket in the CM access unit includes a directory, a substream "BS0", a CRC (1 byte or 2 bytes), a substream "BS1", and a CR.
The substreams "BS0" and "BS1" are composed of C and extra information, and are composed of only PPCM blocks. The subpackets in the second and subsequent PPCM access units are also the directory, the substream "BS0", the CRC, and the substream "BS1".
And the CRC and the extra information, the sub-streams “BS0” and “BS1” have a restart header and P
It is composed of PCM blocks. The extra information has at least a size adjusting function. That is, when the incoming data has a fixed rate (CBR), the sampling number per packet is 40, 80, 1 depending on the sampling frequency fs as described above.
It is set to any of 60, and therefore the data length per packet and the size of the subpacket may not match depending on the determined number of samplings. To match it with the size of the subpacket, for example, ,
Size adjustment is performed by adding 0, 0 ... Further, text data or the like can be used as the size adjusting data.

【００２０】ＰＰＣＭシンク情報（以下、同期情報とも
いう）は次の情報を含む。・１パケット当たりのサンプル数：サンプリング周波数
ｆｓに応じて４０、８０又は１６０が選択される。・データレートがＶＢＲの場合には「０」（サブパケッ
ト内のデータがＶＢＲの圧縮データであることを示す識
別子）、ＣＢＲの場合には「１」（サブパケット内のデ
ータが固定レートであることを示す識別子）・サンプリング周波数ｆｓ及び量子化ビット数Ｑｂ・チャネル割り当て情報The PPCM sync information (hereinafter also referred to as synchronization information) includes the following information. Number of samples per packet: 40, 80 or 160 is selected according to the sampling frequency fs. “0” (data indicating that the data in the subpacket is compressed data of VBR) when the data rate is VBR, and “1” (the data in the subpacket has a fixed rate) when the data rate is CBR Identifier indicating that) -sampling frequency fs and quantization bit number Qb-channel allocation information

【００２１】次に図７を参照して復号化部３’−１、
３’−２について説明する。上記フォーマットの可変レ
ートビットストリームデータＢＳ０、ＢＳ１は、デフォ
ーマット化回路２１により分離される。そして、各ｃｈ
「１」〜「６」の１フレームの先頭サンプルデータと予
測器選択フラグはそれぞれ予測回路２４Ｄ１、２４Ｄ
２、２３Ｄ１〜２３Ｄ４に印加され、各ｃｈ「１」〜
「６」のビット数フラグはアンパッキング回路２２に印
加される。また、ＳＣＲと、ＤＴＳと予測残差データ列
は入力バッファ２２ａに印加され、ＰＴＳは出力バッフ
ァ１１０に印加される。また、データレートがＶＢＲか
ＣＢＲかを示す識別子は各予測器２４Ｄ１、２４Ｄ２、
２３Ｄ１、２３Ｄ２、２３Ｄ３、２３Ｄ４に印加され、
これらにおいて識別子に応じた入出力データの処理プロ
グラムが決定されて処理されることになる。ＶＢＲであ
る場合には処理プログラムを切り換えると共に入力デー
タを毎回ロードする必要があり処理に時間を要すること
になるが、ＣＢＲの場合には固定レートであることから
処理プログラムを切り換える必要がなく処理が速くな
る。また、サンプリング周波数ｆｓ及び量子化ビット数
ＱｂはＤ／Ａ変換器１０２に印加される。ここで、予測
回路２４Ｄ１、２４Ｄ２、２３Ｄ１〜２３Ｄ４内の複数
の予測器（不図示）はそれぞれ、符号化側の予測回路１
３Ｄ１、１３Ｄ２、１５Ｄ１〜１５Ｄ４内の複数の予測
器と同一の特性であり、予測器選択フラグにより同一特
性のものが選択される。Next, referring to FIG. 7, the decoding unit 3'-1,
3′-2 will be described. The variable rate bit stream data BS0 and BS1 in the above format are separated by the deformatting circuit 21. And each ch
The leading sample data of one frame of "1" to "6" and the predictor selection flag are prediction circuits 24D1 and 24D, respectively.
2, 23D1 to 23D4 are applied to each channel “1” to
The bit number flag of “6” is applied to the unpacking circuit 22. The SCR, DTS, and prediction residual data string are applied to the input buffer 22 a, and PTS is applied to the output buffer 110. The identifiers indicating whether the data rate is VBR or CBR are the predictors 24D1 and 24D2,
Applied to 23D1, 23D2, 23D3, 23D4,
In these, the processing program of the input / output data according to the identifier is determined and processed. In the case of VBR, it is necessary to switch the processing program and load the input data every time, and thus it takes time to process, but in the case of CBR, since the fixed rate is used, the processing program does not need to be switched and the processing is performed. Get faster Further, the sampling frequency fs and the quantization bit number Qb are applied to the D / A converter 102. Here, the plurality of predictors (not shown) in the prediction circuits 24D1, 24D2, and 23D1 to 23D4 are the prediction circuits 1 on the encoding side, respectively.
The characteristics are the same as those of the plurality of predictors in 3D1, 13D2, 15D1 to 15D4, and those having the same characteristics are selected by the predictor selection flag.

【００２２】デフォーマット化回路２１により、最初オ
ーディオパックからオーディオパケットが分離され、次
にオーディオパケットからストリームデータ（予測残差
データ列）が分離されてビットストリームＢＳ０とＢＳ
１が取り出される。またＳＣＲが取り出され、図８に示
すようにＳＣＲによるタイミングにしたがってアクセス
ユニット毎に入力バッファ２２ａに取り込まれて蓄積さ
れる。ここで、１つのアクセスユニットのデータ量は、
例えばｆｓ＝９６ｋＨｚの場合には（１／９６ｋＨｚ）
秒分であるが、図９、図１０（ａ）に詳しく示すように
可変長である。そして、入力バッファ２２ａに蓄積され
たストリームデータはＤＴＳに基づいてＦＩＦＯで読み
出されてアンパッキング回路２２に印加される。The deformatting circuit 21 first separates the audio packet from the audio pack, and then separates the stream data (prediction residual data string) from the audio packet to form the bit streams BS0 and BS.
1 is taken out. Further, the SCR is taken out, and taken in and stored in the input buffer 22a for each access unit according to the timing by the SCR as shown in FIG. Here, the data amount of one access unit is
For example, in the case of fs = 96 kHz (1/96 kHz)
Although it is seconds, it has a variable length as shown in detail in FIGS. 9 and 10A. Then, the stream data accumulated in the input buffer 22 a is read by the FIFO based on the DTS and applied to the unpacking circuit 22.

【００２３】アンパッキング回路２２は各ｃｈ「１」〜
「６」の予測残差データ列をビット数フラグ毎に基づい
て分離してそれぞれ予測回路２４Ｄ１、２４Ｄ２、２３
Ｄ１〜２３Ｄ４に出力する。予測回路２４Ｄ１、２４Ｄ
２、２３Ｄ１〜２３Ｄ４ではそれぞれ、アンパッキング
回路２２からの各ｃｈ「１」〜「６」の今回の予測残差
データと、内部の複数の予測器の内、予測器選択フラグ
により選択された各１つにより予測された前回の予測値
が加算されて今回の予測値が算出され、次いで１フレー
ムの先頭サンプルデータを基準として各サンプルのＰＣ
Ｍデータが算出されて出力バッファ１１０に蓄積され
る。出力バッファ１１０に蓄積されたＰＣＭデータはＰ
ＴＳに基づいて読み出されて出力され、したがって、図
１０（ａ）に示す可変長のアクセスユニットが伸長され
て、図１０（ｂ）に示す一定長のプレゼンテーションユ
ニットが出力される。The unpacking circuit 22 has channels "1" to
The prediction residual data string of "6" is separated based on each bit number flag, and the prediction circuits 24D1, 24D2, and 23 are respectively separated.
Output to D1 to D4. Prediction circuits 24D1 and 24D
2, 23D1 to 23D4, the current prediction residual data of each channel “1” to “6” from the unpacking circuit 22 and each of the plurality of internal predictors selected by the predictor selection flag. The previous predicted value predicted by one is added to calculate the current predicted value, and then the PC of each sample is based on the first sample data of one frame as a reference.
M data is calculated and stored in the output buffer 110. The PCM data stored in the output buffer 110 is P
It is read and output based on the TS, and thus the variable-length access unit shown in FIG. 10A is expanded and the constant-length presentation unit shown in FIG. 10B is output.

【００２４】また、ＰＰＣＭシンク情報内のサンプリン
グ周波数ｆｓ及び量子化ビット数Ｑｂに基づいて、ＰＣ
ＭデータがＤ／Ａ変換器１０２によりアナログ信号に変
換される。また、同時にＰＰＣＭシンク情報においてＣ
ＢＲの識別子が検出され、ディレクトリ内のエクストラ
データの位置が検出されて、更に例えば０，０…のデー
タや、テキストデータ等のサイズ調整用のエクストラデ
ータが検出されると、それがテキストデータである場合
にはエクストラデータをこのアンパッキング回路２２か
ら図示しないテキストデータデコード回路に供給し、そ
こで、デコード処理をしてテキストデータとして取り出
し、出力バッファ１１０を通じて出力されることにな
る。また一方、エクストラデータが０，０…データであ
った場合には、何の処理も施されないようになってい
る。また、テキストデータデコーダ回路が用意されてい
ない場合には、この処理はパスされる。また、ここで、
操作部１０１を介してサーチ再生が指示された場合に
は、制御部１００により図５に示す前方アクセスユニッ
ト・サーチポインタ（１秒先）と後方アクセスユニット
・サーチポインタ（１秒前）に基づいてアクセスユニッ
トを再生する。このサーチポインタとしては、１秒先、
１秒前の代わりに２秒先、２秒前のものでよい。Further, based on the sampling frequency fs and the quantization bit number Qb in the PPCM sync information, the PC
The M data is converted into an analog signal by the D / A converter 102. At the same time, C in the PPCM sync information
When the BR identifier is detected, the position of the extra data in the directory is detected, and further, for example, data of 0, 0 ... Or extra data for size adjustment such as text data is detected, it is converted to text data. In some cases, the extra data is supplied from the unpacking circuit 22 to a text data decoding circuit (not shown), where it is subjected to decoding processing to be taken out as text data and output through the output buffer 110. On the other hand, if the extra data is 0, 0 ... Data, no processing is performed. If the text data decoder circuit is not prepared, this process is skipped. Also here
When the search reproduction is instructed through the operation unit 101, based on the front access unit search pointer (1 second ahead) and the rear access unit search pointer (1 second before) shown in FIG. Regenerate the access unit. For this search pointer, 1 second ahead,
Instead of 1 second ago, it may be 2 seconds ahead or 2 seconds ago.

【００２５】図２に示す符号化部２’−１、２’−２に
より予測符号化された可変レートビットストリームデー
タをネットワークを介して伝送する場合には、符号化側
では図１１に示すように伝送用にパケット化し（ステッ
プＳ４１）、次いでパケットヘッダを付与し（ステップ
Ｓ４２）、次いでこのパケットをネットワーク上に送り
出す（ステップＳ４３）。When variable rate bit stream data predictively coded by the coding units 2'-1, 2'-2 shown in FIG. 2 are transmitted via a network, the coding side is as shown in FIG. Is packetized for transmission (step S41), a packet header is added (step S42), and this packet is then sent out on the network (step S43).

【００２６】復号側では図１２（Ａ）に示すようにヘッ
ダを除去し（ステップＳ５１）、次いでデータを復元し
（ステップＳ５２）、次いでこのデータをメモリに格納
して復号を待つ（ステップＳ５３）。そして、復号を行
う場合には図１２（Ｂ）に示すように、デフォーマット
化を行い（ステップＳ６１）、次いで入力バッファ２２
ａの入出力制御を行い（ステップＳ６２）、次いでアン
パッキングを行う（ステップＳ６３）。なお、このと
き、サーチ再生指示がある場合にはサーチポインタをデ
コードする。次いで予測器をフラグに基づいて選択して
デコードを行い（ステップＳ６４）、次いで出力バッフ
ァ１１０の入出力制御を行い（ステップＳ６５）、次い
で元のマルチチャネルを復元し（ステップＳ６６）、次
いでこれを出力し（ステップＳ６７）、以下、これを繰
り返す。On the decoding side, as shown in FIG. 12A, the header is removed (step S51), the data is restored (step S52), this data is then stored in the memory and waits for decoding (step S53). . Then, when decoding is performed, as shown in FIG. 12B, the reformatting is performed (step S61), and then the input buffer 22
Input / output control of a is performed (step S62), and then unpacking is performed (step S63). At this time, if there is a search reproduction instruction, the search pointer is decoded. Next, a predictor is selected based on the flag to perform decoding (step S64), then input / output control of the output buffer 110 is performed (step S65), then the original multi-channel is restored (step S66), and then this is restored. It is output (step S67), and this is repeated thereafter.

【００２７】なお、上記実施形態では、前方グループに
関する２ch「１」、「２」を「１」＝Ｌｆ＋Ｒｆ「２」＝Ｌｆ−Ｒｆにより変換して予測符号化したが、代わりに式（２）に
よりマルチチャネルをダウンミクスしてステレオ２chデ
ータ（Ｌ、Ｒ）を生成し、次いで次式（１）’ 「１」＝Ｌ＋Ｒ「２」＝Ｌ−Ｒ「３」〜「５」は同じ「６」＝Ｌｆｅ−Ｃ …（１）’ により変換して予測符号化するようにしてもよい（第２
の実施形態）。この場合には、復号化側のミクス＆マト
リクス回路４’はチャネル「１」、「２」を加算するこ
とによりチャネルＬを、減算することによりチャネルＲ
を生成することができる。In the above embodiment, 2ch "1" and "2" relating to the front group are converted by "1" = Lf + Rf "2" = Lf-Rf and predictively coded. To down-mix the multi-channel to generate stereo 2ch data (L, R), and then the following equation (1) ′ “1” = L + R “2” = L−R “3” to “5” are the same as “6”. "= Lfe-C (1) 'and the prediction coding may be performed (second).
Embodiment). In this case, the mixing and matrix circuit 4'on the decoding side adds the channels "1" and "2" to the channel L, and subtracts the channel R to the channel R.
Can be generated.

【００２８】また、第３の実施形態として図１３に示す
ように、２ch「１」、「２」の代わりに式（２）により
マルチチャネルをダウンミクスしてステレオ２chデータ
（Ｌ、Ｒ）を生成して、このステレオ２ch（Ｌ、Ｒ）と
４ch「３」〜「６」を予測符号化するようにしてもよ
い。なお、第２、第３の実施形態では、フロントレフト
（Ｌｆ）とフロントライト（Ｒｆ）が復号化側に伝送さ
れないので、復号化側ではこれを式（１）、（２）によ
り生成する。As a third embodiment, as shown in FIG. 13, the stereo 2ch data (L, R) is downmixed by the formula (2) instead of 2ch "1", "2" to downmix the multichannel. The stereo 2ch (L, R) and the 4ch "3" to "6" may be generated and predictively encoded. Note that in the second and third embodiments, the front left (Lf) and the front right (Rf) are not transmitted to the decoding side, so the decoding side generates them by the equations (1) and (2).

【００２９】次に図１４、図１５、図１６を参照して第
４の実施形態について説明する。上記の実施形態では、
１グループの相関性の信号「１」〜「６」を予測符号化
するように構成されているが、この第４の実施形態では
複数グループの相関性のある信号を生成して予測符号化
し、圧縮率が最も高いグループの予測符号化データを選
択するように構成されている。また、この実施例ではそ
の１グループ内における符号化は、前述の各実施例の場
合のように前方グループに関する２ｃｈと他のグループ
に関する４ｃｈに分類して変換するようなことはせず
に、一つにまとめた符号化処理が行われる構成で、図１
４は前述の図１に対応した図として示してある。また、
図１５は符号化部の詳細ブロックを示すものであるが、
本実施例の場合にはｎ個の相関回路１−１〜１−ｎまで
が、ミクス＆マトリクス回路１’側に設けられている。
これらｎ個の相関回路１−１〜１−ｎは例えば６ch（Ｌ
ｆ、Ｃ、Ｒｆ、Ｌｓ、Ｒｓ、Ｌｆｅ）のＰＣＭデータ
を、相関性が異なるｎ種類の６ch信号「１」〜「６」に
変換する。Next, a fourth embodiment will be described with reference to FIGS. 14, 15 and 16. In the above embodiment,
Although it is configured to predictively encode one group of correlated signals “1” to “6”, in the fourth embodiment, a plurality of groups of correlated signals are generated and predictively encoded, It is configured to select the predictive encoded data of the group having the highest compression rate. Further, in this embodiment, the coding within the one group is not performed by classifying and converting into 2ch related to the front group and 4ch related to other groups as in the case of each of the above-described embodiments. 1 is a configuration in which the encoding processing is put together into one.
4 is shown as a diagram corresponding to FIG. 1 described above. Also,
FIG. 15 shows a detailed block of the encoding unit.
In the case of this embodiment, n correlation circuits 1-1 to 1-n are provided on the side of the mix & matrix circuit 1 '.
These n correlation circuits 1-1 to 1-n are, for example, 6ch (L
The PCM data of f, C, Rf, Ls, Rs, and Lfe) is converted into n types of 6ch signals “1” to “6” having different correlations.

【００３０】例えば第１の相関回路１−１は以下のよう
に変換し、「１」＝Ｌｆ「２」＝Ｃ−（Ｌｓ＋Ｒｓ）／２「３」＝Ｒｆ−Ｌｆ「４」＝Ｌｓ−ａ×Ｌｆｅ「５」＝Ｒｓ−ｂ×Ｒｆ「６」＝Ｌｆｅまた、第ｎの相関回路１−ｎは以下のように変換する。「１」＝Ｌｆ＋Ｒｆ「２」＝Ｃ−Ｌｆ「３」＝Ｒｆ−Ｌｆ「４」＝Ｌｓ−Ｌｆ「５」＝Ｒｓ−Ｌｆ「６」＝Ｌｆｅ−ＣFor example, the first correlation circuit 1-1 converts as follows, "1" = Lf "2" = C- (Ls + Rs) / 2 "3" = Rf-Lf "4" = Ls-a × Lfe “5” = Rs−b × Rf “6” = Lfe Further, the nth correlation circuit 1-n performs conversion as follows. "1" = Lf + Rf "2" = C-Lf "3" = Rf-Lf "4" = Ls-Lf "5" = Rs-Lf "6" = Lfe-C

【００３１】また、相関回路１−１〜１−ｎ毎に予測回
路１５とバッファ・選択器１６が設けられ、グループ毎
の予測残差の最小値のデータ量に基づいて圧縮率が最も
高いグループが相関選択信号生成器１７ｂにより選択さ
れる。このとき、フォーマット化回路１９はその選択フ
ラグ（相関回路選択フラグ、その相関回路の相関係数
ａ、ｂ）を追加して多重化する。A prediction circuit 15 and a buffer / selector 16 are provided for each of the correlation circuits 1-1 to 1-n, and the group with the highest compression ratio is based on the data amount of the minimum value of the prediction residual for each group. Are selected by the correlation selection signal generator 17b. At this time, the formatting circuit 19 adds the selection flag (correlation circuit selection flag, correlation coefficients a and b of the correlation circuit) and multiplexes them.

【００３２】そして、図１６は前述の図６に対応したデ
ータエリアを示し、この実施例ではサブストリーム「Ｂ
Ｓ１」を用いず、サブストリーム「ＢＳ０」のみで構成
することになる。FIG. 16 shows a data area corresponding to the above-mentioned FIG. 6, and in this embodiment, the substream “B
S1 "is not used and only substream" BS0 "is used.

【００３３】また、図１７に示す復号化側では、符号化
側の相関回路１−１〜１−ｎに対してｎ個の相関回路４
−１〜４−ｎ（又は係数ａ、ｂが変更可能な図示省略の
１つの相関回路）が設けられる。なお、図１５に示すｎ
グループの予測回路が同一の構成である場合、復号装置
では図１７に示すようにｎグループ分の予測回路を設け
る必要はなく、１つのグループ分の予測回路でよい。そ
して、符号化装置から伝送された選択フラグに基づいて
相関回路４−１〜４−ｎの１つを選択、又は係数ａ、ｂ
を設定して元の６ch（Ｌｆ、Ｃ、Ｒｆ、Ｌｓ、Ｒｓ、Ｌ
ｆｅ）を復元し、また、式（２）によりマルチチャネル
をダウンミクスしてステレオ２chデータ（Ｌ、Ｒ）を生
成する。Also, on the decoding side shown in FIG. 17, n correlation circuits 4 are provided for the correlation circuits 1-1 to 1-n on the encoding side.
-1 to 4-n (or one correlation circuit (not shown) whose coefficients a and b can be changed) are provided. Note that n shown in FIG.
When the prediction circuits of the groups have the same configuration, it is not necessary to provide the prediction circuits for n groups in the decoding device as shown in FIG. 17, and the prediction circuits for one group may be used. Then, one of the correlation circuits 4-1 to 4-n is selected based on the selection flag transmitted from the encoding device, or the coefficients a and b are selected.
To set the original 6ch (Lf, C, Rf, Ls, Rs, L
fe) is restored, and the multi-channel is downmixed by the equation (2) to generate stereo 2ch data (L, R).

【００３４】また、上記の第１の実施形態では、１種類
の相関性の信号「１」〜「６」を予測符号化するように
構成されているが、この信号「１」〜「６」のグループ
と原信号（Ｌｆ、Ｃ、Ｒｆ、Ｌｓ、Ｒｓ、Ｌｆｅ）のグ
ループを予測符号化し、圧縮率が高い方のグループを選
択するようにしてもよい。In the first embodiment, the signals "1" to "6" having one type of correlation are configured to be predictively coded, but the signals "1" to "6" are used. It is also possible to predictively code the group of and the group of the original signals (Lf, C, Rf, Ls, Rs, Lfe) and select the group having the higher compression rate.

【００３５】[0035]

【発明の効果】以上説明したように本発明によれば、圧
縮データを含むサブパケットと、そのサンプリング周波
数及び量子化ビット数を含む同期情報部を有するデータ
構造にフォーマット化するようにしたので、マルチチャ
ネルの音声信号を可変の圧縮率で符号化する場合に再生
側の復号効率を改善することができる。As described above, according to the present invention, the data packet is formatted into the data structure having the sub-packet containing the compressed data and the synchronization information part containing the sampling frequency and the number of quantization bits. When a multi-channel audio signal is encoded with a variable compression rate, the decoding efficiency on the reproducing side can be improved.

[Brief description of drawings]

【図１】本発明が適用される音声符号化装置及び音声復
号装置の第１の実施形態を示すブロック図である。FIG. 1 is a block diagram showing a first embodiment of a speech coding apparatus and speech decoding apparatus to which the present invention is applied .

【図２】図１の符号化部を詳しく示すブロック図であ
る。FIG. 2 is a block diagram showing in detail a coding unit of FIG.

【図３】図１、図２の符号化部により符号化されたビッ
トストリームを示す説明図である。FIG. 3 is an explanatory diagram showing a bitstream encoded by the encoding unit in FIGS. 1 and 2.

【図４】ＤＶＤのパックのフォーマットを示す説明図で
ある。FIG. 4 is an explanatory diagram showing a format of a DVD pack.

【図５】ＤＶＤのオーディオパックのフォーマットを示
す説明図である。FIG. 5 is an explanatory diagram showing a format of a DVD audio pack.

【図６】図５のオーディオデータエリアのフォーマット
を詳しく示す説明図である。6 is an explanatory diagram showing a format of an audio data area in FIG. 5 in detail.

【図７】図１の復号化部を詳しく示すブロック図であ
る。FIG. 7 is a block diagram showing in detail the decoding unit of FIG. 1.

【図８】図７の入力バッファの書き込み／読み出しタイ
ミングを示すタイミングチャートである。8 is a timing chart showing write / read timing of the input buffer of FIG.

【図９】アクセスユニット毎の圧縮データ量を示す説明
図である。FIG. 9 is an explanatory diagram showing a compressed data amount for each access unit.

【図１０】アクセスユニットとプレゼンテーションユニ
ットを示す説明図である。FIG. 10 is an explanatory diagram showing an access unit and a presentation unit.

【図１１】音声伝送方法を示すフローチャートである。FIG. 11 is a flowchart showing a voice transmission method.

【図１２】音声伝送方法を示すフローチャートである。FIG. 12 is a flowchart showing a voice transmission method.

【図１３】本発明が適用される音声符号化装置及び音声
復号装置の第３の実施形態を示すブロック図である。FIG. 13 is a block diagram showing a third embodiment of a speech coding apparatus and speech decoding apparatus to which the present invention is applied .

【図１４】本発明が適用される音声符号化装置及び音声
復号装置の第４の実施形態を示すブロック図である。FIG. 14 is a block diagram showing a fourth embodiment of a speech coding apparatus and speech decoding apparatus to which the present invention is applied .

【図１５】第４の実施形態の音声符号化装置を示すブロ
ック図である。FIG. 15 is a block diagram showing a speech coding apparatus according to a fourth embodiment.

【図１６】図６に対応した別の実施例の説明図である。FIG. 16 is an explanatory diagram of another embodiment corresponding to FIG.

【図１７】第４の実施形態の音声復号装置を示すブロッ
ク図である。FIG. 17 is a block diagram showing a speech decoding apparatus according to a fourth embodiment.

[Explanation of symbols]

１’ ６chミクス＆マトリクス回路１３Ｄ１，１３Ｄ２，１５Ｄ１〜１５Ｄ４予測回路
（バッファ・選択器１４Ｄ１，１４Ｄ２，１６Ｄ１〜１
６Ｄ４と共に圧縮手段を構成する。）１４Ｄ１，１４Ｄ２，１６Ｄ１〜１６Ｄ４バッファ・
選択器１７選択信号／ＤＴＳ生成器（タイミング生成手段）１７ｃＰＴＳ生成器（タイミング生成手段）１９フォーマット化回路（フォーマット化手段）２１デフォーマット化回路（分離手段）２２アンパッキング回路２２ａ入力バッファ２４Ｄ１，２４Ｄ２，２３Ｄ１〜２３Ｄ４予測回路
（伸長手段）１００制御部１０２Ｄ／Ａ変換器１１０出力バッファ1'6ch mix & matrix circuit 13D1, 13D2, 15D1 to 15D4 Prediction circuit (buffer / selector 14D1, 14D2, 16D1 to 1)
A compression means is configured with 6D4. ) 14D1, 14D2, 16D1 to 16D4 buffers
Selector 17 Select signal / DTS generator (timing generating means) 17c PTS generator (timing generating means) 19 Formatting circuit (formatting means) 21 Deformatting circuit (separating means) 22 Unpacking circuit 22a Input buffer 24D1, 24D2, 23D1 to 23D4 Prediction circuit (expansion means) 100 Control unit 102 D / A converter 110 Output buffer

───────────────────────────────────────────────────── フロントページの続き (56)参考文献特開平８−272393（ＪＰ，Ａ) 特開平３−24834（ＪＰ，Ａ) 特開平10−233058（ＪＰ，Ａ) 特開平10−320928（ＪＰ，Ａ) 特開昭64−44499（ＪＰ，Ａ) (58)調査した分野(Int.Cl.⁷，ＤＢ名) G10L 19/04 ─────────────────────────────────────────────────── ─── Continuation of the front page (56) Reference JP-A-8-272393 (JP, A) JP-A-3-24834 (JP, A) JP-A-10-233058 (JP, A) JP-A-10- 320928 (JP, A) JP-A 64-44499 (JP, A) (58) Fields investigated (Int. Cl. ⁷ , DB name) G10L 19/04

Claims

(57) [Claims]

1. A multichannel audio signal having a certain sampling frequency and a quantized bit number is obtained in response to an input audio signal for a channel as it is or for each channel correlated with each other, and a head sample value is obtained. Linear prediction values of the current signal are predicted from the past in the time domain by a plurality of linear prediction methods having different characteristics, and the prediction residual obtained from the predicted linear prediction value and the speech signal is minimized. A step of predictively encoding by selecting a linear prediction method; a linear prediction method for each channel selected in the step, a subpacket for storing prediction coded data including a prediction residual and a predetermined head sample value, and reproduction Side, a synchronization information part including a sampling frequency and a quantization bit number used when the original analog audio signal is restored. Formatting into a data structure according to the above, the data formatted in the data structure is recorded, and among the recorded data, the prediction coded data is a prediction value used for restoring the original audio signal. It is recorded as data for calculating be characterized Tei Rukoto
Recording medium .

2. The sub-packet and the synchronization information part according to claim 1, and a pack header including SCR information are further formatted and recorded as one pack, and the SCR information is recorded in the pack. A recording medium characterized by being used as time management information when reproducing.

3. A voice decoding device for decoding an original multi-channel voice signal from the data recorded on the recording medium according to claim 1, comprising means for separating the data structure into a subpacket and a synchronization information part. Decompressing means for decompressing the compressed data in the subpacket for each channel, means for converting the decompressed audio data for each channel into an original decompressed multichannel audio signal, and the converted multichannel Means for converting the audio data of 1. into an analog audio signal based on the sampling frequency and the number of quantization bits in the synchronization information part.

4. An audio decoding device for decoding an original multi-channel audio signal from the data recorded on the recording medium according to claim 2, wherein the first separating means separates the SCR information contained in the header. A buffer for holding the sub-packet and the synchronization information part based on the separated SCR information, and a second separating means for separating the sub-packet and the synchronization information part held in the buffer, Decompression means for decompressing the compressed data in the subpacket for each channel based on the identifier in the synchronization information part, and means for converting the decompressed audio data for each channel into the original multichannel decompressed audio signal. And converting the converted multi-channel audio data into an analog audio signal based on the sampling frequency and the number of quantization bits in the synchronization information section. And a means for converting into a speech signal.