Abstract: The present invention relates to an image processing device and a method that make it possible to minimize increases in the memory capacity required for an orthogonal transformation and an inverse orthogonal transformation. A partial matrix constituting part of a transformation matrix is used to derive said transformation matrix, the derived transformation matrix is used to perform an orthogonal transformation of the predicted residual of an image, the predicted residual is subjected to an orthogonal transformation and coefficient data thus obtained is encrypted, and a bit stream is generated. The present invention can be applied, for example, to an image processing device, an image encoding device, an image decoding device, or the like.
Title of invention: Image processing apparatus and method
Technical field
[0001]
The present disclosure relates to an image processing device and method, and more particularly to an image processing device and method capable of suppressing an increase in memory capacity required for orthogonal transform and inverse orthogonal transform.
Background technology
[0002]
Conventionally, with respect to luminance, a plurality of different PTs (Transform Units) are adaptively adaptive for each of a horizontal primary conversion PThor (also referred to as a primary horizontal conversion) and a vertical primary conversion PTver (also referred to as a primary vertical conversion). An adaptive primary transform (AMT) that selects a primary transform from orthogonal transforms is disclosed (for example, see Non-Patent Document 1).
[0003]
In Non-Patent Document 1, there are five one-dimensional orthogonal transforms DCT-II, DST-VII, DCT-VIII, DST-I, and DCT-VI as candidates for primary transform. Further, it has been proposed to add two 1-dimensional orthogonal transforms of DST-IV and IDT (Identity Transform: skipping one-dimensional transform), and make a total of seven 1-dimensional orthogonal transforms as candidates for the primary transform ( See, for example, Non-Patent Document 2).
Prior art documents
Non-patent literature
[0004]
Non-Patent Document 1: Jianle Chen, Elena Alshina, Gary J. Sullivan, Jens-Rainer, Jill Boyce, "Algorithm Description of Joint Exploration Test Model 4", JVET-G1001_v1, Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11 7th Meeting: Torino, IT, 13-21 July 2017
Non-Patent Document 2: V. Lorcy, P. Philippe, "Proposed improvements to the Adaptive multiple Core transform" , JVET-C0022, Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11 3rd Meeting: Geneva, CH, 26 May-1 June 2016
Summary of the invention
Problems to be Solved by the Invention
[0005]
However, in the case of these methods, there is a risk that the size of the LUT (Look Up Table) required to hold the entire conversion matrix of the primary conversion may increase. That is, considering the hardware implementation of this primary conversion, there is a risk that the memory size required to hold the coefficients of the conversion matrix may increase.
[0006]
The present disclosure has been made in view of such circumstances, and it is possible to suppress an increase in memory capacity required for orthogonal transform and inverse orthogonal transform.
Means for solving the problems
[0007]
An image processing device according to an aspect of the present technology uses a derivation unit that derives the transformation matrix using a submatrix that forms a part of a transformation matrix, and the transformation matrix that is derived by the derivation unit. An image processing apparatus including an orthogonal transformation unit that orthogonally transforms a prediction residual, and an encoding unit that encodes coefficient data obtained by orthogonally transforming the prediction residual by the orthogonal transformation unit to generate a bitstream. is there.
[0008]
An image processing method according to one aspect of the present technology is to derive the transformation matrix using a submatrix forming a part of the transformation matrix, and orthogonally transform the prediction residual of the image using the derived transformation matrix. An image processing method for encoding coefficient data obtained by orthogonally transforming the prediction residual to generate a bit stream.
[0009]
An image processing device according to another aspect of the present technology uses a decoding unit that decodes a bitstream and obtains coefficient data in which a prediction residual of an image is orthogonally transformed, and a submatrix that forms a part of a transformation matrix. An image processing device comprising: a deriving unit that derives a transform matrix; and an inverse orthogonal transform unit that performs an inverse orthogonal transform of the coefficient data obtained by the decoding unit using the transform matrix derived by the deriving unit. ..
[0010]
An image processing method according to another aspect of the present technology is to decode a bitstream, obtain coefficient data in which a prediction residual of an image is orthogonally transformed, and use the submatrix forming a part of the transformation matrix to convert the transformation matrix. Is an image processing method in which the coefficient data obtained is inversely orthogonally transformed using the derived transformation matrix.
[0011]
In an image processing device and method according to one aspect of the present technology, a submatrix that forms a part of a transformation matrix is used to derive the transformation matrix, and the derived transformation matrix is used to predict a residual image. The difference is orthogonally transformed, and the coefficient data obtained by orthogonally transforming the prediction residual is encoded to generate a bitstream.
[0012]
In an image processing device and method according to another aspect of the present technology, a bitstream is decoded, coefficient data obtained by orthogonally transforming a prediction residual of an image is obtained, and a submatrix forming a part of a transformation matrix is used. Then, the transformation matrix is derived, and the derived transformation matrix is used to inverse orthogonally transform the obtained coefficient data.
Effect of the invention
[0013]
According to the present disclosure, images can be processed. In particular, it is possible to suppress an increase in memory capacity required for orthogonal transform and inverse orthogonal transform.
Brief description of the drawings
[0014]
FIG. 1 is a diagram showing a correspondence relationship between a transform set and a selected orthogonal transform.
FIG. 2 is a diagram showing a correspondence relationship between orthogonal transformation types and used functions.
FIG. 3 is a diagram showing a correspondence relationship between transform sets and prediction modes.
FIG. 4 is a diagram showing an example of types of orthogonal transforms stored in a LUT.
FIG. 5 is a diagram showing an example of a LUT size required to hold a conversion matrix in HEVC.
FIG. 6 is a diagram showing an example of a LUT size required to hold a conversion matrix.
FIG. 7 is a diagram illustrating an example of similarity between conversion matrices.
FIG. 8 is a diagram illustrating an example of conversion types that can be replaced by flipping.
FIG. 9 is a diagram illustrating an example of a conversion type that can be replaced by transposition.
FIG. 10 is a diagram showing a list of main alternative specific examples of a transformation matrix.
FIG. 11 is a block diagram showing a main configuration example of an image encoding device.
FIG. 12 is a block diagram showing a main configuration example of an orthogonal transformation unit.
FIG. 13 is a flowchart illustrating an example of the flow of image encoding processing.
FIG. 14 is a flowchart illustrating an example of the flow of orthogonal transform processing.
FIG. 15 is a block diagram illustrating a main configuration example of an image decoding device.
FIG. 16 is a block diagram showing a main configuration example of an inverse orthogonal transform unit.
FIG. 17 is a flowchart illustrating an example of the flow of image decoding processing.
FIG. 18 is a flowchart illustrating an example of the flow of inverse orthogonal transform processing.
FIG. 19 is a diagram showing an example of conversion type derivation.
FIG. 20 is a diagram showing a specific example of conversion type derivation.
FIG. 21 is a diagram showing an example of a LUT size required to hold a conversion matrix.
FIG. 22 is a block diagram showing a main configuration example of a primary conversion unit.
FIG. 23 is a block diagram showing a main configuration example of a primary horizontal conversion unit.
FIG. 24 is a block diagram showing a main configuration example of a conversion matrix derivation unit.
FIG. 25 is a block diagram showing a main configuration example of a primary vertical conversion unit.
FIG. 26 is a block diagram showing a main configuration example of a conversion matrix derivation unit.
FIG. 27 is a flowchart illustrating an example of the flow of primary conversion processing.
FIG. 28 is a flowchart illustrating an example of the flow of primary horizontal conversion processing.
FIG. 29 is a diagram showing an example of an arithmetic expression for each element.
FIG. 30 is a flowchart illustrating an example of the flow of a conversion matrix derivation process.
FIG. 31 is a diagram showing an example of assigning conversion types to conversion type identifiers.
FIG. 32 is a flowchart illustrating an example of the flow of primary vertical conversion processing.
FIG. 33 is a diagram showing an example of an arithmetic expression for each element.
FIG. 34 is a block diagram showing a main configuration example of an inverse primary conversion unit.
FIG. 35 is a block diagram showing a main configuration example of an inverse primary vertical conversion section.
FIG. 36 is a block diagram showing a main configuration example of a conversion matrix derivation unit.
FIG. 37 is a block diagram showing a main configuration example of an inverse primary horizontal conversion unit.
FIG. 38 is a block diagram showing a main configuration example of a conversion matrix derivation unit.
FIG. 39 is a flowchart illustrating an example of the flow of reverse primary conversion processing.
FIG. 40 is a flowchart illustrating an example of the flow of reverse primary conversion selection processing.
FIG. 41 is a flowchart illustrating an example of the flow of reverse primary vertical conversion processing.
FIG. 42 is a flowchart illustrating an example of the flow of reverse primary horizontal conversion processing.
FIG. 43 is a diagram showing an example of conversion type derivation.
FIG. 44 is a diagram showing a specific example of conversion type derivation.
FIG. 45 is a diagram showing an example of a LUT size required to hold a conversion matrix.
FIG. 46 is a block diagram showing a main configuration example of a conversion matrix derivation unit.
FIG. 47 is a block diagram showing a main configuration example of a conversion matrix derivation unit.
FIG. 48 is a block diagram showing a main configuration example of a transformation matrix deriving unit.
FIG. 49 is a diagram showing an example of assigning conversion types to conversion type identifiers.
FIG. 50 is a block diagram showing a main configuration example of a conversion matrix derivation unit.
FIG. 51 is a block diagram showing a main configuration example of a conversion matrix derivation unit.
FIG. 52 is a diagram showing an example of conversion type derivation.
FIG. 53 is a diagram showing an example of a LUT size required for holding a conversion matrix.
FIG. 54 is a flowchart illustrating an example of the flow of a conversion matrix derivation process.
FIG. 55 is a diagram showing an example of assigning conversion types to conversion type identifiers.
FIG. 56 is a diagram showing an example of conversion type derivation.
FIG. 57 is a diagram showing an example of a LUT size required to hold a conversion matrix.
FIG. 58 is a diagram showing an example of assigning conversion types to conversion type identifiers.
FIG. 59 is a diagram showing an example of conversion type derivation.
FIG. 60 is a diagram showing an example of a LUT size required to hold a conversion matrix.
FIG. 61 is a flowchart illustrating an example of the flow of a conversion matrix derivation process.
FIG. 62 is a diagram showing an example of assigning conversion types to conversion type identifiers.
FIG. 63 is a diagram illustrating the spatial symmetry of two-dimensional orthogonal transformation.
FIG. 64 is a diagram for explaining symmetry in the horizontal direction.
FIG. 65 is a diagram for explaining symmetry in the vertical direction.
FIG. 66 is a diagram for explaining symmetry in the horizontal and vertical directions.
FIG. 67 is a diagram showing a list of main specific examples of substitution of a transformation matrix that involves transformation of prediction residuals.
FIG. 68 is a diagram showing an example of a LUT size required to hold a conversion matrix.
FIG. 69 is a block diagram showing a main configuration example of a primary conversion unit.
FIG. 70 is a flowchart illustrating an example of the flow of primary conversion processing.
FIG. 71 is a flowchart illustrating an example of the flow of a prediction residual replacement operation process.
FIG. 72 is a diagram showing an example of assigning conversion types to conversion type identifiers.
FIG. 73 is a block diagram showing a main configuration example of an inverse primary conversion unit.
FIG. 74 is a flowchart illustrating an example of the flow of reverse primary conversion processing.
FIG. 75 is a diagram showing an example of a LUT size required to hold a conversion matrix.
FIG. 76 is a block diagram showing a main configuration example of a conversion matrix derivation unit.
FIG. 77 is a block diagram showing a main configuration example of a conversion matrix derivation unit.
FIG. 78 is a flowchart illustrating an example of the flow of a transformation matrix derivation process.
FIG. 79 is a diagram showing an example of conversion type assignments to conversion type identifiers.
FIG. 80 is a block diagram showing a main configuration example of a conversion matrix derivation unit.
FIG. 81 is a block diagram showing a main configuration example of a conversion matrix derivation unit.
FIG. 82 is a diagram showing an example of assigning conversion types to conversion type identifiers.
FIG. 83 is a diagram showing an example of assigning conversion types to conversion type identifiers.
FIG. 84 is a diagram illustrating derivation of a conversion matrix using a partial matrix.
FIG. 85 is a diagram showing a list of main specific examples of the transformation matrix derived from the partial matrix.
FIG. 86 is a diagram showing an example of how a transformation matrix is derived.
FIG. 87 is a diagram showing an example of how a transformation matrix is derived.
FIG. 88 is a diagram showing an example of how a transformation matrix is derived.
FIG. 89 is a diagram showing an example of how a transformation matrix is derived.
FIG. 90 is a diagram showing an example of how a transformation matrix is derived.
FIG. 91 is a diagram showing an example of how a transformation matrix is derived.
FIG. 92 is a block diagram illustrating a main configuration example of a conversion matrix derivation unit.
FIG. 93 is a block diagram showing a main configuration example of a conversion matrix derivation unit.
FIG. 94 is a flowchart illustrating an example of the flow of a transformation matrix derivation process.
FIG. 95 is a diagram showing an example of assigning conversion types to conversion type identifiers.
FIG. 96 is a block diagram showing a main configuration example of a conversion matrix derivation unit.
FIG. 97 is a block diagram showing a main configuration example of a conversion matrix derivation unit.
FIG. 98 is a diagram showing an example of a LUT size required to hold a conversion matrix.
FIG. 99 is a block diagram illustrating a main configuration example of a conversion matrix derivation unit.
FIG. 100 is a flowchart illustrating an example of the flow of a transformation matrix derivation process.
FIG. 101 is a diagram showing an example of allocation of conversion types to conversion type identifiers.
FIG. 102 is a block diagram illustrating a main configuration example of a computer.
MODE FOR CARRYING OUT THE INVENTION
[0015]
Hereinafter, modes for implementing the present disclosure (hereinafter referred to as embodiments) will be described. The description will be given in the following order.
1. Adaptive primary conversion [0016]
2. First embodiment (conversion matrix derivation from conversion matrix) [0041]
2-1. Common concept
2-2. Example 1-1
2-3. Example 1-2
2-4. Example 1-3
2-5. Example 1-4
2-6. Example 1-5
3. Second embodiment (prediction residual conversion) [0561]
3-1. Common concept
3-2. Example 2-1
3-3. Example 2-2
3-4. Example 2-3
3-5. Example 2-4
4. Third embodiment (conversion matrix derivation from partial matrix) [0734]
4-1. Common concept
5. Fourth embodiment (combination of embodiments) [0840]
5-1. Common concept
6. Note
[0016]
<1. Adaptive primary conversion>
The scope disclosed in the present technology is not limited to the contents described in the embodiments, but also the following non-patent documents known at the time of application. It also includes the contents described in.
[0017]
Non-Patent Document 1: (above)
Non-Patent Document 3: TELECOMMUNICATION STANDARDIZATION SECTOR OF ITU (International Telecommunication Union), "Advanced video coding for generic audiovisual services", H.264, 04/2017
Non-Patent Document 4: TELECOMMUNICATION STANDARDIZATION SECTOR OF ITU (International Telecommunication Union), "High efficiency video coding", H.265, 12/2016
[0018]
That is, the contents described in the above non-patent documents also serve as the basis for determining the support requirement. For example, even if the Quad-Tree Block Structure described in Non-Patent Document 4 and the QTBT (Quad Tree Plus Binary Tree) Block Structure described in Non-Patent Document 1 are not directly described in the examples, the present invention It is within the technical disclosure range and satisfies the support requirements of the claims. Further, for example, technical terms such as Parsing, Syntax, and Semantics are also within the scope of the disclosure of the present technology even when there is no direct description in the examples. Shall meet the support requirements in the range of.
[0019]
Further, in the present specification, a “block” (not a block indicating a processing unit) used as a partial region of an image (picture) or a processing unit for description indicates an arbitrary partial region in a picture unless otherwise specified, The size, shape, characteristics, etc. are not limited. For example, “block” includes TB (Transform Block), TU (Transform Unit), PB (Prediction Block), and PU (Prediction Unit) described in Non-Patent Document 1, Non-Patent Document 3, and Non-Patent Document 4 described above. ), SCU (Smallest Coding Unit), CU (Coding Unit), LCU (Largest Coding Unit), CTB (Coding Tree Block), CTU (Coding Tree Unit), conversion block, subblock, macroblock, tile, or slice , An arbitrary partial area (processing unit) is included.
[0020]
When designating the size of such a block, not only the block size may be designated directly but the block size may be designated indirectly. For example, the block size may be specified using identification information for identifying the size. Further, for example, the block size may be designated by a ratio or a difference from the size of a reference block (for example, LCU or SCU). For example, when transmitting the information designating the block size as the syntax element or the like, the information indirectly designating the size as described above may be used as the information. By doing so, the amount of information can be reduced, and the coding efficiency may be improved in some cases. The block size designation also includes designation of a block size range (for example, designation of an allowable block size range).
[0021]
In
the test model (JEM4 (Joint Exploration Test Model 4)) described in Non-Patent Document 1, the luminance transform block has a horizontal primary transform PThor (also called primary horizontal transform) and a vertical transform. The adaptive primary transform (AMT (Adaptive Multiple Core Transforms)) which adaptively selects the primary transform from a plurality of different one-dimensional orthogonal transforms is disclosed for each primary transform PTver (also referred to as primary vertical transform).
[0022]
Specifically, when the adaptive primary conversion flag apt_flag indicating whether or not to perform adaptive primary conversion is 0 (false) for the luminance conversion block, as in the table (LUT_TrSetToTrTypIdx) shown in FIG. 1, for example, , DCT (Discrete Cosine Transform)-II or DST (Discrete Sine Transform)-VII is uniquely determined by the mode information as the primary transformation (TrSetIdx = 4).
[0023]
When the adaptive primary conversion flag apt_flag is 1 (true) and the current CU (Coding Unit) including the conversion block of the luminance to be processed is an intra CU, as shown in the table in FIG. The transformation set TrSet including orthogonal transformations that are candidates for primary transformation in each of the (x direction) and the vertical direction (y direction) is one of the three transformation sets TrSet (TrSetIdx = 0,1,2) shown in FIG. Selected from. Note that DST-VII, DCT-VIII, and the like shown in FIG. 1 indicate the types of orthogonal transformation, and the functions shown in the table of FIG. 2 are used for each.
[0024]
This transform set TrSet is uniquely determined based on (the intra prediction mode information of) the correspondence table of the transform sets and the mode information shown in FIG. For example, as in the following Expressions (1) and (2), the conversion set identifier TrSetIdx that specifies the corresponding conversion set TrSet is set for each conversion set TrSetH, TrSetV.
[0025]
[Number 1]
[0026]
Here, TrSetH represents the conversion set of the primary horizontal conversion PThor, TrSetV represents the conversion set of the primary vertical conversion PTver, and the lookup table LUT_IntraModeToTrSet is the correspondence table of FIG. The first array of the lookup table LUT_IntraModeToTrSet[][] has the intra prediction mode IntraMode as an argument, and the second array has {H=0, V=1} as an argument.
[0027]
For example, in the case of the intra prediction mode number 19 (IntraMode == 19), as the conversion set TrSetH (also referred to as the primary horizontal conversion set) of the primary horizontal conversion PThor, the conversion set with the conversion set identifier TrSetIdx = 0 shown in the table of FIG. Is selected, and the conversion set having the conversion set identifier TrSetIdx=2 shown in the table of FIG. 1 is selected as the conversion set TrSetV (also referred to as the primary vertical conversion set) of the primary vertical conversion PTver.
[0028]
In addition, when the adaptive primary conversion flag apt_flag is 1 (true) and the current CU including the conversion block of the luminance to be processed is an inter CU, the conversion set TrSetH of the primary horizontal conversion and the conversion set of the primary vertical conversion are set. A conversion set InterTrSet (TrSetIdx = 3) dedicated to inter CU is assigned to TrSetV.
[0029]
Then, for each of the horizontal direction and the vertical direction, which orthogonal transformation is applied in the selected transformation set TrSet is determined by the corresponding one of the primary horizontal transformation designation flag pt_hor_flag and the primary vertical transformation designation flag pt_ver_flag. select.
[0030]
For example, a primary {horizontal, vertical} conversion set TrSet{H,V} and a primary {horizontal, vertical} conversion designation flag pt_{hor,ver}_flag are set as in the following Expressions (3) and (4). Is derived from the conversion set definition table (LUT_TrSetToTrTypeIdx) shown in FIG.
[0031]
[Number 2]
[0032]
For example, when the intra prediction mode number 34 (IntraMode == 34) (that is, the primary horizontal conversion set TraSetH is 0) and the primary horizontal conversion specification flag pt_hor_flag is 0, the conversion set definition table (LUT_TrSetToTrTypeIdx) in FIG. The value of the conversion type identifier TrTypeIdxH in Expression (3) is 4, and the conversion type TrTypeH corresponding to the value of the conversion type identifier TrTypeIdxH is DST-VII with reference to FIG. That is, DST-VII of the conversion set whose conversion set identifier TrSetIdx is 0 is selected as the conversion type of the primary horizontal conversion PThor. When the primary horizontal conversion designation flag pt_hor_flag is 1, DCT-VIII is selected as the conversion type. Note that selecting the conversion type TrType includes selecting the conversion type specified by the conversion type identifier TrTypeIdx via the conversion type identifier TrTypeIdx.
[0033]
The primary conversion identifier pt_idx is derived from the primary horizontal conversion designation flag pt_hor_flag and the primary vertical conversion designation flag pt_ver_flag based on the following equation (5). That is, the upper 1 bit of the primary conversion identifier pt_idx corresponds to the value of the primary vertical conversion designation flag, and the lower 1 bit corresponds to the value of the primary horizontal conversion designation flag.
[0034]
[Number 3]
[0035]
Encoding is performed by applying arithmetic encoding to the bin sequence of the derived primary conversion identifier pt_idx to generate a bit sequence. The adaptive primary conversion flag apt_flag and the primary conversion identifier pt_idx are signaled in the brightness conversion block.
[0036]
As described above, in Non-Patent Document 1, as candidates for primary conversion, DCT-II (DCT2), DST-VII (DST7), DCT-VIII (DCT8), DST-I (DST1), DCT-V (DCT5 ) Five one-dimensional orthogonal transforms have been proposed. In addition, in Non-Patent Document 2, in addition to them, two one-dimensional orthogonal transforms of DST-IV (DST4) and IDT (Identity Transform: one-dimensional transform skip) are added, and a total of seven one-dimensional orthogonal transforms are added. It has been proposed to make the transform a candidate for the primary transform.
[0037]
That is, in the case of Non-Patent Document 1, as shown in FIG. 4, the one-dimensional orthogonal transform is stored in the LUT as a candidate for the primary transform. In addition, in the case of Non-Patent Document 2, in addition to them, DST-IV (DST4) and IDT are also stored in the LUT (see FIG. 4).
[0038]
In the case of HEVC (High Efficiency Video Coding), the size of the LUT (Look Up Table) required to hold the conversion matrix is as shown in the table of FIG. In other words, the LUT size is about 1.3KB in total. On the other hand, in the case of the method described in Non-Patent Document 1, for example, the DCT2 needs to hold a conversion matrix for each size of 2/4/8/16/32/64/128 points on the LUT. Further, for other one-dimensional conversion (DST7/DST1/DCT8), it is necessary to hold a conversion matrix for each size of 4/8/16/32/64 points on the LUT. In this case, assuming that the bit precision of each coefficient of the conversion matrix is 10 bits, the size of the LUT required to hold the entire conversion matrix of the primary conversion is as shown in A of FIG. That is, the size of the LUT in this case is about 53 KB in total. That is, the size of the LUT in this case increases about 50 times as compared with the case of HEVC.
[0039]
Similarly, in the case of the method described in Non-Patent Document 2, the size of the LUT required to hold the entire conversion matrix of the primary conversion is as shown in the table shown in B of FIG. In other words, the LUT size in this case is about 67KB in total. That is, the size of the LUT in this case is increased by about 60 times as compared with the case of HEVC.
[0040]
Considering the hardware implementation of the primary conversion, the size of the LUT is reflected in the storage capacity (memory capacity). That is, in the case of the methods described in Non-Patent Document 1 and Non-Patent Document 2, the circuit scale (memory capacity required to hold the coefficients of the transformation matrix) is increased by about 50 to 60 times as compared with the case of HEVC. There was a risk of
[0041]
<2. First Embodiment>
<2-1. Common Concept>
Then, the second transform matrix is derived using the first transform matrix, and the prediction residual of the image is orthogonally transformed using the derived second transform matrix. , And the coefficient data obtained by orthogonally transforming the prediction residual is encoded to generate a bit stream.
[0042]
For example, in the image processing apparatus, the derivation unit that derives the second transformation matrix using the first transformation matrix and the second transformation matrix derived by the derivation unit are used to orthogonalize the prediction residual of the image. An orthogonal transform unit for transforming, and a coding unit for encoding coefficient data obtained by orthogonally transforming the prediction residual by the orthogonal transform unit to generate a bit stream are provided.
[0043]
By doing so, the transformation matrix can be derived from another transformation matrix, so that it is possible to suppress an increase in the number of transformation matrices prepared for orthogonal transformation, and to reduce the memory capacity required for orthogonal transformation. The increase can be suppressed.
[0044]
Also, the bitstream is decoded, the coefficient data in which the prediction residual of the image is orthogonally transformed is obtained, the second transformation matrix is derived using the first transformation matrix, and the derived second transformation matrix is obtained. Is used to perform inverse orthogonal transform on the obtained coefficient data.
[0045]
For example, in an image processing device, a decoding unit that decodes a bitstream and obtains coefficient data in which the prediction residual of an image is orthogonally transformed, and a deriving unit that derives a second transformation matrix using the first transformation matrix. , And an inverse orthogonal transform unit that performs an inverse orthogonal transform of the coefficient data obtained by the decoding unit using the second transform matrix derived by the deriving unit.
[0046]
By doing so, the transformation matrix can be derived from another transformation matrix, so that the increase in the number of transformation matrices prepared for the inverse orthogonal transformation can be suppressed, and the memory required for the inverse orthogonal transformation can be suppressed. It is possible to suppress an increase in capacity.
[0047]
One of the main roles of the conversion matrix is to bias the signals of low-order (especially 0th-order) frequency components toward the direction of the DC component. It becomes an important characteristic. The waveform component of the low-order (especially 0th-order) basis vector (row vector) is important for how the frequency components are biased. In other words, transform matrices that have similar tendencies in the waveform components of the basis vectors can be expected to have similar performance with respect to orthogonal transform and inverse orthogonal transform (how the frequency components are biased is similar. Exist).
[0048]
Therefore, attention is paid to the waveform of the low-order (particularly 0-order) basis vector (row vector) of the transformation matrix. For example, in the conversion matrix 30 of FIG. 7, it is assumed that a low-order (particularly 0-order) row vector waveform (trend of the value of each element) in the frame 31 is shown as a graph 32.
[0049]
The graph 32 shows (the tendency of) the element of the low frequency component on the left side of the figure, and shows the value (trend of) the element of the high frequency component on the right side of the figure. Further, the graph 32 shows that the higher the value in the figure, the larger the value, and the lower the figure, the smaller the value. Note that the vertical center of the graph 32 indicates 0, the upper side of the center indicates a positive value, and the lower side of the center indicates a negative value.
[0050]
The waveform 32A of the graph 32 indicates the waveform of the 0th-order row vector of the conversion matrix 30. As shown by the waveform 32A, in this case, the 0th-order row vector of the conversion matrix 30 tends to increase in value from the low-frequency component to the high-frequency component.
[0051]
Further, for example, in the conversion matrix 30 of FIG. 7, it is assumed that the waveform of the low-order (particularly 0-order) column vector (trend of the value of each element) in the frame 33 is shown as a graph 34. The graph 34 shows (the tendency of) the element of the low frequency component in the upper part of the figure, and shows the value (trend of the element) of the high frequency component in the lower part of the figure. Further, in the graph 34, the value is larger on the left side in the figure, and the value is smaller on the right side in the figure. The center of the graph 34 in the left-right direction indicates 0, the left side of the center indicates a positive value, and the right side of the center indicates a negative value.
[0052]
The waveform 34A of the graph 34 indicates the waveform of the 0th-order column vector of the conversion matrix 30. As shown in the waveform 34A, in this case, the 0th-order column vector of the conversion matrix 30 has a peak in the intermediate frequency component (that is, the value decreases toward the low frequency component on the low frequency side, On the high frequency side, the value tends to decrease toward higher frequency components).
[0053]
In the present specification, the waveform of the 0th-order column vector of the conversion matrix 30 may be represented in a transposed state as shown in the graph 35. The structure of the graph 35 is similar to that of the graph 32. Waveform 35A is equivalent to waveform 34A.
[0054]
As described above, such transformation matrices having similar waveforms of low-order (particularly zero-order) basis vectors (row vectors) have similar performance. In other words, the transformation matrix can be replaced by another transformation matrix having a low-order (particularly zero-order) basis vector (row vector) having a similar waveform. Therefore, by using this, it is possible to suppress an increase in the number of conversion matrices stored in the LUT.
[0055]
Here, focusing on the conversion types described in Non-Patent Document 1 and Non-Patent Document 2, the waveforms of the zero-order row vector and the zero-order column vector of the conversion matrix of these conversion types are four types (types). Can be classified into. FIG. 8 shows an example thereof.
[0056]
The first type is a flat type. The flat type is a waveform type in which the values are substantially uniform in each frequency component. The second type is the augmented type. This increasing type is a type of waveform in which the value tends to increase from the low frequency component to the high frequency component. The third type is the reduction type. This decreasing type is a type of waveform in which the value tends to decrease from the low frequency component to the high frequency component. The fourth type is a mountain type. The mountain shape is a type of waveform that tends to have a peak (maximum value) in the middle. That is, in the case of this mountain shape, the waveform tends to decrease in value toward the low frequency component on the low frequency component side and decrease in value toward the high frequency component on the high frequency component side.
[0057]
It should be noted that each of these types represents the approximate shape of the waveform and does not have to match exactly. For example, in the case of the increase type, it is sufficient that the waveform as a whole tends to increase in value from the low frequency side to the high frequency side, and even if it does not strictly increase monotonically from the low frequency side to the high frequency side. Good.
[0058]
Similarly, in the case of the reduction type, it is sufficient that the waveform as a whole tends to decrease in value from the low frequency side to the high frequency side, and it does not strictly decrease monotonically from the low frequency side to the high frequency side. Good.
[0059]
Similarly, in the case of the mountain shape, the waveform has a peak (maximum value) as a whole near the center, and it is sufficient that the value tends to decrease in the direction away from the peak on both sides of the peak. The value does not have to monotonically decrease in the direction away from. Further, the peak does not have to be formed by one component, and may be such that the approximate position or value of the peak can be specified from a plurality of components. The position of the peak does not have to be strictly in the center.
[0060]
Similarly, in the case of the flat type, the waveform need only have substantially uniform values as a whole, and does not need to be strictly flat. That is, the values may have some variations. In other words, a flat type may be used if it cannot be classified into the other three types.
[0061]
The above classification of waveforms is an example, and the present invention is not limited to the above example. That is, the waveforms may be classified into types other than those described above, the number of types to be classified is arbitrary, and is not limited to the above four types. Note that this classification is performed for convenience of description of the present technology, and is not performed as an actual process.
[0062]
According to this classification, as shown in FIG. 8, the waveform of the 0th-order row vector of the transform matrix of DCT2 is classified into the flat type, and the waveform of the 0th-order column vector is classified into the decreasing type. Further, the waveform of the 0th-order row vector of the DST7 conversion matrix is classified into the increasing type, and the waveform of the 0th-order column vector is classified into the mountain type. Further, the waveform of the 0th-order row vector of the transform matrix of the DCT8 is classified into the decreasing type, and the waveform of the 0th-order column vector is classified into the decreasing type. Further, the waveform of the zero-order row vector of the transform matrix of DCT5 is classified into the flat type, and the waveform of the zero-order column vector is classified into the flat type. The waveform of the 0th-order row vector of the transformation matrix of DST4 is classified into the increasing type, and the waveform of the 0th-order column vector is classified into the increasing type.
[0063]
As described above, the transformation matrix can be replaced by another transformation matrix having similar zero-order row vector waveforms. That is, the conversion types having the same waveform of the 0th-order row vector can substitute the conversion matrix with each other.
[0064]
That is, when the derivation unit derives the second transformation matrix using the above-mentioned first transformation matrix, the lowest-order row vector derives the second transformation matrix having a desired type of waveform. Good. By doing so, it is possible to suppress an increase in the number of transform matrices prepared for the orthogonal transform/inverse orthogonal transform, and to suppress an increase in the memory capacity required for the orthogonal transform/inverse orthogonal transform. ..
[0065]
For example, the deriving unit may use the first conversion matrix to derive the second conversion matrix whose lowest-order row vector has a flat waveform. By doing so, the derived second conversion matrix can replace the conversion matrix in which the waveform of the lowest-order row vector is flat. Further, for example, the deriving unit may use the first conversion matrix to derive the second conversion matrix in which the lowest-order row vector has an increasing waveform. By doing so, the derived second conversion matrix can replace the conversion matrix in which the waveform of the lowest-order row vector increases.
[0066]
Further, for example, the deriving unit may use the first conversion matrix to derive the second conversion matrix in which the lowest-order row vector has a decreasing waveform. By doing so, the derived second conversion matrix can replace the conversion matrix in which the waveform of the lowest-order row vector is decreasing. Further, for example, the derivation unit may use the first transformation matrix to derive the second transformation matrix in which the lowest-order row vector has a mountain-shaped waveform. By doing so, the derived second transformation matrix can replace the transformation matrix in which the waveform of the lowest-order row vector is mountain-shaped.
[0067]
For example, in FIG. 8, since DST7, DST4, DST8, and DST3 all have the increasing waveform of the 0th-order row vector, they can mutually substitute the conversion matrix. That is, even conversion matrices of different conversion types can be replaced.
[0068]
That is, the deriving unit may derive a second conversion matrix of a conversion type different from the first conversion matrix when deriving the second conversion matrix using the first conversion matrix described above. Good. By doing so, it is possible to suppress an increase in the number of conversion types prepared for the orthogonal transform/inverse orthogonal transform, and to suppress an increase in the memory capacity required for the orthogonal transform/inverse orthogonal transform. ..
[0069]
In the derivation, the derivation unit may derive the second conversion matrix having the same number of rows and the same number of columns as the first conversion matrix. When changing the number of rows and columns, the type of waveform may change unintentionally. Therefore, by setting the number of rows and the number of columns to be the same as those of the first conversion matrix, it is possible to suppress the possibility of an unintended change in the waveform type, and it is easier to derive the second conversion matrix. be able to.
[0070]
It should be noted that the matrix can be easily operated on the elements. The operations on the elements of the matrix include, for example, rearrangement of the elements. More specifically, for example, in a matrix, the arrangement order of the element groups can be flipped (inverted) in a predetermined direction, or the element groups can be transposed to replace rows and columns. Note that transposition is equivalent to flip (reverse) with a diagonal line connecting the upper left end and the lower right end of the matrix as an axis. That is, transposition can be said to be a part of flip. It is also easy to invert the sign of each element (positive to negative, negative to positive).
[0071]
By using such an operation, the waveform (type) of the zero-order row vector can be intentionally changed. For example, when a zero-order row vector waveform has an increasing matrix flipped in the row direction, the zero-order row vector waveform changes to a decreasing waveform. On the contrary, when the matrix of the 0th-order row vector waveform is flipped in the row direction, the waveform of the 0th-order row vector changes to the increasing type.
[0072]
That is, as shown in FIG. 8, by flipping the conversion matrix of the conversion type (for example, DST7, DST4, DST8, DST3) in which the waveform of the 0th-order row vector is an increasing type, the 0th-order row vector is flipped. It is possible to obtain a conversion matrix that can replace a conversion matrix of a conversion type in which the waveform of the vector is a reduction type (for example, DCT7, DCT4, DCT8, DCT3).
[0073]
Further, for example, when the matrix is transposed, the waveform type of the 0th-order row vector and the waveform type of the 0th-order column vector are exchanged. That is, by transposition, the waveform of the zero-order row vector of the matrix becomes the same type as the waveform of the zero-order column vector of the matrix before transposition.
[0074]
For example, as shown in FIG. 9, by transposing the transformation matrix of DCT2 (DCT6) in which the waveform of the 0th-order column vector is a decreasing type, the transformation type of the 0th-order row vector is a decreasing type (DCT3). , DCT7, DCT4, DCT8) can be obtained as an alternative transformation matrix. Further, for example, as shown in FIG. 9, by transposing the transformation matrix of DST7 (DST3) in which the waveform of the 0th-order column vector is mountain-shaped, the waveform of the 0th-order row vector is mountain-shaped. A conversion matrix that can replace the conversion matrix of (DST2, DST6, DST1, DST5) can be obtained.
[0075]
That is, the derivation unit may derive the second conversion matrix by performing an operation on the elements of the first conversion matrix as described above. The operation for this element may include rearrangement of elements (change of arrangement order) as described above. By doing so, the type of the waveform can be intentionally changed, and thus more diverse second conversion matrices can be derived from the first conversion matrix. Therefore, it is possible to suppress an increase in the number of conversion types prepared for the orthogonal transform/inverse orthogonal transform, and to suppress an increase in the memory capacity required for the orthogonal transform/inverse orthogonal transform.
[0076]
Note that, of course, the derivation unit may perform such an operation multiple times to derive the second conversion matrix. For example, operations such as flip and transpose can be arbitrarily combined. The same operation may be repeated multiple times. By doing so, it is possible to derive more various second conversion matrices from the first conversion matrix.
[0077]
As described above, the transformation matrix used for the orthogonal transformation/inverse orthogonal transformation is stored in the LUT. Therefore, the derivation unit may derive the second conversion matrix using the first conversion matrix stored in the lookup table (LUT). By doing so, an increase in the size of the LUT can be suppressed. Therefore, it is possible to suppress an increase in the memory capacity required for orthogonal transformation/inverse orthogonal transformation.
[0078]
FIG. 10 shows a list of derivation examples of the transformation matrix involving the above operations. The conversion matrix (first conversion matrix) used for derivation is also referred to as a base conversion matrix T base . The conversion type of the base conversion matrix is also referred to as a base conversion type or a first conversion type. Further, the conversion type of the derived conversion matrix (second conversion matrix) is also referred to as a second conversion type.
[0079]
In the table shown in FIG. 10, except for the row of the item name in the top row, the derivation of the example from the first row from the top shows the waveform of the lowest-order row vector of the first transformation type and the transformation type of the transformation matrix to substitute. It focuses on the similarity with the waveform of the lowest order row vector.
[0080]
In this case, the derivation unit flips the first conversion matrix and derives the second conversion matrix. More specifically, the deriving unit uses the transformation matrix of DST7 as the base transformation matrix T base , flips it in the row direction, and derives the transformation matrix of FlipDST7. Since the waveform of the 0th-order row vector of the DST7 conversion matrix is the increasing type, the waveform of the 0th-order row vector of the derived FlipDST7 conversion matrix is the decreasing type. Therefore, the conversion matrix of FlipDST7 can replace the conversion matrix of DCT8 in which the waveform of the zero-order row vector is reduced.
[0081]
By applying such derivation, it is not necessary to prepare a transform matrix of DCT8 as a candidate of transform matrix used for orthogonal transform/inverse orthogonal transform. That is, the number of unique conversion types can be reduced. That is, an increase in LUT size can be suppressed. In addition, by performing orthogonal transformation/inverse orthogonal transformation using the derived transformation matrix of the second transformation type (FlipDST7), the same degree of coding as when the DCT8 transformation matrix is used for orthogonal transformation/inverse orthogonal transformation Efficiency is obtained. Also, in this case, the second conversion matrix (reduction conversion matrix) can be derived by a single operation.
[0082]
The derivation of the example one step below (second from the top) is similar to the waveform of the lowest-order row vector of the transformation matrix of the alternative transformation matrix and the waveform of the lowest-order column vector of the first transformation type. It focuses on sex.
[0083]
In this case, the derivation unit transposes the first transformation matrix to derive the second transformation matrix. More specifically, the deriving unit uses the transformation matrix of DST7 as the base transformation matrix T base, and transposes it to derive the transformation matrix of TrDST7. Since the waveform of the 0th-order column vector of the DST7 conversion matrix is mountain-shaped, the waveform of the 0th-order row vector of the derived TrDST7 conversion matrix is mountain-shaped. Therefore, the conversion matrix of TrDST7 can replace the conversion matrix of DST1 in which the waveform of the 0th row vector is mountain-shaped.
[0084]
By applying such derivation, it is not necessary to prepare a DST1 transform matrix as a candidate for a transform matrix used for orthogonal transform/inverse orthogonal transform. That is, the number of unique conversion types can be reduced. That is, an increase in LUT size can be suppressed. Also, by performing orthogonal transformation/inverse orthogonal transformation using the derived transformation matrix of the second transformation type (TrDST7), the same degree of coding as when the transformation matrix of DST1 is used for orthogonal transformation/inverse orthogonal transformation is performed. Efficiency is obtained. Further, in this case, the second conversion matrix (mountain type alternative conversion matrix) can be derived by one operation.
[0085]
The derivation of the example one step below (third step from the top) focuses on the characteristics between the paired DCT/DST. More specifically, between the paired DCT/DST (for example, DST7 and DCT8), the even-numbered row vector is axisymmetric, and the odd-numbered row vector is point-symmetrical.
[0086]
In this case, the derivation unit flips the first conversion matrix, inverts the sign of the odd-numbered row vector in the flipped first conversion matrix, and derives the second conversion matrix. More specifically, the deriving unit takes the transformation matrix of DST7 as the base transformation matrix T base , flips it in the row direction, and further inverts the sign of the odd-ordered row vector to derive the transformation matrix of DCT8. .. The sign inversion of the row vector may be performed by converting the most significant bit of each element of the row vector. Of course, the derived transformation matrix of DCT8 can replace the transformation matrix of DST8 in which the waveform of the zero-order row vector is reduced.
[0087]
By applying such derivation, it is not necessary to prepare a transform matrix of DCT8 as a candidate of transform matrix used for orthogonal transform/inverse orthogonal transform. That is, the number of unique conversion types can be reduced. That is, an increase in LUT size can be suppressed. Also, of course, the same code as when the DCT8 transform matrix is used for orthogonal transform/inverse orthogonal transform by performing orthogonal transform/inverse orthogonal transform using the derived transform matrix of the second transform type (DCT8). Efficiency can be obtained. Further, in this case, the second conversion matrix (paired alternative conversion matrix) can be derived by performing the operation twice.
[0088]
Further, the derivation of the example one step below (fourth step from the top) is similar to the case of the first step from the top, and conversion of a conversion matrix that substitutes for the waveform of the lowest-order row vector of the first conversion type. It focuses on the similarity with the waveform of the lowest order row vector of the type.
[0089]
In this case, the derivation unit flips the first conversion matrix and derives the second conversion matrix. More specifically, the deriving unit uses the transform matrix of DCT8 as the base transform matrix T base , flips it in the row direction, and derives the transform matrix of FlipDCT8. Since the waveform of the 0th-order row vector of the transform matrix of DCT8 is a decreasing type, the waveform of the 0th-order row vector of the derived transform matrix of FlipDCT8 is an increasing type. Therefore, the conversion matrix of FlipDCT8 can replace the conversion matrix of DST7 in which the waveform of the 0th-order row vector increases.
[0090]
By applying such derivation, there is no need to prepare a DST7 transform matrix as a candidate for a transform matrix used for orthogonal transform/inverse orthogonal transform. That is, the number of unique conversion types can be reduced. That is, an increase in LUT size can be suppressed. Also, by performing orthogonal transformation/inverse orthogonal transformation using the derived transformation matrix of the second transformation type (FlipDCT8), the same degree of encoding as when the transformation matrix of DST7 is used for orthogonal transformation/inverse orthogonal transformation is performed. Efficiency is obtained. Also, in this case, the second conversion matrix (reduction conversion matrix) can be derived by a single operation.
[0091]
Also, the derivation of the example one step below (fifth step from the top) is similar to the waveform of the highest-order column vector of the first transformation type and the waveform of the lowest-order row vector of the transformation type of the transformation matrix to substitute. It focuses on sex.
[0092]
In this case, the derivation unit flips the first conversion matrix and transposes the flipped first conversion matrix to derive the second conversion matrix. More specifically, the deriving unit derives the transformation matrix of TrFlipDCT8 by using the transformation matrix of DCT8 as the base transformation matrix T base , flipping it in the row direction, and further transposing it. The waveform of the highest-order column vector of the transformation matrix of DCT8 is mountain-shaped, and thus the waveform of the 0th-order row vector of the derived transformation matrix of TrFlipDCT8 is mountain-shaped. Therefore, the transformation matrix of TrFlipDCT8 can replace the transformation matrix of DST1 in which the waveform of the zero-order row vector is mountain-shaped.
[0093]
By applying such derivation, it is not necessary to prepare a DST1 transform matrix as a candidate for a transform matrix used for orthogonal transform/inverse orthogonal transform. That is, the number of unique conversion types can be reduced. That is, an increase in LUT size can be suppressed. Also, by performing orthogonal transformation/inverse orthogonal transformation using the derived transformation matrix of the second transformation type (TrFlipDCT8), the same degree of encoding as when the transformation matrix of DST1 is used for orthogonal transformation/inverse orthogonal transformation is performed. Efficiency is obtained. Further, in this case, the second conversion matrix (mountain-shaped alternative conversion matrix) can be derived by two operations.
[0094]
Further, the derivation of the example one stage below (the sixth stage from the top) focuses on the characteristics between the paired DCT/DST, as in the case of the third stage from the top. More specifically, between the paired DCT/DST (for example, DCT8 and DST7), the even-numbered row vector is axisymmetric and the odd-numbered row vector is point-symmetrical.
[0095]
In this case, the derivation unit flips the first conversion matrix, inverts the sign of the odd-numbered row vector in the flipped first conversion matrix, and derives the second conversion matrix. More specifically, the derivation unit derives the transformation matrix of DST7 by setting the transformation matrix of DCT8 as the base transformation matrix T base , flipping it in the row direction, and further inverting the sign of the odd-ordered row vector. .. Of course, the derived DST7 conversion matrix can replace the DST7 conversion matrix in which the waveform of the 0th-order row vector increases.
[0096]
By applying such derivation, there is no need to prepare a DST7 transform matrix as a candidate for a transform matrix used for orthogonal transform/inverse orthogonal transform. That is, the number of unique conversion types can be reduced. That is, an increase in LUT size can be suppressed. In addition, naturally, the same code as when the DST7 transform matrix is used for orthogonal transform/inverse orthogonal transform is performed by performing orthogonal transform/inverse orthogonal transform using the derived transform matrix of the second transform type (DST7). Efficiency can be obtained. Further, in this case, the second conversion matrix (paired alternative conversion matrix) can be derived by performing the operation twice.
[0097]
Further, the derivation of the example one step below (the seventh step from the top) is similar to the waveform of the lowest-order column vector of the first transformation type and the waveform of the lowest-order row vector of the transformation type of the transformation matrix to substitute. It focuses on sex.
[0098]
In this case, the derivation unit transposes the first transformation matrix to derive the second transformation matrix. More specifically, the deriving unit uses the transform matrix of DCT2 as the base transform matrix T base, and transposes it to derive the transform matrix of DCT3. Since the waveform of the 0th-order column vector of the transform matrix of DCT2 is the decreasing type, the waveform of the 0th-order row vector of the derived transform matrix of DCT3 is the decreasing type. Therefore, the DCT3 transform matrix can replace the DCT8 transform matrix in which the waveform of the zero-order row vector is reduced.
[0099]
By applying such derivation, it is not necessary to prepare a transform matrix of DCT8 as a candidate of transform matrix used for orthogonal transform/inverse orthogonal transform. That is, the number of unique conversion types can be reduced. That is, an increase in LUT size can be suppressed. In addition, by performing orthogonal transformation/inverse orthogonal transformation using the derived transformation matrix of the second transformation type (DCT3), the same degree of encoding as when the transformation matrix of DCT8 is used for orthogonal transformation/inverse orthogonal transformation is performed. Efficiency is obtained. Also, in this case, the second conversion matrix (reduction conversion matrix) can be derived by a single operation.
[0100]
In addition, the derivation of the example one step lower (8th step from the top) is similar to the waveform of the highest-order column vector of the first transformation type and the waveform of the lowest-order row vector of the transformation type of the transformation matrix to substitute. It focuses on sex.
[0101]
In this case, the derivation unit transposes the first transformation matrix, flips the transposed first transformation matrix, and derives the second transformation matrix. More specifically, the deriving unit derives the transformation matrix of FlipDCT3 by using the transformation matrix of DCT2 as the base transformation matrix T base , transposing it, and further flipping it in the row direction. Since the waveform of the highest-order column vector of the transform matrix of DCT2 is the increasing type, the waveform of the 0th-order row vector of the derived transform matrix of FlipDCT3 is the increasing type. Therefore, the conversion matrix of FlipDCT3 can replace the conversion matrix of DST7 in which the waveform of the 0th-order row vector increases.
[0102]
By applying such derivation, there is no need to prepare a DST7 transform matrix as a candidate for a transform matrix used for orthogonal transform/inverse orthogonal transform. That is, the number of unique conversion types can be reduced. That is, an increase in LUT size can be suppressed. In addition, by performing orthogonal transformation/inverse orthogonal transformation using the derived transformation matrix of the second transformation type (FlipDCT3), the same degree of encoding as when the transformation matrix of DST7 is used for orthogonal transformation/inverse orthogonal transformation is performed. Efficiency is obtained. Further, in this case, the second conversion matrix (increasing type alternative conversion matrix) can be derived by two operations.
[0103]
It should be noted that each of the above-described derivation examples may be implemented independently, or a plurality of derivation examples may be combined and implemented.
[0104]
Next, a configuration for deriving the above transformation matrix will be described. FIG. 11 is a block diagram showing an example of the configuration of an image encoding device that is one aspect of an image processing device to which the present technology is applied. The image encoding device 100 shown in FIG. 11 is a device that encodes image data of a moving image. For example, the image encoding device 100 implements the technology described in Non-Patent Document 1, Non-Patent Document 3, or Non-Patent Document 4, and uses a method in conformity with the standard described in any of those documents. The image data of a moving image is encoded.
[0105]
Note that FIG. 11 shows main components such as a processing unit and a data flow, and the components shown in FIG. 11 are not necessarily all. That is, in the image encoding device 100, a processing unit not shown as a block in FIG. 11 may exist, or a process or data flow not shown as an arrow or the like in FIG. 11 may exist. This also applies to other drawings that describe the processing unit and the like in the image encoding device 100.
[0106]
As shown in FIG. 11, the image encoding device 100 includes a control unit 101, a rearrangement buffer 111, a calculation unit 112, an orthogonal transformation unit 113, a quantization unit 114, an encoding unit 115, an accumulation buffer 116, and an inverse quantization unit. 117, an inverse orthogonal transform unit 118, a calculation unit 119, an in-loop filter unit 120, a frame memory 121, a prediction unit 122, and a rate control unit 123.
[0107]
The
control unit 101 controls the moving image data held by the rearrangement buffer 111 based on the block size of an external unit or a pre-specified processing unit to be a processing unit block (CU, PU, conversion block, etc.). ). Further, the control unit 101 determines the coding parameters (header information Hinfo, prediction mode information Pinfo, conversion information Tinfo, filter information Finfo, etc.) to be supplied to each block, for example, based on RDO (Rate-Distortion Optimization). To do.
[0108]
Details of these encoding parameters will be described later. When the control unit 101 determines the above encoding parameters, it supplies them to each block. Specifically, it is as follows.
[0109]
The header information Hinfo is supplied to each block.
The prediction mode information Pinfo is supplied to the encoding unit 115 and the prediction unit 122.
The transform information Tinfo is supplied to the encoding unit 115, the orthogonal transform unit 113, the quantization unit 114, the inverse quantization unit 117, and the inverse orthogonal transform unit 118.
The filter information Finfo is supplied to the in-loop filter unit 120.
[0110]
Each field (input image) of moving image data is input to the image encoding device 100 in the reproduction order (display order). The rearrangement buffer 111 acquires each input image in the reproduction order (display order) and holds (stores) it. Under the control of the control unit 101, the rearrangement buffer 111 rearranges the input image in the encoding order (decoding order) or divides it into blocks of processing units. The rearrangement buffer 111 supplies each processed input image to the calculation unit 112. The rearrangement buffer 111 also supplies each input image (original image) to the prediction unit 122 and the in-loop filter unit 120.
[0111]
The
calculation unit 112 receives the image I corresponding to the block of the processing unit and the predicted image P supplied from the prediction unit 122, and calculates the predicted image P from the image I by the following formula (6). Thus, the prediction residual D is derived and supplied to the orthogonal transform unit 113.
[0112]
[Number 4]
[0113]
The
orthogonal transform unit 113 receives the prediction residual D supplied from the arithmetic unit 112 and the conversion information Tinfo supplied from the control unit 101 as input, and based on the conversion information Tinfo, the prediction residual The orthogonal transformation is performed on D to derive the transformation coefficient Coeff. The orthogonal transform unit 113 supplies the obtained transform coefficient Coeff to the quantization unit 114.
[0114]
The
quantization unit 114 receives the transform coefficient Coeff supplied from the orthogonal transform unit 113 and the transform information Tinfo supplied from the control unit 101, and based on the transform information Tinfo, the transform coefficient Coeff. Is scaled (quantized). The quantization rate is controlled by the rate controller 123. The quantization unit 114 supplies the quantized transform coefficient obtained by such quantization, that is, the quantized transform coefficient level level, to the encoding unit 115 and the dequantization unit 117.
[0115]
The
encoding unit 115 includes the quantized transform coefficient level level supplied from the quantization unit 114 and various encoding parameters (header information Hinfo, prediction mode information Pinfo, conversion information) supplied from the control unit 101. Tinfo, filter information Finfo, etc.), information about filters such as filter coefficients supplied from the in-loop filter unit 120, and information about the optimum prediction mode supplied from the prediction unit 122. The encoding unit 115 performs variable-length encoding (for example, arithmetic encoding) on the quantized transform coefficient level level to generate a bit string (encoded data).
[0116]
Also, the encoding unit 115 derives residual information Rinfo from the quantized transform coefficient level level, encodes the residual information Rinfo, and generates a bit string.
[0117]
Further, the encoding unit 115 includes the information about the filter supplied from the in-loop filter unit 120 in the filter information Finfo and the information about the optimum prediction mode supplied from the prediction unit 122 in the prediction mode information Pinfo. Then, the encoding unit 115 encodes the above-described various encoding parameters (header information Hinfo, prediction mode information Pinfo, conversion information Tinfo, filter information Finfo, etc.) to generate a bit string.
[0118]
Also, the encoding unit 115 multiplexes the bit strings of various information generated as described above to generate encoded data. The encoding unit 115 supplies the encoded data to the accumulation buffer 116.
[0119]
The
accumulation buffer 116 temporarily holds the encoded data obtained by the encoding unit 115. The accumulation buffer 116 outputs the coded data retained therein at a predetermined timing, for example, as a bit stream to the outside of the image coding apparatus 100. For example, this encoded data is transmitted to the decoding side via an arbitrary recording medium, an arbitrary transmission medium, an arbitrary information processing device, or the like. That is, the accumulation buffer 116 is also a transmission unit that transmits encoded data (bit stream).
[0120]
The
inverse quantization unit 117 performs processing related to inverse quantization. For example, the inverse quantization unit 117 receives the quantized conversion coefficient level level supplied from the quantization unit 114 and the conversion information Tinfo supplied from the control unit 101, and quantizes based on the conversion information Tinfo. The value of the transform coefficient level level is scaled (dequantized). The inverse quantization is an inverse process of the quantization performed by the quantization unit 114. The inverse quantization unit 117 supplies the transform coefficient Coeff_IQ obtained by such inverse quantization to the inverse orthogonal transform unit 118.
[0121]
The
inverse orthogonal transform unit 118 performs a process related to the inverse orthogonal transform. For example, the inverse orthogonal transform unit 118 receives the transform coefficient Coeff_IQ supplied from the dequantization unit 117 and the transform information Tinfo supplied from the control unit 101, and based on the transform information Tinfo, converts the transform coefficient Coeff_IQ into a transform coefficient Coeff_IQ. Inverse orthogonal transformation is performed for the prediction residual D'. The inverse orthogonal transform is an inverse process of the orthogonal transform performed by the orthogonal transform unit 113. The inverse orthogonal transformation unit 118 supplies the prediction residual D′ obtained by such inverse orthogonal transformation to the calculation unit 119. Since the inverse orthogonal transform unit 118 is the same as the inverse orthogonal transform unit on the decoding side (described later), the description on the decoding side (described later) can be applied to the inverse orthogonal transform unit 118.
[0122]
The
calculation unit 119 receives the prediction residual D′ supplied from the inverse orthogonal transform unit 118 and the prediction image P supplied from the prediction unit 122 as inputs. The calculation unit 119 adds the prediction residual D′ and the prediction image P corresponding to the prediction residual D′ to derive the local decoded image R local . The calculation unit 119 supplies the derived locally decoded image R local to the in-loop filter unit 120 and the frame memory 121.
[0123]
The
in-loop filter section 120 performs processing relating to in-loop filter processing. For example, the in-loop filter unit 120 receives the locally decoded image R local supplied from the calculation unit 119, the filter information Finfo supplied from the control unit 101, and the input image (original image) supplied from the rearrangement buffer 111. Is input. The information input to the in-loop filter unit 120 is arbitrary, and information other than this information may be input. For example, prediction mode, motion information, code amount target value, quantization parameter QP, picture type, block (CU, CTU, etc.) information, etc. may be input to the in-loop filter unit 120 as necessary. Good.
[0124]
The in-loop filter unit 120 appropriately performs filter processing on the local decoded image R local based on the filter information Finfo . The in-loop filter unit 120 also uses the input image (original image) and other input information for the filtering process as necessary.
[0125]
For example, as described in Non-Patent Document 1, the in-loop filter unit 120 includes a bilateral filter, a deblocking filter (DBF (DeBlocking Filter)), an adaptive offset filter (SAO (Sample Adaptive Offset)), and an adaptive loop filter. The four in-loop filters (ALF (Adaptive Loop Filter)) are applied in this order. Note that which filter is applied and in what order is arbitrary, and can be appropriately selected.
[0126]
Of course, the filtering process performed by the in-loop filter unit 120 is arbitrary and is not limited to the above example. For example, the in-loop filter unit 120 may apply a Wiener filter or the like.
[0127]
The in-loop filter unit 120 supplies the filtered locally decoded image R local to the frame memory 121. Note that, for example, when transmitting information regarding a filter such as a filter coefficient to the decoding side, the in-loop filter unit 120 supplies the information regarding the filter to the encoding unit 115.
[0128]
The
frame memory 121 performs processing related to storage of data related to images. For example, the frame memory 121, local decoded image R is supplied from the operation unit 119 local or filtered local decoded image R is supplied from the in-loop filter unit 120 local as input, hold it (memory) .. Further, the frame memory 121 reconstructs the decoded image R for each picture using the locally decoded image R local and holds it (stores it in the buffer in the frame memory 121). The frame memory 121 supplies the decoded image R (or a part thereof) to the prediction unit 122 in response to a request from the prediction unit 122.
[0129]
The
prediction unit 122 performs a process related to generation of a predicted image. For example, the prediction unit 122 includes the prediction mode information Pinfo supplied from the control unit 101, the input image (original image) supplied from the rearrangement buffer 111, and the decoded image R (or part thereof) read from the frame memory 121. Is input. The prediction unit 122 performs prediction processing such as inter prediction and intra prediction using the prediction mode information Pinfo and the input image (original image), performs prediction by referring to the decoded image R as a reference image, and based on the prediction result. Motion compensation processing is performed to generate a predicted image P. The prediction unit 122 supplies the generated predicted image P to the calculation unit 112 and the calculation unit 119. In addition, the prediction unit 122 supplies the prediction mode selected by the above processing, that is, information on the optimum prediction mode, to the encoding unit 115 as necessary.
[0130]
The
rate control unit 123 performs processing relating to rate control. For example, the rate control unit 123 controls the rate of the quantization operation of the quantization unit 114 based on the code amount of the encoded data accumulated in the accumulation buffer 116 so that overflow or underflow does not occur.
[0131]
In the image encoding device 100 having the above configuration, the orthogonal transformation unit 113 performs the processing to which the above-described present technology is applied as the derivation unit and the orthogonal transformation unit. In addition, the encoding unit 115, as an encoding unit, performs processing to which the above-described present technology is applied. Further, the inverse orthogonal transform unit 118 performs processing to which the above-described present technology is applied as an inverse orthogonal transform unit and a derivation unit. Therefore, the image coding apparatus 100 can suppress an increase in the memory capacity required for orthogonal transform/inverse orthogonal transform.
[0132]
FIG. 12 is a block diagram showing a main configuration example of the orthogonal transform unit 113 in FIG. As shown in FIG. 12, the orthogonal transform unit 113 has a switch 151, a primary transform unit 152, and a secondary transform unit 153.
[0133]
The switch 151 receives the conversion skip flag ts_flag[compID] corresponding to the prediction residual D and the component identifier compID, and the value of the conversion skip flag ts_flag[compID] is NO_TS(=0) (when the conversion skip is not applied). ), and supplies the prediction residual D to the primary conversion part 152. Further, when the value of the conversion skip flag ts_flag[compID] is 2D_TS(=1) (indicating that the two-dimensional conversion skip is applied), the primary conversion unit 152 and the secondary conversion unit 153 are skipped, and the prediction residual D Is output to the outside of the orthogonal transform unit 113 as a transform coefficient Coeff (supplied to the quantization unit 114).
[0134]
The primary conversion unit 152 performs processing related to primary conversion, which is a predetermined conversion processing such as orthogonal conversion. For example, the primary conversion unit 152 uses the component identifier compID, the adaptive primary conversion flag apt_flag[compID] of the component identifier compID, the primary conversion identifier pt_idx[compID] of the component identifier compID, the prediction mode information PInfo, the size of the conversion block (width pair Numerical value log2TBWSize, logarithmic logarithmic value log2TBHSize), and prediction residual D are input. The horizontal width TBWSize of the conversion block is also called TBWidth, and its logarithmic value is also called log2TBWidth. Similarly, the vertical width TBHSize of the conversion block is also called TBHeight, and its logarithmic value is also called log2TBHeight.
[0135]
The primary conversion unit 152 refers to the prediction mode information PInfo, the component identifier compID, the adaptive primary conversion flag apt_flag[compID] of the component identifier compID, and the primary conversion identifier pt_idx[compID] of the component identifier compID to obtain the component identifier compID. The corresponding conversion type TrTypeH of the primary horizontal conversion (and the primary horizontal conversion type identifier TrTypeIdxH indicating the conversion type) and the conversion type TrTypeV of the primary vertical conversion (and the primary vertical conversion type identifier TrTypeIdxV indicating the conversion type) are selected.
[0136]
Further, the primary conversion unit 152, for the prediction residual D, the primary horizontal conversion type identifier TrTypeIdxH (or the primary horizontal conversion type TrTypeH) and the primary horizontal conversion determined by the horizontal width log2TBWSize of the conversion block, and the primary vertical conversion type identifier. The primary vertical conversion determined by TrTypeIdxV (or the primary vertical conversion type TrTypeV) and the vertical width log2TBHSize of the conversion block is performed to derive the conversion coefficient Coeff_P after the primary conversion. The primary horizontal transform is a horizontal one-dimensional orthogonal transform, and the primary vertical transform is a vertical one-dimensional orthogonal transform.
[0137]
The primary conversion unit 152 supplies the derived conversion coefficient Coeff_P to the secondary conversion unit 153.
[0138]
The secondary conversion unit 153 performs processing related to secondary conversion, which is a predetermined conversion processing such as orthogonal conversion. For example, the secondary conversion unit 153 receives the secondary conversion identifier st_idx, the scan identifier scanIdx indicating the scan method of the conversion coefficient, and the conversion coefficient Coeff_P. The secondary conversion unit 153 performs secondary conversion on the conversion coefficient Coeff_P based on the secondary conversion identifier st_idx and the scan identifier scanIdx, and derives the conversion coefficient Coeff_S after the secondary conversion.
[0139]
More specifically, when the secondary conversion identifier st_idx indicates that the secondary conversion is applied (st_idx>0), the secondary conversion unit 153, for the conversion coefficient Coeff_P, of the secondary conversion corresponding to the secondary conversion identifier st_idx. The processing is executed to derive the conversion coefficient Coeff_S after the secondary conversion.
[0140]
The secondary transform unit 153 outputs the secondary transform coefficient Coeff_S to the outside of the orthogonal transform unit 113 as a transform coefficient Coeff (supplies to the quantization unit 114).
[0141]
When the secondary conversion identifier st_idx indicates that the secondary conversion is not applied (st_idx==0), the secondary conversion unit 153 skips the secondary conversion and converts the conversion coefficient Coeff_P after the primary conversion to the conversion coefficient Coeff (secondary conversion. The latter transform coefficient Coeff_S) is output to the outside of the orthogonal transform unit 113 (supplied to the quantization unit 114).
[0142]
In the orthogonal transform unit 113 having the above-described configuration, the primary transform unit 152 performs the process to which the above-described present technology is applied as the derivation unit and the orthogonal transform unit. That is, the deriving unit derives the second transform matrix using the first transform matrix, and the orthogonal transform unit performs the primary transform on the prediction residual using the second transform matrix derived by the deriving unit. .. Therefore, it is possible to suppress an increase in the memory capacity required for the primary conversion.
[0143]
As described above, the primary conversion unit 152 performs the primary horizontal conversion and the primary vertical conversion as the primary conversion. That is, the derivation unit derives the second conversion matrix for horizontal one-dimensional orthogonal transformation and the second conversion matrix for vertical one-dimensional orthogonal transformation, and the orthogonal transformation unit performs the primary transformation as Using the second transformation matrix for horizontal one-dimensional orthogonal transformation derived by the deriving unit, performing horizontal one-dimensional orthogonal transformation, and further for vertical one-dimensional orthogonal transformation derived by the deriving unit. Using the second transformation matrix of, the vertical one-dimensional orthogonal transformation is performed. Therefore, it is possible to suppress an increase in the memory capacity required for the primary conversion in which the horizontal one-dimensional orthogonal transformation and the vertical one-dimensional orthogonal transformation are performed.
[0144]
Next, the flow of each process executed by the image coding apparatus 100 as described above will be described. First, an example of the flow of image coding processing will be described with reference to the flowchart in FIG.
[0145]
When the image encoding process is started, in step S101, the rearrangement buffer 111 is controlled by the control unit 101 to rearrange the order of the frames of the input moving image data from the display order to the encoding order.
[0146]
In step S102, the control unit 101 sets a processing unit (block division) for the input image held by the rearrangement buffer 111.
[0147]
In step S103, the control unit 101 determines (sets) the coding parameter for the input image held by the rearrangement buffer 111.
[0148]
In step S104, the prediction unit 122 performs a prediction process to generate a predicted image and the like in the optimum prediction mode. For example, in this prediction process, the prediction unit 122 performs intra prediction to generate a predicted image or the like in the optimal intra prediction mode, performs inter prediction to generate a predicted image or the like in the optimal inter prediction mode, and The optimum prediction mode is selected from among them based on the cost function value and the like.
[0149]
In step S105, the calculation unit 112 calculates the difference between the input image and the predicted image of the optimum mode selected by the prediction process of step S104. That is, the calculation unit 112 generates the prediction residual D between the input image and the predicted image. The data amount of the prediction residual D thus obtained is smaller than that of the original image data. Therefore, the data amount can be compressed as compared with the case where the image is encoded as it is.
[0150]
In step S106, the orthogonal transform unit 113 performs an orthogonal transform process on the prediction residual D generated by the process of step S105 to derive a transform coefficient Coeff.
[0151]
In step S107, the quantization unit 114 quantizes the transform coefficient Coeff obtained by the process of step S106 by using the quantization parameter calculated by the control unit 101, and derives the quantized transform coefficient level level. ..
[0152]
In step S108, the dequantization unit 117 dequantizes the quantized transform coefficient level level generated by the process of step S107 with a characteristic corresponding to the quantization characteristic of step S107, and derives a transform coefficient Coeff_IQ. ..
[0153]
In step S109, the inverse orthogonal transform unit 118 performs inverse orthogonal transform on the transform coefficient Coeff_IQ obtained by the process of step S108 by a method corresponding to the orthogonal transform process of step S106 to derive the prediction residual D′. Since this inverse orthogonal transform process is the same as the inverse orthogonal transform process (described later) performed on the decoding side, the description (described below) applied to the decoding side is applied to the inverse orthogonal transform process of step S109. can do.
[0154]
In step S110, the calculation unit 119 adds the prediction image obtained by the prediction process of step S104 to the prediction residual D′ derived by the process of step S109 to obtain a locally decoded decoded image. To generate.
[0155]
In step S111, the in-loop filter unit 120 performs the in-loop filter process on the locally decoded decoded image derived by the process of step S110.
[0156]
In step S112, the frame memory 121 stores the locally decoded decoded image derived by the process of step S110 and the locally decoded decoded image filtered in step S112.
[0157]
In step S113, the encoding unit 115 encodes the quantized transform coefficient level level obtained in the process of step S107. For example, the encoding unit 115 encodes the quantized transform coefficient level level, which is information about an image, by arithmetic encoding or the like to generate encoded data. In addition, at this time, the encoding unit 115 encodes various encoding parameters (header information Hinfo, prediction mode information Pinfo, conversion information Tinfo). Further, the encoding unit 115 derives residual information RInfo from the quantized transform coefficient level level and encodes the residual information RInfo.
[0158]
In step S114, the accumulation buffer 116 accumulates the encoded data thus obtained and outputs it as a bit stream to the outside of the image encoding device 100, for example. This bit stream is transmitted to the decoding side, for example, via a transmission line or a recording medium. The rate control unit 123 also performs rate control as needed.
[0159]
When the process of step S114 ends, the image coding process ends.
[0160]
In the image encoding process having the above flow, the process to which the above-described present technology is applied is performed as the process of step S106. Further, as the processing of step S109, processing to which the above-described present technology is applied is performed. Furthermore, as the processing of step S113, the processing to which the above-described present technology is applied is performed. Therefore, by executing this image coding process, it is possible to suppress an increase in the memory capacity required for orthogonal transform/inverse orthogonal transform.
[0161]
Next, an example of the flow of the orthogonal transform processing executed in step S106 of FIG. 13 will be described with reference to the flowchart of FIG.
[0162]
When the orthogonal transform process is started, in step S131, the switch 151 sets the transform skip flag ts_flag to 2D_TS (in the case of indicating a two-dimensional transform skip) (for example, 1 (true)) or the transform quantization bypass flag transquant_bypass_flag to 1. (True), it is determined whether or not. When it is determined that the transform skip flag ts_flag is 2D_TS (for example, 1 (true)) or the transform quantization bypass flag is 1 (true), the orthogonal transform process ends, and the process returns to FIG. 13. In this case, the orthogonal transform process (primary transform or secondary transform) is omitted, and the input prediction residual D is used as the transform coefficient Coeff.
[0163]
Further, in step S131 of FIG. 14, it is determined that the conversion skip flag ts_flag is not 2D_TS (not two-dimensional conversion skip) (for example, 0 (false)) and the conversion quantization bypass flag transquant_bypass_flag is 0 (false). If so, the process proceeds to step S132. In this case, the primary conversion process and the secondary conversion process are performed.
[0164]
In step S132, the primary conversion unit 152 performs the primary conversion process on the input prediction residual D based on the adaptive primary conversion information specified by the component identifier compID, and derives the conversion coefficient Coeff_P after the primary conversion. To do.
[0165]
In step S133, the secondary conversion unit 153 performs a secondary conversion process on the conversion coefficient Coeff_P to derive a conversion coefficient Coeff_S (conversion coefficient Coeff) after the secondary conversion.
[0166]
When the process of step S133 ends, the orthogonal transform process ends.
[0167]
In the orthogonal transformation process as described above, the process to which the above-described present technology is applied is performed as the process of step S132. Therefore, by executing this orthogonal transformation process, it is possible to suppress an increase in the memory capacity required for the primary transformation.
[0168]
FIG. 15 is a block diagram showing an example of a configuration of an image decoding device which is one mode of an image processing device to which the present technology is applied. The image decoding device 200 shown in FIG. 15 is a device for decoding encoded data in which the prediction residual between an image and its predicted image is encoded, such as AVC and HEVC. For example, the image decoding device 200 implements the technology described in Non-Patent Document 1, Non-Patent Document 3, or Non-Patent Document 4, and a moving image is generated by a method that conforms to the standard described in any of those documents. The encoded data in which the image data of the image is encoded is decoded. For example, the image decoding device 200 decodes the encoded data (bit stream) generated by the image encoding device 100 described above.
[0169]
Note that FIG. 15 shows main components such as a processing unit and a data flow, and the components shown in FIG. 15 are not necessarily all. That is, in the image decoding apparatus 200, a processing unit not shown as a block in FIG. 15 may exist, or a process or data flow not shown as an arrow or the like in FIG. 15 may exist. This also applies to other drawings that explain the processing unit and the like in the image decoding apparatus 200.
[0170]
In FIG. 15, the image decoding device 200 includes a storage buffer 211, a decoding unit 212, an inverse quantization unit 213, an inverse orthogonal transformation unit 214, a calculation unit 215, an in-loop filter unit 216, a rearrangement buffer 217, a frame memory 218, and The prediction unit 219 is provided. The prediction unit 219 includes an intra prediction unit and an inter prediction unit (not shown). The image decoding device 200 is a device for generating moving image data by decoding encoded data (bit stream).
[0171]
The
accumulation buffer 211 acquires and holds (stores) the bitstream input to the image decoding device 200. The accumulation buffer 211 supplies the accumulated bitstream to the decoding unit 212 at a predetermined timing or when a predetermined condition is satisfied.
[0172]
The
decoding unit 212 performs processing related to image decoding. For example, the decoding unit 212 receives the bitstream supplied from the accumulation buffer 211 as input, performs variable-length decoding on the syntax value of each syntax element from the bit string according to the definition of the syntax table, and derives the parameter. To do.
[0173]
The parameters derived from the syntax elements and the syntax values of the syntax elements include information such as header information Hinfo, prediction mode information Pinfo, conversion information Tinfo, residual information Rinfo, and filter information Finfo. That is, the decoding unit 212 parses (analyzes and acquires) these pieces of information from the bitstream. These pieces of information will be described below.
[0174]
The
header information Hinfo includes header information such as VPS (Video Parameter Set)/SPS (Sequence Parameter Set)/PPS (Picture Parameter Set)/SH (Slice Header). The header information Hinfo includes, for example, image size (width PicWidth, height PicHeight), bit depth (luminance bitDepthY, color difference bitDepthC), color difference array type ChromaArrayType, CU size maximum value MaxCUSize/minimum value MinCUSize, and quadtree partitioning ( Quad-tree split) maximum depth MaxQTDepth/minimum depth MinQTDepth, binary tree split (Binary-tree split) maximum depth MaxBTDepth/minimum depth MinBTDepth, maximum value of transform skip block MaxTSSize (also called max transform skip block size) ), information that defines an on/off flag (also referred to as a valid flag) of each encoding tool, and the like.
[0175]
For example, as the on/off flag of the encoding tool included in the header information Hinfo, there are on/off flags related to the conversion and quantization processing described below. The on/off flag of the coding tool can also be interpreted as a flag indicating whether or not the syntax related to the coding tool is present in the coded data. Further, when the value of the on/off flag is 1 (true), it indicates that the coding tool is usable, and when the value of the on/off flag is 0 (false), the coding tool is unusable. Show. The interpretation of the flag value may be reversed.
[0176]
Inter-component prediction valid flag (ccp_enabled_flag): This is flag information indicating whether inter-component prediction (also called CCP (Cross-Component Prediction) or CC prediction) can be used. For example, if the flag information is “1” (true), it indicates that the flag can be used, and if the flag information is “0” (false), it indicates that the flag cannot be used.
[0177]
This CCP is also called inter-component linear prediction (CCLM or CCLMP).
[0178]
The
prediction mode information Pinfo includes, for example, size information PBSize (prediction block size) of the processing target PB (prediction block), intra prediction mode information IPinfo, motion prediction information MVinfo, and the like.
[0179]
The intra prediction mode information IPinfo includes, for example, prev_intra_luma_pred_flag, mpm_idx, rem_intra_pred_mode in JCTVC-W1005, 7.3.8.5 Coding Unit syntax, and luminance intra prediction mode IntraPredModeY derived from the syntax.
[0180]
Also, the intra prediction mode information IPinfo includes, for example, inter-component prediction flag (ccp_flag (cclmp_flag)), multi-class linear prediction mode flag (mclm_flag), color difference sample position type identifier (chroma_sample_loc_type_idx), color difference MPM identifier (chroma_mpm_idx), and , Luminance intra prediction mode (IntraPredModeC) and the like derived from these syntaxes are included.
[0181]
The inter-component prediction flag (ccp_flag (cclmp_flag)) is flag information indicating whether to apply inter-component linear prediction. For example, when ccp_flag==1, it indicates that inter-component prediction is applied, and when ccp_flag==0, it indicates that inter-component prediction is not applied.
[0182]
The multi-class linear prediction mode flag (mclm_flag) is information about the mode of linear prediction (linear prediction mode information). More specifically, the multi-class linear prediction mode flag (mclm_flag) is flag information indicating whether to set the multi-class linear prediction mode. For example, "0" indicates that it is a one-class mode (single class mode) (for example, CCLMP), and "1" indicates that it is a two-class mode (multi-class mode) (for example, MCLMP). ..
[0183]
The color difference sample position type identifier (chroma_sample_loc_type_idx) is an identifier for identifying the type of pixel position of the color difference component (also referred to as color difference sample position type). For example, when the color difference array type (ChromaArrayType), which is information about the color format, indicates the 420 format, the color difference sample position type identifier is assigned as shown in the following expression (7).
[0184]
[Number 5]
[0185]
The color difference sample position type identifier (chroma_sample_loc_type_idx) is transmitted (stored in) as information (chroma_sample_loc_info()) regarding the pixel position of the color difference component.
[0186]
The color difference MPM identifier (chroma_mpm_idx) is an identifier indicating which prediction mode candidate in the color difference intra prediction mode candidate list (intraPredModeCandListC) is designated as the color difference intra prediction mode.
[0187]
The motion prediction information MVinfo includes, for example, information such as merge_idx, merge_flag, inter_pred_idc, ref_idx_LX, mvp_lX_flag, X={0,1}, mvd (see, for example, JCTVC-W1005, 7.3.8.6 Prediction Unit Syntax). ..
[0188]
Of course, the information included in the prediction mode information Pinfo is arbitrary, and information other than these information may be included.
[0189]
The
conversion information Tinfo includes, for example, the following information. Of course, the information included in the conversion information Tinfo is arbitrary, and information other than these information may be included.
[0190]
The width size TBWSize and the height width TBHSize of the conversion block to be processed (or each TBWSize with base 2 and log2TBWSize, log2TBHSize of logarithmic value of TBHSize may be used).
Conversion skip flag (ts_flag): This flag indicates whether (reverse) primary conversion and (reverse) secondary conversion are skipped.
Scan identifier (scanIdx)
quantization parameter (qp)
quantization matrix (scaling_matrix (eg JCTVC-W1005, 7.3.4 Scaling list data syntax))
[0191]
The
residual information Rinfo (see, for example, 7.3.8.11 Residual Coding syntax of JCTVC-W1005) includes the following syntax, for example.
[0192]
cbf (coded_block_flag): Residual data presence flag
last_sig_coeff_x_pos: Last non-zero coefficient X coordinate
last_sig_coeff_y_pos: Last non-zero coefficient Y coordinate
coded_sub_block_flag: Sub-block non-zero coefficient presence flag
sig_coeff_flag: Non-zero coefficient presence flag
gr1_flag: Non-zero coefficient level Flag indicating whether it is greater than 1 (also called GR1 flag)
gr2_flag: Flag indicating whether the level of non-zero coefficient is greater than 2 (also called GR2 flag)
sign_flag: Sign indicating the sign of non-zero coefficient (also called sign code)
coeff_abs_level_remaining: Non-zero coefficient residual level (also called non-zero coefficient residual level),
etc.
[0193]
Of course, the information included in the residual information Rinfo is arbitrary, and information other than these information may be included.
[0194]
The
filter information Finfo includes, for example, control information regarding each of the following filter processes.
[0195]
Deblocking filter (DBF) control information
Pixel adaptive offset (SAO) control information
Adaptive loop filter (ALF) control information
Other linear/nonlinear filter control information
[0196]
More specifically, for example, a picture to which each filter is applied, information designating an area within the picture, filter On/Off control information for each CU, filter On/Off control information regarding a slice or tile boundary, and the like are included. included. Of course, the information included in the filter information Finfo is arbitrary, and information other than these information may be included.
[0197]
Returning to the explanation of the decoding unit 212, the decoding unit 212 derives the quantized transform coefficient level level at each coefficient position in each transform block by referring to the residual information Rinfo. The decoding unit 212 supplies the quantized transform coefficient level level to the inverse quantization unit 213.
[0198]
The decoding unit 212 also supplies the parsed header information Hinfo, prediction mode information Pinfo, quantized transform coefficient level level, transform information Tinfo, and filter information Finfo to each block. Specifically, it is as follows.
[0199]
The header information Hinfo is supplied to the dequantization unit 213, the inverse orthogonal transform unit 214, the prediction unit 219, and the in-loop filter unit 216.
The prediction mode information Pinfo is supplied to the inverse quantization unit 213 and the prediction unit 219.
The transform information Tinfo is supplied to the inverse quantization unit 213 and the inverse orthogonal transform unit 214.
The filter information Finfo is supplied to the in-loop filter unit 216.
[0200]
Of course, the above-mentioned example is an example, and the present invention is not limited to this example. For example, each coding parameter may be supplied to an arbitrary processing unit. Further, other information may be supplied to any processing unit.
claims
[Claim 1]
A derivation unit that derives the transformation matrix using a partial matrix that forms a part of the transformation matrix, and
an orthogonal transformation unit that orthogonally transforms the prediction residual of the image using the transformation matrix derived by the derivation unit. ,
An encoding unit that encodes coefficient data obtained by orthogonally transforming the prediction residual by the orthogonal transform unit and generates a bitstream
.
[Claim 2]
The image processing apparatus according to claim 1, wherein the derivation unit derives the conversion matrix using the partial matrix stored in a lookup table .
[Claim 3]
The image according to claim 1, wherein the orthogonal transform unit performs a primary transform on the prediction residual using the transform matrix derived by the deriving unit, and further performs a secondary transform on a result of the primary transform. Processing equipment.
[Claim 4]
The derivation unit derives the transformation matrix for horizontal one-dimensional orthogonal transformation and the transformation matrix for vertical one-dimensional orthogonal transformation, and the
orthogonal transformation unit uses the
derivation unit as the primary transformation. The horizontal one-dimensional orthogonal transformation is performed by using the transformation matrix for the horizontal one-dimensional orthogonal transformation, which is
derived by the above, and the vertical one-dimensional orthogonal transformation is performed by the deriving unit.
The image processing apparatus according to claim 3 , wherein the vertical one-dimensional orthogonal transformation is performed using a transformation matrix .
[Claim 5]
The deriving unit, the partial matrix, by deriving the remaining submatrix of the transformation matrix flips about the axis of the predetermined direction passing through the center of the transformation matrix, deriving the transformation matrix
claim 1. The image processing device according to 1.
[Claim 6]
The image processing apparatus according to claim 5, wherein the derivation unit further inverts the sign of the element of the flipped partial matrix to derive the conversion matrix .
[Claim 7]
The submatrix is a left-half submatrix or a right-half submatrix of the transformation matrix, the
derivation unit flips the submatrix in the row direction of the transformation matrix, and further, the flipped submatrix. inverts the sign of the odd-numbered row vector of the by deriving a transformation matrix right half of the matrix or the left half of the submatrix arranged to derive the transformation matrix
according to claim 1 Image processing device.
[Claim 8]
The submatrix is an upper half submatrix or a lower half submatrix of the transformation matrix, the
derivation unit flips the submatrix in a column direction of the transformation matrix, and further, the flipped submatrix. inverts the sign of the odd-numbered column vector of the conversion by deriving the partial matrix or upper half of the sub-matrix of the lower half of the matrix, is configured to derive the transformation matrix
according to claim 1 Image processing device.
[Claim 9]
The submatrix is an upper half submatrix or a lower half submatrix of the transformation matrix, and the
deriving unit flips the submatrix in a rotation direction around the center of the transformation matrix, and further The submatrix of the conversion matrix lower half or the upper half of the transformation matrix is inverted by reversing the signs of the elements of which the row number and the column number are both even and the elements of which the row number and the column number are both odd.
The image processing device according to claim 1 , wherein the conversion matrix is derived by deriving a half submatrix .
[Claim 10]
The submatrix is a submatrix in the upper right triangular portion or a submatrix in the lower left triangular portion of the transformation matrix, and the
deriving unit transposes the submatrix to obtain a submatrix in the lower left triangular portion or an upper right portion of the transformation matrix.
The image processing apparatus according to claim 1 , wherein the conversion matrix is derived by deriving a partial matrix of a triangular portion .
[Claim 11]
The submatrix is a submatrix in the upper left triangular portion or a submatrix in the lower right triangular portion of the conversion matrix, and the
derivation unit sets the submatrix as a diagonal line connecting the upper right end and the lower left lower end of the conversion matrix. Flip diagonally, and further invert the sign of the element of which both the row number and the column number of the flipped sub-matrix are even, and the element of which the row number and the column number are both odd,
The image processing apparatus according to claim 1 , wherein the conversion matrix is derived by deriving a partial matrix of a lower right triangular portion or a partial matrix of an upper left triangular portion of the conversion matrix .
[Claim 12]
The submatrix is a submatrix of the upper left quarter of the conversion matrix, the
derivation unit
flips the submatrix in the row direction of the conversion matrix, and further, the odd number of the flipped submatrix. The sign of the th row vector is inverted to derive a submatrix in the upper right quarter of the
transformation matrix, the submatrix is flipped in the column direction of the transformation matrix, and the flipped submatrix Of the odd-numbered column vector is inverted to derive a submatrix of the lower left quarter of the
transformation matrix, the submatrix is flipped in a rotation direction about the center of the transformation matrix, and , The element of which the row number and the column number are both even and the element of which the row number and the column number are both odd in the flipped submatrix are inverted, and the lower right quadrant of the conversion matrix The image processing apparatus according to claim 1 ,
wherein the conversion matrix is derived by deriving a partial matrix of
1.
[Claim 13]
The deriving unit derives a first transform matrix using the partial matrix, further derives a second transform matrix using the derived first transform matrix, and the
orthogonal transform unit uses the deriving unit.
The image processing apparatus according to claim 1 , wherein the prediction residual is orthogonally transformed using the second transformation matrix derived by the unit.
[Claim 14]
The orthogonal transformation unit performs the orthogonal transformation of the prediction residual using a TU (Transform Unit) of a Quad-Tree Block Structure or a QTBT (Quad Tree Plus Binary Tree) Block Structure as a processing unit
. Image processing device.
[Claim 15]
The image according to claim 1 , wherein the encoding unit encodes the coefficient data using a CU (Coding Unit) of a Quad-Tree Block Structure or a QTBT (Quad Tree Plus Binary Tree) Block Structure as a processing unit. Processing equipment.
[Claim 16]
The transformation matrix is derived using a submatrix that forms a part of the transformation matrix
, the prediction residual of the image is orthogonally transformed using the derived transformation matrix, and the
prediction residual is obtained by orthogonal transformation. An
image processing method for encoding the obtained coefficient data and generating a bit stream .
[Claim 17]
A decoding unit that decodes a bitstream to obtain coefficient data in which the prediction residual of an image is orthogonally transformed,
a deriving unit that derives the transformation matrix using a partial matrix that forms a part of the transformation matrix, and the
deriving unit An
image processing apparatus comprising: an inverse orthogonal transform unit that performs an inverse orthogonal transform on the coefficient data obtained by the decoding unit, using the transform matrix derived by .
[Claim 18]
The inverse orthogonal transform unit performs an inverse secondary transform on the coefficient data obtained by the decoding unit, and further, using the transform matrix derived by the deriving unit, inverses the result of the inverse secondary transform.
The image processing apparatus according to claim 17, which performs primary conversion .
[Claim 19]
The deriving unit derives the transform matrix for horizontal one-dimensional inverse orthogonal transform and the transform matrix for vertical one-dimensional orthogonal transform, and the
inverse orthogonal transform unit performs the inverse primary transform as the inverse primary transform. ,
The inverse one-dimensional orthogonal transformation in the horizontal direction is performed using the transformation matrix for the inverse one-dimensional orthogonal transformation in the horizontal direction
derived by the deriving unit , and the inverse in the vertical direction derived by the deriving unit.
The image processing apparatus according to claim 18 , wherein the inverse one-dimensional orthogonal transformation in the vertical direction is performed using the transformation matrix for one-dimensional orthogonal transformation .
[Claim 20]
The deriving unit, the partial matrix, by deriving the remaining submatrix of the transformation matrix flips about the axis of the predetermined direction passing through the center of the transformation matrix, deriving the transformation matrix
claim 17. The image processing device according to item 17.
[Claim 21]
The deriving unit derives a first transform matrix using the sub-matrix, further derives a second transform matrix using the derived first transform matrix, and the
inverse orthogonal transform unit
The image processing device according to claim 17 , wherein the coefficient data is subjected to inverse orthogonal transform using the second transform matrix derived by the deriving unit.
[Claim 22]
The image processing device according to claim 17 , wherein the decoding unit decodes the bitstream using a CU (Coding Unit) of a Quad-Tree Block Structure or a QTBT (Quad Tree Plus Binary Tree) Block Structure as a processing unit. ..
[Claim 23]
The inverse orthogonal transform unit, Quad-Tree Block Structure or, as QTBT (Quad Tree Plus Binary Tree) Block processing units TU (Transform Unit) of Structure,, performs inverse orthogonal transform of the coefficient data
according to claim 17 Image processing device.
[Claim 24]
Decoding the bitstream, obtaining coefficient data in which the prediction residual of the image is orthogonally transformed
,
deriving the transformation matrix using a partial matrix forming a part of the transformation matrix, and using the derived transformation matrix And an
image processing method for performing inverse orthogonal transform on the obtained coefficient data .
| # | Name | Date |
|---|---|---|
| 1 | 202017020563-TRANSLATIOIN OF PRIOIRTY DOCUMENTS ETC. [15-05-2020(online)].pdf | 2020-05-15 |
| 2 | 202017020563-STATEMENT OF UNDERTAKING (FORM 3) [15-05-2020(online)].pdf | 2020-05-15 |
| 3 | 202017020563-PRIORITY DOCUMENTS [15-05-2020(online)].pdf | 2020-05-15 |
| 4 | 202017020563-POWER OF AUTHORITY [15-05-2020(online)].pdf | 2020-05-15 |
| 5 | 202017020563-FORM 1 [15-05-2020(online)].pdf | 2020-05-15 |
| 6 | 202017020563-DRAWINGS [15-05-2020(online)].pdf | 2020-05-15 |
| 7 | 202017020563-DECLARATION OF INVENTORSHIP (FORM 5) [15-05-2020(online)].pdf | 2020-05-15 |
| 8 | 202017020563-COMPLETE SPECIFICATION [15-05-2020(online)].pdf | 2020-05-15 |
| 9 | 202017020563-Proof of Right [01-09-2020(online)].pdf | 2020-09-01 |
| 10 | 202017020563-FORM 18 [11-10-2021(online)].pdf | 2021-10-11 |
| 11 | 202017020563.pdf | 2021-10-19 |
| 12 | 202017020563-FER.pdf | 2022-03-09 |
| 13 | 202017020563-FER_SER_REPLY [09-09-2022(online)].pdf | 2022-09-09 |
| 14 | 202017020563-DRAWING [09-09-2022(online)].pdf | 2022-09-09 |
| 15 | 202017020563-CORRESPONDENCE [09-09-2022(online)].pdf | 2022-09-09 |
| 16 | 202017020563-COMPLETE SPECIFICATION [09-09-2022(online)].pdf | 2022-09-09 |
| 17 | 202017020563-CLAIMS [09-09-2022(online)].pdf | 2022-09-09 |
| 18 | 202017020563-ABSTRACT [09-09-2022(online)].pdf | 2022-09-09 |
| 19 | 202017020563-PatentCertificate10-11-2025.pdf | 2025-11-10 |
| 20 | 202017020563-IntimationOfGrant10-11-2025.pdf | 2025-11-10 |
| 1 | SearchStrategyE_09-03-2022.pdf |