Abstract: A video decoding method according to the present disclosure may comprise the steps of: determining whether an inverse transform is skipped in a current block; decoding a residual coefficient of the current block; and selectively applying the inverse transform to the residual coefficient on the basis of the determination. When decoding the residual coefficient, either a first syntax indicating whether the residual coefficient is greater than 0 or a second syntax indicating the absolute value of the residual coefficient may be selectively decoded.
Title of Invention: Video signal processing method and apparatus
technical field
[One]
The present disclosure relates to a video signal processing method and apparatus.
background
[2]
Recently, the demand for high-resolution and high-quality images such as HD (High Definition) images and UHD (Ultra High Definition) images is increasing in various application fields. As the image data becomes higher resolution and higher quality, the amount of data increases relatively compared to the existing image data. The storage cost will increase. High-efficiency image compression techniques can be used to solve these problems that occur as image data becomes high-resolution and high-quality.
[3]
Inter-screen prediction technology that predicts pixel values included in the current picture from pictures before or after the current picture with image compression technology, intra-picture prediction technology that predicts pixel values included in the current picture using pixel information in the current picture, Various techniques exist, such as entropy encoding technology in which a short code is assigned to a value with a high frequency of occurrence and a long code is assigned to a value with a low frequency of occurrence.
[4]
Meanwhile, as the demand for high-resolution images increases, the demand for stereoscopic image content as a new image service is also increasing. A video compression technology for effectively providing high-resolution and ultra-high-resolution stereoscopic image content is being discussed.
DETAILED DESCRIPTION OF THE INVENTION
technical challenge
[5]
An object of the present disclosure is to provide a method and apparatus for efficiently encoding/decoding a residual coefficient in encoding/decoding a video signal.
[6]
An object of the present disclosure is to provide a method and apparatus for additionally applying a second transform to a result of a first transform in encoding/decoding a video signal.
[7]
The technical problems to be achieved in the present disclosure are not limited to the technical problems mentioned above, and other technical problems not mentioned will be clearly understood by those of ordinary skill in the art to which the present disclosure belongs from the description below. will be able
means of solving the problem
[8]
A video signal decoding method according to the present disclosure includes determining whether an inverse transform is skipped in a current block, decoding a residual coefficient of the current block, and selectively applying the inverse transform to the residual coefficient based on whether the determination is made may include the step of applying In this case, when decoding the residual coefficient, any one of a first syntax indicating whether the residual coefficient is greater than 0 and a second syntax indicating an absolute value of the residual coefficient may be alternatively decoded.
[9]
A video signal encoding method according to the present disclosure includes the steps of determining whether to skip a transform in a current block, quantizing a result to which a transform is applied or a result in which a transform is skipped, and a residual coefficient output as a result of the quantization It may include the step of encoding. In this case, when the residual coefficient is encoded, any one of a first syntax indicating whether the residual coefficient is greater than 0 and a second syntax indicating an absolute value of the residual coefficient may be alternatively encoded.
[10]
In the video signal decoding method according to the present disclosure, it may be determined whether to decode the first syntax or the second syntax by comparing the number of decoded bins and a threshold value using context information.
[11]
In the video signal decoding method according to the present disclosure, at least one of the first syntax, at least one gt_N_flag indicating whether an absolute value has a value greater than (2N-1), or a parity flag indicating whether an absolute value is an even number When is decoded, the number of bins decoded using the context information may increase.
[12]
In the video signal decoding method according to the present disclosure, when the first syntax is decoded and the first syntax indicates that the residual coefficient has a non-zero value, the absolute value of the residual coefficient is greater than 1. gt_1_flag indicating whether or not to have it may be further decoded.
[13]
In the video signal decoding method according to the present disclosure, when the gt_1_flag indicates that the absolute value has a value greater than 1, a parity flag indicating whether the absolute value is an even number and a parity flag indicating whether the absolute value is greater than 3 gt_2_flag may be further decoded.
[14]
In the video signal decoding method according to the present disclosure, the threshold value may be determined based on the size of the current block.
[15]
The features briefly summarized above with respect to the present disclosure are merely exemplary aspects of the detailed description of the present disclosure that follows, and do not limit the scope of the present disclosure.
Effects of the Invention
[16]
According to the present disclosure, encoding/decoding efficiency can be improved by differently setting the encoding method of a residual coefficient according to the number of bins to be encoded using context information.
[17]
According to the present disclosure, encoding/decoding efficiency can be improved by additionally applying the second transform to the first transformed result.
[18]
Effects obtainable in the present disclosure are not limited to the above-mentioned effects, and other effects not mentioned may be clearly understood by those of ordinary skill in the art to which the present disclosure belongs from the description below. will be.
Brief description of the drawing
[19]
1 is a block diagram illustrating an image encoding apparatus according to an embodiment of the present disclosure.
[20]
2 is a block diagram illustrating an image decoding apparatus according to an embodiment of the present disclosure.
[21]
3 is a flowchart illustrating an intra prediction method according to an embodiment of the present disclosure.
[22]
4 illustrates types of intra prediction modes.
[23]
5 is a diagram for explaining an example of deriving a prediction sample under a planar mode.
[24]
6 shows an example in which prediction samples are generated under horizontal mode and vertical mode.
[25]
7 and 8 are diagrams illustrating examples to which the second transform is applied.
[26]
9 and 10 illustrate a second transform based on a second transform kernel of an asymmetric form.
[27]
11 illustrates an example in which whether information indicating whether the second transform is applied is determined based on the position of the last non-zero coefficient.
[28]
12 illustrates constraint region candidates for a 4x4 block.
[29]
13 is a diagram illustrating an example in which a second transform kernel having a predefined size is applied.
[30]
14 illustrates a scan method.
[31]
15 is a flowchart illustrating a process of encoding residual coefficients in an encoder.
[32]
16 is a flowchart illustrating a process of encoding size information of residual coefficients.
[33]
17 is a flowchart illustrating a process of decoding a residual coefficient in a decoder.
[34]
18 is a diagram illustrating a decoding process of size information of residual coefficients.
[35]
19 and 20 are diagrams illustrating examples of counting the number of bins using context information.
[36]
21 to 23 show examples of different priorities between syntaxes encoded using context information.
[37]
24 and 25 show a surrounding reconstruction area referenced to determine context information.
[38]
26 exemplifies the number of referable context information when encoding the flag sig_flag.
[39]
27 exemplifies the number of referenceable context information when encoding gt_N_flag or par_flag.
Modes for carrying out the invention
[40]
Since the present disclosure can make various changes and can have various embodiments, specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the present disclosure to specific embodiments, and should be understood to include all modifications, equivalents and substitutes included in the spirit and scope of the present disclosure. In describing each figure, like reference numerals have been used for like elements.
[41]
Terms such as first, second, etc. may be used to describe various elements, but the elements should not be limited by the terms. The above terms are used only for the purpose of distinguishing one component from another. For example, without departing from the scope of the present disclosure, a first component may be referred to as a second component, and similarly, a second component may also be referred to as a first component. and/or includes a combination of a plurality of related listed items or any of a plurality of related listed items.
[42]
When a component is referred to as being “connected” or “connected” to another component, it is understood that the other component may be directly connected or connected to the other component, but other components may exist in between. it should be On the other hand, when it is mentioned that a certain element is "directly connected" or "directly connected" to another element, it should be understood that no other element is present in the middle.
[43]
The terms used in the present application are used only to describe specific embodiments, and are not intended to limit the present disclosure. The singular expression includes the plural expression unless the context clearly dictates otherwise. In the present application, terms such as “comprise” or “have” are intended to designate that a feature, number, step, operation, component, part, or combination thereof described in the specification exists, but one or more other features It should be understood that this does not preclude the existence or addition of numbers, steps, operations, components, parts, or combinations thereof.
[44]
Hereinafter, with reference to the accompanying drawings, preferred embodiments of the present disclosure will be described in more detail. Hereinafter, the same reference numerals are used for the same components in the drawings, and repeated descriptions of the same components are omitted.
[45]
[46]
1 is a block diagram illustrating an image encoding apparatus according to an embodiment of the present disclosure.
[47]
Referring to FIG. 1 , the image encoding apparatus 100 includes a picture division unit 110 , prediction units 120 and 125 , a transform unit 130 , a quantization unit 135 , a rearrangement unit 160 , and an entropy encoding unit ( 165 ), an inverse quantization unit 140 , an inverse transform unit 145 , a filter unit 150 , and a memory 155 .
[48]
Each of the constituent units shown in FIG. 1 is independently illustrated to represent different characteristic functions in the image encoding apparatus, and does not mean that each constituent unit is composed of separate hardware or one software constituent unit. That is, each component is listed as each component for convenience of description, and at least two components of each component are combined to form one component, or one component can be divided into a plurality of components to perform a function, and each Integrated embodiments and separate embodiments of components are also included in the scope of the present disclosure without departing from the essence of the present disclosure.
[49]
In addition, some components are not essential components to perform an essential function in the present disclosure, but may be optional components for merely improving performance. The present disclosure may be implemented by including only essential components to implement the essence of the present disclosure, except for components used for performance improvement, and a structure including only essential components excluding optional components used for performance improvement Also included in the scope of the present disclosure.
[50]
The picture divider 110 may divide the input picture into at least one processing unit. In this case, the processing unit may be a prediction unit (PU), a transform unit (TU), or a coding unit (CU). The picture splitter 110 divides one picture into a combination of a plurality of coding units, prediction units, and transformation units, and combines one coding unit, prediction unit, and transformation unit according to a predetermined criterion (eg, a cost function). can be selected to encode the picture.
[51]
For example, one picture may be divided into a plurality of coding units. In order to split a coding unit in a picture, a recursive tree structure such as a quad tree structure may be used. A coding in which one image or a largest coding unit is used as a root and is divided into other coding units. A unit may be divided having as many child nodes as the number of divided coding units. A coding unit that is no longer split according to certain restrictions becomes a leaf node. That is, when it is assumed that only square splitting is possible for one coding unit, one coding unit may be split into up to four different coding units.
[52]
Hereinafter, in an embodiment of the present disclosure, a coding unit may be used as a unit for performing encoding or may be used as a meaning for a unit for performing decoding.
[53]
A prediction unit may be split in the form of at least one square or rectangle having the same size within one coding unit, and one prediction unit among the split prediction units within one coding unit is a prediction of another. It may be divided to have a shape and/or size different from that of the unit.
[54]
When a prediction unit for performing intra prediction based on a coding unit is generated, if it is not the smallest coding unit, intra prediction may be performed without dividing the prediction unit into a plurality of prediction units NxN.
[55]
The prediction units 120 and 125 may include an inter prediction unit 120 performing inter prediction and an intra prediction unit 125 performing intra prediction. Whether to use inter prediction or to perform intra prediction for a prediction unit may be determined, and specific information (eg, intra prediction mode, motion vector, reference picture, etc.) according to each prediction method may be determined. In this case, a processing unit in which prediction is performed and a processing unit in which a prediction method and specific content are determined may be different. For example, a prediction method and a prediction mode may be determined in a prediction unit, and prediction may be performed in a transformation unit. A residual value (residual block) between the generated prediction block and the original block may be input to the transform unit 130 . Also, prediction mode information, motion vector information, etc. used for prediction may be encoded by the entropy encoder 165 together with a residual value and transmitted to a decoding apparatus. When a specific encoding mode is used, it is also possible to encode the original block as it is without generating a prediction block through the prediction units 120 and 125 and transmit it to the decoder.
[56]
The inter prediction unit 120 may predict a prediction unit based on information on at least one of a picture before or after a picture of the current picture, and in some cases, prediction based on information of a partial region in the current picture that has been encoded Units can also be predicted. The inter prediction unit 120 may include a reference picture interpolator, a motion prediction unit, and a motion compensator.
[57]
The reference picture interpolator may receive reference picture information from the memory 155 and generate pixel information of integer pixels or less in the reference picture. In the case of luminance pixels, a DCT-based 8-tap interpolation filter in which filter coefficients are different to generate pixel information of integer pixels or less in units of 1/4 pixels may be used. In the case of a color difference signal, a DCT-based 4-tap interpolation filter in which filter coefficients are different to generate pixel information of integer pixels or less in units of 1/8 pixels may be used.
[58]
The motion prediction unit may perform motion prediction based on the reference picture interpolated by the reference picture interpolator. As a method for calculating the motion vector, various methods such as Full search-based Block Matching Algorithm (FBMA), Three Step Search (TSS), and New Three-Step Search Algorithm (NTS) may be used. The motion vector may have a motion vector value of 1/2 or 1/4 pixel unit based on the interpolated pixel. The motion prediction unit may predict the current prediction unit by using a different motion prediction method. Various methods, such as a skip method, a merge method, an AMVP (Advanced Motion Vector Prediction) method, an intra block copy method, etc., may be used as the motion prediction method.
[59]
The intra prediction unit 125 may generate a prediction unit based on reference pixel information around the current block, which is pixel information in the current picture. When the neighboring block of the current prediction unit is a block on which inter prediction is performed, and thus the reference pixel is a pixel on which inter prediction is performed, the reference pixel included in the block on which the inter prediction is performed is a reference pixel of the block on which the intra prediction is performed. information can be used instead. That is, when the reference pixel is not available, the unavailable reference pixel information may be replaced with at least one reference pixel among the available reference pixels.
[60]
In intra prediction, the prediction mode may have a directional prediction mode in which reference pixel information is used according to a prediction direction and a non-directional mode in which directional information is not used when prediction is performed. A mode for predicting luminance information and a mode for predicting chrominance information may be different, and intra prediction mode information used for predicting luminance information or predicted luminance signal information may be utilized to predict chrominance information.
[61]
When intra prediction is performed, if the size of the prediction unit and the size of the transformation unit are the same, intra prediction of the prediction unit based on the pixel present on the left side, the pixel present on the upper left side, and the pixel present on the upper side of the prediction unit can be performed. However, when the size of the prediction unit is different from the size of the transformation unit when intra prediction is performed, intra prediction may be performed using a reference pixel based on the transformation unit. Also, intra prediction using NxN splitting may be used only for the smallest coding unit.
[62]
The intra prediction method may generate a prediction block after applying an adaptive intra smoothing (AIS) filter to a reference pixel according to a prediction mode. The type of AIS filter applied to the reference pixel may be different. In order to perform the intra prediction method, the intra prediction mode of the current prediction unit may be predicted from the intra prediction mode of the prediction unit existing around the current prediction unit. When the prediction mode of the current prediction unit is predicted using the mode information predicted from the neighboring prediction unit, if the intra prediction mode of the current prediction unit and the neighboring prediction unit are the same, the current prediction unit and the neighboring prediction unit are used using predetermined flag information It is possible to transmit information indicating that the prediction modes of , and if the prediction modes of the current prediction unit and the neighboring prediction units are different from each other, entropy encoding may be performed to encode prediction mode information of the current block.
[63]
In addition, a residual block including residual information, which is a difference value between a prediction unit and an original block of the prediction unit, in which prediction is performed based on the prediction unit generated by the prediction units 120 and 125 may be generated. The generated residual block may be input to the transform unit 130 .
[64]
The transform unit 130 converts the original block and the residual block including residual information of the prediction unit generated by the prediction units 120 and 125 to DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), KLT and It can be converted using the same conversion method. Whether to apply DCT, DST, or KLT to transform the residual block may be determined based on intra prediction mode information of a prediction unit used to generate the residual block.
[65]
The quantization unit 135 may quantize values transformed in the frequency domain by the transform unit 130 . The quantization coefficient may vary according to blocks or the importance of an image. The value calculated by the quantization unit 135 may be provided to the inverse quantization unit 140 and the rearrangement unit 160 .
[66]
The reordering unit 160 may rearrange the coefficient values on the quantized residual values.
[67]
The reordering unit 160 may change the two-dimensional block form coefficient into a one-dimensional vector form through a coefficient scanning method. For example, the rearranging unit 160 may use a Zig-Zag Scan method to scan from DC coefficients to coefficients in a high-frequency region and change them into a one-dimensional vector form. A vertical scan that scans a two-dimensional block shape coefficient in a column direction and a horizontal scan that scans a two-dimensional block shape coefficient in a row direction may be used instead of the zig-zag scan according to the size of the transform unit and the intra prediction mode. That is, it may be determined whether any of the zig-zag scan, the vertical scan, and the horizontal scan is used according to the size of the transform unit and the intra prediction mode.
[68]
The entropy encoding unit 165 may perform entropy encoding based on the values calculated by the reordering unit 160 . For entropy encoding, various encoding methods such as Exponential Golomb, Context-Adaptive Variable Length Coding (CAVLC), and Context-Adaptive Binary Arithmetic Coding (CABAC) may be used.
[69]
The entropy encoder 165 receives the residual value coefficient information and block type information, prediction mode information, division unit information, prediction unit information and transmission unit information, motion of the coding unit from the reordering unit 160 and the prediction units 120 and 125 . Various information such as vector information, reference frame information, interpolation information of a block, and filtering information may be encoded.
[70]
The entropy encoder 165 may entropy-encode the coefficient values of the coding units input from the reordering unit 160 .
[71]
The inverse quantizer 140 and the inverse transform unit 145 inversely quantize the values quantized by the quantizer 135 and inversely transform the values transformed by the transform unit 130 . The residual values generated by the inverse quantizer 140 and the inverse transform unit 145 are combined with the prediction units predicted through the motion estimation unit, the motion compensator, and the intra prediction unit included in the prediction units 120 and 125 and restored. You can create a Reconstructed Block.
[72]
The filter unit 150 may include at least one of a deblocking filter, an offset correcting unit, and an adaptive loop filter (ALF).
[73]
The deblocking filter may remove block distortion caused by the boundary between blocks in the reconstructed picture. In order to determine whether to perform deblocking, it may be determined whether to apply the deblocking filter to the current block based on pixels included in several columns or rows included in the block. When a deblocking filter is applied to a block, a strong filter or a weak filter can be applied according to the required deblocking filtering strength. In addition, in applying the deblocking filter, horizontal filtering and vertical filtering may be concurrently processed when vertical filtering and horizontal filtering are performed.
[74]
The offset corrector may correct the offset of the deblocked image with respect to the original image in units of pixels. In order to perform offset correction on a specific picture, a method of dividing pixels included in an image into a certain number of regions, determining the region to be offset and applying the offset to the region, or taking edge information of each pixel into consideration can be used to apply
[75]
Adaptive loop filtering (ALF) may be performed based on a value obtained by comparing the filtered reconstructed image and the original image. After dividing the pixels included in the image into a predetermined group, one filter to be applied to the corresponding group is determined, and filtering can be performed differentially for each group. As for information on whether to apply ALF, the luminance signal may be transmitted for each coding unit (CU), and the shape and filter coefficients of the ALF filter to be applied may vary according to each block. Also, the ALF filter of the same type (fixed type) may be applied regardless of the characteristics of the target block.
[76]
The memory 155 may store the reconstructed block or picture calculated through the filter unit 150 , and the stored reconstructed block or picture may be provided to the predictors 120 and 125 when inter prediction is performed.
[77]
[78]
2 is a block diagram illustrating an image decoding apparatus according to an embodiment of the present disclosure.
[79]
Referring to FIG. 2 , the image decoding apparatus 200 includes an entropy decoding unit 210, a reordering unit 215, an inverse quantization unit 220, an inverse transform unit 225, prediction units 230 and 235, and a filter unit ( 240) and a memory 245 may be included.
[80]
When an image bitstream is input by the image encoding apparatus, the input bitstream may be decoded by a procedure opposite to that of the image encoding apparatus.
[81]
The entropy decoding unit 210 may perform entropy decoding in a procedure opposite to that performed by the entropy encoding unit of the image encoding apparatus. For example, various methods such as Exponential Golomb, Context-Adaptive Variable Length Coding (CAVLC), and Context-Adaptive Binary Arithmetic Coding (CABAC) may be applied corresponding to the method performed by the image encoding apparatus.
[82]
The entropy decoder 210 may decode information related to intra prediction and inter prediction performed by the encoding apparatus.
[83]
The reordering unit 215 may perform rearrangement based on a method of rearranging the entropy-decoded bitstream by the entropy decoding unit 210 by the encoder. Coefficients expressed in the form of a one-dimensional vector may be restored and rearranged as coefficients in the form of a two-dimensional block. The reordering unit 215 may receive information related to coefficient scanning performed by the encoder and perform the reordering by performing a reverse scanning method based on the scanning order performed by the corresponding encoder.
[84]
The inverse quantization unit 220 may perform inverse quantization based on the quantization parameter provided by the encoding apparatus and the reordered coefficient values of the blocks.
[85]
The inverse transform unit 225 may perform inverse transforms, ie, inverse DCT, inverse DST, and inverse KLT, on the transforms performed by the transform unit, ie, DCT, DST, and KLT, on the quantization result performed by the image encoding apparatus. Inverse transform may be performed based on a transmission unit determined by the image encoding apparatus. The inverse transform unit 225 of the image decoding apparatus may selectively perform transformation techniques (eg, DCT, DST, KLT) according to a plurality of pieces of information such as a prediction method, a size of a current block, and a prediction direction.
[86]
The prediction units 230 and 235 may generate a prediction block based on the prediction block generation related information provided from the entropy decoding unit 210 and previously decoded block or picture information provided from the memory 245 .
[87]
As described above, when intra prediction is performed in the same manner as in the operation of the image encoding apparatus, when the size of the prediction unit and the size of the transformation unit are the same, the pixel present at the left side of the prediction unit, the pixel present at the upper left side, and the upper side Intra prediction is performed on the prediction unit based on existing pixels, but when the size of the prediction unit and the size of the transformation unit are different when performing intra prediction, intra prediction is performed using the reference pixel based on the transformation unit can do. Also, intra prediction using NxN splitting may be used only for the smallest coding unit.
[88]
The prediction units 230 and 235 may include a prediction unit determiner, an inter prediction unit, and an intra prediction unit. The prediction unit determining unit receives various information such as prediction unit information input from the entropy decoder 210, prediction mode information of the intra prediction method, and motion prediction related information of the inter prediction method, and divides the prediction unit from the current coding unit, and predicts It may be determined whether the unit performs inter prediction or intra prediction. The inter prediction unit 230 uses information required for inter prediction of the current prediction unit provided from the image encoding apparatus based on information included in at least one picture before or after the current picture including the current prediction unit. Inter prediction may be performed on the prediction unit. Alternatively, inter prediction may be performed based on information on a pre-restored partial region in the current picture including the current prediction unit.
[89]
In order to perform inter prediction, a motion prediction method of a prediction unit included in a corresponding coding unit based on a coding unit is selected from among skip mode, merge mode, AMVP mode, and intra block copy mode. You can decide which way to go.
[90]
The intra prediction unit 235 may generate a prediction block based on pixel information in the current picture. When the prediction unit is a prediction unit on which intra prediction is performed, intra prediction may be performed based on intra prediction mode information of the prediction unit provided by the image encoding apparatus. The intra prediction unit 235 may include an Adaptive Intra Smoothing (AIS) filter, a reference pixel interpolator, and a DC filter. The AIS filter is a part that performs filtering on the reference pixel of the current block, and may be applied by determining whether to apply the filter according to the prediction mode of the current prediction unit. AIS filtering may be performed on the reference pixel of the current block by using the prediction mode and AIS filter information of the prediction unit provided by the image encoding apparatus. When the prediction mode of the current block is a mode in which AIS filtering is not performed, the AIS filter may not be applied.
[91]
When the prediction mode of the prediction unit is a prediction unit in which intra prediction is performed based on a pixel value obtained by interpolating the reference pixel, the reference pixel interpolator may interpolate the reference pixel to generate a reference pixel of a pixel unit having an integer value or less. When the prediction mode of the current prediction unit is a prediction mode that generates a prediction block without interpolating the reference pixel, the reference pixel may not be interpolated. The DC filter may generate the prediction block through filtering when the prediction mode of the current block is the DC mode.
[92]
The reconstructed block or picture may be provided to the filter unit 240 . The filter unit 240 may include a deblocking filter, an offset correction unit, and an ALF.
[93]
Information on whether a deblocking filter is applied to a corresponding block or picture and information on whether a strong filter or a weak filter is applied when the deblocking filter is applied may be provided from the image encoding apparatus. The deblocking filter of the image decoding apparatus may receive deblocking filter-related information provided from the image encoding apparatus, and the image decoding apparatus may perform deblocking filtering on the corresponding block.
[94]
The offset correction unit may perform offset correction on the reconstructed image based on the type of offset correction applied to the image during encoding, information on the offset value, and the like.
[95]
ALF may be applied to a coding unit based on information on whether ALF is applied, ALF coefficient information, etc. provided from the encoding apparatus. Such ALF information may be provided by being included in a specific parameter set.
[96]
The memory 245 may store the reconstructed picture or block to be used as a reference picture or reference block, and may also provide the reconstructed picture to an output unit.
[97]
As described above, hereinafter, in the embodiments of the present disclosure, a coding unit is used as a term for a coding unit for convenience of description, but may be a unit for performing decoding as well as coding.
[98]
In addition, the current block denotes an encoding/decoding target block, and depending on the encoding/decoding step, a coding tree block (or coding tree unit), a coding block (or a coding unit), a transform block (or a transform unit), or a prediction block (or prediction unit) and the like. In this specification, a 'unit' may indicate a basic unit for performing a specific encoding/decoding process, and a 'block' may indicate a pixel array of a predetermined size. Unless otherwise specified, 'block' and 'unit' may be used interchangeably. For example, in the embodiments to be described later, it may be understood that the coding block (coding block) and the coding unit (coding unit) are mutually equivalent.
[99]
[100]
An image may be encoded/decoded in units of blocks. A coding block may be recursively partitioned based on a tree structure. As an example, the coding block may be divided by at least one of quad tree division, binary tree division, or ternary tree division.
[101]
In addition, the coding block may be divided into a plurality of prediction blocks or a plurality of transform blocks.
[102]
3 is a flowchart illustrating an intra prediction method according to an embodiment of the present disclosure.
[103]
Referring to FIG. 3 , the index of the reference sample line of the current block may be determined ( S301 ). The index may specify one of a plurality of reference sample line candidates. The plurality of reference sample line candidates may include an adjacent reference sample line adjacent to the current block and at least one non-adjacent reference sample line not adjacent to the current block.
[104]
As an example, an adjacent reference sample line composed of an adjacent row having a y-axis coordinate smaller by 1 than the top row of the current block and an adjacent column having an x-axis coordinate smaller by 1 than the leftmost column of the current block may be used as a reference sample line candidate. .
[105]
A first non-adjacent reference sample line including a non-adjacent row whose y-axis coordinate is 2 less than the top row of the current block and a non-adjacent column whose x-axis coordinate is 2 less than the leftmost column of the current block is selected as a reference sample line candidate. can be used.
[106]
A second non-adjacent reference sample line including a non-adjacent row whose y-axis coordinate is 3 less than the top row of the current block and a non-adjacent column whose x-axis coordinate is 3 less than the leftmost column of the current block is selected as a reference sample line candidate. can be used.
[107]
The index may indicate one of an adjacent reference sample line, a first non-adjacent reference sample line, or a second non-adjacent reference sample line. For example, an index of 0 means that the adjacent reference sample line is selected, an index of 1 means that the first non-adjacent reference sample line is selected, and an index of 2 means that the second non-adjacent reference sample line is selected. means selected.
[108]
An index specifying one of the plurality of reference sample line candidates may be signaled through the bitstream.
[109]
Alternatively, the index may be signaled for the luma component block, and the signaling of the index may be omitted for the chroma component block. If the signaling of the index is omitted, it may be considered that the index is 0. That is, with respect to the chroma component block, intra prediction may be performed using an adjacent reference sample line.
[110]
Reconstructed samples included in the selected reference sample line may be derived as reference samples.
[111]
Next, an intra prediction mode of the current block may be determined (S302).
[112]
4 illustrates types of intra prediction modes. As in the example shown in FIG. 4 , the intra prediction modes include a non-directional prediction mode (DC and Planar) and a directional prediction mode. In FIG. 4 , it is exemplified that 65 directional prediction modes are defined.
[113]
A flag indicating whether the intra prediction mode of the current block is the same as the Most Probable Mode (MPM) may be signaled through the bitstream. For example, the value of the MPM flag being 1 indicates that the same MPM as the intra prediction mode of the current block exists. On the other hand, the value of the MPM flag being 0 indicates that the same MPM as the intra prediction mode of the current block does not exist.
[114]
When the value of the MPM flag is 1, a flag indicating whether the intra prediction mode of the current block is the same as the default intra prediction mode may be signaled. The default intra prediction mode may be at least one of DC, a planar, a vertical prediction mode, and a horizontal prediction mode. As an example, a flag intra_not_planar_flag indicating whether the intra prediction mode of the current block is the planar mode may be signaled. A value of the flag intra_not_planar_flag of 0 indicates that the intra prediction mode of the current block is planar. On the other hand, the value of the flag intra_not_planar_flag of 1 indicates that the intra prediction mode of the current block is not the planner. When the value of the flag intra_not_planar_flag is 1, an index specifying one of the MPM candidates may be signaled. The intra prediction mode of the current block may be set to be the same as the MPM indicated by the MPM index.
[115]
A prediction sample may be derived based on the reference samples belonging to the reference sample line and the intra prediction mode (S303).
[116]
When the intra prediction mode of the current block is the directional prediction mode, a prediction sample may be derived using a reference sample positioned on a line along the angle of the directional prediction mode.
[117]
When the intra prediction mode of the current block is the planar mode, a prediction sample may be derived using a reference sample positioned in a vertical direction and a reference sample positioned in a horizontal direction of the prediction target sample.
[118]
5 is a diagram for explaining an example of deriving a prediction sample under a planar mode.
[119]
In FIG. 5 , T denotes a reference sample adjacent to the upper right corner of the current block, and L denotes a reference sample adjacent to the lower left corner of the current block.
[120]
Under the planar mode, a horizontal direction prediction sample P1 and a vertical direction prediction sample P2 may be derived for a prediction target sample.
[121]
The horizontal direction prediction sample P1 may be generated by linearly interpolating the reference sample H and the upper right reference sample T positioned on the same horizontal line as the prediction target sample.
[122]
The vertical direction prediction sample P2 may be generated by linearly interpolating the reference sample V and the lower left reference sample L positioned on the same vertical line as the prediction target sample.
[123]
Thereafter, a prediction sample may be derived based on a weighted sum operation of the horizontal direction prediction sample P1 and the vertical direction prediction sample P2. Equation 1 shows an example in which the prediction sample P is derived by a weighted sum operation of the horizontal direction prediction sample P1 and the vertical direction prediction sample P2.
[124]
[Formula 1]
[125]
In Equation 1, α denotes a weight applied to the horizontal direction prediction sample P1, and β denotes a weight applied to the vertical direction prediction sample P2.
[126]
The weights α and β may be determined based on the size or shape of the current block. Specifically, weights α and β may be determined in consideration of at least one of a width or a height of the current block. For example, when the width and height of the current block are the same, the weights α and β may be set to the same value. When the weights α and β are the same, the prediction sample may be derived as an average value of the horizontal direction prediction sample P1 and the vertical direction prediction sample P2. On the other hand, when the width and height of the current block are different, the weights α and β may be set differently. For example, when the width of the current block is greater than the height, the weight β can be set to a value greater than the weight α, and when the height of the current block is greater than the width, the weight α can be set to a value greater than the weight β. have. Or, conversely, if the width of the current block is greater than the height, set the weight α to a value greater than the weight β, and if the height of the current block is greater than the width, set the weight β to a value greater than the weight α. can be set.
[127]
As another example, weights α and β may be derived from one of a plurality of weight set candidates. For example, when weight candidate sets (1, 1), (3, 1), and (1, 3) representing a combination of weights α and β are predefined, the weights α and β are one of the weight candidate sets. can be selected in the same way as
[128]
An index indicating one of the plurality of weight set candidates may be signaled through the bitstream. The index may be signaled at the block level. As an example, the index may be signaled in units of coding blocks or transform blocks.
[129]
Alternatively, the index may be signaled at the coding tree unit, slice, picture, or sequence level. Blocks included in the index transmission unit may determine weights α and β with reference to an index signaled at a higher level. That is, for blocks included in the index transmission unit, the weights α and β may be identically set.
[130]
In the example of FIG. 5 , it is shown that the upper right reference sample T is used to derive the horizontal direction prediction sample P1, and the lower left reference sample L is used to derive the vertical direction prediction sample P2.
[131]
A horizontal direction prediction sample P1 may be derived using a reference sample other than the upper right reference sample, or a vertical direction prediction sample P2 may be derived using a reference sample other than the lower left reference sample. As an example, reference sample set candidates for the first reference sample used to derive the horizontal direction prediction sample P1 and the second reference sample used to derive the vertical direction prediction sample P2 are configured, and selected from among the plurality of reference sample set candidates Using one, the horizontal direction prediction sample P1 and the vertical direction prediction sample P2 may be derived.
[132]
An index identifying one of the plurality of reference sample set candidates may be signaled through the bitstream. The index may be signaled in units of blocks, sub-blocks, or samples.
[133]
Alternatively, a reference sample set candidate may be selected based on the location of the prediction target sample.
[134]
Under the directional prediction mode, a prediction sample may be generated using reconstructed pixels around the current block.
[135]
6 shows an example in which prediction samples are generated under horizontal mode and vertical mode.
[136]
As in the example shown in FIG. 6 , in the horizontal mode, a prediction sample may be generated using a reconstructed sample positioned in a horizontal direction of the prediction target sample.
[137]
In the vertical mode, a prediction sample may be generated using a reconstructed sample positioned in a vertical direction of the prediction target sample.
[138]
After intra prediction is performed based on the intra prediction mode, a residual block may be derived by differentiating the prediction block from the original block.
[139]
In this case, a prediction method using one of modes 0 to 66 may be used, or a limited prediction method may be used. In the limited prediction method, only the intra prediction mode in the horizontal direction (No. 18) or the intra prediction mode in the vertical direction (No. 50) can be used. In this case, the intra prediction mode may be specified by 1-bit information. Alternatively, in addition to the vertical and horizontal directions, diagonal prediction modes, for example, the lower left diagonal direction (No. 2) and the upper right diagonal direction (No. 66) may be added as available candidates. In this case, the intra prediction mode may be specified with 2-bit information. Alternatively, two of three diagonal modes, for example, the lower left diagonal direction (No. 2), the upper left diagonal direction (No. 34), and the upper right diagonal direction (No. 66) may be added as available candidates.
[140]
The number of available intra prediction modes may be encoded and transmitted to a decoder. Alternatively, the number of intra prediction modes available in the encoder and the decoder may be fixed. Alternatively, the number of available intra prediction modes may be determined based on the size or shape of the current block.
[141]
[142]
After prediction is performed, a residual block may be obtained by differentiating the original block and the prediction block. When the residual block is obtained, the residual coefficient may be obtained by performing at least one of transform or quantization on the residual block.
[143]
Information indicating whether a transform is applied to the current block may be encoded and signaled. As an example, transform_skip_flag may be encoded and signaled. When transform_skip_flag is 1, it indicates that no transform is applied to the current block. Here, the transform may include not only a first transform, which will be described later, but also a second transform. When transform_skip_flag is 0, it indicates that transform is applied to the current block. When transform_skip_flag is 0, the first transform is necessarily applied to the current block, while the second transform may be selectively applied.
[144]
Transformation may be performed based on at least one of a DCT-based transform kernel and a DST-based transform kernel. Here, the DCT-based transform kernel may include at least one of DCT-2 and DCT-8, and the DST-based transform kernel may include DST-7. An additional transform may be applied to the result of transforming the residual samples. Hereinafter, for convenience of description, a transform performed by a DCT or DST-based transform kernel will be referred to as a first transform, and a transform additionally applied to a result of the first transform will be referred to as a second transform. Also, transform coefficients generated as a result of the first transform will be referred to as first transform coefficients, and transform coefficients generated as a result of the second transform will be referred to as second transform coefficients.
[145]
The second transform may be applied to at least some of the first transform coefficients. For example, the second transform may be applied to 16, 48, or 64 first transform coefficients according to the size of the second transform kernel. The shape of the region including the first transform coefficients to which the second transform is applied may be a square, non-square, or polygonal shape.
[146]
Equation 2 shows an application aspect of the second transform.
[147]
[Formula 2]
[148]
In Equation 2, B_Rx1 represents second transform coefficients composed of R rows and 1 column. T_RXN represents a second transform kernel composed of R rows and N columns. A_Nx1 represents first transform coefficients composed of N rows and 1 column.
[149]
7 and 8 are diagrams illustrating examples to which the second transform is applied.
[150]
7 shows an example in which the second transform kernel is 64x64 in size. The first transform coefficients generated as a result of the first transform in the 8x8 block may be arranged in one dimension. In this case, a one-dimensional array may be generated by scanning the first transform coefficients in a predetermined scan method. The predetermined scan method may include at least one of a diagonal scan, a horizontal scan, a vertical scan, and a raster scan.
[151]
When the 64x1 size input force is generated through the rearrangement, the second transform coefficient may be derived through matrix multiplication between the 64x64 second transform kernel and the 64x1 size input matrix.
[152]
As a result of performing the second transform, 64 second transform coefficients may be generated, and the second transform coefficients in the 8x8 block may be rearranged. After the 8x8 block in which the second transform coefficients are rearranged is quantized, the quantized transform block may be encoded.
[153]
8 shows an example in which the second transform kernel has a size of 48x48. 48 first transform coefficients among the first transform coefficients generated as a result of the first transform in the 8x8 block may be rearranged in one dimension. In this case, the 48 first transform coefficients may be included in a polygonal region except for a sub-block having a size of 4x4 at the lower right of the 8x8 block.
[154]
When the 48x1 sized input matrix is generated by rearranging the 48 first transform coefficients in one dimension, the second transform coefficient may be derived through matrix multiplication between the 48x48 sized second transform kernel and the 48x1 sized input matrix.
[155]
As a result of performing the second transform, 48 second transform coefficients may be generated, and the second transform coefficients in the 8x8 block may be rearranged. For example, 48 second transform coefficients may be rearranged in a polygonal region except for a sub-block having a size of 4x4 at the lower right of the 8x8 block.
[156]
In a region where the second transform coefficients are not disposed, the first transform coefficients may be maintained as they are. After quantization is applied to the block including the second transform coefficients and the first transform coefficients, the quantized transform block may be encoded.
[157]
Alternatively, transform coefficients in a region in which the second transform coefficients are not disposed may be set to zero. That is, values of transform coefficients in a region to which the second transform is not applied are set to 0, and then quantization and encoding may be performed.
[158]
The size of the second transform kernel may be determined based on the size of the current block. For example, when at least one of a width or a height of the current block is 4, the second transform may be applied to 16 first transform coefficients. On the other hand, when the width and height of the current block are 8 or more, the second transform may be applied to 48 or 64 first transform coefficients.
[159]
Alternatively, information indicating the size and type of the second transform kernel may be encoded and signaled. The information may be signaled at the block level. As an example, information specifying at least one of the number of rows and the number of columns of the transform size may be encoded. Alternatively, after allocating a different index to each of the combinations of the number of rows and the number of columns, an index specifying one of the combinations may be encoded. Alternatively, after allocating a different index to each of the plurality of second transform kernel candidates, an index specifying one of the second transform kernel candidates may be encoded. Here, each of the plurality of second transform kernel candidates may have different sizes and/or different coefficients.
[160]
Alternatively, after determining the size of the second transform kernel based on the size of the current block, an index specifying one of a plurality of second transform kernel candidates having the determined size may be encoded.
[161]
In the examples shown in FIGS. 7 and 8 , the second transform kernel having the same number of rows and the same number of columns is used. In order to simplify the second transformation, it is also possible to set the number of rows and the number of columns differently.
[162]
9 and 10 illustrate a second transform based on a second transform kernel of an asymmetric form.
[163]
The number R of rows of the second transform kernel may be set to a value smaller than the number N of columns. For example, the number R of rows may be set to 8, and the number of columns N may be set to 48.
[164]
When the number of rows of the second transform kernel decreases, the number of second transform coefficients output as a result of the second transform also decreases. For example, when a matrix multiplication between a second transform kernel having a size of 8x48 and an input matrix having a size of 48x1 is performed, second transform coefficients having a size of 8x1 are generated.
[165]
Eight second transform coefficients may be rearranged within an 8x8 block. In this case, values of the transform coefficients may be set to 0 in a region to which the second transform coefficients are not allocated within the region to which the second transform is applied (ie, the region including the first transform coefficients to which the second transform is applied). For example, when a polygonal shape region including 48 samples is an application region of the second transform, values of transform coefficients are set to 0 in the residual region except for a region to which 8 second transform coefficients are allocated among the polygonal shape regions. can
[166]
In a region to which the second transform is not applied, the first transform coefficients may be maintained as they are.
[167]
Alternatively, at least some of the first transform coefficients in a region to which the second transform is not applied may be transformed into 0 for encoding. 10 , an example in which at least a portion of a region to which the second transform is not applied is transformed into 0 is exemplified.
[168]
As in the example shown in (a) of FIG. 10 , values of first transform coefficients corresponding to a high-frequency region in a region in which the second transformation is not performed may be transformed to zero. As an example, values of the first transform coefficients in which the sum of the x-axis and y-axis coordinates is equal to or greater than a threshold value may be transformed into 0.
[169]
Alternatively, the first transform coefficients that are converted to 0 may be selected according to a specific form. As an example, as in the example shown in FIG. 10B , the first transform coefficients included in the lower n rows in the region where the second transform is not performed may be transformed to zero. Alternatively, as in the example shown in (c) of FIG. 10 , the first transform coefficients included in the right n columns in the region where the second transform is not performed may be transformed into zero.
[170]
Alternatively, as in the example shown in (d) of FIG. 10 , all first transform coefficients in a region in which the second transform is not performed may be transformed into zero.
[171]
The shape of the region including the first transform coefficients transformed into 0 may be determined based on at least one of a size and shape of a current block, an intra prediction mode, and a transform kernel. Alternatively, an index for specifying one matching the region among a plurality of candidate shapes may be encoded and signaled.
[172]
Whether the second transform is allowed may be determined based on at least one of the encoding mode of the current block and the first transform kernel. Here, the encoding mode refers to intra prediction or inter prediction. For example, when the current block is encoded by intra prediction, the second transform is allowed, whereas when the current block is encoded by inter prediction, the second transform may not be allowed.
[173]
Information indicating whether the second transform has been applied may be encoded and signaled. The information may be a 1-bit flag. Whether the second transform is applied to the current block may be determined according to whether the flag is true or false. Alternatively, the information may be index information. A value of 0 indicates that the second transform is not applied to the current block. On the other hand, a value of the index greater than 0 indicates that the second transform is applied to the current block. When the value of the index is greater than 0, the second transformation kernel can be specified by the index.
[174]
Information indicating whether the second transform has been performed on the current block may be individually encoded for each color component. For example, with respect to each of the luma component (Y), the first chrominance component (Cb), and the second chrominance component (Cr), information indicating whether the second transformation has been performed may be encoded.
[175]
Alternatively, information indicating whether the second transformation has been performed on the chrominance components may be integrated-encoded. For example, to each of the color difference components Cb and Cr, whether to apply the second transform may be jointly determined. That is, the first color difference component Cb and the second color difference component Cr may share information indicating whether the second transformation has been performed.
[176]
Alternatively, based on the tree structure, it may be determined whether the information is encoded for each color component. For example, when the luma component and the chroma component have the same tree structure, three color components (ie, Y, Cb, Cr) may share information indicating whether the second transformation is performed. On the other hand, when the luma component and the chroma component have different tree structures, information indicating whether the second transformation is performed on each of the luma component and the chroma component may be signaled.
[177]
The plurality of second transform kernel candidates may be grouped into at least a plurality of groups. One group may be specified based on at least one of the size, shape, and intra prediction modes of the current block among the plurality of groups. When a group is specified, at least one of a plurality of second transform kernel candidates included in the specified group may be specified using the index information.
[178]
Whether to encode information indicating whether the second transform is applied may be determined based on the position of the last non-zero coefficient in the current block.
[179]
11 illustrates an example in which whether information indicating whether the second transform is applied is determined based on the position of the last non-zero coefficient.
[180]
For convenience of description, when the second transform is performed, it is assumed that values of the transform coefficients are set to 0 in residual regions other than the region in which the second transform coefficients are rearranged.
[181]
As a result of performing the second transform, the number of second transform coefficients equal to the number of rows R of the second transform kernel is generated. As in the above example, all values of the residual transform coefficients except for the second transform coefficients are set to 0, so that non-zero coefficients do not exist in the residual region except for the region in which the R second transform coefficients are rearranged. . According to the above principle, a region in which the R second transform coefficients are rearranged may be set as a limiting region.
[182]
11 (a) to (c), it is exemplified that the upper left 4x4 block in the 8x8 size block is set as the restricted area.
[183]
The presence of a non-zero coefficient outside the limit region indicates that the second transform is not applied to the current block. Accordingly, when the last non-zero coefficient exists outside the limited region, encoding of information indicating whether the second transform is applied to the current block may be omitted. As an example, as in the example shown in FIG. 11A , when the last non-zero coefficient exists outside the restriction region, encoding of information indicating whether the second transform is applied may be omitted.
[184]
The decoder may also determine that the second inverse transform is not applied to the current block without decoding the information when the last non-zero coefficient exists outside the limited region.
[185]
As in the example shown in (b) of FIG. 11 , when the second transform is applied to the current block, a non-zero transform coefficient may exist only in the limited region. Alternatively, as in the example shown in FIG. 11C , even when the second transform is not applied to the current block, a non-zero transform coefficient may exist only in the limited region. Accordingly, when the last non-zero coefficient exists in the restricted region, information indicating whether the second transform is applied may be encoded.
[186]
The decoder may determine whether to apply the second inverse transform to the current block based on the information.
[187]
Alternatively, when the last non-zero coefficient exists in the restricted region, the encoding of the information may be omitted and the second transform may be necessarily applied.
[188]
The size of the restriction region may be determined based on the size of the second transform kernel. As an example, when the second transformation kernel is a matrix having a size of RxN, a rectangular region having a width and a height of Log 2 R, respectively, may be set as a limiting region.
[189]
Alternatively, a region to which the second transform is applied may be set as a limited region.
[190]
Alternatively, information indicating at least one of the size and shape of the restricted region may be encoded and signaled. The information may be signaled through a higher level such as a sequence, a picture header, or a slice header.
[191]
Alternatively, at least one of the size and shape of the restriction region may be predefined in the encoder and the decoder. For example, between the encoder and the decoder, it may be previously agreed to set the upper left 4x4 block in the current block as the limited region.
[192]
Alternatively, at least one of the size and shape of the restriction region may be adaptively determined based on at least one of the size and shape of the current block, the first transform kernel, and the intra prediction mode.
[193]
Alternatively, after defining a plurality of restriction region candidates, an index specifying one of the plurality of restriction region candidates may be encoded and signaled.
[194]
12 illustrates constraint region candidates for a 4x4 block.
[195]
When the current block has a size of 4x4, index information specifying one of the plurality of restriction region candidates shown in FIG. 12 may be encoded.
[196]
At least one of the constrained region candidates shown in FIG. 12 may be applied not only to a block having a size of 4x4 but also to a block having a size larger than 4x4. As an example, at least one of the constrained area candidates shown in FIG. 12 may be applied to a block having at least one of a width or a height of 4 and the other of which is greater than 4.
[197]
Alternatively, the size or number of restriction region candidates may be set differently according to the size of the current block.
[198]
Instead of encoding an index specifying one of the restriction region candidates, one of the restriction region candidates may be specified based on the size or shape of the current block.
[199]
In the above-described example, it has been described that the size of the second transform kernel can be adaptively selected. As another example, a second transform kernel of a predefined size may be applied to all blocks. For example, for all blocks, a second transform kernel having a size of 16x48 may be used. In this case, the second transform may be applied to the 48 first transform coefficients.
[200]
13 is a diagram illustrating an example in which a second transform kernel having a predefined size is applied.
[201]
When the 16x48 second transform kernel is used, as in the example shown in FIG. 13A , the second transform may be applied to an area in which the lower right 4x4 sub-block in the 8x8 size block is excluded. . As an example, a second transform may be applied to the first transform coefficients included in the polygonal shape region shown in FIG. 13A .
[202]
In this case, when at least one of a width or a height of the current block is less than 8, the second transform may not be applied to the current block.
[203]
Alternatively, when at least one of a width or a height of the current block is less than 8, the second transform may be performed after transforming the region to which the second transform is applied to a rectangular shape such as 4x12 or 12x4.
[204]
Alternatively, when at least one of the width or height of the current block is less than 8, the current block and the area to which the second transform is applied after matching the upper left position of the current block with the upper left area of the area to which the second transform is applied The second transform may be applied only to the overlapping region between the two regions.
[205]
13B shows an example in which the second transformation is performed only on the overlapping region.
[206]
After arranging the first transform coefficients included in the overlapping region in one dimension, it may be set as an input matrix for the second transform.
[207]
When the first transform and the second transform are applied to the current block, the decoder performs an inverse transform (second inverse transform) on the second transform, and then inverse transforms the first transform (second inverse transform) on the result of the second inverse transform. 1 inverse transform) to derive residual samples.
[208]
The second inverse transform may be performed based on a transpose matrix of the second transform kernel. As an example, when the second transform kernel has a size of 8x48, the second inverse transform may be performed by a transform kernel having a size of 48x8.
[209]
Second transform coefficients generated by the second transform may be set as an input matrix of the second inverse transform. As an example, when the second transform is performed by a second transform kernel having a size of 8x48, an input matrix having a size of 8x1 including 8 coefficients may be used when the second inverse transform is performed. Thereafter, transform coefficients on which the second inverse transform is performed may be output by a matrix product between the transpose matrix of the second transform kernel and the input matrix. As an example, 48 transform coefficients may be output through matrix multiplication between a 48×8 transform kernel and an 8×1 input matrix.
[210]
After rearranging transform coefficients in the current block, a first inverse transform may be applied to the rearranged block.
[211]
[212]
After quantizing a transform coefficient generated by transforming the residual sample, the quantized transform coefficient may be encoded. Alternatively, quantization may be omitted and the transform coefficients may be encoded.
[213]
When no transform is applied to the current block, the residual sample may be quantized and then the quantized residual sample may be encoded.
[214]
When the transform is skipped, quantization information may be additionally encoded for each block. For example, quantization information for transform skip encoded by applying Differential Pulse-Coded Modulation (DPCM) to quantization information transmitted through a sequence, a picture header, or a slice header may be additionally signaled.
[215]
After quantization is performed, run length encoding may be applied. That is, quantized coefficients generated as a result of quantization may be encoded in a run-length manner. Here, the run means that the same data is continuous, and the run length means the length of the continuous data. As an example, if there is a character string aaaaaabbccccccc, a is continuous 6 times, b is continuous 2 times, and c is 7 consecutive times, so that it can be expressed and encoded as 6a2b7c or a6b2c7.
[216]
The above encoding method may be defined as a run length encoding method.
[217]
For efficient run-length coding, an optimal scan scheme may be determined.
[218]
14 illustrates a scan method.
[219]
According to the scan method shown in FIG. 14 , coefficients are scanned according to a specific directionality. Among the scan schemes shown in FIG. 14 , the one in which the same values are consecutively aligned may be determined as the optimal scan scheme.
[220]
When transform skip is applied to the current block, information specifying a scan method of the current block may be encoded and signaled. The information may be an index specifying one of a plurality of scan methods.
[221]
The number or types of available scan method candidates may be differently set based on at least one of the size, shape, and intra prediction mode of the current block. For example, when the intra prediction mode of the current block is a horizontal direction or a vertical direction, only the two scan methods shown in FIG. 4 may be set as candidates. On the other hand, when the intra prediction mode of the current block is in a diagonal direction (eg, 2, 34, or 66), all four scan methods shown in FIG. 14 may be set as candidates. Accordingly, a bit length allocated to an index for specifying a scan method may be different according to the intra prediction mode of the current block. For example, when the intra prediction mode of the current block is in a horizontal direction or a vertical direction, the index may have a length of 1 bit. On the other hand, when the intra prediction mode of the current block is diagonal, the index may have a length of 2 bits.
[222]
Alternatively, the scan method may be determined based on at least one of the size, shape, and intra prediction mode of the current block. For example, when the intra prediction mode is in the horizontal direction, the horizontal scanning method or the vertical scanning method shown in FIG. 14A may be applied.
[223]
Instead of encoding the quantized coefficients in the run-length method, an encoding method in which additional prediction is applied to the quantized coefficients may be applied. For example, when a residual block is generated through intra prediction within the current block, the residual block may be skipped transform and quantized. When a quantized coefficient is output through the quantization, DPCM may be applied to the output value.
[224]
One of the plurality of direction candidates may be used for DPCM. As an example, horizontal DPCM or vertical DPCM may be applied to the quantized coefficients.
[225]
In the encoder, information for specifying the DPCM direction applied to the quantized residual coefficient may be encoded and signaled. Alternatively, the prediction direction used to generate the prediction block may be set to the DPCM direction.
[226]
The DPCM direction may be used in intra prediction mode prediction. As an example, when the MPM candidate of the current block is derived, if horizontal DPCM is applied to the neighboring block, the MPM may be derived by considering that the intra prediction mode of the neighboring block is in the horizontal direction. Alternatively, when the MPM candidate of the current block is derived, if the vertical direction DPCM is applied to the neighboring block, the MPM may be derived by considering that the intra prediction mode of the neighboring block is the vertical direction. Alternatively, when diagonal DPCM is applied to a neighboring block, it is assumed that the intra prediction mode of the neighboring block is diagonal (eg, 2, 34 or 66) or a non-directional mode (eg, planar or DC) and MPM can induce
[227]
In this case, the MPM directly derived from the DPCM direction may have the highest priority or the lowest priority among MPM candidates. Here, the highest priority means that the lowest index among MPM candidates is allocated (that is, set as the first MPM), and the lowest priority means that the highest index among MPM candidates is allocated. (ie, the one set as the last MPM).
[228]
[229]
For convenience of description, the residual signal related data encoded by the encoder will be referred to as a residual coefficient. For example, the residual coefficient may mean at least one of a quantized transform coefficient, a transform coefficient, or a quantized residual sample according to whether transform or quantization is applied.
[230]
A flag indicating whether a non-zero residual coefficient exists in the current block may be encoded and signaled. When a non-zero residual coefficient exists in the current block, the position of the last non-zero residual coefficient in the scan order may be encoded.
[231]
In addition, a sub-block flag indicating whether a non-zero residual coefficient exists in a sub-block may be encoded in units of sub-blocks in the current block. When a non-zero residual coefficient exists in a subblock, information on each residual coefficient may be additionally encoded according to a scan order.
[232]
In this case, for a subblock having a scan order earlier than a subblock including the last non-zero residual coefficient, encoding of the subblock flag may be omitted. Since non-zero residual coefficients are not included in the sub-block, the value of the sub-block flag may be regarded as zero.
[233]
Also, for a subblock including the last non-zero residual coefficient, encoding of the subblock flag may be omitted. Since a non-zero residual coefficient is necessarily included in the sub-block, the value of the sub-block flag may be regarded as 1.
[234]
As another example, encoding of location information of the last non-zero residual coefficient may be omitted. When encoding of position information of the last non-zero residual coefficient is omitted, sub-block flags may be encoded for all sub-blocks in the current block.
[235]
In this case, when it is determined that non-zero residual coefficients are not included in the remaining sub-blocks except for the last sub-block having the scan order number, it may be understood that the non-zero residual coefficients are necessarily included in the last sub-block. Accordingly, for the last sub-block, encoding of the sub-block flag may be omitted and the value may be regarded as 1.
[236]
Information indicating whether the position information of the last non-zero coefficient is encoded may be additionally encoded. When the position information of the last non-zero coefficient is encoded, the value of the information may be set to 1. In this case, the sub-block flag may be coded from the sub-block in which the last non-zero coefficient exists. On the other hand, when the location information of the last non-zero coefficient is not encoded, the value of the information may be set to 0. In this case, the sub-block flag may be coded from the sub-block having the first scan order.
[237]
When a non-zero residual coefficient exists in the current block, it may be assumed that a non-zero residual coefficient is necessarily included in the first sub-block in the current block. Accordingly, encoding of a sub-block flag indicating whether a non-zero residual coefficient exists in the first sub-block may be omitted.
[238]
The information about each residual coefficient may include at least one of a flag indicating whether the residual coefficient has a non-zero value, information indicating a size of the residual coefficient, and information indicating a sign of the residual coefficient.
[239]
Residual coefficients may be encoded according to a predetermined scan order. In this case, the encoding order of the residual coefficients may be different based on whether the transform is skipped in the current block. For example, when the transform is not skipped in the current block, a residual coefficient located at the lower right of the subblock may be encoded first, and the residual coefficient located at the upper left of the subblock may be encoded last. That is, a scan order between residual coefficients may be determined according to an inverse-diagonal scan, an inverse-horizontal scan, or an inverse-vertical scan. On the other hand, when the transform is skipped in the current block, the residual coefficient located at the upper left of the subblock may be encoded first, and the residual coefficient located at the lower right of the subblock may be encoded last. That is, a scan order between residual coefficients may be determined along a diagonal scan, a horizontal scan, or a vertical scan.
[240]
Alternatively, even when a transform is skipped in the current block, a scan order between residual coefficients may be determined according to an inverse-diagonal scan, an inverse-horizontal scan, or an inverse-vertical scan.
[241]
The scan order of the residual coefficients may be predefined in the encoder and the decoder. Alternatively, information indicating a scan order of residual coefficients may be encoded and signaled. Alternatively, the scan order may be determined based on at least one of a size, a shape, an intra prediction mode of the current block, whether a transform is skipped, or whether a secondary transform is performed.
[242]
15 is a flowchart illustrating a process of encoding residual coefficients in an encoder.
[243]
First, a flag indicating whether the residual coefficient has a non-zero value, significant_flag, may be encoded (S1510). When the value of the residual coefficient is 0, encoding may be performed by setting the value of the flag sig_flag to 0. On the other hand, when the value of the residual coefficient is not 0, encoding may be performed by setting the value of the flag sig_flag to 1. When the value of the residual coefficient is not 0, information on the size of the residual coefficient may be further encoded ( S1520 ).
[244]
16 is a flowchart illustrating a process of encoding size information of residual coefficients.
[245]
The absolute value of the residual coefficient may be encoded using at least one or more gt_N_flags. Here, N may be a natural number of 1 or more. The flag gt_N_flag may indicate whether the absolute value of the residual coefficient has a value greater than 2 (N-1). The number of gt_N_flags used to encode the absolute value of the residual coefficient may be determined based on whether a transform is skipped in the current block. For example, when a transform is not skipped in the current block, two gt_N_flags (N is 1 to 2) may be used. On the other hand, when a transform is skipped in the current block, three or more gt_N_flags (eg, three, four, or five) may be used. In this embodiment, it is assumed that two gt_N_flags are used.
[246]
A flag gt1_flag indicating whether the absolute value of the residual coefficient is greater than 1 may be encoded ( S1610 ). When the absolute value of the residual coefficient is 1, encoding can be performed by setting the value of the flag gt1_flag to 0. On the other hand, when the absolute value of the residual coefficient is greater than 1, encoding can be performed by setting the value of the flag gt1_flag to 1.
[247]
When the absolute value of the residual coefficient is greater than 1, a flag par_flag indicating whether the absolute value of the residual coefficient is an even number or an odd number may be encoded ( S1620 ). When the absolute value of the residual coefficient is an even number, the flag par_flag may be set to 0 to perform encoding. On the other hand, when the absolute value of the residual coefficient is an odd number, encoding may be performed by setting the flag par_flag to 1. Alternatively, on the contrary, the flag par_flag may be set to 1 when the absolute value of the residual coefficient is an even number, and the flag par_flag may be set to 0 when the absolute value of the residual coefficient is an odd number.
[248]
Next, a flag gt_2_flag indicating whether the absolute value of the residual coefficient is greater than 3 may be encoded ( S1630 ). When the absolute value of the residual coefficient is 3 or less, the value of the flag gt_2_flag may be set to 0. On the other hand, when the absolute value of the residual coefficient is greater than 3, the value of the flag gt_2_flag may be set to 1.
[249]
When the absolute value of the residual coefficient is greater than 3, rem_level indicating the residual size may be encoded (S1640). The syntax rem_level may be derived by shifting a value derived by subtracting 4 from the absolute value of the residual coefficient to the right by 1.
[250]
In addition to the gt_1_flag and gt_2_flag flags shown in FIG. 16, gt_N_flag, such as gt_3_flag, gt_4_flag, or gt_5_flag, may be further encoded. In this case, when the value of gt_(N-1)_flag is 1, gt_N_flag may be additionally encoded.
[251]
gt_N_flag may indicate whether the absolute value of the residual coefficient has a value greater than (2N-1). When gt_N_flag is additionally used, rem_level may be derived by shifting a value derived by subtracting 2N from the absolute value of the residual coefficient by right shifting by 1.
[252]
In the above example, it has been exemplified that the absolute value of the residual coefficient is encoded using sig_flag, gt_1_flag, par_flag, gt_2_flag, and rem_level. As another example, the absolute value of the residual coefficient may be encoded as it is. As an example, the syntax abs_level indicating the absolute value of the residual coefficient may be encoded. A method of selecting the encoding method of the absolute value of the residual coefficient will be described later.
[253]
After encoding the size information of the residual coefficient, a flag sign_flag indicating the sign of the residual coefficient may be encoded (S1030). A value of the flag sign_flag of 0 indicates that the residual coefficient is positive. On the other hand, the value of the flag sing_flag being 1 indicates that the residual coefficient is negative.
[254]
Table 1 shows values assigned to each syntax when the value of the residual coefficient is -21 and two gt_N_flags are used.
[255]
[Table 1]
division Equation value
Residual Coefficient (Coeff) coeff -21
sig_flag coeff != 0 One
gt_1_flag !! (|coeff|-1) One
par_flag (|coeff|-2) & 1 One
gt_2_flag (|coeff|-2) >> 1 One
rem_level (|coeff|-4) >> 1 8
sign_flag One
[256]
In Table 1, coeff represents the value of the residual coefficient, and the 'Formula' item represents the formula used to derive the value of each syntax.
[257]
Table 2 shows the values assigned to each syntax when the value of the residual coefficient is -21 and five gt_N_flags are used.
[258]
[Table 2]
division Equation value
Residual Coefficient (Coeff) coeff -21
sig_flag coeff != 0 One
gt_1_flag !! (|coeff|-1) One
par_flag (|coeff|-2) & 1 One
gt_2_flag |coeff| >= 4 One
gt_3_flag |coeff| >= 6 One
gt_4_flag |coeff| >= 8 One
gt_5_flag |coeff| >= 10 One
rem_level (|coeff|-10) >> 1 5
sign_flag One
[259]
17 is a flowchart illustrating a process of decoding a residual coefficient in a decoder.
[260]
When it is determined that non-zero residual coefficients are included in the subblock, the residual coefficients may be reconstructed based on a predetermined scan order.
[261]
First, a flag indicating whether the residual coefficient has a non-zero value, sig_flag, may be decoded (S1710). A value of the flag sig_flag of 0 indicates that the value of the residual coefficient is 0. On the other hand, when the value of the flag sig_flag is 1, it indicates that the value of the residual coefficient is not 0. When the value of the flag sig_flag is 1, information on the magnitude of the residual coefficient may be further decoded (S1720).
[262]
18 is a diagram illustrating a decoding process of size information of residual coefficients.
[263]
For convenience of description, it is assumed that the residual coefficients are coded using a maximum of two gt_N_flags.
[264]
A flag gt1_flag indicating whether the absolute value of the residual coefficient is greater than 1 may be decoded (S1810). A value of the flag gt_1_flag of 0 indicates that the absolute value of the residual coefficient is 1. On the other hand, the value of the flag gt_1_flag equal to 1 indicates that the absolute value of the residual coefficient is greater than 1.
[265]
When the value of the flag gt_1_flag is 1, the flag par_flag indicating whether the absolute value of the residual coefficient is an even number or an odd number may be decoded (S1820). A value of the flag par_flag of 0 indicates that the absolute value of the residual coefficient is an even number, and a value of the flag par_flag of 1 indicates that the absolute value of the residual coefficient is an odd number.
[266]
Next, a flag gt_2_flag indicating whether the absolute value of the residual coefficient is greater than 3 may be decoded ( S1830 ). A value of the flag gt_2_flag of 0 indicates that the absolute value of the residual coefficient is less than 3. When the value of the flag gt_2_flag is 0, the absolute value of the residual coefficient may be determined to be 2 or 3 according to the value of the flag par_flag.
[267]
A value of the flag gt_2_flag of 1 indicates that the absolute value of the residual coefficient is greater than 3.
[268]
When the value of the flag gt_2_flag is 1, rem_level indicating the residual size may be decoded (S1840). An absolute value of the residual coefficient may be derived by adding 3 or 4 to a value derived by shifting the value representing the syntax rem_level to the left by 1.
[269]
In addition to the gt_1_flag and gt_2_flag flags shown in FIG. 18 , gt_N_flag, such as gt_3_flag, gt_4_flag, or gt_5_flag, may be further decoded. In this case, when the value of gt_(N-1)_flag is 1, gt_N_flag may be further decoded.
[270]
gt_N_flag may indicate whether the absolute value of the residual coefficient has a value greater than (2N-1). When gt_N_flag is additionally used, rem_level may be set to a value derived by shifting a value derived by differentiating 2N from the absolute value of the residual coefficient by 1 to the right.
[271]
In the above-described example, it is exemplified that the absolute value of the residual coefficient is decoded using sig_flag, gt_1_flag, par_flag, gt_2_flag, and rem_level. As another example, the absolute value of the residual coefficient may be decoded as it is. As an example, the syntax abs_level indicating the absolute value of the residual coefficient may be decoded. A method of selecting a decoding method for the absolute value of the residual coefficient will be described later.
[272]
After decoding the size information of the residual coefficient, a flag sign_flag indicating the sign of the residual coefficient may be decoded ( S1230 ). A value of the flag sign_flag of 0 indicates that the residual coefficient is positive. On the other hand, the value of the flag sing_flag being 1 indicates that the residual coefficient is negative.
[273]
Table 3 shows an example of decoding a residual coefficient having a value of -21 using two gt_N_flags.
[274]
[Table 3]
division value Equation
sig_flag One -
gt_1_flag One -
par_flag One -
gt_2_flag One -
tmp_coeff 5 1+gt_1_flag+par_flag+(gt_2_flag <<1)
rem_level 8
sign_flag One sign = (sign_flag == 1 ? -1 : 1)
residual coefficient (coeff) -21 tmp_coeff + (rem_level<<1) * sign
[275]
In Table 3, the variable tmp_coeff represents a temporary restoration coefficient. When the value of gt_2_flag is 0, the temporary restoration coefficient tmp_coeff may be set as an absolute value of the residual coefficient. On the other hand, when the value of gt_2_flag is 1, the absolute value of the residual coefficient can be derived by updating the temporary restoration coefficient tmp_coeff based on the syntax rem_level.
[276]
Table 4 shows an example of decoding a residual coefficient having a value of -21 using five gt_N_flags.
[277]
[Table 4]
division value Equation
sig_flag One -
gt_1_flag One -
par_flag One -
tmp_coeff 5 1+gt_1_flag+par_flag
gt_2_flag One tmp_coeff += (sig_2_flag <<1)
gt_3_flag One tmp_coeff += (sig_3_flag <<1)
gt_4_flag One tmp_coeff += (sig_4_flag <<1)
gt_5_flag One tmp_coeff += (sig_5_flag <<1)
rem_level 5 tmp_coeff += (rem_level << 1)
sign_flag One sign = (sign_flag == 1 ? -1 : 1)
residual coefficient (coeff) -21 tmp_coeff + (rem_level<<1) * sign
[278]
In Table 3, the variable tmp_coeff represents a temporary restoration coefficient. When gt_N_flag is 0, the temporary restoration coefficient tmp_coeff may be set as an absolute value of the residual coefficient. On the other hand, when gt_N_flag is 1, the temporary restoration coefficient may be updated (eg, tmp_coeff += sig_N_flag << 1), and the following syntax may be parsed.
[279]
As described, the residual coefficient may be encoded by at least one syntax. A residual coefficient may be changed into a plurality of bins through a binarization process of the syntax(s), and the changed bins may be encoded through entropy encoding.
[280]
Entropy encoding may be divided into encoding using context information and encoding not using context information. The context indicates the probability that the value of the bin is 0 or 1.
[281]
In order to limit the number of bins to be encoded using context information, a threshold may be set. Among the generated bins, bins having a count value less than a threshold value are encoded using context information. When the count value is greater than or equal to the threshold, encoding using context information may no longer be used.
[282]
The threshold value may be determined based on the number of non-zero residual coefficients in the current block. For example, a value obtained by multiplying the number of non-zero residual coefficients in the current block by a real number or a value obtained by adding or subtracting an offset may be set as the threshold value.
[283]
Alternatively, the threshold value may be determined based on the number of pixels included in the current block. For example, a value obtained by multiplying the number of pixels in the current block by a real number or by adding or subtracting an offset may be set as the threshold value.
[284]
Alternatively, information indicating a threshold may be signaled through a bitstream. The information may be encoded through an upper header such as a sequence, picture header, or slice header.
[285]
Alternatively, the threshold value may be determined based on at least one of the size and shape of the current block.
[286]
Alternatively, the threshold value may be determined based on at least one of whether transform skip is applied, a transform kernel applied to the current block, or a quantization parameter.
[287]
When counting the number of bins to be encoded using context information, when information indicating the position of the last non-zero residual coefficient is encoded, the counter may be set not to operate. That is, the information may be excluded from counting.
[288]
Alternatively, when encoding a flag indicating whether a non-zero residual coefficient exists for each subblock in the current block, the counter may be set not to operate. That is, the flag may be excluded from counting.
[289]
According to an embodiment of the present invention, in order to limit the number of bins encoded using context information, when the number of bins encoded using context information is equal to or greater than a threshold value, gt_N_flag or the like is used to limit the residual coefficient. Instead of stepwise encoding , the absolute value of the residual coefficient may be encoded as it is. For example, when the number of bins encoded using context information is smaller than a threshold value, at least one of sig_flag, sign_flag, gt_1_flag, par_flag, gt_2_flag, gt_3_flag, gt_4_flag, gt_5_flag, or rem_level illustrated in Tables 1 to 4 can be used to encode the absolute value of the residual coefficient. On the other hand, when the number of bins encoded using context information is greater than or equal to a threshold value, the syntax abs_level indicating the absolute value of the residual coefficient may be encoded.
[290]
Even in the decoder, a counter may be operated whenever an encoded bin is decoded using context information. When the value of the counter is smaller than the threshold value, the absolute value of the residual coefficient may be reconstructed using at least one of sig_flag, sign_flag, gt_1_flag, par_flag, gt_2_flag, gt_3_flag, gt_4_flag, gt_5_flag, or rem_level. On the other hand, when the value of the counter is equal to or greater than the threshold value, the absolute value of the residual coefficient may be restored using the syntax abs_level.
[291]
19 is a diagram illustrating an example of counting the number of bins using context information.
[292]
For convenience of description, it is assumed that there are 16 residual coefficients in the sub-block, and it is assumed that each of the coefficients is C0 to C15. Here, C15 denotes a residual coefficient located in the lower right corner in the subblock, and C0 denotes a residual coefficient located in the upper left corner in the subblock.
[293]
In addition, it is assumed that the residual coefficients are generated through transformation, and accordingly, the scan order is determined in the order of C15 - C0.
[294]
In addition, it is assumed that the maximum number of bins encoded using context information is 36, and information indicating the position of the last non-zero residual coefficient and a flag indicating whether a non-zero residual coefficient exists in a sub-block are counted assumed to be excluded from
[295]
In FIG. 19 , one pass indicates syntaxes encoded using context information. Except for pass 1, passes 2-1, 2-2, and 3 indicate syntaxes that are encoded without using context information.
[296]
A path indicates an encoding order and a decoding order. For example, the decoder may decode all syntaxes belonging to path 1 and then decode syntaxes belonging to path 2-1. Also, after all syntaxes belonging to the 2-1 path are decoded, the syntaxes belonging to the 3rd path may be decoded.
[297]
In the illustrated example, 2-2 passes represent alternate paths of 1 pass, 2-1 pass and 3 pass.
[298]
When a coefficient of a bin encoded using context information is smaller than a threshold value, an absolute value of a residual coefficient may be encoded through one pass and a 2-1 pass. On the other hand, when a coefficient of a bin encoded using context information is equal to or greater than a threshold value, the absolute value of the residual coefficient may be encoded through 2-2 passes.
[299]
For example, when the first residual coefficient C15 is -21, as in the example shown in Table 2, flags sig_flag, gt_1_flag, par_flag, gt_2_flag, and rem_level may be encoded. When the first residual coefficient C15 is encoded, syntaxes (ie, sig_flag, gt_1_flag, par_flag, and gt_2_flag) that are encoded using context information are encoded using a total of four bins, so the counter increases to 4.
[300]
Since the counter value after encoding the first residual coefficient C15 is less than the threshold value 36, syntaxes encoded using context information may be used also for the second residual coefficient C14. Assuming that four syntaxes encoded using context information are used for each of C15 to C7, after the absolute value of the residual coefficient C7 is encoded, the counter value is set to 36, which is equal to the threshold value.
[301]
Accordingly, when the next residual coefficient C6 is encoded, the absolute value of the residual coefficient C6 can be directly encoded through the syntax abs_level without using syntaxes encoded using context information. That is, for the residual coefficients C6 to C0, the syntax abs_level belonging to the 2-2 path is used instead of the syntax rem_level belonging to the four syntaxes belonging to the 1 pass (ie, sig_flag, gt_1_flag, par_flag, gt_2_flag) and the syntax rem_level belonging to the 2-1 path. An absolute residual coefficient value may be coded.
[302]
In this case, even when the number of bins encoded using the context information is smaller than the threshold value, when the difference between the number and the threshold value is smaller than the maximum number of syntaxes set to be encoded using the context information, the context information It can be set so that abs_level is encoded without using . For example, when sig_flag, gt_1_flag, par_flag, and gt_2_flag are set to be encoded using context information, the syntaxes may be encoded only when the difference between the number and the threshold is greater than 4. On the other hand, when the difference is less than 4, abs_level may be encoded.
[303]
In the illustrated example, it is shown that only four syntaxes belonging to one pass are encoded using context information. Unlike the described example, at least one of the syntax rem_level belonging to the 2-1 path and the syntax sign_flag belonging to the 3rd path may be encoded using context information. For example, when rem_level is encoded using context information, the counter may increase by the number of bins allocated to the syntax rem_level.
[304]
In the example shown in FIG. 19 , the flag par_flag may be set not to be coded using context information. 20 shows an example of this.
[305]
In FIG. 20 , one pass represents syntaxes encoded using context information. Except for pass 1, passes 2, 3-1, 3-2, and 4 indicate syntaxes that are encoded without using context information. When a coefficient of a bin encoded using context information is smaller than a threshold value, the absolute value of the residual coefficient may be encoded through one pass, two passes, and a 3-1 pass. On the other hand, when a coefficient of a bin encoded using context information is equal to or greater than a threshold value, the absolute value of the residual coefficient may be encoded through 3-2 passes.
[306]
When the flag par_flag is set to be encoded without using context information, the counter may be set not to increase with respect to the number of bins (ie, one) allocated to the flag par_flag.
[307]
Accordingly, for each residual coefficient, the counter is incremented only for the bins assigned to the three syntaxes, sig_flag, gt_1_flag and gt_2_flag.
[308]
Assuming that for each of the residual coefficients C15 to C4, the syntax sig_flag, gt_1_flag and gt_2_flag are encoded, after coding the syntax for the residual coefficient C4, the counter is set to 36 equal to the threshold.
[309]
Accordingly, when encoding the residual coefficient C3, the absolute value of the residual coefficient C3 can be encoded as it is through the syntax abs_level included in the 3-2 path. That is, with respect to the residual coefficients C3 to C0, the absolute values of the residual coefficients may be coded using the syntax abs_level belonging to the 3-2 pass instead of the syntaxes belonging to the 1st pass and the 2-1 pass.
[310]
Priorities may also be set between syntaxes encoded using context information. In this case, after counting the number of bins allocated to syntaxes with high priority, the number of bins allocated to syntaxes with low priority may be counted.
[311]
21 illustrates an example in which priorities are different between syntaxes encoded using context information.
[312]
It is assumed that the transformation is skipped in the current block, and accordingly, the scan order is determined in the order of C0 - C15.
[313]
In the example of FIG. 21 , syntaxes belonging to pass 1 and pass 2 may be encoded using context information. In this case, when syntaxes belonging to path 1 have a higher priority than syntaxes belonging to path 2, when counting the number of encoded bins using context information, counting the number of bins allocated to syntaxes belonging to path 1 After this, the number of bins allocated to syntaxes belonging to 2 passes may be counted.
[314]
For example, assuming that 16 residual coefficients are encoded and the threshold value is 96, after syntax belonging to one pass is encoded in all 16 residual coefficients, the counter value is set to 64. Since the counter value is smaller than the threshold value, syntaxes belonging to the second pass can also be coded using context information.
[315]
In the illustrated example, after encoding syntaxes belonging to 2 passes for the residual coefficient C7, the value of the counter is set to 96. Accordingly, syntaxes belonging to 2 passes for the next residual coefficient C8 can be encoded without using context information.
[316]
In the example shown in FIG. 21 , the flag par_flag may be set not to be coded using context information. 22 shows an example of this.
[317]
In FIG. 22 , one pass and three pass indicate syntaxes encoded using context information. 2 pass and 4 pass indicate syntaxes that are encoded without using context information.
[318]
The flag par_flag may be coded without using context information. Accordingly, when the flag par_flag is encoded, the counter may be set not to increase. After encoding syntaxes belonging to 3 passes for the residual coefficient C11, the counter value is set to 96, which is the same as the threshold value. Accordingly, from the next residual coefficient C12, context information may not be used when encoding syntaxes belonging to 3 passes.
[319]
21 and 22 , it is shown that gt_N_flags are distributed in different paths. For example, in the example shown in FIG. 21 , gt_1_flag belongs to one pass, while gt_2_flag is exemplified as belonging to two passes.
[320]
It is also possible to set the distribution of syntaxes to be different from those of FIGS. 21 and 22 . For example, all gt_N_flags may be assigned to one pass, or gt_1_pass and gt_2_pass may be assigned to one pass.
[321]
Instead of setting the flag par_flag as a separate path, par_flag may be set to the same path as gt_1_flag or gt_2_flag, but context information may not be used when encoding the flag par_flag. 23 is an example of this. In the example shown in FIG. 23, the flag par_flag is shown to be assigned to the same path as gt_2_flag, gt_3_flag, gt_4_flag and gt_5_flag.
[322]
The flag par_flag may be assigned to a lower path than gt_N_flag. As an example, in the example shown in FIG. 22 , 2 paths including par_flag may be changed to 3 paths, and the existing 3 paths may be changed to 2 paths. In this case, syntaxes using context information may be encoded first, and then syntaxes not using context information may be encoded.
[323]
It is also possible to encode par_flag and rem_level in three passes.
[324]
[325]
As in the above example, the absolute value of the residual coefficient may be encoded using at least one of sig_flag, par_flag, gt_N_flag, and rem_level. In this case, the remaining syntax except for rem_level may be coded with reference to various context information according to the properties of the neighboring coefficients. As an example, a flag sig_flag representing whether the residual coefficient is 0 or not may be encoded with reference to various context information according to properties of neighboring residual coefficients. In this case, the number of referenceable context information may be determined according to the position of the pixel.
[326]
24 and 25 show a surrounding reconstruction area referenced to determine context information.
[327]
24 is an example of a case in which residual coefficients are encoded according to a scan order from the lower right residual coefficient to the upper left residual coefficient. As an example, FIG. 24 may be applied to a case in which a transform is not skipped in the current block.
[328]
25 is an example in which residual coefficients are encoded according to a scan order from the upper left residual coefficient to the lower right residual coefficient. As an example, FIG. 25 may be applied when a transform is skipped in the current block.
[329]
24 and 25 , a maximum of two or a maximum of five restoration coefficients may be referred to. For example, when the position of the residual coefficient is (x, y), a region including restoration coefficients having an absolute value of 1 or less of the sum of the x-coordinate difference with the residual coefficient and the y-coordinate difference or a restoration coefficient having an absolute value of 2 or less A region including the ? may be set as a peripheral restoration region.
[330]
Alternatively, if the reconstruction coefficients in the reconstruction area with the residual coefficients are out of a block boundary or there is a reconstruction coefficient that has not yet been restored in the scan order, the unavailable reconstruction coefficients may be excluded from the reference target.
[331]
Alternatively, with respect to an unavailable restoration coefficient, context information may be selected by assuming that the information on the corresponding position is a default value. For example, as in the example shown in (a) of FIG. 25 , when a reconstruction region is set and a residual coefficient to be currently encoded is included in the leftmost column in the current block, the sig_flag value of the reconstruction coefficient left of the current residual coefficient is 0. or 1 can be considered.
[332]
Alternatively, when a reconstruction coefficient that deviates from a block boundary is included in a sub-block different from the current residual coefficient but is included in the same coding block, the corresponding reconstruction coefficient may be set to be available.
[333]
It is also possible to determine the context information by referring to a larger number or a smaller number of reconstruction coefficients than the illustrated example. As an example, although not shown in FIG. 24 , only the reconstruction coefficients located on the right side of the residual coefficients and the reconstruction coefficients located below the residual coefficients may be used to determine the context information.
[334]
Alternatively, after allocating an index to each of the plurality of restoration region candidates, an index specifying one of them may be encoded and transmitted to the decoder. Alternatively, the restoration area may be adaptively determined according to the size or shape of the current block. Alternatively, the restoration region may be determined based on the quantization state information QState. As an example, when the variable QState is 0 or 1, a restoration region including up to two restoration coefficients may be used. Alternatively, when QState is 2 or 3, a restoration region including up to 5 restoration coefficients may be used.
[335]
Alternatively, it may be set to refer to one of the N pieces of fixed context information instead of setting the surrounding restoration area. For example, N may be 1. Alternatively, the value of N may be determined according to the position of the residual system. For example, when the sum of x and y is less than the threshold, N may be set to 1, and when the sum of x and y is equal to or greater than the threshold, N may be set to 2. The threshold value may be transmitted to the encoder through an upper header. Alternatively, the threshold value may be previously agreed upon in the encoder and the decoder.
[336]
During sig_flag encoding/decoding, values of sig_flag of the reconstructed coefficients included in the reconstructed area around the residual coefficients may be summed.
[337]
Alternatively, an absolute value of a reconstructed coefficient or a partially reconstructed coefficient included in the reconstructed area around the residual coefficient may be calculated. Here, the absolute value of the partially reconstructed coefficient may mean a temporary reconstruction coefficient derived based on syntaxes included in one pass, for example, (sig_flag + gt_1_flag + par_flag + (gt_2_flag<<1)).
[338]
One of a plurality of context information may be specified using the derived value.
[339]
26 exemplifies the number of referable context information when encoding the flag sig_flag.
[340]
Figure 26 (a) is an example of the luma component, Figure 26 (b) is an example of the chroma component.
[341]
The current block may be divided into a plurality of regions, and types of referable context information may be set differently for each region. As an example, reference is possible in each of a first region in which the sum of x and y coordinates is less than 2, a second region in which the sum of x and y coordinates is 2 or more and less than 5, and a third region in which the sum of x and y coordinates is 5 or more. The types of context information may be different.
[342]
Also, the number of referable context information in each region may be different. As an example, the number of referenceable context information in each region may be fixed.
[343]
Alternatively, the number of referenceable context information may be different for each region. As an example, reference is possible in each of a first region in which the sum of x and y coordinates is less than 2, a second region in which the sum of x and y coordinates is 2 or more and less than 5, and a third region in which the sum of x and y coordinates is 5 or more. The number of context information may be different.
[344]
As another example, the types of referable context information may be set differently according to the quantization state information. The variable QState representing the quantization state information may have a value of 0 to 3. For example, when the variable QState is 0 and 1, the context information of the first type is referred to, while when the QState is 2, the context information of the second type is referred to, and when the QState is 3, the context information of the third type is referred. can
[345]
In the example shown in (a) of FIG. 26, it is exemplified that the luma block is divided into three regions, and the number of referenceable context information for each region is four. If it is assumed that three types of context information are available according to the quantization state information, a total of 36 (3x4x3) context information in the luma block may be set in a referable state.
[346]
In the example shown in (b) of FIG. 26, it is exemplified that the chroma block is divided into two regions, and the number of referenceable context information for each region is four. If it is assumed that three types of context information are available according to the quantization state information, a total of 24 (2x4x3) context information in a chroma block may be set to a referable state.
[347]
Depending on the size of the reconstruction area (ie, the number of reconstruction coefficients included in the reconstruction area), the number of referable context information may be set differently.
[348]
Alternatively, the number of referenceable context information may be different depending on whether transform skip is applied to the current block. For example, when transform skip is applied to the current block, three or five pieces of context information can be referenced, whereas when transform is applied to the current block, four context information can be referenced.
[349]
In the case of using the sum of the values of sig_flag of the surrounding reconstruction coefficients, the derived sum may be compared with a threshold value. Here, the threshold value may be set equal to the number of referenceable context information. Depending on the number of restoration coefficients included in the restoration area, the sum value may have a range of 0 to 2 or a range of 0 to 5. When the sum value is greater than the threshold value, the sum value may be converted into a threshold value. For example, when the sum value is 5 and the threshold value is 4, the sum value may be changed to 4. Then, based on the sum value, one of the plurality of context information may be specified. That is, the sum value may be used as an index for specifying one of a plurality of context information.
[350]
When the sum of absolute values of the surrounding reconstruction coefficients is used, after calculating the sum of the absolute values of the reconstruction coefficients in the surrounding reconstruction area, the derived value may be divided by a predefined value. For example, the predefined value may be a natural number such as 2, 3, 4 or 5. Alternatively, the sum of the absolute values may be divided by the number of reconstruction coefficients in the surrounding reconstruction area.
[351]
A result value derived by the division operation may be compared with a threshold value. In this case, when the result value is greater than the threshold value, the derived value may be converted into the threshold value. For example, when the threshold value is 3 and the result value is greater than 3, the result value may be converted to 3. Accordingly, the result value may be set to a value between 0 and 3.
[352]
According to the derived result value, context information to be referred to when encoding/decoding a corresponding residual coefficient may be specified. That is, the result value may function as an index for specifying one of a plurality of context information. Accordingly, the threshold value may be determined based on the number of referenceable context information in the region including the residual coefficient.
[353]
Residual coefficients may be encoded/decoded using simplified context information. For example, when the variable QState is 0 or 1, the number of referable context information may be four. On the other hand, when the variable QState is 2 or 3, the number of referable context information may be two.
[354]
Alternatively, the referenceable number for each region may be set differently. For example, when the variable QState is 2 or 3, the number of referenceable context information may be set to four in the third region, while the number of referable context information may be set to two in the first region and the second region. .
[355]
When encoding/decoding residual coefficients, when the number of referenceable context information is two, the threshold value may be set to 1. In this case, since the result value is a value of 0 or 1, one of the two context information can be specified by the result value.
[356]
Alternatively, when the number of referenceable context information is two, the context information may be specified based on whether a value derived by a modulo operation between the sum of absolute values of the reconstruction coefficients in the surrounding reconstruction area and 2 is 0 or not.
[357]
Even when encoding/decoding gt_N_flag or par_flag, context information may be determined with reference to a neighboring reconstruction region.
[358]
27 exemplifies the number of referenceable context information when encoding gt_N_flag or par_flag.
[359]
27 (a) illustrates a luma block, and FIG. 27 (b) illustrates a chroma block.
[360]
In the case of a luma block, the block may be divided into a plurality of regions. In the example shown in (a) of FIG. 27, the luma block includes a first area including the residual coefficient of the (0, 0) position, a second area in which the sum of the x-axis and y-axis coordinates is 1 or more and less than 3, It is exemplified that the sum of the x-axis and y-axis coordinates is 3 or more and is divided into a third region smaller than 10 and a fourth region excluding the first to third regions.
[361]
In this case, the last non-zero residual coefficient may be set as the fifth region. Since the last non-zero residual coefficient is encoded/decoded first in the scan order, only one context information for the last non-zero residual coefficient may be set to be referenceable.
[362]
In the case of the chroma block, the number of divided regions may be smaller than that of the luma block. As an example, in (b) of FIG. 27 , it is exemplified that the chroma block is divided into a first region including a residual coefficient of (0, 0) and a second region excluding the first region.
[363]
In this case, the third region including the last non-zero residual coefficient may be set.
[364]
In the example shown in FIG. 27 , it is exemplified that the number of referenceable context information is 5 in each region except for the region including the last non-zero residual coefficient. In this case, the types of referable context information for each region may be different. Accordingly, the total number of referrable context information in the luma block may be 21 (4x5+1), and the total number of referable context information in the chroma block may be 11 (2x5+1).
[365]
As described above, the number and/or type of referenceable context information may be set differently according to region or quantization state information.
[366]
During encoding/decoding of gt_N_flag or par_flag, at least one of the sum of absolute values of residual coefficients included in the peripheral reconstruction region and the sum of sig_flag may be derived. Thereafter, a result value derived by differentiating the sum of sig_flag from the sum of absolute values may be compared with a threshold value. In this case, when the result value is greater than the threshold value, the result value may be converted into the threshold value. For example, when the threshold value is 4, a result value greater than 4 may be converted to 4. As a result, the result value is set to a value ranging from 0 to 4. Based on the result value, one of the five context information may be specified. That is, the result value may function as an index for specifying one of a plurality of context information.
[367]
In a simplified way, it is also possible to determine which contextual information should be referred to. For example, when encoding the residual coefficients in each region, only one predefined context information may be referred to while omitting the process of deriving the sum of the absolute values of the restoration coefficients in the surrounding reconstructed region. The simplified method for determining context information may be applied to both the luma block and the chroma block, or may be applied to only one of the luma block and the chroma block.
[368]
For example, in the luma block, a method of specifying one of a plurality of context information based on a result value may be used, while in the chroma block, a method of using predefined context information may be used.
[369]
The division method and number of regions are not limited to the illustrated example. As an example, at least one of a partitioning method or the number may be determined in consideration of at least one of the size, shape, and transform skip application of the current block and the location of the last non-zero residual coefficient. Alternatively, information specifying any one of the partitioning method or the number of regions in the current block may be encoded and signaled through a higher header.
[370]
According to the size of the restoration area (ie, the number of restoration coefficients), the number of referenceable context information may be differently set.
[371]
For example, when context information is determined using five reconstruction coefficients, one of up to five pieces of context information may be specified during gt_N_flag or par_flag encoding. On the other hand, when context information is determined using two restoration coefficients, one of up to three pieces of context information may be specified. When the number of referenceable context information is reduced, the threshold value may also be reduced.
[372]
In addition, the size of the reconstructed area may be set to be different for each area, or the size of the reconstructed area may be set to be different for each color component.
[373]
Alternatively, the context information may be derived by comparing information on each of the reconstruction coefficients rather than summing the information on each of the reconstruction coefficients in the surrounding reconstruction area. As an example, when par_flag encoding/decoding is performed, it is assumed that a reconstructed region is set as shown in (a) of FIG. 27 . In this case, with respect to the left reconstruction coefficient and the upper reconstruction coefficient of the residual coefficient, context information to be referenced when encoding/decoding par_flag for the current residual coefficient may be determined by referring to par_flag of each reconstruction coefficient.
[374]
For example, when par_flag is not coded for both the left and top reconstruction coefficients (that is, when at least one of sig_flag or gt_1_flag is 0 for each of the left and top reconstruction coefficients), or the value of par_flag of the two reconstruction coefficients This different case may be defined as the first case. A case in which par_flag values of both restoration coefficients are all 1 may be defined as a second case, and a case in which par_flag values of both restoration coefficients are all 0 may be defined as a third case. When par_flag encoding of the current residual coefficient is performed, it may be set to refer to different context information for each case. That is, after allocating an index (0 to 2) to each case, one piece of context information can be specified based on the index.
[375]
[376]
The names of syntaxes used in the above-described embodiments are merely named for convenience of description.
[377]
It is included in the scope of the present disclosure to apply the decoding process or the embodiments described based on the encoding process to the encoding process or the decoding process. It is also within the scope of the present disclosure to change the embodiments described in a certain order in an order different from that described.
[378]
Although the above-described embodiment has been described based on a series of steps or a flowchart, this does not limit the time-series order of the invention, and may be performed simultaneously or in a different order, if necessary. In addition, each of the components (eg, unit, module, etc.) constituting the block diagram in the above-described embodiment may be implemented as a hardware device or software, or a plurality of components may be combined to form one hardware device or software. may be implemented. The above-described embodiment may be implemented in the form of program instructions that can be executed through various computer components and recorded in a computer-readable recording medium. The computer-readable recording medium may include program instructions, data files, data structures, etc. alone or in combination. Examples of the computer-readable recording medium include a hard disk, a magnetic medium such as a floppy disk and a magnetic tape, an optical recording medium such as a CD-ROM, a DVD, and a magneto-optical medium such as a floppy disk. media), and hardware devices specially configured to store and execute program instructions such as ROM, RAM, flash memory, and the like. The hardware device may be configured to operate as one or more software modules to perform processing according to the present disclosure, and vice versa.
Industrial Applicability
[379]
The present invention can be applied to an electronic device capable of encoding/decoding an image.
Claims
[Claim 1]
determining whether an inverse transform is skipped in the current block; decoding the residual coefficients of the current block; and selectively applying the inverse transform to the residual coefficients based on whether the determination is made, wherein when decoding the residual coefficients, a first syntax indicating whether the residual coefficients are greater than 0 and an absolute value of the residual coefficients An image decoding method, characterized in that any one of the second syntax representing
[Claim 2]
The image decoding method according to claim 1, wherein it is determined whether to decode the first syntax or the second syntax by comparing the number of decoded bins and a threshold value using context information. Way.
[Claim 3]
The method of claim 2, wherein when at least one of the first syntax, at least one gt_N_flag indicating whether the absolute value has a value greater than (2N-1), or a parity flag indicating whether the absolute value is an even number, is decoded, An image decoding method, characterized in that the number of bins decoded using the context information increases.
[Claim 4]
The gt_1_flag of claim 1, wherein when the first syntax is decoded and the first syntax indicates that the residual coefficient has a non-zero value, gt_1_flag indicating whether the absolute value of the residual coefficient has a value greater than 1. is further decoded, the image decoding method.
[Claim 5]
5. The method of claim 4, wherein when the gt_1_flag indicates that the absolute value has a value greater than 1, a parity flag indicating whether the absolute value is an even number and gt_2_flag indicating whether the absolute value is greater than 3 are further decoded A video decoding method, characterized in that.
[Claim 6]
The method of claim 2 , wherein the threshold value is determined based on a size of the current block.
[Claim 7]
determining whether to skip transform in the current block; quantizing a result to which a transform is applied or a result in which a transform is skipped; and encoding a residual coefficient output as a result of the quantization, wherein when encoding the residual coefficient, a first syntax indicating whether the residual coefficient is greater than 0 and a second syntax indicating an absolute value of the residual coefficient An image encoding method, characterized in that any one is alternatively encoded.
[Claim 8]
The image encoding method according to claim 8, wherein it is determined whether to encode the first syntax or the second syntax by comparing the number of encoded bins and a threshold value using context information. Way.
[Claim 9]
The method according to claim 8, wherein when at least one of the first syntax, at least one gt_N_flag indicating whether an absolute value has a value greater than (2N-1), or a parity flag indicating whether an absolute value is an even number, is encoded, An image encoding method, characterized in that the number of bins encoded using the context information increases.
[Claim 10]
The gt_1_flag of claim 7, wherein the first syntax is encoded, and when the first syntax indicates that the residual coefficient has a non-zero value, gt_1_flag indicating whether an absolute value of the residual coefficient has a value greater than 1. is additionally encoded, the video encoding method.
[Claim 11]
11. The method of claim 10, wherein when the gt_1_flag indicates that the absolute value has a value greater than 1, a parity flag indicating whether the absolute value is an even number and gt_2_flag indicating whether the absolute value is greater than 3 are additionally encoded. An image encoding method, characterized in that.
[Claim 12]
The method of claim 8 , wherein the threshold value is determined based on a size of the current block.
[Claim 13]
A computer-readable recording medium storing a bitstream encoded by a video encoding method, the video encoding method comprising the steps of: determining whether to skip transform in a current block; quantizing a result to which a transform is applied or a result in which a transform is skipped; and encoding a residual coefficient output as a result of the quantization, wherein when encoding the residual coefficient, a first syntax indicating whether the residual coefficient is greater than 0 and a second syntax indicating an absolute value of the residual coefficient A computer-readable recording medium, characterized in that any one is alternatively encoded.
| # | Name | Date |
|---|---|---|
| 1 | 202217023001.pdf | 2022-04-19 |
| 2 | 202217023001-TRANSLATIOIN OF PRIOIRTY DOCUMENTS ETC. [19-04-2022(online)].pdf | 2022-04-19 |
| 3 | 202217023001-STATEMENT OF UNDERTAKING (FORM 3) [19-04-2022(online)].pdf | 2022-04-19 |
| 4 | 202217023001-NOTIFICATION OF INT. APPLN. NO. & FILING DATE (PCT-RO-105-PCT Pamphlet) [19-04-2022(online)].pdf | 2022-04-19 |
| 5 | 202217023001-FORM 1 [19-04-2022(online)].pdf | 2022-04-19 |
| 6 | 202217023001-DRAWINGS [19-04-2022(online)].pdf | 2022-04-19 |
| 7 | 202217023001-DECLARATION OF INVENTORSHIP (FORM 5) [19-04-2022(online)].pdf | 2022-04-19 |
| 8 | 202217023001-COMPLETE SPECIFICATION [19-04-2022(online)].pdf | 2022-04-19 |
| 9 | 202217023001-FORM-26 [10-05-2022(online)].pdf | 2022-05-10 |
| 10 | 202217023001-Proof of Right [13-05-2022(online)].pdf | 2022-05-13 |
| 11 | 202217023001-FORM 3 [06-10-2022(online)].pdf | 2022-10-06 |
| 12 | 202217023001-FORM 18 [29-08-2023(online)].pdf | 2023-08-29 |