Abstract: The present invention relates to an image processing device and method with which decoding processing can be performed at the correct timing. A HRD-type setting unit determines, on the basis of a user instruction, a HRD-parameter type, i.e. parameters for managing a decoder buffer, and acquires, from a storage buffer and a storage buffer (subordinate tier) of a base-layer image encoder, information indicating a storage state, in accordance with the determined HRD-parameter type. The HRD-type setting unit calculates HRD parameters on the basis of the acquired information, and causes the HRD parameters and flags indicating the HRD-parameter type to be encoded in a reversible encoder. The present invention is application to image processing devices.
The present disclosure relates to an image processing device and method, and
more particularly, to an image processing device and method which are capable of
performing a decoding process at a proper timing in scalable video coding.
[Background Art]
[0002]
Recently, devices for compressing and encoding an image by adopting a
encoding scheme of handling image information digitally and performing compression
by an orthogonal transform such as a discrete cosine transform and motion compensation
using image information-specific redundancy for the purpose of information
transmission and accumulation with high efficiency when the image information is
handled digitally have become widespread. Moving Picture Experts Group (MPEG),
H.264, MPEG-4 Part 10 (Advanced Video Coding) (hereinafter referred to as
H.264/AVC), and the like are examples of such encoding schemes.
[0003]
Therefore, for the purpose of improving encoding efficiency compared to
H.264/AVC, standardization of a encoding scheme referred to as high efficiency video
coding (HEVC) by Joint Collaboration Team-Video Coding (JCTVC), which is a joint
standardizing organization of International Telecommunication Union
Telecommunication Standardization Sector (ITU-T) and International Organization for
Standardization (ISO)/International Electrotechnical Commission (IEC), is currently in
progress, and Non-Patent Literature 1 has been issued as a draft of the scheme.
[0004]
Meanwhile, the existing image encoding schemes such as MPEG-2 and AVC
have a scalability function of dividing an image into a plurality of layers and encoding
the plurality of layers.
4
[0005]
In other words, for example, for a terminal having a low processing capability
such as a mobile phone, image compression informat ion of only a base layer is
transmitted, and a moving image of low spatial and temporal resolutions or a low quality
is reproduced, and for a terminal having a high processing capability such as a television
or a personal computer, image compression information of an enhancement layer as well
as a base layer is transmitted, and a moving image of high spatial and temporal
resolutions or a high quality is reproduced. That is, image compression information
according to a capability of a terminal or a network can be transmitted from a server
without performing the transcoding process.
[0006]
In the HEVC, it is possible to designate a hypothetical reference decoder (HRD)
parameter so that an overflow or an underflow of a buffer does not occur when a
decoding process for image compression information is performed. Particularly, it is
possible to designate an HRD parameter for each layer when scalable video coding is
performed (see Non-Patent Literature 2).
[Prior Art Literatures]
[Non-Patent Literatures]
[0007]
[Non-Patent Literature 1] Benjamin Bross, Woo-Jin Han, Jens-Rainer
Ohm, Gary J. Sullivan, Thomas Wiegand," High efficiency video coding (HEVC) text
specification draft 9," JCTVC-K1003, 2012,10,21
[Non-Patent Literature 2] Jill Boyce, Ye-Kui Wang, "NAL unit header and
parameter set designs for HEVC extensions," JCTVC-K1007, 2012,10,19
[Summary of the Invention]
[Problem to Be Solved by the Invention]
[0008]
However, when the HRD parameter is designated for each layer or a time layer
5
serving as one of sublayers, it is difficult to detect whether the decoding process is
performed by a single decoding device or a plurality of decoding devices.
[0009]
The present disclosure was made in light of the foregoing, and it is desirable to
perform a decoding process at a proper timing.
[Means for Solving the Problem]
[0010]
According to one aspect of the present disclosure, there is provided an image
processing device including a receiving section configured to receive a bitstream
obtained by encoding an image having at least one layer and buffer management
parameter information of each layer indicating at least one of that a parameter for
managing a decoder buffer is a parameter for performing a decoding process of only a
corresponding layer and that the parameter for managing the decoder buffer is a
parameter for performing a decoding process of a corresponding layer and a lower layer ,
and a decoding section configured to decode the bitstream received by the receiving
section and generate an image.
[0011]
The layer can include a layer and a sublayer.
[0012]
The layer is a view of multi-view coding.
[0013]
The layer is a layer of scalable video coding.
[0014]
According to one aspect of the present disclosure, there is provided an image
processing method including receiving, by an image processing device, a bitstream
obtained by encoding an image having at least one layer and buffer management
parameter information of each layer indicating at least one of that a parameter for
managing a decoder buffer is a parameter for performing a decoding process of only a
6
corresponding layer and that the parameter for managing the decoder buffer is a
parameter for performing a decoding process of a corresponding layer and a lower layer,
receiving, by the image processing device, a bitstream obtained by encoding an image
including at least one layer using a parameter corresponding to the buffer management
parameter information, and decoding, by the image processing device, the received
bitstream and generating an image.
[0015]
According to another aspect of the present disclosure, there is provided an image
processing device including a setting section configured to set buffer management
parameter information of each layer indicating at least one of that a parameter for
managing a decoder buffer is a parameter for performing a decoding process of only a
corresponding layer and that the parameter for managing the decoder buffer is a
parameter for performing a decoding process of a corresponding layer and a lower layer,
an encoding section configured to encode an image having at least one layer and
generate a bitstream, and a transmitting section configured to transmit the buffer
management parameter information set by the setting section and the bitstream generated
by the encoding section.
[0016]
The layer can include a layer and a sublayer.
[0017]
The layer is a view of multi-view coding.
[0018]
The layer is a layer of scalable video coding.
[0019]
According to another aspect of the present disclosure, there is provided an image
processing method including setting, by an image processing device, buffer management
parameter information of each layer indicating at least one of that a parameter for
managing a decoder buffer is a parameter for performing a decoding process of only a
corresponding layer and that the parameter for managing the decoder buffer is a
7
parameter for performing a decoding process of a corresponding layer and a lower layer,
encoding, by the image processing device, an image having at least one layer and
generating a bitstream, and transmitting, by the image processing device, the set buffer
management parameter information and the generated bitstream.
[0020]
According to one aspect of the present disclosure, a bitstream obtained by
encoding an image having at least one layer and buffer management parameter
information of each layer indicating at least one of that a parameter for managing a
decoder buffer is a parameter for performing a decoding process of only a corresponding
layer and that the parameter for managing the decoder buffer is a parameter for
performing a decoding process of a corresponding layer and a lower layer are received.
Then, the received bitstream is decoded to generate an image.
[0021]
According to another aspect of the present disclosure, buffer management
parameter information of each layer indicating at least one of that a parameter for
managing a decoder buffer is a parameter for performing a decoding process of only a
corresponding layer and that the parameter for managing the decoder buffer is a
parameter for performing a decoding process of a corresponding layer and a lower layer
is set, and an image having at least one layer is encoded to generate a bitstream. Then,
the set buffer management parameter information and the generated bitstream are
transmitted.
[0022]
Also, the above-described image processing device may be an independent
device or an inner block constituting one image encoding device or image decoding
device.
[Effects of the Invention]
[0023]
According to one aspect of the present disclosure, it is possible to decode an
8
image. Particularly, it is possible to perform a decoding process at a proper timing.
[0024]
According to another aspect of the present disclosure, it is possible to encode an
image. Particularly, it is possible to perform a decoding process at a proper timing.
[Brief Description of the Drawings]
[0025]
[FIG. 1] FIG. 1 is a diagram for describing an example of a configuration of a
coding unit.
[FIG. 2] FIG. 2 is a diagram for describing an example of spatial scalable video
coding.
[FIG. 3] FIG. 3 is a diagram for describing an example of temporal scalable
video coding.
[FIG. 4] FIG. 4 is a diagram for describing an example of scalable video coding
of a signal to noise ratio.
[FIG. 5] FIG. 5 is a diagram illustrating an example of syntax of an HRD
parameter of a HEVC.
[FIG. 6] FIG. 6 is a diagram for describing a parallel process of scalability video
coding.
[FIG. 7] FIG. 7 is a diagram illustrating an example of syntax of an HRD
parameter according to the present technology.
[FIG. 8] FIG. 8 is a diagram illustrating another example of syntax of an HRD
parameter according to the present technology.
[FIG. 9] FIG. 9 is a diagram illustrating another example of syntax of an HRD
parameter according to the present technology.
[FIG. 10] FIG. 10 is a block diagram illustrating an example of a main
configuration of a scalable encoding device.
[FIG. 11] FIG. 11 is a block diagram illustrating an example of a main
configuration of an enhancement layer image encoding section.
9
[FIG. 12] FIG. 12 is a block diagram illustrating an example of a main
configuration of an accumulation buffer and an HRD type setting section.
[FIG. 13] FIG. 13 is a diagram for describing an example of a layer structure.
[FIG. 14] FIG. 14 is a flowchart for describing an example of a flow of an
encoding process.
[FIG. 15] FIG. 15 is a flowchart for describing an example of a layer encoding
process.
[FIG. 16] FIG. 16 is a flowchart for describing an HRD parameter encoding
process.
[FIG. 17] FIG. 17 is a flowchart for describing an HRD parameter calculation
process.
[FIG. 18] FIG. 18 is a flowchart for describing an HRD parameter of the time
layer calculation process.
[FIG. 19] FIG. 19 is a flowchart for describing another example of an HRD
parameter encoding process.
[FIG. 20] FIG. 20 is a block diagram illustrating an example of a main
configuration of a scalable decoding device.
[FIG. 21] FIG. 21 is a block diagram illustrating an example of a main
configuration of an enhancement layer image decoding section.
[FIG. 22] FIG. 22 is a block diagram illustrating an example of a main
configuration of an accumulation buffer and an HRD type decoding section.
[FIG. 23] FIG. 23 is a flowchart for describing an example of a flow of a
decoding process.
[FIG. 24] FIG. 24 is a flowchart for describing an example of a flow of a layer
decoding process.
[FIG. 25] FIG. 25 is a flowchart for describing another example of a flow of an
HRD parameter decoding process.
[FIG. 26] FIG. 26 is a flowchart for describing another example of a flow of an
accumulation buffer monitoring process.
10
[FIG. 27] FIG. 27 is a diagram illustrating an example of a multi-view image
encoding scheme.
[FIG. 28] FIG. 28 is a diagram illustrating an example of a main configuration of
a multi-view image encoding device to which the present disclosure is applied.
[FIG. 29] FIG. 29 is a diagram illustrating an example of a main configuration of
a multi-view image decoding device to which the present disclosure is applied.
[FIG. 30] FIG. 30 is a block diagram illustrating an example of a main
configuration of a computer.
[FIG. 31] FIG. 31 is a block diagram illustrating an example of a schematic
configuration of a television device.
[FIG. 32] FIG. 32 is a block diagram illustrating an example of a schematic
configuration of a mobile phone.
[FIG. 33] FIG. 33 is a block diagram illustrating an example of a schematic
configuration of a recording/reproduction device.
[FIG. 34] FIG. 34 is a block diagram illustrating an example of a schematic
configuration of an image capturing device.
[FIG. 35] FIG. 35 is a block diagram illustrating an example of using scalable
video coding.
[FIG. 36] FIG. 36 is a block diagram illustrating another example of using
scalable video coding.
[FIG. 37] FIG. 37 is a block diagram illustrating another example of using
scalable video coding.
[Modes for Carrying out the Invention]
[0026]
Hereinafter, modes (hereinafter referred to as “embodiments”) for carrying out
the present disclosure will be described. The description will proceed in the following
order:
0. Overview
1. First embodiment (image encoding device)
11
2. Second embodiment (image decoding device)
3. Third embodiment (multi-view image encoding device/multi-view image
decoding device)
4. Fourth embodiment (computer)
5. Applications
6. Applications of scalable video coding
[0027]
<0. Overview>
Hereinafter, the present technology will be described in connection with an
application to image encoding and decoding of a High Efficiency Video Coding (HEVC)
scheme.
[0028]
In an Advanced Video Coding (AVC) scheme, a hierarchical structure based on a
macroblock and a sub macroblock is defined. However, a macroblock of 16×16 pixels
is not optimal for a large image frame such as a Ultra High Definition (UHD)
(4000×2000 pixels) serving as a target of a next generation encoding scheme.
[0029]
On the other hand, in the HEVC scheme, a coding unit (CU) is defined as
illustrated in FIG. 1.
[0030]
A CU is also referred to as a coding tree block (CTB), and serves as a partial
area of an image of a picture unit undertaking the same role of a macroblock in the AVC
scheme. The latter is fixed to a size of 16×16 pixels, but the former is not fixed to a
certain size but designated in image compression information in each sequence.
[0031]
For example, a largest coding unit (LCU) and a smallest coding unit (SCU) of a
CU are specified in a sequence parameter set (SPS) included in encoded data to be
12
output.
[0032]
As split-flag=1 is set in a range in which each LCU is not smaller than an SCU,
a coding unit can be divided into CUs having a smaller size. In the example of FIG. 1,
a size of an LCU is 128, and a largest scalable depth is 5. A CU of a size of 2N×2N is
divided into CUs having a size of N×N serving as a layer that is one-level lower when a
value of split_flag is 1.
[0033]
Further, a CU is divided in prediction units (PUs) that are areas (partial areas of
an image of a picture unit) serving as processing units of intra or inter prediction, and
divided into transform units (TUs) that are areas (partial areas of an image of a picture
unit) serving as processing units of orthogonal transform. Currently, in the HEVC
scheme, in addition to 4×4 and 8×8, orthogonal transform of 16×16 and 32×32 can be
used.
[0034]
As in the HEVC scheme, in the case of an encoding scheme in which a CU is
defined and various kinds of processes are performed in units of CUs, in the AVC
scheme, a macroblock can be considered to correspond to an LCU, and a block (sub
block) can be considered to correspond to a CU. Further, in the AVC scheme, a motion
compensation block can be considered to correspond to a PU. Here, since a CU has a
hierarchical structure, a size of an LCU of a topmost layer is commonly set to be larger
than a macroblock in the AVC scheme, for example, such as 128×128 pixels.
[0035]
Thus, hereinafter, an LCU is assumed to include a macroblock in the AVC
scheme, and a CU is assumed to include a block (sub block) in the AVC scheme. In
other words, a “block” used in the following description indicates an arbitrary partial
area in a picture, and, for example, a size, a shape, and characteristics thereof are not
limited. In other words, a “block” includes an arbitrary area (a processing unit) such as
a TU, a PU, an SCU, a CU, an LCU, a sub block, a macroblock, or a slice. Of course, a
13
“block” includes other partial areas (processing units) as well. When it is necessary to
limit a size, a processing unit, or the like, it will be appropriately described.
[0036]
Meanwhile, in the AVC and HEVC encoding schemes, in order to achieve high
encoding efficiency, it is important to select an appropriate prediction mode.
[0037]
As an example of such a selection method, there is a method implemented in
reference software (found at http://iphome.hhi.de/suehring/tml/index.htm) of
H.264/MPEG-4 AVC called a joint model (JM).
[0038]
In the JM, as will be described later, it is possible to select two mode
determination methods, that is, a high complexity mode and a low complexity mode. In
both modes, cost function values related to respective prediction modes are calculated,
and a prediction mode having a smaller cost function value is selected as an optimal
mode for a corresponding block or macroblock.
[0039]
A cost function in the high complexity mode is represented as in the following
Formula (1):
[0040]
Cost(ModeΩ)=D+λ*R ... (1)
[0041]
Here, Ω indicates a universal set of candidate modes for encoding a
corresponding block or macroblock, and D indicates differential energy between a
decoded image and an input image when encoding is performed in a corresponding
prediction mode. λ indicates Lagrange’s undetermined multiplier given as a function of
a quantization parameter. R indicates a total coding amount including an orthogonal
transform coefficient when encoding is performed in a corresponding mode.
[0042]
14
In other words, in order to perform encoding in the high complexity mode, it is
necessary to perform a temporary encoding process once by all candidate modes in order
to calculate the parameters D and R, and thus a large computation amount is required.
[0043]
A cost function in the low complexity mode is represented by the following
Formula (2):
[0044]
Cost(ModeΩ)=D+QP2Quant(QP)*HeaderBit ... (2)
[0045]
Here, D is different from that of the high complexity mode and indicates
differential energy between a prediction image and an input image. QP2Quant (QP) is
given as a function of a quantization parameter QP, and HeaderBit indicates a coding
amount related to information belonging to a header such as a motion vector or a mode
including no orthogonal transform coefficient.
[0046]
In other words, in the low complexity mode, it is necessary to perform a
prediction process for respective candidate modes, but since a decoded image is not
necessary, it is unnecessary to perform an encoding process. Thus, it is possible to
implement a computation amount smaller than that in the high complexity mode.
[0047]
Meanwhile, the existing image encoding schemes such as MPEG2 and AVC have
a scalability function as illustrated in FIGS. 2 to 4. Scalable video coding refers to a
scheme of dividing (hierarchizing) an image into a plurality of layers and per forming
encoding for each layer.
[0048]
In hierarchization of an image, one image is divided into a plurality of images
(layers) based on a certain parameter. Basically, each layer is configured with
differential data so that redundancy is reduced. For example, when one image is
15
hierarchized into two layers, that is, a base layer and an enhancement layer, an image of
a lower quality than an original image is obtained using only data of the base layer, and
an original image (that is, a high-quality image) is obtained by combining data of the
base layer with data of the enhancement layer.
[0049]
As an image is hierarchized as described above, it is possible to obtain images
of various qualities according to the situation. For example, for a terminal havin g a
low processing capability such as a mobile phone, image compression information of
only a base layer is transmitted, and a moving image of low spatial and temporal
resolutions or a low quality is reproduced, and for a terminal having a high processing
capability such as a television or a personal computer, image compression information of
an enhancement layer as well as a base layer is transmitted, and a moving image of high
spatial and temporal resolutions or a high quality is reproduced. In other words, image
compression information according to a capability of a terminal or a network can be
transmitted from a server without performing the transcoding process.
[0050]
As a parameter having scalability, for example, there is spatial resolution
(spatial scalability) as illustrated in FIG. 2. When the spatial scalability differs,
respective layers have different resolutions. In other words, each picture is
hierarchized into two layers, that is, a base layer of a resolution spatially lower than that
of an original image and an enhancement layer that is combined with an image of the
base layer to obtain an original image (an original spatial resolution) as illustrated in FIG.
2. Of course, the number of layers is an example, and each picture can be hierar chized
into an arbitrary number of layers.
[0051]
As another parameter having such scalability, for example, there is a temporal
resolution (temporal scalability) as illustrated in FIG. 3. In the case of the temporal
scalability, respective layers have different frame rates. In other words, in this case,
each picture is hierarchized into layers having different frame rates, a moving image of a
16
high frame rate can be obtained by combining a layer of a high frame rate with a layer of
a low frame rate, and an original moving image (an original frame rate) can be obtained
by combining all the layers as illustrated in FIG. 3. The number of layers is an example,
and each picture can be hierarchized into an arbitrary number of layers.
[0052]
Further, as another parameter having such scalability, for example, there is a
signal-to-noise ratio (SNR) (SNR scalability). In the case of the SNR scalability,
respective layers having different SNRs. In other words, in this case, each picture is
hierarchized into two layers, that is, a base layer of an SNR lower than that of an
original image and an enhancement layer that is combined with an image of the base
layer to obtain an original SNR as illustrated in FIG. 4. In other words, for base layer
image compression information, information related to an image of a low PSNR is
transmitted, and a high PSNR image can be reconstructed by combining the information
with the enhancement layer image compression information. Of course, the number of
layers is an example, and each picture can be hierarchized into an arbitrary number of
layers.
[0053]
A parameter other than the above-described examples may be applied as a
parameter having scalability. For example, there is bit -depth scalability in which the
base layer includes an 8-bit image, and a 10-bit image can be obtained by adding the
enhancement layer to the base layer.
[0054]
Further, there is chroma scalability in which the base layer includes a
component image of a 4:2:0 format, and a component image of a 4:2:2 format c an be
obtained by adding the enhancement layer to the base layer.
[0055]
Further, as a parameter having scalability, there is a multi -view. In this case,
an image is hierarchized into layers of different views.
[0056]
17
For example, layers described in the present embodiment include spatial,
temporal, SNR, bit depth, color, and view of scalability video coding described above.
[0057]
Further, a term “layer” used in this specification includes a layer of scalable
video coding and each view when a multi-view of a multi-view is considered.
[0058]
Further, the term “layer” used in this specification is assumed to include a main
layer (corresponding to sub) and a sublayer. As a specific example, a main layer may
be a layer of spatial scalability, and a sublayer may be configured with a layer of
temporal scalability.
[0059]
In the present embodiment, a layer (Japanese) and a layer have the same
meaning, a layer (Japanese) will be appropriately described as a layer.
[0060]
Meanwhile, in the HEVC, when the decoding process of image compression
information is performed, it is possible to designate a Hypothetical Reference Decoder
(HDR) parameter illustrated in FIG. 5 so that no overflow or underflow of a buffer
occurs. In other words, the HRD parameter is a parameter used to manage a decoder
buffer. Particularly, when scalable video coding is performed, it is possible to
designate the HRD parameter for each layer in a video parameter set (VPS).
[0061]
In an example of FIG. 6, two examples (ex1 and ex2) using a sequence including
an I picture, a b picture, a B picture, a b picture, and a B picture are illustrated on the
left side of FIG. 6 as an example of temporal scalability. In this sequence, the I picture,
the B picture, and the B picture are a lower time layer, and the b picture and the b
picture are an upper time layer.
[0062]
18
Here, the B picture indicates a picture that is referred to, and the b picture
indicates a picture that is not referred to.
[0063]
ex1 is an example in which all the pictures are decoded by a decoding device #0.
On the other hand, ex2 is an example in which the lower time layer of the I picture, the
B picture, and the B picture is decoded by the decoding device #0, and the upper time
layer of the b picture and the b picture is decoded by a decoding device #1.
[0064]
On the right side of FIG. 6, as a scalable HEVC example, two examples (ex11
and ex12) using a sequence including the I picture, the B picture, and the B pi cture of
the EL (enhancement layer) serving as the upper layer and the I picture, the B picture,
and the B picture of the BL (base layer) serving as the lower layer are illustrated. The
scalable HEVC means scalable video coding defined in the HEVC.
[0065]
ex11 is an example in which all the pictures are decoded by the decoding device
#0. On the other hand, ex12 is an example in which the lower layer of the I picture, the
B picture, and the B picture of the BL is decoded by the decoding device #0, and the
upper layer of the I picture, the B picture, and the B picture of the EL is decoded by the
decoding device #1.
[0066]
For each layer of the scalable HEVC, each time layer of temporal scalability at
the right side is configured as a sublayer.
[0067]
As described above, in the temporal scalability of the related art, in the scalable
HEVC, a process may be performed by a single decoding device, and a parallel process
may be performed by a plurality of decoding devices. Further, through the syntax of
FIG. 5, it is possible to designate the HRD parameter serving as the parameter used to
manage the decoder buffer for each layer or a time layer that is one of sublayers.
[0068]
19
However, as illustrated in FIG. 6, it is difficult to detect whether the decoding
process is performed by a single decoding device or a plurality of decoding devices.
[0069]
In this regard, in the present technology, the HRD parameter is transmitted
through syntax illustrated in FIG. 7. In other words, in the present technology,
information indicating whether the HRD parameter is a parameter for performing a
decoding process only in a corresponding layer or a parameter for performing a decoding
process of a corresponding layer and a lower layer is set. Thus, as illustrated in FIG. 6,
it is clearly defined whether the decoding process is performed by a single decoding
device or a plurality of decoding devices, and thus it is possible to perform a decoding
process at a proper timing.
[0070]
FIG. 7 is a diagram illustrating an example of syntax of the HRD parameter
according to the present technology. Numbers on the left of each row are row numbers
added for description.
[0071]
In an example of FIG. 7, hrd_parameters_type_flag is defined in a 10th row.
When a value of hrd_parameters_type_flag is 1, a value for performing a decoding
process of only a corresponding layer is set as an HRD parameter of a subsequent
paragraph. When a value of hrd_parameters_type_flag is 0, a value for performing a
decoding process of a corresponding layer and a lower layer is set as an HRD parameter
of a subsequent paragraph.
[0072]
Further, hrd_parameters_type_flag may be included in an if statement starting
from an 11th row.
[0073]
sub_hrd_parameters_type[i]_flag is defined in a 25th row. When a value of
sub_hrd_parameters_type[i]_flag is 1, a value for performing a decoding process of only
20
a corresponding time layer is set as a sub HRD parameter of a subsequent paragraph.
When a value of hrd_parameters_type_flag is 0, a value for performing a decoding
process of a corresponding time layer and a lower time layer is set as a sub HRD
parameter of a subsequent paragraph.
[0074]
FIG. 7 illustrates the example in which the designating is performed by any one
method (either only a corresponding layer is included or a lower layer is also included)
for each layer and each time layer, but the present technology is not limited to this
example. For example, the HRD parameter in which the HRD parameter is designated
by both methods may be included.
[0075]
FIG. 8 is a diagram illustrating another example of syntax of the HRD parameter
according to the present technology. Numbers on the left of each row are row numbers
added for description.
[0076]
In an example of FIG. 8, hrd_parameters_type1_present_flag is defined in an
11th row. When a value of hrd_parameters_type1_present_flag is 1, a value for
performing a decoding process of only a corresponding layer is set as an HRD parameter
of type1 set in 13th to 24th rows. When a value of hrd_parameters_type1_present_flag
is 0, a value for performing a decoding process of a corresponding layer and a lower
layer is set as the HRD parameter of type1.
[0077]
hrd_parameters_type2_present_flag is defined in a 12th row. When a value of
hrd_parameters_type2_present_flag is 1, a value for performing a decoding process of
only a corresponding layer is set as an HRD parameter of type2 defined in 25th to 36th
rows. When a value of hrd_parameters_type1_present_flag is 0, a value for performing
a decoding process of a corresponding layer and a lower layer is set as the HRD
parameter of type2.
21
[0078]
Similarly to the example described above with reference to FIG. 7, the flags of
the 11th and 12th rows may be described before an if statement starting from a 10th row.
[0079]
sub_hrd_parameters_type1_present_flag is defined in a 40th row. When a
value of sub_hrd_parameters_type1_present_flag is 1, a value for performing a decoding
process of only a corresponding time layer is set as an HRD parameter of type1 set in
45th to 52nd rows. When a value of sub_hrd_parameters_type1_present_flag is 0, a
value for performing a decoding process of a corresponding time layer and a lower time
layer is set as the HRD parameter of type1.
[0080]
sub_hrd_parameters_type2_present_flag is defined in a 41st row. When a
value of sub_hrd_parameters_type2_present_flag is 1, a value for performing a decoding
process of only a corresponding time layer is set as an HRD parameter of type2 set in
53rd to 60th rows. When a value of sub_hrd_parameters_type1_present_flag is 0, a
value for performing a decoding process of a corresponding time layer and a lower time
layer is set as an HRD parameter of type2.
[0081]
As described above, in the present technology, the parameters of type1 and
type2 serving as the HRD parameter for the decoding process of only the corresponding
layer and the HRD parameter for the decoding process of the corresponding layer and
the lower layer are set at the encoding side. Thus, the decoding side can select the
parameter according to a device or a received bitstream.
[0082]
Further, when the image compression information includes only one layer, that
is, a scalability layer or a temporal scalability layer, hrd_parameter_type_f lag and
sub_hrd_parameter_type_flag may have any value, and the decoding process is not
affected.
[0083]
22
Next, the present technology will be described in connection with applications
to a specific device. For the sake of convenience of description, the following
description will proceed with a case of ex12 of the scalable HEVC and ex1 of temporal
scalability in FIG. 6. Here, the present technology is not limited to this case. For
example, there may be a case of ex12 of the scalable HEVC and ex2 of temporal
scalability in FIG. 6, a case of ex11 of the scalable HEVC and ex2 of temporal
scalability in FIG. 6, and a case of ex11 of the scalable HEVC and ex1 of temporal
scalability in FIG. 6.
[0084]
<1. First embodiment>
FIG. 10 is a block diagram illustrating an example of a main configuration of a
scalable encoding device.
[0085]
A scalable encoding device 100 illustrated in FIG. 10 encodes each layer of
image data hierarchized into a base layer and an enhancement layer.
[0086]
The scalable encoding device 100 is configured to include a base layer image
encoding section 101-1, an enhancement layer image encoding section 101-2, and an
encoding control section 102.
[0087]
The base layer image encoding section 101-1 acquires image information (base
layer image information) of the base layer. The base layer image encoding section 101 -
1 encodes the base layer image information without referring to other layers, generates
encoded data (base layer encoded data) of the base layer, and outputs the generated
encoded data.
[0088]
The enhancement layer image encoding section 101-2 acquires image
information (enhancement layer image information) of the enhancement layer. The
23
enhancement layer image encoding section 101-2 encodes the enhancement layer image
information. At this time, the enhancement layer image encoding section 101-2
performs inter-layer prediction with reference to information related to encoding of the
base layer as necessary.
[0089]
Further, the enhancement layer image encoding section 101-2 sets the HRD
parameter type for each layer, and calculates the HRD parameter serving as the
parameter used to manage the decoder buffer based on state information of an
accumulation buffer according to the set HRD parameter type. The enhancement layer
image encoding section 101-2 encodes the calculated HRD parameter.
[0090]
Specifically, the HRD parameter type indicates whether the HRD parameter is
the parameter for decoding only a corresponding layer or the parameter for performi ng
the decoding process of a corresponding layer and a lower layer. This type may set
both of the parameters as well as any one of the parameters. The encoding side
calculates the HRD parameter according to the flag (information) indicating the set type,
and transmits the flag indicating the set type and the calculated HRD parameter to the
decoding side. Hereinafter, the flag indicating the HRD parameter type is appropriately
referred to as an “HRD parameter type flag.”
[0091]
When the flag indicating the HRD parameter type is 1, the enhancement layer
image encoding section 101-2 calculates the HRD parameter based on the state
information of its own accumulation buffer. When the flag indicating the HRD
parameter type is 0, the enhancement layer image encoding section 101-2 acquires state
information of the whole accumulation buffer of the base layer image encoding section
101-1, and calculates the HRD parameter based on the state information of the base
layer image encoding section 101-1 and its own accumulation buffer. This process is
performed on a layer and a sublayer (time layer). In the base layer image encoding
section 101-1, this process is performed only on the sublayer.
24
[0092]
The enhancement layer image encoding section 101-2 generates encoded data
(enhancement layer encoded data) of the enhancement layer through the above encoding,
and outputs the generated encoded data.
[0093]
The base layer image encoding section 101-1 and the enhancement layer image
encoding section 101-2 are appropriately referred to collectively as a “layer image
encoding section 101.”
[0094]
The encoding control section 102 controls the encoding process of the layer
image encoding sections 101, for example, in view of the reference relation of the layer
image encoding sections 101.
[0095]
In the example of FIG. 10, one enhancement layer image encoding section 101-2
is illustrated, but when there is an upper layer, enhancement layer image encoding
sections 101-3, 101-4, ... that encode the upper layer are provided for each of the upper
layers.
[0096]
FIG. 11 is a block diagram illustrating an example of a main configuration of the
enhancement layer image encoding section 101-2. The base layer image encoding
section 101-1 of FIG. 10 has basically the same configuration as the enhancement layer
image encoding section 101-2 of FIG. 11 except that a type of an image serving as a
target is different. For the sake of convenience of description, in the example of FIG.
11, a configuration of the enhancement layer image encoding section 101-2 will be
described as an example.
[0097]
As illustrated in FIG. 11, the enhancement layer image encoding section 101-2
includes an A/D converting section 111, a screen reordering buffer 112, an operation
25
section 113, an orthogonal transform section 114, a quantization section 115, a lossless
encoding section 116, an accumulation buffer 117, an inverse quantization section 118,
and an inverse orthogonal transform section 119. The enhancement layer image
encoding section 101-2 further includes an operation section 120, a loop filter 121, a
frame memory 122, a selecting section 123, an intra prediction section 124, a motion
prediction/compensation section 125, a predictive image select ing section 126, and a
rate control section 127. The enhancement layer image encoding section 101-2 further
includes an HRD type setting section 128.
[0098]
The A/D converting section 111 performs A/D conversion on input image data
(the enhancement layer image information), and supplies the converted image data
(digital data) to be stored in the screen reordering buffer 112. The screen reordering
buffer 112 reorders images of frames stored in a display order in a frame order for
encoding according to a Group Of Pictures (GOP), and supplies the images in which the
frame order is reordered to the operation section 113. The screen reordering buffer 112
also supplies the images in which the frame order is reordered to the intra prediction
section 124 and the motion prediction/compensation section 125.
[0099]
The operation section 113 subtracts a predictive image supplied from the intra
prediction section 124 or the motion prediction/compensation section 125 via the
predictive image selecting section 126 from an image read from the screen reordering
buffer 112, and outputs differential information thereof to the orthogonal transform
section 114. For example, in the case of an image that has been subjected to intra
coding, the operation section 113 subtracts the predictive image supplied from the intra
prediction section 124 from the image read from the screen reordering buffer 112.
Further, for example, in the case of an image that has been subjected to inter coding, the
operation section 113 subtracts the predictive image supplied from the motion
prediction/compensation section 125 from the image read from the screen reordering
buffer 112.
26
[0100]
The orthogonal transform section 114 performs an orthogonal transform such as
a discrete cosine transform or a Karhunen-Loève Transform on the differential
information supplied from the operation section 113. The orthogonal transform section
114 supplies transform coefficients to the quantization section 115.
[0101]
The quantization section 115 quantizes the transform coefficients supplied from
the orthogonal transform section 114. The quantization section 115 sets a quantization
parameter based on information related to a target value of a coding amount supplied
from the rate control section 127, and performs the quantizing. The quantization
section 115 supplies the quantized transform coefficients to the lossless encoding section
116.
[0102]
The lossless encoding section 116 encodes the transform coefficients quantized
in the quantization section 115 according to an arbitrary encoding scheme. Since
coefficient data is quantized under control of the rate control section 127, the coding
amount becomes a target value (or approaches a target value) set by the rate control
section 127.
[0103]
The lossless encoding section 116 acquires information indicating an intra
prediction mode or the like from the intra prediction section 124, and acquires
information indicating an inter prediction mode, differential motion vector information,
or the like from the motion prediction/compensation section 125. Further, the lossless
encoding section 116 appropriately generates an NAL unit of the enhancement layer
including a sequence parameter set (SPS), a picture parameter set (PPS), and the like.
[0104]
The lossless encoding section 116 encodes various kinds of information
according to an arbitrary encoding scheme, and sets (multiplexes) the encoded
information as part of encoded data (also referred to as an “encoded stream”). The
27
lossless encoding section 116 supplies the encoded data obtained by the encoding to be
accumulated in the accumulation buffer 117.
[0105]
Examples of the encoding scheme of the lossless encoding section 116 include
variable length coding and arithmetic coding. As the variable length coding, for
example, there is Context-Adaptive Variable Length Coding (CAVLC) defined in the
H.264/AVC scheme. As the arithmetic coding, for example, there is Context -Adaptive
Binary Arithmetic Coding (CABAC).
[0106]
The accumulation buffer 117 temporarily holds the encoded data (enhancement
layer encoded data) supplied from the lossless encoding section 116. The accumulation
buffer 117 outputs the held enhancement layer encoded data to a recording device
(recording medium), a transmission path, or the like (not illustrated) at a subsequent
stage at a certain timing. In other words, the accumulation buffer 117 serves as a
transmitting section that transmits the encoded data as well. Further, when there is a
request from the HRD type setting section 128, the accumulation buffer 117 supplies
information indicating a state of the accumulation buffer 117. Further, for example,
when there is an enhancement layer image encoding section 101-3 of an upper layer, and
there is a request from its HRD type setting section 128 as indicated by a dotted line, the
accumulation buffer 117 supplies the information indicating the state of the
accumulation buffer 117 to the enhancement layer image encoding section 101-3 of the
upper layer.
[0107]
The transform coefficients quantized by the quantization section 115 are also
supplied to the inverse quantization section 118. The inverse quantization section 118
inversely quantizes the quantized transform coefficients according to a method
corresponding to the quantization performed by the quant ization section 115. The
inverse quantization section 118 supplies the obtained transform coefficients to the
inverse orthogonal transform section 119.
28
[0108]
The inverse orthogonal transform section 119 performs an inverse orthogonal
transform on the transform coefficients supplied from the inverse quantization section
118 according to a method corresponding to the orthogonal transform process performed
by the orthogonal transform section 114. An output (restored differential information)
that has been subjected to the inverse orthogonal transform is supplied to the operation
section 120.
[0109]
The operation section 120 obtains a locally decoded image (a decoded image) by
adding the predictive image supplied from the intra prediction section 124 or the motion
prediction/compensation section 125 via the predictive image selecting section 126 to
the restored differential information serving as an inverse orthogonal transform result
supplied from the inverse orthogonal transform section 119. The decoded image is
supplied to the loop filter 121 or the frame memory 122.
[0110]
The loop filter 121 includes a deblock filter, an adaptive offset filter, an adaptive
loop filter, or the like, and appropriately performs a filter process on the reconstructed
image supplied from the operation section 120. For example, the loop filter 121
performs the deblock filter process on the reconstructed image, and removes block
distortion of the reconstructed image. Further, for example, the loop filter 121
improves the image quality by performing the loop filter process on the deblock filter
process result (the reconstructed image from which the block distortion has been
removed) using a Wiener filter. The loop filter 121 supplies the filter process result
(hereinafter referred to as a “decoded image”) to the frame memory 122.
[0111]
The loop filter 121 may further perform any other arbitrary filter process on the
reconstructed image. The loop filter 121 may supply information used in the filter
process such as a filter coefficient to the lossless encoding section 116 as necessary so
that the information can be encoded.
29
[0112]
The frame memory 122 stores the reconstructed image supplied from the
operation section 120 and the decoded image supplied from the loop filter 121. The
frame memory 122 supplies the stored reconstructed image to the intra prediction section
124 via the selecting section 123 at a certain timing or based on an external request, for
example, from the intra prediction section 124. Further, the frame memory 122
supplies the stored decoded image to the motion prediction/compensation section 125
via the selecting section 123 at a certain timing or based on an external request, for
example, from the motion prediction/compensation section 125.
[0113]
The frame memory 122 stores the supplied decoded image, and supplies the
stored decoded image to the selecting section 123 as a reference image at a certain
timing.
[0114]
The selecting section 123 selects a supply destination of the reference image
supplied from the frame memory 122. For example, in the case of the intra prediction,
the selecting section 123 supplies the reference image (a pixel value of a current picture)
supplied from the frame memory 122 to the motion prediction/compensation section 125.
Further, for example, in the case of the inter prediction, the selecting section 123
supplies the reference image supplied from the frame memory 122 to the motion
prediction/compensation section 125.
[0115]
The intra prediction section 124 performs the int ra prediction (intra-screen
prediction) for generating the predictive image using the pixel value of the current
picture serving as the reference image supplied from the frame memory 122 via the
selecting section 123. The intra prediction section 124 performs the intra prediction in
a plurality of intra prediction modes that are prepared in advance.
[0116]
The intra prediction section 124 generates predictive images in all the intra
30
prediction modes serving as the candidates, evaluates cost function value s of the
predictive images using the input image supplied from the screen reordering buffer 112,
and selects an optimal mode. When the optimal intra prediction mode is selected, the
intra prediction section 124 supplies the predictive image generated in the optimal mode
to the predictive image selecting section 126.
[0117]
As described above, the intra prediction section 124 appropriately supplies, for
example, the intra prediction mode information indicating the employed intra prediction
mode to the lossless encoding section 116 so that the information is encoded.
[0118]
The motion prediction/compensation section 125 performs the motion prediction
(the inter prediction) using the input image supplied from the screen reordering buffer
112 and the reference image supplied from the frame memory 122 via the selecting
section 123. Although not illustrated, in the motion prediction/compensation section
125, the reference image supplied from the frame memory 122 of the base layer image
encoding section 101-1 is also referred to as necessary. The motion
prediction/compensation section 125 performs a motion compensation process according
to a detected motion vector, and generates a predictive image (inter -predictive image
information). The motion prediction/compensation section 125 performs the inter
prediction in a plurality of inter prediction modes that are prepared in advance.
[0119]
The motion prediction/compensation section 125 generates predictive images in
all the inter prediction modes serving as a candidate. The motion
prediction/compensation section 125 evaluates cost function values of the predictive
images using the input image supplied from the screen reordering buffer 112,
information of the generated differential motion vector, and the like, and selects an
optimal mode. When the optimal inter prediction mode is selected, the motion
prediction/compensation section 125 supplies the predictive image generated in the
optimal mode to the predictive image selecting section 126.
31
[0120]
The motion prediction/compensation section 125 supplies information indicating
the employed inter prediction mode, information necessary for performing processing in
the inter prediction mode when the encoded data is decoded, and the like to the lossless
encoding section 116 so that the information is encoded. For example, as the necessary
information, there is information of a generated differential motion vector, and as
prediction motion vector information, there is a flag indicating an index of a prediction
motion vector.
[0121]
The predictive image selecting section 126 selects a supply source of the
prediction image to be supplied to the operation section 113 and the operation section
120. For example, in the case of the intra coding, the predictive image selecting
section 126 selects the intra prediction section 124 as the supply source of the predictive
image, and supplies the predictive image supplied from the intra prediction section 124
to the operation section 113 and the operation section 120. For example, in the case of
the inter coding, the predictive image selecting section 126 selects the motion
prediction/compensation section 125 as the supply source of the predictive image, and
supplies the predictive image supplied from the motion prediction/compensation section
125 to the operation section 113 and the operation section 120.
[0122]
The rate control section 127 controls a rate of a quantization operation of the
quantization section 115 based on the coding amount of the encoded data accumulated in
the accumulation buffer 117 such that no overflow or underflow occurs.
[0123]
The HRD type setting section 128 decides the HRD parameter type according to
the user’s instruction, and acquires information indicating an accumulation state from
the accumulation buffer 117 or the accumulation buffer (the lower layer) 117 of the base
layer image encoding section 101-1 according to the decided HRD parameter type. The
HRD type setting section 128 calculates the HRD parameter based on the acquired
32
information, and causes the lossless encoding section 116 to encode the flag indicating
the HRD parameter type and the HRD parameter.
[0124]
Further, when the image compression information (encoded data) to be output
includes one layer, the value of the flag indicating the HRD parameter type is arbitrary,
and does not affect the process at the decoding side.
[0125]
FIG. 12 is a block diagram illustrating an example of a configuration of the
accumulation buffer and the HRD type setting section of FIG. 11.
[0126]
In an example of FIG. 12, the accumulation buffer 117 is configured to include a
partial accumulation buffer 131 and a whole accumulation buffer 132.
[0127]
The HRD type setting section 128 is configured to include a layer HRD
parameter calculating section 141, a time layer HRD parameter calculating section 142,
an HRD parameter type setting section 143, and a time HRD parameter type setting
section 144.
[0128]
The partial accumulation buffer 131 is configured with each accumulation buffer
that accumulates encoded data related to each upper time layer among the encoded data
(codes) accumulated in the whole accumulation buffer 132. The information indicating
the state of each accumulation buffer is supplied to the time layer HRD parameter
calculating section 142 on request.
[0129]
The whole accumulation buffer 132 accumulates the encoded data (codes)
encoded by the lossless encoding section 116. Further, information indicating a state of
the whole accumulation buffer of the whole accumulation buffer 132 is supplied to the
33
layer HRD parameter calculating section 141 and the time layer HRD parameter
calculating section 142 on request. Further, there are cases in which there is the
enhancement layer image encoding section 101-3 of the upper layer as indicated by a
dotted line. In this case, the information indicating the state of the whole accumulation
buffer of the whole accumulation buffer 132 is also supplied to the HRD type setting
section (upper layer) 127 according to the request of the HRD type setting section (upper
layer) 127 of the enhancement layer image encoding section 101-3.
[Claim 1]
An image processing device comprising:
a receiving section configured to receive a bitstream obtained by encoding an
image having at least one layer and buffer management parameter information of each
layer indicating at least one of that a parameter for managing a decoder buffer is a
parameter for performing a decoding process of only a corresponding layer and that the
parameter for managing the decoder buffer is a parameter for performing a decoding
process of a corresponding layer and a lower layer; and
a decoding section configured to decode the bitstream received by the receiving
section and generate an image.
[Claim 2]
The image processing device according to claim 1,
wherein the layer includes a layer and a sublayer.
[Claim 3]
The image processing device according to claim 2,
wherein the layer is a view of multi-view coding.
[Claim 4]
The image processing device according to claim 2,
wherein the layer is a layer of scalable video coding.
[Claim 5]
An image processing method comprising:
receiving, by an image processing device, a bitstream obtained by encoding an
image having at least one layer and buffer management parameter information of each
layer indicating at least one of that a parameter for managing a decoder buffer is a
parameter for performing a decoding process of only a corresponding layer and that the
parameter for managing the decoder buffer is a parameter for performing a decoding
process of a corresponding layer and a lower layer; and
decoding, by the image processing device, the received bitstream and generating
98
an image.
[Claim 6]
An image processing device comprising:
a setting section configured to set buffer management parameter information of
each layer indicating at least one of that a parameter for managing a decoder buffer is a
parameter for performing a decoding process of only a corresponding layer and that the
parameter for managing the decoder buffer is a parameter for performing a decoding
process of a corresponding layer and a lower layer;
an encoding section configured to encode an image having at least one layer and
generate a bitstream; and
a transmitting section configured to transmit the buffer management parameter
information set by the setting section and the bi tstream generated by the encoding
section.
[Claim 7]
The image processing device according to claim 6,
wherein the layer includes a layer and a sublayer.
[Claim 8]
The image processing device according to claim 7,
wherein the layer is a view of multi-view coding.
[Claim 9]
The image processing device according to claim 7,
wherein the layer is a layer of scalable video coding.
[Claim 10]
An image processing method comprising:
setting, by an image processing device, buffer management parameter
information of each layer indicating at least one of that a parameter for managing a
decoder buffer is a parameter for performing a decoding process of only a corresponding
layer and that the parameter for managing the decoder buffer is a parameter for
performing a decoding process of a corresponding layer and a lower layer;
99
encoding, by the image processing device, an image having at least one layer
and generating a bitstream; and
transmitting, by the image processing device, the set buffer management
parameter information and the generated bitstream.
| # | Name | Date |
|---|---|---|
| 1 | Priority Document [10-11-2016(online)].pdf | 2016-11-10 |
| 2 | Power of Attorney [10-11-2016(online)].pdf | 2016-11-10 |
| 3 | Form 5 [10-11-2016(online)].pdf | 2016-11-10 |
| 4 | Form 3 [10-11-2016(online)].pdf | 2016-11-10 |
| 5 | Form 20 [10-11-2016(online)].pdf | 2016-11-10 |
| 6 | Form 1 [10-11-2016(online)].pdf | 2016-11-10 |
| 7 | Drawing [10-11-2016(online)].pdf | 2016-11-10 |
| 8 | Description(Complete) [10-11-2016(online)].pdf | 2016-11-10 |
| 9 | abstract.jpg | 2017-01-13 |
| 10 | Other Patent Document [22-02-2017(online)].pdf | 2017-02-22 |
| 11 | Form 18 [25-02-2017(online)].pdf | 2017-02-25 |
| 12 | Form 3 [28-04-2017(online)].pdf | 2017-04-28 |
| 13 | 201618038475-FER.pdf | 2021-10-17 |
| 1 | SearchStrategyE_21-08-2020.pdf |