Abstract: The present disclosure relates to an image decoding device and method which allows more accurate determination of the performance required for decoding. Both encoded data produced from image data and decoding load definition information which defines the amount of load for decode-processing a partial region of the image represented by the image data, are obtained, and the decoding of the obtained encoded data is controlled on the basis of the obtained decoding load definition information, decoding the obtained encoded data in accordance with the control. The present disclosure can be applied to image processing devices such as, e.g., image encoding devices which scalably encode image data and image decoding devices which decode encoded data obtained by scalably encoding image data.
The present disclosure relates to an image decoding device and method, and
particularly, relates to an image decoding device capable of recognizing performance
10 necessary for decoding more accurately and a method.
Background Art
[0002]
In recent years, in order to further improve coding efficiency over MPEG-4
15 Part10 (Advanced Video Coding, hereinafter referred to as “AVC”), Joint
Collaboration Team-Video Coding (JCTVC), which is a joint standardization
organization of International Telecommunication Union Telecommunication
Standardization Sector (ITU-T) and International Organization for
Standardization/International Electrotechnical Commission (ISO/IEC), has
20 proceeded with standardization of a coding scheme called High Efficiency Video
Coding (HEVC) (for example, refer to Non-Patent Literature 1).
[0003]
In HEVC, it is possible to decode only a region whose decoding is
necessary by an application using a tile structure. In order to indicate the fact that a
25 tile region is independently decodable, second and later versions (including MVHEVC,
SHVC, Range Ext. and the like) of HEVC are supported by motionconstrained
tile sets SEI.
Citation List
Patent Literature
30 Non-Patent Literature
[0004]
SP353442WO00
2/125
[Non-Patent Literature 1] Benjamin Bross, Woo-Jin Han, Jens-Rainer Ohm,
Gary J. Sullivan, Ye-Kui Wang, Thomas Wiegand, “High Efficiency Video Coding
(HEVC) text specification draft 10 (for FDIS & Last Call)”, JCTVC-L1003_v34,
Joint Collaborative Team on Video Coding (JCT-VC) of ITU-T SG 16 WP 3 and
ISO/IEC JTC 1/5 SC 29/WG 11 12th Meeting: Geneva, CH, 14-23 Jan. 2013.
Summary of Invention
Technical Problem
[0005]
10 However, as information on a level, which serves as a reference for
determining whether a decoder can decode a stream, and a buffer capacity, only a
value of the entire stream or a value of a layer unit is defined.
[0006]
Therefore, even in an application that decodes only a part of an entire image,
15 determination of whether decoding is possible is performed by assuming a load when
an entire screen is decoded. Accordingly, there is concern of an unnecessarily high
level decoder being necessary. In addition, there is concern of applications to be
delivered being unnecessarily limited accordingly.
[0007]
20 The present disclosure has been made in view of the above-mentioned
problems and can recognize performance necessary for decoding more accurately.
Solution to Problem
[0008]
25 An aspect of the present technology is an image decoding device including:
an acquisition unit configured to acquire coded data of image data and decoding load
definition information for defining a magnitude of a load of a decoding process of a
partial region of an image of the image data; a control unit configured to control
decoding of the coded data acquired by the acquisition unit based on the decoding
30 load definition information acquired by the acquisition unit; and a decoding unit
configured to decode the coded data acquired by the acquisition unit under control of
SP353442WO00
3/125
the control unit.
[0009]
The partial region may be independently decodable.
[0010]
The decoding 5 load definition information may include information for
defining a magnitude of a load of a decoding process of the partial region according
to a level indicating a magnitude of a load of the decoding process.
[0011]
The decoding load definition information may include information for
10 defining a magnitude of a load of a decoding process of the partial region according
to information indicating a size of the partial region.
[0012]
The decoding load definition information may include information for
defining a magnitude of a load of a decoding process of the partial region according
15 to information indicating a length in a vertical direction and information indicating a
length in a horizontal direction of the partial region.
[0013]
The decoding load definition information may be included in supplemental
enhancement information (SEI) of an independently decodable partial region.
20 [0014]
The image data may include a plurality of layers, and the decoding load
definition information of the plurality of layers may be included in the SEI.
[0015]
The decoding load definition information may include information
25 indicating a size of the partial region serving as a reference, and a level indicating a
magnitude of a load of a decoding process of the partial region.
[0016]
The partial region may be a tile.
[0017]
30 The partial region may be a set of a plurality of tiles.
[0018]
SP353442WO00
4/125
The decoding load definition information may include information for
defining a maximum magnitude of a load of a decoding process among a plurality of
partial regions included in a picture of the image data according to a level indicating
a magnitude of a load of the decoding process.
5 [0019]
The decoding load definition information may include information for
defining a magnitude of a load common in a plurality of partial regions included in a
picture of the image data according to a level indicating a magnitude of a load of the
decoding process.
10 [0020]
When the plurality of partial regions included in the picture have an L shape,
a magnitude of the load may be defined for a rectangular region including the L
shape.
[0021]
15 The acquisition unit may further acquire information indicating whether the
decoding load definition information is set, and when the acquired information
indicates that the decoding load definition information is set, acquires the the
decoding load definition information.
[0022]
20 An aspect of the present technology is an image decoding method including:
acquiring coded data of image data and decoding load definition information for
defining a magnitude of a load of a decoding process of a partial region of an image
of the image data; controlling decoding of the acquired coded data based on the
acquired decoding load definition information; and decoding the acquired coded data
25 according to the controlling.
[0023]
In an aspect of the present technology, coded data of image data and
decoding load definition information for defining a magnitude of a load of a
decoding process of a partial region of an image of the image data are acquired;
30 decoding of the acquired coded data is controlled based on the acquired decoding
load definition information; and the acquired coded data is decoded according to the
SP353442WO00
5/125
controlling.
Advantageous Effects of Invention
[0024]
According to the present disclosure, 5 it is possible to code and decode an
image. In particular, it is possible to recognize performance necessary for decoding
more accurately.
Brief Description of Drawings
10 [0025]
[FIG. 1] FIG. 1 is a diagram describing an exemplary configuration of a coding unit.
[FIG. 2] FIG. 2 is a diagram illustrating an example of a layered image encoding
scheme.
[FIG. 3] FIG. 3 is a diagram for describing an example of spatial scalable coding.
15 [FIG. 4] FIG. 4 is a diagram for describing an example of temporal scalable coding.
[FIG. 5] FIG. 5 is a diagram for describing an example of scalable coding of a signalto-
noise ratio.
[FIG. 6] FIG. 6 is a diagram describing an exemplary application that performs partial
display.
20 [FIG. 7] FIG. 7 is a diagram describing another exemplary application that performs
partial display.
[FIG. 8] FIG. 8 is a diagram describing an exemplary method of defining a decoding
load to which the present technology is applied.
[FIG. 9] FIG. 9 is a diagram illustrating an extension example of MCTS SEI.
25 [FIG. 10] FIG. 10 is a diagram describing an overview of MCTS SEI.
[FIG. 11] FIG. 11 is a diagram describing an overview of MCTS SEI.
[FIG. 12] FIG. 12 is a diagram describing an overview of MCTS SEI.
[FIG. 13] FIG. 13 is a diagram describing an overview of MCTS SEI.
[FIG. 14] FIG. 14 is a diagram describing an overview of MCTS SEI.
30 [FIG. 15] FIG. 15 is a diagram describing an overview of MCTS SEI.
[FIG. 16] FIG. 16 is a diagram describing an overview of MCTS SEI.
SP353442WO00
6/125
[FIG. 17] FIG. 17 is a diagram describing an overview of MCTS SEI.
[FIG. 18] FIG. 18 is a diagram illustrating exemplary transmission of a syntax for
each ROI.
[FIG. 19] FIG. 19 is a diagram illustrating an extension example of MCTS SEI.
[FIG. 20] FIG. 20 is a diagram illustrating 5 an exemplary syntax of MCTS SEI.
[FIG. 21] FIG. 21 is a diagram illustrating an extension example of MCTS SEI.
[FIG. 22] FIG. 22 is a diagram describing a state of parameter mapping.
[FIG. 23] FIG. 23 is a diagram describing syntax elements.
[FIG. 24] FIG. 24 is a diagram illustrating an extension example of MCTS SEI.
10 [FIG. 25] FIG. 25 is a diagram describing a state of parameter mapping.
[FIG. 26] FIG. 26 is a diagram describing syntax elements.
[FIG. 27] FIG. 27 is a block diagram illustrating a main configuration example of an
image coding device.
[FIG. 28] FIG. 28 is a block diagram illustrating a main configuration example of a
15 base layer image coding unit.
[FIG. 29] FIG. 29 is a block diagram illustrating a main configuration example of an
enhancement layer image coding unit.
[FIG. 30] FIG. 30 is a block diagram illustrating a main configuration example of a
header information generating unit.
20 [FIG. 31] FIG. 31 is a flowchart describing an exemplary flow of image coding
processes.
[FIG. 32] FIG. 32 is a flowchart describing an exemplary flow of base layer coding
processes.
[FIG. 33] FIG. 33 is a flowchart describing an exemplary flow of enhancement layer
25 coding processes.
[FIG. 34] FIG. 34 is a flowchart describing an exemplary flow of header information
generating processes.
[FIG. 35] FIG. 35 is a block diagram illustrating a main configuration example of an
image decoding device.
30 [FIG. 36] FIG. 36 is a block diagram illustrating a main configuration example of a
base layer image decoding unit.
SP353442WO00
7/125
[FIG. 37] FIG. 37 is a block diagram illustrating a main configuration example of an
enhancement layer image decoding unit.
[FIG. 38] FIG. 38 is a block diagram illustrating an exemplary configuration of a
header information analyzing unit.
[FIG. 39] FIG. 39 is a flowchart describing an exemplary 5 flow of image decoding
processes.
[FIG. 40] FIG. 40 is a flowchart describing an exemplary flow of header information
analyzing processes.
[FIG. 41] FIG. 41 is a flowchart describing an exemplary flow of base layer decoding
10 processes.
[FIG. 42] FIG. 42 is a flowchart describing an exemplary flow of enhancement layer
decoding processes.
[FIG. 43] FIG. 43 is a diagram illustrating an example of a multi-view image coding
scheme.
15 [FIG. 44] FIG. 44 is a diagram illustrating a main configuration example of a multiview
image coding device to which the present technology is applied.
[FIG. 45] FIG. 45 is a diagram illustrating a main configuration example of a multiview
image decoding device to which the present technology is applied.
[FIG. 46] FIG. 46 is a block diagram illustrating an example of a main configuration
20 of a computer.
[FIG. 47] FIG. 47 is a block diagram illustrating an example of a schematic
configuration of a television device.
[FIG. 48] FIG. 48 is a block diagram illustrating an example of a schematic
configuration of a mobile telephone.
25 [FIG. 49] FIG. 49 is a block diagram illustrating an exemplary schematic
configuration of a recording and reproduction device.
[FIG. 50] FIG. 50 is a block diagram illustrating an exemplary schematic
configuration of an imaging device.
[FIG. 51] FIG. 51 is a block diagram illustrating an example of scalable coding use.
30 [FIG. 52] FIG. 52 is a block diagram illustrating another example of scalable coding
use.
SP353442WO00
8/125
[FIG. 53] FIG. 53 is a block diagram illustrating still another example of scalable
coding use.
[FIG. 54] FIG. 54 is a block diagram illustrating an example of a schematic
configuration of a video set.
[FIG. 55] FIG. 55 is a block diagram illustrating 5 an example of a schematic
configuration of a video processor.
[FIG. 56] FIG. 56 is a block diagram illustrating another example of the schematic
configuration of the video processor.
10 Description of Embodiments
[0026]
Hereinafter, aspects (hereinafter referred to as “embodiments”) for
implementing the present disclosure will be described. The descriptions will
proceed in the following order.
15 1. First embodiment (decoding load definition of partial image)
2. Second embodiment (image coding device)
3. Third embodiment (image decoding device)
4. Fourth embodiment (multi-view image coding and multi-view image decoding
device)
20 5. Fifth embodiment (computer)
6. Sixth embodiment (application example)
7. Seventh embodiment (application example of scalable coding)
8. Eighth embodiment (set, unit, module, and processor)
[0027]
25 <1. First embodiment>
In recent years, devices in which image information is digitally handled, and
in this case, in order to transmit and accumulate information with high efficiency,
image information-specific redundancy is used, and an image is compression-coded
30 employing a coding scheme in which an orthogonal transform such as a discrete
cosine transform and motion compensation are used for compression have
SP353442WO00
9/125
proliferated. As the coding scheme, Moving Picture Experts Group (MPEG) is
exemplified.
[0028]
In particular, MPEG2 (ISO/IEC 13818-2) is a standard that is defined as a
general-purpose image coding scheme, and generally 5 supports both an interlaced
scanning image and a progressive scanning image as well as a standard resolution
image and a high-definition image. For example, MPEG2 is currently being widely
used for a wide range of applications including professional applications and
consumer applications. When an MPEG2 compression scheme is used, for
10 example, an interlaced scanning image having a standard resolution of 720x480
pixels may be assigned a code amount (bit rate) of 4 to 8 Mbps. In addition, when
the MPEG2 compression scheme is used, for example, an interlaced scanning image
having a high resolution of 1920x1088 pixels may be assigned a code amount (bit
rate) of 18 to 22 Mbps. Therefore, it is possible to implement a high compression
15 rate and good image quality.
[0029]
MPEG2 is mainly designed for high image quality coding suitable for
broadcast, but does not correspond to a lower code amount (bit rate) than that of
MPEG1, that is, a coding scheme of a higher compression rate. With the
20 proliferation of mobile terminals, it is assumed that needs for such a coding scheme
will increase in the future. Accordingly, MPEG4 coding schemes have been
standardized. A standard of image coding schemes was approved as an
international standard ISO/IEC14496-2 in December 1998.
[0030]
25 Further, in recent years, for the initial purpose of image coding for
television conferencing, a standard called H.26L (ITU-T (International
Telecommunication Union Telecommunication Standardization Sector) Q6/16
VCEG (Video Coding Expert Group)) has been standardized. It is known that
H.26L requests a greater amount of computation for coding and decoding than
30 coding schemes of the related art such as MPEG2 or MPEG4, but has a higher
coding efficiency. In addition, currently, as a part of MPEG4 activities, based on
SP353442WO00
10/125
H.26L, standardization in which functions that are not supported in H.26L are also
incorporated to implement higher coding efficiency is being performed as Joint
Model of Enhanced-Compression Video Coding.
[0031]
As schedules of standardization, H.264 5 and MPEG-4 Part10 (Advanced
Video Coding, hereinafter referred to as “AVC”) became international standards in
March 2003.
[0032]
Further, as extensions of H.264/AVC, standardization of Fidelity Range
10 Extension (FRExt) including coding tools necessary for professional use such as
RGB, 4:2:2, or 4:4:4, and 8x8 DCT or a quantization matrix defined in MPEG-2 was
completed in February 2005. Therefore, when H.264/AVC is used, the coding
scheme is also able to appropriately represent film noise included in a movie and is
used for a wide range of applications such as a Blu-Ray Disc (trademark).
15 [0033]
However, in recent years, needs for higher compression rate coding
including compression of an image of about 4000x2000 pixels, four times that of a
high definition image, or delivery of a high definition image in an environment
having a limited transmission capacity such as the Internet, are increasing.
20 Therefore, in previously described VCEG under ITU-T, study for increasing coding
efficiency continues.
[0034]
Therefore, currently, in order to further increase coding efficiency over that
of AVC, Joint Collaboration Team-Video Coding (JCTVC), which is a joint
25 standardization organization of ITU-T and International Organization for
Standardization/International Electrotechnical Commission (ISO/IEC), proceeding
with a standardization of a coding scheme called High Efficiency Video Coding
(HEVC). As a standard of HEVC, a committee draft, which is a draft specification,
has been issued in January 2013 (for example, refer to Non-Patent Literature 1).
30 [0035]
SP353442WO00
11/125
Hereinafter, the present technology will be described with application
examples of image coding and decoding of a High Efficiency Video Coding (HEVC)
scheme.
[0036]
5
In the Advanced Video Coding (AVC) scheme, a layered structure of
macroblocks and sub-macroblocks is defined. However, a macroblock of 16x16
pixels is not optimal for a large image frame provided in the next generation coding
scheme Ultra High Definition (UHD, 4000 pixelsx2000 pixels).
10 [0037]
On the other hand, in the HEVC scheme, as illustrated in FIG. 1, a coding
unit (CU) is defined.
[0038]
The CU is also called a coding tree block (CTB) and is a partial region of an
15 image of a picture unit, which similarly serves as the macroblock in the AVC
scheme. The latter is fixed to a size of 16x16 pixels. On the other hand, the
former has a size that is not fixed, but is designated in image compression
information in respective sequences.
[0039]
20 For example, in the sequence parameter set (SPS) included in coded data to
be output, a maximum size (largest coding unit (LCU)) and a minimum size
(smallest coding unit (SCU)) of the CU are defined.
[0040]
In each LCU, in a range equal to or greater than a size of the SCU, when
25 split-flag=1 is set, the unit may be divided into CUs having a smaller size. In an
example of FIG. 1, the LCU has a size of 128 and a maximum level depth of 5.
When a value of split_flag is set to “1,” a CU having a size of 2Nx2N is divided into
the next lowest level of CUs having a size of NxN.
[0041]
30 Further, the CU is divided into a prediction unit (PU) that is a region (a
partial region of an image of a picture unit) serving as a processing unit of intra or
SP353442WO00
12/125
inter prediction, and is divided into a transform unit (TU)) that is a region (a partial
region of an image of a picture unit) serving as a processing unit of an orthogonal
transform. Currently, in the HEVC scheme, it is possible to use 16x16 and 32x32
orthogonal transform in addition to 4x4 and 8x8.
5 [0042]
In a coding scheme in which the CU is defined and various processes are
performed in units of CUs as in the HEVC scheme described above, the macroblock
in the AVC scheme may be considered to correspond to the LCU and the block (subblock)
may be considered to correspond to the CU. In addition, a motion
10 compensation block in the AVC scheme may be considered to correspond to the PU.
However, since the CU has a layered structure, the LCU of the topmost level has a
size that is generally set to be greater than a macroblock of the AVC scheme, for
example, 128x128 pixels.
[0043]
15 Accordingly, hereinafter, the LCU may include the macroblock in the AVC
scheme, and the CU may include the block (sub-block) in the AVC scheme. That
is, the term “block” used in the following description refers to any partial region in
the picture and has a size, a shape, a characteristic and the like that are not limited.
In other words, the “block” includes any region (processing unit), for example, a TU,
20 a PU, an SCU, a CU, an LCU, a sub-block, a macroblock, or a slice. It is needless
to say that a partial region (processing unit) other than these is included. When
there is a need to limit a size, a processing unit or the like, it will be appropriately
described.
[0044]
25 In addition, in this specification, a coding tree unit (CTU) is a unit including
a parameter when processing is performed in the coding tree block (CTB) of the
LCU (a maximum number of the CU) and an LCU base (level) thereof. In addition,
the coding unit (CU) of the CTU is a unit including a parameter when processing is
performed in a coding block (CB) and a CU base (level) thereof.
30 [0045]
SP353442WO00
13/125
Meanwhile, in the AVC and HEVC coding schemes, in order to achieve
higher coding efficiency, it is important to select an appropriate prediction mode.
[0046]
As an example of such a selection scheme, a method implemented in
reference software (disclosed 5 in http://iphome.hhi.de/suehring/tml/index.htm) of
H.264/MPEG-4 AVC called Joint Model (JM) may be exemplified.
[0047]
In JM, it is possible to select a method of determining two modes, a high
complexity mode and a low complexity mode, to be described below. In both, a
10 cost function value for each prediction mode Mode is calculated, and a prediction
mode minimizing the value is selected as an optimal mode for the block or the
macroblock.
[0048]
A cost function in the high complexity mode is represented as the following
15 Equation (1).
[0049]
[Math. 1]
[0050]
20 Here, denotes an entire set of candidate modes for coding the block or the
macroblock, and D denotes difference energy between a decoded image and an input
image when coding is performed in the prediction mode. λ denotes a Lagrange
undetermined multiplier provided as a function of a quantization parameter. R
denotes a total amount of codes when coding is performed in the mode including an
25 orthogonal transform coefficient.
[0051]
That is, when coding is performed in the high complexity mode, in order to
calculate the parameters, D and R, it is necessary to perform a provisional encoding
process once in all candidate modes. Therefore, a higher amount of computation is
30 necessary.
[0052]
SP353442WO00
14/125
A cost function in the low complexity mode is represented as the following
Equation (2).
[0053]
[Math. 2]
5
[0054]
Here, D denotes difference energy between a prediction image and an input
image unlike the high complexity mode. QP2Quant (QP) is provided as a function
of a quantization parameter QP. HeaderBit denotes a code amount of information
10 belonging to a header in a motion vector or a mode having no orthogonal transform
coefficient.
[0055]
That is, in the low complexity mode, it is necessary to perform a prediction
process in respective candidate modes, but is not necessary for the decoded image.
15 Therefore, it is not necessary to perform a coding process. For this reason, a lower
amount of computation than that in the high complexity mode may be implemented.
[0056]
Incidentally, the image coding schemes such as MPEG2 and AVC described
20 above have a scalability function. Scalable coding (layered coding) refers to a
scheme in which an image is divided into a plurality of layers (layered), and coding
is performed for each layer. FIG. 2 is a diagram illustrating an example of a layered
image coding scheme.
[0057]
25 As illustrated in FIG. 2, in layering of the image, using a predetermined
parameter having a scalability function as a reference, one image is divided into a
plurality of levels (layers). That is, the layered image (hierarchical image) includes
an image of a plurality of layers whose predetermined parameter values are different
from each other. The plurality of layers of the layered image include a base layer in
30 which coding and decoding are performed using only an image of its own layer
without using an image of another layer, and a non-base layer (also referred to as an
SP353442WO00
15/125
“enhancement layer”) in which coding and decoding are performed using an image
of another layer. The non-base layer may use an image of the base layer or use an
image of another non-base layer.
[0058]
In general, the non-base layer includes data (difference 5 data) of a difference
image between its own image and an image of another layer so that redundancy is
reduced. For example, when one image is divided into two layers, the base layer
and the non-base layer (also referred to as an “enhancement layer”), an image having
lower quality than an original image may be obtained when only data of the base
10 layer is used. Therefore, when data of the base layer and data of the non-base layer
are synthesized, the original image (that is, a high quality image) may be obtained.
[0059]
When the image is layered in this manner, it is possible to easily obtain an
image of a variety of levels of quality according to circumstances. For example, in
15 terminals having a low processing capacity such as a mobile telephone, image
compression information of only the base layer is transmitted, and a moving image
having a low spatial and temporal resolution or having low image quality is played
and in terminals having a high processing capacity such as a television or a personal
computer, image compression information of the enhancement layer in addition to
20 the base layer is transmitted, and a moving image having a high spatial and temporal
resolution or high image quality is played, so that a transcoding process is not
performed and image compression information may be transmitted from a server
according to a capability of a terminal or a network.
[0060]
25
In such layered image coding and layered image decoding (scalable coding
and scalable decoding), a parameter having a scalability function is arbitrary. For
example, a spatial resolution illustrated in FIG. 3 may be used as a parameter thereof
(spatial scalability). In the spatial scalability, a resolution of an image is different
30 for each layer. That is, as illustrated in FIG. 3, each picture is divided into two
layers, and the base layer having a lower spatial resolution than the original image
SP353442WO00
16/125
and the enhancement layer in which synthesizing with the image of the base layer is
performed and the original image (original spatial resolution) may be obtained. It is
needless to say that the number of levels is only an example, and the image may be
layered into any number of levels.
5 [0061]
In addition, as a parameter enabling such a scalable property, another
example, for example, a temporal resolution may also be applied (temporal
scalability) as illustrated in FIG. 4. In the temporal scalability, a frame rate is
different for each layer. That is, in this case, as illustrated in FIG. 4, the image is
10 divided into layers having different frame rates, and when a layer having a high
frame rate is added to a layer having a low frame rate, a moving image having a
higher frame rate may be obtained, and when all layers are added, an original moving
image (original frame rate) may be obtained. The number of levels is only an
example, and the image may be layered into any number of levels.
15 [0062]
Further, as a parameter enabling such a scalable property, in another
example, for example, as illustrated in FIG. 5, a signal to noise ratio (SNR) may be
applied (SNR scalability). In the SNR scalability, the SNR is different for each
layer. That is, in this case, as illustrated in FIG. 5, each picture is layered into two
20 levels, a base layer having a lower SNR than the original image and an enhancement
layer that may be synthesized with an image of the base layer to obtain the original
image (original SNR). That is, in base layer image compression information,
information on an image of a low PSNR is transmitted, and enhancement layer image
compression information is added thereto. Therefore, it is possible to reconstruct a
25 high PSNR image. It is needless to say that the number of levels is only an example
and the image may be layered into any number of levels.
[0063]
It is needless to say that the parameter enabling such a scalable property
may be a parameter other than the above-described example. For example, in bit30
depth scalability, the base layer is an image of 8 bits, the enhancement layer is added
thereto, and thus an image of 10 bits is obtained.
SP353442WO00
17/125
[0064]
In addition, in chroma scalability, the base layer is a component image of a
4:2:0 format, the enhancement layer is added thereto, and thus a component image of
a 4: 2: 2 format is obtained.
5 [0065]
In HEVC, it is possible to decode only a region whose decoding is
necessary by an application using a tile structure. In order to indicate the fact that a
tile region is independently decodable, second and later versions (including MV10
HEVC, SHVC, Range Ext. and the like) of HEVC are supported by motionconstrained
tile sets SEI.
[0066]
Application examples to which the present technology is applied will be
15 described.
[0067]
In a system configured to deliver an image from a server to a terminal, for
example, as exemplified in FIG. 6, there is an application in which a single screen is
divided into a plurality of screens and delivery is performed while a display region is
20 switched. In addition, for example, as exemplified in FIG. 7, there is an application
in which a partial region to be displayed (delivered) is selected in order to select an
aspect ratio or a resolution of an image.
[0068]
In the application of FIG. 6, partial images are segmented from an entire
25 image using a tile as a unit in coding and decoding of the image, and delivered to
terminals. Positions of the segmented partial images in the entire image may be
designated by, for example, a user of the terminal. Therefore, in the terminal, it is
possible to display a partial image of a desired position of the entire image. For
example, in a service such as sports broadcast, in a wide angle image that is provided
30 from a server or the like and obtained by capturing an image of an entire venue, an
entire field or the like, focusing on a desired part (for example, a favorite player, a
SP353442WO00
18/125
coach, in front of a goal, a bench, and an audience seat) of the user, the partial image
may be segmented and downloaded (or streamed), and displayed on the terminal.
That is, the user of the terminal can focus on the desired part of the entire image.
[0069]
In the application of FIG. 7, by simply selecting 5 a tile, a resolution of a
display image may be set to HD or a cinema size.
[0070]
However, as information on a level, which serves as a reference for
determining whether a decoder can decode a stream, and a buffer capacity, only a
10 value of the entire stream or a value of a layer unit is defined.
[0071]
Therefore, even in an application that decodes only a part of an entire image,
determination of whether decoding is possible is performed by assuming a load when
an entire screen is decoded. Accordingly, there is concern of an unnecessarily high
15 level decoder being necessary. In addition, there is concern of applications to be
delivered being unnecessarily limited accordingly.
[0072]
Therefore, decoding load definition information for defining a magnitude of
a load of a decoding process of an independently decodable partial region of an
20 image of image data to be coded is set, and the decoding load definition information
is transmitted. For example, the decoding load definition information is transmitted
from a coding side to a decoding side together with coded data of image data.
[0073]
In this manner, the decoder may recognize performance necessary for
25 decoding the partial region according to the decoding load definition information,
and determine whether decoding is possible. That is, it is possible to recognize
performance necessary for decoding more accurately. Therefore, it is possible to
select the decoder having appropriate performance for the image data. Therefore, it
is possible to suppress a situation of applying the decoder having an unnecessarily
30 high level with respect to a decoding load of image data from occurring. In
addition, it is possible to suppress applications to be delivered from being
SP353442WO00
19/125
unnecessarily limited accordingly.
[0074]
The decoding load definition information is defined according to, for
example, FIG. 8. For example, as exemplified in 5 A of FIG. 8, the decoding load
definition information for defining a magnitude of a load of a decoding process of
the partial region may be set for an independently decodable partial region of a single
layer. In the example of A of FIG. 8, level 4.0 is set for decoding an entire image of
a single layer, and level 2.0 is set for decoding an independently decodable tile
10 (partial region) of the image.
[0075]
In addition, for example, as exemplified in B of FIG. 8, the decoding load
definition information may be set for an independently decodable partial region of
each layer of an image including a plurality of layers. In the example of B of FIG.
15 8, level 5.0 is set for decoding an image of all layers, and level 4.0 is set for decoding
an entire image of the base layer (layer 0). Further, level 2.0 is set for decoding an
independently decodable tile (partial region) of the image of the base layer (layer 0).
Further, level 4.0 is set for decoding (that is, decoding of a tile of the base layer
(layer 0) to be referred to and a tile of the enhancement layer (layer 1) referring
20 thereto) a tile that refers to only an independently decodable tile of an image of the
base layer (layer 0) of an image of an enhancement layer (layer 1).
[0076]
Further, for example, as exemplified in C of FIG. 8, the decoding load
definition information may be set for the entire image of the layer referring to only
25 the independently decodable partial region and the referred partial region. That is, a
side that refers to the independently decodable tile (partial region) may be the entire
image rather than the partial region. In the example of C of FIG. 8, basically, the
same level as in the example of B of FIG. 8 is set. However, in the example of B of
FIG. 8, the level is set for decoding a tile referring to only an independently
30 decodable tile of the image of the base layer (layer 0) in the image of the
enhancement layer (layer 1). However, in the example of C of FIG. 8, alternatively,
SP353442WO00
20/125
level 4.0 is set for decoding (that is, decoding the tile of the base layer (layer 0) to be
referred to and the entire image of the enhancement layer (layer 1) referring thereto)
the entire image of the enhancement layer (layer 1) referring to only the
independently decodable tile of the image of the base layer (layer 0).
5 [0077]
Also, in this case, in order to identify a partial region (tile) (a position
thereof) referenced by the entire image of the enhancement layer (layer 1), position
information of the tile of the base layer (layer 0) serving as a reference source may
be associated with (mapped with) the entire image of the enhancement layer (layer 1)
10 serving as a reference source. In the example of C of FIG. 8, coordinates of the
same position in the tile of the base layer (layer 0) are mapped with upper-left corner
coordinates in the entire image of the enhancement layer (layer 1).
[0078]
15 Note that, a parameter defined by a level includes a maximum pixel number
(MaxLumaPs), a maximum buffer capacity (MaxCPB Size), a maximum number of
pixels (MaxLumaSr) of an image per second, a maximum bit rate (MaxBR) of an
image or the like.
[0079]
20
Definition of a magnitude of a load necessary for decoding is performed by
extending, for example, motion constrained tile set supplemental enhancement
information (MCTS SEI).
[0080]
25 For example, as in a described syntax A of FIG. 9, in MCTS SEI, as the
decoding load definition information for defining a magnitude of a load of a
decoding process of the independently decodable partial region, a level (mcts_level
_idc[i]) indicating a magnitude of a load of a decoding process of the partial region
may be set. Here, “i” denotes a set (also referred to as a “tile set”), which is the
30 partial region composed of a single tile or a plurality of tiles. That is, in the
example of A of FIG. 9, a value of level information (mcts_level _idc) necessary for
SP353442WO00
21/125
decoding is set for each set. In this case, semantics may be as described in, for
example, B of FIG. 9.
[0081]
In MCTS SEI, the independently decodable partial region is set for each
rectangular set. For example, when an upper-left shaded 5 part of A of FIG. 10 is the
independently decodable partial region, the partial region is set for each set, in
MCTS SEI, as illustrated in B of FIG. 10. Also, as exemplified in C of FIG. 10, a
tile included in the set may overlap another set. The number of pixels of the partial
region may be calculated from the number of pixels of each set, for example, as
10 exemplified in D of FIG. 10.
[0082]
In addition, for example, as in a described syntax A of FIG. 11, in MCTS
SEI, as the decoding load definition information for defining a magnitude of a load
of a decoding process of the independently decodable partial region of a plurality of
15 layers, a level (mcts_level _idc[i][j]) indicating a magnitude of a load of a decoding
process of the partial region of each layer may be set. Here, “i” denotes a set and
“j” denotes a layer. That is, in the example of A of FIG. 11, a value of level
information (mcts_level _idc) necessary for decoding is set for each set and for each
layer. In this case, semantics may be as described in, for example, B of FIG. 11.
20 [0083]
Further, for example, as in a described syntax A of FIG. 12, in MCTS SEI,
as the decoding load definition information for defining a magnitude of a load of a
decoding process of the independently decodable partial region, information
(maxLumaP S_in_set[i]) indicating a size of the partial region may be set. Here, “i”
25 denotes a set. That is, in the example of A of FIG. 12, a value of information
(maxLumaP S_in_set) indicating a size of the set (partial region) is set for each set.
In this case, semantics may be as described in, for example, B of FIG. 12.
[0084]
In addition, for example, as in a described syntax A of FIG. 13, in MCTS
30 SEI, as the decoding load definition information for defining a magnitude of a load
of a decoding process of the independently decodable partial region, information
SP353442WO00
22/125
(mcts_height_in_luma_samples[i]) indicating a length in a vertical direction and
information (mcts_width_in_luma_samples[i]) indicating a length in a horizontal
direction of the partial region may be set. Here, “i” denotes a set. That is, in the
example of A of FIG. 13, a value of information (mcts_height_in_luma_samples)
indicating a length in a vertical 5 direction and a value of information
(mcts_width_in_luma_samples) indicating a length in a horizontal direction of the
set (partial region) are set for each set. In this case, semantics may be as described
in, for example, B of FIG. 13.
[0085]
10 Further, for example, as in a syntax described in FIG. 14, in MCTS SEI, as
the decoding load definition information for defining a magnitude of a load of a
decoding process of the independently decodable partial region, a parameter
(mcts_hrd_parameters ()) of a virtual reference decoder configured to decode the
partial region may be set.
15 [0086]
In this case, for example, as in a described syntax A of FIG. 15, as the
parameter (mcts_hrd_parameters ()) of the virtual reference decoder, a maximum
input bit rate (mcts_bit_rate_value_minus1) and a buffer capacity
(mcts_cpb_size_value_minus1) of the virtual reference decoder may be set. In this
20 case, semantics may be as described in, for example, B of FIG. 15.
[0087]
In addition, for example, as illustrated in A of FIG. 16, in addition to the
extension of MCTS SEI described above, in the sequence parameter set (SPS),
information (mcts_present_flag) indicating whether the decoding load definition
25 information described above is set in MCTS SEI may be set. In this case, semantics
may be as described in, for example, B of FIG. 16.
[0088]
Further, for example, as illustrated in A of FIG. 17, in the sequence
parameter set (SPS), instead of the information indicating whether the decoding load
30 definition information is set, the same decoding load definition information as the
decoding load definition information set in MCTS SEI may be set. In this case,
SP353442WO00
23/125
semantics may be as described in, for example, B of FIG. 17.
[0089]
Also, information set in the sequence parameter set (SPS) may be set in a
video parameter set (VPS), instead of the sequence parameter set (SPS).
5 [0090]
It is needless to say that a method of setting the decoding load definition
information is arbitrary, and is not limited to the above-described example. In
addition, the above-described plurality of methods may be combined. Further, the
above-described method may be combined with other methods.
10 [0091]
As described above, when the decoding load definition information for
defining a magnitude of a load of a decoding process of the independently decodable
partial region is set, it is possible to recognize performance necessary for decoding
more accurately based on the decoding load definition information. In addition,
15 when the decoding load definition information is transmitted to the decoding side, it
is possible to recognize performance necessary for decoding more accurately even on
the decoding side.
[0092]
20 For example, in a use case of DASH as illustrated in FIG. 6, it is preferable
that a certain number of tile regions be moved and reproduced. However, when all
partial images (a combination of tiles) to be reproduced are registered as a tile set
(tile_set), and the decoding load definition information is set therefor, there is a
possibility of an amount of information being increased when the number of tile sets
25 is great.
[0093]
For example, as exemplified in FIG. 18, when a partial image (tile set) of
2x2 tiles is reproduced (displayed), if a position of the partial image is moved (when
a part to be displayed in the entire image is changed), all partial images (tile sets)
30 during a movement thereof are displayed (reproduced). The number of such partial
images may become enormous, as shown in the equation on the right side of the
SP353442WO00
24/125
drawing. Accordingly, when the decoding load definition information is set for all
such partial images, there is a possibility of an amount of information being
impractically increased. In addition, there is a possibility of redundancy of
decoding load information of each tile set in this case becoming extremely high.
That is, there is a possibility of an amount 5 of unnecessary information being
increased.
[0094]
Here, for example, MCTS SEI is extended, and as the decoding load
information, information indicating a size of a partial image serving as a reference
10 and a level indicating a magnitude of a load of a decoding process of the partial
image are set. That is, information from which it is possible to estimate a
magnitude of a load of a decoding process according to the size of the partial image
is stored in MCTS SEI.
[0095]
15 Therefore, for example, when such information is referred to (when a size of
a region to be decoded is compared with a size of the partial image serving as a
reference thereof), it is possible recognize to a magnitude of a load of a decoding
process of the region more accurately.
[0096]
20 Here, the size of the partial image serving as the reference may be indicated
by any information, and may be indicated, for example, in units of tiles obtained by
uniformly dividing the entire image. In addition, the number of sizes serving as a
reference thereof may be arbitrary, but it is preferable that the number of sizes be
plural in order for a magnitude of a load of a decoding process to be recognized more
25 accurately. In addition, a magnitude of a load of a decoding process may be
indicated by any information, and may be indicated by, for example, level
information (level).
[0097]
Incidentally, in an existing ROI, it is assumed that an application can
30 designate any position, or in other words, it is absolutely necessary that some region
(ROI) be defined.
SP353442WO00
25/125
[0098]
However, an application such as tiled streaming of DASH has a concept of
segmenting and displaying a region selected by the user (also including switching to
a stream having a different resolution). That is, since the user can arbitrarily
determine a region to be selected, it is 5 assumed that all tiles are independently
decodable tiles and are further uniformly divided and the number of tiles to be
selected differing according to a capability (level) of a device to be reproduced is
assumed to correspond to operation of the most common service.
[0099]
10 Therefore, assumption of an application for the existing ROI and
assumption of an application for tiled streaming of DASH have slightly different
directions.
[0100]
Accordingly, a mode of an application is introduced such that a level of a
15 region (tile) unit decoded by the application may be defined while both of the
assumptions are satisfied in one SEI, and information to be defined may be changed
for each mode for extension.
[0101]
For example, a concept of a mode is defined for each application to change
20 information to be defined as necessary. For example, a DASH mode is set.
Therefore, it is defined that the DASH mode == “uniform division and independence
of all tiles.” That is, in the DASH mode, the screen is assumed to be uniformly
divided into tiles (uniform_spacing_flag=1@PPS). In addition, individual tiles are
assumed to be independently decodable.
25 [0102]
Therefore, in the DASH mode, the number of tiles to be decoded at a time
and level information corresponding thereto are described (defined). For example,
in FIG. 19, as the number of tiles to be decoded at a time, three cases, 4, 12, and 30,
are set. Level information (decoding load information) is set for each of the three
30 cases.
[0103]
SP353442WO00
26/125
In this manner, it is possible to set the decoding load information more
appropriately for both of the applications described above.
[0104]
FIG. 20 illustrates an extension example (an exemplary syntax) of MCTS
SEI in this case. In the example of FIG. 20, a 5 mode (mcts_mode) of an application
is set in a 2nd row from the top. Therefore, as described in a 4th row from the top,
when the mode is a mode (mode for current ROI application) of an application for
the existing ROI (mcts_mode==0), the decoding load information is set for each
independently decodable partial image, similarly to each example described above.
10 For example, level information (mcts_level _idc[i]) necessary for decoding is set in a
13th row from the top.
[0105]
In addition, when the mode of the application is the DASH mode (mode for
DASH application) (mcts_mode==1), the number of tiles of a region to be decoded
15 at a time and level information corresponding thereto are set. For example, in the
17th row from the top, identification information (mcts_id[i]) of the region is set.
In the next row, information (num_of_tiles_minus1[i]) indicating the number of tiles
included in the region indicated by the identification information is set. Further, in
the next row, level information (mcts_level _idc[i]) necessary for decoding the
20 region is set.
[0106]
Here, “i” denotes a set. That is, in the example of FIG. 20, a value of level
information (mcts_level _idc) necessary for decoding, a value of identification
information (mcts_id), and a value of information (num_of_tiles_minus1) indicating
25 the number of tiles included in a region indicated by the identification information
are set for each set.
[0107]
When the decoding load information is set as described above, it is possible
to recognize performance necessary for decoding more accurately according to a size
30 (the number of tiles) of the region to be decoded based on the decoding load
information. In addition, when the decoding load definition information is
SP353442WO00
27/125
transmitted to the decoding side, it is possible to recognize performance necessary
for decoding more accurately according to a size (the number of tiles) of the region
to be decoded, even on the decoding side.
[0108]
A rectangular tile set has been described above. In this case, for example,
when the MCTS (independently decodable tile group (partial region)) is assumed to
have an “L” shape, it is necessary to define two sets, a tile set in which a vertical
direction indicating a vertical line part of the letter “L” is defined as a longitudinal
10 direction, and a tile set in which a horizontal direction indicating a horizontal line
part of the letter “L” is defined as a longitudinal direction.
[0109]
In addition, it has been described above that, as the decoding load definition
information, a value of information (mcts_level _idc[i]) indicating a level (level
15 necessary for decoding) indicating a magnitude of a load of a decoding process of the
partial region is set for each rectangular set. That is, in this case, in the L-shaped
partial region, it is necessary to set two pieces of information (mcts_level _idc[i])
indicating a level necessary for decoding. It is needless to say that it is possible to
deal with a case in which levels necessary for decoding sets are different in this
20 manner. However, when levels necessary for decoding sets are the same, it
becomes redundant and there is a possibility of coding efficiency decreasing.
[0110]
Therefore, one level may be set for the independently decodable partial
region rather than for each set. An independently decodable set may be the
25 independently decodable partial region. That is, a common level may be defined
for a plurality of independently decodable partial regions. An exemplary syntax in
this case is described in A of FIG. 21. In addition, an example of semantics in this
case is described in B of FIG. 21.
[0111]
30 In the example of A of FIG. 21, in a 3rd row from the top, information
(each_tile_one_tile_set_flag) indicating whether all tiles form an independently
SP353442WO00
28/125
decodable set is set. In a 4th row from the top, information (mcts_level
_idc_present_flag) indicating whether information (mcts_level _idc) indicating a
level necessary for decoding is included is set in motion-constrained tile sets defined
in an SEI message.
5 [0112]
Therefore, in a 5th row from the top, when it is determined that all tiles do
not form a uniquely decodable set (!each_tile_one_tile_set_flag), in a loop of a 7th
row to a 16th row from the top, settings are performed for each set, and a level
(mcts_level _idc) necessary for decoding is set in an 18th row from the top other than
10 the loop.
[0113]
That is, as exemplified in FIG. 22, when all tiles are independently
decodable, a set in which “0”th identification information (mcts_id[0]) is assigned
and a set in which “1”st identification information (mcts_id[1]) is assigned are
15 adjacent to each other, and levels necessary for decoding sets are the same, level
information (mcts_level _idc[i]) is not set for each set, but level information
(mcts_level _idc) common to both sets may be set.
[0114]
In this manner, only one (common) piece of level information may be set
20 for a plurality of independently decodable partial regions (for example, partial
regions necessary to be represented by a plurality of sets). Therefore, it is possible
to reduce redundancy and increase coding efficiency.
[0115]
Also, in HEVC described in Non-Patent Literature 1, it is assumed that the
25 level is defined in an entire image (picture) unit, and definition of a parameter of the
level is also performed in the entire image (picture) unit. Therefore, when the level
is defined in a partial region unit described above, definition of the parameter of the
level for the partial region is also performed, and the definition may be assigned
(parameter mapping is performed) to definition of the level of the entire image
30 (picture) unit.
[0116]
SP353442WO00
29/125
For example, when the independently decodable partial region has a
rectangular shape, the parameter of the level for the partial region is defined. The
definition of the level of the entire image (picture) unit may be replaced with the
definition. In addition, for example, when the independently decodable partial
region has an “L” shape, a rectangle including the 5 L-shaped partial region is set, the
parameter of the level for the rectangular region is defined, and definition of the level
of the entire image (picture) unit may be replaced with the definition.
[0117]
For example, in FIG. 22, when the independently decodable partial region
10 includes the set of identification information (mcts_id[0]) and the set of identification
information (mcts_id[1]), a rectangular region including both of the sets is set, and is
assigned to a picture (a unit of definition of the parameter of the level). That is, a
size W of the rectangular region in a horizontal direction is estimated as
pic_width_in_luma_samples, a size H of the rectangular region in a vertical direction
15 is estimated as pic_height_in_luma_samples, a size WxH of the rectangular region is
estimated as PicSizeInSamplesY, and the parameter of the level for the rectangular
region may be defined.
[0118]
In this manner, a value of the parameter defined for the rectangular region
20 may be mapped with the parameter defined in the picture unit. Therefore, when
such parameter mapping is performed, it is possible to employ more appropriate
definition for the independently decodable partial region.
[0119]
Incidentally, when all tiles are independently decodable, it is possible to set
25 a level (mcts_level _idc[i]) necessary for decoding the partial region according to a
size of the partial region. In this case, the size of the partial region may be
represented by the number of tiles (the number of rows) in a vertical direction and
the number of tiles (the number of columns) in a horizontal direction of the partial
region. Here, “i” denotes the number of correspondence relations between the size
30 and the level of the partial region.
[0120]
SP353442WO00
30/125
In the example of A of FIG. 21, in a loop of a 20th row to a 28th row from
the top, the size and the level of the partial region are associated. In the loop,
information (num_mc_tile_columns_minus1[i]) indicating the number of tiles of the
partial region in a vertical direction is set (a 24th row from the top), and information
(num_mc_tile_rows_minus1[i]) indicating the number 5 of tiles of the partial region in
a horizontal direction is set (a 25th row from the top).
[0121]
For example, in FIG. 23, in a “0”th correspondence relation in which
identification information (mcts_level _id[0]) is set, the partial region of 2x2 tiles is
10 associated with a level (meta_level _idc[0]). That is, in the correspondence
relation, a value of information (num_mc_tile_columns_minus1[0]) indicating the
number of tiles of the partial region in a vertical direction is set to “1,” and
information (num_mc_tile_rows_minus1[0]) indicating the number of tiles of the
partial region in a horizontal direction is set to “1.” When such information is set,
15 in addition to the fact that the number of tiles of the partial region corresponding to
the level (meta_level _idc[0]) is 4, a shape (a rectangle of two vertical tiles x two
horizontal tiles) of the partial region is also shown.
[0122]
For example, in FIG. 23, in a “1”st correspondence relation in which
20 identification information (mcts_level _id[1]) is set, the partial region of 4x4 tiles is
associated with a level (meta_level _idc[1]). That is, in the correspondence
relation, a value of information (num_mc_tile_columns_minus1[1]) indicating the
number of tiles of the partial region in a vertical direction is set to “3,” and
information (num_mc_tile_rows_minus1[1]) indicating the number of tiles of the
25 partial region in a horizontal direction is set to “3.” When such information is set,
in addition to the fact that the number of tiles of the partial region corresponding to
the level (meta_level _idc[1]) is 16, a shape (a rectangle of four vertical tiles x four
horizontal tiles) of the partial region is also shown.
[0123]
30 In this manner, it is possible to increase convenience of information for
associating the partial region with the level. For example, a terminal configured to
SP353442WO00
31/125
acquire and display a partial image thereof determines whether the partial region is
horizontally long based on the information. When the partial region is horizontally
long, it is possible to adjust an aspect ratio of a display image such as display by
inserting a black band on the top and bottom more easily.
5 [0124]
When all tiles are independently decodable, a maximum value (max_level
_idc) of a level of a tile unit in the picture may be set. That is, one tile is set as one
set, a maximum value in the picture of the level set for each set may be set. An
10 exemplary syntax in this case is described in A of FIG. 24. In addition, an example
of semantics in this case is described in B of FIG. 24.
[0125]
In the example of A of FIG. 24, in a 21st row from the top, a maximum
value (max_level _idc) of the level in the picture is set. As exemplified in FIG. 26,
15 the level is a level that is set for the set composed of one tile. For example, in FIG.
26, a maximum value of the level set for each set of 60 sets (60 tiles) in total
including 6 vertical sets (6 tiles) x 10 horizontal sets (10 tiles) in the picture is set as
max_level _idc.
[0126]
20 The number of tiles segmented from the entire image is determined by the
application. Although the number of tiles corresponding to each level is not
entirely defined, the application can sufficiently accurately determine the number of
tiles that can be segmented (can perform a process such as decoding) based on the
maximum value (max_level _idc) of the level.
25 [0127]
That is, instead of entirely defining the number of tiles corresponding to
each level, when only a maximum value (max_level _idc) of the level set for each set
(that is, for each tile) composed of one tile is set in the picture, the application may
control the number of segmented tiles based on the setting so that the process such as
30 decoding does not fail.
[0128]
SP353442WO00
32/125
Therefore, in this manner, compared to when the number of tiles
corresponding to each level is entirely defined, a syntax of all tiles can be simplified
and a load of the process can be reduced. In addition, compared to when the
number of tiles corresponding to each level is entirely defined, an amount of
information to be transmitted can be reduced and coding 5 efficiency can increase.
[0129]
Also, as exemplified in A of FIG. 24, in a 17th row from the top, a level
(mcts_level _idc[i]) for each partial region is set. That is, as exemplified in A of
FIG. 25, when there are two independently decodable sets, a set to which
10 identification information (mcts_id[0]) is assigned and a set to which identification
information (mcts_id[1]) is assigned, it is possible to set a level (mcts_level _idc[i])
for each set.
[0130]
In this case, a rectangular region including the set is set for each set, and
15 may be assigned to each picture (a unit of defining the parameter of the level). For
example, in A of FIG. 25, the rectangular region including the set of identification
information (mcts_id[0]) is set, a size W of the rectangular region in a horizontal
direction is estimated as pic_width_in_luma_samples, a size H of the rectangular
region in a vertical direction is estimated as pic_height_in_luma_samples, a size
20 WxH of the rectangular region is estimated as PicSizeInSamplesY, and the parameter
of the level for the rectangular region may be defined. Similarly, a rectangular
region including the set of identification information (mcts_id[1]) is set, a size W of
the rectangular region in a horizontal direction is estimated as
pic_width_in_luma_samples, a size H of the rectangular region in a vertical direction
25 is estimated as pic_height_in_luma_samples, a size WxH of the rectangular region is
estimated as PicSizeInSamplesY, and the parameter of the level for the rectangular
region may be defined.
[0131]
In this manner, a value of the parameter defined for each rectangular region
30 may be mapped with the parameter defined in the picture unit. Therefore, when
such parameter mapping is performed, it is possible to employ more appropriate
SP353442WO00
33/125
definition for each independently decodable partial region.
[0132]
Also, when the independently decodable partial region is formed of a
plurality of rectangular regions, a rectangle including all these rectangular regions
may be set. For example, in B of 5 FIG. 25, a set of identification information
(mcts_id[0]) includes two rectangular regions, a rectangular region including all of
the set is set, a size W of the rectangular region in a horizontal direction is estimated
as pic_width_in_luma_samples, a size H of the rectangular region in a vertical
direction is estimated as pic_height_in_luma_samples, a size WxH of the rectangular
10 region is estimated as PicSizeInSamplesY, and the parameter of the level for the
rectangular region may be defined.
[0133]
In this manner, even when the independently decodable partial region is
formed of a plurality of rectangular regions, a value of the parameter defined for the
15 rectangular region may be mapped with the parameter defined in the picture unit.
Therefore, when such parameter mapping is performed, it is possible to employ more
appropriate definition of the partial region even if the independently decodable
partial region is formed of a plurality of rectangular regions.
[0134]
20 <2. Second embodiment>
Next, a device configured to implement the present technology described
above and a method thereof will be described. FIG. 27 is a diagram illustrating an
image coding device, which is an aspect of an image processing device to which the
25 present technology is applied. An image coding device 100 illustrated in FIG. 27 is
a device configured to perform layered image coding (scalable coding). As
illustrated in FIG. 27, the image coding device 100 includes a base layer image
coding unit 101, an enhancement layer image coding unit 102, a multiplexing unit
103, and a control unit 104.
30 [0135]
The base layer image coding unit 101 codes a base layer image and
SP353442WO00
34/125
generates a base layer image coding stream. The enhancement layer image coding
unit 102 codes an enhancement layer image and generates an enhancement layer
image coding stream. The multiplexing unit 103 multiplexes the base layer image
coding stream generated in the base layer image coding unit 101 and the
enhancement layer image coding stream generated 5 in the enhancement layer image
coding unit 102, and generates a layered image coding stream. The multiplexing
unit 103 transmits the generated layered image coding stream to the decoding side.
[0136]
The control unit 104 performs settings related to all image data, controls the
10 base layer image coding unit 101 and the enhancement layer image coding unit 102
based on the settings, and thus controls coding of each of the layers. In addition,
the control unit 104 generates the video parameter set (VPS) using the settings,
supplies the parameter to the multiplexing unit 103, and transmits the parameter to
the decoding side. In this case, the video parameter set may be transmitted to be
15 included in the layered image coding stream or may be transmitted as data separate
from the layered image coding stream.
[0137]
In addition, when the decoding load definition information or information
indicating whether the decoding load definition information is set is set in the video
20 parameter set (VPS), the control unit 104 collects the decoding load definition
information or the like from the base layer image coding unit 101 and the
enhancement layer image coding unit 102, and sets the decoding load definition
information or information indicating whether the decoding load definition
information is set in the video parameter set (VPS) based on the information.
25 [0138]
In addition, the base layer image coding unit 101 and the enhancement layer
image coding unit 102 may exchange decoding load related information, which is
information on a magnitude of a load of a decoding process, with each other. For
example, as exemplified in FIG. 11, when the decoding load definition information
30 of a plurality of layers is set, a coding unit of the layer collects decoding load related
information of other layers.
SP353442WO00
35/125
[0139]
FIG. 28 is a block diagram illustrating a main configuration example of the
base layer image coding unit 101 of FIG. 27. As illustrated in FIG. 28, the base
layer image coding unit 101 includes an 5 A/D conversion unit 111, a screen
rearrangement buffer 112, a computation unit 113, an orthogonal transform unit 114,
a quantization unit 115, a reversible coding unit 116, an accumulation buffer 117, an
inverse quantization unit 118, and an inverse orthogonal transform unit 119. In
addition, the base layer image coding unit 101 includes a computation unit 120, a
10 loop filter 121, a frame memory 122, a selection unit 123, an intra prediction unit
124, an inter prediction unit 125, a prediction image selection unit 126, and a rate
control unit 127.
[0140]
The A/D conversion unit 111 performs A/D conversion of input image data
15 (base layer image information), and supplies and stores the converted image data
(digital data) in the screen rearrangement buffer 112. The screen rearrangement
buffer 112 rearranges images of frames of a stored display order according to an
order of frames for coding, depending on a group of picture (GOP), and supplies the
image whose frame order is rearranged to the computation unit 113. In addition,
20 the screen rearrangement buffer 112 supplies the image in which the frame order is
rearranged to the intra prediction unit 124 and the inter prediction unit 125.
[0141]
The computation unit 113 subtracts a prediction image supplied from the
intra prediction unit 124 or the inter prediction unit 125 through the prediction image
25 selection unit 126 from the image read from the screen rearrangement buffer 112,
and outputs difference information thereof to the orthogonal transform unit 114.
For example, in an image on which intra coding is performed, the computation unit
113 subtracts a prediction image supplied from the intra prediction unit 124 from the
image read from the screen rearrangement buffer 112. In addition, for example, in
30 an image on which inter coding is performed, the computation unit 113 subtracts a
prediction image supplied from the inter prediction unit 125 from the image read
SP353442WO00
36/125
from the screen rearrangement buffer 112.
[0142]
The orthogonal transform unit 114 performs an orthogonal transform such
as a discrete cosine transform or a Karhunen-Loève transform on the difference
information supplied from the computation unit 5 113. The orthogonal transform unit
114 supplies a conversion coefficient thereof to the quantization unit 115.
[0143]
The quantization unit 115 quantizes the conversion coefficient supplied
from the orthogonal transform unit 114. The quantization unit 115 sets the
10 quantization parameter based on information on a target value of a code amount
supplied from the rate control unit 127, and performs quantization thereof. The
quantization unit 115 supplies the quantized conversion coefficient to the reversible
coding unit 116.
[0144]
15 The reversible coding unit 116 codes the conversion coefficient quantized in
the quantization unit 115 using an arbitrary coding scheme. Since coefficient data
is quantized under control of the rate control unit 127, the code amount becomes a
target value (or approximates a target value) set by the rate control unit 127.
[0145]
20 In addition, the reversible coding unit 116 acquires information indicating a
mode of intra prediction from the intra prediction unit 124, and acquires information
indicating a mode of inter prediction or differential motion vector information from
the inter prediction unit 125. Further, the reversible coding unit 116 appropriately
generates a network abstraction layer (NAL) unit of a base layer including a
25 sequence parameter set (SPS), a picture parameter set (PPS) and the like.
[0146]
The reversible coding unit 116 codes various pieces of information using an
arbitrary coding scheme and sets the information as a part of coded data (also
referred to as a “coding stream”) (multiplexes). The reversible coding unit 116
30 supplies and accumulates the coded data obtained by coding to the accumulation
buffer 117.
SP353442WO00
37/125
[0147]
Examples of the coding scheme of the reversible coding unit 116 include
variable-length coding and arithmetic coding. Examples of the variable-length
coding include context-adaptive variable length coding (CAVLC) defined in an
H.264 scheme or the AVC scheme. Examples 5 of the arithmetic coding include
context-adaptive binary arithmetic coding (CABAC).
[0148]
The accumulation buffer 117 temporarily maintains the coding stream (base
layer coding stream) supplied from the reversible coding unit 116. The
10 accumulation buffer 117 outputs the maintained base layer coding stream to the
multiplexing unit 103 (FIG. 27) at a predetermined timing. That is, the
accumulation buffer 117 also serves as a transmission unit configured to transmit the
base layer coding stream.
[0149]
15 In addition, the conversion coefficient quantized in the quantization unit 115
is also supplied to the inverse quantization unit 118. The inverse quantization unit
118 performs inverse quantization of the quantized conversion coefficient using a
method corresponding to quantization by the quantization unit 115. The inverse
quantization unit 118 supplies the obtained conversion coefficient to the inverse
20 orthogonal transform unit 119.
[0150]
The inverse orthogonal transform unit 119 performs an inverse orthogonal
transform of the conversion coefficient supplied from the inverse quantization unit
118 using a method corresponding to an orthogonal transform process by the
25 orthogonal transform unit 114. An inverse orthogonal-transformed output (restored
difference information) is supplied to the computation unit 120.
[0151]
The computation unit 120 adds the prediction image from the intra
prediction unit 124 or the inter prediction unit 125 through the prediction image
30 selection unit 126 to the restored difference information, which is an inverse
orthogonal transform result supplied from the inverse orthogonal transform unit 119,
SP353442WO00
38/125
and obtains a locally decoded image (decoded image). The decoded image is
supplied to the loop filter 121 or the frame memory 122.
[0152]
The loop filter 121 includes a deblocking filter, an adaptation loop filter or
the like, and performs an appropriate filter process 5 on a reconstructed image supplied
from the computation unit 120. For example, the loop filter 121 performs a
deblocking filter process on the reconstructed image and thus removes block
distortion of the reconstructed image. In addition, for example, the loop filter 121
performs a loop filter process on the deblocking filter process result (the
10 reconstructed image whose block distortion is removed) using a Wiener filter to
improve image quality. The loop filter 121 supplies the filter process result
(hereinafter also referred to as a “decoded image”) to the frame memory 122.
[0153]
Also, the loop filter 121 may further perform any other filter process on the
15 reconstructed image. In addition, the loop filter 121 may supply information on a
filter coefficient or the like used in the filter process to the reversible coding unit 116
as necessary, and code the information.
[0154]
The frame memory 122 stores the supplied decoded image, and supplies the
20 stored decoded image to the selection unit 123 as a reference image at a
predetermined timing.
[0155]
More specifically, the frame memory 122 stores the reconstructed image
supplied from the computation unit 120 and the decoded image supplied from the
25 loop filter 121. The frame memory 122 supplies the stored reconstructed image to
the intra prediction unit 124 through the selection unit 123 at a predetermined timing
or based on a request from the outside such as the intra prediction unit 124. In
addition, the frame memory 122 supplies the stored decoded image to the inter
prediction unit 125 through the selection unit 123 at a predetermined timing or based
30 on a request from the outside such as the inter prediction unit 125.
[0156]
SP353442WO00
39/125
The selection unit 123 selects a supply destination of the reference image
supplied from the frame memory 122. For example, in intra prediction, the
selection unit 123 supplies the reference image (a pixel value in a current picture or a
base layer decoded image) supplied from the frame memory 122 to the intra
prediction unit 124. In addition, for example, 5 in inter prediction, the selection unit
123 supplies the reference image (a decoded image other than the current picture of
an enhancement layer or the base layer decoded image) supplied from the frame
memory 122 to the inter prediction unit 125.
[0157]
10 The intra prediction unit 124 performs a prediction process on a current
picture, which is an image of frames of a process target, and generates a prediction
image. The intra prediction unit 124 performs the prediction process for each
predetermined block (using a block as a processing unit). That is, the intra
prediction unit 124 generates a prediction image of a current block, which is a
15 process target of the current picture. In this case, the intra prediction unit 124
performs a prediction process (in-screen prediction (also referred to as “intra
prediction”)) using the reconstructed image supplied from the frame memory 122
through the selection unit 123 as the reference image. That is, the intra prediction
unit 124 generates a prediction image using a pixel value in the periphery of the
20 current block included in the reconstructed image. The peripheral pixel value used
in the intra prediction is a pixel value of a pixel, which has been processed, of the
current picture. In the intra prediction (that is, in a method of generating a
prediction image), a plurality of methods (also referred to as “intra prediction
modes”) are prepared in advance as candidates. The intra prediction unit 124
25 performs intra prediction in the plurality of intra prediction modes prepared in
advance.
[0158]
The intra prediction unit 124 generates a prediction image in all of the intra
prediction modes serving as candidates, evaluates a cost function value of each
30 prediction image using an input image supplied from the screen rearrangement buffer
112, and selects an optimal mode. When an optimal intra prediction mode is
SP353442WO00
40/125
selected, the intra prediction unit 124 supplies the prediction image generated in the
optimal mode to the prediction image selection unit 126.
[0159]
In addition, as described above, the intra prediction unit 124 appropriately
supplies intra prediction mode information indicating 5 an employed intra prediction
mode or the like to the reversible coding unit 116 and codes it.
[0160]
The inter prediction unit 125 performs a prediction process on the current
picture and generates a prediction image. The inter prediction unit 125 performs
10 the prediction process for each predetermined block (using a block as a processing
unit). That is, the inter prediction unit 125 generates a prediction image of a current
block, which is a process target of the current picture. In this case, the inter
prediction unit 125 performs the prediction process using image data of the input
image supplied from the screen rearrangement buffer 112 and image data of the
15 decoded image supplied from the frame memory 122 as the reference image. The
decoded image is an image (a picture other than the current picture) of frames that
are processed before the current picture. That is, the inter prediction unit 125
performs a prediction process (inter-screen prediction (also referred to as “inter
prediction”)) of generating a prediction image using an image of another picture.
20 [0161]
The inter prediction is performed by motion prediction and motion
compensation. More specifically, the inter prediction unit 125 uses the input image
and the reference image, performs a motion prediction on the current block, and
detects a motion vector. Therefore, the inter prediction unit 125 uses the reference
25 image, performs a motion compensation process according to the detected motion
vector, and generates a prediction image (inter prediction image information) of the
current block. In the inter prediction (that is, in a method of generating a prediction
image), a plurality of methods (also referred to as “inter prediction modes) are
prepared in advance as candidates. The inter prediction unit 125 performs such
30 inter prediction in the plurality of inter prediction modes prepared in advance.
[0162]
SP353442WO00
41/125
The inter prediction unit 125 generates a prediction image in all of the inter
prediction modes serving as candidates. The inter prediction unit 125 uses the input
image supplied from the screen rearrangement buffer 112 and information of a
generated difference motion vector, evaluates a cost function value of each
prediction image and selects an 5 optimal mode. When an optimal inter prediction
mode is selected, the inter prediction unit 125 supplies the prediction image
generated in the optimal mode to the prediction image selection unit 126.
[0163]
When information indicating an employed inter prediction mode or the
10 coded data is decoded, the inter prediction unit 125 supplies information necessary
for performing a process in the inter prediction mode or the like to the reversible
coding unit 116 and codes it. As necessary information, for example, information
of a generated difference motion vector or a flag indicating an index of a predicted
motion vector as predicted motion vector information is exemplified.
15 [0164]
The prediction image selection unit 126 selects a supply source of the
prediction image supplied to the computation unit 113 or the computation unit 120.
For example, in intra coding, the prediction image selection unit 126 selects the intra
prediction unit 124 as the supply source of the prediction image, and supplies the
20 prediction image supplied from the intra prediction unit 124 to the computation unit
113 or the computation unit 120. In addition, for example, in inter coding, the
prediction image selection unit 126 selects the inter prediction unit 125 as the supply
source of the prediction image, and supplies the prediction image supplied from the
inter prediction unit 125 to the computation unit 113 or the computation unit 120.
25 [0165]
The rate control unit 127 controls a rate of the quantization operation of the
quantization unit 115 based on the code amount of the coded data accumulated in the
accumulation buffer 117 so that no overflow or underflow occurs.
[0166]
30 In addition, the frame memory 122 supplies the stored base layer decoded
image to the enhancement layer image coding unit 102.
SP353442WO00
42/125
[0167]
In addition, as illustrated in FIG. 28, the base layer image coding unit 101
further includes a header information generating unit 128.
[0168]
The header information generating 5 unit 128 generates header information
such as the sequence parameter set (SPS) or MCTS SEI. In this case, as described
in the first embodiment, the header information generating unit 128 performs a
process of setting the decoding load definition information for defining a magnitude
of a load of a decoding process of the independently decodable partial region. For
10 example, the header information generating unit 128 may acquire decoding load
related information of the base layer from the reversible coding unit 116 and generate
the decoding load definition information of the independently decodable partial
region of the base layer based on the decoding load related information. In
addition, the header information generating unit 128 may acquire, for example,
15 decoding load related information of the enhancement layer from the enhancement
layer image coding unit 102, and generate the decoding load definition information
of the independently decodable partial region of the enhancement layer based on the
decoding load related information.
[0169]
20 Further, the header information generating unit 128 may supply the
decoding load definition information or the like to the control unit 104, and enable
settings for the decoding load definition information to be performed in the video
parameter set.
[0170]
25
FIG. 29 is a block diagram illustrating a main configuration example of the
enhancement layer image coding unit 102 of FIG. 27. As illustrated in FIG. 29, the
enhancement layer image coding unit 102 basically has the same configuration as the
base layer image coding unit 101 of FIG. 28.
30 [0171]
That is, as illustrated in FIG. 29, the enhancement layer image coding unit
SP353442WO00
43/125
102 includes an A/D conversion unit 131, a screen rearrangement buffer 132, a
computation unit 133, an orthogonal transform unit 134, a quantization unit 135, a
reversible coding unit 136, an accumulation buffer 137, an inverse quantization unit
138, and an inverse orthogonal transform unit 139. In addition, the enhancement
layer image coding unit 102 includes a 5 computation unit 140, a loop filter 141, a
frame memory 142, a selection unit 143, an intra prediction unit 144, an inter
prediction unit 145, a prediction image selection unit 146, and a rate control unit 147.
[0172]
The A/D conversion unit 131 to the rate control unit 147 correspond to and
10 perform the same processes as the A/D conversion unit 111 to the rate control unit
127 of FIG. 28. However, respective units of the enhancement layer image coding
unit 102 perform a process of coding enhancement layer image information rather
than the base layer. Therefore, when processes of the A/D conversion unit 131 to
the rate control unit 147 are described, the above descriptions of the A/D conversion
15 unit 111 to the rate control unit 127 of FIG. 28 may be applied. However, in this
case, it is necessary that data to be processed be data of the enhancement layer rather
than data of the base layer. In addition, it is necessary to appropriately replace a
processing unit of an input source or an output destination of data with a
corresponding processing unit among the A/D conversion unit 131 to the rate control
20 unit 147 and read it.
[0173]
The enhancement layer image coding unit 102 further includes a header
information generating unit 148.
[0174]
25 The header information generating unit 148 corresponds to the header
information generating unit 128 of FIG. 28 and performs the same process as the
header information generating unit 128. However, the header information
generating unit 148 performs a process of the enhancement layer rather than the base
layer.
30 [0175]
Alternatively, when the decoding load definition information of the
SP353442WO00
44/125
enhancement layer is also created in the header information generating unit 128 of
the base layer, the header information generating unit 148 of the enhancement layer
may be omitted.
[0176]
<5 Header information generating unit>
FIG. 30 is a diagram illustrating an exemplary configuration of function
blocks of the header information generating unit 128 of FIG. 28. In the header
information generating unit 128, for example, when a program read by a CPU from a
ROM or the like is executed using a RAM, the above-described process is executed
10 and thus various function blocks illustrated in FIG. 30 are implemented.
[0177]
As illustrated in FIG. 30, the header information generating unit 128
includes a decoding load related information acquisition unit 151, an MCTS SEI
generating unit 152, and an SPS generating unit 153.
15 [0178]
The decoding load related information acquisition unit 151 acquires
information on a load of the decoding process that is used to generate the decoding
load definition information of the independently decodable partial region. As long
as it is used to generate the decoding load definition information, any content of
20 information on a load of the decoding process may be used.
[0179]
As described in the first embodiment, the MCTS SEI generating unit 152
generates MCTS SEI including the decoding load definition information of the
independently decodable partial region. That is, the MCTS SEI generating unit 152
25 sets the decoding load definition information of the independently decodable partial
region in MCTS SEI. Content of the decoding load definition information is
arbitrary. For example, in the first embodiment, any one or more of various pieces
of information described with reference to FIGS. 9 to 26 may be included in the
decoding load definition information. Also, when a level is defined in a partial
30 region unit, the MCTS SEI generating unit 152 also defines a parameter of the level
for the partial region, as described in the first embodiment. The definition may be
SP353442WO00
45/125
assigned (parameter mapping is performed) to definition of the level of the entire
image (picture) unit.
[0180]
As described in the first embodiment, the SPS generating unit 153 generates
the sequence parameter set (SP 5 S) including the decoding load definition information
of the independently decodable partial region or information indicating whether the
decoding load definition information of the independently decodable partial region is
set in MCTS SEI (also referred to collectively as “information on definition of a
decoding load”). That is, the SPS generating unit 153 sets information on definition
10 of a decoding load of the independently decodable partial region in the sequence
parameter set (SPS). Content of the information on definition of a decoding load is
arbitrary. For example, in the first embodiment, any one or more of various pieces
of information described with reference to FIGS. 9 to 26 may be included in the
decoding load definition information.
15 [0181]
Also, as described in the first embodiment, the decoding load definition
information of the independently decodable partial region is set only in MCTS SEI,
and this information may not be set in the sequence parameter set (SPS). In this
case, the SPS generating unit 153 may be omitted.
20 [0182]
Next, a flow of processes executed by the image coding device 100
described above will be described. First, an exemplary flow of image coding
processes will be described with reference to a flowchart of FIG. 31.
25 [0183]
When the image coding process starts, the control unit 104 of the image
coding device 100 performs settings of entire scalable coding in Step S101.
[0184]
In Step S102, the control unit 104 controls respective units of the base layer
30 image coding unit 101 to the multiplexing unit 103 according to settings performed
in Step S101.
SP353442WO00
46/125
[0185]
In Step S103, the control unit 104 generates a video parameter set (VPS) by
applying the settings performed in Step S101.
[0186]
In Step S104, the base layer image coding unit 101 5 codes image data of the
base layer.
[0187]
In Step S105, the enhancement layer image coding unit 102 codes image
data of the enhancement layer.
10 [0188]
In Step S106, the multiplexing unit 103 multiplexes the base layer image
coding stream generated in Step S104 and the enhancement layer image coding
stream generated in Step S105 (that is, bitstreams of the layers), and generates the
layered image coding stream of one system. Also, the multiplexing unit 103
15 includes the video parameter set (VPS) generated in Step S103 in the layered image
coding stream as necessary. The multiplexing unit 103 outputs the layered image
coding stream and transmits the stream to the decoding side.
[0189]
When the process of Step S106 ends, the image coding device 100 ends the
20 image coding process. One picture is processed by such image coding processes.
Therefore, the image coding device 100 repeatedly executes such image coding
processes for each picture of layered moving image data. However, processes that
are not necessarily performed for each picture, for example, the processes of Steps
S101 to S103, are appropriately omitted.
25 [0190]
Next, in Step S104 of FIG. 31, an exemplary flow of base layer coding
processes executed by the base layer image coding unit 101 will be described with
reference to a flowchart of FIG. 32.
30 [0191]
When the base layer coding process starts, the A/D conversion unit 111 of
SP353442WO00
47/125
the base layer image coding unit 101 performs A/D conversion of an image of frames
(picture) of the input moving image in Step S121.
[0192]
In Step S122, the screen rearrangement buffer 112 stores the image on
which A/D conversion is performed in Step 5 S121, and performs rearrangement
according to an order of coding from a display order of pictures.
[0193]
In Step S123, the intra prediction unit 124 performs an intra prediction
process in an intra prediction mode.
10 [0194]
In Step S124, the inter prediction unit 125 performs an inter prediction
process in which motion prediction or motion compensation is performed in the inter
prediction mode.
[0195]
15 In Step S125, the prediction image selection unit 126 selects the prediction
image based on a cost function value or the like. That is, the prediction image
selection unit 126 selects any of the prediction image generated by intra prediction of
Step S123 and the prediction image generated by inter prediction of Step S124.
[0196]
20 In Step S126, the computation unit 113 computes a difference between the
input image whose frame order is rearranged in the process of Step S122 and the
prediction image selected in the process of Step S125. That is, the computation unit
113 generates image data of a difference image between the input image and the
prediction image. The image data of the difference image obtained in this manner
25 has a smaller amount of data than original image data. Therefore, compared to
when the image is directly coded, it is possible to compress an amount of data.
[0197]
In Step S127, the orthogonal transform unit 114 performs orthogonal
transform of the image data of the difference image generated in the process of Step
30 S126.
[0198]
SP353442WO00
48/125
In Step S128, the quantization unit 115 uses the quantization parameter
calculated by the rate control unit 127 and quantizes the orthogonal transform
coefficient obtained in the process of Step S127.
[0199]
In Step S129, the inverse 5 quantization unit 118 performs inverse
quantization of the coefficient (also referred to as a “quantization coefficient”)
generated and quantized in the process of Step S128 using a characteristic
corresponding to a characteristic of the quantization unit 115.
[0200]
10 In Step S130, the inverse orthogonal transform unit 119 performs an inverse
orthogonal transform of the orthogonal transform coefficient obtained in the process
of Step S129.
[0201]
In Step S131, the computation unit 120 adds the prediction image selected
15 in the process of Step S125 to the difference image restored in the process of Step
S130, and thus generates image data of the reconstructed image.
[0202]
In Step S132, the loop filter 121 performs the loop filter process of the
image data of the reconstructed image generated in the process of Step S131.
20 Therefore, block distortion of the reconstructed image or the like is removed.
[0203]
In Step S133, the frame memory 122 stores data such as the decoded image
(the base layer decoded image) obtained in the process of Step S132 or the
reconstructed image obtained in the process of Step S131.
25 [0204]
In Step S134, the reversible coding unit 116 codes the coefficient obtained
and quantized in the process of Step S128. That is, reversible coding such as
variable-length coding or arithmetic coding is performed on data corresponding to
the difference image.
30 [0205]
In addition, in this case, the reversible coding unit 116 codes information on
SP353442WO00
49/125
a prediction mode of the prediction image selected in the process of Step S125, and
adds the difference image to the coded data obtained by coding. That is, the
reversible coding unit 116 also codes optimal intra prediction mode information
supplied from the intra prediction unit 124 or information corresponding to the
optimal inter prediction mode supplied 5 from the inter prediction unit 125 and adds
the result to the coded data.
[0206]
In Step S135, the header information generating unit 128 generates header
information of various null units or the like. The generated header information is
10 supplied to the reversible coding unit 116 and added to the coded data.
[0207]
In Step S136, the accumulation buffer 117 accumulates the coded data (the
base layer image coding stream) obtained in the processes of Step S134 and Step
S135. The base layer image coding stream accumulated in the accumulation buffer
15 117 is appropriately read, supplied to the multiplexing unit 103, multiplexed with the
enhancement layer image coding stream, and then is transmitted to the decoding side
through a transmission path or a recording medium.
[0208]
In Step S137, the rate control unit 127 controls a rate of the quantization
20 operation of the quantization unit 115 based on the code amount (an amount of
generated codes) of the coded data accumulated in the accumulation buffer 117 in the
process of Step S136 so that no overflow or underflow occurs. In addition, the rate
control unit 127 supplies information on the quantization parameter to the
quantization unit 115.
25 [0209]
When the process of Step S137 ends, the base layer coding process ends,
and the process returns to FIG. 31.
[0210]
30 Next, in Step S105 of FIG. 31, an exemplary flow of enhancement layer
coding processes executed by the enhancement layer image coding unit 102 will be
SP353442WO00
50/125
described with reference to a flowchart of FIG. 33.
[0211]
Respective processes (Steps S141 to S157) of the enhancement layer coding
processes correspond to respective processes (Steps S121 to S137) of the base layer
coding processes, and are executed basically in the 5 same manner as these processes.
While the processes of the base layer coding processes are performed on the base
layer, respective processes (Steps S141 to S157) of the enhancement layer coding
processes are performed on the enhancement layer.
[0212]
10 In addition, when the decoding load definition information of the
independently decodable partial region is set only in the base layer, settings of the
decoding load definition information may be omitted in Step S155.
[0213]
When the process of Step S157 ends, the enhancement layer coding process
15 ends, and the process returns to FIG. 31.
[0214]
Next, an exemplary flow of header generation processes executed in Step
S135 of FIG. 32 will be described with reference to a flowchart of FIG. 34.
20 [0215]
When the header generation process starts, the header information
generating unit 128 generates various pieces of header information, for example, a
sequence parameter set (SPS), SEI, a picture parameter set (PPS), and a slice header
(SliceHeader) in Step S161.
25 [0216]
In Step S162, the decoding load related information acquisition unit 151
acquires decoding load related information, which is information on a load of a
decoding process of the partial region, necessary for generating the decoding load
definition information of the independently decodable partial region. Additionally,
30 when the decoding load definition information of the enhancement layer is set,
decoding load related information is also acquired from the enhancement layer. As
SP353442WO00
51/125
long as it is used to generate the decoding load definition information, any content of
the decoding load related information may be used.
[0217]
In Step S163, the MCTS SEI generating unit 152 sets the decoding load
definition information of the partial region 5 in MCTS SEI of the independently
decodable partial region generated in Step S161. For example, the MCTS SEI
generating unit 152 sets the decoding load definition information, as described with
reference to syntaxes of FIGS. 9 to 26 in the first embodiment.
[0218]
10 For example, as illustrated in FIG. 10, when the independently decodable
partial region includes a plurality of sets (a plurality of tiles), the MCTS SEI
generating unit 152 uses decoding load related information of each tile, and sets
decoding load related information for each set. In addition, the MCTS SEI
generating unit 152 may use decoding load related information of each tile and set
15 the decoding load definition information of the entire partial region.
[0219]
Also, when there are a plurality of independently decodable partial regions,
the MCTS SEI generating unit 152 may set the decoding load definition information
for each partial region. Content of the decoding load definition information is
20 arbitrary. For example, in the first embodiment, any one or more of various pieces
of information described with reference to FIGS. 9 to 26 may be included in the
decoding load definition information.
[0220]
In Step S164, the SPS setting unit 153 sets the decoding load definition
25 information of the independently decodable partial region, or information (also
referred to as “information on definition of a decoding load”) indicating whether the
decoding load definition information of the partial region is set in MCTS SEI in the
sequence parameter set (SPS) generated in Step S161. Content of the information
on definition of a decoding load is arbitrary. For example, in the first embodiment,
30 any one or more of various pieces of information described with reference to FIGS. 9
to 26 may be included in the decoding load definition information.
SP353442WO00
52/125
[0221]
When the process of Step S164 ends, the header information generating
process ends and the process returns to FIG. 32.
[0222]
Also, as described in the 5 first embodiment, the decoding load definition
information of the independently decodable partial region is set only in MCTS SEI,
and this information may not be set in the sequence parameter set (SPS). In this
case, the process of Step S164 may be omitted.
[0223]
10 In addition, in Step S163, when a level is defined in a partial region unit, the
MCTS SEI generating unit 152 also defines a parameter of the level for the partial
region, as described in the first embodiment. The definition may be assigned to
definition of the level of the entire image (picture) unit (parameter mapping is
performed).
15 [0224]
The header information set as described above is supplied to the reversible
coding unit 116 and included in the coded data.
[0225]
When the respective processes are executed in this manner, the image
20 coding device 100 can recognize performance necessary for decoding more
accurately.
[0226]
Also, when the decoding load definition information of the independently
decodable partial region is set in the enhancement layer, the header information
25 generating unit 148 may be executed as described with reference to the flowchart of
FIG. 34. On the other hand, when no decoding load definition information is set,
only the process of Step S161 of FIG. 34 may be executed.
[0227]
<3. Third embodiment>
30
Next, decoding of the coded data coded as described above will be
SP353442WO00
53/125
described. FIG. 35 is a block diagram illustrating a main configuration example of
an image decoding device corresponding to the image coding device 100, which is an
aspect of an image processing device to which the present technology is applied.
An image decoding device 200 illustrated in FIG. 35 decodes coded data generated
by the image coding device 100 using 5 a decoding method corresponding to the
coding method (that is, hierarchically decodes coded data that is hierarchically
coded). As illustrated in FIG. 35, the image decoding device 200 includes a
demultiplexing unit 201, a base layer image decoding unit 202, an enhancement layer
image decoding unit 203, and a control unit 204.
10 [0228]
The demultiplexing unit 201 receives the layered image coding stream in
which the base layer image coding stream and the enhancement layer image coding
stream transmitted from the coding side are multiplexed, demultiplexes the received
stream, and extracts the base layer image coding stream and the enhancement layer
15 image coding stream. The base layer image decoding unit 202 decodes the base
layer image coding stream extracted by the demultiplexing unit 201 and obtains the
base layer image. The enhancement layer image decoding unit 203 decodes the
enhancement layer image coding stream extracted by the demultiplexing unit 201
and obtains the enhancement layer image.
20 [0229]
The control unit 204 analyzes the video parameter set (VPS) supplied from
the demultiplexing unit 201, and controls the base layer image decoding unit 202 and
the enhancement layer image decoding unit 203 based on the information (controls
coding of each of the layers).
25 [0230]
In addition, the control unit 204 acquires the analysis result of the decoding
load definition information of the header information from the base layer image
decoding unit 202 and the enhancement layer image decoding unit 203, and controls
operations of respective processing units of the image decoding device 200 according
30 to the analysis result.
[0231]
SP353442WO00
54/125
FIG. 36 is a block diagram illustrating a main configuration example of the
base layer image decoding unit 202 of FIG. 35. As illustrated in FIG. 36, the base
layer image decoding unit 202 includes an accumulation buffer 211, a reversible
decoding unit 212, an inverse quantization 5 unit 213, an inverse orthogonal transform
unit 214, a computation unit 215, a loop filter 216, a screen rearrangement buffer
217, and a D/A conversion unit 218. In addition, the base layer image decoding
unit 202 includes a frame memory 219, a selection unit 220, an intra prediction unit
221, an inter prediction unit 222, and a prediction image selection unit 223.
10 [0232]
The accumulation buffer 211 also serves as a reception unit configured to
receive the transmitted coded data (the base layer image coding stream supplied from
the demultiplexing unit 201). The accumulation buffer 211 receives and
accumulates the transmitted coded data, and supplies the coded data to the reversible
15 decoding unit 212 at a predetermined timing. Information necessary for decoding
such as prediction mode information is added to the coded data.
[0233]
The reversible decoding unit 212 decodes the information that is supplied by
the accumulation buffer 211 and coded by the reversible coding unit 116 using a
20 decoding scheme corresponding to the coding scheme. The reversible decoding
unit 212 supplies the quantized coefficient data of the difference image obtained by
decoding to the inverse quantization unit 213.
[0234]
In addition, the reversible decoding unit 212 determines whether the intra
25 prediction mode or the inter prediction mode is selected as an optimal prediction
mode, and supplies information on the optimal prediction mode in a mode that is
determined to have been selected between the intra prediction unit 221 and the inter
prediction unit 222. That is, for example, when the intra prediction mode is
selected as the optimal prediction mode on the coding side, the information on the
30 optimal prediction mode (intra prediction mode information) is supplied to the intra
prediction unit 221. In addition, for example, when the inter prediction mode is
SP353442WO00
55/125
selected as the optimal prediction mode on the coding side, the information on the
optimal prediction mode (inter prediction mode information) is supplied to the inter
prediction unit 222.
[0235]
Further, the reversible decoding unit 212 5 extracts information necessary for
inverse quantization, for example, a quantization matrix or a quantization parameter,
from the coded data, and supplies the information to the inverse quantization unit
213.
[0236]
10 The inverse quantization unit 213 performs inverse quantization of the
quantized coefficient data obtained by decoding of the reversible decoding unit 212
using a scheme corresponding to a quantization scheme of the quantization unit 115.
Also, the inverse quantization unit 213 is the same processing unit as the inverse
quantization unit 118. The inverse quantization unit 213 supplies the obtained
15 coefficient data (the orthogonal transform coefficient) to the inverse orthogonal
transform unit 214.
[0237]
The inverse orthogonal transform unit 214 performs an inverse orthogonal
transform of the orthogonal transform coefficient supplied from the inverse
20 quantization unit 213 using a scheme corresponding to an orthogonal transform
scheme of the orthogonal transform unit 114 as necessary. Also, the inverse
orthogonal transform unit 214 is the same processing unit as the inverse orthogonal
transform unit 119.
[0238]
25 According to the inverse orthogonal transform process, the image data of
the difference image is restored. The restored image data of the difference image
corresponds to the image data of the difference image before an orthogonal transform
is performed on the coding side. In the following, the restored image data of the
difference image obtained by the inverse orthogonal transform process of the inverse
30 orthogonal transform unit 214 is also referred to as “decoded residual data.” The
inverse orthogonal transform unit 214 supplies the decoded residual data to the
SP353442WO00
56/125
computation unit 215. In addition, the image data of the prediction image is
supplied from the intra prediction unit 221 or the inter prediction unit 222 to the
computation unit 215 through the prediction image selection unit 223.
[0239]
The computation unit 215 uses the decoded 5 residual data and the image data
of the prediction image, and obtains the image data of the reconstructed image in
which the difference image and the prediction image are added. The reconstructed
image corresponds to the input image before the prediction image is subtracted by
the computation unit 113. The computation unit 215 supplies the reconstructed
10 image to the loop filter 216.
[0240]
The loop filter 216 appropriately performs the loop filter process including
the deblocking filter process or an adaptive loop filter process on the supplied
reconstructed image, and generates the decoded image. For example, the loop filter
15 216 performs the deblocking filter process on the reconstructed image and thus
removes block distortion. In addition, for example, the loop filter 216 performs the
loop filter process on the deblocking filter process result (the reconstructed image
whose block distortion is removed) using a Wiener filter to improve image quality.
[0241]
20 Also, a type of the filter process performed by the loop filter 216 is
arbitrary, and a filter process other than the above-described process may be
performed. In addition, the loop filter 216 may perform the filter process using the
filter coefficient supplied from the image coding device. Further, the loop filter 216
may omit such a filter process and output input data without the filter process.
25 [0242]
The loop filter 216 supplies the decoded image (or the reconstructed image),
which is the filter process result, to the screen rearrangement buffer 217 and the
frame memory 219.
[0243]
30 The screen rearrangement buffer 217 rearranges an order of frames of the
decoded image. That is, the screen rearrangement buffer 217 rearranges the image
Claim 1 (currently amended)
An image decoding device comprising:
a control unit configured to control decoding of coded data in which the
image is coded based on information indicating a size of a partial region serving as a
5 reference of an image and a level indicating a magnitude of a load of a decoding
process of the partial region; and
a decoding unit configured to decode the coded data under control of the
control unit.
10 Claim 2
15
The image decoding device according to claim 1,
wherein the partial region is independently decodable.
Claim 3 (currently amended)
The image decoding device according to claim 1,
wherein the control unit controls decoding of the coded data using
information for defining a magnitude of a load of a decoding process of the partial
region according to a level indicating a magnitude of a load of the decoding process.
20 Claim 4 (currently amended)
25
The image decoding device according to claim 1,
wherein the control unit controls decoding of the coded data usmg
information for defining a magnitude of a load of a decoding process of the partial
region according to information indicating a size of the partial region.
Claim 5 (currently amended)
The image decoding device according to claim I,
wherein the control unit controls decoding of the coded data using
information for defining a magnitude of a load of a decoding process of the partial
30 region according to information indicating a length in a vertical direction and
information indicating a length in a horizontal direction of the partial region.
SP353442WOOO
Claim 6 (currently amended)
The image decoding device according to claim l, further comprising
an acquisition unit configured to acquire information indicating a size of the partial
5 region and a level indicating a magnitude of a load of a decoding process of the
partial region as decoping load definition information indicating a magnitude of a
load of a decoding process of the partial region,
10
15
wherein the control unit controls decoding of the coded data using the
decoding load definition information acquired by the acquisition unit.
Claim 7 (currently amended)
The image decoding device according to claim 6,
wherein the acquisition unit acquires the decoding load definition
information as auxiliary information of the coded data.
Claim 8 (currently amended)
The image decoding device according to claim 7,
wherein the acquisition unit acquires the decoding load definition
information as supplemental enhancement information (SEI) of an independently
20 decodable partial region.
Claim 9 (currently amended)
The image decoding device according to claim 8,
wherein the image data includes a plurality of layers, and
25 wherein the acquisition unit acquires the decoding load definition
30
information of the plurality oflayers as the SEI.
Claim 1 0 (currently amended)
The image decoding device according to claim l,
wherein the partial region is a tile.
'!
Claim II (currently amended)
The image decoding device according to claim 1,
wherein the partial region is a set of a plurality of tiles.
SP353442WOOO
5 Claim 12 (currently amended)
The image decoding device according to claim 1,
wherein the control unit controls decoding of the coded data using
information for defining a maximum magnitude of a load of a decoding process
among a plurality of partial regions included in a picture of the image data according
10 to a level indicating a magnitude of a load of the decoding process.
Claim 13 (currently amended)
The image decoding device according to claim I,
wherein the control unit controls decoding of the coded data using
15 information for defining a magnitude of a load common in a plurality of partial
regions included in a picture of the image data according to a level indicating a
magnitude of a load of the decoding process.
20
Claim 14 (currently amended)
The image decoding device according to claim 13,
wherein, when a plurality of the partial regions included in the picture have
an L shape, the control unit controls decoding of the coded data using information for
defining a magnitude of the load for a rectangular region including the L shape.
25 Claim 15 (currently amended)
The image decoding device according to claim 1,
wherein the control unit controls decoding of the coded data usmg
information indicating whether the decoding load definition information is set.
30 Claim 16 (new)
An image decoding method comprising:
5
10
15
20
25
SP353442WOOO
based on information indicating a size of a partial region serving as a
reference of an image and a level indicating a magnitude of a load of a decoding
process of the partial region, controlling decoding of coded data in which the image
is coded; and
decoding the coded data according to the control.
Claim 17 (new)
An image coding device comprising:
a coding unit configured to code an image; and
a setting unit configured to set information indicating a size of a partial
region serving as a reference of the image and a level indicating a magnitude of a
load of a decoding process of the partial region.
Claim 18 (new)
The image coding device according to claim 17,
wherein the partial region is independently decodable.
Claim 19 (new)
The image coding device according to claim 17,
wherein the setting unit generates information for defining a magnitude of a
load of a decoding process of the partial region according to a level indicating a
magnitude of a load of the decoding process.
Claim 20 (new)
The image coding device according to claim 17,
wherein the setting unit generates information for defining a magnitude of a
load of a decoding process of the partial region according to information indicating a
size of the partial region.
30 Claim 21 (new)
The image coding device according to claim 17,
5
SP353442WOOO
wherein the setting unit generates information for defining a magnitude of a
load of a decoding process of the partial region according to information indicating a
length in a vertical direction and information indicating a length in a horizontal
direction of the partial region.
Claim 22 (new)
The image coding device according to claim 17, further comprising
a transmission unit configured to transmit information indicating a size of
the partial region and a level indicating a magnitude of a load of a decoding process
10 of the partial region as decoding load definition information indicating a magnitude
of a load of a decoding process of the partial region.
15
20
Claim 23 (new)
The image coding device according to claim 22,
wherein the transmission unit transmits the decoding load definition
information as auxiliary information of coded data in which the image obtained by
the coding unit is coded.
Claim 24 (new)
The image coding device according to claim 23,
wherein the transmission unit transmits the decoding load definition
information as supplemental enhancement infmmation (SEI) of an independently
decodable partial region
25 Claim 25 (new)
30
The image coding device according to claim 24,
wherein the image data includes a plurality of layers, and
wherein the transmission unit transmits the decoding load definition
information of the plurality of layers as the SEI.
')'Yl
Claim 26 (new)
The image coding device according to claim 17,
wherein the partial region is a tile.
SP353442WOOO
5 Claim 27 (new)
10
15
The image coding device according to claim 17,
wherein the partial region is a set of a plurality of tiles.
Claim 28 (new)
The image coding device according to claim 17,
wherein the setting unit generates information for defining a maximum
magnitude of a load of a decoding process among a plurality of partial regions
included in a picture of the image data according to a level indicating a magnitude of
a load of the decoding process.
Claim 29 (new)
The image coding device according to claim 17,
wherein the setting unit generates information for defining a magnitude of a
load common in a plurality of partial regions included in a picture of the image data
20 according to a level indicating a magnitude of a load of the decoding process.
Claim 30 (new)
The image coding device according to claim 29,
wherein, when a plurality of the partial regions included in the picture have
25 an L shape, the setting unit generates information for defining a magnitude of the
load for a rectangular region including the L shape.
Claim 31 (new)
The image coding device according to claim l 7,
30 wherein the setting unit generates information indicating whether the
decoding load definition information is set.
5
Claim 32 (new)
An image coding method comprising:
coding an image; and
SP353442WOOO
setting information indicating a size of a partial region serving as a
reference of the image and a level indicating a magnitude of a load of a decoding
process of the partial region.
| # | Name | Date |
|---|---|---|
| 1 | Priority Document [04-01-2016(online)].pdf | 2016-01-04 |
| 2 | Power of Attorney [04-01-2016(online)].pdf | 2016-01-04 |
| 3 | Form 5 [04-01-2016(online)].pdf | 2016-01-04 |
| 4 | Form 3 [04-01-2016(online)].pdf | 2016-01-04 |
| 5 | Form 1 [04-01-2016(online)].pdf | 2016-01-04 |
| 6 | Drawing [04-01-2016(online)].pdf | 2016-01-04 |
| 7 | Description(Complete) [04-01-2016(online)].pdf | 2016-01-04 |
| 9 | 201617000122-Form-1-(14-01-2016).pdf | 2016-01-14 |
| 10 | 201617000122-Correspondence Others-(14-01-2016).pdf | 2016-01-14 |
| 11 | 201617000122-Form-3-(29-04-2016).pdf | 2016-04-29 |
| 12 | 201617000122-Correspondence Others-(29-04-2016).pdf | 2016-04-29 |
| 13 | abstract.jpg | 2016-06-22 |
| 14 | Form 3 [31-08-2016(online)].pdf | 2016-08-31 |
| 15 | Form 18 [19-05-2017(online)].pdf | 2017-05-19 |
| 16 | 201617000122-OTHERS [02-03-2021(online)].pdf | 2021-03-02 |
| 17 | 201617000122-FER_SER_REPLY [02-03-2021(online)].pdf | 2021-03-02 |
| 18 | 201617000122-DRAWING [02-03-2021(online)].pdf | 2021-03-02 |
| 19 | 201617000122-CORRESPONDENCE [02-03-2021(online)].pdf | 2021-03-02 |
| 20 | 201617000122-CLAIMS [02-03-2021(online)].pdf | 2021-03-02 |
| 21 | 201617000122-FER.pdf | 2021-10-17 |
| 22 | 201617000122-PatentCertificate14-08-2023.pdf | 2023-08-14 |
| 23 | 201617000122-IntimationOfGrant14-08-2023.pdf | 2023-08-14 |
| 1 | SearchStrategyfor201617000122E_26-08-2020.pdf |