Abstract: The present invention pertains to an image processing device and method that enable large reductions in encoding efficiency to be prevented. Provided are: a receiving unit that receives an encoded bit stream which includes syntax elements pertaining to processing between images; and an analysis unit that analyzes the syntax elements received by the receiving unit in a state in which values are limited when the bit stream has been encoded according to a profile for encoding static images. The present invention can be applied to an image processing device for example.
IMAGE PROCESSING DEVICE AND METHOD
TECHNICAL FIELD [0001]
The present disclosure relates to an image processing device and method, and more particularly, an image processing device and method vjhich are capable of suppressing a reduction in coding efficiency.
BACKGROUND ART [0002]
In recent years, for the purpose of digitalizing image information and transmitting and accumulating information at high efficiency at that time, devices that compress and encode images by using image information-specific redundancy employing a coding scheme that performs compression through an orthogonal transform such as a discrete cosine transform and motion compensation have been spread. As such a coding scheme, for example, there is Moving Picture Experts Group (MPEG). [0003]
Particularly, MPEG 2 (International Organization for Standardization/International Electrotechnical Commission (ISO/IEC) 13818-2) is a standard that is defined as a general-purpose image coding scheme, and covers interlaced scan images, progressive scan images, standard resolution images, and high definition images. Currently, MPEG 2 is being vjidely used for a v/ide range of applications such as professional use and consumer use Using the MPEG 2 compression scheme, for example, in the
case of an interlaced scan image of a standard resolution
having 720 x 480 pixels, a coding amount (bit rate) of 4 Mbps to 8 Mbps is allocated. Further, using the MPEG 2 compression scheme, for example, in the case of an
interlaced scan image of a high resolution having 1920 x 1088 pixels, a coding amount {bit rate) of 18 Mbps to 22 Mbps is allocated. Thus, it is possible to implement a high compression rate and an excellent image quality. [0004]
MPKG 2 is mainly intended for high definition coding suitable for broadcasting but does not support a coding scheme having a coding amount (bit rate) lov^er than that of MPEG 1, that is, a coding scheme of a higher compression rate. With the spread of mobile terminals, the need for such a coding scheme is considered to be increased in the future, and thus an MPEG 4 coding scheme has been standardized. In connection v/ith an image coding scheme, an international standard thereof has been approved as ISO/IEC 14496-2 in December, 1998. [0005J
Further, in recent years, standardization of a standard such as H.26L (International Telecommunication Union Telecommunication Standardization Sector Q6/16 Video Coding Expert Group (ITU-T Q6/16 VCEG)) for the purpose of image coding for video conference has been conducted. H.26L requires a larger computation amount for coding and decoding than in an existing coding scheme such as MPEG 2 or MPEG 4, but is knov/n to implement higher coding efficiency. Further/ currently, as one of activities of MPEG 4, standardization of incorporating even a function that is not supported in H.2 6L and
implementing higher coding efficiency based on H.26L has been performed as a Joint Model of Enhanced-Compression Video Coding. [0006]
As a standardization schedule, an international standard called H.264 and MPEG-4 PartlO (which is also hereinafter referred to as "Advanced Video Coding (AVC)") has been established in March, 2003. [00073
Furthermore, as an extension of H.264/AVC, Fidelity Range Extension (FRExt) including an encoding tool necessary for professional use such as RGB or 4:2:2 or 4:4:4 or 8x8 DCT and a quantization matrix which are specified in MPEG-2 has been standardized in February, 2005. As a result, the H.264/AVC has become a coding scheme capable of also expressing a film noise included in a movie vjell and is being used in a wide range of applications such as a Blu-Ray Disc (a trademark), [0008]
Hov^ever, in recent years, there is an increasing need for further high compression rate coding capable of compressing an image (v^hich is also referred to as a "4K image") of about 4000 x 2000 pixels v/hich are four times as high as a high-definition image or delivering a high--definition image in a limited transmission capacity environment such as the Internet. To this end, an improvement in coding efficiency has been under continuous revievj by Video Coding Expert Group (VCEG) under ITU-T. [0009]
In this regard, currently, in order to further
improve coding efficiency to be higher than in the AVC, standardization of a coding scheme called High Efficiency Video Coding (HEVC) has been being conducted by Joint Collaboration Team-Video Coding (JCTVC) v/hich is a joint standardization organization of ITU-T and ISO/IEC. In the HEVC standard/ a committee draft that is a first draft specification has been issued in February, 2012
(see Non-Patent Document 1).
[0010]
In the case of the HEVC, generally, information that is transmitted from an encoding side to a decoding side includes syntax elements for P slices and B slices, that is, syntax elements related to an inter-screen process as v;ell as syntaxes for I slices .
[0011]
Meanv/hile, a still picture profile serving as a profile for using the HEVC as a still image codec has been proposed {for example, see Non-Patent Document 2).
[0012]
Since the still picture profile is a profile for encoding and decoding still images, v^hen this profile is applied, syntax elements related to the inter-screen process are unnecessary.
CITATION LIST
NON-PATENT DOCUMENT
[0013]
Non-Patent Document 1: Benjamin Bross, Woo-Jin Han, Jens-
Rainer Ohm, Gary J. Sullivan, Thomas Wiegand, "High
efficiency video coding (HEVC) text specification draft
8," JCTVC-H1003 d7. Joint Collaborative Team on Video
Coding (JCT-VC) of ITU-T SG16 WP3 and ISO/IEC JTC1/SC29/WG11, 10th Meeting: Stockholm, SE, 11-20 July 2012
Non-Patent Document 2: Kemal Ugur, Jani Lainema, Miska Hannuksela, "On still picture profile," JCTVC-J0037, Joint Collaborative Team on Video Coding (JCT-VC) of ITU-T SG16 WP3 and ISO/IEC JTC1/SC29/WG11, 10th Meeting: Stockholm, SE, 11-20 July 2012
SUMMARY OF THE INVENTION
PROBLEMS TO BE SOLVED BY THE INVENTION
[0014]
Hovjever, in the case of the method disclosed in Non-Patent Document 2, control on the syntax elements related to the inter-screen process is not performed. In other v/ords, similarly to the case of the moving image profile, the syntax elements related to the inter-screen process are transmitted from the encoding side to the decoding side. Thus, the coding efficiency is likely to be lov/ered since unnecessary information is transmitted. [0015]
The present disclosure vjas made in light of the foregoing, and it is desirable to be capable of suppressing a reduction in coding efficiency.
SOLUTIONS TO PROBLEMS [0016]
An image processing device according to an aspect of the present technology is an image processing device including a receiving unit that receives an encoded bitstream including a syntax element related to an inter-
image process and an analyzing unit that analyzes the syntax element received by the receiving unit in a state in which a value of the syntax element is restricted v/hen the bitstream is a bitstream encoded based on a profile for encoding a still image. [0017]
The image processing device may further include a decoding unit that decodes the bitstream using the syntax element analyzed by the analyzing unit. [0018]
The image processing device may further include a determining unit that determines v/hether or not the value of the syntax element received by the receiving unit is properly restricted when the bitstream is the bitstream encoded based on the profile for encoding the still image. [0019]
The image processing device may further include an abnormality processing unit that performs an abnormality process vjhen the determining unit determines that the value of the syntax element is not properly restricted. [0020]
The syntax element may be stored in a sequence parameter set of the bitstream. [0021]
The syntax element may be a syntax related to a reference picture. [0022]
The syntax element may be a syntax related to the number of reference pictures. [0023]
The image processing device may further include a
determining unit that determines whether or not the value of the syntax element received by the receiving unit is a predetermined value. [0024]
The predetermined value may be 0. [0025]
An image processing method according to an aspect of the present technology is an image processing method including receiving an encoded bitstream including a syntax element related to an inter-image process and analyzing the received syntax element in a state in v;hich a value of the syntax element is restricted v;hen the bitstream is a bitstream encoded based on a profile for encoding a still image. [0026]
Further, the bitstream may be decoded using the analyzed syntax element. [0027]
It may be determined v/hether or not the value of the received syntax element is properly restricted v/hen the bitstream is the bitstream encoded based on the profile for encoding the still image. [0028]
An abnormality process may be performed v/hen the value of the syntax element is determined to be not properly restricted. [0029]
The syntax element may be stored in a sequence parameter set of the bitstream. [0030]
The syntax element may be a syntax related to a
reference picture. [0031]
The syntax element may be a syntax related to the number of reference pictures. 5 [0032]
Further, it may be determined v/hether or not the value of the received syntax element is a predetermined value. [0033] 10 The predetermined value may be 0. [0034]
According to an aspect of the present technology, an encoded bitstream including a syntax element related to an inter-image process is received, and the received 15 syntax element is analyzed in a state in vihich a value of the syntax element is restricted v/hen the bitstream is a bitstream encoded based on a profile for encoding a still image.
2 0 EFFECTS OF THE INVENTION [0035]
According to the present disclosure, it is possible to encode and decode an image. Particularly, it is possible to suppress a reduction in coding efficiency. 25
BRIEF DESCRIPTION OF DRAWINGS [0036]
Fig. 1 is a diagram for describing an exemplary configuration of a coding unit, 30 Fig. 2 is a diagram for describing a tile.
Fig. 3 is a diagram illustrating an exemplary
profile tier level (Profile_tier_level).
Fig. 4 is a diagram illustrating an exemplary video parameter set (VPS).
Fig. 5 is a diagram illustrating an exemplary seguence parameter set (SPS).
Fig. 6 is a diagram illustrating an exemplary seguence parameter set (SPS)/ subseguently to Fig. 5.
Fig. 7 is a diagram illustrating another exemplary profile tier level (Profile_tier_level).
Fig. 8 is a diagram illustrating another exemplary video parameter set (VPS).
Fig. 9 is a diagram illustrating another exemplary seguence parameter set (SPS).
Fig. 10 is a diagram illustrating another exemplary seguence parameter set (SPS), subseguently to Fig. 9.
Fig. 11 is a diagram illustrating still another exemplary video parameter set (VPS).
Fig. 12 is a diagram illustrating still another exemplary seguence parameter set (SPS).
Fig. 13 is a diagram illustrating still another exemplary seguence parameter set (SPS), subseguently to Fig. 12.,
Fig. 14 is a diagram illustrating an exemplary picture parameter set (PPS).
Fig. 15 is a diagram illustrating an exemplary picture parameter set (PPS), subsequently to Fig. 14.
Fig. 16 is a diagram illustrating an exemplary scaling list data (scaling_list_data ()).
Fig. 17 is" a diagram illustrating an exemplary semantics of a size ID (SizelD).
Fig. 18 is a diagram illustrating an exemplary
semantics of a size ID (SizelD) and a matrix ID (MatrixID).
Fig. 19 is a diagram illustrating another exemplary scaling list data {scaling_list_data ()).
Fig. 20 is a diagram illustrating an exemplary slice header (slice_header ()).
Fig. 21 is a diagram illustrating an exemplary slice header (slice^header ()), subsequently to Fig. 20.
Fig. 22 is a diagram illustrating an exemplary slice header {slice_header ()), subsequently to Fig. 21.
Fig. 23 is a diagram illustrating an exemplary allocation of nal_unit_type.
Fig. 24 is a diagram for describing an exemplary Region of Interest (ROI) region.
Fig. 25 is a block diagram illustrating an exemplary main configuration of an image coding device.
Fig. 2 6 is a block diagram illustrating an exemplary main configuration of a lossless encoding unit.
Fig. 27 is a flov/chart for describing an exemplary f lov/ of a coding process .
Fig. 28 is a f lov/chart for describing an exemplary flov/ of a coding process, subsequently to Fig. 27-
Fig. 2 9 is a flov/chart for describing an exemplary flov? of a lossless coding process.
Fig. 30 is a flovjchart for describing an exemplary flov/ a syntax element setting process.
Fig. 31 is a block diagram illustrating another exemplary configuration of an image coding device.
Fig. 32 is a flowchart for describing another exemplary flow of a coding process.
Fig. 33 is a flovjchart for describing another
exemplary flov/ of a coding process, subsequently to Fig. 32.
Fig. 34 is a block diagram illustrating an exemplary main configuration of an image decoding device. 5 Fig. 35 is a block diagram illustrating an
exemplary main configuration of a lossless decoding unit.
Fig. 36 is a flowchart for describing an exemplary flow of a decoding process.
Fig. 37 is a flowchart for describing an exemplary 10 flow of a decoding process, subsequently to Fig. 36.
Fig. 38 is a flovjchart for describing an exemplary flov/ of a syntax element analysis process.
Fig. 39 is a block diagram illustrating another exemplary configuration of an image decoding device. 15 Fig. 40 is a block diagram illustrating another exemplary configuration of a lossless decoding unit.
Fig. 41 is a flov/chart for describing another exemplary flow of a decoding process.
Fig. 42 is a flowchart for describing another 20 exemplary flow of a decoding process, subsequently to Fig. 41.
Fig. 43 is a flowchart for describing an exemplary flovj of a syntax element inspection process.
Fig. 44 is a diagram illustrating an exemplary 25 multi-viev/ image coding scheme.
Fig. 45 is a diagram illustrating an exemplary main configuration of a multi-viev/ image coding device to vjhich the present technology is applied.
Fig. 4 6 is a diagram illustrating an exemplary main 30 configuration of a multi-viev/ image decoding device to v/hich the present technology is applied.
12
SP351163WO00
Fig. 47 is a diagram illustrating an exemplary scalable image coding scheme.
Fig. 48 is a diagram for describing an exemplary spatial scalable coding. 5 Fig. 4 9 is a diagram for describing an exemplary temporal scalable coding.
Fig. 50 is a diagram for describing an exemplary scalable coding of a signal to noise ratio.
Fig. 51 is a diagram illustrating an exemplary main 10 configuration of a scalable image coding device to v/hich the present technology is applied.
Fig. 52 is a diagram illustrating an exemplary main configuration of a scalable image decoding device to v/hich the present technology is applied. 15 Fig. 53 is a block diagram illustrating an exemplary main configuration of a computer.
Fig. 54 is a block diagram illustrating an exemplary schematic configuration of a television device.
Fig. 55 is a block diagram illustrating an 20 exemplary schematic configuration of a mobile telephone.
Fig. 56 is a block diagram illustrating an exemplary schematic configuration of a recording/reproducing device.
Fig. 57 is a block diagram illustrating an 25 exemplary schematic configuration of an imaging device.
Fig. 58 is a block diagram illustrating a utilization example of scalable coding.
Fig. 59 is a block diagram illustrating another utilization example of scalable coding. 30 Fig. 60 is a block diagram illustrating still another utilization example of scalable coding.
13
SP351163WO00
Fig. 61 is a block diagram illustrating an exemplary schematic configuration of a video set.
Fig. 62 is a block diagram illustrating an exemplary schematic configuration of a video processor. 5 Fig. 63 is a block diagram illustrating another exemplary schematic configuration of a video processor.
Fig. 64 is an explanatory diagram illustrating a configuration of a content reproducing system.
Fig. 65 is an explanatory diagram illustrating a 10 data flovi in a content reproducing system.
Fig. 66 is an explanatory diagram illustrating a specific example of MPD.
Fig. 67 is a functional block diagram illustrating a configuration of a content server of a content 15 reproducing system.
Fig. 68 is a functional block diagram illustrating a configuration of a content reproducing device of a content reproducing system.
Fig. 69 is a functional block diagram illustrating 20 a configuration of a content server of a content reproducing system.
Fig. 70 is a sequence chart illustrating an exemplary communication process performed by devices of a v/ireless communication system. 25 Fig. 71 is a sequence chart illustrating an
exemplary communication process performed by devices of a wireless communication system.
Fig. 72 is a diagram schematically illustrating an exemplary configuration of a frame format transceived in 30 a communication process performed by devices of a v/ireless communication system.
J
14
SP351163WO00
Fig. 73 is a sequence chart illustrating an exemplary communication process performed by devices of a wireless communication system.
5 MODE FOR CARRYING OUT THE INVENTION [0037]
Hereinafter, modes (hereinafter, referred to as "embodiments") of carrying out the present disclosure v/ill be described. The description will proceed in the 10 follov/ing order.
0. Overview
1. First embodiment (image coding device)
2. Second embodiment (image coding device)
3. Third embodiment (image decoding device)
15 4. Fourth embodiment (image decoding device)
5. Fifth embodiment (multi-viev/ image coding device and multi-viev/ image decoding device)
6. Sixth embodiment (scalable image coding device and scalable image decoding device)
20 7. Seventh embodiment (computer)
8. Application examples
9. Application examples of scalable coding
10. Eighth embodiment (set/unit/module/processor)
11. Ninth embodiment (application examples of 25 content reproducing system of MPEG-DASH)
12. Tenth embodiment (application examples of
wireless communication system of Wi-Fi standard)
[0038]
<0. Overview> 30
Hereinafter, the present technology v;ill be
1
15
SP351163WO00
described in connection v/ith an application to image
coding and decoding of a HEVC scheme.
[0039]
5 In the AVC, a hierarchical structure based on a macroblock and a sub macroblock is defined. However, a macroblock of 16 x 16 pixels is not optimal for a large image frame such as an Ultra High Definition (UHD) (4000
X 2000 pixels) serving as a target of a next generation 10 coding scheme. [0040]
On the other hand, in the HEVC, a coding unit (CU) is defined as illustrated in Fig. 1. [0041] 15 A CU is also referred to as a coding tree block (CTB), and the CU is a partial area of an image of a picture unit undertaking the same role of a macroblock in the AVC. The macroblock of the AVC is fixed to a size of 16 x 16 pixels, but a size of the CU of the HEVC is not 20 fixed and designated in image compression information in each s equence. [0042]
For example^ a largest coding unit (LCU) and a smallest coding unit (SOU) of a CU are specified in a 25 sequence parameter set (SPS) included in encoded data to be output. [0043]
As split_flag = 1 is set in a range in which each LCU is not smaller than a SCU, a coding unit can be 30 divided into CUs having a smaller size. In the example of Fig. 1, a size of an LCU is 128 x 128, and a largest
16
SP351163WO00
scalable depth is 5. A CU of a size of 2N x 2N is divided into CUs having a size of N x N serving as the hierarchy that is one-level lov/er v;hen a value of split^flag is "1." 5 [0044]
Further, a CU is divided in prediction units (PUs) that are areas (partial areas of an image of a picture unit) serving as processing units of intra or inter prediction, and divided into transform units (TUs) that
10 are areas (partial areas of an image of a picture unit) serving as processing units of orthogonal transform. In the HEVC, any of 4 x 4, 8x8, 16 x 16, and 32 x 32 can be used as a processing unit of orthogonal transform. [0045]
15 In the case of the coding scheme in v/hich a CU is defined, and various kinds of processes are performed in units of CUs such as the HEVC, a macroblock in the AVC can be considered to correspond to an LCU, and a block (sub block) can be considered to correspond to a CU.
20 Further, a motion compensation block in the AVC can be
considered to correspond to a PU. Here, since a CU has a hierarchical structure, a size of an LCU of a topmost layer is commonly set to be larger than a macroblock in the AVC, for example, such as 128 x 128 pixels.
25 [0046]
Thus, hereinafter, an LCU is assumed to include a macroblock in the AVC, and a CU is assumed to include a block (sub block) in the AVC. In other v/ords, a "block" used in the follov/ing description indicates an arbitrary
30 partial area in a picture, and, for example, a size,
shape, and characteristics of a block are not limited.
17
SP351163WO00
In other words, a "block" includes an arbitrary area (a processing unit) such as a TU, a PU, an SCU, a CU, an LCU, a sub block, a macroblock, or a slice. Of course, a "block" includes any other partial area (processing unit) 5 as well. When it is necessary to limit a size, a
processing unit, or the like, it v/ill be appropriately
described.
[0047]
10 Meanv/hile, in the coding schemes such as the AVC and the HEVC, in order to achieve higher coding efficiency, it is important to select an appropriate prediction mode. [0048]
15 As an example of such a selection method, there is a method implemented in reference softv/are (opened at http://iphome.hhi.de/suehring/tml/index.htm) of H.264/MPEG-4 AVC called a joint model (JM), [0049]
20 In the JM, it is possible to select tv/o mode
determination methods, that is, a high complexity mode and a lov; complexity mode described below. In both modes, cost function values related to the respective prediction modes are calculated, and a prediction mode having a
25 smallest cost function value is selected as an optimal
mode for a current block serving as a processing target. [0050]
A cost function in the high complexity mode is represented as in the following Formula (1):
30 [0051]
[Mathematical Formula 1]
SP351163WO00
Cost(Mode e H) = D + X*R ... (1)
[0052]
Here^ Q, indicates a universal set of a candidate mode for encoding a current blocks and D indicates 5 differential energy betv/een a decoded image and an input image when encoding is performed in a corresponding prediction mode. X indicates Lagrange's undetermined multiplier given as a function of a quantization parameter. R indicates a total coding amount including 10 an orthogonal transform coefficient v/hen encoding is performed in a corresponding prediction mode. [0053]
In other viords, in order to perform encoding in the high complexity mode, it is necessary to perform a 15 temporary encoding process once by all candidate modes in order to calculate the parameters D and R, and thus a larger computation amount is required. [0054]
A cost function in the low complexity mode is 20 represented by the follov/ing Formula (2) : [0055] [Mathematical Formula 2]
Cost(Mode e Q) = D + QP2Quant(QP)*HeaderBit ... (2) [0056]
25 Here, D indicates differential energy betv/een a predicted image and an input image unlike the high complexity mode. QP2Quant(QP) is given as a function of a quantization parameter QP, and HeaderBit indicates a coding amount related to information belonging to a
30 header such as a motion vector or a mode including no orthogonal transform coefficient.
i
J
19
SP351163WO00
[0057]
In other v/ords, in the lov; complexity mode, it is necessary to perform a prediction process for respective â– candidate modes, but since up to a decoded image is not 5 necessary, it is unnecessary to perform up to a coding process. Thus, it can be implemented with a computation amount smaller than that in the high complexity mode. [0058]
10 Meanv/hile, in the HEVC, a tile illustrated in Fig. 2 is specified as a unit of a parallel process in addition to a slice specified in the AVC. [0059]
A v/idth and a height of each tile are designated in 15 image compression information, and a decoding process can be independently performed on each tile. [0060]
Further, in Non-Patent Document 2, the still 20 picture profile serving as the profile for using the HEVC as the still image codec is proposed. [00611
Hov^ever, in the case of the HEVC, generally, information that is transmitted from an encoding side to 25 a decoding side includes syntax elements for P slices and B slices, that is, syntax elements related to an inter-screen process as v;ell as syntaxes for I slices. [0062]
Since the still picture profile is the profile for 30 encoding and decoding still images, when this profile is applied, the above-described syntax elements related to
4
20
SP351163WO00
the inter-screen process are unnecessary. Hov/ever, in the case of the method disclosed in Non-Patent Document 2, control on the syntax elements related to the inter-screen process is not performed. In other v/ords, 5 similarly to the case of the moving image profile, the syntax elements related to the inter-screen process are transmitted from the encoding side to the decoding side. Thus, the coding efficiency is likely to be lov/ered since unnecessary information is transmitted. 10 [0063]
In this regard, in the present technology, v/hen the coding process is performed based on the profile for encoding the still images, values of syntax elements related to the inter-image process are restricted, and 15 transmission of unnecessary information is suppressed. As a result, it is possible to suppress a reduction in coding efficiency. [0064]
Specific examples of the restriction v;ill be 20 described belov;. [0065]
<0-l: Restriction of syntax elem.ent related to sub layer>
Fig. 3 is a diagram illustrating an exemplary 25 syntax of a profile tier level (profile_tier_level()) in the HEVC. Numbers at the left end are given as line numbers for description and not included in an actual syntax. Similarly, numbers at the left end illustrated in Figs. 4 to 23 v;hich v;ill be described belov/ are given 30 as line numbers for description of figures and not actually included in a syntax.
^
21
SP351163WO00
[0066]
A syntax element general_profile_idc shov/n in a 5-th line in the profile tier level (profile_tier_level()) illustrated in Fig. 3 specifies that a profile of a 5 sequence (current sequence) of a processing target is a still picture profile. [00673
The profile tier level (profile_tier_level {)} of Fig. 3 is called by a "video parameter set (VPS)" or a 10 "sequence parameter set (SPS)." [00683
Fig. 4 is a diagram illustrating an exemplary video parameter set (VPS) in the HEVC. Fiqs. 5 and 6 are diagrams illustrating an exemplary sequence parameter set 15 (SPS) in the HEVC. [00693
As illustrated in Fig. 4, in the video parameter set (VPS), a profile tier level {profile_tier_level{)) in a 7-th line is called. Further, as illustrated in Fig. 5, 20 in the sequence parameter set (SPS), the profile tier level {profile_tier__level () ) in a 5-th line is called. [00703
Here, when encoding is performed based on the still picture profile, there is no temporal layer depth (which 25 is also referred to as a "sub layer"). In other words,
syntax elements related to the sub layer are unnecessary. [0071]
In this regard, in the profile tier level (profile_tier_level{)) of Fig. 3, before the still 30 picture profile is specified by the syntax element
general profile idc, 0 may be designated as a value of a
22
SP351163WO00
parameter vps_max_sub_layers__minusl (a 6-th line) related
to the sub layer in the video parameter set (VPS) of Fig.
4, and 0 may be designated as a value of a parameter
sps_max_sub_layers_minusl (a 3-rd line) related to the 5 sub layer in the sequence parameter set (SPS) (Figs. 5
and 6), t0072]
In other words, v/hen the syntax element
general__profile_idc in the profile tier level 10 (prof ile_tier__level () ) of Fig. 3 specifies that the
profile is the still picture profile, 0 may have to be
designated as the value of the parameter
vps_max_sub_layers_minusl (the 6~th line) related to the
sub layer in the video parameter set (VPS) of Fig. 4, and 15 0 may have to be designated as the value of the parameter
sps_max__sub_layers_minusl (the 3-rd line) related to the
sub layer in the sequence parameter set (SPS) (Figs. 5
and 6).
[0073] 20 As a result, it is possible to prevent unnecessary
portions of the profile tier level (profile_tier__level () )
from being read. In other v/ords, it is possible to
prevent an increase in a Joad caused by the reading and
prevent reading and transmission of unnecessary 25 parameters. Thus, it is possible to suppress a reduction
in coding efficiency.
[0074]
Further, as a result, it is unnecessary to change
the syntaxes of the profile tier level 30 (prof ile_tier__level 0 ) , the video parameter set (VPS),
and the sequence parameter set (SPS), and it is possible
23
SP351163WO00
to suppress a reduction in coding efficiency through control by semantics. When the syntax is changed, for example, it is likely to be difficult to maintain syntax compatibility v/ith an encoder and a decoder of a related 5 art that do not support the still picture profile. Particularly, in the case of encoders and decoders implemented by hardv/are, there are cases in v?hich it is difficult to update the syntax. A decrease in syntax compatibility is likely to reduce versatility. Hovjever,
10 as described above, v/hen the values of the syntax element are restrigtQcl by the gemanticg/- it is possible to. maintain syntax compatibility and prevent a reduction in versatility. [0075]
15 Further, since the syntax compatibility is
maintained as described above, it is possible to easily apply a common syntax even to both encoding of still images and encoding of moving images, and thus it is possible to easily implement an encoder and a decoder
20 that process both a still image and a moving image through a common circuit. In other words, it can contribute to a size reduction of a device, suppression of an increase in cost, and the like. [0076]
25 <0-2: Restriction of syntax element related to calling of profile tier level>
Here, it is not impossible to implement such a restriction by changing the syntax. [0077]
30 As described above, when the profile tier level
(profile_tier_level()) is called from the video parameter
24
SP351163WO00
set (VPS) or the sequence parameter set (SPS), a value of a syntax element ProfilePresentFlag related to the calling of the profile tier level v/hich is designated at the time of the calling is consistently 1. 5 [0078]
In other v/ords, it is redundant to transmit this syntax element. Further, an if statement of a 2-nd line of the profile tier level (profile_tier_level()) of (Fig. 3) is unnecessary as well.
10 [0079]
In this regard, the syntax of the profile tier level (profile_tier_level0) of Fig. 3 may be changed as in an example illustrated in Fig. 7, the syntax of the video parameter set (VPS) of Fig. 4 may be changed as in
15 an example illustrated in Fig, 8, and the syntax of the seguence parameter set (SPS) of Figs. 5 and 6 may be changed as in an example illustrated in Figs. 9 and 10. [0080]
In other words, the profile tier level
20 (profile_tier_level0) may be designated by designating only a syntax element MaxNumSubLayersMinusl related to the sub layer as in the examples of Fig. 7 (a 1-st line), Fig. 8 (a 7-th line), and Fig. 9 (a 5-th line) without designating the syntax element ProfilePresentFlag.
25 [0081]
Further, as illustrated in the 1-st to 7-th lines of Fig. 7, in the profile tier level
(profile_tier_level()), the if statement using the syntax element ProfilePresentFlag may be omitted.
30 [0082]
As a result, it is possible to prevent transmission
25
SP351163WO00
of an unnecessary parameter and suppress a reduction in coding efficiency. Further, it is possible to suppress an increase in the load of the process of reading the profile tier level {prof ile_tier_level () } v/hich is caused 5 by the reading of the unnecessary if statement. [0083]
In other v^ords, the value of the syntax element ProfilePresentFlag related to the calling of the profile tier level may have to be fixed to 1.
10 [0084]
<0-3: Restriction of syntax element related to profile tier level>
In the above method, when the syntax element max_sub_layers_minusl related to the sub layer is encoded,
15 a setting has to be performed after detecting information as to whether or not encoding is performed based on the still picture profile in a subsequent profile tier level (prof ile__tier_level () ) . [0085]
20 In this regard, the syntax of the video parameter set (VPS) of Fig. 4 may be changed as in an example illustrated in Fig. 11/ and the syntax of the sequence parameter set (SPS) of Figs. 5 and 6 may be changed in an example illustrated in Figs. 12 and 13.
25 [0086]
In other v/ords, in the video parameter set (VPS), as in 6-th to 8-th lines {Fig. 11}, a syntax element profile_tier_level(1,0} related to the profile tier level may be designated, a value of a syntax element
30 vps_max_sub__layers_minusl related to the sub layer may be designated, and a syntax element profile tier level
26
SP351163WO00
(0,vps_max_sub_layers_minusl) related to the profile tier
level may be designated.
[0087]
Similarly, in the sequence parameter set (SPS), as 5 in 3-rd to 6-th lines (Fig. 12), the syntax element
profile_tier_level(1,0) related to the profile tier level may be designated, a value of a syntax element sps_max_sub_layers_minusl related to the sub layer may be designated, and the syntax element profile_tier_level 10 (0,sps_max_sub_layers_minusl} related to the profile tier level may be designated. [0088]
Further, for example, when a parameter set to v/hich a syntax element belongs is not discriminated unlike 15 vps_max_sub_layers_minusl and sps_max_sub_layers_minusl, the parameter sets are simply referred to as max__sub_layers_minusl. In other v/ords,
vps_max_sub_layers_minusl is max_sub_layers_minusl in the video parameter set (VPS) , and sps__max__sub_layers_minusl 20 is max_sub_layers_minusl in the sequence parameter set (SPS) . Other syntax elements are assumed to have a similar relation. [0089]
As the syntax element profile_tier_level(1,0} is 25 designated before the syntax element
max_sub_layers_minusl related to the sub layer is designated, information related to v/hen all temporal layers (temporal^layer) of image data are encoded or decoded is transmitted. 30 [0090]
Here, when the still picture profile is applied, 0
27
SP351163WO00
is designated as the value of the syntax element max_sub_layers_minusl related to the sub layer. [0091]
After the syntax element max__sub_layers_minusl 5 related to the sub layer is encoded/ information related to when some temporal layers (temporal_layer) of image data are encoded or decoded is transmitted through profile_tier_level(0,nax_sub_layers_mxnusl). [0092] 10 By changing the syntax as described above/ it is possible to designate the value of the syntax element max_sub_layers_minusl related to the sub layer after it is designated v/hether or not the profile is the still picture profile. 15 [0093]
<0-4: Restriction of syntax element related to virtual reference decoder>
Further/ v/hen the still picture profile is applied, it is unnecessary to control the virtual reference 20 decoder. In this regard/ as illustrated in Fig. 4/ 0 may be designated as a value of a syntax element vps_num_hrd_parameters (a 14-th line) related to the virtual reference decoder in the video parameter set (VPS). 25 [0094]
In other vrords/ v?hen the still picture profile is applied/ the value of the syntax element vps_num_hrd_parameters (the 14-th line) related to the virtual reference decoder in the video parameter set 30 (VPS) may have to be fixed to 0. [0095]
i
28
SP351163WO00
The syntax element vps_num_hrd_parameters related to the virtual reference decoder is a positive value that is encoded by an extended Golomb coding and then transmitted. Thus, when this value is 0 or 1, the coding 5 amount becomes minimum. In other words, as 0 is designated as the value of the syntax element vps_num_hrd_parameters related to the virtual reference decoder, it is possible to suppress a reduction in coding efficiency. 10 [0096]
Further, as 0 is designated as the value of the syntax element vps_num_hrd_parameters related to the virtual reference decoder, it is possible to skip a loop process (a for statement) of 15-th to 19-th lines. Thus, 15 as a result, it is possible to suppress an increase in a load. Further, it is possible to prevent transmission of the unnecessary syntax elements in the loop process (the for statement) and suppress a reduction in coding efficiency. 20 [0097]
Further, it is possible to restrict the value of the syntax element by the semantics v/ithout changing the syntax and suppress a reduction in syntax versatility. [0098] 25 Further, this restriction can be applied even v/hen the video parameter set (VPS) is the example of Fig. 8 or the example of Fig. 11. [0099]
<0-5: Restriction of syntax element related to P 30 slice and B slice>
Further, v/hen the still picture profile is applied,
29
SP351163WO00
there is neither a P slice nor a B slice. In this regard, as illustrated in Fig. 5/ 0 may be designated as a value of a syntax element restricted_ref_pic_lists_flag {a 32-nd line) related to the P slice and the B slice in the 5 sequence parameter set (SPS) . [0100]
In other words, when the still picture profile is applied/ the value of the syntax element restricted_ref__pic_lists_flag {the 32-nd line) related to
10 the P slice and the B slice in the sequence parameter set (SPS) may have to be fixed to 0. [0101]
Further/ as 0 is designated as the value of the syntax element restricted_ref_pic_lists_flag related to
15 the P slice and the B slice/ it is possible to skip 33-rd to 42-nd lines. ThuS/ as a result/ it is possible to suppress an increase in a load. Further/ it is possible to prevent transmission of unnecessary syntax elements of the 33-rd to 42-nd lines and suppress a reduction in
20 coding efficiency. [0102]
Further/ it is possible to restrict the value of the syntax element by the semantics without changing the syntax and suppress a reduction in syntax versatility.
25 [0103]
Further/ the restriction can be applied even v/hen the sequence parameter set (SPS) is the example of Figs. 9 and 10 or the example of Figs. 12 and 13. [0104]
30 <0-6: Restriction of syntax element related to short term>
30
SP351163WO00
Further, v/hen the still picture profile is applied,
there is no concept of time (there is no other picture).
In this regard, as illustrated in Fig. 6, 0 may be
designated as a value of a syntax element 5 num_short_term_ref_pic_sets (a 56-th line) related to the
short term in the sequence parameter set (SPS).
[0105]
In other words, when the still picture profile is
applied, the value of the syntax element 10 num_short_term_ref_pic_sets (the 56-th line) related to
the short term in the sequence parameter set (SPS) may
have to be fixed to 0.
[0106]
The syntax element num_short_term_ref_pic_sets 15 related to the short term is a positive value that is
encoded by extended Golomb coding and then transmitted.
Thus, when the value is 0 or 1, the coding amount becomes
minimum. In other words, as 0 is designated as the value
of the syntax element num_short_term_ref_pic_^sets related 20 to the short term, it is possible to suppress a reduction
in coding efficiency.
[0107]
Further, as 0 is designated as the value of the
syntax element num_short_term_ref_pic_sets related to the 25 short term, it is possible to skip 57-th and 58-th lines.
Thus, as a result, it is possible to suppress an increase
in a load.
[0108]
Further, it is possible to restrict the value of 30 the syntax element by the semantics v/ithout changing the
syntax and suppress a reduction in syntax versatility.
^
31
SP351163WO00
[0109]
Further, this restriction can be applied even v/hen
the sequence parameter set (SPS) is the example of Figs.
9 and 10 or the example of Figs. 12 and 13. 5 [0110]
<0-7: Restriction of syntax element related to long
term>
Further, when the still picture profile is applied,
there is no concept of time (there is no other picture). 10 In this regard, as illustrated in Fig. 6, 0 may be
designated as a value of a syntax element
long_term_ref_pics_present__flag {a 59-th line) related to
the long term in the sequence parameter set (SPS).
[Olllj 15 In other words, when the still picture profile is
applied, the value of the syntax element
long_term_ref_pics_present_flag (the 59-th line) related
to the long term in the seguence parameter set (SPS) may
have to be fixed to 0. 20 [0112]
As 0 is designated as the value of the syntax
element long_term._ref_pics_present_flag related to the
long term, it is possible to skip 60-th to 66-th lines.
Thus, as a result, it is possible to suppress an increase 25 in a load. Further, it is possible to prevent
transmission of an unnecessary syntax element in the 60-th to 66-th lines and suppress a reduction in coding
efficiency.
[0113] 30 Further, it is possible to restrict the value of
the syntax element by the semantics v/ithout changing the
^
32
SP351163WO00
syntax and suppress a reduction in syntax versatility. [0114]
Further, this restriction can be applied even when the sequence parameter set (SPS) is the example of Figs. 5 9 and 10 or the example of Figs. 12 and 13. [0115]
<0-8: Restriction of syntax element related to motion vector>
Further, v/hen the still picture profile is applied, 10 there is no concept of time {there is no other picture). In this regard, as illustrated in Fig. 6, 0 may be designated as a value of a syntax element sps_temporal_mvp_enable_flag (a 67-th line) related to the motion vector in the sequence parameter set (SPS). 15 [0116]
In other words, v/hen the still picture profile is applied, the value of the syntax element
sps_temporal_mvp__enable__flag (the 67-th line) related to the motion vector in the sequence parameter set (SPS) may 20 have to be fixed to 0. [0117]
As 0 is designated as the value of the syntax element sp3_temporal__mvp_enablc__flag related to the motion vector, it is possible to minimize the coding 25 amount, and it is possible to suppress a reduction in coding efficiency. [0118]
Further, it is possible to restrict the value of the syntax element by the semantics without changing the 30 syntax and suppress a reduction in syntax versatility. [0119]
33
SP351163WO00
Further, this restriction can be applied even when the sequence parameter set (SPS) is the example of Figs. 9 and 10 or the example of Figs. 12 and 13. [0120] 5 <0-9: Restriction of syntax element of picture parameter set>
Figs. 14 and 15 are diagrams illustrating an exemplary picture parameter set (PPS) in the HEVC. [0121] 10 When the still picture profile is applied, there is no concept of time (there is no other picture}. In this regard, as illustrated in Fig. 14, 0 or 1 may be designated as both of values of the syntax element num_ref_xdx_10_default_active minusl (a 6-th line) 15 related to LO and a syntax element
num_ref_idx_ll_default_active minusl (a 7-th line) related to LI in the picture parameter set (PPS). [0122]
In other words, v/hen the still picture profile is 20 applied, both of the values of the syntax element
num _ref_idx_10_default_active minusl (the 6-th line) related to T.O and the syntax element
num_ref_idx_ll_defaultractive minusl (the 7-th line) related to LI in the picture parameter set (PPS) may have 25 to be fixed to 0 or 1. [0123]
The syntax elements are positive values that are encoded by extended Golomb coding and then transmitted. Thus, v/hen the values are 0 or 1, the coding amount 30 becomes minimum. In other v/ords, as 0 is designated as the values of the syntax element
34
SP351163WO00
num_ref__idx__10_default_active minusl related to LO and the syntax element num_ref_idx_ll_default_active minusl related to LI, it is possible to suppress a reduction in coding efficiency. 5 [0124]
Further, v;hen the still picture profile is applied, there is no concept of time (there is no other picture). In this regard, as illustrated in Fig. 15, 0 may be designated as a value of a syntax element (flag)
10 lists_modification__present_flag (a 49-th line) of the picture parameter set (PPS) indicating v/hether or not there is a syntax element ref_pic_list_modification in a current slice header. [0125]
15 In other v/ords, when the still picture profile is applied, the value of the syntax element (flag) lists_modification_present_flag (the 49-th line) of the picture parameter set (PPS) indicating whether or not there is a syntax element ref_pic_list_modification in
20 the current slice header may have to be fixed to 0. [0126]
When the value of the syntax element lists_modification_present_flag is 1, the syntax element ref_pic_li3t_modification related to the reference image
25 list is transmitted for the current slice as illustrated in 53-rd and 54-th lines of Fig. 21, but in the case of the still picture profile, this syntax element is unnecessary. In other v/ords, as 0 is designated as the value of the syntax element
30 lists_modification_present_flag, it is possible to omit transmission of the syntax element
-1. 35
J SP351163WO00
ref__pic_list_modification of the reference image list that is unnecessary information for the current slice, and it is possible to suppress a reduction in coding efficiency. 5 [0127]
Further, v/hen the still picture profile is applied/ there is no concept of time (there is no other picture). In this regard, as illustrated in Fig. 15, 0 may be designated as a value of a syntax element
10 log2_parallel_merge_level_minus2 (a 50-th line) of the picture parameter set (PPS) designating a parallel process level of a merge mode and a skip mode in a prediction process. [0128]
15 In other v/ords, v^hen the still picture profile is applied/ the value of the syntax element
log2_parallel_merge_level_minus2 (the 50-th line) of the picture parameter set (PPS) designating the parallel process level of the merge mode and the skip mode in the
20 prediction process may have to be fixed to 0. [0129]
This syntax elem.ent is a positive value that is encoded by extended Golomb coding and then transmitted. Thus, v/hen this value is 0, the coding amount becomes
25 minimum. In other v/ords, as 0 is designated as the value of the syntax element log2_parallel_merge_level_minus2 designating the parallel process level of the merge mode and the skip mode in the prediction process, it is possible to suppress a reduction in coding efficiency.
30 [0130]
Further, vjhen the still picture profile is applied,
36
SP351163WO00
there is no concept of time {there is no other picture). In this regard, as illustrated in Fig. 14, 0 may be designated as both of values of a syntax element (flag) v/eighted_pred_flag (a 18-th line) related to the v/eighted 5 prediction of the P slice and a syntax element (flag) v/eighted_bipred_flag {a 19-th line) related to the v;eighted prediction of the B slice in the picture parameter set (PPS). [0131]
10 In other words, v/hen the still picture profile is applied, the values of the syntax element (flag) weighted_pred_flag {the 18-th line) related to the v/eighted prediction of the P slice and the syntax element (flag) weighted_bipred_flag (the 19-th line) related to
15 the weighted prediction of the B slice in the picture parameter set (PPS) may have to be fixed to 0. [0132]
When the value of the syntax element v/eighted__pred_flag or v/eighted_bipred_flag is 1, as
20 illustrated in 65-th to 68-th lines of Fig. 21, the
v/eighted prediction table or the like is transmitted for the current slice, but in the case of the still picture profile, this information is unnecessary. In other v/ords, as 0 is designated as the value of the syntax element
25 v/eighted_pred_flag or v/eighted_bipred__flag, it is
possible to omit transmission of unnecessary information for the current slice, and it is possible to suppress a reduction in coding efficiency. [0133]
30 Here, as v/ill be described later, even v/hen the
slice type is restricted to the I slice, transmission of
37
SP351163WO00
the information can be omitted, and thus the above-described restriction related to the syntax element weighted_pred_flag or v/eighted_bipred_flag can be omitted [0134] 5 Further, it is possible to restrict the value of the syntax element by the semantics without changing the syntax and suppress a reduction in syntax versatility. [0135]
<0-10; Restriction of syntax element related to 10 prediction mode of scaling list>
Meanv/hile, in the sequence parameter set (SPS) illustrated in Figs. 5 and 6, a scaling list (scaling_list_data ()) is called (a 49-th line). Similarly, even in the picture parameter set (PPS) 15 illustrated in Figs. 14 and 15, the scaling list (scaling_list_data ()) is called (a 48-th line). [0136]
Fig. 16 illustrates an exemplary scaling list (scaling_list_data ()). In the scaling list 20 (scaling_list_data {)) illustrated in Fig. 16, semantics of a size ID (sizelD) and a matrix ID (matrixID) are as illustrated in Figs. 17 and 18. [0137]
Here, when the value of the size ID (sizelD) is 0, 25 1, or 2, and the value of the matrix ID (matrixID) is 3, 4, or 5 or when the value of the size ID (sizelD) is 3, and the value of the matrix ID (matrixID) is 1, in the still picture profile, it is redundant to transmit the scaling list. 30 [0138]
In this regard, in this case, 0 may be designated
^
SP351163WO00
as a value of a syntax element
scaling_list_pred__mode_flag related to the prediction
mode of the scaling list, and 0 or 1 may be designated as
a value of a syntax element 5 scaling_list_pred__matrix_id_delta related to a prediction
matrix of the scaling list. [0139]
In other v/ords, when the still picture profile is
applied, the value of the syntax element 10 scaling__list_pred_mode_flag (a 4-th line) related to the
prediction mode of the scaling list may have to be fixed
to 0, and the value of the syntax element
scaling_list__pred_matrix_id_delta {a 6-th line) related
to the prediction matrix of the scaling list may have to 15 be fixed to 0 or 1.
[0140]
Further, as 1 is designated as the value of the
syntax element scaling__list_pred_matrix_id_delta related
to the prediction matrix of the scaling list, it is 20 possible to reduce the necessity of inserting start code
emulation prevention while keeping a code length to a
minimum.
[0141]
<0-ll: Change of syntax related to prediction mode 25 of scaling list>
Further, instead of restricting the value of the
syntax element by the semantics as described above, the
syntax may be changed as illustrated in Fig. 19.
[0142] 30 In the syntax of Fig. 19, as illustrated in a 4-th
line, the condition of the for statement is designated in
i
39
SP351163WO00
detail using the size ID (sizelD) and the matrix ID
(matrixID) as described above.
[0143]
As a result, the same effects can be obtained. 5 [0144]
<0-12: Restriction of syntax element related to slice type>
Figs. 20 to 22 illustrate an exemplary syntax of a slice header. 10 [0145]
When the still picture profile is applied, the slice type is the I slice. In this regard, as illustrated in Fig. 20, a value indicating the I slice may be designated as a value of a syntax element 15 slice_type {an 11-th line) related to the slice type in the slice header (slice_header{)). [0146]
In other v/ords, v/hen the still picture profile is applied, the value of the syntax element slice_type (the 20 11-th line) related to the slice type in the slice header (slice_header{)) may have to be fixed to the I slice. [0147]
As a result, it is possible to slcip 44-th to 68-th lines of the slice header (slice_header()). Thus, as a 25 result, it is possible to suppress an increase in a load. Further, it is possible to prevent transmission of unnecessary syntax elements of the 44-th to 68-th lines, and it is possible to suppress a reduction in coding efficiency. 30 [0148]
Further, it is possible to restrict the value of
40
SP351163WO00
the syntax element by the semantics without changing the syntax and suppress a reduction in syntax versatility. [0149]
<0-13: Restriction of syntax element related to NAL 5 unit type>
Meanv/hile/ in Figs. 20 to 22, IdrPicFlag and RapPicFlag are calculated based on nal_unit_type as in the following Formulas (3) and (4): [0150] 10 [Mathematical Formula 3]
IdrPicFlag = (nal_unit_type -^= IDR_W_LP | | nal_unit_type == IDR_N_LP) ... (3)
RapPicFlag ^ (nal_unit_type >= 7 && nal_unit_type <= 12) ... (4) 15 [0151]
Here, a syntax element nal_unit_type related to the NAL unit type is allocated as illustrated in Fig. 23. [0152]
In this regard/ v/hen the still picture profile is 20 applied, IDR_W_LP or IDR_N_LP may be designated as the NAL unit type (nal_unit_type) for the VCL. [0153]
In other v/ords, v/hen the still picture profile is applied, the NAL unit type (nal_unit_type) for the VCL 25 may have to be fixed to IDR_W_LP or IDR_N_LP. [0154]
By performing the above process^ it is possible to prevent transmission of redundant information and improve the coding efficiency of image compression information to 30 be output when encoding is performed based on the still picture profile.
1
41
SP351163WO00
[0155]
Further, the syntax elements related to the inter-image process in v;hich the values are restricted when the coding process is performed based on the profile for 5 encoding the still image is not limited to the above example. The values of the syntax elements other than the above-described syntax elements may be restricted. At this time, the syntax may be changed by restricting the semantics.
10 [0156]
Further, as illustrated in Fig. 24, v/hen there is a Region of Interest (ROI) in a part of a still image, for example, v/hen a person's face is shown in a part of a still image, the image may be divided into tiles of Fig,
15 2, a tile serving as a ROI region may be designated by metadata such as Supplemental Enhancement Information (SEI), and the decoding process of only the ROI region may be performed for a decoded image. [0157]
20 Next/ application examples of the present
technology to specific devices v/ill be described. [0158]
<1. First embodiment>
25 Fig. 25 is a blocJc diagram illustrating an
exemplary configuration of an image coding device as an example of an image processing device to v/hich the present technology is applied. For example, an image coding device 100 illustrated in Fig. 25 encodes image
30 data of a moving image using the prediction process of
the HEVC or a prediction process of a scheme based on the
42
SP351163WO00
HEVC. [0159]
Further, the image coding device 100 can encode image data of a still image as v/ell as image data of a 5 moving image. In this case, a still image profile (for example, the still picture profile) is set to the image coding device 100. [0160]
As illustrated in Fig. 25, the image coding device 10 100 includes an A/D converter 101, a screen rearrangement buffer 102, an operation unit 103, an orthogonal transform unit 104, a quantization unit 105, a lossless encoding unit 106, an accumulation buffer 107, an inverse quantization unit 108, and an inverse orthogonal 15 transform unit 109. The image coding device 100 further includes an operation unit 110, a loop filter 111, a frame memory 112, an intra prediction unit 113, an inter prediction unit 114, a predicted image selecting unit 115, and a rate control unit 116. 20 [0161]
The A/D converter 101 performs A/D conversion on image data (analog data) of an input image input to the image coding device 100. The A/D converter 101 supplies the converted image data (digital data) to the screen 25 rearrangement buffer 102. [0162]
The screen rearrangement buffer 102 stores each frame image data of the input image supplied in a display order in the order. The screen rearrangement buffer 102 30 rearranges the order of the frames of the input image in an encoding order according to a Group Of Picture (GOP).
4
43
SP351163WO00
In other v/ords, the screen rearrangement buffer 102 reads the image data of the frames stored in the display order in the encoding order. The screen rearrangement buffer
102 supplies the read image data to the operation unit
5 103. Further, the screen rearrangement buffer 102
supplies the read image data to the intra prediction unit 113 and the inter prediction unit 114 as vAell. In other words, the image data of the frames are supplied to the operation unit 103, the intra prediction unit 113, and
10 the inter prediction unit 114 in the encoding order.
Further, v/hen the input image is the still image, since there is no concept of time {since the number of frames is 1), the rearranging is omitted. [0163]
15 The operation unit 103 generates image data of a
differential image obtained by subtracting the predicted image from the input image using the image data of the input image read from the screen rearrangement buffer 102 and the image data of the predicted image supplied from
20 the intra prediction unit 113 or the inter prediction
unit 114 through the predicted image selecting unit 115. For example, in the case of an image on which the intra coding is performed, the operation unit 103 generates a differential image betv/een the input image and the
25 predicted image generated by the intra prediction unit 113. Further, for example, in the case of an image on which the inter coding is performed, the operation unit
103 generates a differential image betv/een the input
image and the predicted image generated by the inter
30 prediction unit 114. The operation unit 103 outputs the generated image data of the differential image to the
44
SP351163WO00
orthogonal transform unit 104.
[0164]
The orthogonal transform unit 104 performs
orthogonal transform such as discrete cosine transform or 5 Karhunen-Loeve transform on the image data of the
differential image supplied from the operation unit 103.
The orthogonal transform unit 104 supplies obtained
transform coefficients to the quantization unit 105.
[0165] 10 The quantization unit 105 quantizes the transform
coefficients supplied from the orthogonal transform unit
104. The quantization unit 105 sets a quantization
parameter based on information related to a target value
of a coding amount supplied from the rate control unit 15 116, and performs the quantization. The quantization
unit 105 supplies the quantized transform coefficients to
the lossless encoding unit 106.
[0166]
The lossless encoding unit 106 encodes the 20 transform coefficients quantized by the quantization unit
105 according to an arbitrary coding scheme, and
generates encoded data. Since coefficient data is
quantized under control of the rate control unit 116, the
data amount (the coding amount) of the encoded data 25 becomes the target value set by the rate control unit 116
(or approximates to the target value).
[0167]
The lossless encoding unit 106 acquires, for
example, information indicating an intra prediction mode 30 from the intra prediction unit 113, and acquires, for
example, information indicating an inter prediction mode
7
45
SP351163WO00
and differential motion vector information from the inter prediction unit 114. The lossless encoding unit 106 encodes various kinds of information according to an arbitrary coding scheme, and sets (multiplexes) the 5 encoded information as part of header information of encoded data. The lossless encoding unit 106 supplies the obtained encoded data to be accumulated in the accumulation buffer 107. [0168]
10 Examples of the coding scheme of the lossless encoding unit 106 include variable length coding and arithmetic coding. As the variable length coding, for example, there is Context-Adaptive Variable Length Coding (CAVLC) defined in the H.264/AVC scheme. As the
15 arithmetic coding, for example, there is Context-Adaptive Binary Arithmetic Coding (CABAC). [0169]
The accumulation buffer 107 temporarily holds the encoded data supplied from the lossless encoding unit 106
20 The accumulation buffer 107 outputs the held encoded data to the outside of the image coding device 100 at a predetermined timing. In other words, the accumulation buffer 107 also serves as a transmitting unit that transmits the encoded data.
25 [0170]
The transform coefficients quantized by the guantization unit 105 are also supplied to the inverse guantization unit 108. The inverse quantization unit 108 inversely guantizes the guantized transform coefficients
30 by a method corresponding to the guantization performed by the guantization unit 105. The inverse guantization
=^'
46
SP351163WO00
unit 108 supplies the obtained transform coefficients to
the inverse orthogonal transform unit 109.
[0171]
The inverse orthogonal transform unit 109 performs 5 inverse orthogonal transform on the transform
coefficients supplied from the inverse guantization unit 108 by a method corresponding to the orthogonal transform process performed by the orthogonal transform unit 104. The image data of the differential image is restored by
10 the inverse orthogonal transform. The inverse orthogonal transform unit 109 supplies the restored image data of the differential image to the operation unit 110 as the inverse orthogonal transform result. [0172]
15 The operation unit 110 generates image data of an image obtained by adding the restored differential image and the predicted image using the inverse orthogonal transform result supplied from the inverse orthogonal transform unit 109 and the image data of the predicted
20 image supplied from the intra prediction unit 113 or the inter prediction unit 114 through the predicted image selecting unit 115. In other v/ords, a locally reconstructed image (hereinafter, referred to as a "the reconstructed image") is obtained by the addition process,
25 The operation unit 110 supplies the image data of the
reconstructed image to the loop filter 111 or the intra prediction unit 113. [0173]
The loop filter 111 includes a deblocking filter,
30 an adaptive loop filter, or the like, and performs an appropriate filter process on the image data of the
47
SP351163WO00
reconstructed image supplied from the operation unit 110. For example, the loop filter 111 performs the deblocking filter process on the image data of the reconstructed image, and removes block distortion of the reconstructed 5 image. Further, for example, the loop filter 111
improves the image quality of the reconstructed image by performing the loop filter process on the deblocking filter process result (the image data of the reconstructed image from v/hich the block distortion has
10 been removed) using the Wiener Filter. [0174]
The loop filter 111 may perform another arbitrary filter process on the reconstructed image. The loop filter 111 may supply information used in the filter
15 process such as a filter coefficient to the lossless
encoding unit 106 as necessary so that the information
can be encoded.
[0175]
The loop filter 111 supplies the image data of the
20 reconstructed image (hereinafter, referred to as a
"decoded image") that has been subjected to the filter process as described above to the frame memory 112. [0176]
The frame memory 112 stores the supplied image data
25 of the decoded image. Further, the frame memory 112
supplies the stored image data of the decoded image to the inter prediction unit 114 as a reference image at a predetermined timing. [0177J
30 The intra prediction unit 113 performs the
prediction process on the current picture that is the
48
SP351163WO00
image of the frame of the processing target, and generates the predicted image. The intra prediction unit 113 performs the prediction process in units of predetermined blocks (using a block as a processing unit). 5 In other v/ords, the intra prediction unit 113 generates the predicted image of the current block serving as the processing target in the current picture. At this time, the intra prediction unit 113 performs the prediction process (intra-screen prediction {v/hich is also referred
10 to as intra prediction)) using the reconstructed image supplied from the operation unit 110 as the reference image. In other v?ords, the intra prediction unit 113 generates the predicted image using pixel values of pixels neighboring the current block which are included
15 in the reconstructed image. The pixel values of the neighboring pixels used for the intra prediction are pixel values of previously processed pixels of the current picture. In intra prediction (that is, in the scheme of generating the predicted image), a plurality of
20 methods (v/hich are also referred to as "intra prediction modes") are prepared as candidates in advance. The intra prediction unit 113 performs the intra prediction in a plurality of intra prediction modes v/hich are prepared in advance.
25 [0178]
The intra prediction unit 113 generates predicted images in all the intra prediction modes serving as the candidates, evaluates cost function values of the predicted images using the input image supplied from the
30 screen rearrangement buffer 102, and selects an optimal
mode. When the optimal intra prediction mode is selected,
49
SP351163WO00
the intra prediction unit 113 supplies the predicted image generated in the optimal mode to the predicted image selecting unit 115. [0179] 5 Further, as described above, the intra prediction unit 113 appropriately supplies, for example, the intra prediction mode information indicating the employed intra prediction mode to the lossless encoding unit 106 so that the information is encoded.
10 [0180]
The inter prediction unit 114 performs the prediction process on the current picture, and generates the predicted image. The inter prediction unit 114 performs the prediction process in units of predetermined
15 blocks (using a block as a processing unit). In other words, the inter prediction unit 114 generates the predicted image of the current block serving as the processing target in the current picture. At this time, the inter prediction unit 114 performs the prediction
20 process using the image data of the input image supplied from the screen rearrangement buffer 102 and the image data of the decoded image supplied from the frame memory 112 as the reference image. The decoded image is an image (another picture other than the current picture) of
25 the frame processed before the current picture. In other v/ords, the inter prediction unit 114 performs the prediction process (inter-screen prediction (which is also referred to as "inter prediction")} of generating the predicted image using an image of another picture.
30 [0181]
The inter prediction includes motion prediction and
50
SP351163WO00
motion compensation. More specifically, the inter prediction unit 114 performs the motion prediction on the current block using the input image and the reference image, and detects a motion vector. Then, the inter 5 prediction unit 114 performs the motion compensation
process according to the detected motion vector using the reference image, and generates the predicted image (inter predicted image information) of the current block. In the inter prediction (that is, in the scheme of
10 generating the predicted image), a plurality of methods (which are also referred to as "inter prediction modes") are prepared as candidates in advance. The inter prediction unit 114 performs the inter prediction in a plurality of inter prediction modes which are prepared in
15 advance. [0182]
The inter prediction unit 114 generates the predicted images in all the inter prediction modes serving as the candidate. The inter prediction unit 114
20 evaluates cost function values of the predicted images using the input image supplied from the screen rearrangement buffer 102, information of the generated differential motion vector, and the like, and selects an optimal mode. When the optimal inter prediction mode is
25 selected, the inter prediction unit 114 supplies the predicted image generated in the optimal mode to the predicted image selecting unit 115. [0183]
The inter prediction unit 114 supplies the
30 information indicating the employed inter prediction mode, information necessary for performing processing in the
51
SP351163WO00
inter prediction mode v/hen the encoded data is decoded, and the like to the lossless encoding unit 106 so that the information is encoded. As the necessary information, there is information of a generated differential motion 5 vector, and as prediction motion vector information, there is a flag indicating an index of a prediction motion vector, [0184]
The predicted image selecting unit 115 selects a
10 supply source of the predicted image to be supplied to the operation unit 103 and the operation unit 110. For example, in the case of the intra coding, the predicted image selecting unit 115 selects the intra prediction unit 113 as the supply source of the predicted image, and
15 supplies the predicted image supplied from the intra prediction unit 113 to the operation unit 103 and the operation unit 110. Further, for example, in the case of the inter coding, the predicted image selecting unit 115 selects the inter prediction unit 114 as the supply
20 source of the predicted image, and supplies the predicted image supplied from the inter prediction unit 114 to the operation unit 103 and the operation unit 110. [0185J
The rate control unit 116 control a rate of a
25 quantization operation of the quantization unit 105 based on the coding amount of the encoded data accumulated in the accumulation buffer 107 such that neither an overflov/ nor an underflov/ occurs. [0186]
30 The image coding device 100 further includes a profile setting unit 121.
52
SP351163WO00
[0187]
The profile setting unit 121 sets a profile that is applied to encoding of image data. For example, when the still image is encoded, the profile setting unit 121 sets 5 the still picture profile. For example, the profile setting unit 121 sets the profile according to an external instruction such as a user instruction. Of course, this method is arbitrary, and the profile may be set based on any information. When the profile is set, 10 the profile setting unit 121 supplies the information to the lossless encoding unit 106. [0188]
The lossless encoding unit 106 performs encoding according to the profile set by the profile setting unit 15 121. For example, v/hen the still picture profile is set by the profile setting unit 121, the lossless encoding unit 106 restricts the value of the syntax element related to the inter-image process. The specific example of this restriction has been described in <0. Overviev/>. 20 [0189]
Fig. 26 is a block diagram illustrating an exemplary main configuration of the lossless encoding unit 106 of Fig. 25 related to a syntax element setting. 25 As illustrated in Fig. 25, the lossless encoding unit 106 includes a syntax element setting unit 131 and an encoding unit 132. [0190]
The syntax element setting unit 131 performs 30 processing related to a syntax element setting. For example, the syntax element setting unit 131 sets the
syntax elements of various NAL units (nal^unit) such as the video parameter set {VPS}, the sequence parameter set (SPS), the picture parameter set (PPS), and the slice header. The syntax element setting unit 131 acquires necessary information from the processing units of the image coding device 100 such as the intra prediction unit 113 and the inter prediction unit 114 as necessary, and sets the value of the syntax element based on the acquired information. [0191]
Further, the syntax element setting unit 131 receives the designated profile from the profile setting unit 121. The syntax element setting unit 131 sets the syntax element according to the profile designated by the profile setting unit 121. [0192]
For example, v/hen the still picture profile is set by the profile setting unit 121, the syntax element setting unit 131 restricts the value of the syntax element related to the inter-image process. The specific example of this restriction has been described in <0. Overview>. [0193]
The syntax element setting unit 131 supplies the set syntax element to the encoding unit 132. [0194]
For example, the encoding unit 132 acquires the quantization coefficients of the image data from the quantization unit 105, encodes the acquired quantization coefficients according to the arithmetic coding or the lil^e, and obtains the encoded data. The encoding unit
132 supplies the obtained encoded data to the
accumulation buffer 107.
[0195]
Further, the encoding unit 132 also encodes information other than the image data, and supplies the encoded information to the accumulation buffer 107. For example, the encoding unit 132 encodes the syntax element set by the syntax element setting unit 131, and supplies the encoded syntax element to the accumulation buffer 107. [0196]
As described above, when the still picture profile is set/ the lossless encoding unit 106 restricts the value of the syntax element related to the inter-image process, and generates the encoded data. Thus, the image coding device 100 can suppress transmission of redundant information and suppress a reduction in coding efficiency. [0197]
In addition, the image coding device 100 can suppress an increase of an unnecessary load in the decoding process by reducing the redundant information of the encoded data. [01981
Further, the image coding device 100 can sv/itch the profile and easily encode both the still image and the moving image through a single circuit. [0199]
Next, an exemplary flov/ of the coding process performed by the image coding device 100 illustrated in Fig. 25 v/ill be described v/ith reference to flov/charts of Figs. 27 and 28.
[0200]
When the coding process starts, in step SlOl of Fig. 27, the profile setting unit 121 sets the profile according to the image data to be encoded, for example, according to the external instruction such as the user instruction. [0201]
in step S102, the lossless encoding unit 106 determines v/hether or not the profile set in step SlOl is the still image profile {the profile suitable for encoding of the image data of the still image). When the set profile is determined to be the moving image profile (the profile suitable for encoding of the image data of the moving image), the process proceeds to step S103. [0202]
When the image data of the moving image is encoded, in step S103, the A/D converter 101 performs A/D conversion on image of frames (pictures) of an input moving image. [0203]
In step S10 4, the screen rearrangement buffer 102 stores the images that have been subjected to the A/D conversion in step Si03, and rearranges the respective pictures arranged in the display order in the encoding order. [0204]
In step S105, the intra prediction unit 113 performs the intra prediction process of the intra prediction mode. [0205]
In step S106, the inter prediction unit 114
performs the inter prediction process in which the motion prediction and the motion compensation of the inter prediction mode are performed. [0206]
In step S107/ the predicted image selecting unit 115 selects the predicted image based on the cost function value or the like. In other v^ords^ the predicted image selecting unit 115 selects either of the predicted image generated by the intra prediction in step S105 and the predicted image generated by the inter prediction in step S106. [0207]
In step S108, the operation unit 103 calculates a difference betv^een the input image in v/hich the frame order is rearranged by the process of step S104 and the predicted image selected by the process of step S107. In other v/ords, the operation unit 103 generates image data of a differential image between the input image and the predicted image. The obtained image data of the differential image is smaller in a data amount than original image data. Thus, the data amount can be compressed to be smaller than when the image is encoded v/ithout change. [0208]
In step S109, the orthogonal transform unit 104 performs the orthogonal transform on the image data of the differential image generated by the process of step S108. [0209]
In step SllO, the quantization unit 105 quantizes the orthogonal transform coefficients obtained by the
process of step S109 using the quantization parameter
calculated by the rate control unit 116.
[0210]
In step Sill, the inverse quantization unit 108 inversely quantizes the quantized coefficients {which are also referred to as "quantization coefficients") generated by the process of step SllO according to characteristics corresponding to characteristics of the quantization unit 105. [0211]
In step S112, the inverse orthogonal transform unit 109 performs the inverse orthogonal transform on the orthogonal transform coefficients obtained by the process of step Sill. [0212]
In step S113, the operation unit 110 generates image data of a reconstructed image by adding the predicted image selected by the process of step S107 to the differential image restored by the process of step S112. [0213]
In step 3114, the loop filter 111 performs the loop filter process on the image data of the reconstructed image generated by the process of step S113. As a result, for example, the block distortion of the reconstructed image is removed. [0214]
In step S115, the frame memory 112 stores the decoded image obtained by the process of step S114. [0215]
In step S116, the lossless encoding unit 106
encodes the quantized coefficients obtained by the process of step SllO. In other words, lossless coding such as variable length coding or arithmetic coding is performed on data corresponding to the differential image. [0216]
At this time, the lossless encoding unit 106 encodes information related to the prediction mode of the predicted image selected by the process of step 3107/ and adds the encoded information to the encoded data obtained by encoding the differential image. In other v/ords, the lossless encoding unit 106 also encodes the optimal intra prediction mode information supplied from the intra prediction unit 113, information according to the optimal inter prediction mode supplied from the inter prediction unit 114, or the like, and adds the encoded information to the encoded data. [0217]
Further, the lossless encoding unit 106 also sets and encodes the syntax elements of various kinds of NAL units or the like, and adds the encoded syntax elements to the encoded data. [0218]
In step S117, the accumulation buffer 107 accumulates the encoded data obtained by the process of step S115. The encoded data accumulated in the accumulation buffer 107 is appropriately read, and transmitted to the decoding side via a transmission path or a recording medium. [0219]
In step S118, the rate control unit 116 controls the rate of the quantization operation of the quantization unit 105 based on the coding amount (the generated coding amount) of the encoded data accumulated
in the accumulation buffer 107 by the process of step 5117 such that neither an overflow nor an underflow
5 occurs. Further, the rate control unit 116 supplies
information related to the quantization parameter to the
quantization unit 105. When the process of step 5118
ends, the coding process ends.
CLAIMS
1. (Amended). An image processing device, comprising: a receiving unit that receives an encoded bitstream including a syntax element related to processing of a sub layer; and an analyzing unit that analyzes the syntax element received by the receiving unit in a state in which a value of the syntax element is restricted when the bitstream is a bitstream encoded based on a profile for encoding a still image.
2. The image processing device according to claim1, further comprising a decoding unit that decodes the bitstream using the syntax element analyzed by the analyzing unit.
3. The image processing device according to claim1, further comprising a determining unit that determines whether or not the value of the syntax element received by the reciving unit is property restricted when the bitstream is the bitstream encoded based on the profile for encoding the still image.
4. The image processing device according to claim3, further comprising an abnormality processing unit that performs an abnormality process when the determining unit determines that the value of the syntax element is not property restricted.
5. The image processing device according to claim1, wherein the syntax element is stored in a sequence parameter set of the bitstream.
6. (Amended). The image processing device according to claim5, wherein the syntax element is a syntax related to a maximum of the number of sub layers.
7. The image processing device according to claim1, further comprising a determining unit that determines whether or not the value of the syntax element received by the receiving unit is a predetermined value.
8. The image processing device according to claim7, wherein the predetermined value is 0.
9. (Amended). An image processing method, comprising to processing of a sub layer; and analyzing the receiving syntax element in a state in which a value of the syntax element is restricted when the bitstream is a bitstream encoded based on a profile for encoding a still image.
10. The image processing method according to claim9, further comprising decoding the bitstream using the analyzed syntax element.
11. The image processing method according to claim9, further comprising determining whether or not the value of the received syntax element is properly restricted when the bitstream is the bitstream encoded based on the profile for encoding the still image.
12. The image processing method according to claim11, further comprising performing an abnormality process when the value of the syntax element is determined to be not property restricted.
13. The image processing method according to claim9, wherein the syntax element is stored in a sequence parameter set of the bitstream.
14. (Amended). The image processing method according to claim13, wherein the syntax element is syntax related to a maximum of the number of sub layers.
15. The image processing method according to claim9, further comprising determining wherther or not the value of the received syntax element is a predetermined value.
16. The image processing method according to claim15, wherein the predetermined value is 0.
| # | Name | Date |
|---|---|---|
| 1 | 4389-DELNP-2015.pdf | 2015-05-25 |
| 2 | PCT-IB-304.pdf | 2015-06-04 |
| 3 | OTHER RELEVANT DOCUMENTS.pdf | 2015-06-04 |
| 4 | FORM 3.pdf | 2015-06-04 |
| 5 | FORM 2 + SPECIFICATION-1.pdf | 2015-06-04 |
| 6 | DRAWING.pdf | 2015-06-04 |
| 7 | POWER OF AUTHORITY.pdf | 2015-06-25 |
| 8 | PCT-IB-304.pdf_1231.pdf | 2015-06-25 |
| 9 | OTHER RELEVANT DOCUMENTS.pdf_1232.pdf | 2015-06-25 |
| 10 | FORM 5.pdf | 2015-06-25 |
| 11 | FORM 3.pdf_1230.pdf | 2015-06-25 |
| 12 | FORM 2 + SPECIFICATION-1.pdf_1228.pdf | 2015-06-25 |
| 13 | DRAWING.pdf_1229.pdf | 2015-06-25 |
| 14 | 4389-delnp-2015-Form-3-(26-08-2015).pdf | 2015-08-26 |
| 15 | 4389-delnp-2015-Correspondence Others-(26-08-2015).pdf | 2015-08-26 |
| 16 | 4389-delnp-2015-Form-3-(04-04-2016).pdf | 2016-04-04 |
| 17 | 4389-delnp-2015-Correspondence Others-(04-04-2016).pdf | 2016-04-04 |
| 18 | Form 18 [06-10-2016(online)].pdf | 2016-10-06 |
| 19 | 4389-DELNP-2015-FER.pdf | 2019-12-16 |
| 1 | SerachStrategy_GP_TPO_12-12-2019.pdf |