Sign In to Follow Application
View All Documents & Correspondence

Image Processing Device And Method

Abstract: This disclosure relates to an image processing device and method which make it possible to suppress an increase in the load of decoding processing of coded data in a video-based approach to a point cloud. Parameters relating to a plurality of point group models of a point cloud are transformed, a two-dimensional planar image on which the plurality of point group models relating to which the parameters have been transformed are projected is coded, and a bit stream including coded data of the two-dimensional image and transformation information that is information relating to the transformation of the parameters is generated. This disclosure is applicable, for example, to an information processing device, an image processing device, electronic equipment, an information processing method, or a program.

Get Free WhatsApp Updates!
Notices, Deadlines & Correspondence

Patent Information

Application #
Filing Date
24 December 2020
Publication Number
11/2021
Publication Type
INA
Invention Field
ELECTRONICS
Status
Email
Parent Application

Applicants

SONY CORPORATION
1-7-1, Konan, Minato-ku, Tokyo 1080075

Inventors

1. KUMA Satoru
c/o SONY CORPORATION, 1-7-1, Konan, Minato-ku, Tokyo 1080075
2. NAKAGAMI Ohji
c/o SONY CORPORATION, 1-7-1, Konan, Minato-ku, Tokyo 1080075
3. YANO Koji
c/o SONY CORPORATION, 1-7-1, Konan, Minato-ku, Tokyo 1080075
4. KATO Tsuyoshi
c/o SONY CORPORATION, 1-7-1, Konan, Minato-ku, Tokyo 1080075

Specification

Title of Invention: Image Processing Device and Method
Technical field
[0001]
 The present disclosure relates to an image processing device and a method, and more particularly to an image processing device and a method capable of suppressing an increase in the load of decoding processing of encoded data in a video-based approach of a point cloud.
Background technology
[0002]
 Conventionally, as a method for encoding 3D data representing a three-dimensional structure such as a point cloud, there has been coding using a voxel such as Octtree (for example, Non-Patent Document 1). reference).
[0003]
 In recent years, as another coding method, for example, an approach in which the position and color information of a point cloud is projected onto a two-dimensional plane for each small area and encoded by a coding method for a two-dimensional image (hereinafter, video-based). An approach (also referred to as a video-based approach) has been proposed (see, for example, Non-Patent Documents 2 to 4).
[0004]
 In such coding, when a plurality of point cloud models exist in the point cloud, in the conventional method, each point cloud model is coded independently and bitstreams different from each other.
Prior art literature
Non-patent literature
[0005]
Non-Patent Document 1: R. Mekuria, Student Member IEEE, K. Blom, P. Cesar., Member, IEEE, "Design, Implementation and Evaluation of a Point Cloud Codec for Tele-Immersive Video", tcsvt_paper_submitted_february.pdf
Non-Patent Document 2: Tim Golla and Reinhard Klein, "Real-time Point Cloud Compression," IEEE, 2015
Non-Patent Document 3: K. Mammou, “Video-based and Hierarchical Approaches Point Cloud Compression”, MPEG m41649, Oct. 2017
Non-Patent Document 4: K. Mammou, “PCC Test Model Category 2 v0,” N17248 MPEG output document, October 2017
Outline of the invention
Problems to be solved by the invention
[0006]
 However, in the case of this method, as the number of point cloud models increases, the number of instances required for the decoder may increase.
[0007]
 The present disclosure has been made in view of such a situation, and suppresses an increase in the number of instances required for decoding the coded data of the point cloud video-based approach, and suppresses an increase in the load of the decoding process. Is what allows you to do it.
Means to solve problems
[0008]
 The image processing device on one aspect of the present technology is a two-dimensional image processing device in which a conversion unit that converts parameters related to a plurality of point group models of a point cloud and the plurality of point group models whose parameters are converted by the conversion unit are projected. It is an image processing apparatus including a coding unit that encodes a plane image and generates a bit stream including the coding data of the two-dimensional plane image and conversion information that is information related to the conversion of the parameters by the conversion unit. ..
[0009]
 In the image processing method of one aspect of the present technology, parameters relating to a plurality of point group models of a point cloud are converted, and a two-dimensional plane image on which the plurality of point group models to which the parameters are converted are projected is encoded. This is an image processing method for generating a bit stream including encoded data of a two-dimensional plane image and conversion information which is information related to the conversion of the parameters.
[0010]
 The image processing device of the other aspect of the present technology decodes the bit stream and is information on the conversion of the two-dimensional plane image on which the plurality of point group models are projected and the respective parameters of the plurality of point group models. Each of the plurality of point group models is reconstructed from the decoding unit that generates information and the two-dimensional plane image generated by the decoding unit, and each of the plurality of point group models is reconstructed based on the transformation information. It is an image processing apparatus including a reconstruction unit that reversely transforms the parameters of.
[0011]
 The image processing method of the other aspect of the present technology is a transformation which is information about the transformation of the two-dimensional plane image in which the bit stream is decoded and the plurality of point group models are projected and the respective parameters of the plurality of point group models. An image that generates information, reconstructs each of the plurality of point group models from the generated two-dimensional plane image, and inversely transforms each parameter of the plurality of point group models based on the conversion information. It is a processing method.
[0012]
 The image processing device on the other side of the present technology projects a plurality of partial point cloud models constituting the point cloud model onto projection planes set independently of each other, and each projection by the projection unit. A coding unit that encodes a two-dimensional plane image in which patches of the partial point cloud model projected on a surface are arranged and an occupancy map including projection surface information that is information about the projection surface to generate a bit stream. It is an image processing apparatus including.
[0013]
 In the image processing method of still another aspect of the present technology, a plurality of partial point cloud models constituting the point cloud model are projected onto projection planes set independently of each other, and the portion projected on each projection plane. This is an image processing method for generating a bit stream by encoding a two-dimensional plane image in which a patch of a point cloud model is arranged and an occupancy map including projection plane information which is information about the projection plane.
[0014]
 The image processing device on the other side of the present technology decodes the bit stream and projects the point cloud model into the two-dimensional plane image and the projection planes of the plurality of partial point cloud models included in the point cloud model. Based on a decoding unit that generates an occupancy map including projection plane information which is information about, the two-dimensional plane image generated by the decoding unit, and the projection plane information included in the occupancy map. It is an image processing apparatus including a reconstruction unit for reconstructing the point cloud model.
[0015]
 Another aspect of the image processing method of the present technology is to decode a bit stream and project a two-dimensional plane image on which a point group model is projected, and a projection surface of each of a plurality of partial point group models included in the point group model. An occupancy map including projection plane information which is information about the above is generated, and the point group model is reconstructed based on the generated two-dimensional plane image and the projection plane information included in the occupancy map. This is an image processing method.
[0016]
 In the image processing apparatus and method of one aspect of the present technology, the parameters related to the plurality of point cloud models of the point cloud are converted, and the two-dimensional plane image on which the plurality of point cloud models to which the parameters are converted are projected is encoded. Then, a bit stream including the coded data of the two-dimensional plane image and the conversion information which is the information regarding the conversion of the parameters is generated.
[0017]
 In image processing devices and methods of other aspects of the present technology, information on a two-dimensional plane image in which a bit stream is decoded and a plurality of point group models are projected, and conversion of each parameter of the plurality of point group models. The transformation information is generated, each of the plurality of point group models is reconstructed from the generated two-dimensional plane image, and each parameter of the plurality of point group models is inversely transformed based on the transformation information. Will be done.
[0018]
 In the image processing apparatus and method of still another aspect of the present technology, a plurality of partial point cloud models constituting the point cloud model are projected onto projection planes set independently of each other, and projected onto each projection plane. The two-dimensional plane image in which the patch of the partial point cloud model is arranged and the occupancy map including the projection plane information which is the information about the projection plane are encoded to generate a bit stream.
[0019]
 In the image processing apparatus and method of yet another aspect of the present technology, a two-dimensional plane image in which a bit stream is decoded and a point cloud model is projected, and a plurality of partial point cloud models included in the point cloud model are respectively. An occupancy map including projection plane information, which is information about the projection plane of, is generated, and the point cloud model is regenerated based on the generated two-dimensional plane image and the projection plane information included in the occupancy map. Will be built.
Effect of the invention
[0020]
 According to the present disclosure, images can be processed. In particular, it is possible to suppress an increase in the load of decoding processing of the coded data of the point cloud video-based approach.
A brief description of the drawing
[0021]
[Fig. 1] Fig. 1 is a diagram summarizing the main features of this technology.
[Fig. 2] Fig. 2 is a diagram illustrating an outline of coding / decoding of a point cloud to which this technology is applied.
[Fig. 3] Fig. 3 is a diagram illustrating an outline of coding / decoding of a point cloud to which this technology is applied.
[Fig. 4] Fig. 4 is a block diagram showing a main configuration example of a coding device.
[Fig. 5] Fig. 5 is a flowchart illustrating an example of a flow of coding processing.
[Fig. 6] Fig. 6 is a flowchart illustrating an example of a flow of conversion processing.
[Fig. 7] Fig. 7 is a flowchart illustrating an example of a flow of a bounding box setting process.
FIG. 8 is a flowchart illustrating an example of a packing process flow.
[Fig. 9] Fig. 9 is a flowchart illustrating an example of a flow of BB information generation processing.
FIG. 10 is a block diagram showing a main configuration example of a decoding device.
FIG. 11 is a flowchart illustrating an example of a flow of decoding processing.
[Fig. 12] Fig. 12 is a diagram summarizing the main features of this technology.
[Fig. 13] Fig. 13 is a diagram illustrating an outline of coding / decoding of a point cloud to which this technology is applied.
FIG. 14 is a block diagram showing a main configuration example of a coding device.
[Fig. 15] Fig. 15 is a block diagram showing a main configuration example of a patch disassembly unit.
FIG. 16 is a flowchart illustrating an example of a flow of coding processing.
[Fig. 17] Fig. 17 is a flowchart illustrating an example of a flow of patch disassembly processing.
[Fig. 18] Fig. 18 is a flowchart illustrating an example of a flow of division processing.
FIG. 19 is a flowchart illustrating an example of a packing process flow.
FIG. 20 is a block diagram showing a main configuration example of a decoding device.
FIG. 21 is a flowchart illustrating an example of a flow of decoding processing.
[Fig. 22] Fig. 22 is a block diagram showing a main configuration example of a computer.
Mode for carrying out the invention
[0022]
 Hereinafter, embodiments for carrying out the present disclosure (hereinafter referred to as embodiments) will be described. The explanation will be given in the following order.
 1. 1. Conversion of point cloud model
 2. Supports partial decoding
 3. First Embodiment (encoding device)
 4. Second embodiment (decoding device)
 5. Local projection plane control
 6. Third Embodiment (encoding device)
 7. Fourth Embodiment (decoding device)
 8. Addendum
[0023]
 <1. Conversion of point cloud model>
  
 The scope disclosed in this technology is not limited to the contents described in the embodiments, but is as follows, which is known at the time of filing. The contents described in non-patent documents are also included.
[0024]
 Non-Patent Document 1: (above)
 Non-Patent Document 2: (above)
 Non-Patent Document 3: (above)
 Non-Patent Document 4: (above)
 Non-Patent Document 5: TELECOMMUNICATION STANDARDIZATION SECTOR OF ITU (International Telecommunication Union), "Advanced video coding for generic audiovisual services ", H.264, 04/2017
 Non-Patent Document 6: TELECOMMUNICATION STANDARDIZATION SECTOR OF ITU (International Telecommunication Union)," High efficiency video coding ", H.265, 12/2016
 Non-Patent Document 7: Jianle Chen, Elena Alshina, Gary J. Sullivan, Jens-Rainer, Jill Boyce, "Algorithm Description of Joint Exploration Test Model 4", JVET-G1001_v1, Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29/WG 11 7th Meeting: Torino, IT, 13-21 July 2017
[0025]
 In other words, the contents described in the above-mentioned non-patent documents are also the basis for determining the support requirements. For example, even if the Quad-Tree Block Structure described in Non-Patent Document 6 and the QTBT (Quad Tree Plus Binary Tree) Block Structure described in Non-Patent Document 7 are not directly described in the embodiment. It is within the scope of disclosure of this technology and shall meet the support requirements of the scope of claims. Similarly, technical terms such as Parsing, Syntax, and Semantics are also within the scope of disclosure of the present technology even if there is no direct description in the embodiment. Shall meet the support requirements of the scope of the claim.
[0026]
  
 Conventionally, a point cloud that represents a three-dimensional structure based on the position information and attribute information of a point cloud, a mesh that is composed of vertices, edges, and faces and defines a three-dimensional shape using polygonal representation, etc. The data existed.
[0027]
 For example, in the case of a point cloud, a three-dimensional structure is expressed as a set (point cloud) of a large number of points. That is, the point cloud data is composed of position information and attribute information (for example, color, etc.) of each point in this point cloud. Therefore, the data structure is relatively simple, and an arbitrary three-dimensional structure can be expressed with sufficient accuracy by using a sufficiently large number of points.
[0028]
   A video-based approach (Video-based approach) in which
 the position and color information of such a point cloud is projected onto a two-dimensional plane for each small area and encoded by a coding method for a two-dimensional image. approach) has been proposed.
[0029]
 In this video-based approach, the input point cloud is divided into a plurality of segmentations (also referred to as regions), and each region is projected onto a two-dimensional plane. The data for each position of the point cloud (that is, the data for each point) is composed of position information (Geometry (also referred to as Depth)) and attribute information (Texture) as described above, and each of them is 2 for each area. Projected onto a dimensional plane.
[0030]
 Then, each segmentation (also referred to as a patch) projected on the two-dimensional plane is arranged in the two-dimensional image, and is used for a two-dimensional plane image such as AVC (Advanced Video Coding) or HEVC (High Efficiency Video Coding). It is encoded by the coding method of.
[0031]
   In such a point cloud, there may be a plurality of point cloud models composed of dense point clouds. For example, when a plurality of people are scattered in a plaza and the three-dimensional space including the entire plaza is converted into a point cloud, a dense point cloud (point cloud model) is formed in each person's part. That is, a plurality of point cloud models are formed.
[0032]
 In such a case, if the entire point cloud is regarded as one large point cloud model and coded, there is a risk that the coding efficiency will be reduced because there are many parts where the points are sparse. Therefore, a method has been considered in which each of a plurality of point cloud models included in the point cloud is encoded independently of each other (for example, for each person). However, in this case, since a bit stream is generated for each point cloud model, if the number of point cloud models increases, the number of bit streams may increase and the number of instances required for the decoder may increase. For example, when there are hundreds of point cloud models in the point cloud, there is a possibility that hundreds of times as many instances as the number of instances required for decoding one bitstream will be required. As the load of the decoding process increases in this way, the cost of the decoder may increase and the processing time may increase.
[0033]
  
 Therefore, a plurality of point cloud models are stored in one bitstream. For example, as shown in the top row of the table in FIG. 1, a plurality of point cloud models are converted into one group and encoded to combine the multiple point cloud models into one bitstream. To store in.
[0034]
 By doing so, even when a plurality of point cloud models exist in the point cloud, it is possible to suppress an increase in the number of bitstreams while suppressing a decrease in coding efficiency. Therefore, it is possible to suppress an increase in the number of instances required for decoding. That is, it is possible to suppress an increase in the load of the decoding process. This makes it possible to suppress an increase in the cost of the decoder. In addition, it is possible to suppress an increase in the processing time of the decoding process.
[0035]
  
 More specifically, as shown in row # 1 of the table in FIG. 1, one of the parameters related to each point cloud model is transformed to reduce the sparse points. Group them together (conversion of each point cloud model). That is, the groups are collectively encoded into a single bitstream.
[0036]
 For example, as shown on the far left of FIG. 2, there are point cloud models 11-1 to point cloud models 11-4 in the point cloud, and bounding boxes 12-1 to 12-4 are provided for each of them. It is assumed that it is set. When the point cloud model 11-1 to the point cloud model 11-4 are not described separately from each other, they are referred to as a point cloud model 11. When the bounding boxes 12-1 to 12-4 are not described separately from each other, they are referred to as bounding boxes 12. That is, these point cloud models 11 are separated from each other, and a bounding box 12 is set for each.
[0037]
 At the time of coding, as shown second from the left in FIG. 2, the parameters of these point cloud models 11 are converted and collected to form a group with few sparse points. In the example of FIG. 2, a group consisting of bounding boxes 12-1 to 12-4 is formed, and a bounding box 13 for coding is set for the entire group.
[0038]
 The content of this conversion (parameters to be converted) is arbitrary. For example, as shown in the table of FIG. 1, the coordinates of the point cloud model 11 may be transformed. For example, as the coordinate transformation, the coordinates of the point cloud model 11 may be shifted (the position of the point cloud model 11 may be moved). That is, the positions of the point cloud model 11-1 to the point cloud model 11-4 in the point cloud shown on the far left of FIG. 2 may be spatially separated from the others. By such a conversion, for example, the positions of the point cloud models can be brought closer, and the sparse part of the group can be reduced.
[0039]
 Further, the coordinates of the point cloud model 11 may be rotated (the posture of the point cloud model 11 may be rotated). That is, the postures (orientations) of the point cloud model 11-1 to the point cloud model 11-4 in the point cloud shown on the far left of FIG. 2 may be different from the others. By such a conversion, for example, the postures of the point cloud models can be aligned, and the sparse portion of the group can be reduced.
[0040]
 Further, for example, as shown in the table of FIG. 1, the time (time stamp (TimeStump)) of the point cloud model 11 may be converted. That is, the positions of the point cloud model 11-1 to the point cloud model 11-4 in the point cloud shown on the far left of FIG. 2 may be separated from each other in time (existing at a time different from the others). It may include a point cloud model to be used). By such conversion, for example, the times of the point cloud models can be aligned, and the point cloud models existing at different times can be grouped into one group.
[0041]
 Further, for example, as shown in the table of FIG. 1, the size (scale) of the point cloud model 11 may be converted. That is, the scales of the point cloud model 11-1 to the point cloud model 11-4 in the point cloud shown on the far left of FIG. 2 may be different from the others. By such conversion, for example, the size (scale) of each point cloud model can be made uniform, and the resolution in the spatial direction can be made uniform.
[0042]
 Further, for example, as shown in the table of FIG. 1, the frame rate of the point cloud model 11 may be converted. That is, the frame rates of the point cloud model 11-1 to the point cloud model 11-4 in the point cloud shown on the far left of FIG. 2 may be different from the others. By such conversion, for example, the frame rate (that is, the resolution in the time direction) of each point cloud model can be made uniform.
[0043]
 In FIG. 2, at the time of coding, a group of point cloud models in the bounding box 13 for coding shown second from the left is projected onto a two-dimensional plane, packed in a video frame, and third from the left. As shown, a color video frame 14 containing the attribute information (Texture), a geometry video frame 15 containing the position information (Depth), and an occupancy map 16 are generated.
[0044]
 Then, it is encoded using a coding method for a two-dimensional image, and one 2D bit stream 17 (including an occupancy map 18 and a header (Header) 19) is inserted as shown on the far right of FIG. Generate. That is, the coded data of the point cloud model 11-1 to the point cloud model 11-4 is stored in this one 2D bit stream 17.
[0045]
 By doing so, it is possible to suppress an increase in the number of bit streams and suppress an increase in the load of the decoding process.
[0046]
   In
 the case of decoding, the above reverse processing is performed. That is, the inverse conversion process of returning each point cloud model from the state of the second group from the left in FIG. 2 to the original state of the first from the left is performed. Therefore, as shown in row # 1 of the table in FIG. 1, the conversion of parameters related to the point cloud model performed during encoding (in FIG. 2) so that this inverse conversion can be performed during decoding. The conversion information indicating the contents of the conversion from the first state to the second state from the left) is generated and transmitted to the decoding side (for example, included in the bit stream) (that is, the conversion information is signaled).
[0047]
 This conversion information may be any information as long as it indicates the content of conversion (the amount of change in the converted parameters). For example, as shown in the table of FIG. 1, it may be information indicating the ratio of the converted parameters (for example, Shift, Rotate, Time Stump, Scale, Frame Rate, etc.) before and after the conversion. Further, for example, as shown in the table of FIG. 1, it may be information indicating the difference between the converted parameters (for example, Shift, Rotate, Time Stump, Scale, Frame Rate, etc.) before and after the conversion.
[0048]
 Further, the method of transmitting this conversion information is arbitrary. For example, it may be transmitted in association with a bit stream containing the coded data of the point cloud model. For example, it may be included in the bitstream. For example, as shown in FIG. 2, such conversion information may be included in the header 19 of the 2D bit stream 17 as BB information 21.
[0049]
 For example, the BB information 21 includes conversion information such as shift amount information (Shift x / y / z), rotation amount information (Rotate x / y / z), and time stamp change amount information (Time). Stump), information indicating the amount of change in scale (Scale x / y / z), information indicating the amount of change in frame rate (Frame Rate), etc. are included. At the time of coding, the BB information 21 is generated and stored in the header 19, and at the time of decoding, the BB information 21 is read from the header 19 and used for the inverse conversion.
[0050]
   As
 shown in row # 2 of the table in FIG. 1, when projecting the group of the point cloud model shown second from the left in FIG. 2 onto a two-dimensional plane. , The projection plane may be set for each point cloud model. By doing so, each point cloud model can be projected onto a more appropriate projection plane, so that the reduction in coding efficiency due to the inefficiency of the projection of the point cloud model can be suppressed (code). The efficiency of conversion can be improved).
[0051]
 For example, as shown in the table of FIG. 1, it may be possible to rotate the projection planes of each point cloud model in six orthogonal directions. By doing so, it is possible to relatively easily make the projection surface a more appropriate surface (a surface capable of performing efficient projection) with respect to the point cloud model.
[0052]
 Further, as shown in the table of FIG. 1, an arbitrary projection plane may be set (added). By doing so, the degree of freedom of the projection surface of each point cloud model can be improved, and more efficient projection can be expected (further improvement in coding efficiency can be expected).
[0053]
   When the
 projection plane is set for each point cloud model in this way (when the projection plane is not common to the group), the projection is performed when decoding (when reconstructing the point cloud). Surface information is also needed. Therefore, as shown in row # 2 of the table of FIG. 1, the projection plane information which is the information about the projection plane set as described above is transmitted to the decoding side (signaling of the projection plane information).
[0054]
 As shown in the table of FIG. 1, the projection plane information may include any information as long as it is information necessary for specifying the projection plane. For example, as shown in FIG. 2, the projection plane information 22 may be stored in the occupancy map 16 (that is, the occupancy map 18).
[0055]
 For example, the projection surface information 22 indicates information indicating the amount of shift of the projection surface in the six orthogonal directions (Shift x / y / z), information indicating the amount of change in scale (Scale x / y / z), and the amount of rotation. Information (Rotate x / y / z) etc. is included. Further, for example, the projection plane information 22 may include information indicating the added arbitrary projection plane.
[0056]
 At the time of coding, this projection surface information 22 is generated and stored in the occupancy map 16, and at the time of decoding, the projection surface information 22 is read out from the occupancy map 16 to reconstruct the point cloud. It is used for.
[0057]
 <2. Partial decoding support>
  
 Also, as shown in row # 3 of the table in Fig. 1, patches for each point cloud model are placed in a two-dimensional image and packed as a video frame. In some cases, the area where the patch is placed may be controlled for each point cloud model. For example, patches belonging to the same point cloud model may be placed in the same area. This area is optional. For example, it may be an independently decodable coding unit. That is, patches belonging to the same point cloud model may be arranged in the same code unit that can be independently decoded.
[0058]
 This independently decodable coding unit is arbitrary. For example, as shown in the table of FIG. 1, it may be a frame, a slice, or a tile.
[0059]
 For example, as shown in A of FIG. 3, it is assumed that the point cloud model 31-1 and the point cloud model 31-2 exist in the point cloud. A bounding box 32-1 is set in the point cloud model 31-1 and a bounding box 32-2 is set in the point cloud model 31-2. In other words, these are point cloud models that are independent of each other. When the patches of the point cloud model 31-1 and the point cloud model 31-2 are arranged and packed in the two-dimensional image 33, in the conventional method, both patches do not divide the area as shown in B of FIG. Was placed in. In FIG. 3B, patches 34-1 to 34-4 are patches of the point cloud model 31-1 and patches 35-1 to 35-5 are patches of the point cloud model 31-2. As shown in FIG. 3C, patches 34-1 to 34-4 of the point cloud model 31-1 are arranged on, for example, slice 36-2 of the two-dimensional image 33, and the point cloud model 31- Patch 35-1 to patch 35-5 of 2 are arranged on the slice 36-1 of the two-dimensional image 33.
[0060]
 The independently decodable coding unit for arranging the patches in this way is controlled for each point cloud model (patches belonging to the same point cloud model are arranged in the same and independently decodable coding unit. By doing so, it is possible to realize so-called "partial decoding" in which only a part of the point cloud model is decoded and reconstructed. For example, in the case of C in FIG. 3, since slices 36-1 and slices 36-2 are coding units that can be decoded independently of each other, only patches 35-1 to 35-5 of slices 36-1 can be decoded. , Only patch 34-1 to patch 34-4 of slice 36-2 can be decoded. That is, only the point cloud model 31-1 can be decoded and reconstructed, or only the point cloud model 31-2 can be decoded and reconstructed.
[0061]
   In
 order to perform such partial decoding, the decoding side grasps which area (which independently decodable coding unit) the patch of which point cloud model is arranged. There is a need. Therefore, as shown in row # 3 of the table in FIG. 1, model information, which is information about the point cloud model, is generated and transmitted (signaled) from the coding side to the decoding side.
[0062]
 The content of this model information is arbitrary. For example, as shown in the table of FIG. 1, the model information may include information indicating the number of point cloud models packed in the video frame. Further, for example, as shown in the table of FIG. 1, the model information may include information indicating an area (a coding unit that can be independently decoded) in which patches of each point cloud model are arranged.
[0063]
 The method of transmitting this model information is arbitrary. For example, it may be transmitted in association with a bit stream containing the coded data of the point cloud model. For example, it may be included in the bitstream. For example, as shown in FIG. 2, such model information may be included in the header 19 of the 2D bit stream 17 as BB information 21.
[0064]
 For example, the BB information 21 includes the number of models and information on the patch placement area of ​​each point cloud model as model information. At the time of encoding, the BB information 21 is generated and stored in the header 19, and at the time of decoding, the BB information 21 is read from the header 19 and used for partial decoding.
[0065]
 <3. First Embodiment>
  
 Next, a configuration for realizing each of the above methods will be described. FIG. 4 is a block diagram showing an example of a configuration of a coding device, which is an aspect of an image processing device to which the present technology is applied. The coding device 100 shown in FIG. 4 is a device (a coding device to which a video-based approach is applied) that projects 3D data such as a point cloud onto a two-dimensional plane and encodes it by a coding method for a two-dimensional image. ).
[0066]
 Note that FIG. 4 shows the main things such as the processing unit and the data flow, and not all of them are shown in FIG. That is, in the coding apparatus 100, there may be a processing unit that is not shown as a block in FIG. 4, or there may be a processing or data flow that is not shown as an arrow or the like in FIG. This also applies to other figures for explaining the processing unit and the like in the coding apparatus 100.
[0067]
 As shown in FIG. 4, the coding device 100 includes a model conversion unit 111, a patch decomposition unit 112, a packing unit 113, an auxiliary patch information compression unit 114, a video coding unit 115, a video coding unit 116, and an OMap coding unit. It has 117, a multiplexer 118, and a BB information generator 119.
[0068]
 The model transformation unit 111 performs processing related to parameter transformation related to the point cloud model. For example, the model transformation unit 111 acquires 3D data (for example, Point Cloud) representing a three-dimensional structure input to the coding apparatus 100. Further, the model conversion unit 111 converts the parameters related to the point cloud model included in the acquired point cloud. At that time, the model conversion unit 111 transforms the parameters related to each point cloud model by the method described above in and , and combines the plurality of point cloud models into one group. Summarize. The model conversion unit 111 supplies the patch decomposition unit 112 with a point cloud model whose parameters have been converted, that is, a point cloud including the point cloud model grouped into one group.
[0069]
 Further, the model transformation unit 111 generates transformation information about the transformation by the method described above in . The model transformation unit 111 supplies the generated transformation information to the BB information generation unit 119.
[0070]
 The patch decomposition unit 112 performs processing related to decomposition of 3D data. For example, the patch decomposition unit 112 acquires a point cloud (a point cloud including a plurality of point cloud models grouped together) supplied from the model transformation unit 111. Further, the patch decomposition unit 112 decomposes the acquired point cloud into a plurality of segmentations, projects the point cloud onto a two-dimensional plane for each segmentation, and generates a patch of position information and a patch of attribute information. The patch disassembling unit 112 supplies information about each generated patch to the packing unit 113. Further, the patch disassembling unit 112 supplies auxiliary patch information, which is information related to the disassembling, to the auxiliary patch information compression unit 114.
[0071]
 The packing unit 113 performs processing related to data packing. For example, the packing unit 113 acquires information on a patch of position information (Geometry) indicating the position of a point from the patch decomposition unit 112 and information on a patch of attribute information (Texture) such as color information added to the position information. To do.
[0072]
 Further, the packing unit 113 arranges each acquired patch on the two-dimensional image and packs it as a video frame. For example, the packing unit 113 arranges a patch of position information on a two-dimensional image and packs it as a video frame of position information (also referred to as a geometry video frame). Further, for example, the packing unit 113 arranges a patch of attribute information on a two-dimensional image and packs it as a video frame (also referred to as a color video frame) of the attribute information.
[0073]
 At the time of these packing, the packing unit 113 controls the area (encoded unit that can be independently decoded) for arranging the patch for each point cloud model by the method described above in . .. Then, the packing unit 113 generates model information by the method described above in . The packing unit 113 supplies the generated model information to the BB information generation unit 119.
[0074]
 In addition, the packing unit 113 generates an occupancy map corresponding to these video frames. Further, the packing unit 113 performs a dilation process on the color video frame.
[0075]
 The packing unit 113 supplies the geometry video frame thus generated to the video coding unit 115. Further, the packing unit 113 supplies the color video frame thus generated to the video coding unit 116. Further, the packing unit 113 supplies the occupancy map thus generated to the OMap coding unit 117. Further, the packing unit 113 supplies control information regarding such packing to the multiplexer 118.
[0076]
 The auxiliary patch information compression unit 114 performs processing related to compression of auxiliary patch information. For example, the auxiliary patch information compression unit 114 acquires the data supplied from the patch decomposition unit 112. The auxiliary patch information compression unit 114 encodes (compresses) the auxiliary patch information included in the acquired data. The auxiliary patch information compression unit 114 supplies the obtained coded data of the auxiliary patch information to the multiplexer 118.
[0077]
 The video coding unit 115 performs processing related to coding of a video frame of position information (Geometry). For example, the video coding unit 115 acquires a geometry video frame supplied from the packing unit 113. Further, the video coding unit 115 encodes the acquired geometry video frame by a coding method for an arbitrary two-dimensional image such as AVC or HEVC. The video coding unit 115 supplies the coded data (coded data of the geometry video frame) obtained by the coding to the multiplexer 118.
[0078]
 The video coding unit 116 performs processing related to coding of a video frame of attribute information (Texture). For example, the video coding unit 116 acquires a color video frame supplied from the packing unit 113. Further, the video coding unit 116 encodes the acquired color video frame by a coding method for an arbitrary two-dimensional image such as AVC or HEVC. The video coding unit 116 supplies the coded data (color video frame coded data) obtained by the coding to the multiplexer 118.
[0079]
 The OMap coding unit 117 performs processing related to coding the occupancy map. For example, the OMap coding unit 117 acquires the occupancy map supplied from the packing unit 113. Further, the OMap coding unit 117 encodes the acquired occupancy map by an arbitrary coding method such as arithmetic coding. The OMap coding unit 117 supplies the coded data (coded data of the occupancy map) obtained by the coding to the multiplexer 118.
[0080]
 The multiplexer 118 performs processing related to multiplexing. For example, the multiplexer 118 acquires the encoded data of the auxiliary patch information supplied from the auxiliary patch information compression unit 114. Further, the multiplexer 118 acquires control information regarding packing supplied from the packing unit 113. Further, the multiplexer 118 acquires the coded data of the geometry video frame supplied from the video coding unit 115. Further, the multiplexer 118 acquires the coded data of the color video frame supplied from the video coding unit 116. Further, the multiplexer 118 acquires the coded data of the occupancy map supplied from the OMap coding unit 117. Further, the multiplexer 118 acquires the BB information supplied from the BB information generation unit 119.
[0081]
 The multiplexer 118 multiplexes the acquired information to generate a bitstream. The multiplexer 118 outputs the generated bit stream to the outside of the coding device 100.
[0082]
 The BB information generation unit 119 performs processing related to the generation of BB information. For example, the BB information generation unit 119 acquires the conversion information supplied from the model transformation unit 111. Further, the BB information generation unit 119 acquires the model information supplied from the packing unit 113. The BB information generation unit 119 generates BB information including the acquired conversion information and model information. The BB information generation unit 119 supplies the generated BB information to the multiplexer 118. That is, the BB information is transmitted to the decoding side.
[0083]
 With such a configuration, the coding device 100 can convert a plurality of point cloud models of the point cloud, combine them, encode them, and store them in one bit stream. That is, even when a plurality of point cloud models exist in the point cloud, it is possible to suppress an increase in the number of bit streams while suppressing a decrease in coding efficiency. Therefore, the coding device 100 can suppress an increase in the number of instances required for decoding. That is, it is possible to suppress an increase in the load of the decoding process. This makes it possible to suppress an increase in the cost of the decoder. In addition, it is possible to suppress an increase in the processing time of the decoding process.
[0084]
  
 Next, an example of the flow of coding processing executed by the coding apparatus 100 will be described with reference to the flowchart of FIG.
[0085]
 When the coding process is started, the model conversion unit 111 of the coding device 100 executes the conversion process in step S101 to convert a plurality of point cloud models included in the point cloud into one group. Further, the model transformation unit 111 generates transformation information about the transformation.
[0086]
 In step S102, the patch decomposition unit 112 projects (a group of) a plurality of point cloud models collected in step S101 onto a two-dimensional plane and decomposes them into patches. At that time, the patch disassembling unit 112 may set the projection plane of each point cloud model by the method as described above in, for example, . By doing so, the patch decomposition unit 112 can project each point cloud model onto a more appropriate projection plane, so that the reduction in coding efficiency due to the inefficiency of the projection of the point cloud model is suppressed. (The coding efficiency can be improved).
[0087]
 Further, the patch disassembling unit 112 may generate projection plane information related to the setting of the projection plane by the method described above in, for example, and transmit it to the decoding side. By doing so, the decoding side can correctly decode (correctly reconstruct the point cloud).
[0088]
 In addition, the patch disassembly unit 112 generates auxiliary patch information about the disassembly.
[0089]
 In step S103, the auxiliary patch information compression unit 114 compresses (encodes) the auxiliary patch information generated in step S102.
[0090]
 In step S104, the packing unit 113 executes the packing process, arranges each patch of the position information and the attribute information generated in step S102 on the two-dimensional image, and packs it as a video frame. In addition, the packing unit 113 generates model information and an occupancy map. Further, the packing unit 113 performs a dilation process on the color video frame. In addition, the packing unit 113 generates control information regarding such packing.
[0091]
 In step S105, the BB information generation unit 119 executes the BB information generation process to generate BB information including the conversion information generated in step S101, the model information generated in step S104, and the like.
[0092]
 In step S106, the video coding unit 115 encodes the geometry video frame generated in step S104 by a coding method for a two-dimensional image. At that time, the video coding unit 115 encodes the geometry video frame according to the setting of the coding unit (coding unit area setting) that can be independently decoded in which the patch of each point cloud model is arranged in step S104. That is, for example, when slices or tiles are set, the geometry video frame is encoded independently for each slice or tile.
[0093]
 In step S107, the video coding unit 116 encodes the color video frame generated in step S104 by the coding method for the two-dimensional image. At that time, the video coding unit 116 encodes the color video frame according to the setting of the coding unit (coding unit area setting) that can be independently decoded in which the patch of each point cloud model is arranged in step S104. That is, for example, when slices or tiles are set, the color video frame is encoded independently for each slice or tile.
[0094]
 In step S108, the OMap coding unit 117 encodes the occupancy map generated in step S104 by a predetermined coding method.
[0095]
 In step S109, the multiplexer 118 is generated in the various information generated as described above (for example, encoded data of the auxiliary patch information generated in step S103, control information regarding packing generated in step S104, and step S105). (BB information, coded data of geometry video frame generated in step S106, coded data of color video frame generated in step S107, coded data of occupancy map generated in step S108, etc.) are multiplexed. , Generate a bit stream containing this information.
[0096]
 In step S110, the multiplexer 118 outputs the bitstream generated in step S109 to the outside of the coding device 100.
[0097]
 When the process of step S110 is completed, the coding process is completed.
[0098]
  
 Next, an example of the flow of conversion processing executed in step S101 of FIG. 5 will be described with reference to the flowchart of FIG.
[0099]
 When the conversion process is started, the model conversion unit 111 executes the bounding box setting process in step S131 to set the bounding box of each point cloud model.
[0100]
 In step S132, the model conversion unit 111 converts the parameters of each bounding box (each point cloud model) set in step S131. For example, the model conversion unit 111 converts parameters such as shift, rotation, time stamp, scale, and frame rate as described above in and the like. By doing so, the model transformation unit 111 can suppress an increase in the number of bit streams while suppressing a decrease in coding efficiency even when a plurality of point cloud models exist in the point cloud. Therefore, the model transformation unit 111 can suppress an increase in the number of instances required for decoding. That is, it is possible to suppress an increase in the load of the decoding process. This makes it possible to suppress an increase in the cost of the decoder. In addition, it is possible to suppress an increase in the processing time of the decoding process.
[0101]
 In step S133, the model transformation unit 111 generates transformation information of the parameters of each bounding box as described above in and the like. For example, the model transformation unit 111 generates transformation information indicating the ratio, the difference, etc. before and after the transformation. By doing so, the model transformation unit 111 can correctly perform the inverse transformation on the decoding side.
[0102]
 When the process of step S133 is completed, the conversion process is completed.
[0103]
  
 Next, an example of the flow of the bounding box setting process executed in step S131 of FIG. 6 will be described with reference to the flowchart of FIG. 7.
[0104]
 When the bounding box setting process is started, the model transformation unit 111 derives the normal vector of the part of each point cloud model included in the point cloud in step S141.
[0105]
 In step S142, the model transformation unit 111 sets the bounding box so that the one having a large distribution in the normal direction derived in step S141 matches the orthogonal projection vector. When the process of step S142 is completed, the bounding box setting process is completed, and the process returns to FIG.
[0106]
  
 Next, an example of the flow of packing process executed in step S104 of FIG. 5 will be described with reference to the flowchart of FIG.
[0107]
 When the packing process is started, the packing unit 113 sets the coding unit area according to the point cloud model in step S151, and allocates different coding unit areas to each point cloud model.
[0108]
 In step S152, the packing unit 113 searches for and arranges the optimum position in the coding unit region assigned to the point cloud model in step S151 for each patch of each point cloud model.
[0109]
 That is, the packing unit 113 converts patches belonging to the same point cloud model into the same coding units that can be independently decoded by the method described above in and the like. Place to generate geometry video frames and color video frames. By doing so, it is possible to realize so-called "partial decoding" in which only a part of the point cloud model is decoded and reconstructed on the decoding side.
[0110]
 In step S153, the packing unit 113 generates model information about the arrangement of the point cloud model in step S152 by the method described above in . By doing so, it is possible to realize so-called "partial decoding" in which only a part of the point cloud model is decoded and reconstructed on the decoding side.
[0111]
 In step S154, the packing unit 113 generates an occupancy map.
[0112]
 In step S155, the packing unit 113 performs a dilation process on the color video frame.
[0113]
 When the process of step S155 is completed, the packing process is completed and the process returns to FIG.
[0114]
  
 Next, an example of the flow of the BB information generation process executed in step S105 of FIG. 5 will be described with reference to the flowchart of FIG.
[0115]
 When the BB information generation process is started, the BB information generation unit 119 acquires the conversion information generated in step S133 of FIG. 6 in step S161.
[0116]
 In step S162, the BB information generation unit 119 acquires the model information generated in step S153 of FIG.
[0117]
 In step S163, the BB information generation unit 119 generates BB information including the conversion information and the model information. This BB information is multiplexed with the coded data and the like by the multiplexer 118 in step S109 (FIG. 5).
[0118]
 When the process of step S163 is completed, the BB information generation process is completed, and the process returns to FIG.
[0119]
 By executing each process as described above, it is possible to suppress an increase in the number of bit streams and suppress an increase in the load of the decoding process.
[0120]
 <4. Second Embodiment>
  
 FIG. 10 is a block diagram showing an example of a configuration of a decoding device which is an aspect of an image processing device to which the present technology is applied. The decoding device 200 shown in FIG. 10 decodes the encoded data obtained by projecting 3D data such as a point cloud onto a two-dimensional plane by a decoding method for a two-dimensional image and projects it into a three-dimensional space. (Decoding device to which a video-based approach is applied). For example, the decoding device 200 decodes the bit stream generated by encoding the point cloud by the coding device 100 (FIG. 4), and reconstructs the point cloud.
[0121]
 It should be noted that FIG. 10 shows the main things such as the processing unit and the data flow, and not all of them are shown in FIG. That is, in the decoding device 200, there may be a processing unit that is not shown as a block in FIG. 10, or there may be a processing or data flow that is not shown as an arrow or the like in FIG. This also applies to other figures illustrating the processing unit and the like in the decoding device 200.
[0122]
 As shown in FIG. 10, the decoding device 200 includes a demultiplexer 211, an auxiliary patch information decoding unit 212, a point cloud model selection unit 213, a video decoding unit 214, a video decoding unit 215, an OMap decoding unit 216, and an unpacking unit 217. And has a 3D reconstruction unit 218.
[0123]
 The demultiplexer 211 performs processing related to data demultiplexing. For example, the demultiplexer 211 acquires a bitstream input to the decoding device 200. This bit stream is supplied from, for example, the encoding device 100. The demultiplexer 211 demultiplexes this bit stream, extracts the encoded data of the auxiliary patch information, and supplies it to the auxiliary patch information decoding unit 212. Further, the demultiplexer 211 extracts the coded data of the geometry video frame from the bit stream by demultiplexing and supplies it to the video decoding unit 214. Further, the demultiplexer 211 extracts the coded data of the color video frame from the bit stream by demultiplexing and supplies it to the video decoding unit 215. Further, the demultiplexer 211 extracts the encoded data of the occupancy map from the bit stream by demultiplexing and supplies it to the OMap decoding unit 216. Further, the demultiplexer 211 extracts control information regarding packing from the bit stream by demultiplexing and supplies it to the unpacking unit 217. Further, the demultiplexer 211 extracts BB information from the bit stream by demultiplexing and supplies it to the point cloud model selection unit 213 and the 3D reconstruction unit 218.
[0124]
 The auxiliary patch information decoding unit 212 performs processing related to decoding the coded data of the auxiliary patch information. For example, the auxiliary patch information decoding unit 212 acquires the encoded data of the auxiliary patch information supplied from the demultiplexer 211. Further, the auxiliary patch information decoding unit 212 decodes (extends) the encoded data of the auxiliary patch information included in the acquired data. The auxiliary patch information decoding unit 212 supplies the auxiliary patch information obtained by decoding to the 3D reconstruction unit 218.
[0125]
 The point cloud model selection unit 213 performs processing related to selection of the point cloud model to be partially decoded. For example, the point cloud model selection unit 213 acquires BB information from the demultiplexer 211. Further, the point cloud model selection unit 213 accepts the designation of the point cloud model input by the user or the like based on the model information included in the BB information. For example, the point cloud model selection unit 213 presents the point cloud model included in the model information to the user or the like as an option, and causes the user or the like to select the point cloud model to be decoded. The point cloud model selection unit 213 provides information for designating an area (an independently decodable coding unit) corresponding to the selected point cloud model (where the patch of the point cloud model is arranged) to the video decoding unit 214. , Video decoding unit 215, and OMap decoding unit 216.
[0126]
 The video decoding unit 214 performs processing related to decoding the coded data of the geometry video frame. For example, the video decoding unit 214 acquires the coded data of the geometry video frame supplied from the demultiplexer 211. The video decoding unit 214 decodes the coded data of the geometry video frame by a decoding method for an arbitrary two-dimensional image such as AVC or HEVC.
[0127]
 The video decoding unit 214 can partially decode the region (encoding unit that can be independently decoded) designated by the point cloud model selection unit 213. For example, when the area to be decoded is specified by the point cloud model selection unit 213, the video decoding unit 214 sets the specified area of ​​the coded data of the geometry video frame in any two dimensions such as AVC and HEVC. Decrypt by the decoding method for images. For example, the video decoding unit 214 decodes the frame / slice / tile or the like specified by the point cloud model selection unit 213 of the coded data of the geometry video frame.
[0128]
 As described above, the video decoding unit 214 can partially decode the geometry video frame. The video decoding unit 214 supplies the decoded geometry video frame (or a part of the region thereof) to the unpacking unit 217.
[0129]
 The video decoding unit 215 performs processing related to decoding the coded data of the color video frame. For example, the video decoding unit 215 acquires the coded data of the color video frame supplied from the demultiplexer 211. The video decoding unit 215 decodes the coded data of the color video frame by a decoding method for an arbitrary two-dimensional image such as AVC or HEVC.
[0130]
 The video decoding unit 215 can partially decode the region (encoding unit that can be independently decoded) specified by the point cloud model selection unit 213. For example, when the area to be decoded is specified by the point cloud model selection unit 213, the video decoding unit 215 can use the specified area of ​​the coded data of the color video frame in any two dimensions such as AVC and HEVC. Decrypt by the decoding method for images. For example, the video decoding unit 215 decodes the frame / slice / tile or the like specified by the point cloud model selection unit 213 of the coded data of the color video frame.
[0131]
 As described above, the video decoding unit 215 can partially decode the color video frame. The video decoding unit 215 supplies the color video frame (or a part of the region thereof) obtained by decoding to the unpacking unit 217.
[0132]
 The OMap decoding unit 216 performs processing related to decoding the coded data of the occupancy map. For example, the OMap decoding unit 216 acquires the encoded data of the occupancy map supplied from the demultiplexer 211. The OMap decoding unit 216 decodes the coded data of the occupancy map by an arbitrary decoding method corresponding to the coding method.
[0133]
 The OMap decoding unit 216 can partially decode the area (encoding unit that can be independently decoded) specified by the point cloud model selection unit 213. For example, when the area to be decoded is specified by the point cloud model selection unit 213, the OMap decoding unit 216 decodes the specified area of ​​the coded data of the occupancy map arbitrarily according to the coding method. Decrypt by method. For example, the OMap decoding unit 216 decodes the frame / slice / tile or the like specified by the point cloud model selection unit 213 of the encoded data of the occupancy map.
[0134]
 As described above, the OMap decoding unit 216 can partially decode the occupancy map. The OMap decoding unit 216 supplies the Occupancy map (or a part of the region thereof) obtained by decoding to the unpacking unit 217.
[0135]
 The unpacking unit 217 performs a process related to unpacking. For example, the unpacking unit 217 acquires a geometry video frame from the video decoding unit 214, a color video frame from the video decoding unit 215, and an occupancy map from the OMap decoding unit 216. Further, the unpacking unit 217 unpacks the geometry video frame and the color video frame based on the control information regarding the packing. The unpacking unit 217 is a 3D reconstruction unit that displays position information (Geometry) data (geometry patch, etc.) and attribute information (Texture) data (texture patch, etc.) obtained by unpacking, and an occupancy map. Supply to 218.
[0136]
 The 3D reconstruction unit 218 performs processing related to reconstruction of the point cloud. For example, the 3D reconstruction unit 218 has BB information supplied from the demultiplexer 211, auxiliary patch information supplied from the auxiliary patch information decoding unit 212, and position information (Geometry) data supplied from the unpacking unit 217. Reconstruct the point cloud based on (geometry patch, etc.), attribute information (Texture) data (texture patch, etc.), and occupancy map.
[0137]
 For example, the 3D reconstruction unit 218 identifies the projection plane corresponding to each point cloud model in the group based on the projection plane information, and reconstructs the point cloud from a patch or the like using the projection plane. Therefore, since the decoding device 200 can project each point cloud model onto a more appropriate projection plane, it is possible to suppress the reduction in coding efficiency due to the inefficiency of the projection of the point cloud model (reference numeral). The efficiency of conversion can be improved).
[0138]
 Further, for example, the 3D reconstruction unit 218 can reversely convert the reconstructed point cloud model by using the conversion information included in the BB information. Therefore, the 3D reconstruction unit 218 can perform the inverse conversion so as to correctly correspond to the conversion process performed on the coding side.
[0139]
 The 3D reconstruction unit 218 outputs the reconstructed point cloud to the outside of the decoding device 200. This point cloud is, for example, supplied to a display unit and imaged, and the image is displayed, recorded on a recording medium, or supplied to another device via communication.
[0140]
 With such a configuration, the decoding device 200 can correctly decode a plurality of point cloud models collected in one bit stream. Therefore, the decoding device 200 can suppress an increase in the number of bit streams while suppressing a decrease in coding efficiency even when a plurality of point cloud models exist in the point cloud. Therefore, the decoding device 200 can suppress an increase in the number of instances required for itself. That is, it is possible to suppress an increase in the load of the decoding process. This makes it possible to suppress an increase in the cost of the decoder. In addition, it is possible to suppress an increase in the processing time of the decoding process.
[0141]
  
 Next, an example of the flow of the decoding process executed by the decoding device 200 will be described with reference to the flowchart of FIG.
[0142]
 When the decoding process is started, the demultiplexer 211 of the decoding device 200 demultiplexes the bitstream in step S201.
[0143]
 In step S202, the auxiliary patch information decoding unit 212 decodes the auxiliary patch information extracted from the bit stream in step S201.
[0144]
 In step S203, the point cloud model selection unit 213 accepts the designation of the point cloud model to be decoded.
[0145]
 In step S204, the point cloud model selection unit 213 arranges an independently decodable coding unit (that is, a patch of the designated point cloud model) corresponding to the designation of the received point cloud model based on the BB information. Select an independently decodable coding unit).
[0146]
 When partial decoding is not performed, that is, when the entire video frame is decoded, the processes of steps S203 and S204 may be omitted.
[0147]
 In step S205, the video decoding unit 214 decodes the coded data of the geometry video frame (position information video frame) extracted from the bit stream in step S201.
[0148]
 When performing partial decoding, that is, when the designation of the point cloud model to be decoded is accepted in step S203 and the coding unit to be decoded is selected in step S204, the video decoding unit 214 of the geometry video frame, step S204 Decodes the coding unit (eg, frame / slice / tile, etc.) selected in.
[0149]
 In step S206, the video decoding unit 215 decodes the coded data of the color video frame (video frame of the attribute information) extracted from the bit stream in step S201.
[0150]
 When performing partial decoding, the video decoding unit 215 decodes the coding unit (for example, frame / slice / tile, etc.) selected in step S204 of the color video frame.
[0151]
 In step S207, the OMap decoding unit 216 decodes the encoded data of the occupancy map extracted from the bit stream in step S201.
[0152]
 When performing partial decoding, the OMap decoding unit 216 decodes the coding unit (for example, frame / slice / tile, etc.) selected in step S204 of the occupancy map.
[0153]
 In step S208, the unpacking unit 217 unpacks. For example, the unpacking unit 217 unpacks the geometry video frame obtained by decoding the coded data in step S205 to generate a geometry patch. Further, the unpacking unit 217 unpacks the color video frame obtained by decoding the coded data in step S206 to generate a texture patch. Further, the unpacking unit 217 unpacks the occupancy map obtained by decoding the coded data in step S207, and extracts the occupancy map corresponding to the geometry patch and the texture patch.
[0154]
 In step S209, the 3D reconstruction unit 218 uses the point cloud (point cloud) based on the auxiliary patch information obtained in step S202, the geometry patch, texture patch, occupancy map, etc. obtained in step S208. Rebuild the model).
[0155]
 In step S210, the 3D reconstruction unit 218 performs an inverse conversion process, which is an inverse process of the conversion process performed on the coding side for each reconstructed point cloud model based on the conversion information included in the BB information. I do.
[0156]
 When the process of step S210 is completed, the decoding process is completed.
[0157]
 By executing each process as described above, the decoding device 200 can suppress an increase in the number of bit streams and suppress an increase in the load of the decoding process.
[0158]
 <5. Local projection plane control> If the
 point cloud model has sparse points, the projection plane set for the entire point cloud model is the part where the points included in the point cloud model are dense (partial point cloud). It is not always the optimum projection direction for the model). That is, there is a possibility that the partial point cloud model is projected in an inefficient direction and the coding efficiency is reduced.
[0159]
  
 Therefore, as shown in the first stage from the top of the table in FIG. 12, the projection plane of the point cloud model may be locally controlled. For example, as shown in the second row from the top of the table in FIG. 12, a point cloud model including a portion with sparse points is divided into a plurality of partial point cloud models, and a projection plane is set for each partial point cloud model. You may do so.
[0160]
 The method of dividing this projection model is arbitrary. For example, as shown in the table of FIG. 12, it may be divided into dense point clouds included in the point cloud model and modeled as a partial point cloud.
[0161]
 For example, as shown on the far left of FIG. 13, the point cloud model corresponding to the bounding box 312 includes the partial point cloud model 311-1 and the partial point cloud model 311 in which the points are sparse and the points are dense. Suppose it has -2. In such a case, the partial point cloud model 311-1 and the partial point cloud model 311-2 are divided, and the projection planes are set independently for each. For example, as shown second from the left in FIG. 13, a projection plane 313-1 in 6 directions orthogonal to the partial point cloud model 311-1 is set, and 6 orthogonal to the partial point cloud model 311-2. The projection plane 313-2 in the direction is set. Then, each of them is projected onto the projection plane to generate a patch, arranged in a two-dimensional image, and a color video frame (Texture) 314 and a geometry video frame (Depth) are arranged so as to be displayed third from the left in FIG. Generate 315, and Occupancy Map 315. They are then encoded to generate a 2D bitstream 317 and an occupancy map 318. At the time of decoding, these reverse processes are performed.
[0162]
 By doing so, the projection plane can be set for each partial point cloud model, so that each partial point cloud model can be projected onto a more appropriate projection plane. Therefore, it is possible to suppress the reduction of the coding efficiency due to the inefficiency of the projection of the point cloud model (the coding efficiency can be improved).
[0163]
 For example, as shown in the table of FIG. 12, the projection planes in the six orthogonal directions of each partial point cloud model may be rotated. By doing so, it is possible to relatively easily make the projection surface a more appropriate surface (a surface capable of performing efficient projection) with respect to the partial point cloud model.
[0164]
 Further, for example, as shown in the table of FIG. 12, an arbitrary projection plane may be set (added). By doing so, the degree of freedom of the projection surface of each partial point cloud model can be improved, and more efficient projection can be expected (a further improvement in coding efficiency can be expected). ..
[0165]
   When the
 projection plane is set for each partial point cloud model in this way (when it is not the projection plane of the entire point cloud model), when decoding (when reconstructing the point cloud) , Information on the projection plane is also needed. Therefore, as shown in the second row from the top of the table of FIG. 12, the projection plane information which is the information about the projection plane set as described above is generated and transmitted to the decoding side (signaling of the projection plane information).
[0166]
 For example, when projection planes (projection plane 313-1 and projection plane 313-2) are set for each partial point cloud model as shown in the second from the left in FIG. 13, projection plane information 321 is generated for those projection planes. To.
[0167]
 This projection plane information may include any information as long as it is information necessary for specifying the projection plane, as shown in the table of FIG. 12, for example. For example, in the case of FIG. 13, the projection surface information 321 includes information (Rotate x / y / z) indicating the amount of rotation of the projection surface. Of course, in addition to this, for example, information indicating the amount of shift of the projection plane in the six orthogonal directions (Shift x / y / z), information indicating the amount of change in scale (Scale x / y / z), etc. are included. You may. Further, for example, the projection plane information 321 may include information indicating the added arbitrary projection plane.
[0168]
 Further, for example, as shown in the table of FIG. 12, this projection plane information may be stored in the occupancy map. For example, in the case of FIG. 13, the projection plane information 321 is stored in the occupancy map 316 (that is, the occupancy map 318). At the time of coding, this projection surface information 321 is generated and stored in the occupancy map 16, and at the time of decoding, the projection surface information 321 is read out from the occupancy map 316 to reconstruct the point cloud. It is used for.
[0169]
 <6. Third Embodiment>
  
 Next, a configuration for realizing the above method will be described. FIG. 14 is a block diagram showing an example of a configuration of a coding device, which is an aspect of an image processing device to which the present technology is applied. The coding device 400 shown in FIG. 14 is the same device as the coding device 100 (FIG. 4), and is obtained by projecting 3D data such as a point cloud onto a two-dimensional plane and using a coding method for a two-dimensional image. It is a device that performs coding (a coding device that applies a video-based approach).
[0170]
 It should be noted that FIG. 14 shows the main things such as the processing unit and the data flow, and not all of them are shown in FIG. That is, in the coding apparatus 400, there may be a processing unit that is not shown as a block in FIG. 14, or there may be a processing or data flow that is not shown as an arrow or the like in FIG. This also applies to other figures for explaining the processing unit and the like in the coding apparatus 400.
[0171]
 As shown in FIG. 14, the coding apparatus 400 includes a patch decomposition unit 411, a packing unit 412, an auxiliary patch information compression unit 413, a video coding unit 414, a video coding unit 415, an OMap coding unit 416, and a multiplexer 417. Has.
[0172]
 The patch decomposition unit 411 performs processing related to decomposition of 3D data. For example, the patch decomposition unit 411 acquires the point cloud input to the coding device 400. Further, the patch decomposition unit 411 decomposes the acquired point cloud into a plurality of segmentations, projects the point cloud onto a two-dimensional plane for each segmentation, and generates a patch of position information and a patch of attribute information. The patch disassembling unit 411 supplies information about each generated patch to the packing unit 412. Further, the patch disassembly unit 411 supplies auxiliary patch information, which is information related to the disassembly, to the auxiliary patch information compression unit 413.
[0173]
 The packing unit 412 performs processing related to data packing. For example, the packing unit 412 acquires information on a patch of position information (Geometry) indicating the position of a point from the patch disassembly unit 411 and information on a patch of attribute information (Texture) such as color information added to the position information. To do.
[0174]
 Further, the packing unit 412 arranges each acquired patch on the two-dimensional image and packs it as a video frame. For example, the packing unit 412 arranges a patch of position information on a two-dimensional image and packs it as a geometry video frame. Further, for example, the packing unit 412 arranges a patch of attribute information on a two-dimensional image and packs it as a color video frame. In addition, the packing unit 412 generates an occupancy map corresponding to these video frames. Further, the packing unit 412 performs a dilation process on the color video frame.
[0175]
 The packing unit 412 supplies the geometry video frame thus generated to the video coding unit 414. Further, the packing unit 412 supplies the color video frame thus generated to the video coding unit 415. Further, the packing unit 412 supplies the occupancy map thus generated to the OMap coding unit 416. Further, the packing unit 412 supplies control information regarding such packing to the multiplexer 417.
[0176]
 The auxiliary patch information compression unit 413 performs processing related to compression of auxiliary patch information. For example, the auxiliary patch information compression unit 413 acquires the data supplied from the patch decomposition unit 411. The auxiliary patch information compression unit 413 encodes (compresses) the auxiliary patch information included in the acquired data. The auxiliary patch information compression unit 413 supplies the obtained coded data of the auxiliary patch information to the multiplexer 417.
[0177]
 The video coding unit 414 performs processing related to coding of a video frame of position information (Geometry). For example, the video coding unit 414 acquires a geometry video frame supplied from the packing unit 412. Further, the video coding unit 414 encodes the acquired geometry video frame by a coding method for an arbitrary two-dimensional image such as AVC or HEVC. The video coding unit 414 supplies the coded data (coded data of the geometry video frame) obtained by the coding to the multiplexer 417.
[0178]
 The video coding unit 415 performs processing related to coding of a video frame of attribute information (Texture). For example, the video coding unit 415 acquires a color video frame supplied from the packing unit 412. Further, the video coding unit 415 encodes the acquired color video frame by a coding method for an arbitrary two-dimensional image such as AVC or HEVC. The video coding unit 415 supplies the coded data (color video frame coded data) obtained by the coding to the multiplexer 417.
[0179]
 The OMap coding unit 416 performs processing related to coding the occupancy map. For example, the OMap coding unit 416 acquires the occupancy map supplied from the packing unit 412. Further, the OMap coding unit 416 encodes the acquired occupancy map by an arbitrary coding method such as arithmetic coding. The OMap coding unit 416 supplies the coding data (coded data of the occupancy map) obtained by the coding to the multiplexer 417.
[0180]
 The multiplexer 417 performs processing related to multiplexing. For example, the multiplexer 417 acquires the encoded data of the auxiliary patch information supplied from the auxiliary patch information compression unit 413. In addition, the multiplexer 417 acquires control information regarding packing supplied from the packing unit 412. Further, the multiplexer 417 acquires the coded data of the geometry video frame supplied from the video coding unit 414. Further, the multiplexer 417 acquires the coded data of the color video frame supplied from the video coding unit 415. Further, the multiplexer 417 acquires the encoded data of the occupancy map supplied from the OMap coding unit 416.
[0181]
 The multiplexer 417 multiplexes the acquired information to generate a bitstream. The multiplexer 417 outputs the generated bit stream to the outside of the coding apparatus 400.
[0182]
  
 FIG. 15 is a block diagram showing a main configuration example of the patch disassembly unit 411 of FIG. As shown in FIG. 15, the patch disassembling unit 411 includes a model dividing unit 431, a partial point cloud model projection surface setting unit 432, a projection surface information generation unit 433, a partial point cloud model projection unit 434, and a projection surface information addition unit. It has 435.
[0183]
 The model division unit 431 performs processing related to division of the point cloud model. For example, the model dividing unit 431 acquires the point cloud (point cloud model) input to the coding device 400. Further, the model dividing unit 431 may be described in, for example, <5. By the method as described above in of , the part where the points included in the point cloud model are dense is set as the partial point cloud model, and the partial point cloud model is used. Divide the point cloud model for each. The model dividing unit 431 supplies the set information of each partial point cloud model to the partial point cloud model projection plane setting unit 432.
[0184]
 The partial point cloud model projection plane setting unit 432 performs processing related to setting the projection plane for each partial point cloud model. For example, the partial point cloud model projection plane setting unit 432 acquires information about the partial point cloud model supplied from the model dividing unit 431. Further, the partial point cloud model projection plane setting unit 432 may be set to, for example, <5. The projection plane is set for each of the partial point cloud models by the method as described above in of . The partial point cloud model projection surface setting unit 432 supplies information regarding the projection surface for each partial point cloud model set in this way to the projection surface information generation unit 433 and the partial point cloud model projection unit 434.
[0185]
 The projection surface information generation unit 433 performs processing related to the generation of projection surface information. For example, the projection surface information generation unit 433 acquires information on the projection surface set for each partial point cloud model, which is supplied from the partial point cloud model projection surface setting unit 432. Further, the projection plane information generation unit 433 is described in, for example, <5. The projection plane information for each projection plane is generated by the method as described above in of Local projection plane control>. The projection surface information generation unit 433 supplies the generated projection surface information to the projection surface information addition unit 435.
[0186]
 The partial point cloud model projection unit 434 performs processing related to projection for each partial point cloud model. For example, the partial point cloud model projection unit 434 acquires information on the projection plane set for each partial point cloud model, which is supplied from the partial point cloud model projection plane setting unit 432. Further, the partial point cloud model projection unit 434 uses the acquired information on the projection surface, for example, <5. Each subpoint cloud model is projected onto the projection plane for each small area by the method described above in of , and a patch is generated. The partial point cloud model projection unit 434 supplies the generated patch to the projection surface information addition unit 435.
[0187]
 The projection surface information addition unit 435 performs processing related to addition of projection surface information. For example, the projection surface information addition unit 435 acquires the projection surface information supplied from the projection surface information generation unit 433. Further, the projection surface information addition unit 435 acquires a patch or the like supplied from the partial point cloud model projection unit 434. The projection surface information addition unit 435 may be described in, for example, <5. The patch acquired from the partial point cloud model projection unit 434 by the method described above in of Local projection surface control>, and the projection surface related to the projection surface used to generate the patch. Add information. Further, the projection surface information addition unit 435 supplies a patch to which the projection surface information is added to the packing unit 412. The packing unit 412 stores the projection plane information added to the patch in the occupancy map.
[0188]
 With such a configuration, the coding apparatus 400 can set the projection plane independently of the others for each of the plurality of partial point cloud models of the point cloud and project the projection plane. By doing so, each partial point cloud model can be projected onto a more appropriate projection plane. Therefore, it is possible to suppress the reduction of the coding efficiency due to the inefficiency of the projection of the point cloud model (the coding efficiency can be improved).
[0189]
  
 Next, an example of the flow of coding processing executed by the coding apparatus 400 will be described with reference to the flowchart of FIG.
[0190]
 When the coding process is started, the patch decomposition unit 411 of the coding device 400 executes the patch decomposition process in step S401 to decompose the point cloud model into patches. In addition, the patch disassembly unit 112 generates auxiliary patch information about the disassembly.
[0191]
 In step S402, the auxiliary patch information compression unit 413 compresses (encodes) the auxiliary patch information generated in step S401.
[0192]
 In step S403, the packing unit 412 executes the packing process, arranges each patch of the position information and the attribute information generated in step S401 on the two-dimensional image, and packs it as a video frame. In addition, the packing unit 412 generates model information and an occupancy map. Further, the packing unit 412 performs a dilation process on the color video frame. In addition, the packing unit 412 generates control information regarding such packing.
[0193]
 In step S404, the video coding unit 414 encodes the geometry video frame generated in step S403 by a coding method for a two-dimensional image.
[0194]
 In step S405, the video coding unit 415 encodes the color video frame generated in step S403 by the coding method for the two-dimensional image.
[0195]
 In step S406, the OMap coding unit 416 encodes the occupancy map generated in step S403 by a predetermined coding method.
[0196]
 In step S407, the multiplexer 417 is generated in the various information generated as described above (for example, encoded data of the auxiliary patch information generated in step S402, control information regarding packing generated in step S403, and step S404). The coded data of the geometry video frame, the coded data of the color video frame generated in step S405, the coded data of the occupancy map generated in step S406, etc.) are multiplexed to obtain a bit stream containing such information. Generate.
[0197]
 In step S408, the multiplexer 417 outputs the bitstream generated in step S407 to the outside of the coding device 400.
[0198]
 When the process of step S408 is completed, the coding process is completed.
[0199]
  
 Next, an example of the flow of patch disassembly process executed in step S401 of FIG. 16 will be described with reference to the flowchart of FIG.
[0200]
 When the patch disassembly process is started, the model dividing unit 431 may, for example, <5. The point cloud model to be processed including the sparsely pointed portion is divided into a plurality of partial point cloud models by the method described above in of . ..
[0201]
 In step S422, the partial point cloud model projection plane setting unit 432 is, for example, <5. The projection plane of each partial point cloud model set in step S421 is set by the method as described above in of .
[0202]
 In step S423, the projection plane information generation unit 433, for example, <5. The projection plane information about the projection plane of each partial point cloud model set in step S422 is generated by the method as described above in of Local projection plane control>.
[0203]
 In step S424, the partial point cloud model projection unit 434 is, for example, <5. Each subpoint cloud model is projected onto the projection plane set in step S422 by the method described above in of , and a patch is generated.
[0204]
 In step S425, the projection plane information addition unit 435 is described in, for example, <5. The projection plane information generated in step S423 is added (added) to the patch generated in step S424 by the method described above in of Local projection plane control>. The projection plane information given to this patch is stored in the occupancy map by the packing unit 412.
[0205]
 When the process of step S425 is completed, the patch disassembly process is completed, and the process returns to FIG.
[0206]
  
 Next, an example of the flow of division processing executed in step S421 of FIG. 17 will be described with reference to the flowchart of FIG.
[0207]
 When the division process is started, the model division unit 431 identifies a dense point cloud from the histogram of the point cloud model to be processed in step S441.
[0208]
 In step S442, the model dividing unit 431 estimates the projection direction of each of the dense point cloud specified in step S441 based on the normal vector of the points.
[0209]
 In step S443, the model dividing unit 431 uses a dense point cloud whose projection directions estimated in step S442 are different from each other as a partial point cloud model.
[0210]
 When the process of step S443 is completed, the division process is completed, and the process returns to FIG.
[0211]
  
 Next, an example of the flow of packing process executed in step S403 of FIG. 16 will be described with reference to the flowchart of FIG.
[0212]
 When the packing process is started, the packing unit 412 arranges each patch of each partial point cloud model on the two-dimensional image in step S461.
[0213]
 In step S462, the packing unit 412 generates an occupancy map including the projection plane information generated in step S423 of FIG.
[0214]
 In step S463, the packing unit 412 performs a dilation process on the color video frame.
[0215]
 When the process of step S463 is completed, the packing process is completed, and the process returns to FIG.
[0216]
 By executing each process as described above, the projection plane can be set independently for each of the plurality of partial point cloud models of the point cloud, and each partial point cloud model is projected onto a more appropriate projection plane. be able to. Therefore, it is possible to suppress the reduction of the coding efficiency due to the inefficiency of the projection of the point cloud model (the coding efficiency can be improved).
[0217]
 <7. Fourth Embodiment>
  
 FIG. 20 is a block diagram showing an example of a configuration of a decoding device which is an aspect of an image processing device to which the present technology is applied. The decoding device 500 shown in FIG. 20 is the same device as the decoding device 200 (FIG. 10), and two-dimensionally encodes coded data in which 3D data such as a point cloud is projected onto a two-dimensional plane and encoded. It is a device (decoding device to which a video-based approach is applied) that decodes by a decoding method for images and projects it into a three-dimensional space. For example, the decoding device 500 decodes the bit stream generated by encoding the point cloud by the coding device 400 (FIG. 14), and reconstructs the point cloud.
[0218]
 It should be noted that FIG. 20 shows the main things such as the processing unit and the data flow, and not all of them are shown in FIG. 20. That is, in the decoding device 500, there may be a processing unit that is not shown as a block in FIG. 20, or there may be a processing or data flow that is not shown as an arrow or the like in FIG. This also applies to other figures illustrating the processing unit and the like in the decoding device 500.
[0219]
 As shown in FIG. 20, the decoding device 500 includes a demultiplexer 511, an auxiliary patch information decoding unit 512, a video decoding unit 513, a video decoding unit 514, an OMap decoding unit 515, an unpacking unit 516, and a 3D reconstruction unit 517. Have.
[0220]
 The demultiplexer 511 performs processing related to data demultiplexing. For example, the demultiplexer 511 acquires a bitstream input to the decoding device 500. This bitstream is supplied, for example, by the encoding device 400. The demultiplexer 511 demultiplexes this bit stream, extracts the encoded data of the auxiliary patch information, and supplies it to the auxiliary patch information decoding unit 512. Further, the demultiplexer 511 extracts the coded data of the geometry video frame from the bit stream by demultiplexing and supplies it to the video decoding unit 513. Further, the demultiplexer 511 extracts the coded data of the color video frame from the bit stream by demultiplexing and supplies it to the video decoding unit 514. Further, the demultiplexer 511 extracts the encoded data of the occupancy map from the bit stream by demultiplexing and supplies it to the OMap decoding unit 515. Further, the demultiplexer 511 extracts control information regarding packing from the bit stream by demultiplexing and supplies it to the unpacking unit 516.
[0221]
 The auxiliary patch information decoding unit 512 performs processing related to decoding the coded data of the auxiliary patch information. For example, the auxiliary patch information decoding unit 512 acquires the encoded data of the auxiliary patch information supplied from the demultiplexer 511. Further, the auxiliary patch information decoding unit 512 decodes (extends) the encoded data of the auxiliary patch information included in the acquired data. The auxiliary patch information decoding unit 512 supplies the auxiliary patch information obtained by decoding to the 3D reconstruction unit 517.
[0222]
 The video decoding unit 513 performs processing related to decoding the coded data of the geometry video frame. For example, the video decoding unit 513 acquires the coded data of the geometry video frame supplied from the demultiplexer 511. The video decoding unit 513 decodes the coded data of the geometry video frame by a decoding method for an arbitrary two-dimensional image such as AVC or HEVC. The video decoding unit 513 supplies the decoded geometry video frame (or a part of the region thereof) to the unpacking unit 516.
[0223]
 The video decoding unit 514 performs processing related to decoding the coded data of the color video frame. For example, the video decoding unit 514 acquires the coded data of the color video frame supplied from the demultiplexer 511. The video decoding unit 514 decodes the coded data of the color video frame by a decoding method for an arbitrary two-dimensional image such as AVC or HEVC. The video decoding unit 514 supplies the color video frame (or a part of the region thereof) obtained by decoding to the unpacking unit 516.
[0224]
 The OMap decoding unit 515 performs processing related to decoding the coded data of the occupancy map. For example, the OMap decoding unit 515 acquires the encoded data of the occupancy map supplied from the demultiplexer 511. The OMap decoding unit 515 decodes the coded data of the occupancy map by an arbitrary decoding method corresponding to the coding method.
[0225]
 The OMap decoding unit 515 supplies the Occupancy map (or a part of the region thereof) obtained by decoding to the unpacking unit 516.
[0226]
 The unpacking unit 516 performs a process related to unpacking. For example, the unpacking unit 516 acquires a geometry video frame from the video decoding unit 513, a color video frame from the video decoding unit 514, and an occupancy map from the OMap decoding unit 515. Further, the unpacking unit 516 unpacks the geometry video frame and the color video frame based on the control information regarding the packing. The unpacking unit 516 is a 3D reconstruction unit that displays position information (Geometry) data (geometry patch, etc.) and attribute information (Texture) data (texture patch, etc.) obtained by unpacking, and an occupancy map. Supply to 517.
[0227]
 The 3D reconstruction unit 517 performs processing related to reconstruction of the point cloud. For example, the 3D reconstruction unit 517 includes auxiliary patch information supplied from the auxiliary patch information decoding unit 512, position information (Geometry) data (geometry patch, etc.) supplied from the unpacking unit 516, and attribute information (Texture). ) Data (texture patches, etc.) and occupancy maps, etc., to reconstruct the point cloud.
[0228]
 For example, the 3D reconstruction unit 517 identifies a projection surface corresponding to each partial point cloud model based on the projection surface information, and reconstructs the point cloud from a patch or the like using the projection surface. Therefore, the decoding device 500 can reconstruct each partial point cloud model from the patch projected on the more appropriate projection plane, so that the coding efficiency is reduced due to the inefficiency of the projection of the partial point cloud model. Can be suppressed (encoding efficiency can be improved).
[0229]
 The 3D reconstruction unit 517 outputs the reconstructed point cloud to the outside of the decoding device 500. This point cloud is, for example, supplied to a display unit and imaged, and the image is displayed, recorded on a recording medium, or supplied to another device via communication.
[0230]
 With such a configuration, the decoding device 500 can suppress a reduction in coding efficiency even when a plurality of partial point cloud models exist in the point cloud.
[0231]
  
 Next, an example of the flow of the decoding process executed by the decoding device 500 will be described with reference to the flowchart of FIG.
[0232]
 When the decoding process is started, the demultiplexer 511 of the decoding device 500 demultiplexes the bit stream in step S501.
[0233]
 In step S502, the auxiliary patch information decoding unit 512 decodes the auxiliary patch information extracted from the bit stream in step S501.
[0234]
 In step S503, the video decoding unit 513 decodes the coded data of the geometry video frame (position information video frame) extracted from the bit stream in step S501.
[0235]
 In step S504, the video decoding unit 514 decodes the coded data of the color video frame (video frame of the attribute information) extracted from the bit stream in step S501.
[0236]
 In step S505, the OMap decoding unit 515 decodes the encoded data of the occupancy map extracted from the bitstream in step S501. This occupancy map contains the projection plane information described above.
[0237]
 In step S506, the unpacking unit 516 unpacks. For example, the unpacking unit 516 unpacks the geometry video frame obtained by decoding the coded data in step S503 to generate a geometry patch. Further, the unpacking unit 516 unpacks the color video frame obtained by decoding the coded data in step S504 to generate a texture patch. Further, the unpacking unit 516 unpacks the occupancy map obtained by decoding the coded data in step S505, and extracts the occupancy map corresponding to the geometry patch and the texture patch.
[0238]
 In step S507, the 3D reconstruction unit 517 includes the auxiliary patch information obtained in step S502, the geometry patch, the texture patch, and the occupancy map obtained in step S506, and the projection surface information included in the occupancy map. Reconstruct the point cloud (point cloud model) based on the above.
[0239]
 When the process of step S507 is completed, the decoding process is completed.
[0240]
 By executing each process as described above, the decoding device 500 can suppress a decrease in coding efficiency.
[0241]
 <8. Supplementary note>
   The control information
 related to the present technology described in each of the above embodiments may be transmitted from the coding side to the decoding side. For example, control information (for example, enabled_flag) that controls whether or not the application of the present technology described above is permitted (or prohibited) may be transmitted. Further, for example, a control for designating a range (for example, an upper limit or a lower limit of a block size, or both, a slice, a picture, a sequence, a component, a view, a layer, etc.) for permitting (or prohibiting) the application of the present technology described above. Information may be transmitted.
[0242]
   The
 series of processes described above can be executed by hardware or software. When a series of processes are executed by software, the programs constituting the software are installed on the computer. Here, the computer includes a computer embedded in dedicated hardware, a general-purpose personal computer capable of executing various functions by installing various programs, and the like.
[0243]
 FIG. 22 is a block diagram showing a configuration example of the hardware of a computer that executes the above-mentioned series of processes programmatically.
[0244]
 In the computer 900 shown in FIG. 22, the CPU (Central Processing Unit) 901, the ROM (Read Only Memory) 902, and the RAM (Random Access Memory) 903 are connected to each other via the bus 904.
[0245]
 The input / output interface 910 is also connected to the bus 904. An input unit 911, an output unit 912, a storage unit 913, a communication unit 914, and a drive 915 are connected to the input / output interface 910.
[0246]
 The input unit 911 includes, for example, a keyboard, a mouse, a microphone, a touch panel, an input terminal, and the like. The output unit 912 includes, for example, a display, a speaker, an output terminal, and the like. The storage unit 913 is composed of, for example, a hard disk, a RAM disk, a non-volatile memory, or the like. The communication unit 914 includes, for example, a network interface. The drive 915 drives a removable medium 921 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.
[0247]
 In the computer configured as described above, the CPU 901 loads the program stored in the storage unit 913 into the RAM 903 via the input / output interface 910 and the bus 904 and executes the above-described series. Is processed. The RAM 903 also appropriately stores data and the like necessary for the CPU 901 to execute various processes.
[0248]
 The program executed by the computer (CPU901) can be recorded and applied to the removable media 921 as a package media or the like, for example. In that case, the program can be installed in the storage unit 913 via the input / output interface 910 by mounting the removable media 921 in the drive 915.
[0249]
 The program can also be provided via wired or wireless transmission media such as local area networks, the Internet, and digital satellite broadcasting. In that case, the program can be received by the communication unit 914 and installed in the storage unit 913.
[0250]
 In addition, this program can be pre-installed in the ROM 902 or the storage unit 913.
[0251]
   The
 case where this technology is applied to the coding / decoding of point cloud data has been described above, but this technology is not limited to these examples, and the code of 3D data of any standard is used. It can be applied to conversion / decryption. That is, as long as it does not contradict the present technology described above, various processes such as coding / decoding methods and specifications of various data such as 3D data and metadata are arbitrary. In addition, some of the above-mentioned processes and specifications may be omitted as long as they do not conflict with the present technology.
[0252]
 The present technology can be applied to any configuration. For example, this technology is a transmitter or receiver (for example, a television receiver or mobile phone) for satellite broadcasting, cable broadcasting such as cable TV, distribution on the Internet, and distribution to terminals by cellular communication, or It can be applied to various electronic devices such as devices (for example, hard disk recorders and cameras) that record images on media such as optical disks, magnetic disks, and flash memories, and reproduce images from these storage media.
[0253]
 Further, for example, in the present technology, a processor as a system LSI (Large Scale Integration) or the like (for example, a video processor), a module using a plurality of processors (for example, a video module), a unit using a plurality of modules (for example, a video unit) Alternatively, it can be implemented as a configuration of a part of the device, such as a set (for example, a video set) in which other functions are added to the unit.
[0254]
 Further, for example, the present technology can also be applied to a network system composed of a plurality of devices. For example, the present technology may be implemented as cloud computing that is shared and jointly processed by a plurality of devices via a network. For example, this technology is implemented in a cloud service that provides services related to images (moving images) to arbitrary terminals such as computers, AV (Audio Visual) devices, portable information processing terminals, and IoT (Internet of Things) devices. You may try to do it.
[0255]
 In the present specification, the system means a set of a plurality of components (devices, modules (parts), etc.), and it does not matter whether all the components are in the same housing. Therefore, a plurality of devices housed in separate housings and connected via a network, and a device in which a plurality of modules are housed in one housing are both systems. ..
[0256]
   Systems, equipment, processing departments, etc. to which this technology is applied include, for example, transportation, medical care, crime prevention, agriculture, livestock industry, mining, beauty, factories, home appliances, weather, nature monitoring. It can be used in any field. Moreover, the use is arbitrary.
[0257]
   In
 the present specification, the "flag" is information for identifying a plurality of states, and is not limited to information used for identifying two states of true (1) or false (0). Information that can identify three or more states is also included. Therefore, the value that this "flag" can take may be, for example, 2 values ​​of 1/0 or 3 or more values. That is, the number of bits constituting this "flag" is arbitrary, and may be 1 bit or a plurality of bits. Further, the identification information (including the flag) is assumed to include not only the identification information in the bitstream but also the difference information of the identification information with respect to a certain reference information in the bitstream. In, the "flag" and "identification information" include not only the information but also the difference information with respect to the reference information.
[0258]
 Further, various information (metadata, etc.) regarding the coded data (bit stream) may be transmitted or recorded in any form as long as it is associated with the coded data. Here, the term "associate" means, for example, to make the other data available (linkable) when processing one data. That is, the data associated with each other may be combined as one data or may be individual data. For example, the information associated with the coded data (image) may be transmitted on a transmission path different from the coded data (image). Further, for example, the information associated with the coded data (image) may be recorded on a recording medium (or another recording area of ​​the same recording medium) different from the coded data (image). Good. Note that this "association" may be a part of the data, not the entire data. For example, an image and information corresponding to the image may be associated with each other in an arbitrary unit such as a plurality of frames, one frame, or a part within the frame.
[0259]
 In addition, in this specification, "synthesize", "multiplex", "add", "integrate", "include", "store", "insert", "insert", "insert". A term such as "" means combining a plurality of objects into one, for example, combining encoded data and metadata into one data, and means one method of "associating" described above.
[0260]
 Further, the embodiment of the present technology is not limited to the above-described embodiment, and various changes can be made without departing from the gist of the present technology.
[0261]
 For example, the configuration described as one device (or processing unit) may be divided and configured as a plurality of devices (or processing units). On the contrary, the configurations described above as a plurality of devices (or processing units) may be collectively configured as one device (or processing unit). Further, of course, a configuration other than the above may be added to the configuration of each device (or each processing unit). Further, if the configuration and operation of the entire system are substantially the same, a part of the configuration of one device (or processing unit) may be included in the configuration of another device (or other processing unit). ..
[0262]
 Further, for example, the above-mentioned program may be executed in any device. In that case, the device may have necessary functions (functional blocks, etc.) so that necessary information can be obtained.
[0263]
 Further, for example, each step of one flowchart may be executed by one device, or may be shared and executed by a plurality of devices. Further, when a plurality of processes are included in one step, the plurality of processes may be executed by one device or may be shared and executed by a plurality of devices. In other words, a plurality of processes included in one step can be executed as processes of a plurality of steps. On the contrary, the processes described as a plurality of steps can be collectively executed as one step.
[0264]
 Further, for example, in a program executed by a computer, the processing of the steps for writing the program may be executed in chronological order in the order described in the present specification, or may be executed in parallel or in calls. It may be executed individually at the required timing such as when it is broken. That is, as long as there is no contradiction, the processing of each step may be executed in an order different from the above-mentioned order. Further, the processing of the step for writing this program may be executed in parallel with the processing of another program, or may be executed in combination with the processing of another program.
[0265]
 Further, for example, a plurality of techniques related to this technique can be independently implemented independently as long as there is no contradiction. Of course, any plurality of the present technologies can be used in combination. For example, some or all of the techniques described in any of the embodiments may be combined with some or all of the techniques described in other embodiments. It is also possible to carry out a part or all of any of the above-mentioned techniques in combination with other techniques not described above.
Code description
[0266]
 100 Encoding device, 111 Model conversion unit, 112 Patch decomposition unit, 113 Packing unit, 114 Auxiliary patch information compression unit, 115 Video coding unit, 116 Video coding unit, 117 OMap coding unit, 118 multiplexer, 119 BB information Generation unit, 200 decoding device, 211 demultiplexer, 212 auxiliary patch information decoding unit, 213 point group model selection unit, 214 video decoding unit, 215 video decoding unit, 216 OMap decoding unit, 217 unpacking unit, 218 3D reconstruction unit , 400 Encoding device, 411 patch decomposition section, 412 packing section, 413 auxiliary patch information compression section, 414 video coding section, 415 video coding section, 416 OMap coding section, 417 multiplexer, 431 model division section, 432 section. Point group model projection surface setting unit, 433 projection surface information generation unit, 434 partial point group model projection unit, 435 projection surface information addition unit, 500 decoding device, 511 demultiplexer, 512 auxiliary patch information decoding unit, 513 video decoding unit, 514 video decoding unit, 515 OMap decoding unit, 516 unpacking unit, 517 3D reconstruction unit
The scope of the claims
[Claim 1]
 The conversion unit that converts the parameters related to the plurality of point group models of the point cloud
 and the two-dimensional plane image on which the plurality of point group models whose parameters are converted by the conversion unit are encoded are encoded into the two-dimensional plane image. An
 image processing apparatus including a coding unit for generating a bit stream including the coding data of the above and conversion information which is information related to the conversion of the parameter by the conversion unit.
[Claim 2]

 The image processing apparatus according  to claim 1, wherein the conversion unit converts the coordinates of the point cloud model as the parameter .
[Claim 3]

 The image processing apparatus according to claim 2,  wherein the conversion unit performs shift and / or rotation of the coordinates as conversion of the coordinates .
[Claim 4]

 The image processing apparatus according  to claim 1, wherein the conversion unit converts the time of the point cloud model as the parameter .
[Claim 5]

 The image processing apparatus according  to claim 1, wherein the conversion unit converts the scale of the point cloud model as the parameter .
[Claim 6]

 The image processing apparatus according  to claim 1, wherein the conversion unit converts the frame rate of the point cloud model as the parameter .
[Claim 7]

 The image processing apparatus according to claim 1  , wherein the conversion information includes information indicating the amount of change of the parameter converted by the conversion unit.
[Claim 8]
 The patch of the point cloud model is arranged in a different area for each point cloud model, and further includes a packing unit for packing as a video frame, and the
 coding unit is the video in which the patch is packed by the packing unit.
 The image processing apparatus according to claim 1 , wherein the frame is encoded to generate the bit stream including model information which is information about the point cloud model .
[Claim 9]

 The image processing apparatus according to claim 8,  wherein the region is a coding unit that can be independently decoded .
[Claim 10]

 The image processing apparatus according to claim 8,  wherein the model information includes the number of point cloud models included in the bit stream and information indicating an area in which patches of each point cloud model are arranged .
[Claim 11]
 The parameters related to the plurality of point group models of the point cloud
 are converted, the two-dimensional plane image on which the plurality of point group models to which the parameters are converted are projected is encoded, and the encoded data of the two-dimensional plane image and the said An
 image processing method that generates a bit stream that includes conversion information that is information about parameter conversion .
[Claim 12]
 A decoding unit that decodes a bit stream and generates a two-dimensional plane image on which a plurality of point group models are projected and conversion information that is information regarding conversion of each parameter of the plurality of point group models, and the
 decoding unit. It
 is provided with a reconstruction unit that reconstructs each of the plurality of point group models from the two-dimensional plane image generated by the above-mentioned method and inversely transforms each parameter of the plurality of point group models based on the conversion information. Image processing device.
[Claim 13]
 The bit stream is decoded to generate a two-dimensional plane image on which a plurality of point group models are projected, and conversion information which is information regarding conversion of each parameter of the plurality of point group models, and the generated two-dimensional image is
 generated. An
 image processing method in which each of the plurality of point group models is reconstructed from a plane image, and each parameter of the plurality of point group models is inversely converted based on the conversion information .
[Claim 14]
 A projection unit that projects a plurality of partial point group models constituting the point group model onto projection planes set independently of each other, and
 a patch of the partial point group model projected on each projection plane by the projection unit. An
 image processing device including a coding unit that encodes an arranged two-dimensional plane image and an occupancy map including projection plane information that is information about the projection plane, and generates a bit stream .
[Claim 15]

 The image processing apparatus according to claim 14  , wherein the projection surface information includes information indicating the amount of rotation of the projection surface .
[Claim 16]
 The image processing apparatus according to claim 14, further
 comprising  a division unit that divides the point cloud model into the plurality of partial point cloud models, and a setting unit that sets a projection plane for each of the plurality of partial point cloud models
.
[Claim 17]
 The division portion
  identifies dense point clouds from the histogram of the point cloud model,
  estimates the projection directions of the dense point clouds based on the normal vector, and estimates the projection directions of the
  dense point clouds, and the dense point clouds whose projection directions are different from each other.
 The image processing apparatus according to claim 16, wherein the image processing apparatus is set as a partial point cloud model .
[Claim 18]

 A two-dimensional plane image in which a  plurality of partial point cloud models constituting the point cloud model are projected onto projection planes set independently of each other , and patches of the partial point cloud model projected on each projection plane are arranged. An
 image processing method for generating a bit stream by encoding a and an occupancy map including projection plane information which is information about the projection plane .
[Claim 19]
 A two-dimensional plane image obtained by decoding a bit stream and projecting a point cloud model, and an occupancy map containing projection plane information which is information about each projection plane of a plurality of partial point cloud models included in the point cloud model. a decoding unit that generates,
 and the two-dimensional plane image generated by the decoding unit, based on said projection plane information included in the occupancy map, a reconstruction unit for reconstructing the point cloud model
 the An image processing device provided.
[Claim 20]
 A two-dimensional plane image obtained by decoding a bit stream and projecting a point cloud model, and an occupancy map containing projection plane information which is information about each projection plane of a plurality of partial point cloud models included in the point cloud model.
 An
 image processing method for reconstructing the point cloud model based on the generated two-dimensional plane image and the projected plane information included in the occupancy map .

Documents

Application Documents

# Name Date
1 202017056284-TRANSLATIOIN OF PRIOIRTY DOCUMENTS ETC. [24-12-2020(online)].pdf 2020-12-24
2 202017056284-STATEMENT OF UNDERTAKING (FORM 3) [24-12-2020(online)].pdf 2020-12-24
3 202017056284-PRIORITY DOCUMENTS [24-12-2020(online)].pdf 2020-12-24
4 202017056284-POWER OF AUTHORITY [24-12-2020(online)].pdf 2020-12-24
5 202017056284-FORM 1 [24-12-2020(online)].pdf 2020-12-24
6 202017056284-DRAWINGS [24-12-2020(online)].pdf 2020-12-24
7 202017056284-DECLARATION OF INVENTORSHIP (FORM 5) [24-12-2020(online)].pdf 2020-12-24
8 202017056284-COMPLETE SPECIFICATION [24-12-2020(online)].pdf 2020-12-24
9 202017056284-Proof of Right [29-01-2021(online)].pdf 2021-01-29
10 202017056284-Proof of Right [05-02-2021(online)].pdf 2021-02-05
11 202017056284-FORM-26 [11-02-2021(online)].pdf 2021-02-11
12 202017056284-Proof of Right [19-02-2021(online)].pdf 2021-02-19
13 202017056284-Proof of Right [03-03-2021(online)].pdf 2021-03-03
14 202017056284-FORM 3 [08-04-2021(online)].pdf 2021-04-08
15 202017056284.pdf 2021-10-19