Abstract: The present technology relates to an information processing device and method, and a program which make it possible to reduce the amount of coding. This information processing device is provided with: an acquisition unit which acquires spatial information relating to the position and size within a parent space of a child space included in the parent space, and intra-child space position information indicating the position within the child space of an object included in the child space; and a calculation unit which on the basis of the spatial information and the intra-child space position information, calculates intra-parent space position information indicating the position of the object within the parent space. The present technology is applicable to a signal processing device.
The present technology relates to information processing devices and methods, and programs, and more particularly to information processing devices, methods, and programs capable of reducing the amount of code.
Background technology
[0002]
The MPEG (Moving Picture Experts Group) -H coding standard, which is standardized as the conventional 3D Audio for fixed viewpoints, is based on the idea that the audio object moves in the space around the listener's position as the origin. (See, for example, Non-Patent Document 1).
[0003]
On the other hand, a technique related to a free viewpoint has been proposed in which the listener can freely move in space, that is, the viewpoint position (listening position) of the listener can be freely moved.
Prior art literature
Non-patent literature
[0004]
Non-Patent Document 1: ISO / IEC 23008-3 Information technology - High efficiency coding and media delivery in heterogeneous environments - Part 3: 3D audio
Outline of the invention
Problems to be solved by the invention
[0005]
For example, in audio reproduction for a free viewpoint, when transmitting the position of each object existing in the space, the position information is encoded with respect to the entire space including all the objects. That is, the position information of the object in the coordinate system including the entire space is encoded.
[0006]
However, there is a bias in the location of objects in the entire space, and objects may be concentrated in a specific area in space. In such a case, if the position information of the object is normalized and encoded for the entire space, the amount of information (code amount) for expressing the area where the object does not actually exist is required, and as a result, The amount of transmission code increases.
[0007]
The present technology has been made in view of such a situation, and makes it possible to reduce the amount of code.
Means to solve problems
[0008]
The information processing device of one aspect of the present technology indicates spatial information regarding the position and size of the child space included in the parent space in the parent space and the position of the object included in the child space in the child space. An acquisition unit that acquires the position information in the child space, and a calculation unit that calculates the position information in the parent space indicating the position of the object in the parent space based on the space information and the position information in the child space. Be prepared.
[0009]
The information processing method or program of one aspect of the present technology includes spatial information regarding the position and size of a child space included in the parent space in the parent space and the position of an object contained in the child space in the child space. This includes a step of acquiring the position information in the child space indicating the above and calculating the position information in the parent space indicating the position of the object in the parent space based on the space information and the position information in the child space.
[0010]
In one aspect of the present technology, spatial information regarding the position and size of the child space included in the parent space in the parent space and the position of the object contained in the child space in the child space are indicated in the child space. The position information is acquired, and the position information in the parent space indicating the position of the object in the parent space is calculated based on the space information and the position information in the child space.
The invention's effect
[0011]
According to one aspect of the present technology, the amount of code can be reduced.
[0012]
The effects described here are not necessarily limited, and may be any of the effects described in the present disclosure.
A brief description of the drawing
[0013]
[Fig. 1] It is a figure explaining a parent space and a child space.
[Fig. 2] Fig. 2 is a diagram for explaining the coordinate system of the parent space and the child space.
[Fig. 3] Fig. 3 is a diagram illustrating quantization.
[Fig. 4] Fig. 4 is a diagram showing an example of a division index and the number of divisions.
[Fig. 5] Fig. 5 is a diagram showing an example of quantization of position information in child space.
[Fig. 6] Fig. 6 is a diagram illustrating a difference in the number of bits used.
[Fig. 7] Fig. 7 is a diagram showing a configuration example of an encoder.
FIG. 8 is a flowchart illustrating a coding process.
[Fig. 9] Fig. 9 is a diagram showing a configuration example of a decoder.
FIG. 10 is a flowchart illustrating a decoding process.
[Fig. 11] Fig. 11 is a diagram showing an example of header information.
[Fig. 12] Fig. 12 is a diagram showing an example of child space information.
[Fig. 13] Fig. 13 is a diagram showing an example of an initial quantization position information transmission request.
[Fig. 14] Fig. 14 is a diagram showing an example of initial quantization position information.
FIG. 15 is a diagram showing an example of an additional bit information transmission request.
[Fig. 16] Fig. 16 is a diagram showing an example of additional bit information.
[Fig. 17] Fig. 17 is a diagram showing a configuration example of a content reproduction system.
[Fig. 18] Fig. 18 is a diagram showing a configuration example of an absolute coordinate position information encoder.
[Fig. 19] Fig. 19 is a diagram showing a configuration example of a server.
[Fig. 20] Fig. 20 is a diagram showing a configuration example of a client.
FIG. 21 is a flowchart illustrating a coding process and a file saving process.
FIG. 22 is a flowchart illustrating a position information acquisition process and a position information transmission process.
FIG. 23 is a flowchart illustrating an additional bit information acquisition process and an additional bit information transmission process.
[Fig. 24] Fig. 24 is a diagram showing an example of a child space.
[Fig. 25] Fig. 25 is a diagram showing an example of a child space.
FIG. 26 is a diagram showing an example of header information.
[Fig. 27] Fig. 27 is a diagram showing an example of child space information.
[Fig. 28] Fig. 28 is a diagram illustrating skipping of frames.
FIG. 29 is a diagram showing an example of an initial quantization position information transmission request.
[Fig. 30] Fig. 30 is a diagram showing an example of initial quantization position information.
FIG. 31 is a flowchart illustrating a position information acquisition process and a position information transmission process.
[Fig. 32] Fig. 32 is a diagram showing a configuration example of a computer.
Mode for carrying out the invention
[0014]
Hereinafter, embodiments to which the present technology is applied will be described with reference to the drawings.
[0015]
This technology converts the position information of an object existing in a predetermined space into the position information in a child space smaller than that space, and the converted position information. The amount of code can be reduced by quantizing.
[0016]
In particular, the present technology has the following features T (1) to T (4).
[0017]
Feature T (1)
Forming a scaled space (child space) that includes the existence range of the object to be encoded from the entire space (parent space)
Feature T (2)
Scale conversion from the parent space to the child space After that, encoding and decoding are performed, and the position information of the decoded child space object is inversely scaled from the child space to the parent space and restored on the parent space.
Feature T (3)
Multiple differences The space is arranged as a child space, and the listener selectively receives and reproduces the desired child space.
T (4) The
decoding side skips from the current reproduction position based on the transmitted skip frame information and interpolation information. Generates the position information of the object that is interpolated to the frame immediately before the frame destination
[0018]
The position information of the object to be encoded in the present technology is not limited to the position information indicating the position of the audio object, but any kind of position information such as the position information indicating the position of the video object on the image or in the virtual space. It may be a thing. An example of encoding and decoding the position information of an audio object will be described below.
[0019]
Then, the present technology will be described in more detail below.
[0020]
First, consider the case where there is a locally concentrated object group region in the parent space PA11 which is the entire space of the three-dimensional absolute coordinate system as shown by the arrow Q11 in FIG. In FIG. 1, each circle represents an object (audio object).
[0021]
Here, the parent space PA11 is a space represented by absolute coordinates, and the position in the parent space PA11 is represented by absolute coordinates. In other words, the parent space PA11 is the coordinate space of the absolute coordinate system. The parent space may be a coordinate space of another coordinate system different from the absolute coordinate system such as a polar coordinate system, but the case where the parent space is the coordinate space of the absolute coordinate system will be described below as an example. continue.
[0022]
For example, when expressing the position of an object in the parent space PA11, it is common to express the position of the object seen from the origin O of the three-dimensional absolute coordinate system in the parent space PA11 in three-dimensional absolute coordinates.
[0023]
On the other hand, in the present technology, the device on the coding side, that is, the encoder detects the distribution status of objects in the parent space PA11, and the child is a small space surrounding a group of locally concentrated objects according to the detection result. Make it a space. In this example, the child space CA11 is a scaled small space containing a plurality of objects as shown by arrow Q12, for example.
[0024]
Here, the child space CA11 included in the parent space PA11 is a space represented by absolute coordinates, and in the present technology, the position in the child space CA11 is represented by absolute coordinates. In other words, the child space CA11 is the coordinate space of the absolute coordinate system. The child space may be a coordinate space of another coordinate system different from the absolute coordinate system such as a polar coordinate system, but the case where the child space is the coordinate space of the absolute coordinate system will be described below as an example. continue.
[0025]
In this technology, the position of the object in the parent space PA11 is not expressed by the coordinates of the absolute coordinate system of the parent space PA11, but the child space information regarding the position and size of the child space CA11 in the parent space PA11 and the child space It is expressed by the coordinates of the position of the object in the absolute coordinate system of CA11.
[0026]
Here, the coordinate system of the parent space PA11 and the coordinate system of the child space CA11 will be described with reference to FIG. In FIG. 2, the same reference numerals are given to the portions corresponding to those in FIG. 1, and the description thereof will be omitted as appropriate.
[0027]
For example, as shown by arrow Q21 in FIG. 2, the parent space PA11 is an absolute coordinate system consisting of a total of three axes, the x-axis, the y-axis, and the z-axis, with the center of the cube handled by MPEG-H as the origin O. Is the coordinate space of.
[0028]
In the parent space PA11, in order to express the position in each axial direction, a normalized value from 0 to 1 for expressing the position between the origin O and the end and positive and negative sign information are required. That is, the x-coordinate, y-coordinate, and z-coordinate indicating the positions in the x-axis direction, the y-axis direction, and the z-axis direction are values from -1 to 1, respectively.
[0029]
In the following, the absolute coordinate system of the parent space PA11 with respect to the origin O will also be referred to as the xyz coordinate system, and the coordinates indicating the position in the xyz coordinate system will be described as (x, y, z).
[0030]
Further, in the parent space PA11, for example, the actual length (distance) from the position where the x-coordinate is 0 to the position where the x-coordinate is 1 is the absolute distance absolute distance of the parent space PA11 described later. For example, absolute distance is described in meters.
[0031]
On the other hand, as shown by arrow Q22, in the coordinate representation of the absolute coordinate system of the child space CA11, positive and negative codes are not used from the viewpoint of facilitating the coding of position information.
[0032]
Here, the child space CA11 is in the shape of a cube or a rectangular parallelepiped. In the absolute coordinate system of the child space CA11, the apex of the child space CA11 corresponding to the position (-1, -1, -1) of the parent space PA11 is set as the origin O', and each side is normalized from 0 to 1. It is supposed to use the value.
[0033]
That is, the child space CA11 is a coordinate space of an absolute coordinate system consisting of a total of three axes, the X-axis, the Y-axis, and the Z-axis, with the lower left vertex as the origin O'in the figure of the child space CA11.
[0034]
In the following, the absolute coordinate system of the child space CA11 with respect to the origin O'is also referred to as the XYZ coordinate system, and the coordinates indicating the position in the XYZ coordinate system are described as (X, Y, Z). In particular, here, the coordinate values of X, Y, and Z are values between 0 and 1.
[0035]
Here, the case where the coordinate value in each axial direction of the XYZ coordinate system is 0 or more will be described, but the coordinate value in each axial direction may also take a negative value. Further, the shape of the child space CA11 is not limited to a cube or a rectangular parallelepiped, and may be any shape.
[0036]
In the example shown in FIG. 2, the coordinates of the xyz coordinate system indicating the position of the origin O'of the child space CA11 in the parent space PA11 are (x0, y0, z0), and the pair of the origin O'of the child space CA11. The coordinates of the xyz coordinate system indicating the position of the corner apex are (x1, y1, z1).
[0037]
In the following, the position of the object in the xyz coordinate system, that is, the information indicating the position of the object in the parent space is also referred to as the position information in the parent space. In addition, hereinafter, the position of the object in the XYZ coordinate system, that is, the information indicating the position of the object in the child space is also referred to as the position information in the child space.
[0038]
By the way, for example, when encoding the position information indicating the position of an object, the information indicating the position of the input object is the position information in the parent space indicating the position in the parent space. Therefore, in the present technology, coordinate transformation (scaling transformation) from the parent space to the child space, that is, the position information in the parent space to the position information in the child space is performed.
[0039]
Now, let i be the index indicating the i-th child space in the parent space, and let the coordinates (x, y, z) that are the position information in the parent space be the coordinates that are the position information in the child space of the i-th child space. Assuming that the coordinates are converted to (X, Y, Z), the conversion formula is as shown in the following formula (1).
[0040]
[Number 1]
[0041]
However, in equation (1), 0 ≦ X ≦ 1, 0 ≦ Y ≦ 1, 0 ≦ Z ≦ 1. Further, a_x [i], b_x [i], a_y [i], b_y [i], a_z [i], and b_z [i] in the equation (1) are scaling conversion coefficients of the i-th child space. It is shown in the following equation (2).
[0042]
[Number 2]
[0043]
In equation (2), x0 [i], y0 [i], and z0 [i] are the x-coordinate, y-coordinate, and z-coordinate of the xyz coordinate system that indicates the position of the origin O'in the i-th child space. It is shown and corresponds to the above (x0, y0, z0). Also, x1 [i], y1 [i], and z1 [i] are the x, y, and z coordinates of the xyz coordinate system that indicates the position of the diagonal vertices with respect to the origin O'in the i-th child space. It is shown and corresponds to the above (x1, y1, z1).
[0044]
Therefore, here, as a coordinate transformation, a function transformation using a linear function is performed for each axis of the coordinate system. Then, as the parameters of the linear function, the slopes a_x [i], a_y [i], and a_z [i] of the linear function for each axis, and the intercepts b_x [i], b_y [i] of the linear function, And b_z [i] are used.
[0045]
For example, in the present technology, these six parameters in total can be used as child space information regarding the position and size of the child space in the parent space. The parameters of these linear functions are obtained from the coordinates of the origin O'in the child space (x0, y0, z0) and the coordinates of the diagonal vertices (x1, y1, z1), so that they are substantially parents. It can be said that it is information indicating the position and size of the child space in the space.
[0046]
If you want to directly acquire the information indicating the position of the child space in the parent space from the bit stream, use the child space in the parent space instead of the parameter of the linear function shown in Eq. (2) as the child space information. The coordinates (x0, y0, z0) of the origin O'and the coordinates (x1, y1, z1) of the diagonal vertices of the child space with respect to the origin O'may be transmitted as they are as child space information.
[0047]
However, compared to the case where the coordinates (x0, y0, z0) of the origin O', which is the reference vertex of the child space, and the coordinates of the diagonal vertex (x1, y1, z1) are transmitted as child space information, 1 If the parameters of the next function are transmitted as child space information, the amount of calculation on the decoding side can be reduced.
[0048]
Further, when the object group exists at a plurality of places in the parent space, a plurality of child spaces may be formed for one parent space.
[0049]
In this way, normalizing the position information of an object according to the child space limited to the range in which the object to be encoded exists is an application of a scale factor that matches the size of the spectrum of the existing audio coding. It will be a similar relationship.
[0050]
The position, size, and shape of the child space are determined by the encoder, but it is desirable that the child space contains all the objects of interest and has the smallest volume of space.
[0051]
When the coordinates (X, Y, Z) which are the position information in the child space are obtained by the above coordinate transformation, the position information in the child space is quantized with a predetermined quantization accuracy to obtain the quantized position information.
[0052]
In the quantization process, for example, as shown in FIG. 3, one side of the child space, that is, one axial side is normalized from 0 to 1, and the median value of the divided region of the power of 2 with respect to one side is after quantization. Is the value of. In FIG. 3, for the sake of simplicity, the child space is assumed to be two-dimensional in the X-axis direction and the Y-axis direction.
[0053]
For example, the range in each axis direction of the child space, that is, the range of the coordinates of each axis is from 0 to 1, but this range is divided into 2 n (however, 0 ≤ n) as shown on the right side in the figure. The median value of each range after division is taken as a representative value that can be taken as the value after quantization.
[0054]
In the following, the value of n (exponent part) when dividing the range in each axial direction into 2 n pieces will be referred to as the divided index. This division index indicates how many divisions the range in each axial direction should be by a power of 2.
[0055]
For example, in the example shown in FIG. 3, when the value of the division index is 0, the range in each axial direction is divided into 0 pieces, that is, it is not substantially divided, so the coordinates are the position information in the child space. (X, Y, Z) is (0.5,0.5,0.5).
[0056]
Further, for example, when the value of the division index is 1, the range in each axis direction is divided into two, so that the coordinates of each axis of the position information in the child space are set to either 0.25 or 0.75. Will be done. In FIG. 3, the position of the wedge mark represents the representative value after quantization.
[0057]
When quantization is performed in this way, the larger the power value of 2, that is, the number of divisions determined by the division index, the higher the quantization accuracy.
[0058]
Here, FIG. 4 shows the number of divisions in the range in each axial direction corresponding to the division index when the division index is 0 to 10. In FIG. 4, for example, when the value of the division index is 0, the number of divisions is 0, and when the value of the division index is 10, the number of divisions is 512.
[0059]
Next, the division index for quantizing the position information in the child space with the highest accuracy, that is, the division index for quantization with the smallest quantization step width is shown in the following equation (3). However, here, it is assumed that the rule of not quantizing is applied below the minimum distance MinDist from the listener to the object. The minimum distance MinDist is the minimum possible distance from the listener to the object, and the minimum distance MinDist is such that the listener and the object do not overlap, for example, considering the size of the listener's head. It is decided.
[0060]
[Number 3]
[0061]
In equation (3), max_exp_index_X, max_exp_index_Y, and max_exp_index_Z are the X-axis (X-direction), Y-axis (Y-direction), and Z-axis (Z-direction) divided indexes for quantization with the highest accuracy, respectively. , Also called the highest precision split index). Further, in the equation (3), INT () indicates an INT function that outputs the integer part of the argument, and "^" indicates a power.
[0062]
Further, in the equation (3), absoluteDistanceX, absoluteDistanceY, and absoluteDistanceZ indicate the actual distances (absolute distances) of the sides in each of the X-axis direction, the Y-axis direction, and the Z-axis direction of the child space. The absolute distance of is, for example, in meters (m). Similarly, the minimum distance MinDist from the listener to the object is, for example, in meters (m), and this minimum distance MinDist is a predetermined value such as the size of the head of a general user. NS.
[0063]
When the highest precision division index is obtained by the equation (3), the quantization step width for each axis of the XYZ coordinate system in the child space is obtained by the following equation (4).
[0064]
[Number 4]
[0065]
In equation (4), step_encX, step_encY, and step_encZ indicate the quantization step width with respect to the X-axis, Y-axis, and Z-axis of the XYZ coordinate system, and "^" indicates the power.
[0066]
Furthermore, by calculating the following equation (5) using the quantization step width obtained in this way, the position information (X, Y, Z) of the object in the child space in the child space is quantized with the highest accuracy. The X-coordinate, Y-coordinate, and Z-coordinate Qpx_high [i], Qpy_high [i], and Qpz_high [i] of the quantized position information obtained by conversion can be obtained. In equation (5), i is an index indicating a child space.
[0067]
[Number 5]
[0068]
In the following, among the quantized position information obtained by quantizing the position information in the child space, the quantized position information obtained by quantizing based on the highest precision division index is the highest precision quantized position. It will also be referred to as information.
[0069]
Further, although the method of determining the number of divisions of the region in each axial direction at the highest quantization accuracy is based on the minimum distance MinDist from the listener to the object, other criteria may be used.
[0070]
Furthermore, the highest precision division index may not require a large number of divisions, such as when the absolute distance of the child space is small, so it is unnecessary if the upper limit division index is set according to the absolute distance of the space and the number of divisions is limited. The number of bits can be reduced.
[0071]
In addition, the quantized value of the position information in the child space, that is, the value of the quantized position information is obtained by dividing by the number of divisions of the power of 2. Therefore, the quantization position information of arbitrary quantization accuracy can be obtained by extracting the required number of bits from the MSB (Most Significant Bit) side of the highest precision quantization position information, that is, the most significant bit side.
[0072]
That is, for example, FIG. 5 shows the quantization position information when the coordinate value “0.2” of a predetermined axis of the position information in the child space is quantized with each quantization accuracy, that is, with each division number.
[0073]
In FIG. 5, the characters "division index", "number of divisions on one side of one quadrant", and "quantization step width" are the division index described above, the number of divisions in one axial range, and quantization at the time of quantization. Indicates the step width.
[0074]
The "quantized value", "bit notation", and "number of transmission bits" are bits of one axial coordinate value of the quantization position information (integer value) and one axial coordinate value of the quantization position information. The notation and the number of bits of the coordinate value in one axial direction of the quantization position information are shown. Further, the "decoded value" and the "error with the true value" are one axial coordinate value of the child space position information obtained by decoding the quantization position information and one of the decoded child space position information. It shows the error between the coordinate values in one axial direction and the coordinate values in one axial direction of the position information in the child space before quantization.
[0075]
Here, for example, assuming that the highest precision division index is "11", the bit notation of the quantization value when quantization is performed with the highest precision division index is "00110011001".
[0076]
Further, for example, the bit notation of the quantization value when quantization is performed with a division index "10" having a smaller number of divisions than the highest precision division index, that is, a division index "10" having a lower quantization accuracy is "0011001100".
[0077]
Therefore, if the number of bits (10 bits) indicated by the division index "10" is extracted from the most significant bit side of the quantization value "00110011001" of the highest precision quantization position information, the division index "10" is obtained. It can be seen that the corresponding quantization value "0011001100" is obtained.
[0078]
That is, by extracting the bits of the highest precision quantization position information by the number of bits corresponding to the split index from the most significant bit side of the highest precision quantization position information, an arbitrary split index, that is, a quantum of an arbitrary quantization precision is extracted. Quantization position information can be obtained.
[0079]
As described above, the process of determining the child space, converting the coordinate of the position information in the parent space to the position information in the child space, and quantizing the position information in the child space is performed as the coding process of the position information in the parent space. ..
[0080]
Next, decoding of the quantization position information will be described. The inverse quantization of the quantization position information and the inverse conversion of the position information in the child space obtained by the inverse quantization into the position information in the parent space are performed as the decoding process.
[0081]
For example, let the X, Y, and Z coordinates of the quantized position information be Qpx_dec, Qpy_dec, and Qpz_dec, and let the split indexes of the X, Y, and Z coordinates be exp_index_fin_X, exp_index_fin_Y, and exp_index_fin_Z. In particular, the index indicating the child space is defined as i, and the split indexes exp_index_fin_X, exp_index_fin_Y, and exp_index_fin_Z for the child space are described as the split indexes exp_index_fin_X [i], exp_index_fin_Y [i], and exp_index_fin_Z [i].
[0082]
Further, the X-coordinate, Y-coordinate, and Z-coordinate of the position information in the child space (hereinafter, also referred to as the position information in the decoder space) obtained by dequantizing the quantization position information are referred to as Xdec, Ydec, and Zdec. ..
[0083]
In this case, in the inverse quantization, the calculation of the following equation (6) is performed to obtain the position information in the decoder space. In equation (6), i is an index indicating the child space to which the object belongs, and "^" indicates a power.
[0084]
[Number 6]
[0085]
Subsequently, when the inverse scaling conversion shown in the following equation (7) is performed on the coordinates (Xdec, Ydec, Zdec) of the position information in the decoder space obtained in this way, the original position in the parent space is performed. Information coordinates (xdec, ydec, zdec) can be obtained. The calculation of the equation (7) is the inverse transformation of the equation (1) described above. In equation (7), i is an index indicating the child space to which the object belongs.
[0086]
[Number 7]
[0087]
Further, in equation (7), a'_x [i], b'_x [i], a'_y [i], b'_y [i], a'_z [i], and b'_z [i] are Inverse scaling conversion factor from the i-th child space to the parent space, a_x [i], b_x [i], a_y [i], b_y [i], a_z [i], and b_z [ It is expressed by the following equation (8) using i].
[0088]
[Number 8]
[0089]
By the above decoding process, coordinates (xdec, ydec, zdec) can be obtained as position information in the parent space indicating the position of the object in the parent space. Hereinafter, the position information in the parent space obtained by the decoding process will be appropriately referred to as the position information in the decoding parent space.
[0090]
By converting the position information in the parent space into the position information in the child space and then performing the quantization as in this technology, the code of the quantization position information is compared with the case where the position information in the parent space is quantized as it is. The amount can be reduced. In other words, the amount of line code transmitted for the position information of the object can be reduced.
[0091]
Specifically, for example, as shown in FIG. 6, a plurality of objects are concentrated in a part of the area in the parent space PA11, and the child space CA11 is formed to encode the position information of the objects. think of. In FIG. 6, the same reference numerals are given to the portions corresponding to those in FIG. 1, and the description thereof will be omitted as appropriate.
[0092]
In the example shown in FIG. 6, each circle represents one object, and the length of one side of the parent space PA11 which is a cube is 128 m. A cubic child space CA11 included in the parent space PA11 is formed with respect to this parent space PA11, and the length of one side of the child space CA11 is 10 m.
[0093]
In such a case, for example, the number of coding bits per axis of the coordinate system required to make the position accuracy 0.2 m or less with respect to the true value of the position information in the parent space is 9 bits in the parent space PA11. It becomes 5 bits in the child space CA11. That is, if one side of the coordinate space is 128 m, 9 bits are required, and if one side of the coordinate space is 10 m, 5 bits are required.
[0094]
Therefore, since the coordinates of each coordinate space have three axes, for example, when the position information in the parent space is quantized as it is, 27 bits are required as the number of bits of the quantized position information, whereas in the child space. When the position information is quantized, the number of bits of the quantized position information needs to be 15 bits to obtain the information with the same accuracy. In this way, one object can reduce 12 bits.
[0095]
However, when quantizing the position information in the child space, the child space information regarding the position and size of the child space CA11 is required. For example, the child space information has 32 bits for one linear function parameter. If so, the total is 32 x 6 = 192 bits. Further, when the position and size of the child space do not change in the time direction, it is not necessary to transmit the child space information for each frame.
[0096]
In this way, as the number of objects contained in the child space CA11 increases, the number of bits that can be reduced increases. Therefore, when the number of objects is large, the position information in the child space is not quantized as it is in the parent space. It is possible to transmit the position information of the object with a smaller amount of code (number of bits) by quantizing.
[0097]
Subsequently, a more specific embodiment of the present technology described above will be described.
[0098]
FIG. 7 is a diagram showing a configuration example of an embodiment of an encoder to which the present technology is applied.
[0099]
The encoder 11 shown in FIG. 7 has an acquisition unit 21, a coding unit 22, and a communication unit 23.
[0100]
The acquisition unit 21 acquires the number of objects indicating the number of objects to be encoded from an external device or the like, the absolute distance absolute distance in the parent space, and the position information in the parent space of each object and supplies the information to the coding unit 22. ..
[0101]
The coding unit 22 encodes the position information in the parent space of each object based on the number of objects supplied from the acquisition unit 21, the absolute distance absolute distance in the parent space, and the position information in the parent space of the objects, and obtains the result. The encoded data is supplied to the communication unit 23.
[0102]
The coding unit 22 has a child space determination unit 31, a conversion unit 32, and a quantization unit 33.
[0103]
The child space determination unit 31 determines the child space based on the position information in the parent space of each object. The conversion unit 32 converts the position information in the parent space of each object into the position information in the child space based on the determination result of the child space by the child space determination unit 31. The quantization unit 33 quantizes the position information in the child space.
[0104]
The communication unit 23 transmits the coded data supplied from the coding unit 22.
[0105]
Next, the operation of the encoder 11 will be described. That is, the coding process by the encoder 11 will be described below with reference to the flowchart of FIG.
[0106]
In step S11, the acquisition unit 21 acquires the number of objects, the absolute distance in the parent space, and the position information in the parent space of each object and supplies the information to the coding unit 22.
[0107]
In step S12, the child space determination unit 31 determines the child space based on the position information in the parent space of each object supplied from the acquisition unit 21.
[0108]
For example, in step S12, the child space is a cube region (space) that includes all the objects existing in the parent space and has the smallest volume. In addition, for example, when the distance between a plurality of object groups is more than a predetermined distance, the child space may be determined for each of the plurality of object groups.
[0109]
In step S13, the conversion unit 32 converts the position information in the parent space of each object supplied from the acquisition unit 21 into the position information in the child space.
[0110]
For example, in step S13, the calculation of the above equation (1) is performed, and the position information in the child space is calculated by the coordinate transformation. The parameters of the linear function used in the calculation of equation (1) are the determination result of the child space, that is, the coordinates of the origin O'of the child space in the parent space and the diagonal vertices of the child space with respect to the origin O'. It can be obtained from the coordinates of Eq. (2).
[0111]
In step S14, the quantization unit 33 determines the division index.
[0112]
For example, in step S14, the above-mentioned equation (3) is calculated based on the predetermined minimum distance MinDist, the determination result of the child space, and the absolute distances absoluteDistanceX, absoluteDistanceY, and absoluteDistanceZ obtained from the absolute distance absoluteDistance. A precision split index is required.
[0113]
Then, based on the highest precision split index, the split index corresponding to the required quantization precision is determined. Here, the highest-precision partitioning index may be used as it is for quantization, or a partitioning index having a value smaller than the highest-precision partitioning index may be used for quantization.
[0114]
Note that the split index can be determined for each object in each axis of the XYZ coordinate system in the child space.
[0115]
In step S15, the quantization unit 33 obtains the quantization step width by performing the same calculation as the above-mentioned equation (4) based on the division index determined in step S14.
[0116]
In step S16, the quantization unit 33 quantizes the position information in the child space of each object obtained in step S13.
[0117]
For example, in step S16, the quantization step width obtained in step S15 is used to perform the same calculation as in the above equation (5), and the coordinates of each axis of the position information in the parent space are quantized. That is, quantization is performed for each coordinate component of each axis. As a result, the quantization position information of each object can be obtained.
[0118]
The above steps S12 to S16 are performed as a process of encoding the position information in the parent space. It was explained here that the position information in the parent space of all objects is converted into the position information in the child space and quantized, but only some of the position information in the parent space is converted into the position information in the child space. It may be quantized and the remaining part of the position information in the parent space may be quantized as it is. In such a case, for example, for an object whose distance to all other objects is equal to or greater than a predetermined distance and no other object exists nearby, the position information in the parent space is quantized as it is. May be good.
[0119]
In step S17, the coding unit 22 includes the quantization position information obtained in step S16, the number of objects and the absolute distance absoluteDistance acquired in step S11, the child space information obtained from the determination result of the child space in step S12, and each object. Encoded data including information indicating the child space to which the object belongs (hereinafter, also referred to as object affiliation information) and the division index determined in step S14 is generated.
[0120]
The coding unit 22 supplies the obtained coded data to the communication unit 23.
[0121]
In step S18, the communication unit 23 transmits the coded data supplied from the coding unit 22, and the coding process ends.
[0122]
As described above, the encoder 11 converts the position information in the parent space into the position information in the child space by coordinate conversion, and quantizes the position information in the child space to generate coded data. By doing so, it is possible to reduce the amount of code of the quantization position information as compared with the case where the position information in the parent space is encoded (quantized) as it is.
[0123]
Further
, the coded data output from the encoder 11 is supplied to the decoder directly or via another device or recording medium and decoded.
[0124]
Here, a decoder that decodes the coded data output by the encoder 11 will be described. FIG. 9 shows a configuration example of an embodiment of a decoder to which the present technology is applied.
[0125]
The decoder 61 shown in FIG. 9 has an acquisition unit 71, a decoding unit 72, and an output unit 73.
[0126]
The acquisition unit 71 acquires the coded data by receiving the coded data transmitted by the encoder 11, acquiring the coded data from another device, or reading the coded data from the recording medium. It is supplied to the decoding unit 72.
[0127]
The decoding unit 72 unpacks the coded data supplied from the acquisition unit 71, decodes the quantized position information extracted from the coded data, and obtains the position information in the parent space of each object. The decoding unit 72 has an inverse quantization unit 81 and a conversion unit 82.
[0128]
The dequantization unit 81 dequantizes the quantization position information of each object and calculates the position information in the child space. The conversion unit 82 converts the position information in the child space into the position information in the parent space by using the child space information extracted from the coded data. Further, the conversion unit 82 calculates the position of the unnormalized absolute object in the parent space by multiplying the position information in the parent space by the absolute distance absoluteDistance extracted from the encoded data, and as a result, The obtained position information is supplied to the output unit 73.
[0129]
The decoding unit 72 having such an inverse quantization unit 81 and a conversion unit 82 determines the position of the object in the parent space based on the spatial information and the quantized position information which is the quantized position information in the child space. It functions as a calculation unit that calculates the position information in the parent space indicating.
[0130]
The output unit 73 outputs the position information indicating the position of the object in the parent space supplied from the decoding unit 72 to the subsequent stage.
[0131]
Next, the operation of the decoder 61 will be described. That is, the decoding process by the decoder 61 will be described below with reference to the flowchart of FIG.
[0132]
In step S41, the acquisition unit 71 acquires the coded data and supplies it to the decoding unit 72. The decoding unit 72 unpacks the coded data supplied from the acquisition unit 71, and extracts the quantization position information, the number of objects, the absolute distance absoluteDistance, the child space information, the object affiliation information, and the division index of each object.
[0133]
In step S42, the inverse quantization unit 81 dequantizes the quantization position information by calculating the above-mentioned equation (6) based on the quantization position information and the division index for the quantization position information of each object. do. As a result, the position information in the child space (position information in the decoder space) can be obtained.
[0134]
In step S43, the conversion unit 82 performs inverse scaling conversion on each object based on the child space position information obtained in step S42 and the child space information, and converts the child space position information into the parent space position information. Convert.
[0135]
For example, in step S43, the inverse scaling conversion coefficient is obtained from the parameter of the linear function as the child space information by the above equation (8), and further, the equation (7) is based on the inverse scaling conversion coefficient and the position information in the decoder space. ) Is calculated, and the position information in the decoding parent space is calculated.
[0136]
When the position information in the decoding parent space of each object is obtained in this way, the conversion unit 82 indicates the position of the object in the parent space by multiplying the position information in the decoding parent space by the absolute distance absoluteDistance. The position information is obtained and supplied to the output unit 73.
[0137]
At this time, the conversion unit 82 also supplies the number of objects and the object affiliation information to the output unit 73 as needed, and the output unit 73 supplies the position information, the number of objects, and the object affiliation information of the objects supplied from the conversion unit 82 to the latter stage. Output to.
[0138]
The process of step S42 and step S43 described above is a process of decoding the quantized position information to obtain the position information in the parent space. When the position information of the object is output in this way, the decoding process ends.
[0139]
As described above, the decoder 61 dequantizes the quantized position information and converts the resulting position information in the child space into the position information in the parent space. By doing so, it is possible to reduce the amount of code of the information indicating the position of the object as compared with the case where the position information in the parent space is encoded (quantized) as it is.
[0140]
By the way, the present technology can also be used for arranging content from a free viewpoint including at least audio data of an audio object. In this case, for example, the server transmits the audio data as the content and the position information of the audio object (object) as the auxiliary information of the content to the client.
[0141]
When the server transmits the quantization position information indicating the position of the object to the client as the auxiliary information of the content, the server holds the above-mentioned most accurate quantization position information in advance as the quantization position information. You can try to keep it.
[0142]
Then, the server can generate the quantization position information of arbitrary quantization accuracy from the highest precision quantization position information according to the conditions of the client side, that is, the listener side, and supply it to the client. ..
[0143]
For example, when there is no limitation on the communication speed, the server can send the highest-precision quantization position information having the highest quantization accuracy to the client as it is. On the other hand, when the communication speed is limited, for example, the quantization position information transmission bit is transmitted to the client with a quantization accuracy lower than the maximum accuracy depending on some conditions. You can lower the rate.
[0144]
The conditions referred to here can be based on, for example, the actual distance from the listener to the object in the parent space, the band condition of the transmission line, and the like.
[0145]
Specifically, for example, in human hearing, it is known that the horizontal direction is recognized as the same angle as the front angle if the angle is within a certain range on the left and right with respect to the front object, and this angle is the perceptual limit. It is called an angle.
[0146]
Therefore, if the condition that the quantization step width is smaller than the horizontal distance corresponding to the perceived limit angle at a certain position of the object as seen from the listener is satisfied, the quantization position information of what kind of quantization accuracy is satisfied. However, when the sound is localized at the position of the object obtained from the quantized position information, the listener does not feel the deviation of the sound image position. In other words, if the quantized position information has a quantum step width equal to or less than the permissible margin of error due to the perceptual limit angle, the listener will not feel the deviation of the sound image position.
[0147]
Therefore, for example, based on the actual distance from the listener to the object and the perceptual limit angle, the quantization position information quantized with the maximum quantization step width that does not exceed the horizontal distance corresponding to the perceptual limit angle. If is selected, the code amount of the quantized position information can be reduced and the transmission bit rate can be reduced without causing the listener to feel the deviation of the sound image position.
[0148]
Further, for example, the quantization position information of the quantization accuracy according to the band state of the transmission line for transmitting the quantization position information may be acquired. In this case, for example, when the transmission band of the transmission line has a margin, the highest accuracy quantization position information is acquired, and when the transmission band of the transmission line is narrow, the lower quantization accuracy determined for the transmission bandwidth is obtained. The quantization position information may be acquired.
[0149]
In addition, it is necessary to transmit information necessary for decoding by the client such as the number of objects and the absolute distance absolute Distance from the server to the client for each frame of audio data.
[0150]
Therefore, by transmitting the header information shown in FIG. 11, for example, from the server to the client, necessary information such as the number of objects and the absolute distance absolute Distance may be transmitted.
[0151]
In the example shown in FIG. 11, the header information is indicated by the absolute distance absoluteDistance indicating the size of the parent space indicated by the character "absoluteDistance", the child space information indicated by the character "ChildAreaInfo ()", and the character "nobjs". Contains the number of objects in the parent space.
[0152]
Further, in the header information, as many indexes as the number of objects, that is, indexes indicating the child spaces to which the objects belong are stored as object affiliation information for each object. Here, the character "ChildAreaIndexForObj [i]" indicates the object affiliation information of the i-th object.
[0153]
Further, the format of the child space information "ChildAreaInfo ()" part included in the header information is as shown in FIG. 12, for example.
[0154]
In the example shown in FIG. 12, the child space information includes the number of child spaces indicating the number of child spaces contained in the parent space, which is indicated by the character "numOfChildArea".
[0155]
Further, in the child space information, as many as the number of child spaces, the highest precision division index in each axial direction and the parameters of the linear function are stored.
[0156]
That is, the letters "max_exp_index_X [i]", "max_exp_index_Y [i]", and "max_exp_index_Z [i]" are in the X-axis, Y-axis, and Z-axis directions for the i-th child space, respectively. Shows the highest precision split index of.
[0157]
Also, the letters "a_x [i]", "a_y [i]", and "a_z [i]" are in the X-axis direction, Y-axis direction, and Z-axis direction for the i-th child space, respectively. It is a parameter indicating the slope of the linear function of. In addition, the letters "b_x [i]", "b_y [i]", and "b_z [i]" are in the X-axis, Y-axis, and Z-axis directions for the i-th child space, respectively. It is a parameter indicating the intercept of the linear function. The slopes and intercepts of these linear functions are the scaling conversion coefficients shown in the above equation (2).
[0158]
The header information may be transmitted to the client for each frame, or may be transmitted when the number, position, or size of the child space changes. In addition, header information may be transmitted at regular frame intervals.
[0159]
Also, when an object moves out of the child space to which it belongs, the object affiliation information may be updated so that the object belongs to another existing child space or a new child space. In this case, the header information including the updated object affiliation information is transmitted from the server to the client.
[0160]
Further, if the server holds the highest-precision quantization position information as described above, the client can acquire the quantization position information of arbitrary quantization accuracy from the server.
[0161]
For example, even after the client has acquired the quantization position information having a predetermined quantization accuracy, it is possible to obtain the quantization position information having a higher quantization accuracy than the already acquired quantization position information.
[0162]
Specifically, first, the quantization position information that the client first acquires from the server is referred to as the initial quantization position information.
[0163]
The initial quantization position information is the quantization position information having a quantization accuracy of the highest accuracy or less corresponding to the highest precision division index included in the header information. Therefore, it is also possible to obtain the highest precision quantization position information as the initial quantization position information.
[0164]
The client acquires the initial quantization position information from the server by, for example, transmitting the initial quantization position information transmission request in the bitstream format shown in FIG. 13 to the server.
[0165]
In the example shown in FIG. 13, in the initial quantization position information transmission request, a division index corresponding to the quantization accuracy for which transmission is requested is stored for each object in each axial direction.
[0166]
That is, the letters "exp_index_initial_x [i]", "exp_index_initial_y [i]", and "exp_index_initial_z [i]" are in the X-axis, Y-axis, and Z-axis directions for the i-th object, respectively. Shows the split index.
[0167]
The server that receives such an initial quantization position information transmission request first quantizes the quantization position information of the quantization accuracy corresponding to the division index included in the initial quantization position information transmission request for each axial direction for each object. Send to the client as quantized location information.
[0168]
For example, in the example shown in FIG. 5, when the division index included in the initial quantization position information transmission request is "3", the quantization position information "001" corresponding to the division index "3" is the initial quantum. It is sent to the client as quantized location information.
[0169]
Specifically, the initial quantization position information is transmitted from the server to the client in the bitstream format shown in FIG. 14, for example.
[0170]
In the example shown in FIG. 14, the bit stream stores the initial quantization position information for each object in each axial direction.
[0171]
Here, the letters "Xpos_initial [i]", "Ypos_initial [i]", and "Zpos_initial [i]" are the X-axis, Y-axis, and Z-axis directions for the i-th object, respectively. The initial quantization position information of the above, and more specifically, the coordinate values of the components in each axial direction constituting the initial quantization position information are shown.
[0172]
After acquiring the initial quantization position information in this way, the client wants to obtain the quantization position information with higher quantization accuracy than the initial quantization position information. Acquires additional bit information which is difference information from the quantized position information.
[0173]
In the client, by adding the additional bit information to the lower bit side of the initial quantization position information, the quantization position information having higher quantization accuracy than the initial quantization position information can be obtained.
[0174]
Specifically, for example, in the example shown in FIG. 5, the quantization position information "001" corresponding to the division index "3" was acquired as the initial quantization position information, but finally corresponds to the division index "5". Suppose you want to obtain the quantization position information "00110".
[0175]
In this case, the client acquires "10", which is the difference information between the initial quantization position information "001" and the finally desired quantization position information "00110", as additional bit information. Then, the client adds the newly acquired additional bit information "10" to the lower bit side of the already acquired initial quantization position information "001" to finally obtain the quantization position information of the quantization accuracy. You can get "00110".
[0176]
When the client acquires the additional bit information from the server in this way, the client selects the request mode for acquiring the additional bit information according to the transmission band status and the like. Then, the client transmits an additional bit information transmission request to the server in a format corresponding to the selected request mode to request the transmission of the additional bit information.
[0177]
Here, for example, the additional bit information transmission request shown in FIG. 15 is transmitted to the server.
[0178]
In the example shown in FIG. 15, the additional bit information transmission request includes information indicating the request mode indicated by the character “request_mode”. For example, there are three request modes: a common mode, a child space common mode, and an individual mode.
[0179]
The common mode is a mode in which additional bit information is requested for each axial direction of all objects so that the number of bits of additional bit information additionally acquired for each axial direction is the same for all objects. For example, in the common mode, the number of bits to be added for each axial direction, that is, the quantization accuracy to be additionally increased can be selected.
[0180]
In addition, the child space common mode is a child space for one or more desired child spaces so that the number of bits of the additional bit information additionally acquired for each axial direction is the same for all objects belonging to the same child space. This mode requires additional bit information for each axis direction of all objects belonging to. For example, even in the child space common mode, the number of bits to be added can be selected for each axial direction. In the child space common mode, there are a child space that requires additional bit information and a child space that does not require additional bit information.
[0181]
Further, the individual mode is a mode in which additional bit information is individually requested for each object. In the individual mode, it is possible to select whether or not to request the additional bit information for each object, and for each object, the additional bit information so that the quantization position information of the desired quantization accuracy can be obtained for each axial direction. Can be requested.
[0182]
In FIG. 15, for example, when the request mode indicated by the character "request_mode" indicates a common mode, the characters "addbit_x" and "addbit_y" are included in the additional bit information transmission request as shown in the part of the character "case 0". , And "addbit_z", information indicating the number of additional bits in each of the X-axis direction, the Y-axis direction, and the Z-axis direction is stored.
[0183]
For example, if the information indicating the number of additional bits is information indicating 3 bits, the additional bit information for 3 bits added to the initial quantization position information is transmitted as the additional bit information.
[0184]
In addition, when the request mode indicated by the character "request_mode" indicates the child space common mode, the additional bit information transmission request is indicated by the character "num_of_addprec_childArea" as shown in the part of the character "case 1". Contains information that indicates the number of child spaces that require additional bit information.
[0185]
In this case, in the additional bit information transmission request, the index indicating the child space requesting the additional bit information indicated by the character "addprec_childArea_index [i]" and the "addbit_x" indicating the number of additional bits in each axial direction in the child space are indicated. "[I]", "addbit_y [i]", and "addbit_z [i]" are stored as many as the number of child spaces requesting additional bit information. In particular, here the letters "addbit_x [i]", "addbit_y [i]", and "addbit_z [i]" are the number of additional bits in each of the X-axis, Y-axis, and Z-axis directions. It is information indicating.
[0186]
Furthermore, if the request mode indicated by the character "request_mode" indicates an individual mode, the additional bit information transmission request is indicated by the character "num_of_addprec_objects", as indicated by the character "case 2". Contains information that indicates the number of objects that request information.
[0187]
In this case, in the additional bit information transmission request, the index indicating the object requesting the additional bit information indicated by the character "addprec_ object_index [i]" and the "add bit_x [" indicating the number of additional bits in each axis direction of the object are indicated. "i]", "addbit_y [i]", and "addbit_z [i]" are stored as many as the number of objects requesting additional bit information. In particular, here the letters "addbit_x [i]", "addbit_y [i]", and "addbit_z [i]" are the number of additional bits in each of the X-axis, Y-axis, and Z-axis directions. It is information indicating.
[0188]
Upon receiving such an additional bit information transmission request shown in FIG. 15, the server transmits the additional bit information to the client in the bitstream format shown in FIG. 16, for example.
[0189]
In the example shown in FIG. 16, when the request mode is the common mode, the bitstream has the characters "Xpos_add [i]", "Ypos_add [i]", and " Contains additional bit information for each of the X-axis, Y-axis, and Z-axis directions for the i-th object indicated by "Zpos_add [i]".
[0190]
Also, when the request mode is the child space common mode, as shown in the character "case 1" part, additional bit information is added to the bit stream for each child space for which bit information is requested, and each object belonging to that child space is added. Bit information is stored. Specifically, the X-axis, Y-axis, and Z for the j-th object in the child space indicated by the letters "Xpos_add [j]", "Ypos_add [j]", and "Zpos_add [j]". Additional bit information for each axial direction is stored.
[0191]
Furthermore, when the request mode is the individual mode, the bitstream stores the additional bit information of each object for which additional bit information is requested, as indicated by the character "case 2". Specifically, the X-axis and Y-axis directions of an object, as indicated by the letters "Xpos_add [addprec_object_index [i]]", "Ypos_add [addprec_object_index [i]]", and "Zpos_add [addprec_object_index [i]]". And additional bit information for each axis in the Z-axis direction is stored.
[0192]
By making it possible to acquire additional bit information in this way, it is possible to obtain quantization position information with sufficient quantization accuracy with a smaller amount of transmission than when acquiring the highest precision quantization position information from the beginning. Become.
[0193]
Next, as described above, a content reproduction system in which a client acquires content and quantization position information from a server and reproduces the content will be described.
[0194]
Such a content reproduction system is configured, for example, as shown in FIG.
[0195]
The content reproduction system shown in FIG. 17 includes an absolute coordinate position information encoder 111, a server 112, a listener position acquisition device 113, a client 114, and an MPEG-H renderer 115. In this content reproduction system, the client 114 calculates the polar coordinate position information of the polar coordinate representation indicating the position of the object as seen from the listener U12.
[0196]
That is, on the content distribution side, the user U11 inputs the position information in the parent space indicating the position of each object in the parent space to the absolute coordinate position information encoder 111.
[0197]
As described above, the position information in the parent space is the absolute coordinate information represented by the coordinates of the three-dimensional absolute coordinate system (Cartesian coordinate system), that is, the xyz coordinate system.
[0198]
The absolute coordinate position information encoder 111 encodes the position information in the parent space input by the user U11, and transmits the coded data file obtained as a result to the server 112.
[0199]
The server 112 is an information processing device having a recording unit 121, and records the coded data file received from the absolute coordinate position information encoder 111 in the recording unit 121. Here, the coded data file contains the highest precision quantization position information of each object.
[0200]
The server 112 generates initial quantization position information and additional bit information from the highest precision quantization position information recorded in the recording unit 121, and transmits the first quantization position information and additional bit information to the client 114.
[0201]
On the client 114 side, the listener position acquisition device 113 acquires the listener position information indicating the position of the listener U12 in the parent space and supplies the listener position information to the client 114. The listener position acquisition device 113 includes, for example, a distance measuring device such as a GPS (Global Positioning System) or a gyro sensor, an input device for acquiring virtual position information of the listener U12 in a virtual space as a parent space, and the like. Will be done.
[0202]
The client 114 is an information processing device having an absolute coordinate position information decoder 131 and a coordinate conversion unit 132, and acquires initial quantization position information and additional bit information from the server 112.
[0203]
The absolute coordinate position information decoder 131 decodes the initial quantization position information acquired from the server 112 and the quantization position information obtained from the initial quantization position information and the additional bit information.
[0204]
The coordinate conversion unit 132 converts the information indicating the position of the object in the parent space obtained by decoding the quantization position information into polar coordinates indicating the relative position of the object as seen from the listener U12 in the parent space. Then, it is used as polar coordinate position information.
[0205]
The coordinate conversion unit 132 supplies the polar coordinate position information of each object to the MPEG-H renderer 115, and the MPEG-H renderer 115 is based on the supplied polar coordinate position information and the audio data of each object acquired from the server 112. And render.
[0206]
Then, the MPEG-H renderer 115 outputs the reproduced audio data obtained by rendering, in which the sound image of the object is localized at each position in the parent space, to a reproduction system such as a speaker, and reproduces the sound. The MPEG-H renderer 115 is a polar coordinate system renderer standardized by MPEG-H.
[0207]
In the content reproduction system as described above, the information indicating the position of the object exchanged between the server 112 and the client 114 is the quantized position information which is an absolute coordinate. Therefore, it is not necessary to consider the position of the listener U12 in the parent space, and there is an advantage that the quantized position information needs to be transmitted to the client 114 only for the moving object.
[0208]
Next, a more detailed configuration example of the absolute coordinate position information encoder 111, the server 112, and the client 114 shown in FIG. 17 will be described.
[0209]
First, a configuration example of the absolute coordinate position information encoder 111 will be described.
[0210]
The absolute coordinate position information encoder 111 is configured as shown in FIG. 18, for example.
[0211]
The absolute coordinate position information encoder 111 shown in FIG. 18 has an acquisition unit 161, an encoding unit 162, and a communication unit 163.
[0212]
The acquisition unit 161 acquires the number of objects of the object to be encoded, the absolute distance absolute distance of the parent space, and the position information in the parent space of each object from an external device or the like and supplies the information to the coding unit 162.
[0213]
The coding unit 162 encodes the position information in the parent space of each object based on the number of objects supplied from the acquisition unit 161, the absolute distance absolute distance in the parent space, and the position information in the parent space of the objects, and obtains the result. The encoded data file is supplied to the communication unit 163.
[0214]
The coding unit 162 includes a child space determination unit 171, a conversion unit 172, and a quantization unit 173. Since the child space determination unit 171 to the quantization unit 173 correspond to the child space determination unit 31 to the quantization unit 33 shown in FIG. 7 and perform the same operation, the description thereof will be omitted.
[0215]
The communication unit 163 transmits the coded data file supplied from the coding unit 162.
[0216]
Further
, the server 112 is configured as shown in FIG. 19, for example. In FIG. 19, the parts corresponding to the case in FIG. 17 are designated by the same reference numerals, and the description thereof will be omitted as appropriate.
[0217]
The server 112 shown in FIG. 19 has a communication unit 201, a control unit 202, and a recording unit 121.
[0218]
The communication unit 201 transmits various information supplied from the control unit 202 to the client 114, and receives various information transmitted from the absolute coordinate position information encoder 111 and the client 114 and supplies the information to the control unit 202. Or
[0219]
The control unit 202 controls the operation of the entire server 112. The control unit 202 has a communication control unit 211 and a transmission information generation unit 212.
[0220]
The communication control unit 211 controls the communication with the absolute coordinate position information encoder 111 and the client 114 by the communication unit 201 by controlling the communication unit 201. The transmission information generation unit 212 generates various information to be transmitted to the client 114 by using the information such as the coded data file recorded in the recording unit 121 as needed.
[0221]
Further
, the client 114 is configured as shown in FIG. 20, for example. In FIG. 20, the parts corresponding to the case in FIG. 17 are designated by the same reference numerals, and the description thereof will be omitted as appropriate.
[0222]
The client 114 shown in FIG. 20 has a communication unit 241, a control unit 242, and an output unit 243.
[0223]
The communication unit 241 transmits various information supplied from the control unit 242 to the server 112, or receives various information transmitted from the server 112 and supplies the information to the control unit 242.
[0224]
The control unit 242 controls the operation of the entire client 114. The control unit 242 includes a communication control unit 251, an absolute coordinate position information decoder 131, and a coordinate conversion unit 132.
[0225]
The communication control unit 251 controls the communication with the server 112 by the communication unit 241 by controlling the communication unit 241. For example, the communication control unit 251 controls the communication unit 241 and functions as an acquisition unit for acquiring the initial quantization position information and the additional bit information from the server 112.
[0226]
The absolute coordinate position information decoder 131 calculates the coded position in the parent space by calculating the position information in the parent space indicating the absolute position of the object in the parent space based on the quantization position information and the additional bit information. It functions as a calculation unit that decodes information.
[0227]
The output unit 243 outputs the polar coordinate position information of each object obtained by the coordinate conversion by the coordinate conversion unit 132 to the MPEG-H renderer 115.
[0228]
Next, specific processing performed in the content reproduction system will be described.
[0229]
First, the coding process by the absolute coordinate position information encoder 111 and the file saving process by the server 112 will be described with reference to the flowchart of FIG.
[0230]
When the coding process is started by the absolute coordinate position information encoder 111, the acquisition unit 161 in step S71 obtains the number of objects input by the user U11, the absolute distance absolute distance in the parent space, and the position information in the parent space of each object. It is acquired and supplied to the coding unit 162.
[0231]
In step S72, the child space determination unit 171 determines the child space based on the position information in the parent space of each object supplied from the acquisition unit 161. In step S72, the same processing as in step S12 of FIG. 8 is performed to determine the child space.
[0232]
In step S73, the conversion unit 172 converts the position information in the parent space of each object supplied from the acquisition unit 161 into the position information in the child space. In step S73, the same process as in step S13 of FIG. 8 is performed.
[0233]
In step S74, the quantization unit 173 determines the highest precision split index.
[0234]
For example, in step S74, the above-mentioned equation (3) is calculated based on the predetermined minimum distance MinDist, the determination result of the child space, and the absolute distances absoluteDistanceX, absoluteDistanceY, and absoluteDistanceZ obtained from the absolute distance absoluteDistance. A precision split index is required.
[0235]
In step S75, the quantization unit 173 obtains the quantization step width by performing the calculation of the above equation (4) based on the highest precision division index determined in step S74.
[0236]
In step S76, the quantization unit 173 quantizes the position information in the child space of each object obtained in step S73.
[0237]
For example, in step S76, the above-mentioned equation (5) is calculated using the quantization step width obtained in step S75, and the coordinates of the position information in the parent space in each axial direction are quantized, and each of them is performed. The highest precision quantization position information of the object can be obtained.
[0238]
By the above processing of steps S72 to S76, the position information in the parent space is encoded.
[0239]
In step S77, the coding unit 162 includes the highest precision quantization position information, the number of objects and the absolute distance absoluteDistance acquired in step S71, and the child space information and the object affiliation information obtained from the determination result of the child space in step S72. A coded data file is generated and supplied to the communication unit 163.
[0240]
Here, for example, the child space information includes information indicating the number of child spaces contained in the parent space, the highest precision division index in each child space, and parameters of a linear function for coordinate conversion for each child space. It is included.
[0241]
In step S78, the communication unit 163 transmits the coded data file supplied from the coding unit 162 to the server 112 by wireless or wired communication or the like, and the coding process ends.
[0242]
As will be described later, when the coded data file is correctly saved in the server 112, a save completion notification to that effect is transmitted from the server 112, so that the absolute coordinate position information encoder 111 has sent the save completion notification. Is received and displayed as appropriate.
[0243]
Further, when the coded data file is transmitted, the server 112 starts the file saving process.
[0244]
That is, in step S101, the communication unit 201 of the server 112 receives the coded data file transmitted from the absolute coordinate position information encoder 111 under the control of the communication control unit 211, and supplies the coded data file to the control unit 202.
[0245]
In step S102, the control unit 202 supplies the coded data file supplied from the communication unit 201 to the recording unit 121 and stores it. As a result, the highest precision quantization position information and the like are stored (recorded) in the recording unit 121.
[0246]
After that, the communication control unit 211 controls the communication unit 201 to send a save completion notification to the effect that the coded data file has been correctly saved to the absolute coordinate position information encoder 111, and the file saving process ends.
[0247]
As described above, the absolute coordinate position information encoder 111 converts the position information in the parent space into the position information in the child space by coordinate conversion, and quantizes the position information in the child space to generate a coded data file. By doing so, it is possible to reduce the amount of code of the quantization position information as compared with the case where the position information in the parent space is encoded (quantized) as it is.
[0248]
Further, the server 112 receives and stores the coded data file transmitted from the absolute coordinate position information encoder 111. As a result, it becomes possible to generate the quantization position information of arbitrary quantization accuracy from the highest precision quantization position information in response to the request of the client 114 and transmit it to the client 114.
[0249]
As a result, the processing load on the content distribution side such as the server 112 and the absolute coordinate position information encoder 111 can be reduced by obtaining the polar coordinate position information of each object in the client 114, and the amount of information transmission can be reduced.
[0250]
When the
coded data file is saved in the server 112, the client 114 can receive the quantized position information of each object for the content from the server 112. become.
[0251]
Hereinafter, the process performed when the client 114 acquires the initial quantization position information from the server 112 will be described. That is, the position information acquisition process by the client 114 and the position information transmission process by the server 112 will be described below with reference to the flowchart of FIG.
[0252]
When the position information acquisition process is started by the client 114, in step S131, the communication unit 241 transmits a header information transmission request to the server 112 under the control of the communication control unit 251.
[0253]
When the header information transmission request is transmitted, the server 112 starts the position information transmission process. That is, in step S161, the communication unit 201 receives the header information transmission request transmitted from the client 114 according to the control of the communication control unit 211, and supplies the header information transmission request to the control unit 202.
[0254]
Then, the transmission information generation unit 212 refers to the coded data file recorded in the recording unit 121, and generates header information including the absolute distance absoluteDistance, the child space information, the number of objects, and the object affiliation information. As a result, for example, the header information shown in FIG. 11 is generated.
[0255]
When the transmission information generation unit 212 generates the header information, the communication control unit 211 supplies the generated header information to the communication unit 201.
[0256]
In step S162, the communication unit 201 transmits the header information supplied from the communication control unit 211 to the client 114 under the control of the communication control unit 211.
[0257]
Then, in the client 114, in step S132, the communication unit 241 receives the header information transmitted from the server 112 according to the control of the communication control unit 251 and supplies it to the control unit 242.
[0258]
When the header information is obtained in this way, the control unit 242 first quantizes it based on the status of the transmission band with the server 112, the number of objects included in the header information, the highest precision division index, and the like. Determines the quantization accuracy of position information, that is, the split index.
[0259]
Then, the control unit 242 generates an initial quantization position information transmission request requesting transmission of the initial quantization position information based on the determined quantization accuracy, and supplies the first quantization position information transmission request to the communication unit 241. For example, here, the initial quantization position information transmission request shown in FIG. 13 is generated.
[0260]
In step S133, the communication unit 241 transmits the initial quantization position information transmission request supplied from the control unit 242 to the server 112 under the control of the communication control unit 251.
[0261]
Then, in the server 112, in step S163, the communication unit 201 receives the initial quantization position information transmission request transmitted from the client 114 and supplies it to the control unit 202 according to the control of the communication control unit 211.
[0262]
In step S164, the transmission information generation unit 212 refers to the coded data file recorded in the recording unit 121 in response to the initial quantization position information transmission request supplied from the communication unit 201, and refers to the initial quantization position. Generate information.
[0263]
For example, in the example of FIG. 5, when the division index requested by the initial quantization position information transmission request is "9", the transmission information generation unit 212 has the highest precision quantization position information included in the coded data file. The upper 9 bits of are extracted and used as the initial quantization position information.
[0264]
In the example of FIG. 5, "001100110" obtained by extracting the same number of bits as the division index "9" from the coordinate value "00110011001" in the predetermined axial direction of the highest precision quantization position information, that is, only 9 bits from the most significant bit side. Is the coordinate value in the predetermined axial direction of the initial quantization position information.
[0265]
When the highest precision division index is requested by the first quantization position information transmission request, the highest precision quantization position information is used as it is as the first quantization position information.
[0266]
When the transmission information generation unit 212 generates the initial quantization position information for each object, the transmission information generation unit 212 stores the initial quantization position information in the bit stream in the format shown in FIG. 14, and the communication control unit 211 generates the generated bit stream. Is supplied to the communication unit 201.
[0267]
In step S165, the communication unit 201 transmits the bit stream supplied from the communication control unit 211, that is, the initial quantization position information of each object to the client 114 under the control of the communication control unit 211, and the position information transmission process ends. ..
[0268]
When the initial quantization position information is transmitted by the server 112, the client 114 performs the process of step S134.
[0269]
In step S134, the communication unit 241 receives the initial quantization position information transmitted from the server 112 under the control of the communication control unit 251 and supplies it to the control unit 242. That is, the communication control unit 251 acquires the initial quantization position information.
[0270]
In step S135, the absolute coordinate position information decoder 131 describes the above-described equation (6) for the initial quantization position information of each object based on the initial quantization position information and the division index stored in the initial quantization position information transmission request. By performing the same calculation as above, the initial quantization position information is inversely quantized. As a result, the position information in the decoder space can be obtained.
[0271]
In step S136, the absolute coordinate position information decoder 131 converts the position information in the child space obtained for each object into the position information in the parent space.
[0272]
That is, the absolute coordinate position information decoder 131 of the above equations (7) and (8) based on the position information in the decoder space obtained in step S135 and the child space information included in the header information. Inverse scaling conversion is performed by performing the calculation. As a result, the position information in the decoder space is converted into the position information in the decoding parent space.
[0273]
Further, the absolute coordinate position information decoder 131 multiplies the obtained position information in the decoding parent space by the absolute distance absoluteDistance included in the header information to indicate the position information of the object in the parent space. Ask for. Then, when the position information indicating the position of the object is requested, the position information acquisition process ends.
[0274]
The absolute coordinate position information decoder 131 performs the above steps S135 and S136 as decoding processes for each object.
[0275]
As described above, the client 114 acquires and decodes the quantization position information having an appropriate quantization accuracy as the initial quantization position information. Further, the server 112 generates the initial quantization position information in response to the request of the client 114 and transmits it to the client 114.
[0276]
By doing so, it is possible to reduce the amount of code of the information indicating the position of the object as compared with the case where the position information in the parent space is encoded (quantized) as it is.
[0277]
By the way, when the position information acquisition process described with reference to FIG. 22 is performed, position information indicating the position of the object in the parent space for each object (hereinafter, , Also called object position information).
[0278]
For example, when the highest-precision quantization position information is acquired as the initial quantization position information, the most accurate object position information can be obtained, but in other cases, the accuracy of the object position information may not be sufficient.
[0279]
Furthermore, although the initial quantization position information with relatively low quantization accuracy was acquired, there was a margin in the transmission band and there was a margin before the content playback time, so the quantization position information with higher quantization accuracy was obtained. Sometimes you want to get additional bit information to get.
[0280]
Therefore, when the object position information is obtained, the client 114 acquires additional bit information from the server 112 as needed. Hereinafter, the processing performed by the client 114 and the server 112 in such a case will be described.
[0281]
That is, the additional bit information acquisition process by the client 114 and the additional bit information transmission process by the server 112 will be described below with reference to the flowchart of FIG. 23. Note that these processes are not performed when the highest-precision quantization position information is acquired as the initial quantization position information.
[0282]
When the additional bit information acquisition process is started in the client 114, the absolute coordinate position information decoder 131 calculates the distance from the listener U12 to the object for each object in step S191.
[0283]
Specifically, from the listener U12 to the object in the parent space based on the listener position information supplied from the listener position acquisition device 113 and the object position information obtained in the process of step S136 of FIG. 22. Euclidean distance is calculated.
[0284]
In step S192, the absolute coordinate position information decoder 131 identifies an object for which additional bit information is to be acquired.
[0285]
For example, as the quantization position information of an object in the parent space, whether or not the initial quantization position information, that is, the position information in the decoding parent space is sufficiently accurate is indicated by the object position information from the listener U12. It can be specified by the distance to the position.
[0286]
Therefore, the absolute coordinate position information decoder 131 obtains the tolerance of the quantum step width from the distance from the listener U12 to the object and the perceived limit angle, and the quantization obtained by the margin of error and the division index of the initial quantization position information. Compare with step width. Then, the absolute coordinate position information decoder 131 identifies an object in which the quantization accuracy of the initial quantization position information is not sufficient, that is, an object for acquiring additional bit information, based on the comparison result.
[0287]
Further, for example, the absolute coordinate position information decoder 131 provides additional bit information based on the bandwidth status of the transmission line with the server 112, the time from the current time to the playback time of the content, the quantization accuracy of the initial quantization position information, and the like. Identify the object to get.
[0288]
Further, the absolute coordinate position information decoder 131 determines the required number of bits of the additional bit information, that is, the required quantization accuracy of the object for which the additional bit information is acquired.
[0289]
Whether or not to acquire the additional bit information and the number of bits of the required additional bit information can be determined for each axial direction of the XYZ coordinate system for each object.
[0290]
In step S193, the absolute coordinate position information decoder 131 selects the request mode based on the specific result of step S192.
[0291]
Here, an appropriate request mode is selected based on, for example, the status of the transmission band, the number of objects for which additional bit information is acquired, the positional relationship of a plurality of objects for which additional bit information is acquired, and the like.
[0292]
In step S194, the absolute coordinate position information decoder 131 generates an additional bit information transmission request based on the specific result of step S192 and the selection result of the request mode in step S193. As a result, for example, the additional bit information transmission request shown in FIG. 15 is generated.
[0293]
The communication control unit 251 supplies the additional bit information transmission request generated in this manner to the communication unit 241 and causes the communication unit 241 to transmit the additional bit information transmission request to the server 112.
[0294]
In step S195, the communication unit 241 transmits the additional bit information transmission request supplied from the control unit 242 to the server 112 under the control of the communication control unit 251.
[0295]
When the additional bit information transmission request is transmitted, the server 112 starts the additional bit information transmission process.
[0296]
That is, in step S221, the communication unit 201 receives the additional bit information transmission request transmitted from the client 114 according to the control of the communication control unit 211, and supplies the additional bit information transmission request to the control unit 202.
[0297]
In step S222, the transmission information generation unit 212 generates additional bit information of the object for which transmission is requested in response to the additional bit information transmission request supplied from the communication unit 201.
[0298]
For example, the transmission information generation unit 212 transmits additional bit information from a position determined by the quantization accuracy (division index) of the initial quantization position information in the highest accuracy quantization position information of the coded data file recorded in the recording unit 121. Additional bit information is generated by extracting the bits for the number of additional bits specified by the request. The generation of such additional bit information can be realized by bit shifting.
[0299]
When the transmission information generation unit 212 generates additional bit information for the required object, the transmission information generation unit 212 stores the additional bit information in the additional bit stream in the format shown in FIG. 16, and the communication control unit 211 stores the generated additional bit stream. It is supplied to the communication unit 201.
[0300]
In step S223, the communication unit 201 transmits an additional bit stream including the additional bit information supplied from the communication control unit 211 to the client 114 under the control of the communication control unit 211, and the additional bit information transmission process ends.
[0301]
Further, when the additional bit stream is transmitted, the client 114 performs the process of step S196.
[0302]
In step S196, the communication unit 241 receives the additional bit stream transmitted from the server 112 and supplies it to the control unit 242 according to the control of the communication control unit 251. That is, the communication control unit 251 acquires an additional bit stream including the additional bit information.
[0303]
In step S197, the absolute coordinate position information decoder 131 generates the quantized position information by adding the additional bit information to the lower bit side of the initial quantization position information for the object for which the transmission of the additional bit information is requested.
[0304]
In step S198, the absolute coordinate position information decoder 131 dequantizes the quantized position information obtained in the process of step S197.
[0305]
Then, in step S199, the absolute coordinate position information decoder 131 converts the position information in the child space obtained by the inverse quantization in step S198 into the position information in the parent space.
[0306]
In step S198 and step S199, the same processing as in step S135 and step S136 of FIG. 22 is performed.
[0307]
Further, in step S199, the absolute coordinate position information decoder 131 obtains the object position information by multiplying the obtained position information in the decoding parent space by the absolute distance absoluteDistance.
[0308]
The above steps S197 to S199 are performed only on the object for which the additional bit information has been acquired.
[0309]
Further, with respect to the object position information thus obtained, the above-described processes of steps S191 to S199 are repeated until the object position information (position information in the parent space) having sufficient quantization accuracy is obtained. You may.
[0310]
In step S200, the coordinate conversion unit 132 determines the position of the object as seen from the listener U12 in the parent space based on the object position information and the listener position information supplied from the listener position acquisition device 113 for each object. Calculate the polar coordinate position information indicating.
[0311]
Then, the coordinate conversion unit 132 supplies the obtained polar coordinate position information to the output unit 243, and the output unit 243 outputs the polar coordinate position information to the MPEG-H renderer 115. When the output unit 243 outputs the polar coordinate position information of all the objects to the MPEG-H renderer 115, the additional bit information acquisition process ends.
[0312]
When the additional bit information acquisition process is completed, the MPEG-H renderer 115 renders. That is, the MPEG-H renderer 115 renders based on the audio data of each object acquired from the server 112 or the like and the polar coordinate position information supplied from the output unit 243, and the sound image of the object at each position in the parent space. Generates playback audio data that is localized and outputs it to a playback system such as a speaker.
[0313]
As described above, the client 114 acquires additional bit information as needed and obtains the final object position information. Further, the server 112 generates and transmits additional bit information in response to the request of the client 114.
[0314]
By doing so, the client 114 can obtain the position information in the parent space with sufficient accuracy with a smaller amount of information transmission than the case where the highest precision quantization position information is acquired from the server 112 from the beginning.
[0315]
Moreover, since the processing for obtaining the required quantization accuracy and the processing for calculating the polar coordinate position information are all performed by the client 114, it is possible to reduce the processing load on the content distribution side such as the server 112 and the absolute coordinate position information encoder 111. can.
[0316]
By the way, in the above, by forming one or more child spaces in the parent space, the position information in the parent space is efficiently coded. An example of conversion has been described.
[0317]
However, as shown in FIG. 24, for example, each of the child spaces originally included in each of the two different parent spaces may be arranged in one virtual parent space.
[0318]
In the example shown in FIG. 24, the child space CA31 and the child space CA32 were originally contained in different parent spaces, but here, the child space CA31 and the child space are contained in one virtual parent space. CA32 is deployed.
[0319]
By doing so, for example, the listener U12 can select only one of the child space CA31 and the child space CA32 according to his / her preference. In this case, the server 112, depending on the selection of the listener U12, the audio data of the objects in the child space selected by the listener U12 of the child space CA31 and the child space CA32, and the quantum of those objects. The quantized position information is transmitted to the client 114.
[0320]
As a result, on the client 114 side, the audio data can be reproduced only in the child space selected by the listener U12.
[0321]
In such an example, for example, the child space CA31 and the child space CA32 may be a space having different reverberation characteristics. That is, the child space CA31 and the child space CA32 may be given different reverberation characteristics. Further, the exact same objects may exist in the child space CA31 and the child space CA32 in the same positional relationship.
[0322]
By doing so, the listener U12 can play audio data with different reverberation characteristics simply by selecting the child space, so that even if the same object is placed, the experience is as if they were in different spaces. can. In addition, the listener U12 can experience as if he / she is going back and forth between a plurality of completely different spaces such as the child space CA31 and the child space CA32. Further, when a plurality of child spaces are formed in one parent space, a child space may be formed for each attribute of each object such as a musical instrument object or a cheering object. In this case, each of the objects contained in each of the different child spaces will have different attributes.
[0323]
In addition, for example, as shown in FIG. 25, a plurality of child spaces may overlap each other. In this example, the child space CA41 and the child space CA42 are formed in one parent space, but the entire child space CA42 is included in the child space CA41.
[0324]
That is, the object existing in the child space CA42 belongs not only to the child space CA42 but also to the child space CA41 at the same time.
[0325]
Even in such a case, the listener U12 can select either the child space CA41 or the child space CA42, and the audio data of the object in the selected child space can be reproduced. can.
[0326]
In such a case, for example, when playing back content related to the parent space, the object to be played can be changed depending on which child space is selected.
[0327]
For example, assuming that the content of a concert is played back, an object that exists only in the child space CA42 is an instrument object, and an object that exists in the child space CA41 but does not exist in the child space CA42 is a cheering object. And.
[0328]
In this case, if the listener U12 selects the child space CA41, both the sound of the musical instrument and the cheers will be reproduced, and if the listener U12 selects the child space CA42, only the sound of the musical instrument will be reproduced. That is, it is possible to freely switch between a space containing only musical instruments and a space containing not only musical instruments but also cheers.
[0329]
When one object can belong to a plurality of child spaces as in the example shown in FIG. 25, the header information described above will be as shown in FIG. 26, for example.
[0330]
In the example shown in FIG. 26, in addition to the case shown in FIG. 11, the header information newly stores information indicating the number of child spaces to which the i-th object indicated by the character "numOfAttChildArea" belongs for each object. Has been done. Then, following the information "numOfAttChildArea" indicating the number of child spaces to which the object belongs, the object belonging information "ChildAreaIndexForObj [i]" is stored as many as the number indicated by the information.
[0331]
Note that the quantization of the position information in the child space of an object that overlaps with multiple child spaces is performed for any of the child spaces to which the object belongs, such as the child space specified by ChildAreaIndexForObj [0]. You just have to make it happen. Further, when the same object belongs to a plurality of child spaces, it is sufficient to determine in advance which child space's position information in the child space is quantized for the object.
[0332]
That is, only the position information in the child space for one child space among the plurality of child spaces to which the object belongs is quantized and transmitted, and the quantized position information for the remaining child spaces is not transmitted. On the client 114 side, the position information in the child space of one child space of an object may be obtained from the position information in the child space of another child space of the same object.
[0333]
Further, the format of the child space information "ChildAreaInfo ()" portion in the header information shown in FIG. 26 is as shown in FIG. 27, for example. Since the format of the child space information shown in FIG. 27 is the same as that shown in FIG. 12, the description thereof will be omitted.
[0334]
Further
, in the free viewpoint audio, the server 112 and the client 114 communicate with each other, whereby information such as quantization position information can be obtained. Be exchanged.
[0335]
However, a delay occurs depending on the congestion status of the network between the server 112 and the client 114, that is, the transmission line, and the client 114 cannot acquire the information necessary for reproduction such as the quantization position information by the reproduction time of the audio data. There is a possibility.
[0336]
Therefore, for example, when the transmission delay becomes large as shown in FIG. 28, the number of frames to be skipped may be transmitted from the client 114 to the server 112.
[0337]
In this case, the server 112 transmits the quantization position information of the object ahead in time for the number of skip frames specified by the client 114 and the interpolation information from the current playback time to the skip frame time.
[0338]
For example, as shown by arrow Q61 in FIG. 28, the object that was in the position (X1, Y1, Z1) in the XYZ coordinate system in frame 1 is in the position (X2, Y2, Z1) in the XYZ coordinate system in frame 8 after 7 frames. It is assumed that the content that seems to be moving to Z2) is played.
[0339]
At this time, the quantized position information could be transmitted and received without delay up to frame 1 and frame 2 as shown by arrow Q62, but at this point, the bandwidth of the transmission line is narrow and there is a high possibility that a delay will occur due to the client 114. It is assumed that it is judged.
[0340]
Then, the client 114 does not request the server 112 to transmit the initial quantization position information of the frame 3, but instead transmits the number of request skip frames for specifying the frame 8 and the interpolation information used for the interpolation processing. Request.
[0341]
Here, the request skip frame number is information indicating the number of frames for skipping the transmission of the quantization position information. Here, for example, since the request skip frame number “5” is transmitted to the server 112, the server 112 skips 5 frames after the frame 2 in which the transmission is completed, and the quantization position is obtained for the next frame 8. Send information and interpolation information.
[0342]
As a result, on the client 114 side, the position information in the parent space can be obtained by the decoding process for the quantized position information of the frame 8. Further, the client 114 can obtain the position information in the parent space in the frames 3 to 7 by the interpolation process based on the position information in the parent space of the frame 2, the position information in the parent space of the frame 8, and the interpolation information.
[0343]
By performing such interpolation processing, it is possible to obtain the position of the object moving from the position (X1, Y1, Z1) to the position (X2, Y2, Z2) at an arbitrary timing (frame).
[0344]
When the client 114 transmits the number of request skip frames to the server 112 in this way, the initial quantization position information transmission request transmitted by the client 114 is as shown in FIG. 29, for example.
[0345]
In the example shown in FIG. 29, in the initial quantization position information transmission request, the divided indexes "exp_index_initial_x [i]", "exp_index_initial_y [i]", and "exp_index_initial_z" in each axis direction of the i-th object shown in FIG. In addition to "[i]", it also contains the number of request skip frames indicated by the letter "SkipFrameRequest".
[0346]
Further, the server 112 that has received the initial quantization position information transmission request shown in FIG. 29 transmits the initial quantization position information in the bitstream format shown in FIG. 30, for example.
[0347]
In the example shown in FIG. 30, in addition to the initial quantization position information “Xpos_initial [i]”, “Ypos_initial [i]”, and “Zpos_initial [i]” shown in FIG. It contains the interpolation information indicated by "[i]".
[0348]
In particular, here, the interpolation information InterpolMode [i] is information indicating whether to perform first-order interpolation or second-order interpolation for the i-th object. However, when the initial quantization position information transmission request shown in FIG. 13 is received from the client 114 instead of FIG. 29, the interpolation information InterpolMode [i] is not stored in the bit stream.
[0349]
For example, if the object is moving in a motion equivalent to a constant velocity motion, the first-order interpolation is performed, and if the object is moving in a motion corresponding to a constant acceleration motion, the second-order interpolation is performed. It should be.
[0350]
Here, an example in which either the first-order interpolation or the second-order interpolation is selectively performed will be described, but other interpolation methods may of course be used. Further, the coefficient used for the interpolation process may be stored as the interpolation information, or the coefficient itself used for the interpolation process may be held in advance by the client 114.
[0351]
Further, the initial quantization position information stored in the bit stream in the example shown in FIG. 30 is that of the frame after skipping by the number of requested skip frames.
[0352]
Here, the process performed when the client 114 transmits the initial quantization position information transmission request shown in FIG. 29 due to a delay in the transmission line or the like will be described.
[0353]
That is, the position information acquisition process by the client 114 and the position information transmission process by the server 112 will be described below with reference to the flowchart of FIG. 31.
[0354]
Note that the control unit 242 uses the quantization position information based on, for example, the bandwidth status of the transmission line between the server 112 and the client 114, that is, the network traffic, the time stamp at the time of content reproduction, that is, the content reproduction status. Decide whether to skip frames for and how many frames to skip.
[0355]
When the position information acquisition process is started by the client 114, in step S251, the communication unit 241 transmits the first quantization position information transmission request supplied from the control unit 242 to the server 112 under the control of the communication control unit 251. ..
[0356]
That is, the control unit 242 generates the initial quantization position information transmission request shown in FIG. 29 based on the network traffic, the reproduction status of the content, and the like, and supplies the first quantization position information transmission request to the communication unit 241. This initial quantization position information transmission request includes the number of request skip frames.
[0357]
Then, in the server 112, in step S291, the communication unit 201 receives the initial quantization position information transmission request transmitted from the client 114 and supplies it to the control unit 202 according to the control of the communication control unit 211.
[0358]
In step S292, the transmission information generation unit 212 refers to the coded data file recorded in the recording unit 121 in response to the initial quantization position information transmission request supplied from the communication unit 201, and refers to the initial quantization position. Generate information.
[0359]
In this case, for example, in the bitstream format shown in FIG. 30, a bitstream including the initial quantization position information and interpolation information of the first frame after skipping is generated. The communication control unit 211 supplies the bit stream thus generated to the communication unit 201. At this time, the transmission information generation unit 212 identifies the movement of the object from the quantization position information in the frames before and after the skip and the skipped frame, that is, the position of the object, and the bit stream includes the interpolation information according to the specific result. To be able to.
[0360]
In step S293, the communication unit 201 transmits the bit stream supplied from the communication control unit 211, that is, the initial quantization position information of each object to the client 114 under the control of the communication control unit 211, and the position information transmission process ends. ..
[0361]
When the initial quantization position information is transmitted by the server 112, in step S252, the communication unit 241 receives the initial quantization position information transmitted from the server 112 according to the control of the communication control unit 251 and sends it to the control unit 242. Supply.
[0362]
When the initial quantization position information is acquired, the processes of steps S253 and S254 are subsequently performed, but since these processes are the same as the processes of steps S135 and S136 of FIG. 22, the description thereof will be omitted. .. However, here, processing is performed on the first frame after skipping, and the position information in the decoding parent space of that frame is obtained.
[0363]
In step S255, the absolute coordinate position information decoder 131 obtains the position information in the decoding parent space of the frame after skipping obtained in step S254 and the position information in the decoding parent space of the frame before skipping obtained immediately before that. Based on this, the interpolation process indicated by the interpolation information received together with the initial quantization position information is performed in step S252.
[0364]
That is, for example, the absolute coordinate position information decoder 131 generates the position information in the decoding parent space of each skipped frame by performing the processing indicated by the interpolation information, that is, the first-order interpolation or the second-order interpolation.
[0365]
Further, the absolute coordinate position information decoder 131 multiplies the position information in the decoding parent space of each frame by the absolute distance absoluteDistance for the skipped frame and the first frame after skipping obtained by the above processing. Obtain the object position information in each frame. When the object position information is obtained in this way, the position information acquisition process ends.
[0366]
As described above, when skipping some frames, the client 114 transmits the first quantization position information transmission request including the number of request skip frames, obtains the position information in the parent space of the skipped frames, and also obtains the position information in the parent space of the skipped frames. Obtain the position information in the parent space of the skipped frame by the interpolation process. Further, the server 112 transmits the initial quantization position information and the interpolation information to the client 114 according to the number of request skip frames.
[0367]
By doing so, even if a delay occurs in the transmission line, it is possible to suppress the delay in reproducing the content.
[0368]
As described above, according to the present technology, when the existence position of an object is biased in the parent space, the required accuracy is maintained with a small amount of transmission by transmitting information indicating the position of the object using the child space. Information indicating the position of the object can be transmitted as it is.
[0369]
Further, as in the first modification of the second embodiment, the listener is required by treating two or more completely different spaces as child spaces and managing those child spaces within one parent space. By decoding only the child space, it is possible to experience the traffic of multiple completely different spaces.
[0370]
Further, as in the second modification of the second embodiment, even when the transmission line is congested and the required bit stream does not reach by the reproduction time, the object can be smoothly moved and reproduced by the look-ahead and interpolation.
[0371]
By the way, the series of processes described above can be executed by hardware or software. When a series of processes are executed by software, the programs that make up the software are installed on the computer. Here, the computer includes a computer embedded in dedicated hardware and, for example, a general-purpose personal computer capable of executing various functions by installing various programs.
[0372]
FIG. 32 is a block diagram showing a configuration example of hardware of a computer that executes the above-mentioned series of processes programmatically.
[0373]
In a computer, a CPU (Central Processing Unit) 501, a ROM (Read Only Memory) 502, and a RAM (Random Access Memory) 503 are connected to each other by a bus 504.
[0374]
An input / output interface 505 is further connected to the bus 504. An input unit 506, an output unit 507, a recording unit 508, a communication unit 509, and a drive 510 are connected to the input / output interface 505.
[0375]
The input unit 506 includes a keyboard, a mouse, a microphone, an image sensor, and the like. The output unit 507 includes a display, a speaker, and the like. The recording unit 508 includes a hard disk, a non-volatile memory, and the like. The communication unit 509 includes a network interface and the like. The drive 510 drives a removable recording medium 511 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.
[0376]
In the computer configured as described above, the CPU 501 loads the program recorded in the recording unit 508 into the RAM 503 via the input / output interface 505 and the bus 504 and executes the above-described series. Is processed.
[0377]
The program executed by the computer (CPU 501) can be recorded and provided on a removable recording medium 511 as a package medium or the like, for example. Programs can also be provided via wired or wireless transmission media such as local area networks, the Internet, and digital satellite broadcasts.
[0378]
In a computer, the program can be installed in the recording unit 508 via the input / output interface 505 by mounting the removable recording medium 511 in the drive 510. Further, the program can be received by the communication unit 509 and installed in the recording unit 508 via a wired or wireless transmission medium. In addition, the program can be pre-installed in the ROM 502 or the recording unit 508.
[0379]
The program executed by the computer may be a program that is processed in chronological order according to the order described in this specification, or may be a program that is processed in parallel or at a necessary timing such as when a call is made. It may be a program in which processing is performed.
[0380]
Further, the embodiment of the present technology is not limited to the above-described embodiment, and various changes can be made without departing from the gist of the present technology.
[0381]
For example, the present technology can have a cloud computing configuration in which one function is shared by a plurality of devices via a network and jointly processed.
[0382]
Further, each step described in the above-mentioned flowchart can be executed by one device or can be shared and executed by a plurality of devices.
[0383]
Further, when a plurality of processes are included in one step, the plurality of processes included in the one step can be executed by one device or shared by a plurality of devices.
[0384]
Further, the present technology can also have the following configurations.
[0385]
(1)
Acquire spatial information regarding the position and size of the child space included in the parent space in the parent space, and position information in the child space indicating the position of the object included in the child space in the child space. An
information
processing device including an acquisition unit for calculating a position information in a parent space indicating the position of the object in the parent space based on the spatial information and the position information in the child space .
(2) The information processing apparatus according
to
(1), wherein the parent space is represented by absolute coordinates .
(3) The information processing apparatus according to (1) or (2),
wherein the spatial information is a conversion coefficient for coordinate conversion
.
(4) The information processing apparatus according to (3),
wherein the conversion coefficient is a parameter of a linear function
.
(5) The
spatial information is information indicating the position of a reference vertex of the child space in the parent space and information indicating the position of a diagonal vertex of the child space with respect to the vertex in the parent space.
The information processing apparatus according to (1) or (2).
(6) The information processing apparatus according to any one of (1) to (5),
wherein the child space is represented by absolute coordinates
.
(7)
The information processing device according to any one of (1) to (6), wherein a plurality of the child spaces are included in the parent space .
(8) The information processing device according to (7),
wherein each of the objects included in each of the plurality of child spaces has different attributes from each other
.
(9) The information processing apparatus according to (7),
wherein a plurality of the child spaces overlap each other
.
(10) The information processing apparatus according to (7),
wherein the plurality of child spaces have different reverberation characteristics
.
(11) The information processing device according to any one of (1) to (10),
wherein the child space includes a plurality of the objects
.
(12) The information processing apparatus according to any one of (1) to (11)
, wherein one object is included in a plurality of the child spaces
.
(13) The
acquisition unit further acquires additional position information for obtaining the position information in the child space, which is more accurate than the acquired position information in the child space.
The calculation unit calculates the position information in the parent space based on the more accurate position information in the child space and the space information obtained based on the acquired position information in the child space and the additional position information. to
(1) through the information processing apparatus according to any one of (12).
(14) The information processing device according to any one of (1) to (13),
wherein the object is an audio object
.
(15) In the child space
, the information processing device indicates the
spatial information regarding the position and size of the child space included in the parent space in the parent space and the position of the object included in the child space in the child space. An information processing method that acquires
position information and calculates position information in the parent space indicating the position of the object in the parent space based on the space information and the position information in the child space
.
(16)
Acquire spatial information regarding the position and size of the child space included in the parent space in the parent space, and position information in the child space indicating the position of the object included in the child space in the child space. and,
based on the spatial information and the terminal space position information, calculates the parent space position information indicating a position of the object in the parent space
a program for executing the processing including a step in the computer.
Code description
[0386]
11 Encoder, 21 Acquisition unit, 22 Coding unit, 23 Communication unit, 31 Child space determination unit, 32 Conversion unit, 33 Quantization unit, 61 Decoder, 71 Acquisition unit, 72 Decoding unit, 73 Output unit, 81 Inverse quantization unit Part, 82 Conversion part
The scope of the claims
[Claim 1]
An acquisition unit that acquires spatial information regarding the position and size of a child space included in the parent space in the parent space and position information in the child space indicating the position of an object included in the child space in the child space. An
information
processing device including a calculation unit that calculates position information in the parent space indicating the position of the object in the parent space based on the space information and the position information in the child space .
[Claim 2]
The information processing device according to claim 1, wherein the parent space is represented by absolute coordinates .
[Claim 3]
The information
processing device according to claim 1, wherein the spatial information is a conversion coefficient for coordinate conversion .
[Claim 4]
The information processing apparatus according to claim 3, wherein the conversion coefficient is a parameter of a linear function .
[Claim 5]
The spatial information, the parent information indicating the position of the vertex serving as a reference of the child space in the space, and the information indicating the position of the opposite corners with respect to the vertices of the child space in the parent space
Claim 1 The information processing device described in.
[Claim 6]
The information processing device according to claim 1, wherein the child space is represented by absolute coordinates .
[Claim 7]
The information processing device according to claim 1, wherein a plurality of the child spaces are included in the parent space .
[Claim 8]
The information processing device according to claim 7, wherein each of the objects included in each of the plurality of child spaces has different attributes from each other .
[Claim 9]
The information processing device according to claim 7, wherein the plurality of child spaces overlap each other .
[Claim 10]
The information processing device according to claim 7, wherein the plurality of child spaces have different reverberation characteristics .
[Claim 11]
The information processing device according to claim 1, wherein the child space includes a plurality of the objects .
[Claim 12]
The information processing device according to claim 1, wherein one object is included in a plurality of child spaces .
[Claim 13]
The acquisition unit further acquires additional position information for obtaining the child space position information having higher accuracy than
the acquired child space position information, and the calculation unit obtains the acquired child space position information and the child space position information.
The information processing device according to claim 1, wherein the information processing apparatus according to claim 1 calculates the position information in the parent space based on the more accurate position information in the child space obtained based on the additional position information and the space information .
[Claim 14]
The information processing device according to claim 1, wherein the object is an audio object .
[Claim 15]
The information processing device has
spatial information regarding the position and size of the child space included in the parent space in the parent space, and position information in the child space indicating the position of the object contained in the child space in the child space. acquires,
on the basis of the spatial information and the terminal space position information, calculates the parent space position information indicating a position of the object in the parent space
information processing method.
[Claim 16]
The spatial information regarding the position and size of the child space included in the parent space in the parent space and the position information in the child space indicating the position of the object included in the child space in the child space are acquired, and the
above-mentioned
A program that causes a computer to execute a process including a step of calculating a position information in a parent space indicating the position of the object in the parent space based on the spatial information and the position information in the child space .
| # | Name | Date |
|---|---|---|
| 1 | 202017053124-TRANSLATIOIN OF PRIOIRTY DOCUMENTS ETC. [07-12-2020(online)].pdf | 2020-12-07 |
| 2 | 202017053124-STATEMENT OF UNDERTAKING (FORM 3) [07-12-2020(online)].pdf | 2020-12-07 |
| 3 | 202017053124-PRIORITY DOCUMENTS [07-12-2020(online)].pdf | 2020-12-07 |
| 4 | 202017053124-POWER OF AUTHORITY [07-12-2020(online)].pdf | 2020-12-07 |
| 5 | 202017053124-FORM 1 [07-12-2020(online)].pdf | 2020-12-07 |
| 6 | 202017053124-DRAWINGS [07-12-2020(online)].pdf | 2020-12-07 |
| 7 | 202017053124-DECLARATION OF INVENTORSHIP (FORM 5) [07-12-2020(online)].pdf | 2020-12-07 |
| 8 | 202017053124-COMPLETE SPECIFICATION [07-12-2020(online)].pdf | 2020-12-07 |
| 9 | 202017053124-Verified English translation [22-12-2020(online)].pdf | 2020-12-22 |
| 10 | 202017053124-Verified English translation [22-12-2020(online)]-1.pdf | 2020-12-22 |
| 11 | 202017053124-FORM-26 [11-02-2021(online)].pdf | 2021-02-11 |
| 12 | 202017053124-Proof of Right [19-02-2021(online)].pdf | 2021-02-19 |
| 13 | 202017053124-Proof of Right [03-03-2021(online)].pdf | 2021-03-03 |
| 14 | 202017053124-FORM 3 [06-04-2021(online)].pdf | 2021-04-06 |
| 15 | 202017053124-Proof of Right [30-04-2021(online)].pdf | 2021-04-30 |
| 16 | 202017053124.pdf | 2021-10-19 |
| 17 | 202017053124-FORM 18 [01-07-2022(online)].pdf | 2022-07-01 |
| 18 | 202017053124-FER.pdf | 2022-11-03 |
| 19 | 202017053124-FER_SER_REPLY [27-04-2023(online)].pdf | 2023-04-27 |
| 20 | 202017053124-CORRESPONDENCE [27-04-2023(online)].pdf | 2023-04-27 |
| 21 | 202017053124-COMPLETE SPECIFICATION [27-04-2023(online)].pdf | 2023-04-27 |
| 22 | 202017053124-CLAIMS [27-04-2023(online)].pdf | 2023-04-27 |
| 23 | 202017053124-US(14)-HearingNotice-(HearingDate-25-02-2025).pdf | 2025-01-08 |
| 24 | 202017053124-Correspondence to notify the Controller [24-02-2025(online)].pdf | 2025-02-24 |
| 25 | 202017053124-Written submissions and relevant documents [12-03-2025(online)].pdf | 2025-03-12 |
| 26 | 202017053124-FORM 3 [12-03-2025(online)].pdf | 2025-03-12 |
| 27 | 202017053124-PatentCertificate19-03-2025.pdf | 2025-03-19 |
| 28 | 202017053124-IntimationOfGrant19-03-2025.pdf | 2025-03-19 |
| 1 | searchE_02-11-2022.pdf |