Sign In to Follow Application
View All Documents & Correspondence

Video Playback Device, Playback Method, And Program

Abstract: The present technology relates to a video playback device, a playback method, and a program whereby it is possible to suppress a reduction in the degree of immersion of a viewer. This video generation device acquires a first 3D shape video which is generated from a plurality of viewpoint videos obtained by capturing an image of a subject from different viewpoints, and a second 3D shape video different from the first 3D shape video, and switches a frame to be played back, from a frame of the first 3D shape video to a frame of the second 3D shape video, on the basis of the state of the viewer viewing the first 3D shape video. The present disclosure can be applied to, for example, a video generation device, a video processing device, and a video playback device.

Get Free WhatsApp Updates!
Notices, Deadlines & Correspondence

Patent Information

Application #
Filing Date
31 December 2020
Publication Number
25/2021
Publication Type
INA
Invention Field
COMMUNICATION
Status
Email
patents@remfry.com
Parent Application
Patent Number
Legal Status
Grant Date
2025-02-24
Renewal Date

Applicants

SONY CORPORATION
1-7-1, Konan, Minato-ku, Tokyo 1080075

Inventors

1. ARAKI Yuichi
c/o SONY CORPORATION, 1-7-1, Konan, Minato-ku, Tokyo 1080075
2. HASEGAWA Yuichi
c/o SONY CORPORATION, 1-7-1, Konan, Minato-ku, Tokyo 1080075
3. IZUMI Nobuaki
c/o SONY CORPORATION, 1-7-1, Konan, Minato-ku, Tokyo 1080075

Specification

Title of invention: Video playback device, playback method, and program
Technical field
[0001]
 The present technology relates to a video playback device, a playback method, and a program, and particularly to a video playback device, a playback method, and a program of a video including a 3D object.
Background technology
[0002]
 Conventionally, using information that senses the actual 3D space, for example, a multi-view image obtained by capturing a subject from a different viewpoint, a 3D object is generated in the viewing space, as if it exists in the viewing space. A method of generating a visible image (also referred to as Volumetric Video) has been considered (for example, Patent Document 1).
Prior art literature
Patent documents
[0003]
Patent Document 1: International Publication No. 2017/082076
Outline of the invention
Problems to be solved by the invention
[0004]
 Since the image generated by the above method is a 3D-shaped image generated by using the actual image of the subject, the viewer can experience an immersive feeling as if the subject is closer.
[0005]
 However, even if the viewer performs an action such as touching or approaching the 3D object of the subject in the viewing space, the 3D object of the subject shows a predetermined movement regardless of the movement of the viewer. It is only, and it is not possible to show a reaction corresponding to the movement of the viewer, that is, an interactive reaction. Therefore, the degree of immersiveness of the viewer may be impaired.
[0006]
 The present disclosure has been made in view of such a situation, and suppresses a decrease in the degree of immersiveness of the viewer.
Means to solve problems
[0007]
 The image reproduction device of one aspect of the present technology is a first 3D shape image generated from a plurality of viewpoint images in which a subject is captured from different viewpoints, and a second image different from the first 3D shape image. Based on the state of the acquisition unit that acquires the 3D shape image and the state of the viewer who is viewing the first 3D shape image, the frame of the first 3D shape image to the frame of the second 3D shape image. It is a video reproduction device including a switching processing unit for switching a frame to be reproduced.
[0008]
 In the reproduction method of one aspect of the present technology, the image reproduction device uses a first 3D shape image generated from a plurality of viewpoint images in which a subject is captured from different viewpoints and an image different from the first 3D shape image. Based on the state of the viewer who has acquired a certain second 3D shape image and is viewing the first 3D shape image, the frame of the first 3D shape image is used to obtain the second 3D shape image. This is a playback method that switches the frame to be played back to the frame.
A brief description of the drawing
[0009]
FIG. 1 is a diagram illustrating an outline from imaging to viewing of a 3D shape image according to the present disclosure.
FIG. 2 is a diagram illustrating an example of a 3D shape image.
FIG. 3 is a diagram illustrating an outline of a switching operation example of the first embodiment according to the present disclosure.
FIG. 4 is a diagram showing a specific example of FIG.
FIG. 5 is a diagram illustrating a specific example of the switching operation of FIG.
FIG. 6 is a block diagram showing a configuration example of the first embodiment according to the present disclosure.
FIG. 7 is a flowchart illustrating a generation process of the video generation device 1 of FIG.
8 is a flowchart illustrating a reproduction process of the video reproduction apparatus 2 of FIG.
9 is a flowchart illustrating the trigger acquisition process of FIG. 8. FIG.
10 is a flowchart illustrating the frame number selection process of FIG. 8. FIG.
FIG. 11 is a diagram illustrating an outline of a switching operation example of the second embodiment according to the present disclosure.
12 is a diagram illustrating an application example of the switching operation of FIG. 11. FIG.
FIG. 13 is a diagram showing a configuration example of switching metadata according to a second embodiment according to the present disclosure.
FIG. 14 is a block diagram showing a configuration example of a second embodiment according to the present disclosure.
15 is a flowchart illustrating processing of the video generator 101 of FIG. 14. FIG.
16 is a flowchart illustrating the switching metadata generation process of FIG. 15. FIG.
FIG. 17 is a flowchart illustrating processing of the video reproduction device 102 of FIG.
FIG. 18 is a flowchart illustrating the trigger acquisition process of FIG.
FIG. 19 is a flowchart illustrating the frame number selection process of FIG.
FIG. 20 is a diagram illustrating an outline of a switching operation example of the third embodiment according to the present disclosure.
FIG. 21 is a diagram illustrating a 3D object and a viewer's line-of-sight vector.
FIG. 22 is a block diagram showing a configuration example of a third embodiment according to the present disclosure.
FIG. 23 is a flowchart illustrating processing of the video generator 201 of FIG. 22.
FIG. 24 is a flowchart illustrating the line-of-sight information generation process of FIG. 23.
FIG. 25 is a flowchart illustrating the frame number selection process of FIG.
[Fig. 26] Fig. 26 is a block diagram showing a configuration example of computer hardware.
Mode for carrying out the invention
[0010]
 Hereinafter, embodiments for carrying out the present disclosure (hereinafter referred to as embodiments) will be described. The explanation will be given in the following order.
 0. Premises of the present disclosure (Figs. 1 and 2)
 1. First Embodiment (FIGS. 3 to 10)
 2. Second Embodiment (FIGS. 11 to 19)
 3. Third Embodiment (FIGS. 20 to 25)
[0011]
 
 (Outline from imaging to viewing of 3D-shaped image)
 FIG. 1 is a diagram showing an outline of processing from imaging to viewing of a 3D-shaped image.
[0012]
 As shown in FIG. 1, the generation of the 3D shape image includes imaging by a plurality of imaging devices 1-1 to 1-3 and a process of generating a 3D object M01 by 3D modeling.
[0013]
 Specifically, as shown in FIG. 1, the plurality of imaging devices 1-1 to 1-3 are arranged outside the subject # Ob1 so as to surround the subject # Ob1. FIG. 1 shows an example in which the number of image pickup devices is three, and the image pickup devices 1-1 to 1-3 are arranged around the subject # Ob1. In the example of FIG. 1, a person performing a predetermined operation is set as the subject # Ob1.
[0014]
 3D modeling is performed using the video frames of the three image pickup devices 1-1 to 1-3 that are simultaneously imaged, and the subject # is performed for each video frame unit of the three image pickup devices 1-1 to 1-3. A 3D object of Ob1 is created.
[0015]
 The format of the 3D object has, for example, mesh data in which the geometry information of the subject # Ob1 is represented by a connection between vertices (Vertex) called a polygon mesh, and color information corresponding to each polygon mesh. It may be in a format, or it may be in a Point Cloud format, which is a set of points in a 3D space having shape information and attribute information (particularly color information).
[0016]
 Then, a 3D shape image is generated using the 3D object of the subject representing a series of movements. A 3D shape image is composed of a time series of a plurality of frames including a 3D object of a subject taking a series of movements. A background image representing a predetermined space is appropriately synthesized with the 3D object.
[0017]
 The content data including the generated 3D shape video data is transmitted to the playback device and reproduced. The data of the 3D shape image is reproduced, and the 3D object of the subject is rendered, so that the 3D shape image is displayed on the viewer's viewing device. In the example of FIG. 1, a display D1 and an HMD (Head Mount Display) D2 are used as viewing devices.
[0018]
 (Example of 3D Shape Image)
 FIG. 2 is a diagram showing an example of a plurality of generated 3D shape images (hereinafter, also referred to as a scene).
[0019]
 In the image generation device, data of a plurality of scenes targeting 3D objects of the same subject is generated. In the example of FIG. 2, scene A, scene B, and scene C, which are three scenes representing different movements independently of each other, are generated.
[0020]
 In the example of FIG. 2, the scene A is composed of three frames A1 to A3, the scene B is composed of five frames B1 to B5, and the scene C is composed of five frames C1 to C5. To.
[0021]
 The illustration of one person shown in black corresponds to one frame constituting the scene. The same applies to the other figures after FIG. The frames A1 to A3 that make up the scene A, the frames B1 to B5 that make up the scene B, and the frames C1 to C5 that make up the scene C show how the same subject takes different movements.
[0022]
 From the video generation device to the video playback device, scene data using the 3D objects of the subjects, which are generated in this way and represent different movements, is transmitted.
[0023]
 In the video reproduction device, the reproduction of the scene data is promoted by switching the reproduction target at a predetermined timing. By switching the playback target, the scene to be viewed by the viewer is switched, and the movement of the 3D object included in the scene changes according to the state of the viewer.
[0024]
 <
 Outline of Operation of First Embodiment>
 FIG. 3 is a diagram for explaining an outline of a scene switching operation example of the first embodiment according to the present disclosure.
[0025]
 As shown in A of FIG. 3, for example, the content is composed of two scenes, a scene B which is a main scene and a scene C which is a sub scene.
[0026]
 When the trigger information requesting the switching of the video is acquired in the frame of the scene B being played by the video playback device, the video playback device switches the scene to be played back to the scene C.
[0027]
 The video reproduction device reproduces the scene C after switching the scene to be reproduced to the scene C. When the reproduction of the scene C is completed, the video reproduction device performs a process of switching the reproduction target to a predetermined frame of the scene B, and continues the reproduction of the scene B thereafter.
[0028]
 The switching process from the scene B to the scene C will be specifically described.
[0029]
 As shown by the downward arrow of B in FIG. 3, for example, it is assumed that the trigger information requesting the switching to the scene C is acquired in the frame B2 of the scene B. In this case, the reproduction target is switched to the frame C3 of the scene C according to the acquisition of the trigger information.
[0030]
 For example, the scene to be switched to and its frame are predetermined. Which frame of the scene C may be used as the switching destination, or a frame of another scene (not shown) may be used as the switching destination.
[0031]
 The above-mentioned trigger information that defines the switching of the scene to be reproduced is acquired, for example, according to the movement of the viewer. The viewer will watch how the movement of the 3D object changes interactively by switching the scene according to his / her movement.
[0032]
 FIG. 4 is a diagram showing a specific example of how the viewer looks.
[0033]
 In the example of FIG. 4, a state in which the main scene is being reproduced in the video reproduction device is shown. Specifically, it is assumed that the scene in which the 3D object M02 of the subject existing in the viewing space walks and moves downward in the figure is reproduced as the main scene. Here, the black-painted ellipse represents the position of the 3D object M02 at each time displayed by the main scene. White numbers represent the time.
[0034]
 The state shown in FIG. 4 is a state in which reproduction is performed from the frame at time 0 to the frame at time 3, and the 3D object M02 is walking and moving to the position corresponding to the frame at time 3.
[0035]
 The position U01 shown in the lower left of FIG. 4 represents the position of the viewer in the viewing space. In the example of FIG. 4, the viewer is viewing at a position away from the position corresponding to the frame at time 3 in which the 3D object M02 exists.
[0036]
 In this case, the trigger information is not acquired and only the main scene continues to be played. That is, the 3D object M02 walks and moves to the position corresponding to the frame at time 7 by performing the reproduction after the frame at time 3.
[0037]
 FIG. 5 is a diagram showing a specific example of how the scenes look when switching occurs, as described with reference to FIG.
[0038]
 As shown in FIG. 5, playback is performed up to the frame at time 3 of the main scene, and the viewer moves from position U01 to position U02 at the timing when the 3D object M02 of the subject is at the position corresponding to the frame at time 3. If so, the video playback device acquires the trigger information.
[0039]
 The video playback device performs processing for switching the playback target from the time 3 frame of the main scene to the time 10 frame of the sub scene in response to the acquisition of the trigger information.
[0040]
 The sub-scene shown in FIG. 5 is composed of frames at times 10 to 14. The sub-scene is a scene in which the 3D object M03 of the subject moves from the position corresponding to the frame at time 10 to the position corresponding to the frame at time 14 so as to avoid the position U02. The 3D object M02 of the subject included in the main scene and the 3D object M03 of the subject included in the sub scene are 3D objects obtained by photographing the same person.
[0041]
 In this example, the content reproduced by the video reproduction device is composed of two scenes, a main scene shown in FIG. 4 and a sub scene to be switched.
[0042]
 By performing such a switching process, the viewer at the position U02 can view the 3D object M02 of the subject so as to avoid the viewer.
[0043]
 The viewer can get an immersive feeling by interactively changing the movement of the 3D object of the subject in response to the viewer changing the position.
[0044]
 (Structure Example of First Embodiment)
 FIG. 6 is a block diagram showing an example of the configuration of the first embodiment according to the present disclosure.
[0045]
 As shown in FIG. 6, the information processing system as the first embodiment according to the present disclosure includes a video generation device 1, a video playback device 2, and a display device 3. In the first embodiment according to the present disclosure, a scene which is a 3D shape image is generated, the generated scene is reproduced, and the scene is displayed.
[0046]
 (Description of Configuration of
 Video Generation Device ) The video generation device 1 is composed of a multi-viewpoint video acquisition unit 11, a 3D model generation unit 12, and a reproduction information generation unit 13.
[0047]
 The multi-viewpoint image acquisition unit 11 of the image generation device 1 acquires subject information (for example, captured 2D image, distance information, and camera parameters) from a plurality of image pickup devices, and causes the 3D model generation unit 12 to acquire the subject information (for example, the captured 2D image, the distance information, and the camera parameters). Supply.
[0048]
 The 3D model generation unit 12 performs modeling using the subject information supplied from the multi-viewpoint video acquisition unit 11 and generates a 3D object of the subject. The 3D model generation unit 12 generates a scene for displaying the 3D object of the subject. For example, a main scene and a sub scene are generated. The 3D model generation unit 12 supplies the scene data, which is the data of the generated scene, to the reproduction information generation unit 13.
[0049]
 The 3D model generation unit 12 generates content data from the reproduction information supplied from the reproduction information generation unit 13 and the scene data generated by the 3D model generation unit 12, and supplies the content data to the video reproduction device 2. The transmission of the content data to the video playback device 2 is performed via a network such as the Internet or a LAN (Local Area Network), for example. Content data may be transmitted via a recording medium such as an HDD (Hard Disk Drive) or a flash memory.
[0050]
 The content data may include a plurality of scene data, or may include only one scene data. Further, the reproduction information may not be included in the content data and may be supplied separately.
[0051]
 The reproduction information generation unit 13 selects and sets, for example, a main scene and a sub scene based on the scene data supplied from the 3D model generation unit 12. Further, the reproduction information generation unit 13 generates reproduction information that defines the reproduction procedure of each of the main scene and the sub scene, and supplies the reproduction information to the 3D model generation unit 12.
[0052]
 (Description of Configuration of
 Video Reproduction Device ) The video reproduction device 2 is composed of an acquisition unit 21, a trigger acquisition processing unit 22, a viewpoint position acquisition unit 23, a frame number selection unit 24, a switching processing unit 25, and a reproduction unit 26.
[0053]
 The acquisition unit 21 of the video reproduction device 2 acquires the content data transmitted from the video generation device 1. The scene data included in the content data is output to the trigger acquisition processing unit 22. Further, among the scene data included in the content data acquired by the acquisition unit 21, the scene data selected based on the reproduction information is supplied to the switching processing unit 25 as the data of the frame to be reproduced.
[0054]
 For example, the acquisition unit 21 supplies the main scene data to the switching processing unit 25 when the reproduction of the content data is started. Further, in the case where the acquisition unit 21 outputs the main scene data to the switching processing unit 25 and reproduces the main scene, when the information for identifying the reproduction target is supplied from the switching processing unit 25, Stop the output of the main scene data. Then, the acquisition unit 21 supplies the data of the frame of the required range of the sub-scene specified by the information specifying the reproduction target to the switching processing unit 25 as the data of the specific scene.
[0055]
 The reproduction information included in the content data acquired by the acquisition unit 21 is output to the frame number selection unit 24. Information indicating the current playback position is also output from the acquisition unit 21 to the frame number selection unit 24. For example, the frame in which the scene data is output from the acquisition unit 21 to the switching processing unit 25 is the frame at the current playback position.
[0056]
 The trigger acquisition processing unit 22 uses the scene data supplied from the acquisition unit 21 and the viewpoint position information supplied from the viewpoint position acquisition unit 23, which is the viewing position information of the viewer in the viewing space, to display the scene. Trigger information indicating whether or not to request switching is supplied to the frame number selection unit 24.
[0057]
 In addition, the trigger information indicating whether or not to request the switching of the scene is repeatedly output from the trigger acquisition processing unit 22 to the frame number selection unit 24 at a predetermined cycle such as each time the frame is reproduced. To. Based on the scene data and the viewpoint position information, the content of the trigger information (requesting / not requesting the switching of the scene) is switched.
[0058]
 The viewpoint position acquisition unit 23 acquires the viewpoint position information and supplies it to the trigger acquisition processing unit 22 and the reproduction unit 26. The viewpoint position of the viewer is specified based on, for example, information regarding the movement of the viewer transmitted from the display device 3. Information regarding the viewer's operation for moving the viewpoint position and the direction of the line of sight is transmitted from the display device 3 to the video reproduction device 2 as described above.
[0059]
 The frame number selection unit 24 uses the reproduction information supplied from the acquisition unit 21 and the trigger information supplied from the trigger acquisition processing unit 22 to reproduce the scene to be reproduced and the frame number to be reproduced of the scene. Is determined, and information for identifying the reproduction target is supplied to the switching processing unit 25.
[0060]
 The switching processing unit 25 supplies the information for specifying the reproduction target, which is supplied from the frame number selection unit 24, to the acquisition unit 21. The acquisition unit 21 supplies main scene data or sub-scene data in response to the supply of information that identifies the reproduction target. The switching processing unit 25 supplies the scene data supplied from the acquisition unit 21 to the reproduction unit 26.
[0061]
 The reproduction unit 26 renders a 3D object of the subject based on the scene data supplied from the switching processing unit 25. Then, the reproduction unit 26 transmits the rendering result to the display device 3 and displays the 3D object.
[0062]
 The video generation device 1 and the video playback device 2 having the above configurations are each composed of devices such as a PC (Personal Computer), a smartphone, a tablet terminal, and a game machine. On the other hand, the display device 3 which is a viewing device is composed of a display, an HMD, and the like as described above.
[0063]
 The video generation device 1 and the video playback device 2 are composed of one device, and the video playback device 2 and the display device 3 are composed of one device. It may be provided in the device. Further, the configuration of one device shown in FIG. 6 may be divided into a plurality of devices.
[0064]
 Next, the operation of each device having the above configuration will be described.
[0065]
 (Operation of Each Device in the First Embodiment)
 (3D Model Data Generation Process of
 Video Generation Device ) FIG. 7 is a flowchart illustrating a generation process of the video generation device 1.
[0066]
 The process of FIG. 7 is started, for example, when the subject information acquired from the plurality of imaging devices is supplied from the multi-viewpoint image acquisition unit 11 of the image generation device 1 to the 3D model generation unit 12.
[0067]
 In step S1 of FIG. 7, the 3D model generation unit 12 generates a 3D object of the subject, which is a 3D model, using the information of the subject.
[0068]
 In step S2, the 3D model generation unit 12 generates a scene using a 3D object of the subject representing a series of movements. When there are a plurality of series of movements, the scenes corresponding to each series of movements are individually generated here.
[0069]
 In step S3, the reproduction information generation unit 13 selects and sets, for example, a main scene and a sub scene based on the scene data supplied from the 3D model generation unit 12. The reproduction information generation unit 13 generates reproduction information that defines the reproduction procedure of each of the main scene and the sub scene, and supplies the reproduction information to the 3D model generation unit 12.
[0070]
 In step S4, the 3D model generation unit 12 sets the reproduction information supplied from the reproduction information generation unit 13 with respect to the scene data, and generates the content data including the scene data in which the reproduction information is set.
[0071]
 In step S5, the 3D model generation unit 12 transmits the content data generated in step S4 to the video reproduction device 2. After that, the generation process of the video generation device 1 is completed.
[0072]
 (3D model data reproduction processing
 of the video reproduction device ) FIG. 8 is a flowchart illustrating the reproduction processing of the video reproduction device 2.
[0073]
 In step S21 of FIG. 8, the acquisition unit 21 of the video reproduction device 2 acquires the content data supplied from the 3D model generation unit 12 of the video generation device 1.
[0074]
 In step S22, the acquisition unit 21 analyzes the reproduction information included in the acquired content data.
[0075]
 In step S23, the trigger acquisition processing unit 22 performs the trigger acquisition process. In the trigger acquisition process, the trigger information indicating whether or not the reproduction target is requested to be switched to the sub scene is acquired based on the content data supplied from the acquisition unit 21 and the state of the viewer. The details of the trigger acquisition process will be described later with reference to the flowchart of FIG.
[0076]
 In step S24, the frame number selection unit 24 performs the frame number selection process. In the frame number selection process, the scene to be reproduced and the reproduction target in the scene are based on the trigger information supplied from the trigger acquisition processing unit 22 and the reproduction information included in the content data supplied from the acquisition unit 21. The frame number to be reproduced is determined as information for identifying the reproduction target. The details of the frame number selection process will be described later with reference to the flowchart of FIG.
[0077]
 In step S25, the switching processing unit 25 supplies the acquisition unit 21 with the information for identifying the reproduction target, which is supplied from the frame number selection unit 24. The acquisition unit 21 supplies the data of the scene to be reproduced to the switching processing unit 25 according to the information for specifying the reproduction target supplied from the switching processing unit 25.
[0078]
 Data of the frame of the main scene or the frame of the sub scene is supplied from the acquisition unit 21 to the switching processing unit 25.
[0079]
 In step S26, the switching processing unit 25 supplies the scene data supplied from the acquisition unit 21 in step S25 to the reproduction unit 26.
[0080]
 In step S27, the reproduction unit 26 renders the 3D object of the subject based on the scene data supplied from the switching processing unit 25. Based on the rendering result, the display device 3 displays the 3D object.
[0081]
 In step S28, the switching processing unit 25 determines whether or not all the frames of the scene to be reproduced have been processed. If it is determined that there is an unprocessed frame, the process returns to step S23, and the subsequent processing is repeated.
[0082]
 On the other hand, if it is determined in step S28 that all frames of the scene to be reproduced have been processed, the process proceeds to step S29.
[0083]
 In step S29, it is determined whether or not to end the reproduction process. If it is determined that the reproduction process is not completed, the process returns to step S23, and the subsequent processes are repeated.
[0084]
 On the other hand, if it is determined in step S29 that the reproduction process is completed, the reproduction process of the video reproduction device 2 is completed.
[0085]
 (Trigger Acquisition Process)
 Next, an example of the flow of the trigger acquisition process executed in step S23 of FIG. 8 will be described with reference to the flowchart of FIG.
[0086]
 When the trigger acquisition process is started, in step S41, the trigger acquisition processing unit 22 acquires the viewpoint position information from the viewpoint position acquisition unit 23. In the viewpoint position acquisition unit 23, the viewing position in the viewing space of the viewer is repeatedly acquired.
[0087]
 In step S42, the trigger acquisition processing unit 22 acquires the scene data supplied by the acquisition unit 21 to the switching processing unit 25. The trigger acquisition processing unit 22 calculates the distance between the viewer's viewing position and the target object in the viewing space based on the acquired data.
[0088]
 In step S43, the trigger acquisition processing unit 22 determines whether or not the distance between the viewer's viewing position and the target object is equal to or less than a predetermined threshold value.
[0089]
 If it is determined in step S43 that the distance between the viewer's viewing position and the target object is equal to or less than the threshold value, the process proceeds to step S44.
[0090]
 In step S44, the trigger acquisition processing unit 22 generates "triggered" trigger information requesting switching to the sub-scene based on the determination in the process of step S43, and supplies the trigger information to the frame number selection unit 24.
[0091]
 On the other hand, if it is determined in step S43 that the distance between the viewer's position and the target object is larger than the threshold value, the process proceeds to step S45.
[0092]
 In step S45, the trigger acquisition processing unit 22 generates “no trigger” trigger information that does not require switching to the sub-scene based on the determination by the processing in step S43, and supplies it to the frame number selection unit 24.
[0093]
 When the trigger information is supplied by the process of step S44 or step S45, the trigger acquisition process ends, and the process returns to step S23 of FIG.
[0094]
 (Flow of Frame Number Selection Process)
 Next, an example of the flow of the frame number selection process executed in step S24 of FIG. 8 will be described with reference to the flowchart of FIG.
[0095]
 When the frame number selection process is started, in step S61, the frame number selection unit 24 determines the trigger information supplied from the trigger acquisition processing unit 22, the reproduction information supplied from the acquisition unit 21, and the current reproduction position. Get the information to represent.
[0096]
 In step S62, the frame number selection unit 24 determines whether or not the trigger information acquired by the process of step S61 includes the information of "with trigger" requesting switching to the sub scene.
[0097]
 In step S62, if it is determined that the trigger information acquired in step S61 includes the information "with trigger", the process proceeds to step S63.
[0098]
 In step S63, the frame number selection unit 24 determines a predetermined frame number of the sub-scene to be switched from the reproduction information acquired in step S61. The frame number selection unit 24 supplies the determined sub-scene and the frame number of the scene to the switching processing unit 25 as information for identifying the reproduction target. The information supplied to the switching processing unit 25 includes, for example, information representing a scene to be reproduced and information representing a frame number of a frame serving as a reproduction start position in the scene.
[0099]
 On the other hand, in step S62, if it is determined that the trigger information acquired in step S61 does not include the "triggered" information, the process proceeds to step S64.
[0100]
 In step S64, the frame number selection unit 24 acquires the frame number of the frame to be reproduced next to the current reproduction position acquired in step S61. Since the trigger information including the information of "no trigger" is acquired, the frame number of the frame constituting the main scene is acquired here. The frame number selection unit 24 supplies the acquired frame number to the switching processing unit 25 as information for identifying the reproduction target.
[0101]
 When the information for identifying the reproduction target is supplied to the switching processing unit 25 by the process of step S63 or step S64, the frame number selection process ends, and the process returns to step S24 of FIG.
[0102]
 After that, when the information including the frame number of the sub scene to be switched is supplied from the frame number selection unit 24 to the switching processing unit 25, the scene to be played back is switched from the main scene to the sub scene. Then, the sub-scene is reproduced (step S25 in FIG. 8).
[0103]
 By the above processing, the movement of the 3D object of the subject changes interactively according to the positional relationship between the state of the viewer who is the viewer and the 3D object of the subject in the viewing space. This allows the viewer to get an immersive feeling.
[0104]
 
 Next, the second embodiment will be described.
[0105]
 As described above, in the first embodiment, as the viewer's viewing position approaches the 3D object, trigger information requesting switching to the sub-scene is immediately output, and the sub-scene is displayed. When the playback is started, the movement of the 3D object changes.
[0106]
 However, the main scene as the switching source and the sub scene as the switching destination are configured by arranging frames including 3D objects representing different series of operations in chronological order.
[0107]
 Therefore, if such scene switching is performed only at the timing of the trigger information, the 3D object of the subject is instantly switched between the frames of the scenes of different movements, which gives a sense of discomfort to the viewer who is the viewer. , There is a risk that the viewer's immersive feeling will be impaired.
[0108]
 For example, if the movement of the 3D object in the switching source frame and the movement of the 3D object in the switching destination frame are completely different, the viewer will feel a sense of discomfort.
[0109]
 Therefore, in the second embodiment, metadata for switching scenes is generated on the video generation device side and provided to the video playback device side. In the video generator, even if the scene is switched, the frame that is considered to have no discomfort in the movement of the 3D object is selected as the switching source and switching destination frame, and the metadata representing the selected frame is generated. ..
[0110]
 As a result, it is possible to seamlessly switch the frames of different motion scenes so that the viewer's immersive feeling is not impaired.
[0111]
 (Outline of Operation of Second Embodiment)
 FIG. 11 is a diagram for explaining an outline of an example of a scene switching operation of the second embodiment according to the present disclosure.
[0112]
 As shown in A of FIG. 11, for example, the content is composed of two scenes, a scene A which is a main scene and a scene B which is a sub scene.
[0113]
 When the trigger information requesting the switching of the video is acquired in a certain frame of the scene A being played by the video playback device, the video playback device can switch the scene to be played back to the scene B. Make a judgment. If it is determined that the frame cannot be switched, the video playback device continues the playback of the scene A until the frame in which the scene can be switched, and then switches the scene to be played back to the scene B.
[0114]
 Whether or not the scene to be played can be switched to another scene is represented by the metadata associated (linked) with each frame. When the trigger information requesting the switching of the video is acquired, the video playback device metadata whether or not the frame of the scene A to be played back is a frame that can be switched to the frame of another scene. Will be determined based on.
[0115]
 The video reproduction device reproduces the scene B after switching the scene to be reproduced to the scene B. When the reproduction of the scene B is completed, the video reproduction device switches the reproduction target to a predetermined frame of the scene A, and continues the reproduction of the scene A thereafter.
[0116]
 The switching process from the scene A to the scene B will be specifically described.
[0117]
 As shown by the downward arrow B in FIG. 11, for example, it is assumed that the trigger information requesting the switching of the image is acquired in the frame A2 of the scene A. In this case, according to the acquisition of the trigger information, the video reproduction device determines whether or not the scene to be reproduced can be switched to another scene having a different movement based on the metadata.
[0118]
 For example, when it is determined that the frame A2 is a frame that cannot be switched to another scene, the video playback device continues the playback of the scene A up to the frame A3 in which the scene can be switched, and then reproduces the scene. , Switch to frame B1 of scene B.
[0119]
 In this way, the determination of whether or not to switch the scene is performed based on the metadata for switching. The frame B1 to be switched is also specified based on the metadata.
[0120]
 FIG. 12 is a diagram illustrating an application example of the switching operation of FIG.
[0121]
 In A of FIG. 12, for example, the content is composed of three scenes, a scene A which is a main scene, a scene B which is a sub scene, and a scene C.
[0122]
 When the trigger information requesting the switching of the video is acquired in a certain frame of the scene A being played by the video playback device, the video playback device can switch the scene to be played back to another scene having a different movement. Judge whether or not. If it is determined that the frame cannot be switched, the video playback device continues playing the scene A until the frame can be switched.
[0123]
 At the same time, the video playback device acquires the viewing experience number N, which is the number of times the scene is switched.
[0124]
 The video reproduction device continues the reproduction of the scene A up to the frame A3, which is a frame in which the scene can be switched, and then switches the scene to be reproduced according to the number of viewing experiences. In the example of A in FIG. 12, when the number of viewing experiences N = 1, the scene is switched to the scene B, and when the number of viewing experiences N = 2, the scene is switched to the scene C.
[0125]
 In this example, the scene (frame) to be switched is switched according to the number of viewing experiences. When the reproduction of the scene A is performed up to the last frame of the scene A, such as via another scene, the process returns to the first frame of the scene A, and the reproduction of the scene A is repeated.
[0126]
 The scene to be switched may not be switched according to the number of viewing experiences, but may be switched according to other criteria such as viewing time and date.
[0127]
 The switching process of A in FIG. 12 will be described as a specific example.
[0128]
 As shown by the downward arrow B in FIG. 12, for example, it is assumed that the trigger information requesting the switching of the image is acquired in the frame A2 of the scene A. In this case, according to the acquisition of the trigger information, the video reproduction device determines whether or not the scene to be reproduced can be switched to another scene having a different movement.
[0129]
 For example, when it is determined that the frame A2 is a frame that cannot be switched to another scene, the video reproduction device continues the reproduction of the scene A until the frame A3 that can switch the scene.
[0130]
 At the same time, the video playback device acquires the number of viewing experiences.
[0131]
 After playing back up to the frame A3 in which the scene can be switched, the video playback device switches the scene to be played back to the frame B1 of the scene B when the number of viewing experiences N = 1, and the number of viewing experiences N = 2. In that case, the scene to be reproduced is switched to the frame C1 of the scene C.
[0132]
 With reference to FIG. 13, a configuration example of metadata for switching (hereinafter, also referred to as switching metadata) will be described.
[0133]
 In the example of FIG. 13, the content is composed of three scenes, scene A, scene B, and scene C described with reference to FIG. The frame configuration of each scene is the same as the description in FIG.
[0134]
 Each piece of information contained in the switching metadata 41 of FIG. 13 is described in a corresponding form for each frame of each scene.
[0135]
 As shown in FIG. 13, the switching metadata 41 is composed of a scene ID, an in-scene frame ID, a global frame ID, a transitionable flag, and a transitionable frame ID (Nth time). Here, the variable N of the transitionable frame ID (Nth time) means the number of viewing experiences.
[0136]
 The scene ID of the switching metadata 41 is identification information indicating a scene to which each frame belongs. Here, scene ID = 0 belongs to scene A, scene ID = 1 belongs to scene B, and scene ID = 2 belongs to scene C.
[0137]
 The in-scene frame ID is identification information in each scene of the target frame. In the example of FIG. 13, the identification information corresponding to the frame number is set as the in-scene frame ID. For example, a frame having a scene ID = 0 and an in-scene frame ID = 0 indicates that the frame belongs to the scene A and corresponds to the frame A1 in the scene A. The same applies to other in-scene frame IDs.
[0138]
 The global frame ID is identification information in scenes A, B, and C, which are all scenes constituting the content of the target frame. Here, values ​​0 to 12 are assigned as global frame IDs to frames A1 to A3 of the scene A, frames B1 to B5 of the scene B, and frames C1 to C5 of the scene C, respectively.
[0139]
 For example, global frame ID = 0 indicates that it is frame A1 of scene A, global frame ID = 3 indicates that it is frame B1 of scene B, and global frame ID = 8 indicates that it is frame C1 of scene C. Indicates that. The same applies to other global frame IDs.
[0140]
 The transitionable flag is flag information indicating whether or not the target frame is a frame that can be transitioned (switchable) with respect to a frame of another scene. The transitionable flag corresponds to the metadata used in the above-mentioned determination as to whether or not the scene to be reproduced can be switched to another scene.
[0141]
 For example, the transitionable flag of the frame A2 of the scene A corresponding to the global frame ID = 1 is set to False. In this case, it is determined that the transition from the frame A2 of the scene A to the frame of another scene is impossible.
[0142]
 Further, the transitionable flag of the frame A3 of the scene A corresponding to the global frame ID = 2 is set to True. In this case, it is determined that it is possible to transition from the frame A3 of the scene A to the frame of another scene.
[0143]
 The value of the transitionable flag (True / False) is determined based on the similarity of the shapes of the 3D objects of the subject. The degree of similarity in shape is calculated by comparing the shapes of the 3D objects of the subject in each frame of each scene in the content with each other.
[0144]
 For example, when the similarity is equal to or greater than the threshold value, the 3D object of the subject of the comparison source frame and the 3D object of the subject of the comparison destination frame are 3D objects having similar shapes. In this case, the transitionable flag in the comparison source frame is set to True.
[0145]
 On the other hand, when the similarity is smaller than the threshold value, the 3D object of the subject of the comparison source frame and the 3D object of the subject of the comparison destination frame are 3D objects having dissimilar shapes. In this case, the transitionable flag in the frame of the comparison source is set to the value of False indicating that the transition is not possible with respect to the frame of another scene.
[0146]
 There are various methods for calculating the similarity of shapes. For example, there is a method of calculating the Hausdorff distance of the vertex coordinates between two meshes.
[0147]
 The transitionable frame ID of the switching metadata 41 is information indicating the global frame ID of the frame to be the transition destination in the frame in which the transitionable flag is True. For example, when the transitionable frame ID of the frame A3 of the scene A corresponding to the global frame ID = 2 in which the transitionable flag is True is the transitionable frame ID = 3, it corresponds to the global frame ID = 2. It is shown that the transition from the frame A3 of the scene A to the frame B1 of the scene B corresponding to the global frame ID = 3 is possible.
[0148]
 The transitionable frame ID indicating the transition destination may be in a format in which a different transition destination is specified depending on the number of viewing experiences N.
[0149]
 For example, the transitionable frame ID (first time) represents the pattern of the transitionable frame when the viewer views the sub-scene B for the first time, that is, when the number of viewing experiences N = 1. Similarly, the transitionable frame ID (second time) represents the pattern of the transitionable frame when the sub-scene is viewed for the second time, that is, when the number of viewing experiences N = 2.
[0150]
 In FIG. 13, the transitionable frame ID (second time) of the frame A3 of the scene A corresponding to the global frame ID = 2 is different from the transitionable frame ID (first time) = 3 and the transitionable frame ID (second time). = 8 is set. This is done so that the switching from the frame A3 of the scene A corresponding to the global frame ID = 2 to another frame is switched to the frame B1 of the scene B when the sub scene is viewed for the first time. When the scene is viewed for the second time, it means that the scene is switched to the frame C1 of the scene C, which is a frame of the scene different from the first viewing.
[0151]
 (Structure Example of the Second Embodiment)
 FIG. 14 is a block diagram showing an example of the configuration of the second embodiment according to the present disclosure.
[0152]
 Of the configurations shown in FIG. 14, the same configurations as those in FIG. 6 are designated by the same reference numerals. Duplicate explanations will be omitted as appropriate. The same applies to other block diagrams described later.
[0153]
 (Explanation of Configuration of Video Generation Device) The configuration
 of the video generation device 101 of FIG. 14 is that a 3D model generation unit 112 is provided instead of the 3D model generation unit 12, and a switching metadata generation unit 111 is newly provided. Is different from the configuration of the image generator 1 of FIG.
[0154]
 The switching metadata generation unit 111 of the video generation device 101 of FIG. 14 generates the switching metadata shown in FIG. 13 based on the scene data supplied from the 3D model generation unit 12 and the reproduction information, and is a 3D model generation unit. Supply to 112. For example, the reproduction information supplied from the 3D model generation unit 112 is used to determine the value to be set in the transitionable flag shown in FIG.
[0155]
 The 3D model generation unit 112 generates content data from the switching metadata supplied from the switching metadata generation unit 111, the reproduction information supplied from the reproduction information generation unit 13, and the scene data generated by the 3D model generation unit 112. , Is transmitted to the video reproduction device 102.
[0156]
 The switching metadata may not be included in the content data and may be supplied separately.
[0157]
 (Explanation of Configuration of Video Playback Device) The configuration
 of the video playback device 102 of FIG. 14 is different from the configuration of the video playback device 2 of FIG. 6 in that a frame number selection unit 121 is provided instead of the frame number selection unit 24. ..
[0158]
 The trigger acquisition processing unit 22 uses the scene data supplied from the acquisition unit 21 and the viewpoint position information supplied from the viewpoint position acquisition unit 23 to generate trigger information indicating whether or not to request scene switching, with a frame number. It is supplied to the selection unit 121.
[0159]
 In addition, the trigger information indicating whether or not to request the switching of the scene is repeatedly output from the trigger acquisition processing unit 22 to the frame number selection unit 121 at a predetermined cycle such as each time the frame is reproduced. To. Based on the scene data and the viewpoint position information, the content of the trigger information (requesting / not requesting the switching of the scene) is switched.
[0160]
 After supplying the trigger information requesting the scene switching to the frame number selection unit 121, the trigger acquisition processing unit 22 supplies the trigger information according to the setting from the frame number selection unit 121 to the information that does not request the scene switching. Set to.
[0161]
 The frame number selection unit 121 uses the information representing the current playback position supplied from the acquisition unit 21, the switching metadata, and the trigger information supplied from the trigger acquisition processing unit 22, to use the scene to be reproduced and the scene to be reproduced. , The frame number to be reproduced in the scene is determined, and the information for identifying the reproduction target is supplied to the switching processing unit 25.
[0162]
 After supplying the information for identifying the reproduction target to the switching processing unit 25, the frame number selecting unit 121 instructs the trigger acquisition processing unit 22 to set the trigger information to the information that does not require the switching of the scene. Provide information to do.
[0163]
 Next, the operation of each device having the above configuration will be described.
[0164]
 (Operation of Each Device in the Second Embodiment)
 (3D Model Data Generation Process of
 Video Generation Device ) FIG. 15 is a flowchart illustrating a generation process of the video generation device 101.
[0165]
 The process shown in FIG. 15 is basically the same process as the process described with reference to FIG. 7, except that a process for generating switching metadata is added. Duplicate explanations will be omitted as appropriate.
[0166]
 In step S81 of FIG. 15, the 3D model generation unit 112 generates a 3D object of the subject, which is a 3D model, using the information of the subject supplied from the multi-viewpoint image acquisition unit 11.
[0167]
 In step S82, the 3D model generation unit 112 generates a scene using a 3D object of the subject representing a series of movements. When there are a plurality of series of movements, the scenes corresponding to each series of movements are individually generated here.
[0168]
 In step S83, the reproduction information generation unit 13 selects and sets, for example, a main scene and a sub scene based on the scene data supplied from the 3D model generation unit 112. The reproduction information generation unit 13 generates reproduction information that defines the reproduction procedure of each of the main scene and the sub scene, and supplies the reproduction information to the 3D model generation unit 112.
[0169]
 In step S84, the 3D model generation unit 112 sets the reproduction information supplied from the reproduction information generation unit 13 with respect to the scene data, and generates the content data including the scene data in which the reproduction information is set.
[0170]
 In step S85, the switching metadata generation unit 111 performs the switching metadata generation process. The switching metadata generation process is appropriately performed in parallel with the processes of steps S83 and S84.
[0171]
 In the switching metadata generation process, switching metadata is generated based on the scene data supplied from the 3D model generation unit 112 and the reproduction information, and is supplied to the 3D model generation unit 112. Details of the switching metadata generation process will be described later with reference to the flowchart of FIG.
[0172]
 In step S86, the 3D model generation unit 112 stores the switching metadata generated in step S85 in the content data generated in step S84, and supplies the switching metadata to the video playback device 102. After that, the generation process of the video generation device 101 is completed.
[0173]
 The switching metadata may be supplied to the video reproduction device 102 separately from the content data.
[0174]
 (Process of Switching Metadata Generation)
 Next, an example of the flow of the switching metadata generation process executed in step S85 of FIG. 15 will be described with reference to the flowchart of FIG.
[0175]
 When the switching metadata generation process is started, in step S100, the switching metadata generation unit 111 identifies each frame of each scene based on the reproduction information and scene data supplied from the 3D model generation unit 112. Information scene IDs, scene frame IDs, and global frame IDs are assigned to each.
[0176]
 In step S101, the switching metadata generation unit 111 sets 0 to the variable s indicating each scene ID.
[0177]
 In step S102, the switching metadata generation unit 111 sets 0 to the variable i indicating the frame number of each scene.
[0178]
 In step S103, the switching metadata generation unit 111 sets the transitionable flag of the i-th frame of the s-th scene to False based on the reproduction information and the scene data supplied from the 3D model generation unit 112.
[0179]
 Further, the switching metadata generation unit 111 sets the i-th frame of the s-th scene as a comparison source target frame as a comparison source for calculating the similarity, and sets a 3D object (3D object) of the subject of the comparison source target frame. (Hereinafter, it is represented by adding OSI)) is acquired.
[0180]
 In step S104, the switching metadata generation unit 111 sets the variable t indicating the target scene to be compared for calculating the similarity to 0 based on the reproduction information and the scene data supplied from the 3D model generation unit 112. Set to.
[0181]
 Further, the switching metadata generation unit 111 sets the variable N indicating the number of viewing experiences of the viewer to 0.
[0182]
 In step S105, the switching metadata generation unit 111 sets the variable j indicating the target frame to be compared for calculating the similarity to 0 based on the reproduction information and the scene data supplied from the 3D model generation unit 112. Set to.
[0183]
 In step S106, the switching metadata generation unit 111 compares the j-th frame of the t-th scene with the comparison destination for calculating the similarity based on the reproduction information and the scene data supplied from the 3D model generation unit 112. It is set as the comparison destination target frame, and the data of the 3D object (hereinafter, represented by adding Otj) of the subject of the comparison destination target frame is acquired.
[0184]
 In step S107, the switching metadata generation unit 111 determines whether or not the degree of shape similarity between the 3D object OSI acquired in step S103 and the 3D object Otj acquired in step S106 is equal to or greater than the threshold value.
[0185]
 If it is determined in step S107 that the shape similarity between the 3D object OSI and the 3D object Otj is equal to or greater than the threshold value, the process proceeds to step S108.
[0186]
 In step S108, the switching metadata generation unit 111 sets the transitionable flag to True for the i-th frame of the s-th scene. Further, the switching metadata generation unit 111 sets a global frame ID indicating that it is the j-th frame of the t-th scene as the transitionable frame ID (Nth time), and is a variable representing the number of viewing experiences of the viewer. Set N + 1 to N.
[0187]
 That is, the transitionable flag of the frame to be compared for calculating the similarity is set to True, and the global frame ID of the frame to be compared is set as the transitionable frame ID (Nth time).
[0188]
 Note that the sub-scene to be switched to does not correspond to the number of viewing experiences, that is, when the sub-scene is fixed, 1 remains set in the variable N representing the number of viewing experiences. ..
[0189]
 If it is determined in step S107 that the shape similarity between the 3D object OSI and the 3D object Otj is smaller than the threshold value, the process of step S108 is skipped.
[0190]
 In step S109, the switching metadata generation unit 111 determines whether or not all the frames of the t-th scene have been processed. If it is determined in step S109 that all frames have not been processed, that is, there are unprocessed frames, the process proceeds to step S110.
[0191]
 In step S110, the switching metadata generation unit 111 sets j + 1 for the variable j. After that, the process returns to step S106, the comparison target frame is changed, and the subsequent processing is repeated.
[0192]
 On the other hand, if it is determined in step S109 that all the frames of the t-th scene have been processed, the process proceeds to step S111.
[0193]
 In step S111, the switching metadata generation unit 111 determines whether or not all the scenes other than the sth scene have been processed. If it is determined in step S111 that all the scenes have not been processed, that is, there are unprocessed scenes, the process proceeds to step S112.
[0194]
 In step S112, the switching metadata generation unit 111 sets t + 1 for the variable t. After that, the process returns to step S105, the target scene is changed, and the subsequent processing is repeated.
[0195]
 On the other hand, if it is determined in step S111 that all the scenes except the sth scene have been processed, the process proceeds to step S113.
[0196]
 In step S113, the switching metadata generation unit 111 determines whether or not all the frames of the sth scene have been processed. If it is determined that all the frames have not been processed, that is, there are unprocessed frames, the process proceeds to step S114.
[0197]
 In step S114, the switching metadata generation unit 111 sets i + 1 for the variable i. After that, the process returns to step S103, the comparison source target frame is changed, and the subsequent processing is repeated.
[0198]
 On the other hand, if it is determined in step S113 that all frames of the sth scene have been processed, the process proceeds to step S115.
[0199]
 In step S115, the switching metadata generation unit 111 determines whether or not all the scenes have been processed. If it is determined in step S115 that all the scenes have not been processed, that is, there are unprocessed scenes, the process proceeds to step S116.
[0200]
 In step S116, the switching metadata generation unit 111 sets s + 1 for the variable s. After that, the process returns to step S102, the target scene is changed, and the subsequent processing is repeated.
[0201]
 On the other hand, if it is determined in step S115 that all the scenes have been processed, the process proceeds to step S117.
[0202]
 In step S117, the switching metadata generation unit 111 supplies the switching metadata generated as described above to the 3D model generation unit 112. When the switching metadata is supplied, the switching metadata generation process ends, and the process returns to step S85 of FIG.
[0203]
 (3D model data reproduction processing
 of the video reproduction device ) FIG. 17 is a flowchart illustrating the reproduction processing of the video reproduction device 102.
[0204]
 The process shown in FIG. 17 is basically the same process as the process described with reference to FIG.
[0205]
 In step S121 of FIG. 17, the acquisition unit 21 of the video reproduction device 102 acquires the content data supplied from the 3D model generation unit 112 of the video generation device 101.
[0206]
 In step S122, the acquisition unit 21 analyzes the reproduction information included in the acquired content data.
[0207]
 In step S123, the trigger acquisition processing unit 22 performs the trigger acquisition process. In the trigger acquisition process, the trigger information indicating whether or not the reproduction target is requested to be switched to the sub scene is acquired based on the content data supplied from the acquisition unit 21 and the state of the viewer. The details of the trigger acquisition process will be described later with reference to the flowchart of FIG.
[0208]
 In step S124, the frame number selection unit 121 performs the frame number selection process. In the frame number selection process, the scene to be reproduced and the scene to be reproduced based on the trigger information supplied from the trigger acquisition processing unit 22, the reproduction information included in the content data supplied from the acquisition unit 21, and the switching metadata. The frame number to be played back in the scene is determined. The details of the frame number selection process will be described later with reference to the flowchart of FIG.
[0209]
 In step S125, the switching processing unit 25 supplies the acquisition unit 21 with the information for identifying the reproduction target, which is supplied from the frame number selection unit 121. The acquisition unit 21 supplies the data of the scene to be reproduced to the switching processing unit 25 according to the information for specifying the reproduction target supplied from the switching processing unit 25.
[0210]
 In step S126, the switching processing unit 25 supplies the scene data supplied from the acquisition unit 21 in step S125 to the reproduction unit 26.
[0211]
 In step S127, the reproduction unit 26 renders the 3D object of the subject based on the scene data supplied from the switching processing unit 25. Based on the rendering result, the display device 3 displays the 3D object.
[0212]
 In step S128, the switching processing unit 25 determines whether or not all the frames of the scene to be reproduced have been processed. If it is determined that there is an unprocessed frame, the process returns to step S123, and the subsequent processing is repeated.
[0213]
 On the other hand, if it is determined in step S128 that all frames of the scene to be reproduced have been processed, the processing proceeds to step S129.
[0214]
 In step S129, it is determined whether or not to end the reproduction process. If it is determined that the reproduction process is not completed, the process returns to step S123, and the subsequent processes are repeated.
[0215]
 On the other hand, if it is determined in step S129 that the reproduction process is to be completed, the reproduction process of the video reproduction device 102 is completed.
[0216]
 (Trigger Acquisition Process)
 Next, an example of the flow of the trigger acquisition process executed in step S123 of FIG. 17 will be described with reference to the flowchart of FIG.
[0217]
 As described above, after the trigger information requesting the switching of the video is acquired, it is the same until the frame set as being able to switch to another frame (the frame in which the transitionable flag is set to True). The scene continues to play. In the process shown in FIG. 18, a process of repeating the output of the trigger information requesting the switching of the video is added until the frame set as being switchable to another frame becomes the playback target. Except for the above, the process is the same as the process shown in FIG.
[0218]
 When the trigger acquisition process is started, in step S141, the trigger acquisition processing unit 22 acquires the viewing position information from the viewpoint position acquisition unit 23.
[0219]
 In step S142, the trigger acquisition processing unit 22 determines whether or not the trigger information supplied immediately before is already in a state with a trigger indicating a request to switch to the sub scene.
[0220]
 If it is determined in step S142 that the trigger is not present, the process proceeds to step S143.
[0221]
 In step S143, the trigger acquisition processing unit 22 acquires the scene data supplied by the acquisition unit 21 to the switching processing unit 25. The trigger acquisition processing unit 22 calculates the distance between the position of the viewer and the target object in the viewing space based on the acquired data.
[0222]
 In step S144, the trigger acquisition processing unit 22 determines whether or not the position of the viewer and the distance between the target objects are equal to or less than a predetermined threshold value.
[0223]
 If it is determined in step S144 that the distance between the viewer's viewing position and the target object is equal to or less than the threshold value, the process proceeds to step S145.
[0224]
 In step S145, the trigger acquisition processing unit 22 generates "triggered" trigger information requesting switching to the sub-scene based on the determination result of the processing in step S144, and supplies the trigger information to the frame number selection unit 121. ..
[0225]
 On the other hand, if it is determined in step S144 that the distance between the viewer's position and the target object is larger than the threshold value, the process proceeds to step S146.
[0226]
 In step S146, the trigger acquisition processing unit 22 generates “no trigger” trigger information that does not require switching to the sub-scene based on the determination result of the processing in step S144, and supplies it to the frame number selection unit 121. ..
[0227]
 When the trigger information is supplied by the process of step S145 or step S146, the trigger acquisition process ends and the process returns to step S123 of FIG. Similarly, when it is determined in step S142 that the trigger is present, the trigger acquisition process ends.
[0228]
 (Frame Number Selection Process)
 Next, an example of the flow of the frame number selection process executed in step S124 of FIG. 17 will be described with reference to the flowchart of FIG.
[0229]
 The process shown in FIG. 19 is different from the process described with reference to FIG. 10 in that the scene switching is performed based on the switching metadata instead of the reproduction information.
[0230]
 When the frame number selection process is started, in step S161, the frame number selection unit 121 includes trigger information supplied from the trigger acquisition processing unit 22, information representing the current playback position supplied from the acquisition unit 21, and switching meta. Get data and the number of viewing experiences.
[0231]
 The frame number selection unit 121 acquires the frame number of the frame to be reproduced by using the information representing the current reproduction position and the switching metadata.
[0232]
 In step S162, the frame number selection unit 121 determines whether or not the trigger information acquired by the process of step S160 includes the information of "with trigger" requesting switching to the sub scene.
[0233]
 If it is determined in step S162 that the trigger information acquired in step S161 includes the information "with trigger", the process proceeds to step S163.
[0234]
 In step S163, the frame number selection unit 121 determines whether or not the transitionable flag in the switching metadata corresponding to the frame number acquired in step S160 is True, based on the switching metadata.
[0235]
 If it is determined in step S163 that the transitionable flag is True, the process proceeds to step S164.
[0236]
 In step S164, the frame number selection unit 121 selects the transitionable frame ID of the corresponding switching metadata based on the number of viewing experiences acquired in step S161. The frame number selection unit 121 acquires a global frame ID, which is information for identifying the reproduction target, from the selected transitionable frame ID, and supplies the global frame ID to the switching processing unit 25.
[0237]
 In step S165, the frame number selection unit 121 supplies the trigger acquisition processing unit 22 with information instructing the setting of the trigger information that is “with trigger” to be changed to “without trigger”. The trigger acquisition processing unit 22 switches the content of the trigger information output to the frame number selection unit 121 based on the information supplied here.
[0238]
 On the other hand, if it is determined in step S162 that the information "with trigger" is not included, or if it is determined in step S163 that the transitionable flag is not True and False, the process proceeds to step S166.
[0239]
 In step S166, the frame number selection unit 121 supplies the frame number to be reproduced, which is acquired in step S161, to the switching processing unit 25 as information for identifying the reproduction target.
[0240]
 When the information for identifying the reproduction target is supplied to the switching processing unit 25 by the process of step S165 or step S166, the frame number selection process ends and the process returns to step S124 of FIG.
[0241]
 By the above processing, the playback target is switched to the frame containing the 3D object of the similar operation, so that the frame of the scene of the different operation can be switched seamlessly as compared with the first embodiment, and the viewer is immersed. It is possible to prevent the feeling from being impaired.
[0242]
 
 Next, the third embodiment will be described.
[0243]
 In the first embodiment, the reaction of the 3D object of the subject can be changed according to the state of the viewer based on the trigger information and the reproduction information.
[0244]
 However, in each scene representing a different series of actions, the reaction of the 3D object may become an undesired reaction of the viewer. For example, if the 3D object takes a reaction that points in a direction different from the viewer's direction, the 3D object will divert the line of sight from the viewer, and such a reaction can be said to be an undesired reaction of the viewer. .. This may give the viewer a sense of discomfort and impair the viewer's immersive feeling.
[0245]
 Therefore, in the third embodiment, the line-of-sight information indicating the line-of-sight direction of the 3D object is provided to the video reproduction device side. In the video playback device, when switching scenes, a scene containing a 3D object that reacts as the viewer thinks is desired by using the line-of-sight information provided by the video generator is selected as the scene to be switched to, and the scene is selected. Is switched.
[0246]
 As a result, it is possible to provide a 3D object of the motion of directing the line of sight to the viewer while seamlessly switching the scenes of different motions according to the state of the viewer, so that the immersive feeling of the viewer is not impaired.
[0247]
 (Outline of Operation of Third Embodiment)
 FIG. 20 is a diagram for explaining an outline of an example of a scene switching operation of the third embodiment according to the present disclosure.
[0248]
 As shown in A of FIG. 20, for example, the content is composed of three scenes, a scene A which is a main scene, and a scene B and a scene C which are sub scenes.
[0249]
 When the trigger information requesting the switching of the video is acquired in a certain frame of the scene A being played by the video playback device, the video playback device selects an appropriate scene from other scenes as the switching destination scene. decide.
[0250]
 In the example of A in FIG. 20, the video playback device determines which scene of scene B or scene C is appropriate for the viewer with respect to the scene to be switched based on the line-of-sight information and the viewpoint position information. decide.
[0251]
 After that, the video reproduction device switches the scene to be reproduced to the scene determined based on the line-of-sight information and the viewpoint position information in the scene B or the scene C.
[0252]
 After switching the scene to be reproduced, the video reproduction device reproduces the scene to be switched. When the playback of the switching destination scene is completed, the video playback device switches to a predetermined scene frame and continues playback.
[0253]
 For example, when the scene to be reproduced is switched to the scene B, the video reproduction device reproduces the scene B. When the playback of the scene B is completed, the video playback device switches to the frame of the scene A, which is a predetermined scene, and continues the playback of the scene A thereafter.
[0254]
 The scene switching process will be specifically described.
[0255]
 As shown by the downward arrow B in FIG. 20, for example, it is assumed that the trigger information requesting the switching of the image is acquired in the frame A3 of the scene A. In this case, according to the acquisition of the trigger information, the video reproduction device determines an appropriate scene to be reproduced from the scene B and the scene C.
[0256]
 Specifically, the video playback device sets the switching destination scene and the scene B as the frame of the scene based on the line-of-sight information associated with each frame of each scene and the viewpoint position information of the viewer. Alternatively, it is determined which frame in scene C is appropriate.
[0257]
 For example, when it is determined that the frame B1 of the scene B is appropriate, the video reproduction device switches the scene to be reproduced to the frame B1 of the scene B as shown at the tip of the white arrow. When it is determined that the frame C1 of the scene C is appropriate, the video reproduction device switches the scene to be reproduced to the frame C1 of the scene C as shown at the tip of the white arrow of the broken line.
[0258]
 FIG. 21 is a diagram illustrating an example of how to obtain the line-of-sight direction for each frame of each scene.
[0259]
 As shown in FIG. 21A, the line-of-sight vector of the object representing the line-of-sight direction of the subject's 3D object # Ob21 in a certain frame of a scene is the position (Ox, Oy, Oz) of the subject's 3D object # Ob21 in space. It is calculated from the difference between the position (Gx, Gy, Gz) in space of the 3D object # Ob31 of the gaze object that the 3D object # Ob21 of the subject is gazing at.
[0260]
 As shown in FIG. 21A, when the viewer views the 3D object # Ob21 of the subject from the virtual viewing position V1 (Vx, Vy, Vz) in the viewing space, the viewer represents the line-of-sight direction of the viewer. The line-of-sight vector is calculated from the difference between the virtual viewing position V1 and the position (Ox, Oy, Oz) of the subject's 3D object # Ob21 in space.
[0261]
 When the viewing line-of-sight angle formed by the line-of-sight vector of the object and the line-of-sight vector of the viewer is large, it can be determined that the 3D object of the subject does not direct the line-of-sight toward the viewer in that frame of the scene. The line-of-sight information representing the line-of-sight direction of the 3D object for use in such a determination is provided to the video reproduction device.
[0262]
 From the above, as shown in B of FIG. 21, it is possible to specify the frame of the scene in which the viewing line-of-sight angle becomes small in the video reproduction device. In the video reproduction device, the frame of the scene in which the viewing line-of-sight angle becomes small is selected as the switching destination frame.
[0263]
 (Structure Example of Third Embodiment)
 FIG. 22 is a block diagram showing an example of the configuration of the third embodiment according to the present disclosure.
[0264]
 Of the configurations shown in FIG. 22, the same configurations as those in FIG. 14 are designated by the same reference numerals. Duplicate explanations will be omitted as appropriate.
[0265]
 (Explanation of Configuration of Video Generation Device) In the configuration
 of the video generation device 201 of FIG. 22, the multi-viewpoint video acquisition unit 211 and the 3D model generation unit 212 are provided instead of the multi-viewpoint video acquisition unit 11 and the 3D model generation unit 12. It differs from the image generation device 101 of FIG. 14 in that a point and a line-of-sight information generation unit 213 are newly provided.
[0266]
 The multi-viewpoint image acquisition unit 211 of the image generation device 201 of FIG. 22 acquires subject information (for example, captured 2D image, distance information, and camera parameters) from a plurality of image pickup devices, and generates a 3D model. It is supplied to the unit 212 and the line-of-sight information generation unit 213.
[0267]
 The 3D model generation unit 212 performs modeling using the subject information supplied from the multi-viewpoint image acquisition unit 211 and the line-of-sight information supplied from the line-of-sight information generation unit 213, and generates a 3D object of the subject having the line-of-sight information. To do.
[0268]
 Further, the 3D model generation unit 212 generates a scene including a 3D object of the subject having line-of-sight information for each frame. For example, a main scene and a sub scene are generated. The 3D model generation unit 212 supplies the generated scene data to the reproduction information generation unit 13.
[0269]
 The 3D model generation unit 212 generates content data from the reproduction information supplied from the reproduction information generation unit 13 and the scene data including the line-of-sight information supplied from the line-of-sight information generation unit 213 generated by the 3D model generation unit 212. Then, it is supplied to the video reproduction device 202.
[0270]
 The line-of-sight information generation unit 213 acquires the subject information supplied from the multi-viewpoint image acquisition unit 211. From the multi-viewpoint image acquisition unit 211, for example, a plurality of 2D images obtained by simultaneously capturing the same subject with each imaging device are supplied.
[0271]
 The line-of-sight information generation unit 213 analyzes for each frame whether or not both the subject and the gaze object are detected in the video frame of each 2D video. The line-of-sight information generation unit 213 selects each video frame in which both the subject and the gaze object are detected as the line-of-sight information generation frame.
[0272]
 For example, a line-of-sight information generation frame selected from the 2D image captured by the image pickup device 1-1, a line-of-sight information generation frame selected from the 2D image captured by the image pickup device 1-2, and the like. The line-of-sight information generation frame selected from the 2D video is selected in association with the frame number in each 2D video.
[0273]
 When the line-of-sight information generation unit 213 has two or more frames (two frames) associated with the same frame number, the line-of-sight information generation unit 213 acquires the line-of-sight information generation frame at that frame number.
[0274]
 The line-of-sight information generation unit 213 calculates the position of the face of the subject and the position of the gaze object from the camera parameters supplied from the multi-viewpoint image acquisition unit 211 and each acquired line-of-sight information generation frame.
[0275]
 The line-of-sight information generation unit 213 calculates the position of the subject's face and the position of the gaze object in the 3D space as 3D coordinates by using the calculated position of the subject's face and the position of the gaze object.
[0276]
 As a method of calculating the position in the 3D space, for example, there is a method of projecting each line-of-sight information generation frame into the 3D space and performing triangulation. When there are two or more video frames in which the subject and the gaze object are shown together as frames at the same time, the position of the subject's face and the gaze object can be determined by using those video frames as the line-of-sight information generation frame. It is possible to calculate the position of the above by a method such as triangulation.
[0277]
 The line-of-sight information generation unit 213 calculates the line-of-sight vector of the subject from the calculated position of the face of the subject and the position of the gaze object, and generates a line-of-sight valid flag indicating that the frame has the information of the line-of-sight vector of the subject. ..
[0278]
 The line-of-sight information generation unit 213 supplies the position of the subject's face in the 3D space, the line-of-sight vector information of the subject, and the line-of-sight valid flag as line-of-sight information to the 3D model generation unit 212.
[0279]
 (Explanation of Configuration of Video Playback Device) In the configuration
 of the video playback device 202 of FIG. 22, a frame number selection unit 221 and a viewpoint position acquisition unit 222 are provided instead of the frame number selection unit 121 and the viewpoint position acquisition unit 23. The point is different from the video reproduction device 102 of FIG.
[0280]
 The frame number selection unit 221 of the video reproduction device 202 acquires information representing the current reproduction position, switching metadata, and line-of-sight information supplied from the acquisition unit 21. Further, the frame number selection unit 221 acquires the trigger information supplied from the trigger acquisition processing unit 22 and the viewpoint position information supplied from the viewpoint position acquisition unit 222.
[0281]
 The frame number selection unit 221 becomes a reproduction target by using the information representing the current reproduction position supplied from the acquisition unit 21, the switching metadata, the line-of-sight information, and the viewpoint position information supplied from the viewpoint position acquisition unit 222. The scene and the frame number to be reproduced in the scene are determined, and the information for identifying the reproduction target is supplied to the switching processing unit 25.
[0282]
 After supplying the information for identifying the reproduction target to the switching processing unit 25, the frame number selection unit 221 instructs the trigger acquisition processing unit 22 to set the trigger information to the information that does not require the switching of the scene. Provide information to do.
[0283]
 The viewpoint position acquisition unit 222 acquires the viewpoint position information and supplies it to the trigger acquisition processing unit 22, the frame number selection unit 221 and the reproduction unit 26.
[0284]
 Next, the operation of each device having the above configuration will be described.
[0285]
 (Operation of Each Device in the Third Embodiment)
 (3D Model Data Generation Process of
 Video Generation Device ) FIG. 23 is a flowchart illustrating a generation process of the video generation device 201.
[0286]
 In step S181 of FIG. 23, the multi-viewpoint image acquisition unit 211 of the image generation device 201 acquires subject information from a plurality of imaging devices and supplies the information to the 3D model generation unit 212 and the line-of-sight information generation unit 213.
[0287]
 In step S182, the line-of-sight information generation unit 213 performs the line-of-sight information generation process. The line-of-sight information generation process generates line-of-sight information for each video frame based on the subject information supplied from the multi-viewpoint video acquisition unit 211. The details of the line-of-sight information generation process will be described later with reference to the flowchart of FIG. 24.
[0288]
 In step S183, the 3D model generation unit 212 generates a 3D object of the subject having the line-of-sight information by 3D modeling using the subject information and the line-of-sight information generated in step S182.
[0289]
 In step S184, the 3D model generation unit 212 generates a scene using a 3D object of the subject representing a series of movements. When there are a plurality of series of movements, the scenes corresponding to each series of movements are individually generated here.
[0290]
 In step S185, the reproduction information generation unit 13 selects and sets, for example, a main scene and a sub scene based on the scene data supplied from the 3D model generation unit 212. The reproduction information generation unit 13 generates reproduction information that defines the reproduction procedure of each of the main scene and the sub scene, and supplies the reproduction information to the 3D model generation unit 212.
[0291]
 In step S186, the 3D model generation unit 212 sets the reproduction information supplied from the reproduction information generation unit 13 with respect to the scene data, and generates content data including the data of the scene in which the reproduction information is set.
[0292]
 In step S187, the switching metadata generation unit 111 performs a switching metadata generation process. In the switching metadata generation process, switching metadata is generated based on the scene data supplied from the 3D model generation unit 212 and the reproduction information, and is supplied to the 3D model generation unit 212. Since the process is the same as the process described with reference to FIG. 16, a detailed description of the switching metadata generation process will be omitted.
[0293]
 In step S188, the 3D model generation unit 212 stores the switching metadata generated in step S187 with respect to the content data generated in step S186 and supplies it to the video reproduction device 202. After that, the generation process of the video generation device 201 is completed.
[0294]
 (Gaze Information Generation Process)
 Next, an example of the flow of the line-of-sight information generation process executed in step S182 of FIG. 23 will be described with reference to the flowchart of FIG. 24.
[0295]
 When the line-of-sight information generation process is started, in step S210, the line-of-sight information generation unit 213 receives subject information (for example, captured 2D image, distance information, and camera parameters) supplied from the multi-viewpoint image acquisition unit 211. ) To get. Here, subject information is acquired for each imaging device that captures the same subject.
[0296]
 In step S211 the line-of-sight information generation unit 213 sets the variable i for designating the frame number of each video frame in the 2D video acquired in step S210 to i = 0. The variable i is a variable that is commonly used for all 2D images acquired by each imaging device that captures the same subject.
[0297]
 In step S212, the line-of-sight information generation unit 213 is a video frame in which the subject and the gaze object are both in the video frame from the i-th video frame in the 2D video for each imaging device acquired in step S210. Is selected as the frame for generating line-of-sight information.
[0298]
 The i-th video frame attracting attention here is a video frame in which both the subject and the gaze object are projected, or a video frame in which only the subject is projected. The former video frame in which both the subject and the gaze object are displayed is selected as the line-of-sight information generation frame.
[0299]
 Such selection of the line-of-sight information generation frame is performed for the i-th video frame of each 2D video captured by a plurality of imaging devices.
[0300]
 In step S213, the line-of-sight information generation unit 213 determines whether or not the number of selected line-of-sight information generation frames is two or more. If it is determined that the number of sheets is two or more, the process proceeds to step S214.
[0301]
 In step S214, the line-of-sight information generation unit 213 acquires the camera parameters of each imaging device corresponding to each line-of-sight information generation frame from the subject information acquired in step S210.
[0302]
 In step S215, the line-of-sight information generation unit 213 refers to each line-of-sight information generation frame acquired in step S212, the position of the subject's face (Ou, Ov) in the line-of-sight information generation frame, and the gaze object (Gu, The position of Gv) is detected. Position detection is performed using camera parameters as appropriate.
[0303]
 In step S216, the line-of-sight information generation unit 213 determines whether or not there are two or more lines-of-sight information generation frames capable of detecting the position of the face of the subject and the position of the gaze object in the same frame. If it is determined that there are two or more sheets, the process proceeds to step S217.
[0304]
 In step S217, the line-of-sight information generation unit 213 uses the position of the subject's face (Ou, Ov) and the position of the gaze object (Gu, Gv) of each line-of-sight information generation frame in the 3D space of the subject's face. (Ox, Oy, Oz) and the position (Gx, Gy, Gz) of the gaze object in 3D space are calculated.
[0305]
 In step S218, the line-of-sight information generation unit 213 determines the position (Ox, Oy, Oz) of the subject's face in the 3D space and the position (Gx, Gy, Gz) of the gaze object in the 3D space calculated in step S217. From the difference between, the line-of-sight vector representing the line-of-sight direction of the subject is calculated. Then, the line-of-sight information generation unit 213 generates a line-of-sight valid flag indicating that the frame has a line-of-sight vector, and sets the line-of-sight valid flag to 1.
[0306]
 In step S219, the line-of-sight information generation unit 213 uses the position of the subject's face in 3D space, the line-of-sight vector, and the line-of-sight valid flag generated in step S218 as line-of-sight information together with the information of the related line-of-sight information generation frame. Generate.
[0307]
 In step S220, the line-of-sight information generation unit 213 determines whether or not all the video frames of each 2D video have been processed.
[0308]
 On the other hand, if it is determined in step S213 or in step S216 that the number of frames for generating line-of-sight information is not two or more, the process proceeds to step S221.
[0309]
 In step S221, the line-of-sight information generation unit 213 sets the line-of-sight valid flag to 0.
[0310]
 If it is determined in step S220 that an unprocessed frame exists, the process proceeds to step S222.
[0311]
 In step S222, the line-of-sight information generation unit 213 sets the variable i representing the frame number of the 2D image to i = i + 1. After that, the process returns to step S212, and the subsequent processes are repeated.
[0312]
 On the other hand, if it is determined in step S220 that all the video frames have been processed, the processing proceeds to step S223.
[0313]
 In step S223, the line-of-sight information generation unit 213 supplies all the line-of-sight information to the 3D model generation unit 212, and ends the line-of-sight information generation process. After that, the process returns to step S182 of FIG. 23, and the subsequent processing is performed.
[0314]
 (3D model data reproduction processing of the video reproduction device
 ) Except for the frame number selection process (step S124), the overall processing flow performed by the video reproduction device 202 of FIG. 22 is shown in FIG. 17 in the second embodiment. It is the same as the processing flow described with reference. Duplicate explanations will be omitted.
[0315]
 (Frame Number Selection Process)
 Next, an example of the flow of the frame number selection process executed in step S124 of FIG. 17 will be described with reference to the flowchart of FIG. 25. The process shown in FIG. 25 is a process performed by the video reproduction device 202 of FIG. 22.
[0316]
 When the frame number selection process is started, in step S240, the frame number selection unit 221 acquires the switching metadata, the information representing the current reproduction position, and the line-of-sight information supplied from the acquisition unit 21.
[0317]
 Further, the frame number selection unit 221 acquires the trigger information supplied from the trigger acquisition processing unit 22 and the viewpoint position information supplied from the viewpoint position acquisition unit 222.
[0318]
 In step S241, the frame number selection unit 221 determines whether or not the trigger information acquired by the process of step S240 includes the information of "with trigger" requesting switching to the sub scene.
[0319]
 If it is determined in step S241 that the information "with trigger" is not included, the process proceeds to step S242.
[0320]
 In step S242, the frame number selection unit 221 supplies the scene to be reproduced and the frame number to be reproduced of the scene acquired in step S240 to the switching processing unit 25 as information for identifying the reproduction target. .. After that, the process returns to step S124 of FIG. 17 and the subsequent processing is performed.
[0321]
 On the other hand, if it is determined in step S241 that the information "with trigger" is included, the process proceeds to step S243.
[0322]
 In step S243, the frame number selection unit 221 sets the variable i indicating the global frame ID for identifying the frame of all the scenes included in the content data based on the switching metadata acquired in step S240, i = 0. Set to.
[0323]
 In step S244, the frame number selection unit 221 determines whether or not the line-of-sight valid flag of the frame whose global frame ID is the i-th frame in the switching metadata is 1 based on the line-of-sight information. If it is determined that the line-of-sight valid flag is 1, the process proceeds to step S245.
[0324]
 In step S245, the frame number selection unit 221 determines the position (Ox, Oy, Oz) of the 3D object of the subject in the viewing space and the line of sight in the i-th frame of the global frame ID from the line-of-sight information acquired in step S240. Get the vector.
[0325]
 In step S246, the frame number selection unit 221 acquires the viewing position of the viewer from the viewpoint position information acquired in step S240. The frame number selection unit 221 sets a vector connecting the viewing position of the viewer and the 3D object of the subject, and calculates the angle formed by the set vector and the line-of-sight vector as the viewing line-of-sight angle. The frame number selection unit 221 stores the calculated viewing line-of-sight angle and information for identifying the frame having the viewing line-of-sight angle.
[0326]
 In step S247, the frame number selection unit 221 determines whether or not the viewing line-of-sight angle newly calculated in step S246 is the smallest among the viewing line-of-sight angles calculated and stored so far.
[0327]
 If it is determined that the new viewing line-of-sight angle is the minimum, the process proceeds to step S248.
[0328]
 In step S248, the frame number selection unit 221 discards the already stored information on the viewing line-of-sight angle, and identifies a new viewing line-of-sight angle having a smaller viewing line-of-sight angle value and a frame having the viewing line-of-sight angle. Memorize information.
[0329]
 If it is determined in step S244 that the line-of-sight valid flag is not 1, the above processes of steps S245 to S248 are skipped.
[0330]
 In step S249, the frame number selection unit 221 determines whether or not processing has been performed on all frames. If it is determined in step S249 that the processing of all frames has not been performed, the processing proceeds to step S250.
[0331]
 In step S250, the frame number selection unit 221 sets the variable i to i = i + 1. After that, the process returns to step S244, and the subsequent processes are repeated.
[0332]
 On the other hand, if it is determined in step S249 that the processing of all the frames has been performed, the processing proceeds to step S251.
[0333]
 In step S251, the frame number selection unit 221 supplies the minimum viewing line-of-sight angle and the frame number having the viewing line-of-sight angle to the switching processing unit 25 as information for identifying the reproduction target.
[0334]
 The information that identifies the playback target may be switched according to the number of viewing experiences acquired in step S240.
[0335]
 After the frame number selection unit 221 supplies the information for identifying the reproduction target to the switching processing unit 25, the frame number selection process ends. After that, the process returns to step S124 of FIG. 17 and the subsequent processing is performed.
[0336]
 By the above processing, a scene including a 3D object of the subject facing the viewer is selected as a playback target, and the scene is switched, so that it is possible to prevent the viewer from feeling uncomfortable. It will be possible. In addition, the viewer can get an immersive feeling.
[0337]
  When
 playing back the content data, the movement of the 3D object of the subject is switched by switching the scene, but instead of the 3D object of the subject, the background image of the viewing space is switched according to the switching metadata. You may do so.
[0338]
  The
 series of processes described above can be executed by hardware or software. When a series of processes are executed by software, the programs constituting the software are installed on the computer. Here, the computer includes a computer embedded in dedicated hardware, a general-purpose personal computer capable of executing various functions by installing various programs, and the like.
[0339]
 FIG. 26 is a block diagram showing a configuration example of hardware of a computer that executes the above-mentioned series of processes programmatically.
[0340]
 In the computer 300 shown in FIG. 26, the CPU (Central Processing Unit) 301, the ROM (Read Only Memory) 302, and the RAM (Random Access Memory) 303 are connected to each other via the bus 304.
[0341]
 The input / output interface 305 is also connected to the bus 304. An input unit 306, an output unit 307, a storage unit 308, a communication unit 309, and a drive 310 are connected to the input / output interface 305.
[0342]
 The input unit 306 includes, for example, a keyboard, a mouse, a microphone, a touch panel, an input terminal, and the like. The output unit 307 includes, for example, a display, a speaker, an output terminal, and the like. The storage unit 308 includes, for example, a hard disk, a non-volatile memory, or the like. The communication unit 309 includes, for example, a network interface. The drive 310 drives a removable medium 311 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.
[0343]
 In the computer configured as described above, the CPU 301 loads the program stored in the storage unit 308 into the RAM 303 via the input / output interface 305 and the bus 304, and executes the above-described series. Is processed. The RAM 303 also appropriately stores data and the like necessary for the CPU 301 to execute various processes.
[0344]
 The program executed by the computer (CPU301) can be recorded and applied to the removable media 311 as a package media or the like, for example. In that case, the program can be installed in the storage unit 308 via the input / output interface 305 by mounting the removable media 311 in the drive 310.
[0345]
 The program can also be provided via wired or wireless transmission media such as local area networks, the Internet, and digital satellite broadcasting. In that case, the program can be received by the communication unit 309 and installed in the storage unit 308.
[0346]
 In addition, this program can be pre-installed in the ROM 302 or the storage unit 308.
[0347]
  This technology can be applied to any image coding / decoding method. That is, as long as it does not contradict the above-mentioned present technology, the specifications of various processes related to image coding / decoding are arbitrary and are not limited to the above-mentioned examples.
[0348]
 Further, in the above, the case where the present technology is applied to the imaging device has been described, but the present technology can be applied not only to the imaging device but also to any device (electronic device). For example, the present technology can be applied to an image processing device or the like that performs image processing on an captured image obtained by high digital gain imaging performed in another device.
[0349]
 In addition, the present technology includes any configuration or a module (for example, video) that uses a processor (for example, a video processor) as a system LSI (Large Scale Integration) or the like, a plurality of processors, or the like to be mounted on an arbitrary device or a device that constitutes the system. It can also be implemented as a module), a unit using a plurality of modules (for example, a video unit), a set in which other functions are added to the unit (for example, a video set), or the like (that is, a part of a device).
[0350]
 Further, the present technology can also be applied to a network system composed of a plurality of devices. For example, it can be applied to a cloud service that provides services related to images (moving images) to arbitrary terminals such as computers, AV (Audio Visual) devices, portable information processing terminals, and IoT (Internet of Things) devices. it can.
[0351]
 Systems, equipment, processing departments, etc. to which this technology is applied should be used in any field such as transportation, medical care, crime prevention, agriculture, livestock industry, mining, beauty, factories, home appliances, weather, nature monitoring, etc. Can be done. Moreover, the use is arbitrary.
[0352]
 For example, the present technology can be applied to systems and devices used for providing ornamental contents and the like. Further, for example, the present technology can be applied to systems and devices used for traffic such as traffic condition supervision and automatic driving control. Further, for example, the present technology can be applied to systems and devices used for security purposes. Further, for example, the present technology can be applied to a system or device used for automatic control of a machine or the like. Further, for example, the present technology can be applied to systems and devices used for agriculture and livestock industry. The present technology can also be applied to systems and devices for monitoring natural conditions such as volcanoes, forests and oceans, and wildlife. Further, for example, the present technology can be applied to systems and devices used for sports.
[0353]
  In the
 present specification, the system means a set of a plurality of components (devices, modules (parts), etc.), and it does not matter whether all the components are in the same housing. Therefore, a plurality of devices housed in separate housings and connected via a network, and a device in which a plurality of modules are housed in one housing are both systems. ..
[0354]
 Further, for example, the configuration described as one device (or processing unit) may be divided and configured as a plurality of devices (or processing units). On the contrary, the configurations described above as a plurality of devices (or processing units) may be collectively configured as one device (or processing unit). Further, of course, a configuration other than the above may be added to the configuration of each device (or each processing unit). Further, if the configuration and operation of the entire system are substantially the same, a part of the configuration of one device (or processing unit) may be included in the configuration of another device (or other processing unit). ..
[0355]
 Further, for example, the present technology can have a cloud computing configuration in which one function is shared and jointly processed by a plurality of devices via a network.
[0356]
 Further, for example, the above-mentioned program can be executed in any device. In that case, the device may have necessary functions (functional blocks, etc.) so that necessary information can be obtained.
[0357]
 For example, each step described in the above-mentioned flowchart can be executed by one device or can be shared and executed by a plurality of devices. Further, when a plurality of processes are included in one step, the plurality of processes included in the one step can be executed by one device or shared by a plurality of devices. In other words, a plurality of processes included in one step can be executed as processes of a plurality of steps. On the contrary, the processes described as a plurality of steps can be collectively executed as one step.
[0358]
 In the program executed by the computer, the processing of the steps for describing the program may be executed in chronological order according to the order described in this specification, or may be called in parallel or called. It may be executed individually at a necessary timing such as time. That is, as long as there is no contradiction, the processing of each step may be executed in an order different from the above-mentioned order. Further, the processing of the step for writing this program may be executed in parallel with the processing of another program, or may be executed in combination with the processing of another program.
[0359]
 It should be noted that the present techniques described in the present specification can be independently implemented independently as long as there is no contradiction. Of course, any plurality of the present technologies can be used in combination. For example, some or all of the techniques described in any of the embodiments may be combined with some or all of the techniques described in other embodiments. It is also possible to carry out a part or all of any of the above-mentioned techniques in combination with other techniques not described above.
[0360]
 The present technology can also have the following configurations.
[0361]
 (1)
 An acquisition unit that acquires a first 3D shape image generated from a plurality of viewpoint images in which a subject is captured from different viewpoints and a second 3D shape image which is an image different from the first 3D shape image. , A
 switching process for switching the frame to be reproduced from the frame of the first 3D shape image to the frame of the second 3D shape image based on the state of the viewer viewing the first 3D shape image. A
 video playback device including a unit.
(2) The
 switching processing unit is based on the positional relationship between the viewpoint position of the viewer in the viewing space as the state of the viewer and the 3D shape of the subject in the first 3D shape image. switching the
 video reproducing apparatus according to (1).
(3)  The video reproduction device according to (2) above, further comprising
 a trigger acquisition processing unit that acquires trigger information for requesting switching of frames to be reproduced based on the positional relationship
.
(4) The
 acquisition unit further acquires switching metadata including a transitionable flag indicating whether or not the frame of the first 3D shape image can be switched to the frame of the second 3D shape image. ,
 The video reproduction device according to (3) above.
(5)
 The switching metadata is generated based on the similarity of the subject of the 3D object in each frame of the first 3D shape the second 3D shape image and 3D objects of an object in each frame of the video,
 the (4 ). The video playback device.
(6) The
 switching processing unit switches a frame to be reproduced from a frame of the first 3D shape image to a frame of the second 3D shape image based on the switching metadata, according to the
 above (5). Video playback device.
(7) When
 the frame of the first 3D shape image being reproduced is a frame in which switching to the frame of the second 3D shape image is not permitted by the switching metadata, the
 switching processing unit may perform the switching processing unit. The
 above (6) , wherein even when the trigger information is acquired, the switching from the first 3D shape image to the second 3D shape image is not performed until the frame in which the switching is recognized by the switching metadata . Video playback device.
(8)  The video playback device according to (7) above
 , wherein the switching metadata further includes a transitionable frame ID according to the number of viewing experiences
.
(9)  The video reproduction according to (8)
 ,
wherein the acquisition unit further acquires line-of-sight information of a 3D object of the subject for each frame of the first 3D shape video and the second 3D shape video. apparatus.
(10) The
 line-of-sight information includes the position of the 3D object of the subject in the 3D shape image, the line-of-sight vector, and the line-of-sight valid flag for each frame of the first 3D shape image and the second 3D shape image. ,
 The video reproduction device according to (9) above.
(11)  The video reproduction device
 according to
(10), wherein the switching processing unit further switches the transitionable frame according to the line-of-sight information and the number of viewing experiences .
(12)
 The first 3D shape video is the main video of the content, and the second 3D shape video is
 any of the above (1) to (11) which is the video of the specific scene of the sub of the content. The video playback device described in Crab.
(13)  The image reproduction device according to any one of (1) to (12 )
 above, wherein the switching processing unit switches a background image of a viewing space instead of a 3D shape image of a subject according to the switching metadata
.
(14)
 video reproduction device,
 different from the first 3D shape image object is generated from a plurality of view images captured from a viewpoint, the first second 3D shape image is a 3D shape image and the different images And get,
 A
 reproduction method in which a frame to be reproduced is switched from a frame of the first 3D shape image to a frame of the second 3D shape image based on a state of a viewer viewing the first 3D shape image .
(15) A  first 3D shape image generated from a plurality of viewpoint images in which a subject is captured from different viewpoints and a second 3D shape image which is an image different from the first 3D shape image are displayed on a
 computer.  Based on the state of the viewer who has
acquired and
is viewing the first 3D shape image, the frame to be reproduced is switched from the frame of the first 3D shape image to the frame of the second 3D shape image.
 A program for executing processing.
Code description
[0362]
 1 Video generator, 2 Video playback device, 3 Display device, 11 Multi-viewpoint video acquisition unit, 12 3D model generation unit, 13 Playback information generation unit, 21 Acquisition unit, 22 Trigger acquisition processing unit, 24 Frame number selection unit, 25 Switching processing unit, 26 playback unit, 101 video generation device, 102 video playback device, 111 switching metadata generation unit, 112 3D model generation unit, 121 frame number selection unit, 201 video generation device, 202 video playback device, 211 multi-viewpoint Video acquisition unit, 212 3D model generation unit, 213 line-of-sight information generation unit, 221 frame number selection unit, 222 viewpoint position acquisition unit
The scope of the claims
[Claim 1]
 An acquisition unit that acquires a first 3D shape image generated from a plurality of viewpoint images in which a subject is captured from different viewpoints and a second 3D shape image that is an image different from the first 3D shape image.
 A switching processing unit that switches the frame to be reproduced from the frame of the first 3D shape image to the frame of the second 3D shape image based on the state of the viewer viewing the first 3D shape image. A
 video playback device including.
[Claim 2]
 The switching processing unit is
 claimed to switch frames based on the positional relationship between the viewpoint position of the viewer in the viewing space as the state of the viewer and the 3D shape of the subject in the first 3D shape image. Item 2. The video playback device according to item 1.
[Claim 3]

 The video playback device according to claim 2, further comprising  a trigger acquisition processing unit that acquires trigger information that requests switching of frames to be reproduced based on the positional relationship .
[Claim 4]
 The acquisition unit further acquires the switching metadata including transition flag from the frame of the first 3D shape image representing whether or not it is possible to switch to a frame of the second 3D shape image,
 claim The video reproduction device according to 3.
[Claim 5]
 The switching metadata is generated based on the similarity of the subject of the 3D object in each frame of the first 3D shape the second 3D shape image and 3D objects of an object in each frame of the video,
 claim 4 The video playback device described in.
[Claim 6]

 The video reproduction device according to claim 5,  wherein the switching processing unit switches a frame to be reproduced from a frame of the first 3D shape image to a frame of the second 3D shape image based on the switching metadata. ..
[Claim 7]
 When the frame of the first 3D shape image being reproduced is a frame in which switching to the frame of the second 3D shape image is not permitted by the switching metadata, the
 switching processing unit uses the trigger information.
 The video reproduction device according to claim 6 , wherein the switching from the first 3D shape image to the second 3D shape image is not performed until the frame in which the switching is recognized by the switching metadata. ..
[Claim 8]

 The video playback device  according to claim 7, wherein the switching metadata further includes a transitionable frame ID according to the number of viewing experiences .
[Claim 9]

 The video reproduction device according to claim 8,  wherein the acquisition unit further acquires line-of-sight information of a 3D object of a subject for each frame of the first 3D shape video and the second 3D shape video .
[Claim 10]
 Wherein the viewing information, for each frame of the first 3D shape image and the second 3D shape image, the position of the subject of the 3D objects in the 3D shape image, line-of-sight vectors include gaze valid flag,
 claim 9. The video playback device according to 9.
[Claim 11]

 The video playback device  according to claim 10, wherein the switching processing unit further switches the transitionable frame according to the line-of-sight information and the number of viewing experiences .
[Claim 12]

 The video playback device according to claim 11,  wherein the first 3D shape video is a video that is the main content, and the second 3D shape video is a video of a specific scene that is a sub of the content .
[Claim 13]

 The image reproduction device according to claim 12,  wherein the switching processing unit switches a background image of a viewing space instead of a 3D shape image of a subject according to the switching metadata .
[Claim 14]
 The image reproduction device
 acquires a first 3D shape image generated from a plurality of viewpoint images in which a subject is captured from different viewpoints and a second 3D shape image which is an image different from the first 3D shape image. Then
 , based on the state of the viewer who is viewing the first 3D shape image, the frame to be reproduced is switched from the frame of the first 3D shape image to the frame of the second 3D shape image, and the
 reproduction is performed. Method.
[Claim 15]
 A computer
 acquires a first 3D shape image generated from a plurality of viewpoint images in which a subject is captured from different viewpoints and a second 3D shape image which is an image different from the first 3D shape image.
 Based on the state of the viewer viewing the first 3D shape image, the
 process of switching the frame to be reproduced from the frame of the first 3D shape image to the frame of the second 3D shape image is executed. Program to make you.
Scope of amended claims (Article 19 of the Convention)
[April 15, 2020 (15.04.2020) Accepted by the International Bureau]
[1]
[After correction] Acquire a first 3D shape image generated from a plurality of viewpoint images in which a subject is captured from different viewpoints, and a second 3D shape image which is an image independent of the first 3D shape image. A
 frame to be reproduced from the frame of the first 3D shape image to the frame of the second 3D shape image based on the state of the acquisition unit and the viewer viewing the first 3D shape image. A
 video playback device including a switching processing unit for switching .
[2]
[After correction] The subject that is the target of the second 3D shape image is the same subject as the subject that is the target of the first 3D shape image, and the first 3D shape image and the second 3D shape are the same.
 The image reproduction device according to claim 1, wherein the image is an image showing different movements of the same subject .
[3]
[After correction] The switching processing unit is based on the positional relationship between the viewpoint position of the viewer in the viewing space as the state of the viewer and the 3D shape of the subject in the first 3D shape image.
 The video playback device according to claim 1, wherein the frame is switched .
[Four]
[After correction]
 The video playback device according to claim 3, further comprising a trigger acquisition processing unit that acquires trigger information that requests switching of frames to be reproduced based on the positional relationship .
[Five]
[After correction] The acquisition unit further acquires switching metadata including a transitionable flag indicating whether or not the frame of the first 3D shape image can be switched to the frame of the second 3D shape image. to,
 the video reproducing apparatus according to claim 4.
[6]
[After correction] The switching metadata is generated based on the similarity between the 3D object of the subject in each frame of the first 3D shape image and the 3D object of the subject in each frame of the second 3D shape image. ,
 The video reproduction device according to claim 5.
[7]
[After correction] The switching processing unit switches a frame to be reproduced from a frame of the first 3D shape image to a frame of the second 3D shape image based on the switching metadata, according to
 claim 6. Video playback device.
[8]
[After correction] When the frame of the first 3D shape image being reproduced is a frame in which switching to the frame of the second 3D shape image is not permitted by the switching metadata, the
 switching processing unit performs the switching processing unit. the even when the trigger information is obtained, up to the frame to be recognized switched by the switching metadata, not switched from the first 3D shape image to the second 3D shape images,
 according to claim 7 Video playback device.
[9]
[After correction]
 The video playback device according to claim 8, wherein the switching metadata further includes a transitionable frame ID according to the number of viewing experiences .
[Ten]
[After correction]
 The video reproduction according to claim 9, wherein the acquisition unit further acquires line-of-sight information of a 3D object of the subject for each frame of the first 3D shape video and the second 3D shape video. apparatus.
[11]
[After correction] The line-of-sight information includes the position of the 3D object of the subject in the 3D shape image, the line-of-sight vector, and the line-of-sight valid flag for each frame of the first 3D shape image and the second 3D shape image. is,
 the video reproducing apparatus according to claim 10.
[12]
[After correction]
 The video playback device according to claim 11, wherein the switching processing unit further switches the transitionable frame according to the line-of-sight information and the number of viewing experiences .
[13]
[After correction]
 The video playback device according to claim 12, wherein the first 3D shape video is a video that is the main content, and the second 3D shape video is a video of a specific scene that is a sub of the content. ..
[14]
[After correction]
 The image reproduction device according to claim 13, wherein the switching processing unit switches a background image of a viewing space instead of a 3D shape image of a subject according to the switching metadata .
[15]
[After correction] The image reproduction device has a
 first 3D shape image generated from a plurality of viewpoint images in which a subject is captured from different viewpoints, and a second 3D image that is independent of the first 3D shape image. The shape image is acquired and
 reproduced from the frame of the first 3D shape image to the frame of the second 3D shape image based on the state of the viewer who is viewing the first 3D shape image.
 Playback method to switch frames .
[16]
[Addition] A
 first 3D shape image generated from a plurality of viewpoint images in which a subject is captured from different viewpoints and a second 3D shape image which is an image independent of the first 3D shape image are added to the computer. Is acquired, and
 the frame to be reproduced is switched from the frame of the first 3D shape image to the frame of the second 3D shape image based on the state of the viewer viewing the first 3D shape image. ,
 A program to execute processing.
[17]
[Addition] A process of generating a first 3D shape image generated from a plurality of viewpoint images in which a subject is imaged from different viewpoints, and a plurality of
 images of the same subject as the first 3D shape image from different viewpoints. 3D shape image generated from the viewpoint image of the above, and the subject of the first 3D shape image generated for switching the image based on the state of the viewer viewing the first 3D shape image.
 A method of generating 3D shape video data including a process of generating a second 3D shape video that moves differently from the above .

Documents

Application Documents

# Name Date
1 202017057405-TRANSLATIOIN OF PRIOIRTY DOCUMENTS ETC. [31-12-2020(online)].pdf 2020-12-31
2 202017057405-STATEMENT OF UNDERTAKING (FORM 3) [31-12-2020(online)].pdf 2020-12-31
3 202017057405-PRIORITY DOCUMENTS [31-12-2020(online)].pdf 2020-12-31
4 202017057405-POWER OF AUTHORITY [31-12-2020(online)].pdf 2020-12-31
5 202017057405-FORM 1 [31-12-2020(online)].pdf 2020-12-31
6 202017057405-DRAWINGS [31-12-2020(online)].pdf 2020-12-31
7 202017057405-DECLARATION OF INVENTORSHIP (FORM 5) [31-12-2020(online)].pdf 2020-12-31
8 202017057405-COMPLETE SPECIFICATION [31-12-2020(online)].pdf 2020-12-31
9 202017057405-Proof of Right [12-01-2021(online)].pdf 2021-01-12
10 202017057405-FORM-26 [11-02-2021(online)].pdf 2021-02-11
11 202017057405-Verified English translation [24-03-2021(online)].pdf 2021-03-24
12 202017057405-FORM 3 [08-04-2021(online)].pdf 2021-04-08
13 202017057405.pdf 2021-10-19
14 202017057405-FORM 18 [20-12-2022(online)].pdf 2022-12-20
15 202017057405-FER.pdf 2023-01-17
16 202017057405-OTHERS [17-07-2023(online)].pdf 2023-07-17
17 202017057405-FER_SER_REPLY [17-07-2023(online)].pdf 2023-07-17
18 202017057405-DRAWING [17-07-2023(online)].pdf 2023-07-17
19 202017057405-CORRESPONDENCE [17-07-2023(online)].pdf 2023-07-17
20 202017057405-COMPLETE SPECIFICATION [17-07-2023(online)].pdf 2023-07-17
21 202017057405-CLAIMS [17-07-2023(online)].pdf 2023-07-17
22 202017057405-ABSTRACT [17-07-2023(online)].pdf 2023-07-17
23 202017057405-Others-300623.pdf 2023-08-07
24 202017057405-GPA-300623.pdf 2023-08-07
25 202017057405-Correspondence-300623.pdf 2023-08-07
26 202017057405-PatentCertificate24-02-2025.pdf 2025-02-24
27 202017057405-IntimationOfGrant24-02-2025.pdf 2025-02-24

Search Strategy

1 SearchE_16-01-2023.pdf

ERegister / Renewals