Sign In to Follow Application
View All Documents & Correspondence

Video Sound Processing Device Video Sound Processing Method And Program

Abstract: The present technology pertains to a video sound processing device a video sound processing method and a program by which a desired object sound can be divided more easily with higher precision. The video sound processing device is provided with: a display control unit that causes displaying of a video object based on a video signal; an object selection unit that selects a prescribed video object from among one or more video objects; and an extraction unit that extracts as an audio object signal an audio signal of the video object selected by the object selection unit. The present technology is applicable to a video sound processing device.

Get Free WhatsApp Updates!
Notices, Deadlines & Correspondence

Patent Information

Application #
Filing Date
23 November 2018
Publication Number
09/2019
Publication Type
INA
Invention Field
COMMUNICATION
Status
Email
ranjna.dutt@remfry.com
Parent Application
Patent Number
Legal Status
Grant Date
2024-04-15
Renewal Date

Applicants

SONY CORPORATION
1-7-1, Konan, Minato-ku, Tokyo 1080075

Inventors

1. HONMA Hiroyuki
c/o SONY CORPORATION, 1-7-1, Konan, Minato-ku, Tokyo 1080075
2. YAMAMOTO Yuki
c/o SONY CORPORATION, 1-7-1, Konan, Minato-ku, Tokyo 1080075

Specification

Technical field
[0001]
 This technology audiovisual processing apparatus and method, and a program, a video sound processing apparatus and method capable of separating the desired object sound easier and accurate, and a program.
BACKGROUND
[0002]
 In recent years, the object audio technology has been used in movies and games, etc., it has also been developed coding system that can handle the object audio. For example, as a standard for coding of object audio, the international standard MPEG (Moving Picture Experts Group) -H Part 3: such 3D audio standards are known (e.g., see Non-Patent Document 1).
[0003]
 In such a coding method, and conventional two-channel stereo system, with a multi-channel stereo system, such as a 5.1 channel is treated as audio objects independent moving sound source or the like, the position of the audio object with the signal data of the audio objects it is possible to encode information as metadata. By doing so, in the conventional coding method it has been difficult, can be easily processed for a particular sound source at the time of reproduction. Specifically, for example, as processing of a particular sound source, it is possible to perform such additional volume control and effects for each audio object.
[0004]
 Such object audio encoding, for example, be used when shooting a moving image in a camcorder or smartphone, sound family that is the shooting target is recorded as object sounds, other sounds of the background sound it is, after shooting, can be independently at the time of playback and editing to manipulate the voice of the family.
[0005]
 However, when picking up the object sound in mobile devices such as camcorders and smart phones, what was the object sounds, nothing to automatically determine what to the background sound, for the sake of computational resources and user interface, easily is not.
[0006]
 Further, as a method for separating the object sounds, conventionally, various sound source separation techniques have been proposed, for example, a technique for separating the sound sources by using an independent component analysis and multiple microphones have been proposed (e.g., Patent references 1).
CITATION
Non-patent literature
[0007]
非特許文献1 : INTERNATIONAL STANDARD ISO/IEC 23008-3 First edition 2015-10-15 Information technology - High efficiency coding and media delivery in heterogeneous environments - Part 3: 3D audio
Patent Document
[0008]
Patent Document 1: JP 2010-233173 JP
Summary of the Invention
Problems that the Invention is to Solve
[0009]
 However, these sound source separation techniques are those for separating the adaptive sound source from the input speech signal from the microphone in practice is often better when would include sounds other than imaging subject, precisely desired it has been difficult to separate the object sound.
[0010]
 This technology has been made in view of such circumstances, it is desirable to make it possible to separate the desired object sound easier and accurate.
Means for Solving the Problems
[0011]
 Audiovisual processing apparatus according to an embodiment of the present technology, a display control unit for displaying an image object based on the image signal, from among one or more of the video object, and the object selection unit for selecting a predetermined the video object, and a extraction unit for extracting an audio signal of the video object selected by the object selection portion as the audio object signal.
[0012]
 The extraction unit may extract the audio object signal from the audio signal.
[0013]
 Wherein the extraction unit from the audio signal, a signal other than the audio object signal of said selected the video object can be extracted as a background sound signal.
[0014]
 The said object selecting unit, to generate an object position information indicating a position in the space of the said selected video object, the extraction section, based on the object position information, thereby extracting the audio object signal can.
[0015]
 The extraction unit may extract the audio object signal by the sound source separation using the object position information.
[0016]
 The extraction unit can perform a fixed beamforming as the sound source separation.
[0017]
 The audiovisual processing apparatus, based on said video signal, said further provided an image object recognition unit for recognizing an image object, the display control unit, together with the image object, an image based on the recognition result of the image object it can be displayed.
[0018]
 Wherein the video object recognition unit, it is possible to recognize the video object by facial recognition.
[0019]
 Wherein the display control unit, as the image can be displayed a frame region of the image object.
[0020]
 Wherein the object selection unit may, depending on the selection operation by the user to select the video object.
[0021]
 Wherein the object selection unit, it is possible to generate metadata of the selected the video object.
[0022]
 Wherein the object selection unit, as the metadata, it is possible to generate an object position information indicating a position in the space of the said selected video object.
[0023]
 Wherein the object selection unit, as the metadata, the processing priority of the selected the video object can be generated.
[0024]
 Wherein the object selection unit, as the metadata, it is possible to produce a spread information indicating the spatial spread of the selected region of the video object has been.
[0025]
 Video The sound processing apparatus, the audio object signal and the metadata may further include an audio encoding unit for encoding.
[0026]
 The audiovisual processing apparatus, a video encoding unit for encoding the video signal, the video bit stream obtained by the encoding of the video signal, obtained by the encoding of the audio object signal and the metadata It may further include a multiplexer for multiplexing the audio bit stream.
[0027]
 The audiovisual processing apparatus may be further provided an imaging unit for obtaining the video signal by performing photographing.
[0028]
 Sound pickup for obtaining the audio signal by performing sound collection can be further provided.
[0029]
 Audiovisual processing method or a program according to an embodiment of the present technology, a display control step of displaying the image object based on the image signal, from among one or more of the video object, the object selecting step of selecting a predetermined said video object When includes an extraction step of extracting an audio signal of the video object selected by the object selecting step as the audio object signal.
[0030]
 In one aspect of the present technology, image is displayed object based on the video signal, one or from a plurality of the video object, given the video object is selected, the audio signal is an audio object in said selected video object It is extracted as a signal.
Effect of the invention
[0031]
 According to the embodiments of the present technology, it is possible to separate the desired object sound easier and accurate.
[0032]
 Here, the advantages described in the present invention is not necessarily limited, it may be any of the effects described in the present disclosure.
BRIEF DESCRIPTION OF THE DRAWINGS
[0033]
FIG. 1 is a diagram showing an example of the configuration of a video sound processing apparatus.
2 is a diagram showing an example of the configuration of a video sound processing apparatus.
Is a flow chart illustrating the FIG. 3] encoding process.
It is a diagram illustrating a display of FIG. 4 audio object selection and rectangular frame image.
Is a diagram illustrating FIG. 5 the coordinate system of the object position information.
Is a diagram illustrating the relationship between FIG. 6 subject and the lens and the imaging surface.
7 is a diagram for describing calculation of object position information.
8 is a diagram showing the syntax of the metadata including the object position information.
It is a diagram illustrating the syntax of the metadata including the FIG. 9 processing priority.
Is a diagram for describing setting of the FIG. 10 process priority.
11 is a diagram showing the syntax of the metadata including the spread information.
12 is a diagram for explaining the spread information.
13 is a diagram for describing calculation of spread information.
14 is a diagram showing a configuration example of a computer.
DESCRIPTION OF THE INVENTION
[0034]
 Hereinafter, with reference to the accompanying drawings, a description will be given of an embodiment according to the present technology.
[0035]

 In the present technique, using an object recognition technology such as the face recognition technology for moving images, the device user of the recognition result on a display device it is presented, to separate the direction of the sound corresponding to the object the device user selects from the recognition result as the object sound is to encode the object audio encoding technique as a background sound other than sounds.
[0036]
 Figure 1 is a diagram showing a configuration example of an embodiment of the audiovisual processing apparatus according to the present technology.
[0037]
 Audiovisual processing apparatus 11 includes an imaging unit 21 shown in FIG. 1, a video object recognition unit 22, a video object recognition result display control unit 23, image display unit 24, the object selection unit 25, the sound pickup unit 26, the sound source separation unit 27, audio encoding unit 28, and a video encoding unit 29, and the multiplexer 30.
[0038]
 Imaging unit 21, for example, from the imaging unit consists of the image pickup element and a lens or the like, by photographing a subject, supplies the video signal of a moving image obtained as a result to the video object recognition unit 22 and the video encoding unit 29 to.
[0039]
 Video object recognition unit 22 based on the video signal supplied from the imaging unit 21, recognizes the video object on the moving image, and supplies the recognition result together with the video signal to the video object recognition result display control unit 23.
[0040]
 The video object recognition unit 22, face recognition and object recognition, such as by motion capture, video objects on a moving image is recognized (detected).
[0041]
 For example, Japanese Patent No. 4492036, have been described in detail object recognition techniques using known image. Such object recognition technique using the known image is adapted to practically available level with the evolution of recent machine learning techniques and cloud network. By using these object recognition technology to recognize any object, extracted by presenting the recognition result to the device user, the audio signal emanating from the position of the object device user has selected as the signal of the audio objects it is possible to become.
[0042]
 Video object recognition result display control unit 23, based on the recognition result and the video signal of the video object that is supplied from the video object recognition unit 22, for controlling display of a moving image in the video display unit 24. That is, the video object recognition result display control unit 23, information indicating with displaying a moving image by supplying a video signal to the video display unit 24, based on the recognition result of the image object, the position of the video object on the video image It is superimposed display.
[0043]
 The video object recognition result display control unit 23 supplies the recognition result of the image object to the object selection unit 25.
[0044]
 Video display unit 24, for example, a liquid crystal display panel, for displaying an image such as a moving image in accordance with control of the video object recognition result display control unit 23.
[0045]
 Also, the moving image on the video display unit 24 (video) the recognition result of the image object with, that is, when the information indicating the position of the video object is displayed, the device user performs an operation for designating a desired video object. Then, a signal indicating the result of the specified operation by the device user are supplied to the object selection unit 25.
[0046]
 Object selecting unit 25, a recognition result of the image object that is supplied from the video object recognition result display control unit 23, based on the signal supplied in accordance with the designation operation of the video object by the device user, the device user generating object position information indicating a position on the three-dimensional space of selected video object. Object selecting unit 25, the object position information obtained is supplied to the sound source separation unit 27 and the audio encoding unit 28.
[0047]
 Sound pickup unit 26 includes, for example, a microphone, picks up ambient sound of the audiovisual processing apparatus 11, and supplies the resulting audio signal to the sound source separation unit 27. In the video sound processing apparatus 11, so that the moving image captured by the imaging unit 21, the content consisting of audio picked up by the sound pickup unit 26 is obtained. That is, the audio signal obtained by the sound pickup unit 26 is an audio signal accompanying the video signal obtained by the imaging unit 21.
[0048]
 The sound source separation unit 27, based on the object position information supplied from the object selection unit 25 performs sound source separation to the supplied audio signal from the sound collection unit 26.
[0049]
 By source separation in the sound source separation unit 27, a video audio signal obtained by the sound pickup unit 26, the audio object signals is an audio signal of the video object selected by the device user, selected by the device user audio signals other than speech object, that is separated into background sound signal and a signal of background sound. Here, the audio signal of the video object selected by the device user, the object sounds, that is separated (extracted) as the audio object signals is an audio signal of an audio object.
[0050]
 The sound source separation unit 27 supplies an audio object signal and the background sound signal obtained by the sound source separation to the audio encoding unit 28.
[0051]
 Audio encoding unit 28, multiplexing the audio object signal and the background sound signal supplied from the sound source separation section 27 encodes the object position information supplied from the object selection unit 25, the resulting audio bitstream supplied to the unit 30.
[0052]
 Video encoding unit 29 encodes the video signal supplied from the imaging unit 21, and supplies the resulting video bit stream to the multiplexing unit 30. Multiplexing unit 30 obtains a video bit stream supplied from the video encoding unit 29, multiplexes the audio bit stream supplied from the audio encoding unit 28, the output bit stream as a final output. Multiplexer 30 outputs the output bit stream thus obtained.
[0053]

 audiovisual processing apparatus shown in FIG. 1 11, while indicating a form of a common practice, in the following, in order to perform a more particular description, image object recognition as the recognition method of the video objects in part 22 will be described a case where the face recognition technology is used as a specific example.
[0054]
 In such a case, the audiovisual processing apparatus is arranged as shown for example in FIG. Note that portions corresponding to the case in FIG. 1 in FIG. 2 are denoted by the same reference numerals, and description thereof will be omitted as appropriate.
[0055]
 Audiovisual processing apparatus 61 shown in FIG. 2, the imaging unit 21, the face recognition unit 71, the face recognition result display control unit 72, image display unit 73, the person selecting section 74, the sound pickup unit 26, the sound source separation unit 27, an audio code section 28, and a video encoding unit 29, and the multiplexer 30.
[0056]
 The configuration of the video audio processing device 61, instead of the video object recognition unit 22 to the object selection unit 25, unlike the configuration of the audiovisual processing apparatus 11 in that the face recognition section 71 to the person selecting section 74 is provided, in other respects it has the same structure as the video sound processing apparatus 11.
[0057]
 Face recognition unit 71 corresponds to the video object recognition unit 22 shown in FIG. Face recognition unit 71 recognizes a face of a person on the moving image by performing face recognition processing on the video signal supplied from the imaging unit 21, the recognition result to the face recognition result display control unit 72 together with the video signal supplies. That is, in the face recognition section 71, based on the video signal, the human face as a video object is recognized (detected).
[0058]
 Face recognition result display control unit 72 corresponds to the video object recognition result display control unit 23 shown in FIG. Face recognition result display control unit 72, based on the face recognition result supplied from the face recognition section 71 and the video signal, for controlling display of a moving image in the video display unit 73. That is, the face recognition result display control unit 72 causes to display a moving image content by supplying a video signal to the video display unit 73, based on the face recognition result, the human face as a video object onto a moving picture position is superimposed displays information indicating a.
[0059]
 The face recognition result display control unit 72 supplies the face recognition result to the person selecting section 74.
[0060]
 Video display unit 73, for example, a liquid crystal display panel, corresponding to the video display unit 24 shown in FIG. Video display unit 73 displays an image such as a moving image in accordance with the control of the face recognition result display control unit 72.
[0061]
 The video display unit 73 has a touch panel 81 provided to be superimposed on the display screen, the touch panel 81 supplies a signal corresponding to an operation by the device user to the person selecting section 74. Here, an example will be described for performing various input by the device user operates the touch panel 81, but other, mouse, buttons, may be input device is used input such as a keyboard is performed .
[0062]
 Result of the face recognition with moving image (video) on the video display unit 73, i.e. the information indicating the position of the face of a person as an image object is displayed, the device user operates the touch panel 81, the desired person to specify. Then, the person selection unit 74 from the touch panel 81, the selection result of the person (face) by the device user is provided.
[0063]
 People selecting unit 74 corresponding to the object selection unit 25 shown in FIG. People selecting unit 74, the result of face recognition supplied from the face recognition result display control unit 72, based on the selection result of the person which is supplied from the touch panel 81, the person selected by the device user, that is a video object select, generates object position information indicating a position on the three-dimensional space of the image object.
[0064]
 Also, a person selecting unit 74 supplies the selection result of the person by the device user on the video display unit 73 to perform a predetermined display, cause the can confirm the selection result of a person the device user performs by itself .
[0065]
 People selecting unit 74, the object position information obtained for the person selected by the device user, and supplies the sound source separation unit 27 and the audio encoding unit 28.
[0066]

 Next, the operation of the audiovisual processing apparatus 61. That is, below with reference to the flowchart of FIG. 3 will be described coding processing performed by the video audio processing device 61.
[0067]
 The encoding process, for example, a video sound processing apparatus 61 is operated by the device user, taking with sound image as a content is started when instructed. Here, although illustrating a case where moving image (video) is captured as an image of the contents as an example, in the following, will be referred to as simply image both of the content moving images.
[0068]
 When photographing is started, the imaging unit 21 photographs a subject and supplies the resulting video signal, sequentially, the face recognition unit 71 and the video coding unit 29. Further, the sound pickup unit 26 picks up sound, and supplies the resulting audio signal, successively, to the sound source separation unit 27.
[0069]
 In step S11, the face recognition unit 71, based on the video signal supplied from the imaging unit 21 detects a human face from the image of the content based on the video signal, the face recognition result and the detection result and the image signal to the display control unit 72.
[0070]
 For example the face recognition unit 71 detects a face from an image of a person by performing a face recognition process using such characteristic amount and dictionaries for the video signal. Here, for example the position of a person's face in the image, such as the size of the area of ​​the face of the person is recognized (detected). In this manner is detected, the human in the image of the content, and more particularly a human face is a video object.
[0071]
 In step S12, the face recognition result display control unit 72, based on the detection result and the image signal of the face supplied from the face recognition section 71, in the region of each face detected in the image content, each of them faces generating a signal of rectangular frame images for displaying a rectangular frame surrounding the.
[0072]
 Incidentally, the rectangular frame image, a region of the human face is a video object, requires long an image to display the frame surrounding at least a portion of the area of ​​the face, not necessarily the entire region of the face is surrounded by a rectangular frame no.
[0073]
 The face recognition result display control unit 72, the display position and size of the human face generated rectangular frame image for each, i.e. the person selects the face position information indicating the position and size of the area of ​​the human face in the image supplied to the part 74.
[0074]
 In step S13, the face recognition result display control unit 72, a signal generated rectangular frame image, supplies the video signal of the content to the video display unit 73, and controls the image display unit 73, a human face area rectangular frame images to display an image of the superimposed display content to.
[0075]
 Thus, on the display screen of the image display unit 73, an image of the content, that is along with the video object based on the image signal, so that the rectangular frame image is an image based on the recognition result of the image object is displayed.
[0076]
 Note that facial recognition technology to detect a human face from the image of the content is proposed various methods, for example, Japanese Patent No. 4264663 (hereinafter, referred to as reference 1) such as is described in detail there. In references 1, perform facial recognition, it is described that is superimposed on a rectangular frame image based on the recognition result.
[0077]
 And methods of face recognition processing in the face recognition section 71, superimposed display technique of the rectangular frame image by the face recognition result display control unit 72 or may be applied as a procedure similar as described in reference 1, the human face recognized, if it is possible to perform display that surrounds the recognized face area, may be any method. Further, for example, here will be described an example of displaying a rectangular frame image, the image of the content of the result of the face recognition, if it is possible in particular to display the position of the face is not limited to the rectangular frame image, the human face outline highlighting and displaying may be or superimpose graphical images of other shapes.
[0078]
 Thus with the image of the content to the video display unit 73 is displayed in the, manipulating the rectangular frame image is displayed in the area of ​​the face of a person on the image, the device user, that is, the audiovisual processing apparatus 61 the user operates the touch panel 81 to select a desired person as an audio object.
[0079]
 For example selection of the desired person it is performed by specifying a rectangular frame image displayed on the face region of the person. The user, from among the one or more persons are displayed on the video display unit 73 (video objects), you can either be selected person only one, it is also possible to select multiple persons.
[0080]
 In the video sound processing apparatus 61, a person is selected in this way, that is, sound emitted from the video object is to be the object sound. The signal of the object sound is extracted from the audio signal as an audio object signal. Therefore, selecting a video object person can be said to be an operation for selecting a desired person as an audio object.
[0081]
 If desired person by the device user is selected, the person selection unit 74 from the touch panel 81, the signal corresponding to the selection operation of the device user are supplied.
[0082]
 In step S14, the person selecting section 74, based on a signal corresponding to the selection operation by the device user supplied from the touch panel 81, on the face position information supplied from the face recognition result display control unit 72, a device user to select a specified audio object by.
[0083]
 In step S15, the person selecting section 74 controls the image display unit 73 based on the selection result of the audio object in the step S14, the rectangular frame displayed in the area of ​​a person's face that has been selected as the audio objects image, i.e. to highlight the selected rectangular frame image.
[0084]
 For example, as shown in FIG. 4, the person HM11 or person HM13 in three on the video display unit 73 is displayed an image of the content that is included as a subject.
[0085]
 In this example, in step S11, the face of the person HM11 to person HM13 is detected, in step S13, in the region of the face of their person, rectangular frame image FR11 to rectangular frame image FR13 are superimposed.
[0086]
 In this state, for example, a device user selects by such as pressing a rectangular frame image FR11 displayed in the portion of the face area of ​​a person HM11 with a finger. This selection operation, in step S14, so that the person rectangular frame image FR11 is displayed HM11 is selected as the audio object. Then, in step S15, the rectangular frame image FR11 displayed in the face area of ​​a person HM11 selected is highlighted.
[0087]
 In this example, the rectangular frame image FR11 is depicted by a solid line, it indicates that the rectangular frame image FR11 is highlighted. In contrast, the rectangular frame image FR12 and rectangular frame image FR13 person that is not selected is depicted with a dotted line, is not highlighted, that is, it indicates that there is a normal display.
[0088]
 Thus the rectangular frame image FR11 is selected, a rectangular frame image FR11 from the state of normal display depicted by dotted lines, the display state is changed to the state in which the highlighted drawn in solid lines.
[0089]
 Check this way a rectangular frame image of the selected person by the device user, by displaying a different display format other rectangular frame image, device user, whether its own selection operation has been correctly reflected easily can do.
[0090]
 Note that, although described with respect to an example in which the rectangular frame image of the selected person is highlighted, the display is not limited thereto, and the rectangular frame image of the selected person, and the other rectangular frame image different it suffices to be displayed in the form. For example a rectangular frame image of the selected person, or to display the other rectangular frame images in different colors and shapes, may be or to blink.
[0091]
 Further, the face recognition unit 71, holds the specific person of the dictionary in advance family members, when a particular person is detected in the face recognition section 71, a person selecting section 74 is device user's selection operation it may be selected that particular person as an audio object without.
[0092]
 In such a case, the person selecting section 74 from the face recognition unit 71 via the face recognition result display control unit 72 is also supplied information indicating a specific person. Then, at the time the image and the rectangular frame image of the content to the video display unit 73 is displayed, the image display unit 73 under the control of the person selecting section 74, to highlight the rectangular frame image of a specific person.
[0093]
 Thus, in advance particular person by leaving to be selected as the audio object, it is possible to save the labor of selection operation by the device user. In this case, in response to the operation of the subsequent device user, it is possible to make selection of a particular person is released.
[0094]
 Further, the rectangular frame image superimposed on the image of the content, for example similar to the example described in reference 1, the movement and the person, the movement of the imaging object region, i.e. automatically with the like change the shooting direction processing, such as to follow is performed.
[0095]
 In a typical camcorder and digital still cameras, but are provided with a mechanism to match the focus in a region touched by the touch panel, it is also possible to perform the selection process and the focus target selection processing of audio objects at the same time . Once audio object selected is made as similarly to the automatic following of the aforementioned rectangular frame images, to automatically follow with the movement of the movement and the imaging region of a person.
[0096]
 Returning to the flowchart of FIG. 3, the person selecting section 74 in step S16, based on the supplied face position information from the face recognition result display control unit 72, the spatial position of the selected audio object in step S14 generating object position information indicating.
[0097]
 Generally, MPEG-H Part 3: object position information of the audio objects in each standard typified 3D audio standards, are encoded in a spherical coordinate system shown in FIG. 5, for example.
[0098]
 In Figure 5, passing through the origin O, which is mutually perpendicular X-axis, Y-axis, and Z-axis and the axis of the three-dimensional orthogonal coordinate system. For example, in three-dimensional orthogonal coordinate system, the position of the audio object OB11 in space is the X-coordinate indicating the position of the X-axis direction x, y is a Y-coordinate indicating the position of the Y-axis direction, and the position of the Z-axis direction the by z is used a Z-coordinate showing represented (x, y, z) and.
[0099]
 Azimuth in a spherical coordinate system relative thereto, elevation, and radius are used in the position of the audio object OB11 in space represented.
[0100]
 Now, a straight line connecting the position of the audio object OB11 on the origin O and the space between the straight line r, is a straight line obtained by projecting the straight line r on the XY plane and straight line L.
[0101]
 In this case, the angle between the X axis and the straight line L theta is the azimuth angle that indicates the position of the audio object OB11. Further, the angle between the straight line r and the XY plane φ is the elevation showing the position of the audio object OB11, the length of the straight line r is the radius indicative of the position of the audio object OB11.
[0102]
 In the following, a represent a position of the audio object in the space in a spherical coordinate, azimuth indicating the position, elevation, and radius, respectively Position_azimuth, and be referred to as Position_elevation, and Position_radius.
[0103]
 Generally, it is difficult to measure the actual size or the distance from the photographer's object from the moving image. In the imaging apparatus such as a camcorder is also present which can measure the distance to the object, such as auto-focus mechanism at a plurality of points, but here the position of the subject using the face recognition process, i.e. the real space of the audio objects It will be described as an example when computing the position.
[0104]
 That is, below with reference to FIG. 6, the shape of the object in a general imaging apparatus will be described a method of measuring the subject distance when known.
[0105]
 In Figure 6, object 103 is photographed and an image pickup device having an imaging surface 101, the imaging unit comprising a lens 102.. Here, the imaging unit comprising the image sensor and lens 102 corresponding to the imaging unit 21, the object 103 corresponds to the person on the selected image as an audio object in the step S14, a person in the real space.
[0106]
 In this example, in the figure of the object 103, the object width is the longitudinal width has become a W2, in the drawing of the image of the object 103 on the imaging surface 101 at the time of photographing the object 103, the vertical width imaging the object width is the W1 is.
[0107]
 Further, in the figure in the space from the lens 102 to the object 103, the lateral distance as a subject distance d, in view of a space outside the imaging surface 101 to the lens 102, the lateral distance between the focal length D.
[0108]
 Incidentally, more specifically, in the drawing from the principal point of the lens 102 to the subject 103, the lateral distance is subject distance d, in the figure from imaging surface 101 to the principal point of the lens 102, the distance in the lateral direction the focal length D. Further, when the lens 102 is composed of a plurality of lenses, the principal point of the composite lens is the principal point of the lens 102.
[0109]
 The person selecting section 74, the focal length D, such as the size (size) and number of pixels of the imaging plane 101, information related to the imaging unit 21 is previously held.
[0110]
 Also, a person selecting section 74, and the face position information supplied from the face recognition result display control unit 72, the information related to the imaging unit 21 held in advance, the position and size of the object on the image, the imaging plane 101 it is possible to identify the relationship between the position and size of the object of the above (the image of). Thus, the person selecting section 74, it is possible to obtain a photographed object width W1 from information about the face position information and the imaging unit 21.
[0111]
 People selecting unit 74, the focal distance D is a known physical quantity, based on the imaging subject widths W1, and the subject width W2 calculates the following equation (1), and calculates the object distance d is an unknown physical quantity.
[0112]
[Number 1]

[0113]
 Although the subject width W2 is strictly an unknown physical quantity, regarded as the variation in the size of the face of each human sufficiently smaller than the object distance d, the average human face size of an object width it can be used as W2. That can be considered an object width W2 and known physical quantity.
[0114]
 For example, the width of a human face is about 16 centimeters average value, since the length of the face is about 23.5 cm are average values, it is possible to use those values ​​as the subject widths. Average length of these faces width and face will vary somewhat with gender differences, age and the like, for example, from the detected face from the image content, the image recognition or the like, sex and the detected person, since it is possible to estimate the approximate age, it may be corrected object width using gender and age of the average value.
[0115]
 Also, in the calculation of the object distance d, using both either one, or vertical width and horizontal width of the longitudinal width (length) and lateral width of a human face as an object can. For the sake of simplicity of explanation, the vertical width of a human face is used as an object width W2, it will be described an example in which only the subject width W2 is used the object distance d is calculated.
[0116]
 In Figure 6, the main point position of the lens 102 corresponds to the origin O shown in FIG. 5, the vertical and horizontal direction in FIG. 6, corresponding to the Z-axis direction and the X-axis direction in FIG. 5, respectively. In particular, in FIG. 6, the right direction corresponds to the front direction, that is the positive direction of the X-axis direction in FIG.
[0117]
 Thus, the object distance d calculated by the calculation described above, viewed from the origin O in the three-dimensional orthogonal coordinate system shown in FIG. 5 shows the distance to the person selected as the audio object in the X-axis direction.
[0118]
 Further, in the space, the same applies to the position of the face of a person is selected as the audio object, by using the average size of the face as a known physical quantity can be determined.
[0119]
 For example, the image of the content displayed on the video display unit 73 as shown in FIG. 7, the person HM11 is selected as the audio object, the rectangular frame image FR11 surrounding the area of ​​the face of the person HM11 is superimposed to. Note that portions corresponding to the case in FIG. 4 in FIG. 7 are denoted by the same reference numerals, and description thereof will be omitted as appropriate.
[0120]
 In Figure 7, the face portion of the person HM11 selected as audio object corresponds to the object 103 shown in FIG.
[0121]
 7, the center position of the image of the contents position A11 is 'a, the center position of the rectangular frame image FR11 displayed position A12 for the person HM11' and. For example, in the person selecting section 74, it is possible to obtain the position A12 'from the face position information supplied from the face recognition result display control unit 72.
[0122]
 Now, in the drawing on the image of the content, the longitudinal and transverse directions, respectively and the Z-axis direction (Z direction) and the Y-axis direction (Y-direction). The position on the content image A11 position of 'position A12 in viewed from the Y-axis direction' 'and, the position A11' image object horizontal position Y1 as viewed from positions in the Z axis direction A12 'of the position image object vertical position Z1 to' to.
[0123]
 In contrast, the center position of the imaging plane 101 and the position A11, and the position A12 of the position corresponding to the position A12 'in the imaging plane 101 on. Furthermore the position of the position A12 in position A11 viewed from the Y axis direction on the imaging plane 101 on the imaging plane object horizontal position Y1, the position of the position A12 in position A11 viewed from the Z-axis direction between the imaging plane object vertical position Z1.
[0124]
 Here, Y-axis and the Z-axis direction in the imaging plane 101 is a front direction and the vertical direction takes in FIG. The person selecting section 74, 'it is possible to obtain a, its position A12' position A12 from the face position information supplied from the face recognition result display control unit 72 images the object horizontal position from the Y1 'and the image object vertical position Z1' it can be determined.
[0125]
 Furthermore, the person selecting section 74, information indicating the size of the imaging plane 101 (sizes) are held in advance as known information from the image object horizontal position Y1 'and the image object vertical position Z1', the imaging surface subject it is possible to obtain the horizontal position Y1 and the imaging plane object vertical position Z1. In other words, the imaging plane object horizontal position Y1 and the imaging plane object vertical position Z1 is known physical quantity.
[0126]
 Thus, a person selecting unit 74, and these imaging surface object horizontal position Y1 and the imaging plane object vertical position Z1, the imaging subject widths W1 and subject width W2 Metropolitan described above, Y-axis of the object 103, i.e. a person HM11 in space it is possible to obtain the position of the direction and the Z-axis direction.
[0127]
 Here, the object 103, that is, the position of each of the Y-axis direction and the Z-axis direction in the space of the face of the person and the object horizontal position y and the object vertical position z. At this time, the object horizontal position y and the object vertical position z is an imaging plane object horizontal position Y1 and the imaging plane object vertical position Z1, the imaging subject widths W1 and subject width W2 Prefecture, each following equation (2) and ( 3) can be obtained by calculating the.
[0128]
[Number 2]

[0129]
[Number 3]

[0130]
 People selecting unit 74, by calculating the equation (2) and (3) to obtain the object horizontal position y and the object vertical position z of the face of the person selected as the audio object in the actual space.
[0131]
 Furthermore, a person selecting unit 74 calculates the subject horizontal position y and the object vertical position z obtained in this manner, the following equations on the basis of the object distance d described above (4) to (6), the actual and calculates the position in the spherical coordinate system of the face of the person.
[0132]
[Formula 4]

[0133]
[Formula 5]

[0134]
[Number 6]

[0135]
 That can be obtained by calculation of equation (4), the radius position_radius indicating the position in space of the person with the audio object.
[0136]
 Also be obtained by calculation of equation (5), the azimuth angle position_azimuth indicating the position in space of the person with the audio object. Furthermore, it can be obtained by calculation of equation (6), the elevation position_elevation in space of the person with the audio object.
[0137]
 Here, atan2 in equation (5) (y, d) is defined by the following equation (7).
[0138]
[Number 7]

[0139]
 However, the object distance d, as shown in the following equation (8), the minimum value d min is assumed to be limited by the. That is, a person selecting unit 74, the absolute value of the object distance d obtained by the equation (1) above the minimum value d min at the time is less than the minimum value d min for the used as the value of the object distance d.
[0140]
[Number 8]

[0141]
 People selecting unit 74, the above process by the determined azimuth Position_azimuth, represented by elevation Position_elevation, and radius Position_radius, information indicating the position in space of the person selected as the audio object and the object position information.
[0142]
 In this way the person selecting section 74, based on the result of face recognition in the face recognition section 71, it is possible to obtain the object position information. In general, however, or cause a change in light in the imaging unit 21, since the low-level noise by camera shake or the like or generation, it is desirable that the smoothing filter or the like to the object position information.
[0143]
 People selecting unit 74 supplies the object location information obtained in the sound source separation unit 27 and the audio encoding unit 28, the process proceeds to step S17. In the process described in the above, when a plurality of audio objects is selected, the object position information is generated for each of those audio objects.
[0144]
 In step S17, the sound source separation unit 27, based on the supplied object position information from the person selection unit 74 performs sound source separation for the supplied audio signal from the sound pickup unit 26, the audio object signal and the background sound of the audio signal It is separated into the signal.
[0145]
 For example, where the spatial position of the audio object, or the sound source position is a position indicated by the object position information, already known.
[0146]
 Therefore, for example, the sound source separation unit 27 by using the fixed beamforming technique, the audio signal, the audio object signal is a signal of audio objects, is separated into a background sound signal and a sound source signal of other than the audio objects.
[0147]
 Specifically, for example, when sound pickup unit 26 is a microphone array constituted by a plurality of microphones, the audio signal obtained by the sound pickup unit 26, and the audio object signal by controlling the directivity it is possible to perform the separation into the background sound signal. That is, it is possible to extract from the audio signal, the audio object signal, the background sound signal.
[0148]
 In other words, the fixed beamforming, audio signal of the position of the video object selected as the audio object in space is extracted as an audio object signal from the audio signal. More specifically, sound signals arriving from the direction of image objects selected as the audio object in space is extracted as an audio object signal from the audio signal. Then, from the audio signal, the signal of the sound other than the voice of the audio object is extracted as the background sound signal.
[0149]
 Incidentally, the method of sound source separation in the sound source separation unit 27 is not limited to the fixed beamforming, for example, using the technique described in JP 2010-233173 JP may be the any method.
[0150]
 Also, here, a video sound processing apparatus 61 has been been described on the assumption that it is a portable device, such as a camcorder, the present technology is not limited to this, television conference system or studio recording etc., it can also be applied to some extent large-scale systems. In such a case, if a large microphone arrays as sound pickup unit 26, it is possible to improve the separation performance of sound sources.
[0151]
 It is also possible to use a plurality of microphones, including a directional microphone, for example, as a sound pickup section 26. In such a case, by changing the direction of the sound source separation section 27 is shown the orientation of the directional microphone by the object position information, it is possible to obtain an audio object signal from the directional microphone, a background sound signal from another microphone it is possible to obtain. That is, it is possible to extract the audio object signal by a directional microphone, it is possible to extract a background sound signal by another microphone.
[0152]
 In step S18, the audio encoding unit 28 encodes the object position information supplied from the person selection unit 74, and an audio object signal and the background sound signal supplied from the sound source separation unit 27, obtained as a result audio It supplies a bit stream to the multiplexing unit 30.
[0153]
 For example, audio object signal and the background sound signal, after being coded independently, the audio bit stream is generated by multiplexing.
[0154]
 As the coding system of the audio object signal and the background sound signal, for example, linear PCM (Pulse Code Modulation) scheme or, "INTERNATIONAL STANDARD ISO / IEC 23008-3 First edition 2015-10-15 Information technology - High efficiency coding and media delivery in heterogeneous environments - Part 3: 3D audio "(hereinafter, referred to as reference 2) irreversible compression method may be used as described in.
[0155]
 Furthermore, the object position information is generally a data called metadata, for each predetermined time interval, are encoded in a format as shown in FIG. 8, for example.
[0156]
 Figure 8 is a diagram showing an example of the syntax of the metadata including the object position information (coding format). In the metadata shown in FIG. 8, "num_objects" indicates the number of audio objects included in the audio bit stream.
[0157]
 Further, "tcimsbf" indicates approximately a is, the complement of 2 code bits beginning of the "Two's complement integer, most significant (sign) bit first". "Uimsbf" is "Unsigned integer, most significant bit first" is an abbreviation of the most significant bit is an integer without leading sign.
[0158]
 Furthermore, "position_azimuth [i]", "position_elevation [i]", and "position_radius [i]" indicates the object position information of the i-th audio object included in the audio bit stream, respectively.
[0159]
 Specifically, "position_azimuth [i]" indicates an azimuth Position_azimuth the position of the audio object in the spherical coordinate system, "position_elevation [i]" is indicated elevation Position_elevation position of the audio object in the spherical coordinate system there. Further, "position_radius [i]" indicates the distance to the position of the audio object in the spherical coordinate system, i.e. the radius Position_radius.
[0160]
 Further, "gain_factor [i]" indicates the gain information of the i-th audio object included in the audio bit stream.
[0161]
 Thus, the person selecting section 74, is generated object position information is meta data of the audio object, the audio encoding unit 28, the object position information and gain information are encoded as metadata.
[0162]
 Returning to the flowchart of FIG. 3, in step S19, the video encoding unit 29 encodes the video signal supplied from the imaging unit 21, and supplies the resulting video bit stream to the multiplexing unit 30.
[0163]
 Incidentally, the encoding system of the video signal, in addition to video coding schemes that are well known, such as MPEG-HEVC or MPEG-AVC, it is possible to make use of the various video coding scheme.
[0164]
 In step S20, the multiplexer 30, the video bit stream supplied from the video encoder unit 29, an audio bit stream supplied from the audio encoding unit 28, together with such as system information for performing these synchronization multiplexed, to generate an output bitstream. Multiplexer 30 outputs the output bit stream obtained in this way, the encoding process ends.
[0165]
 Audiovisual processing apparatus 61 as described above, performs face recognition on the video signal, and displays a rectangular frame image showing the results of face recognition with the image content. The video sound processing unit 61 selects the audio objects in accordance with the selection operation of the device user, and generates an object position information of the audio object.
[0166]
 Thus, it is possible to separate the signals of the desired object sound easier and accurate. That is, device user, whether with which video object (subject) the audio objects in the image content, can be selected simply and intuitively while viewing the display of the image display unit 73. Furthermore, it is possible to separate the signal of the desired audio objects by determining the object position information of the audio object selected, with higher accuracy.
[0167]
 If encoded as audio object signal audio Thus person that is selected can be changed and sound volume at the time of reproduction, the sound source position, sound quality, etc. for each audio object.
[0168]
 In the description given in the above, while the audiovisual processing apparatus 61 taking the image of the content, but device user has described an example of selecting a person to audio objects, selects a person to audio objects after shooting it may be so.
[0169]
 In such a case, at the time of shooting image of the content, for example a video signal obtained by shooting, and the audio signal obtained by sound pickup is recorded without being encoded. After shooting, when the reproduction of the content is instructed, the content is reproduced based on the video signal and audio signal, the encoding process described with reference to FIG. 3 is performed.
[0170]
 Then, above receives the output bit stream obtained by processing the reproduction of the content, in particular a method for reproducing sound of the content will be described.
[0171]
 The output bit streams with respect background sound signal contained in the audio bit stream obtained by demultiplexing, for example, reproduction of a so-called multi-channel stereo system, such as a conventional 2-channel and 5.1-channel is performed.
[0172]
 On the other hand, with respect to an audio object signals, rendering scheme of the audio objects in reference 2 has been described, the rendering is reproduced is performed in accordance with the rendering scheme.
[0173]
 More specifically, audio object signal to the speaker of the reproduction environment in a manner called VBAP (Vector Base Amplitude Panning) is reproduced is mapped.
[0174]
 This is indicated by the object position information, using the output of the nearest three speakers in a position in space of the audio object, a technique of localizing a sound spatial position of the audio object. In reference 2, the number of speaker it is also possible to localize a sound three at a but the number four or more speakers.
[0175]
 Further, in the above example has been described background sound signal to be encoded and decoded in a multi-channel stereo system, viewer space in the vicinity, called the reference is described in the literature 2 standard HOA (Higher Order Ambisonic) and it may be used such a manner as to express the orthogonal base of spherical harmonics.
[0176]
 Further, in the first embodiment, the device user, i.e. the photographer is in most cases to perform photographing to fit subject to be photographed person is always in the screen of the audio object extraction. However, as a countermeasure when the person is selected as the audio object has moved off the screen, and the like commonly known speaker identification technology may generate an object position information using.
[0177]

 In the first embodiment described above, the device user is displayed a person you want to pick up sound as an audio object on the video display unit 73 It had done the operation of selecting from the candidate.
[0178]
 However, some devices which reproduce the content by decoding the output bit stream obtained by the video audio processing device 61, for convenience of computing power, if it is not possible to perform the decoding process and rendering process for all the audio object signal there is.
[0179]
 Meanwhile, it has been proposed a technique for adding processing priority for each audio object in the coding apparatus side "priority" as metadata in the bitstream, this approach has been adopted in references 2 standard.
[0180]
 Therefore, in the second embodiment to which the present technology is applied, together with the device user selects the candidates displayed a person you want to pick up sound as an audio object on the video display unit 73, the processing of the selected person and to be able to set the priority.
[0181]
 In this case, the metadata object position information is stored, for example, as shown in FIG. Figure 9 is a diagram showing an example of the syntax of the metadata.
[0182]
 The metadata shown in FIG. 9, in addition to the information included in the metadata shown in FIG. 8 includes a further i th audio object, that is, the processing priority of the audio object signal "object_priority [i]" .
[0183]
 In this example, the processing priority object_priority [i] is a 3-bit data, are made to be able to take a value from 0 to 7.
[0184]
 For example if the value of the processing priority object_priority [i] is 0, the processing priority of the audio object has the lowest, when the value of the processing priority object_priority [i] is 7, the processing priority of the audio objects highest .
[0185]
 Now, for example, the value of the processing priority object_priority [i] is the output bit stream, respectively 7,3, and an audio object signals of the three audio objects is 0 is assumed to be supplied to the device on the playback side. Further, in the reproducing side of the apparatus, and it is impossible to render all three audio objects.
[0186]
 In such a case, for example, the reproducing side of the apparatus, the audio object value of the processing priority object_priority [i] is 0, or a simple rendering at the expense of sound quality, perform rendering of the audio objects it can be made to no or. Thus, it is possible to realize the real-time reproduction of the content at low processing ability regeneration equipment.
[0187]
 If the processing priority is stored together with the object position information in the meta data as described above, the person selecting section 74 of the video audio processing device 61, on the basis of the signals corresponding to the operation of the equipment user to be supplied from the touch panel 81 , as well as select a person to be an audio object, and sets the processing priority. The person selecting section 74 supplies supplies object position information to the sound source separation section 27 and the audio encoding unit 28, the processing priority obtained in the audio coding unit 28.
[0188]
 Thus, in the audio coding unit 28, the processing priority not only object position information is also encoded, their object position information and the processing priority audio bitstream metadata is stored containing is produced.
[0189]
 If the processing priority is specified by the equipment user to, in step S15 in FIG. 3, for example, as shown in FIG. 10, the rectangular frame image is displayed together with the image content to the video display unit 73. In FIG. 10, portions corresponding to the case in FIG. 4 are given the same reference numerals, and description thereof will be omitted as appropriate.
[0190]
 10, to the person HM11 or person HM13 three people are in the imaging region, shows a state in which equipment user sets the processing priority.
[0191]
 In this example, the rectangular frame image FR11 to rectangular frame image FR13 are highlighted, the person HM11 to persons HM13 corresponding to those of the rectangular frame image, is selected as the audio object.
[0192]
 Further, the rectangular frame image FR11 vicinity of the person HM11, the character information "P = 7" indicating the processing priority set for the person HM11 are displayed, the value of the processing priority of the person HM11 is most It has been shown to be highly 7.
[0193]
 Similarly, the rectangular frame image FR13 vicinity of the person HM13, the person HM13 are displayed character information "P = 0" indicating the processing priority set for, the value of the processing priority of the person HM13 it has been shown that the lowest is 0.
[0194]
 Further, in the example of FIG. 10, in the vicinity of the rectangular frame image FR12 person HM12, process priority setting image CF11 for setting the processing priority of the person HM12 is displayed.
[0195]
 The processing priority setting image CF11, for example equipment user is displayed on contact portions of the rectangular frame image FR12 with a finger.
[0196]
 Further, in the processing priority setting image CF11, processing priority value "0" to "7" is written frame is displayed and equipment user marked value of desired processing priority by touching the frame with a finger, it is possible to set the processing priority of the person HM12.
[0197]
 For example, device user presses the frame "3" is written, a person selecting section 74, based on a signal supplied from the touch panel 81 in accordance with the operation of the device user, processing of a person HM12 as audio objects the value of priority to "3".
[0198]
 Incidentally, the method of determining the processing priority is not limited to the above-described method may be any method. For example the processing priority of the person corresponding to the selected rectangular frame image by the device user and 7, the processing priority of the person rectangular frame image has not been selected is a simple way as a 0 it may be.
[0199]
 Further, for example, in a case like that all persons detected by the face recognition is an audio object, the processing priority of the person other than the person selected as the audio object in step S14 of the encoding process (video object), the the size of the face in the image of a person, i.e. may be selected from among the 6 following values ​​according to the size of the rectangular frame image.
[0200]
 In the encoding process, the person selecting section 74 in step S16, generates the object position information of the audio objects in accordance with the operation of the device user, determines the processing priority of each audio object. In other words, as the metadata of the audio object, the object position information and the processing priority is generated.
[0201]
 The person selecting section 74 supplies the object location information to the sound source separation section 27 supplies the object position information and the processing priority to the audio encoding unit 28.
[0202]
 Further, in step S18, the audio encoding unit 28, the object position information, processing priority, performs encoding of the audio object signals, and background sound signal, generating an audio bit stream.
[0203]
 By generating the processing priority of audio objects as described above, the reproducing side of the apparatus, it is possible to selectively change the process or the like for the audio objects in accordance with its computing power and the like. Thus, for example, to reproduce the content in real time, it is possible to realize a proper content reproduction.
[0204]

 In addition, when rendering the audio objects, rendering is performed by VBAP considered that there is a point sound source in the position indicated by the object position information.
[0205]
 In practice, however, the object has a size, better regarded as sound from the surface with a certain area that is occurring is natural. Therefore, spread in addition to the metadata data called (hereinafter, referred to as spread information), by using the data during rendering, proposed a technique to make a region of the object, that is sound from the surface is generated It is, has been adopted by the references 2 standard.
[0206]
 In a third embodiment according to the present technology, when recognizing an image object, such as a human face, using a size of the rectangular frame is a recognition result of the image object, the spread in the audiovisual processing apparatus 61 side It was to be able to automatically set the information.
[0207]
 In this case, the person selecting section 74, on the basis of the signals corresponding to the operation of the equipment user to be supplied from the touch panel 81, as well as select a person to be an audio object, and sets the spread information. The person selecting section 74 supplies supplies object position information to the sound source separation section 27 and the audio encoding unit 28, a spread information obtained audio encoding unit 28. As in the second embodiment, the person selecting section 74, may be further be processed priority generated.
[0208]
 For example with the object position information, if the processing priority and spreads information included in the metadata, the metadata is shown in FIG. 11. Figure 11 is a diagram showing an example of the syntax of the metadata.
[0209]
 The metadata shown in FIG. 11, in addition to the information included in the metadata shown in FIG. 9 includes a further i th audio object, i.e. spread information of the audio object signal "spread [i]".
[0210]
 In this example, the spread information spread [i] is the 8-bit data is angle information representing the area of ​​the audio objects in the space, i.e. the angle information indicating the spatial spread of the region of the audio objects.
[0211]
 The instrument is a reproducing apparatus for reproducing side to receive the output bit stream, such spread information spread [i] is to render used is performed. It will be described later specific method for calculating the spread information spread [i].
[0212]
 Here, the method of using the playback apparatus spread information in compliance with references 2 standard will be described with reference to FIG. 12.
[0213]
 For example with respect to the origin O of the space, the audio object at the position indicated by the vector p0 is located. Here, vector p0 is a vector whose starting point is the origin O, the position indicated by the vector p0 is the position indicated by the object position information. In the following, it is assumed that the position indicated by the vector p0, also referred to as position p0.
[0214]
 In such a case, the person selecting section 74, the angle information for the vector p0 is generated as spreadsheet information spread [i].
[0215]
 In reproduction apparatus, watches position p0 from the origin O to the front 18 of the angle given by the spread information spread [i] and ending at the position of the resulting circle C11 on the basis, starting from the origin O vector p1 through vector p18 in is obtained as a spread vector.
[0216]
 Circle C11 is a circle around the position p0, in this example, vector p1 and vectors p4 is a vector that is an end point positions on the circumference of the circle C11. In FIG. 12, the point within the circle C11 represents the end position of each spread vectors.
[0217]
 For example a vector p1 to the end point position on the circumference of the circle C11, the angle between the vector p0, the angle indicated by the spread information spread [i].
[0218]
 Incidentally, on the basis of the object position information and the spread information spread [i], for the method of calculating the vector p1 to the vector p18 is a spread vector, since they are described in detail, for example, in references 2, description thereof will be omitted .
[0219]
 When the spread vector is obtained, the reproducing apparatus, for each of these vectors p1 to vector p18 was determined as a spread vector, the gain of the audio object signals are obtained for each speaker, their gain is normalized by VBAP . The audio object signals normalized gain is multiplied is mapped to the speaker, the audio content is played.
[0220]
 By performing the rendering with this way the spread information, it is possible to represent the audio objects that spread in the spatial direction.
[0221]
 If the spread information is calculated in the person selecting section 74 of the video audio processing device 61, a person selecting unit 74, for example, as shown in FIG. 13, the recognition result of the image object by facial recognition, and more particularly the results of face recognition a rectangular frame image obtained from calculating the spread information.
[0222]
 In the example shown in FIG. 13, from the image of the content displayed on the image display unit 73, 2 people person HM31 and the person HM32 are detected. Further, the image of the content, together with the rectangular frame image FR31 is superimposed on the face portion of the person HM31, rectangular frame image FR32 is superimposed on the face portion of the person HM32.
[0223]
 In particular, in FIG. 13, the image of the content represents a state of being projected into the space, the center position of the rectangular frame image FR31, and the center position of the rectangular frame image FR32 is by the object position information of the corresponding audio object and has a position shown.
[0224]
 People selection unit 74, in step S16 in FIG. 3, for each audio object and generating object position information and the processing priority, also generates further spread information of each audio object. That is, the metadata of the audio object, the object position information, the processing priority, and spreads information is generated.
[0225]
 Specifically, for example, will be described a person HM31 as audio object, a person selecting section 74 and starting from the origin O in the space, the position indicated by the object position information, the end point that is the center position of the rectangular frame image FR31 determine the vector VB11. The person selecting section 74, a vector VB11 thus determined, a vector p0 described with reference to FIG. 12. Such vector VB11 can be obtained from the object position information generated for the person HM31.
[0226]
 Also, a person selecting unit 74, and starting from the origin O in space, in the figure in the rectangular frame image FR31, is on the right side, in the center position and FIG rectangular frame image FR31, longitudinal (Z-axis direction) It obtains a vector VB12 height of the end point position is the same, a vector p1 illustrates this vector VB12 with reference to FIG. 12.
[0227]
 Here, the end position of the vector VB12 on the space, using for example the object position information generated for the person HM31, rectangular frame image FR31, relationship between the image of the imaging plane 101 and the content, focal length D, and the object distance d and it can be obtained Te.
[0228]
 Furthermore, a person selecting section 74, a vector VB11 as vector p0 obtained in this way, the angle AG11 with vector VB12 as vector p1, is calculated as the spread information spread [i] for the human HM31. More specifically, a person selecting unit 74 calculates the spread information spread [i] by calculating the following equation (9).
[0229]
[Number 9]

[0230]
 In the expression (9), p0 and p1 represents the vector p0 and vector p1, respectively.
[0231]
 Similarly, a person selecting unit 74, and starting from the origin O in the space, the position indicated by the object position information about the person HM32, i.e. obtains a vector VB13 to ending the center position of the rectangular frame image FR32, a vector p0 to.
[0232]
 Furthermore, the person selecting section 74, and starting from the origin O in space, in the figure in the rectangular frame image FR32, is on the right side, in the center position and FIG rectangular frame image FR32, vertical height same position sought vector VB14 to end point is, the vector p1.
[0233]
 The person selecting section 74 obtains the angle AG12 of vector VB13 and vector VB14 by calculating the equation (9), and spreads the information spread [i] for the human HM32.
[0234]
 Here, I explained how to calculate the spread information using only vector p0 and vector p1. However, when there is a margin in computing power such as audiovisual processing apparatus 61, in the figure in the rectangular frame images, respectively spread vectors upper, lower, left, and right end and end position, the angle between the vector p0 the calculated may be the average value of their angle to used as the spread information spread [i].
[0235]
 When the spread information is calculated as described above, a person selecting section 74 supplies the object location information to the sound source separation unit 27, the object position information, the processing priority, and the spread information to the audio encoding unit 28 supplies.
[0236]
 Further, in step S18, the audio encoding unit 28, the object position information, processing priority, performs encoding of the spread information, audio object signals, and background sound signal, generating an audio bit stream.
[0237]
 By generating a spread information of audio objects as described above, the device user without the designation relating to spread information, the reproducing side of the device, is possible to perform content playback representing the spatial extent of the audio objects become able to.
[0238]
 The series of processes described above can be executed by hardware or can be executed by software. When executing the series of processing by software, a program constituting the software is installed into a computer. Here, the computer includes a computer incorporated in dedicated hardware, by installing various programs, which can execute various functions include, for example, such as a general-purpose personal computer.
[0239]
 Figure 14 is a block diagram showing a configuration example of hardware of a computer that executes the series of processes described above.
[0240]
 In the computer, CPU (Central Processing Unit) 501, ROM (Read Only Memory) 502, RAM (Random Access Memory) 503 are connected to each other via a bus 504.
[0241]
 The bus 504 is further output interface 505 is connected. Output interface 505, an input unit 506, output unit 507, recording unit 508, a communication unit 509, and a drive 510 are connected.
[0242]
 Input unit 506 includes a keyboard, a mouse, a microphone, made of an imaging device. The output unit 507 includes a display and a speaker. Recording unit 508, a hard disk and a nonvolatile memory. Communication unit 509 including a network interface. Drive 510 drives a magnetic disk, an optical disk, a magneto-optical disk, or a removable recording medium 511 such as a semiconductor memory.
[0243]
 Series In the computer configured as described above, CPU 501 is, for example, a program recorded in the recording unit 508 via the input-output interface 505 and the bus 504 and executes the loaded into RAM 503, the above-mentioned processing of is performed.
[0244]
 Program computer (CPU 501) is executed, for example, can be provided by being recorded in a removable recording medium 511 as a package medium or the like. Further, the program may be provided via a local area network, the Internet, or digital satellite broadcasting, a wired or wireless transmission medium.
[0245]
 In the computer, by mounting the removable recording medium 511 into the drive 510, it can be through the input-output interface 505, installed in the recording unit 508. The program via a wired or wireless transmission medium and received by the communication unit 509, can be installed in the recording unit 508. Alternatively, the program may be in the ROM502 and the recording unit 508 installed in advance.
[0246]
 The program which the computer executes may be a program in which processes are performed in time series in the order described herein, at a necessary timing such as when the parallel or call was made processing may be a program to be carried out.
[0247]
 Further, embodiments of the present technology is not limited to the embodiments described above, but various modifications are possible without departing from the scope of the present disclosure.
[0248]
 For example, the present technology, sharing one function by a plurality of devices via a network, it is possible to adopt a configuration of cloud computing which processes jointly.
[0249]
 Further, each step described in the above flowcharts may be executed by one device, it can be performed by allocating a plurality of apparatuses.
[0250]
 Further, when a plurality of processes are included in one step, the plurality of processes included in the one step may be executed by one device, it can be performed by allocating a plurality of apparatuses.
[0251]
 Additionally, the present technology may also be configured as follows.
[0252]
(1)
 a display control unit for displaying an image object based on the image signal,
 from among one or more of the video object, and the object selection unit for selecting a predetermined said video objects,
 wherein selected by the object selection portion an extraction unit that extracts an audio signal of the video object as an audio object signals
 audiovisual processing apparatus comprising a.
(2)
 the extraction unit extracts the audio object signal from the audio signal
 audiovisual processing apparatus according to (1).
(3)
 the extraction unit is configured from the audio signal to extract a signal other than the audio object signal of said selected the video object as a background sound signal
 audiovisual processing apparatus according to (2).
(4)
 the object selection unit generates an object position information indicating a position in the space of the said selected video object,
 wherein the extraction section, based on the object position information, and extracts the audio object signal
 (1) to the audiovisual processing apparatus according to any one of (3).
(5)
 the extracting unit extracts the audio object signal by the sound source separation using the object position information
 Audiovisual processing apparatus according to (4).
(6)
 the extraction unit performs fixed beamforming as the sound source separation
 audiovisual processing apparatus according to (5).
(7)
 based on said video signal, further comprising a video object recognition unit recognizes the image object,
 wherein the display control unit is configured together with the video object, to display an image based on the recognition result of the image object
 (1) or audiovisual processing apparatus according to any one of (6).
(8)
 the video object recognition unit recognizes the video object by the face recognition
 audiovisual processing apparatus according to (7).
(9)
 The display control unit, as the image, the display frame in the region of the video object
 audiovisual processing apparatus according to (7) or (8).
(10)
 the object selection unit, in response to a selection operation by the user, selects the image object
 audiovisual processing apparatus according to any one of (1) to (9).
(11)
 wherein the object selection unit generates the metadata of the selected the video object
 (1) to the audiovisual processing apparatus according to any one of (10).
(12)
 wherein the object selection unit, said as metadata, to generate an object position information indicating a position in the space of the said selected video object
 audiovisual processing apparatus according to (11).
(13)
 wherein the object selection unit, as the metadata, generates the processing priority of the selected the video object
 audiovisual processing apparatus according to (11) or (12).
(14)
 wherein the object selection unit, as the metadata, and generates a spread information indicating a spatial spread in the region of the said selected video object
 audiovisual according to any one of (11) to (13) processing apparatus.
(15)
 further comprising an audio encoding unit for encoding an audio object signal and the metadata
 audiovisual processing apparatus according to any one of (11) to (14).
(16)
 and a video encoding unit for encoding the video signal,
 the video bit stream obtained by the encoding of the video signal, and the audio bit stream obtained by the encoding of the audio object signal and the metadata a multiplexer for multiplexing
 Audiovisual processing apparatus according to, further comprising a (15).
(17)
 performs imaging further comprises an imaging unit for obtaining the video signal
 the video sound processing apparatus according to any one of (1) to (16).
(18)
 conducted sound pickup further comprises a sound pickup for obtaining the audio signal
 (1) to the audiovisual processing apparatus according to any one of (17).
(19)
 and a display control step of displaying the image object based on the image signal,
 from among one or more of the video object, and the object selecting step of selecting a predetermined said video objects,
 wherein selected by the object selecting step an extraction step of extracting the audio signal of the video object as an audio object signals
 audiovisual processing method comprising.
(20)
 and a display control step of displaying the image object based on the image signal,
 from among one or more of the video object, and the object selecting step of selecting a predetermined said video objects,
 wherein selected by the object selecting step an extraction step of extracting the audio signal of the video object as an audio object signal
 program for executing the processing including a computer.
DESCRIPTION OF SYMBOLS
[0253]
 11 audiovisual processing apparatus, 22 a video object recognition unit, a video object recognition result display control unit 23, 24 display unit, 25 an object selection section, 26 sound pickup unit 27 of the sound source separation section, 28 an audio encoding unit, 71 face recognition parts, 72 face recognition result display control unit, 73 image display unit, 74 a person selecting unit, 81 a touch panel

claims
[Requested item 1]
 A display control unit for displaying an image object based on the image signal,
 from among one or more of the video object, and the object selection unit for selecting a predetermined the video object,
 the image object selected by the object selection portion an extraction unit that extracts an audio signal as an audio object signals
 audiovisual processing apparatus comprising a.
[Requested item 2]
 The extraction unit extracts the audio object signal from the audio signal
 audiovisual processing apparatus according to claim 1.
[Requested item 3]
 The extraction unit is configured from the audio signal to extract a signal other than the audio object signal of said selected the video object as a background sound signal
 audiovisual processing apparatus according to claim 2.
[Requested item 4]
 The object selection unit generates an object position information indicating a position in the space of the said selected video object,
 wherein the extraction section, based on the object position information, and extracts the audio object signal
 according to claim 1 audiovisual processing apparatus according to.
[Requested item 5]
 The extraction unit extracts the audio object signal by the sound source separation using the object position information
 audiovisual processing apparatus according to claim 4.
[Requested item 6]
 The extractor performs the fixed beamforming as the sound source separation
 audiovisual processing apparatus according to claim 5.
[Requested item 7]
 Based on the video signal, further comprising a video object recognition unit recognizes the image object,
 wherein the display control unit, together with the image object, to display an image based on the recognition result of the image object
 according to claim 1 audio-video processing apparatus.
[Requested item 8]
 The video object recognition unit recognizes the video object by the face recognition
 audiovisual processing apparatus according to claim 7.
[Requested item 9]
 The display controller as the image, and displays a frame region of the image object
 audiovisual processing apparatus according to claim 7.
[Requested item 10]
 The object selection unit, in response to a selection operation by the user, selects the image object
 audiovisual processing apparatus according to claim 1.
[Requested item 11]
 The object selection unit generates the metadata of the selected the video object
 audiovisual processing apparatus according to claim 1.
[Requested item 12]
 The object selection unit, as the metadata, and generates an object position information indicating a position in the space of the said selected video object
 audiovisual processing apparatus according to claim 11.
[Requested item 13]
 The object selection unit, as the metadata, generates the processing priority of the selected the video object
 audiovisual processing apparatus according to claim 11.
[Requested item 14]
 The object selection unit, said as metadata, and generates a spread information indicating the spatial spread of the selected region of the video object has been
 audiovisual processing apparatus according to claim 11.
[Requested item 15]
 Further comprising an audio encoding unit for encoding the audio object signal and the metadata
 audiovisual processing apparatus according to claim 11.
[Requested item 16]
 A video encoding unit, for encoding the video signal
 multiplexing said video bit stream obtained by the encoding of the video signal, and an audio bit stream obtained by the encoding of the audio object signal and the metadata a multiplexing unit for
 audiovisual apparatus according to claim 15, further comprising a.
[Requested item 17]
 Further comprising an imaging unit for obtaining the video signal by performing photographing
 audiovisual processing apparatus according to claim 1.
[Requested item 18]
 Further comprising a sound pickup for obtaining the audio signal by performing sound collection
 audiovisual processing apparatus according to claim 1.
[Requested item 19]
 A display control step of displaying the image object based on the image signal,
 from among one or more of the video object, and the object selecting step of selecting a predetermined the video object,
 the image object selected by the object selecting step an extraction step of extracting an audio signal as an audio object signals
 audiovisual processing method comprising.
[Requested item 20]
 A display control step of displaying the image object based on the image signal,
 from among one or more of the video object, and the object selecting step of selecting a predetermined the video object,
 the image object selected by the object selecting step an extraction step of extracting an audio signal as an audio object signal
 program for executing the processing including a computer.

Documents

Application Documents

# Name Date
1 201817044157.pdf 2018-11-23
2 201817044157-TRANSLATIOIN OF PRIOIRTY DOCUMENTS ETC. [23-11-2018(online)].pdf 2018-11-23
3 201817044157-STATEMENT OF UNDERTAKING (FORM 3) [23-11-2018(online)].pdf 2018-11-23
4 201817044157-PROOF OF RIGHT [23-11-2018(online)].pdf 2018-11-23
5 201817044157-PRIORITY DOCUMENTS [23-11-2018(online)].pdf 2018-11-23
6 201817044157-POWER OF AUTHORITY [23-11-2018(online)].pdf 2018-11-23
7 201817044157-FORM 1 [23-11-2018(online)].pdf 2018-11-23
8 201817044157-DRAWINGS [23-11-2018(online)].pdf 2018-11-23
9 201817044157-DECLARATION OF INVENTORSHIP (FORM 5) [23-11-2018(online)].pdf 2018-11-23
10 201817044157-COMPLETE SPECIFICATION [23-11-2018(online)].pdf 2018-11-23
11 201817044157-OTHERS-291118.pdf 2018-12-06
12 201817044157-Correspondence-291118.pdf 2018-12-06
13 abstract.jpg 2018-12-26
14 201817044157-FORM 3 [12-04-2019(online)].pdf 2019-04-12
15 201817044157-FORM 18 [14-04-2020(online)].pdf 2020-04-14
16 201817044157-OTHERS [18-10-2021(online)].pdf 2021-10-18
17 201817044157-FER_SER_REPLY [18-10-2021(online)].pdf 2021-10-18
18 201817044157-FER.pdf 2021-10-18
19 201817044157-DRAWING [18-10-2021(online)].pdf 2021-10-18
20 201817044157-CORRESPONDENCE [18-10-2021(online)].pdf 2021-10-18
21 201817044157-CLAIMS [18-10-2021(online)].pdf 2021-10-18
22 201817044157-ABSTRACT [18-10-2021(online)].pdf 2021-10-18
23 201817044157-US(14)-HearingNotice-(HearingDate-01-02-2024).pdf 2023-12-11
24 201817044157-FORM-26 [29-01-2024(online)].pdf 2024-01-29
25 201817044157-Correspondence to notify the Controller [29-01-2024(online)].pdf 2024-01-29
26 201817044157-Written submissions and relevant documents [16-02-2024(online)].pdf 2024-02-16
27 201817044157-PETITION UNDER RULE 137 [16-02-2024(online)].pdf 2024-02-16
28 201817044157-FORM 3 [16-02-2024(online)].pdf 2024-02-16
29 201817044157-PatentCertificate15-04-2024.pdf 2024-04-15
30 201817044157-IntimationOfGrant15-04-2024.pdf 2024-04-15

Search Strategy

1 search4157E_23-03-2021.pdf

ERegister / Renewals

3rd: 18 Jun 2024

From 17/05/2019 - To 17/05/2020

4th: 18 Jun 2024

From 17/05/2020 - To 17/05/2021

5th: 18 Jun 2024

From 17/05/2021 - To 17/05/2022

6th: 18 Jun 2024

From 17/05/2022 - To 17/05/2023

7th: 18 Jun 2024

From 17/05/2023 - To 17/05/2024

8th: 18 Jun 2024

From 17/05/2024 - To 17/05/2025