Sign In to Follow Application
View All Documents & Correspondence

Information Processing Device, Information Processing Method, And Information Processing Program

Abstract: An information processing device (100) according to the present disclosure comprises an acquisition unit (141) which acquires a first image which includes an image of a user's ear, and a computation unit (142) which, using a learned model which is learned so as to output a head-related transfer function corresponding to the ear when the image which includes the image of the ear is inputted, computes a head-related transfer function corresponding to the user on the basis of the first image acquired by the acquisition unit (141).

Get Free WhatsApp Updates!
Notices, Deadlines & Correspondence

Patent Information

Application #
Filing Date
05 March 2021
Publication Number
33/2022
Publication Type
INA
Invention Field
ELECTRONICS
Status
Email
patents@remfry.com
Parent Application
Patent Number
Legal Status
Grant Date
2025-06-11
Renewal Date

Applicants

SONY CORPORATION
1-7-1, Konan, Minato-ku, Tokyo 1080075

Inventors

1. FUKUDA, Kazumi
c/o Sony Corporation, 1-7-1, Konan, Minato-ku, Tokyo 1080075
2. MAGARIYACHI, Tetsu
c/o Sony Corporation, 1-7-1, Konan, Minato-ku, Tokyo 1080075

Specification

Title of invention: Information processing device, information processing method, and information processing program
Technical field
[0001]
The present disclosure relates to an information processing device, an information processing method, and an information processing program. Specifically, it relates to calculation processing of a head-related transfer function.
Background technology
[0002]
Three-dimensional reproduction of sound images in headphones etc. by using head-related transfer function (hereinafter sometimes referred to as HRTF (Head-Related Transfer Function)) that mathematically represents how sound reaches the ear from the sound source technology is used.
[0003]
Since head-related transfer functions vary greatly among individuals, it is desirable to use head-related transfer functions generated for each individual when using them. For example, a three-dimensional digital model of the head (hereinafter referred to as a "3D model") is generated based on an image of the user's auricle, and the head-related transfer function of the user is calculated from the generated 3D model. technology is known.
prior art documents
patent literature
[0004]
Patent Document 1: US Patent No. 9544706
SUMMARY OF THE INVENTION
Problems to be Solved by the Invention
[0005]
According to the conventional technology, head-related transfer functions individually calculated for each user can be used for information processing, so the sense of localization of sound images can be enhanced.
[0006]
However, the conventional technology described above generates a 3D digital model based on an image taken by a user and calculates a head-related transfer function from the generated model, so the computational processing load is relatively large. In other words, the conventional technology described above is not very convenient because it is estimated that it takes a long time to provide the head-related transfer function to the user who has sent the image.
[0007]
Therefore, the present disclosure proposes an information processing device, an information processing method, and an information processing program that can improve user convenience in processing related to head-related transfer functions.
Means to solve problems
[0008]
In order to solve the above problems, an information processing apparatus according to one embodiment of the present disclosure includes an acquisition unit that acquires a first image including a video of a user's ear, and a Calculating a head-related transfer function corresponding to the user based on the first image acquired by the acquisition unit using a trained model trained to output a head-related transfer function corresponding to the ear and
Brief description of the drawing
[0009]
1 is a diagram showing an overview of information processing according to the first embodiment of the present disclosure; FIG.
2 is a conceptual diagram showing the overall flow of information processing according to the present disclosure; FIG.
3 is a diagram showing a configuration example of an information processing device according to the first embodiment of the present disclosure; FIG.
4 is a diagram showing an example of an ear image storage unit according to the present disclosure; FIG.
5 is a diagram showing an example of learning processing regarding an ear model according to the present disclosure; FIG.
6 is a diagram showing an example of learning processing regarding an ear parameter estimation model according to the present disclosure; FIG.
7 is a diagram showing an example of ear image generation processing according to the present disclosure. FIG.
8 is a diagram for explaining an ear parameter estimation model according to the present disclosure; FIG.
9 is a diagram showing an overview of the flow of model generation processing for HRTF; FIG.
10 is a diagram for explaining reconstruction of a 3D model according to the present disclosure; FIG.
FIG. 11 is a diagram for explaining details of model generation processing for HRTF;
12 is a diagram for explaining HRTF compression and decompression according to the present disclosure; FIG.
13 is a diagram showing a flow of HRTF estimation processing according to the present disclosure; FIG.
14 is a diagram showing a configuration example of a user terminal according to the first embodiment of the present disclosure; FIG.
15 is a diagram showing a flow of detection processing according to the present disclosure; FIG.
16 is a diagram showing a configuration example of an information processing device according to a second embodiment of the present disclosure; FIG.
17 is a hardware configuration diagram showing an example of a computer that realizes functions of an information processing apparatus; FIG.
MODE FOR CARRYING OUT THE INVENTION
[0010]
Below, embodiments of the present disclosure will be described in detail based on the drawings. In addition, in each of the following embodiments, the same parts are denoted by the same reference numerals, thereby omitting redundant explanations.
[0011]
(1. First Embodiment)
[1-1. Overview of information processing according to the first embodiment]
First, using FIG. 1, the configuration of the information processing system 1 according to the present disclosure and the outline of the information processing executed by the information processing system 1 will be described. FIG. 1 is a diagram showing an overview of information processing according to the first embodiment of the present disclosure. Information processing according to the first embodiment of the present disclosure is realized by an information processing system 1 shown in FIG. The information processing system 1 includes an information processing device 100 and a user terminal 10 . The information processing apparatus 100 and the user terminal 10 communicate with each other using a wired or wireless network (not shown). Note that the number of devices constituting the information processing system 1 is not limited to that illustrated.
[0012]
The information processing apparatus 100 is an example of an information processing apparatus according to the present disclosure, calculates an HRTF (head-related transfer function) corresponding to each user, and provides the calculated HRTF. The information processing device 100 is implemented by, for example, a server device.
[0013]
The user terminal 10 is an information processing terminal used by a user who wants to receive the provision of HRTF. The user terminal 10 is realized by, for example, a smartphone having a photographing function. In the example of FIG. 1, the user terminal 10 is used by a user U01 who is an example of a user.
[0014]
 HRTF expresses sound changes caused by peripheral objects, including the shape of the human auricle (auricular shell) and head, as a transfer function. In general, measurement data for obtaining the HRTF is acquired by measuring acoustic signals for measurement using a microphone worn in the auricle of a person, a dummy head microphone, or the like.
[0015]
For example, HRTFs used in technologies such as 3D sound are often calculated using measurement data acquired with dummy head microphones, etc., or the average value of measurement data acquired from many people. However, since the HRTF varies greatly among individuals, it is desirable to use the user's own HRTF in order to achieve more effective sound effects. That is, by replacing the general HRTF with the user's own HRTF, it is possible to provide the user with a more realistic acoustic experience.
[0016]
However, there are various problems in measuring the user's HRTF individually. For example, relatively high density measurement data is required to obtain HRTFs that provide good acoustics. Acquiring high-density measurement data requires measurement data of acoustic signals output to the user from various angles surrounding the user. Such measurements take a long time and impose a heavy physical burden on the user. In addition, since accurate measurement requires measurement in an anechoic chamber or the like, the cost burden is large. For this reason, the calculation of the HRTF presents challenges such as reducing the user's burden and reducing the measurement cost.
[0017]
Regarding the above issues, there is a technology that simulates measurements by expressing the user's ears and head in a 3D model and performing acoustic simulation on the 3D model. According to this technology, the user can calculate the HRTF without performing actual measurement in a measurement room by providing scan data of the head or an image of the head.
[0018]
However, 3D model generation processing and acoustic simulation for 3D models have a very large computational load. Therefore, even if a user-specific HRTF is incorporated into, for example, software that utilizes 3D sound using the above technology, there is a risk that a time lag of several tens of minutes or several hours will occur. It is hard to say that this is highly convenient for the user. That is, from the viewpoint of allowing users to utilize HRTFs, there is also the problem of improving the processing speed in calculating the HRTFs.
[0019]
As described above, there are various problems in obtaining HRTFs corresponding to individual users. The information processing apparatus 100 according to the present disclosure solves the above problems by the information processing of the present disclosure.
[0020]
Specifically, when an image including an image of an ear is input, the information processing apparatus 100 receives a trained model (hereinafter simply referred to as “model”) that has been trained to output an HRTF corresponding to the ear. ) to calculate the HRTF corresponding to the user. For example, when the information processing apparatus 100 acquires an image including an image of the ear of the user U01 from the user terminal 10, the information processing apparatus 100 inputs the image into the model and calculates the HRTF unique to the user U01. That is, the information processing apparatus 100 calculates the HRTF without going through the process of generating a 3D model based on the image of the user U01 and performing acoustic simulation.
[0021]
As a result, the information processing apparatus 100 can calculate the HRTF in an extremely short time compared to the case of executing the acoustic simulation. Hereinafter, an outline of information processing executed by the information processing apparatus 100 according to the present disclosure will be described along the flow using FIG. 1 .
[0022]
As shown in FIG. 1, the user U01 shoots himself/herself from the lateral side of the head in order to obtain an image including the image of his/her ears (step S1). For example, the user U01 uses a camera included in the user terminal 10 to photograph his/her head. In the present disclosure, the ear image is generally not limited to a two-dimensional color image that can be captured by the user terminal 10 or the like, but may be a monochrome image, a depth image including depth information, or any combination thereof. may Also, the image used in the processing of the present disclosure may be a plurality of images instead of one.
[0023]
The user terminal 10 performs preprocessing for transmission to the information processing device 100 on the image 5 obtained in step S1 (step S2). Specifically, as preprocessing, the user terminal 10 detects the image of the user U01's ear included in the image 5 and cuts out the range of the detected ear image from the image 5 . Details of preprocessing such as ear detection will be described later.
[0024]
The user terminal 10 creates an image 6 including an image of the ear of the user U01 through preprocessing. Then, the user terminal 10 transmits the image 6 to the information processing device 100 (step S3). A series of processes such as the process of creating the image 6 from the image 5 obtained by shooting and the process of transmitting the image 6 are executed by, for example, a program (for example, an application for smartphones) provided by the information processing apparatus 100. be done. In this way, in the information processing according to the present disclosure, only the image 6 obtained by cutting out only the image of the ear from the image 5 is transmitted on the network, and the image 5 that may identify the individual is not transmitted. Excellent processing can be realized. Further, the information processing apparatus 100 acquires only the image 6 without acquiring the image 5, thereby avoiding the risk of handling personal information. The information processing apparatus 100 may acquire the image 5 from the user terminal 10 and perform processing for creating the image 6 on the information processing apparatus 100 side. This configuration will be described later as a second embodiment.
[0025]
The information processing device 100 acquires the image 6 transmitted from the user terminal 10 . Then, the information processing apparatus 100 inputs the image 6 to the model stored in the storage unit 120 (step S4). This model is a model trained to output an HRTF corresponding to the ear when a two-dimensional image including the image of the ear is input. That is, the information processing apparatus 100 calculates the HRTF corresponding to the ear (in other words, the user U01) included in the image 6 by inputting the image 6 into the model.
[0026]
Then, the information processing device 100 provides the calculated HRTF to the user terminal 10 (step S5). In this way, the user U01 can obtain an HRTF unique to himself/herself without going through actual measurement processing, acoustic simulation of a 3D model, etc., simply by photographing his/her profile and preparing only the image 5. can. That is, the information processing device 1 00 can provide the HRTF to the user U01 in an extremely short time without burdening the user U01 with the measurement. As a result, the information processing apparatus 100 can improve the user's convenience in processing related to HRTF.
[0027]
As described above, in the information processing according to the present disclosure, the HRTF is calculated using the model created through the learning process, thereby speeding up the process. FIG. 1 shows an outline of processing for providing HRTF to the user U01 among the information processing according to the present disclosure. I will explain in detail. Details will be explained in FIG. 2 and below, but the "model" shown in FIG. good.
[0028]
[1-2. Overall flow of information processing according to the present disclosure]
Before describing details such as the configuration of the information processing device 100, FIG. 2 shows the overall flow of information processing executed by the information processing device 100 according to the present disclosure. FIG. 2 is a conceptual diagram showing the overall flow of information processing according to the present disclosure.
[0029]
First, the information processing apparatus 100 collects data on the ear patterns of a plurality of people, and creates an ear model based on the collected ear patterns (step S11). Note that the ear pattern is not necessarily limited to a model of a person's ear made of plaster or the like, and may be any information as long as it indicates the shape of a person's ear. Further, in the present disclosure, an ear model is a model that outputs a corresponding ear shape when parameters indicating characteristics of the ear (hereinafter referred to as “ear parameters”) are input. The ear parameters are obtained, for example, by subjecting the ear shape to principal component analysis based on data relating to the shape of the ear (for example, a CT (Computed Tomography) scan of a collected ear model converted into data). Thus, when the ear parameters are obtained, the information processing apparatus 100 can obtain ear shape data (in other words, a 3D model of an ear) corresponding to the ear parameters.
[0030]
After that, the information processing device 100 generates an ear parameter estimation model based on the ear model (step S12). Information processing apparatus 100 can generate a large number of ear images by inputting ear parameters to the ear model generated in step S11. Ear parameters may be input randomly, or arbitrary rules (for example, if a fact that there is a specific tendency in ear shape for each specific race is found, a rule is derived based on that fact). ), and the generated values ​​may be input. For this reason, information processing apparatus 100 learns the relationship between the generated ear image and the ear parameter that is the source of generation, so that when an image including an ear is input, the ear parameter corresponding to the ear is obtained. You can generate a model that outputs Such a model is an ear parameter estimation model. Thus, if the information processing apparatus 100 obtains a two-dimensional image including the image of the ear, the ear parameter corresponding to the ear can be obtained. Then, when the ear parameters are obtained, the information processing apparatus 100 uses the ear model generated in step S11 to obtain the 3D model of the ear included in the image. In the learning described above, the relationship between an image of the person's ear in which the shape of the ear is converted into data and an ear parameter obtained by converting the ear into data may be learned. In this case, learning is performed using actual captured images instead of CG (Computer Graphics) images, so it is assumed that the accuracy of the generated ear parameter estimation model can be further improved.
[0031]
The information processing apparatus 100 performs an acoustic simulation on a 3D model generated using the ear parameter estimation model, and generates a unique HRTF corresponding to the 3D model (hereinafter, generated corresponding to such individual ear shape). The resulting HRTF is called a "personalized HRTF") is calculated (step S13). That is, through steps S11 to S13, the information processing apparatus 100 can implement a series of processes for calculating a personalized HRTF from an image including the ear by acoustic simulation.
[0032]
Furthermore, the information processing apparatus 100 generates a large number of 3D models from randomly or regularly generated ear parameters, and repeats the process of performing acoustic simulation on the generated 3D models, thereby combining the ear parameters and the personalized HRTFs. learn the relationship between That is, information processing apparatus 100 generates an HRTF learning model based on the calculated personalized HRTF (step S14).
[0033]
In the present disclosure, an HRTF learning model is a model that, when ear parameters are input, outputs a personalized HRTF corresponding to the ear parameters. Accordingly, when the ear parameter is obtained, the information processing apparatus 100 can obtain the personalized HRTF corresponding to the ear parameter.
[0034]
Thereafter, when acquiring an image from the user, the information processing apparatus 100 inputs the image (more precisely, the ear parameter of the ear included in the image) into the HRTF learning model, thereby personalizing the user. HRTF is calculated (step S15). The process shown in step S15 corresponds to the series of processes shown in FIG.
[0035]
As described above, the information processing apparatus 100 can generate a plurality of models and perform information processing using the generated models, thereby calculating a personalized HRTF from an image acquired from the user. Note that the processing shown in FIG. 2 does not necessarily need to be executed in the order of steps S11 to S15. For example, the information processing apparatus 100 only needs to perform the process of step S15 after generating each model required for the process.
[0036]
The outline of the overall flow of information processing according to the present disclosure has been described above. 3 and subsequent figures, configurations of the information processing apparatus 100 and the user terminal 10 will be described, and details of various learning processes will be described in order.
[0037]
[1-3. Configuration of information processing apparatus according to first embodiment]
The configuration of the information processing apparatus 100 according to the first embodiment will be described using FIG. FIG. 3 is a diagram showing a configuration example of the information processing device 100 according to the first embodiment of the present disclosure.
[0038]
As shown in FIG. 3, the information processing device 100 has a communication section 110, a storage section 120, and a control section . Note that the information processing apparatus 100 includes an input unit (for example, a keyboard, a mouse, etc.) that receives various operations from an administrator or the like who manages the information processing apparatus 100, and a display unit (for example, a liquid crystal display, etc.) for displaying various information. ).
[0039]
The communication unit 110 is implemented by, for example, a NIC (Network Interface Card) or the like. The communication unit 110 is connected to a network N (such as the Internet) by wire or wirelessly, and transmits and receives information to and from the user terminal 10 and the like via the network N.
[0040]
The storage unit 120 is implemented by, for example, a semiconductor memory device such as RAM (Random Access Memory) or flash memory, or a storage device such as a hard disk or optical disk. The storage unit 120 stores various data used in the learning process, models generated by the learning process, and the like.
[0041]
As shown in FIG. 3, the storage unit 120 includes an ear shape information storage unit 121, an ear model storage unit 122, an ear image storage unit 123, an ear parameter estimation model storage unit 124, and an HRTF processing model storage unit 125. , an HRTF learning data storage unit 126 and an HRTF learning model storage unit 127 .
[0042]
The ear pattern information storage unit 121 stores information (that is, information about the shape of the ear) in which the ear pattern actually taken from the human body is turned into 3D model data. Specifically, the ear pattern information storage unit 121 stores data representing a three-dimensional shape (3D polygons, etc.) obtained by CT scanning the collected ear pattern.
[0043]
The ear model storage unit 122 stores an ear model according to the present disclosure. The ear model is a model that outputs the shape of the corresponding ear when the ear parameter indicating the characteristics of the ear is input.
[0044]
The ear parameters are obtained by subjecting the data representing the shape of the ear pattern stored in the ear pattern information storage unit 121 to principal component analysis. In other words, the ear parameter is obtained by quantifying the part of the ear that changes greatly (characterizes the shape of the ear) by performing statistical analysis (principal component analysis) on the 3D polygon of the ear. The ear parameters according to the present disclosure are indicated by a combination of ten numerical values, for example, and each numerical value is indicated by a numerical value from -10 to +10, for example. For example, an ear parameter with all numerical values ​​of "0" corresponds to an ear having an average shape of the training data (the ear molds taken). Note that the information processing apparatus 100 may appropriately apply a known technique used in, for example, a person's face generation process or the like to the process of generating a model representing the shape of an ear by principal component analysis. In addition, the information processing apparatus 100 may generate ear parameters by appropriately using known analysis methods such as independent component analysis and other non-linear models, without being limited to principal component analysis. Here, the ear parameter is not limited to the numerical value of the part of the ear that changes greatly, and may be, for example, a parameterized feature related to the shape of the ear that greatly affects the HRTF. .
[0045]
The ear image storage unit 123 stores images including images of ears. For example, the ear image storage unit 123 stores, as the ear image, a CG image obtained by rendering the shape of the ear (3D model of the ear) generated by the ear model. Further, the ear image storage unit 123 may store an image including an image of the ear transmitted from the user as the ear image.
[0046]
Here, FIG. 4 shows an example of the ear image storage unit 123 according to the present disclosure. FIG. 4 is a diagram showing an example of the ear image storage unit 123 of the present disclosure. In the example shown in FIG. 4, the ear image storage unit 123 has items such as "ear parameter", "ear 3D model data", "head 3D model data", "ear image ID", and "image generation parameter". . "Image generation parameter" has small items such as "texture", "camera angle", "resolution", and "brightness".
[0047]
"Ear parameters" are parameters that indicate the characteristics of the shape of the ear. For example, ear parameters are represented by 10-dimensional numerical values ​​or the like. "Ear 3D model data" is data representing the three-dimensional shape of the ear reconstructed based on the ear parameters. "Head 3D model data" is data indicating the three-dimensional shape of the head to which the ear 3D model data is synthesized when reconstructing the 3D model of the person.
[0048]
"Ear image ID" indicates identification information for identifying an ear image obtained by rendering a 3D model. As shown in FIG. 4, a plurality of ear images are generated from one 3D model by variously changing parameters (image generation parameters) set during rendering.
[0049]
"Image generation parameters" indicate setting parameters in rendering for generating images. "Texture" indicates the setting of CG texture. "Camera angle" indicates a pseudo camera shooting angle when rendering a 3D model to obtain a two-dimensional image. "Resolution" indicates the resolution at the time of rendering. "Brightness" indicates the brightness at the time of rendering. The brightness item may include setting data such as the angle of light (incident light) in rendering.
[0050]
In FIG. 4, the data of each item is conceptually described as "A01" and "B01", but in reality, the data of each item includes specific data corresponding to each item. is stored. For example, the item "ear parameter" stores a list of 10 specific numerical values. Similarly for other items, various numerical values ​​​​corresponding to each item Information is stored.
[0051]
That is, in the example shown in FIG. 4, the ear 3D model data generated by the ear parameter "A01" is "B01", and is combined with the ear 3D model data to construct the human 3D model. It indicates that the head 3D model data is "C01". Also, the ear images obtained from the generated 3D model of the person are a plurality of ear images identified by ear image IDs "D01", "D02", "D03", and the like. The ear image identified by the ear image ID "D01" has a texture of "E01", a camera angle of "F01", and a resolution of "G01" as image generation parameters for rendering. It indicates that the brightness is "H01".
[0052]
Return to Figure 3 and continue the explanation. The ear parameter estimation model storage unit 124 stores an ear parameter estimation model. The ear parameter estimation model is a model that outputs ear parameters corresponding to the ear when a two-dimensional image including an image of the ear is input.
[0053]
The HRTF processing model storage unit 125 stores HRTF processing models. Although the details will be described later, the HRTF processing model performs processing for compressing the information amount of the HRTF calculated by acoustic simulation or the like. In the following description, the HRTF compressed by the HRTF processing model may be referred to as HRTF parameters.
[0054]
The HRTF learning data storage unit 126 stores learning data for generating a model (HRTF learning model to be described later) for calculating HRTFs from images including images of ears. Specifically, the HRTF learning data storage unit 126 stores, as learning data, data in which an ear parameter indicating the shape of the ear and an HRTF corresponding to the shape of the ear specified based on the ear parameter are combined. do.
[0055]
The HRTF learning model storage unit 127 stores HRTF learning models. The HRTF learning model is a model that outputs an HRTF corresponding to the ear when an image including a video of the ear is input. For example, when an image including an image of an ear is acquired, the HRTF learning model outputs ear parameters corresponding to the ear using the ear parameter estimation model, and outputs HRTFs corresponding to the ear parameters.
[0056]
The control unit 130 stores a program (for example, an information processing program according to the present disclosure) stored inside the information processing apparatus 100 by a CPU (Central Processing Unit) or MPU (Micro Processing Unit), for example, and stores it in a RAM (Random Access Memory). ) etc. as a work area. Also, the control unit 130 is a controller, and may be realized by an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array).
[0057]
As shown in FIG. 3, the control unit 130 has a learning unit 131 and an estimating unit 140. The learning unit 131 includes a reception unit 132, an ear model learning unit 133, an image generation unit 134, an ear parameter learning unit 135, and an HRTF learning unit 136. Realize or perform an action. The estimation unit 140 also includes an acquisition unit 141, a calculation unit 142, and a provision unit 143, and implements or executes information processing functions and actions described below. Note that the internal configuration of the control unit 130 is not limited to the configuration shown in FIG. 3, and may be another configuration as long as it performs information processing described later.
[0058]
First, the learning unit 131 will be explained. The learning unit 131 performs learning processing on various data and generates various models used by the estimating unit 140 .
[0059]
Note that the learning unit 131 performs learning for generating a model based on various data, but the learning process described below is an example, and the type of learning process executed by the learning unit 131 is specified to some type. not. For example, the learning unit 131 may generate models using various learning algorithms such as neural networks, support vector machines, clustering, and reinforcement learning.
[0060]
The reception unit 132 receives various types of information. For example, the receiving unit 132 receives ear-shaped CT scan data taken from a human body. The receiving unit 132 stores the received data in the ear shape information storage unit 121 .
[0061]
The ear model learning unit 133 performs learning processing regarding the ear model and generates the ear model. Ear model learning section 133 stores the generated ear model in ear model storage section 122 .
[0062]
Here, an example of learning processing executed by the reception unit 132 and the ear model learning unit 133 will be described using FIG. FIG. 5 is a diagram illustrating an example of learning processing regarding an ear model according to the present disclosure.
[0063]
As shown in FIG. 5 , the reception unit 132 receives scanned data collected from the ear mold, and stores the received data in the ear mold information storage unit 121 . The receiving unit 132 also sends the received data to the ear model learning unit 133 (step S16).
[0064]
The ear model learning unit 133 homogenizes the acquired ear pattern data to generate homologous data of the ear pattern (step S17). Here, homogenization refers to unifying the number of vertices and the configuration of polygons of a 3D model so as to be the same as those of a reference 3D model. In this case, care should be taken not to change the shape before and after homologation. Furthermore, the ear model learning unit 133 performs principal component analysis on the homologous data (step S18). Thereby, the ear model learning unit 133 generates a model (ear model) for calculating ear parameters indicating the shape of the ear from the shape of the ear. Ear model learning section 133 stores the generated ear model in ear model storage section 122 .
[0065]
Return to Figure 3 and continue the explanation. The image generator 134 generates an image including the image of the ear. For example, the image generator 134 randomly generates ear parameters, inputs the generated ear parameters into an ear model, and generates a 3D model of the ear. Furthermore, the image generation unit 134 randomly sets parameters (hereinafter referred to as “image generation parameters”) such as the texture (for example, skin color) of the generated 3D model, the rendering quality (image quality, etc.), and the camera angle in CG rendering. to generate Then, the image generation unit 134 performs rendering by appropriately combining the generated 3D model and a plurality of image generation parameters to generate a CG image in which the shape of the ear, the color of the skin, and the like change in various ways.
[0066]
An image transmitted from the user is used for processing in the estimation processing described later, but it is assumed that there is a high possibility that the user's skin color, the angle of the ear at the time of shooting, and the like vary widely in the image transmitted from the user. be. Therefore, in such processing, there is a problem of accurately recognizing the image of the ear in all images transmitted from the user. The image generator 134 generates a large number of images corresponding to various situations as described above, thereby improving the accuracy of image recognition and solving the above problem.
[0067]
The ear parameter learning unit 135 generates an ear parameter estimation model by learning the relationship between an image containing an ear video and ear parameters. The ear parameter learning unit 135 corresponds to the first learning unit according to the present disclosure. The image including the image of the ear may be an image obtained by actually capturing the ear of a person, or may be a CG image generated based on ear parameters, as will be described later.
[0068]
For example, the ear parameter learning unit 135 generates an ear image obtained by rendering three-dimensional data obtained by synthesizing the three-dimensional data of the ear generated based on the ear parameter and the three-dimensional data of the head, and the ear image. An ear parameter estimation model is generated by learning the relationship with the parameters. Specifically, the ear parameter learning unit 135 learns the relationship between the CG image generated by the image generation unit 134 and the ear parameter. As described above, the image generation unit 134 generates the CG image based on the ear parameters set randomly or regularly, so the ear parameters are uniquely determined for the CG image. For this reason, the ear parameter learning unit 135 learns the relationship between the input CG image and the ear parameter. You can generate a model that outputs Note that the ear parameter learning unit 135 does not necessarily have to use ear images obtained by rendering the combined three-dimensional data of the head for learning. That is, the ear parameter learning unit 135 learns the relationship between the ear image obtained by rendering only the three-dimensional data of the ear generated based on the ear parameter and the ear parameter, thereby obtaining an ear parameter estimation model. may be generated.
[0069]
The ear parameter learning unit 135 also determines the relationship between a plurality of ear images obtained by changing the texture of the three-dimensional data of the ear or head, the camera angle in rendering, the brightness in rendering, etc., and the ear parameter common to the plurality of ear images. Generate an ear parameter estimation model by learning the gender. In this way, the ear parameter learning unit 135 performs learning using various types of ear images, so that it can accurately output ear parameters regardless of what kind of image is input (for example, (stable estimations can be made for any variation other than the included ear parameters), and a robust model can be generated.
[0070]
Here, an example of learning processing executed by the image generation unit 134 and the ear parameter learning unit 135 will be described using FIG. FIG. 6 is a diagram illustrating an example of learning processing regarding an ear parameter estimation model according to the present disclosure.
[0071]
As shown in FIG. 6, the image generation unit 134 refers to the ear model storage unit 122 (step S21) and acquires the ear model. The image generator 134 also generates random numbers corresponding to ear parameters, random numbers corresponding to CG textures, rendering camera angles, and the like (step S22). That is, the image generator 134 generates various parameters (image generation parameters) for generating ear images.
[0072]
Then, the image generator 134 acquires ear parameters among the image generation parameters (step S23), inputs the acquired ear parameters into the ear model, and reconstructs a 3D model of the ear (step S24).
[0073]
Subsequently, the image generation unit 134 acquires parameters such as CG texture among the image generation parameters (step S25), inputs the acquired parameters, and applies 3DCG rendering to the 3D model (step S26). As for the head used for rendering, for example, the average head of a plurality of people (for example, a plurality of people whose ear molds were collected), or a 3D model of the head used as a sample is used. . Here, the 3D model of the head may be generated by homogenizing 3D data obtained by 3D scanning the heads of a plurality of persons, similarly to the ear model. In this case, the image generator 134 can randomly generate a 3D head model by generating parameters using random numbers. Note that the image generation unit 134 can also generate various textures using random numbers by creating a texture model generated by a similar method.
[0074]
As a result, the image generation unit 134 generates an image (ear image) including the image of the ear. Note that the image generator 134 can generate a plurality of ear images from one 3D model by variously changing parameters such as texture.
[0075]
Here, an example of the ear image generated by the image generation unit 134 is shown using FIG. FIG. 7 shows an ear image generation process according to the present disclosure.FIG. 4 is a diagram showing an example of the theory;
[0076]
The image generation unit 134 generates a 3D model representing the three-dimensional shape of the ear using randomly generated ear parameters (step S41). The image generation unit 134 also acquires a 3D head model generated based on data representing the average 3D shape of the heads of a plurality of persons (step S42). Then, the image generator 134 synthesizes the ear generated in step S41 and the head 3D model obtained in step S42 to generate a pseudo human 3D model.
[0077]
Subsequently, the image generation unit 134 performs a process (rendering) of taking a pseudo image of the generated 3D model and generating a two-dimensional image from the 3D model. For example, the image generation unit 134 sets the front angle of the ear of the 3D model as a pseudo imaging angle, and generates an image in which the image of the ear is substantially in the center.
[0078]
Here, the image generation unit 134 randomly inputs image generation parameters to the 3D model to obtain CG texture (specifically, skin color, etc.), rendering quality (resolution, etc.), Change the position of the ears, etc. Thereby, the image generator 134 can generate a large number of images with different skin colors (step S43).
[0079]
The image group 20 shows a plurality of ear images generated by the image generation unit 134. In this way, the image generation unit 134 generates a large number of various ear images, thereby improving the recognition accuracy of the ear images, which will be described later.
[0080]
Return to Figure 6 and continue the explanation. The image generation unit 134 stores the generated ear image in the ear image storage unit 123 (step S27). The image generation unit 134 stores the image generation parameters used when generating each image in the ear image storage unit 123 in association with the generated ear images (step S28). As a result, the image generator 134 can hold a large number of ear images associated with ear parameters as learning data. For example, the image generator 134 can hold a large number of ear images associated with ear parameters as learning data.
[0081]
Subsequently, the ear parameter learning unit 135 refers to the ear image storage unit 123 (step S29) to acquire ear images and ear parameters. The ear parameter learning unit 135 then learns the relationship between the ear image and the ear parameter, and generates an ear parameter estimation model. The ear parameter learning unit 135 stores the generated ear parameter estimation model in the ear parameter estimation model storage unit 124 (step S30).
[0082]
The ear parameter estimation model is generated using, for example, a convolutional neural network useful for extracting image feature values. Note that the cost formula (cost function) in learning is represented by the following formula (1), for example.
[0083]
[Number 1]

[0084]
In Equation (1), "α true" indicates the true value of the ear parameter. Also, "α est" indicates the estimated value of the ear parameter. "A ear" indicates an ear model obtained by principal component analysis. Also, the distance function on the right side indicates the L2 norm (Euclidean distance). As the true value of the ear model parameter, for example, a parameter indicating a person's ear actually measured when taking an ear mold can be used. That is, the ear parameter used when generating the ear image is the true value, and the value output when the ear image is input to the ear parameter estimation model under learning is the estimated value. As learning processing, the information processing apparatus 100 updates the coefficients so as to minimize the value of the cost formula for the current estimated value.
[0085]
Here, the ear parameter estimation model generated by the ear parameter learning unit 135 will be described using FIG. FIG. 8 is a diagram for explaining an ear parameter estimation model according to the present disclosure.
[0086]
When acquiring the ear image 30, the information processing apparatus 100 inputs the acquired ear image 30 to the ear parameter estimation model. The ear parameter estimation model has a structure of, for example, a convolutional neural network, and obtains feature values ​​representing the ear image 30 while dividing the input ear image 30 into rectangular portions of several pixels. Finally, the ear parameter estimation model outputs the ear parameter corresponding to the image of the ear included in the ear image 30 as a feature quantity representing the ear image 30 (step S45).
[0087]
The information processing apparatus 100 can reconstruct an ear pattern (3D model) corresponding to the ear included in the ear image 30 by inputting the output ear parameters into the ear model. A 3D model 40 shown in FIG. 8 is a 3D model obtained by reconstructing the ear included in the ear image 30 using CG based on the ear parameters.
[0088]
Return to Figure 3 and continue the explanation. The HRTF learning unit 136 generates various models related to HRTFs by learning the relationship between information related to ear shape and HRTFs. For example, the HRTF learning unit 136 generates a trained model for calculating the HRTF by learning the relationship between an image including the image of the ear and the HRTF corresponding to the ear. The HRTF learning unit 136 corresponds to the second learning unit according to the present disclosure.
[0089]
For example, the HRTF learning unit 136 performs acoustic simulation on three-dimensional data obtained by synthesizing the three-dimensional data of the ear and the three-dimensional data of the head generated based on the ear parameters, and obtains the HRTF obtained by the acoustic simulation. and ear parameters to generate a trained model.
[0090]
Also, the HRTF learning unit 136 may generate a trained model by compressing the information amount of the HRTF obtained by the acoustic simulation and learning the relationship between the compressed HRTF and the ear parameter.
[0091]
Also, the HRTF learning unit 136 may set a listening point for the three-dimensional data of the ear generated based on the ear parameter, and perform acoustic simulation using the set listening point. A listening point is a hypothetical position at which a person is supposed to listen to sound. For example, the position of the listening point corresponds to the position where the microphone is installed in the dummy head microphone (such as the entrance of the ear canal in the dummy head).
[0092]
Regarding each processing of the HRTF learning unit 136 described above, FIG. 9 shows the flow of processing for generating various models related to HRTF. FIG. 9 is a diagram showing an overview of the flow of generation processing of various models related to HRTF.
[0093]
FIG. 9 shows an example in which the HRTF learning unit 136 performs predetermined learning processing based on the image transmitted by the user. In this case, the user uses the user terminal 10 to photograph his/her ears (more precisely, the head including the ears) (step S51). After that, the user terminal 10 specifies a range including the image of the ear from the captured photograph, cuts out the specified range, and performs preprocessing to acquire an ear image (step S52).
[0094]
After that, the HRTF learning unit 136 uses the ear parameter estimation model to calculate the ear parameters of the ear included in the ear image transmitted from the user (step S53). Further, the HRTF learning unit 136 reconstructs a 3D model of the ear based on the ear parameters, and combines the reconstructed ear with the 3D head model to generate a 3D model of the person (step S54).
[0095]
Subsequently, the HRTF learning unit 136 performs an acoustic simulation on the generated 3D model to obtain a personalized HRTF (step S55) for the 3D model. As a result, the HRTF learning unit 136 can obtain learning data in which the ear included in the ear image transmitted from the user is associated with the personalized HRTF.
[0096]
Note that the example of FIG. 9 shows an example of generating learning data in which the personalized HRTF obtained by the acoustic simulation and the ear data are associated with each other, but the HRTF learning unit 136 does not necessarily have to be personalized by the acoustic simulation. In some cases it is not necessary to obtain HRTF. For example, when a personalized HRTF (HRTF obtained using a measuring device in an anechoic room or the like) is obtained for a person whose ear mold was taken, the HRTF learning unit 136 compares the measured personalized HRTF with , and human ear patterns (ear parameters) may be acquired.
[0097]
The HRTF learning unit 136 automatically sets the listening point in the 3D model of the person during the acoustic simulation. This point will be described with reference to FIG. FIG. 10 is a diagram for explaining reconstruction of a 3D model according to the present disclosure.
[0098]
The HRTF learning unit 136 reconstructs a 3D ear model from randomly generated ear parameters (step S71). Subsequently, the HRTF learning unit 136 combines the ear 3D model with the head 3D model to generate a human 3D model (step S72). Further, the HRTF learning unit 136 sets the listening point 60 of the sound source based on the shape of the ear in the 3D model (step S73). For example, the HRTF learning unit 136 can automatically set the listening point by previously learning information that associates the shape of the ear with the position of the listening point of the sound source. For example, when the HRTF learning unit 136 generates the 3D model, it estimates the position of the listening point in the 3D model and automatically sets the listening point. The listening point corresponds, for example, to the ear canal of the ear and can generally be deduced from the shape of the ear.
[0099]
After that, the HRTF learning unit 136 remesh the generated 3D model so as to satisfy the computational constraints of the 3D model in the acoustic simulation (step S74). This is because, in 3D model simulation, for example, the upper limit for the total number of polygons, the length of edges connecting vertices, and the like are sometimes determined by simulation conditions. That is, before the generated 3D model is simulated, the HRTF learning unit 136 appropriately remesh the 3D model so as to satisfy computational constraints so that the simulation can be appropriately performed. The HRTF learning unit 136 then performs an acoustic simulation on the generated 3D model and the set listening point 60 to calculate a personalized HRTF (step S75).
[0100]
Next, with reference to FIG. 11, the detailed flow of the HRTF model generation process will be described. FIG. 11 is a diagram for explaining the details of the model generation process for HRTF.
[0101]
As shown in FIG. 10, the HRTF learning unit 136 performs acoustic simulation (step S82) after performing head synthesis (step S81). The HRTF learning unit 136 analyzes the measurement data obtained by the acoustic simulation (HRTF post-processing) and calculates a numerical value indicating the personalized HRTF (step S83). Note that the HRTF post-processing is, for example, Fourier transformation of HRIF (Head-Related Impulse Response) obtained by acoustic simulation to obtain HRTF.
[0102]
Here, the HRTF learning unit 136 refers to the HRTF processing model storage unit 125 (step S84), and inputs the HRTF to the model (HRTF processing model) for processing the calculated HRTF. As a result, the HRTF learning unit 136 obtains HRTFs with reduced dimensions (step S85). That is, the HRTF learning unit 136 outputs HRTF parameters, which are HRTFs with reduced dimensions, from the HRTF processing model (step S86).
[0103]
In this way, the HRTF learning unit 136 does not directly use the HRTF obtained in the acoustic simulation for processing, but performs processing using HRTF parameters whose dimensions have been reduced. This is because the HRTF is a function with a very large number of dimensions, and if the model generation processing or calculation processing is performed as it is, the calculation processing load increases.
[0104]
The HRTF learning unit 136 collects data related to the head for which the acoustic simulation was performed.Data (data used as a basis for head synthesis, such as ear parameters, for example) are associated with the calculated HRTF parameters and stored in the HRTF learning data storage unit 126 (step S87).
[0105]
After that, the HRTF learning unit 136 randomly generates new ear parameters (step S88), and performs head synthesis using the newly generated ear parameters (step S89). By repeating steps S81 to S89, the HRTF learning unit 136 collects learning data required for learning.
[0106]
After that, when sufficient learning data is accumulated, the HRTF learning unit 136 refers to the HRTF learning data storage unit 126 (step S90), and learns the relationship between ear parameters and HRTFs (HRTF parameters to be exact). (Step S91). Through such learning, the HRTF learning unit 136 generates an HRTF learning model for directly obtaining an HRTF from ear parameters, and stores the generated HRTF learning model in the HRTF learning model storage unit 127 .
[0107]
Next, using FIG. 12, the relationship between the HRTF and the HRTF parameters will be described. FIG. 12 is a diagram for explaining HRTF compression and decompression according to the present disclosure.
[0108]
As shown in FIG. 12, the HRTF learning unit 136 performs FFT (Fast Fourier Transform) on the HRIFHRTF (in the example of FIG. 12, it is assumed to be 1000 directions×500 taps) obtained by acoustic simulation (step S101). Through such processing, the HRTF learning unit 136 extracts amplitude characteristics (step S102), and performs processing for thinning out, for example, frequency components with low auditory sensitivity (step S103). (θ, φ) and a function HRTF (θ, φ, f) related to frequency (f) At this time, if the number of frequency bins is k, the input to the function is The frequencies f are denoted as f k = f 0, f 1, f 2, . Here, since the HRTF after the Nyquist frequency (f k/2) is a fold of the complex conjugate before the frequency f k/2, in information processing, as a frequency bin, from f0=0 to the Nyquist frequency Only (k/2)+1 up to (f k/2) can be used, and absolute values ​​can be used for at least one or more frequency bins, eg f 0 to f k When all frequencies up to /2 are transformed into absolute values, the function H2 after transformation is represented by the following equation (2).
[0109]
[Number 2]

[0110]
That is, the HRTF learning unit 136 can compress the dimension of the original HRTF to the real number (k/2)+1 dimension. Furthermore, the HRTF learning unit 136 can perform frequency compression on H2 in the above equation (2) and further reduce the dimensions to less than (k/2)+1. There are various known techniques for dimensionality compression. For example, the HRTF learning unit 136 uses a technique such as performing cepstrum transform on the function and acquiring only one or more frequency bins and less than (k/2)+1. As an example, the HRTF learning unit 136 obtains the average value of multiple frequency bins and reduces the dimensionality based on the average value. For example, when the frequency bins are represented by the following formula (3) (a l, L, l are each integers equal to or greater than 0), a new function H 3 is represented by the following formula (4).
[0111]
[Number 3]

[0112]
[Number 4]

[0113]
As a result, the HRTF learning unit 136 can reduce the function H2 represented by (K/2)+1 dimensions to L dimensions. It should be noted that the method for obtaining the average value is not limited to the above method, and may be obtained using, for example, the root mean square, weighted average, or the like. This reduces the HRTF to, for example, 1000 directions×50 dimensions. When restoring the dimension reduced by the function H3 (step S110, etc., which will be described later), the HRTF learning unit 136 can restore using various methods such as linear interpolation and spline interpolation. It is expected that the dimension-reverted function H'2 will have smoother characteristics than the function H2. 2(θ, φ, k) is obtained. As an example, it can be devised to select a l such that the frequency interval between f al and f al+1 increases as the frequency increases.
[0114]
The HRTF learning unit 136 further performs spherical harmonic fitting processing on the dimensionally reduced HRTF, thereby compressing the amount of information to about 50 coefficients×50 dimensions (step S104). Here, spherical harmonic fitting refers to fitting in the spatial direction for each compressed frequency using a spherical harmonic function. The relationship between the HRTF and the spherical harmonics is given by the following equation (5).
[0115]
[Number 5]

[0116]
 As shown in the above formula (5), the spherical harmonic function Y is expressed by the coefficient hnm(f). By truncating the number of dimensions n at this time to a certain finite number N, the number of dimensions of the coefficient hnm(f) can be made smaller than the number of dimensions (number of directions) of the original HRTF. This means ignoring spatially too fine amplitudes, which are unnecessary for human perception, and obtaining only smooth shapes. In order to obtain the vector h=(h 00, h 1-1, .
[0117]
[Number 6]

[0118]
That is, when Y in the above equation (6) is a matrix of spherical harmonics and H is a matrix of spherical harmonics, h that minimizes E on the left side is obtained. Since the second term on the right side of the above equation (6) is a regularization term, any value λ may be selected (for example, λ may be 0). Then, the above h is represented by the following formula (7).
[0119]
[Number 7]

[0120]
By using the above equation (7), the HRTF learning unit 136 can obtain each h corresponding to the necessary frequencies. Furthermore, the HRTF learning unit 136 compresses the information amount of the HRTF so that it can be expressed in approximately several hundred dimensions by performing dimension compression by principal component analysis (step S105). Such information becomes the HRTF parameters (step S106).
[0121]
When spherical harmonic fitting is performed after frequency decimation, the value of f above becomes the representative frequency after decimation. Also, the HRTF learning unit 136 may perform frequency thinning after spherical harmonic fitting. Moreover, the method of spatially compressing dimensions is not limited to linear combinations such as spherical harmonics and principal component analysis, and any technique may be used. For example, the HRTF learning unit 136 may use a nonlinear method such as kernel principal component analysis. Also, the HRTF learning unit 136 may change the truncation order N of the spherical harmonic function according to the frequency f, and use a value such as N(f). In addition, there may be coefficients h nm that are not used in the dimensions and orders from 0 to N. Further, the HRTF learning unit 136 may obtain the left and right HRTFs, respectively, or may obtain the left and right HRTFs after conversion into the sum and difference of the left and right HRTFs. Also, the HRTF to be fitted may be obtained by performing various transformations such as the absolute value of the amplitude or its logarithmic representation.
[0122]
Subsequently, the HRTF learning unit 136 can decode the HRTF by performing processing in reverse order from step S101 to step S106. First, the HRTF learning unit 136 acquires HRTF parameters (step S107), and restores dimensional compression by principal component analysis (step S108). Further, the HRTF learning unit 136 performs spherical harmonic reconstruction processing (step S109) and performs frequency interpolation (step S110). Furthermore, the HRTF learning unit 136 obtains amplitude characteristics (step S111) and performs minimum phase reconstruction (step S112). Various known techniques may be used for minimum phase reconstruction. For example, the HRTF learning unit 136 performs an inverse Fast Fourier Transform (IFFT) on the logarithm of the function H'1 (θ, φ, k), which is a function obtained by folding and restoring the function H'2 after the Nyquist frequency. ) and take its real part. Furthermore, minimum phase reconstruction can be performed by performing appropriate windowing in this region, inverse Fourier transforming the exponential function, and taking the real part. For example, the following relational expressions (8) are respectively established.
[0123]
[Number 8]

[0124]
Note that the HRTF learning unit 136 adds an estimated ITD (Interaural Time Difference) or an ITD prepared in advance to the left and right HRIRs (hm) that have undergone minimum phase restoration. may Note that the ITD is obtained by the following equations (9) and (10), for example, based on the difference between the group delays of the left and right HRIRs.
[0125]
[Number 9]

[0126]
[Number 10]

[0127]
Alternatively, the ITD may be calculated by obtaining the cross-correlation on the time axis on the left and right, and defining the time at which the correlation coefficient is maximum as the ITD. In this case, ITD is obtained by, for example, the following formulas (11) and (12).
[0128]
[Number 11]

[0129]
[number 12]

[0130]
For example, the HRTF learning unit 136 uses a relational expression such as the following formula (13) when the left HRIR is delayed from the right by d samples.
[0131]
[number 13]

[0132]
At this time, h L in the above equation (13) is an impulse response that is longer than h m and L by d. delete. At this time, the HRTF learning unit 136 may process an arbitrary window, rectangular window, Hanning window, or the like. Note that the HRTF learning unit 136 may not only add ITD relatively for each direction, but also add a delay including a relative time difference between directions in the entire space. In that case, the HRTF learning unit 136 acquires information indicating the relative time difference between directions as well as the ITD. Further, when the ITD is a function of frequency, the HRTF learning unit 136 may add the ITD in the frequency domain, or may add the ITD after obtaining a representative value or average value. Then, the HRTF learning unit 136 obtains the HRTF by performing inverse Fourier transform after obtaining the HRIR in the original format.
[0133]
In this way, the HRTF learning unit 136 may perform compression to HRTF parameters with less information than the original HRTF, and perform HRTF learning model generation processing and HRTF calculation processing, which will be described later, in the compressed format. . In addition, as described above, HRTF compression utilizes auditory characteristics such as, for example, the fact that human hearing is not sensitive to phase changes, and preferentially thinning out frequencies that are less likely to affect auditory perception. dimensionality reduction. As a result, the HRTF learning unit 136 can speed up the information processing without impairing the auditory localization, which is the characteristic of the HRTF.
[0134]
Return to Figure 3 and continue the explanation. The estimation unit 140 performs estimation processing of the HRTF corresponding to the user based on the image transmitted from the user.
[0135]
The acquisition unit 141 acquires an image including a video of the user's ear. For example, the acquiring unit 141 acquires an ear image in which only the area around the user's ear is cut out from the image captured by the user terminal 10 .
[0136]
Further, the acquisition unit 141 may acquire ear parameters indicating characteristics of the ears included in the image by inputting the acquired ear image into the ear parameter estimation model.
[0137]
The calculation unit 142 is designed to output the HRTF corresponding to the ear when an image including the image of the ear is input.Using the trained model (HRTF learning model), an HRTF (personalized HRTF) corresponding to the user is calculated based on the image acquired by the acquisition unit 141 .
[0138]
Specifically, the calculation unit 142 calculates the personalized HRTF corresponding to the user by inputting the ear parameters acquired by the acquisition unit 141 into the HRTF learning model.
[0139]
When calculating the personalized HRTF, the calculating unit 142 may calculate the HRTF by first calculating the HRTF parameters and then decoding the calculated HRTF parameters. By performing a series of processes in a state in which the information amount of the HRTF is compressed in this way, the calculation unit 142 can speed up the process. In addition, since the calculation unit 142 can avoid outputting strange HRTFs that are not expressed in the HRTF reduction model, stable output can be performed.
[0140]
The provision unit 143 provides the HRTF calculated by the calculation unit 142 to the user via the network N.
[0141]
Here, the flow of processing for estimating HRTF from an image will be described using FIG. FIG. 13 is a diagram illustrating a flow of HRTF estimation processing according to the present disclosure.
[0142]
FIG. 13 shows an example in which the estimating unit 140 performs estimation processing of the HRTF corresponding to the ear included in the image based on the image transmitted from the user. In this case, the user uses the user terminal 10 to photograph his or her ears (more precisely, the head including the ears) (step S131). After that, the user terminal 10 specifies a range including the image of the ear from the captured photograph, cuts out the specified range, and performs preprocessing to obtain an ear image (step S132).
[0143]
When acquiring the ear image transmitted from the user, the acquiring unit 141 inputs the acquired ear image to the trained model. Specifically, the acquisition unit 141 inputs the ear image to the ear parameter estimation model. The ear parameter estimation model outputs an ear parameter corresponding to the image of the ear included in the ear image as a feature amount representing the ear image. Thereby, the acquiring unit 141 acquires ear parameters corresponding to the image (step S133).
[0144]
The calculation unit 142 inputs the acquired ear parameters to the HRTF learning model, and calculates a personalized HRTF corresponding to the ear image (step S133). The providing unit 143 provides (transmits) the calculated personalized HRTF to the user terminal 10, which is the source of the image (step S134).
[0145]
In this way, when various models are generated by the learning unit 131, the information processing apparatus 100 can perform processing from acquisition of ear images to provision of personalized HRTFs as a series of processes. Thereby, the information processing apparatus 100 can improve the user's convenience regarding the provision of the HRTF.
[0146]
In the example of FIG. 13, a combination of an ear parameter estimation model and an HRTF learning model is shown as an example of a trained model, but the combination of trained models is not limited to this example. The trained model may be a combination of the ear parameter estimation model and the HRTF learning model, or a single model that performs processing corresponding to the ear parameter estimation model and the HRTF learning model. may be configured as
[0147]
[1-4. Configuration of user terminal according to first embodiment]
As shown in FIG. 13, in the first embodiment, the user terminal 10 captures the profile of the user, generates an ear image, and the like. Here, the configuration of the user terminal 10 according to the first embodiment will be described. FIG. 14 is a diagram showing a configuration example of the user terminal 10 according to the first embodiment of the present disclosure. As shown in FIG. 14 , the user terminal 10 has a communication section 11 , an input section 12 , a display section 13 , a detection section 14 , a storage section 15 and a control section 16 .
[0148]
The communication unit 11 is implemented by, for example, a NIC. The communication unit 11 is connected to the network N by wire or wirelessly, and transmits and receives information to and from the information processing apparatus 100 and the like via the network N.
[0149]
The input unit 12 is an input device that receives various operations from the user. For example, the input unit 12 is implemented by operation keys or the like provided on the user terminal 10 . The display unit 13 is a display device for displaying various information. For example, the display unit 13 is implemented by a liquid crystal display or the like. Note that when a touch panel is adopted for the user terminal 10, part of the input unit 12 and the display unit 13 are integrated.
[0150]
The detection unit 14 is a general term for various sensors, and detects various types of information regarding the user terminal 10 . Specifically, the detection unit 14 detects user operations on the user terminal 10, location information where the user terminal 10 is located, information on equipment connected to the user terminal 10, environment in the user terminal 10, and the like. do.
[0151]
Also, the detection unit 14 has a lens and an image sensor for photographing as an example of a sensor. That is, the detection unit 14 functions as a camera, for example, when the user activates an application that operates the shooting function.
[0152]
The storage unit 15 stores various information. The storage unit 15 is realized by, for example, a semiconductor memory device such as a RAM or flash memory, or a storage device such as a hard disk or an optical disk. The storage unit 15 stores, for example, images captured by the user.
[0153]
The control unit 16 is a controller, and is implemented, for example, by executing various programs stored in a storage device inside the user terminal 10 using the RAM as a work area by means of a CPU, MPU, or the like. Also, the control unit 16 is a controller, and is realized by an integrated circuit such as ASIC or FPGA, for example.
[0154]
As shown in FIG. 14, the control unit 16 includes an acquisition unit 161, a preprocessing unit 162, a transmission unit 164, and a reception unit 165, and implements or executes the information processing functions and actions described below. do. Further, the preprocessing unit 162 has a posture detection unit 163A and an ear detection unit 163B. Note that the internal configuration of the control unit 16 is not limited to the configuration shown in FIG. 14, and may be another configuration as long as it performs the information processing described later.
[0155]
The acquisition unit 161 acquires various types of information. For example, the acquisition unit 161 acquires an image captured by the detection unit 14 .
[0156]
The posture detection unit 163A reads the image acquired by the acquisition unit 161 and detects the user's posture included in the image.
[0157]
The ear detection unit 163B detects a range (ear image) including the user's ears included in the image based on the user's posture detected by the posture detection unit 163A. Specifically, ear detection unit 163B identifies an image of the user's ears from an image including an image of the entire head of the user, and detects the identified range as an ear image.
[0158]
For example, the ear detection unit 163B identifies the range including the ear image based on the relationship between the feature points of the user's head included in the entire image and the user's posture.
[0159]
In addition, when posture detection unit 163A or ear detection unit 163B cannot specify the range including the image of the ears based on the relationship between the feature points of the user's head included in the entire image and the user's posture, The user may be newly requested to obtain an image that is different from the image and includes an image of the entire head of the user. Specifically, the posture detection unit 163A or the ear detection unit 163B displays a message on the display unit 13 indicating that there is a possibility that the information processing according to the present disclosure may not be appropriately performed in the profile image captured by the user. and prompts the user to redo the shooting. Posture detection unit 163A or ear detection unit 163B is not limited to the case where the range in which the image of the ear is included cannot be specified. Even when the distance exceeds the threshold, the user may be prompted to retake the image. Alternatively, the posture detection unit 163A or the ear detection unit 163B may generate correction information for correcting the user's posture and position in the image instead of detecting the user's ear image as preprocessing. The correction information is, for example, information that instructs the amount of rotation of the range including the ear image according to the inclination and rotation of the feature points of the user's head. Such information is generated based on the user's posture, the positional relationship between the profile of the user and the positions of the detected ears, and the like, as will be described later. In this case, the posture detection unit 163A or the ear detection unit 163B may specify the image of the user's ears by correcting the rotation of the entire image based on the correction information, and detect the specified range as the ear image. Also, posture detection section 163A or ear detection section 163B may transmit the entire image to information processing apparatus 100 together with the generated correction information. In this case, the information processing apparatus 100 performs preprocessing to specify the image of the user's ear by correcting the rotation of the entire image based on the correction information transmitted together with the entire image, and to detect the specified range as the ear image. Perform on own device.
[0160]
Here, the flow of preprocessing executed by the preprocessing unit 162 (posture detection unit 163A and ear detection unit 163B) will be described with reference to FIG. FIG. 15 is a diagram showing the flow of detection processing according to the present disclosure.
[0161]
As shown in FIG. 15, the acquisition unit 161 acquires the entire image 50 when the user captures the profile of the user (step S141).
[0162]
The posture detection unit 163A detects the profile of the user in the acquired whole image 50 (step S142). For example, the posture detection unit 163A uses a known technique such as human face detection processing to identify a range in which the image of the user's profile is included in the entire image 50 .
[0163]
Here, as shown in image 51, the posture detection unit 163A detects feature points included in the profile of the user. For example, the posture detection unit 163A detects a portion of the side face that protrudes in the horizontal direction (specifically, the top of the user's nose, etc.), the top of the head, the position of the mouth, the position of the chin, and the like. Detect feature points. Also, the posture detection unit 163A detects the positions of the user's ears and sideburns, etc. from the information on the boundary between the hair and the skin. Also, the posture detection unit 163A detects the position of the user's eyes, etc. from the color information of the image of the side face.
[0164]
Then, the posture detection unit 163A detects the user's posture based on the detected feature points (step S143). For example, the posture detection unit 163A detects the posture of the user's head from the three-dimensional arrangement of feature points as shown in the image 54 .
[0165]
Such posture detection processing is to prevent the posture in the ear image sent by the user from significantly inconsistent with the posture of the 3D model used during learning. That is, when an image whose posture is significantly different from that of the 3D model is transmitted from the user terminal 10, there is a possibility that the information processing apparatus 100 cannot appropriately recognize the image of the ear due to discrepancies between the learning data and the transmitted ear image. because there is
[0166]
For this reason, the posture detection unit 163A determines whether or not the average value of the angles at the time of rendering in the head 3D model 55 used for learning and the angle obtained in the image 54 are within a predetermined threshold. It is determined whether or not the image was properly captured (step S144). For example, when learning an ear parameter estimation model, the angle φ is assumed to be within a predetermined numerical value. Similarly, when learning the ear parameter estimation model, it is assumed that the angle θ between the direction of the camera and the line segment connecting the tip of the nose and the predetermined position of the ear is within a predetermined numerical value. This is because the ear image used for learning does not greatly deviate from the image showing the profile of a person in order to improve the image recognition accuracy. In other words, posture detection section 163A detects that the image transmitted by the user is the same as the image at the time of learning.It is determined whether or not the image maintains an angle such that it can be recognized as an image showing a face in profile.
[0167]
When the posture detection unit 163A determines that the user has not properly captured the image (for example, when the user's face exceeds a predetermined threshold and the tip of the nose is pointing downward), the posture detection unit 163A outputs a message instructing to redo the image capturing. is displayed, and a newly captured image is obtained (step S145).
[0168]
On the other hand, when it is determined that the user is appropriately photographing (step S146), the ear detection unit 163B identifies the range 57 including the image of the ear from the image 56, and cuts out the range 57 (step S147). ). Thereby, the ear detection unit 163B acquires the ear image 58. FIG.
[0169]
By performing the detection process shown in FIG. 15, the information processing apparatus 100 determines whether the user's ears are tilted due to poor shooting conditions or whether the user's ears are actually tilted. , HRTF can be calculated.
[0170]
In addition, as described above, the user terminal 10 cuts out the ear image from the full image of the profile of the user, so that only the ear image can be sent and processed instead of the full image including the user's face. As a result, the user terminal 10 can prevent personal information from leaking, and can improve the security of information processing. Note that the user terminal 10 detects the user's ears included in the image by using not only the above-described detection method but also image recognition technology based on machine learning, etc., and performs a process of cutting out the ear image from the entire profile image. you can go
[0171]
Return to Fig. 14 and continue the explanation. Transmitter 164 transmits to information processing apparatus 100 an ear image generated based on the range detected by ear detector 163B.
[0172]
The receiving unit 165 receives the personalized HRTF provided from the information processing device 100 . For example, the receiving unit 165 can realize 3D sound or the like optimized for an individual user by convolving the received personalized HRTF with music or voice using a voice reproduction application or the like.
[0173]
(2. Second embodiment)
Next, a second embodiment will be described. In the above-described first embodiment, an example has been described in which the user terminal 10 generates an ear image by cutting out only the image of the ear from the image captured by the user. The information processing apparatus 100A according to the second embodiment performs a process of clipping out only the image of the ear in its own apparatus instead of the user terminal 10 .
[0174]
The configuration of an information processing apparatus 100A according to the second embodiment will be described using FIG. FIG. 16 is a diagram illustrating a configuration example of an information processing device 100A according to the second embodiment of the present disclosure. As shown in FIG. 16, the information processing apparatus 100A further includes a preprocessing section 144 (a posture detection section 145A and an ear detection section 145B) as compared with the first embodiment.
[0175]
The posture detection unit 145A performs the same processing as the posture detection unit 163A according to the first embodiment. Also, the ear detection unit 145B performs the same processing as the ear detection unit 163B according to the first embodiment. In other words, the information processing apparatus 100A according to the second embodiment executes the preprocessing that was executed by the user terminal 10 according to the first embodiment.
[0176]
In the second embodiment, the acquisition unit 141 acquires from the user terminal 10 an entire image of the profile of the user. Then, the posture detection unit 145A and the ear detection unit 145B perform processing similar to that described with reference to FIG. 15 to generate an ear image from the entire image. The calculator 142 calculates a personalized HRTF from the ear images generated by the posture detector 145A and the ear detector 145B.
[0177]
As described above, according to the information processing apparatus 100A according to the second embodiment, the user can receive the personalized HRTF simply by photographing and transmitting an image. In addition, according to the configuration of the second embodiment, it is not necessary to execute preprocessing in the user terminal 10, so it is possible to reduce the processing load of the user terminal 10, for example. In general, it is assumed that the processing speed of the server device (information processing device 100) is faster than that of the user terminal 10. Therefore, according to the configuration of the second embodiment, the entire information processing according to the present disclosure speed can be improved. When the correction information is transmitted together with the entire image, the posture detection unit 145A and the ear detection unit 145B specify the image of the user's ears by correcting the rotation of the entire image based on the correction information included in the entire image. and the specified range may be detected as the ear image.
[0178]
(3. Other embodiments)
The processing according to each embodiment described above may be implemented in various different forms other than the above embodiments.
[0179]
Further, among the processes described in each of the above embodiments, all or part of the processes described as being performed automatically can be performed manually, or the processes described as being performed manually can be performed manually. can also be performed automatically by known methods. In addition, information including processing procedures, specific names, various data and parameters shown in the above documents and drawings can be arbitrarily changed unless otherwise specified. For example, the various information shown in each drawing is not limited to the illustrated information.
[0180]
Also, each component of each device illustrated is functionally conceptual and does not necessarily need to be physically configured as illustrated. In other words, the specific form of distribution and integration of each device is not limited to the one shown in the figure, and all or part of them can be functionally or physically distributed and integrated in arbitrary units according to various loads and usage conditions. Can be integrated and configured.
[0181]
In addition, the above-described embodiments and modifications can be appropriately combined within a range that does not contradict the processing content.
[0182]
In addition, the effects described in this specification are only examples and are not limited, and other effects may be provided.
[0183]
(4. Hardware configuration)
The information equipment such as the information processing apparatus 100 and the user terminal 10 according to each embodiment described above is implemented by a computer 1000 configured as shown in FIG. 17, for example. The information processing apparatus 100 according to the first embodiment will be described below as an example. FIG. 17 is a hardware configuration diagram showing an example of a computer 1000 that implements the functions of the information processing apparatus 100. As shown in FIG. The computer 1000 has a CPU 1100 , a RAM 1200 , a ROM (Read Only Memory) 1300 , a HDD (Hard Disk Drive) 1400 , a communication interface 1500 and an input/output interface 1600 . Each part of computer 1000 is connected by bus 1050 .
[0184]
The CPU 1100 operates based on programs stored in the ROM 1300 or HDD 1400 and controls each section. For example, the CPU 1100 loads programs stored in the ROM 1300 or HDD 1400 into the RAM 1200 and executes processes corresponding to various programs.
[0185]
The ROM 1300 stores a boot program such as BIOS (Basic Input Output System) executed by the CPU 1100 when the computer 1000 is started, and programs dependent on the hardware of the computer 1000.
[0186]
The HDD 1400 is a computer-readable recording medium that non-temporarily records programs executed by the CPU 1100 and data used by such programs. Specifically, HDD 1400 is a recording medium that records an information processing program according to the present disclosure, which is an example of program data 1450 .
[0187]
A communication interface 1500 is an interface for connecting the computer 1000 to an external network 1550 (for example, the Internet). For example, CPU 1100 receives data from another device via communication interface 1500, and transmits data generated by CPU 1100 to another device.
[0188]
The input/output interface 1600 is an interface for connecting the input/output device 1650 and the computer 1000 . For example, the CPU 1100 receives data from input devices such as a keyboard and mouse via the input/output interface 1600 . The CPU 1100 also transmits data to an output device such as a display, speaker, or printer via the input/output interface 1600 . Also, the input/output interface 1600 may function as a media interface for reading a program or the like recorded on a predetermined recording medium. Media include, for example, optical recording media such as DVD (Digital Versatile Disc) and PD (Phase change rewritable disk), magneto-optical recording media such as MO (Magneto-Optical disk), tape media, magnetic recording media, semiconductor memories, etc. is.
[0189]
For example, when the computer 1000 functions as the information processing apparatus 100 according to the first embodiment, the CPU 1100 of the computer 1000 implements the functions of the control unit 130 and the like by executing the information processing program loaded on the RAM 1200. do. The HDD 1400 also stores an information processing program according to the present disclosure and data in the storage unit 120 . Although CPU 1100 reads and executes program data 1450 from HDD 1400 , as another example, these programs may be obtained from another device via external network 1550 .
[0190]
This technology can also be configured as follows.
(1)
an acquisition unit that acquires the first image including the image of the user's ear;
Based on the first image acquired by the acquisition unit using a trained model that is trained to output a head-related transfer function corresponding to the ear when an image including an image of the ear is input, a calculation unit that calculates a head-related transfer function corresponding to a user;
Information processing device equipped with
(2)
The acquisition unit
Acquire an ear parameter, which is a variable indicating the characteristics of the ear included in the first image,
The calculation unit
By inputting the ear parameters into the trained model, a head-related transfer function corresponding to the user is calculated
The information processing device according to (1) above.
(3)
The acquisition unit
Acquire the ear parameters of the ear included in the first image using an ear parameter estimation model trained to output ear parameters corresponding to the ear when an image including the image of the ear is input.
The information processing device according to (2) above.
(4)
A first learning unit that generates the ear parameter estimation model by learning the relationship between the image containing the image of the ear and the ear parameter of the ear
The information processing apparatus according to (3), further comprising:
(5)
The first learning unit
The ear parameter estimation model is generated by learning the relationship between the ear image obtained by rendering the three-dimensional data of the ear generated based on the ear parameter and the ear parameter.
The information processing device according to (4) above.
(6)
The first learning unit
By learning the relationship between a plurality of ear images obtained by changing the texture of the three-dimensional data of the ear or head, the camera angle in rendering, or the brightness in rendering, and the ear parameter common to the plurality of ear images, generate an ear parameter estimation model
The information processing device according to (5) above.
(7)
A second learning unit that generates the trained model by learning the relationship between the image containing the image of the ear and the head-related transfer function corresponding to the ear
The information processing apparatus according to any one of (1) to (6), further comprising:
(8)
The second learning unit
Acoustic simulation is performed on 3D data that is a combination of the 3D data of the ear generated based on the ear parameters and the 3D data of the head.and generate the trained model by learning the relationship between the head-related transfer function obtained by the acoustic simulation and the ear parameter.
The information processing device according to (7) above.
(9)
The second learning unit
The learned model is generated by compressing the amount of information in the head-related transfer function obtained by the acoustic simulation and learning the relationship between the compressed head-related transfer function and the ear parameter.
The information processing device according to (8) above.
(10)
The second learning unit
 Set the listening point of the three-dimensional data of the ear generated based on the ear parameter, and perform the acoustic simulation using the set listening point
The information processing device according to (8) or (9) above.
(11)
further comprising a preprocessing unit that identifies an image of the user's ears from a second image that includes an image of the entire head of the user, and detects the identified range as the first image,
The acquisition unit
Acquire the first image detected by the preprocessing unit
The information processing device according to any one of (1) to (10) above.
(12)
The preprocessing unit
The range is specified based on the relationship between the feature points of the user's head included in the second image and the user's posture
The information processing device according to (11) above.
(13)
The preprocessing unit
If the range cannot be specified based on the relationship between the feature points of the user's head included in the second image and the user's posture, an image different from the second image, the user's head Make a new request to acquire an image that includes the entire video
The information processing device according to (12) above.
(14)
The preprocessing unit
By correcting the rotation of the second image based on the correction information included in the second image, the image of the user's ears is specified, and the specified range is detected as the first image.
The information processing device according to any one of (11) to (13) above.
(15)
"The computer
 Acquire the first image containing the image of the user's ear,
Corresponding to the user based on the first image acquired using a trained model that is trained to output a head-related transfer function corresponding to the ear when an image including an image of the ear is input. Calculate the head-related transfer function for
Information processing method.
(16)
the computer,
an acquisition unit that acquires the first image including the image of the user's ear;
Based on the first image acquired by the acquisition unit using a trained model that is trained to output a head-related transfer function corresponding to the ear when an image including an image of the ear is input, a calculation unit that calculates a head-related transfer function corresponding to a user;
Information processing program to function as
(17)
An information processing system including an information processing device and a user terminal,
 The user terminal
a preprocessing unit that identifies the image of the user's ears from the second image that includes the image of the entire head of the user and detects the identified range as the first image;
a transmission unit that transmits the first image detected by the preprocessing unit to the information processing device,
The processing device is
an acquisition unit that acquires the first image including the image of the user's ear;
Based on the first image acquired by the acquisition unit using a trained model that is trained to output a head-related transfer function corresponding to the ear when an image including an image of the ear is input, a calculation unit that calculates a head-related transfer function corresponding to a user;
An information processing system equipped with
Code explanation
[0191]
1 Information processing system
10 User terminal
100 Information processing device
110 communication department
120 storage unit
130 control unit
131 Learning Department
 132 Reception department
133 Ear model learning unit
134 image generator
135 ear parameter learning unit
136 HRTF learning department
140 estimation unit
141 Acquisition unit
142 calculator
143 Provision Department
144 preprocessing unit
 145A Attitude detection part
 145B Ear detector
161 Acquisition unit
162 preprocessing unit
 163A Attitude detection part
 163B Ear detector
164 transmission unit
165 receiver
The scope of the claims
[Claim 1]
an acquisition unit that acquires the first image including the image of the user's ear;
Based on the first image acquired by the acquisition unit using a trained model that is trained to output a head-related transfer function corresponding to the ear when an image including an image of the ear is input, a calculation unit that calculates a head-related transfer function corresponding to a user;
Information processing device equipped with
[Claim 2]
The acquisition unit
Acquire an ear parameter, which is a variable indicating the characteristics of the ear included in the first image,
The calculation unit
By inputting the ear parameters into the trained model, a head-related transfer function corresponding to the user is calculated
The information processing device according to claim 1.
[Claim 3]
The acquisition unit
Acquire the ear parameters of the ear included in the first image using an ear parameter estimation model trained to output ear parameters corresponding to the ear when an image including the image of the ear is input.
The information processing device according to claim 2.
[Claim 4]
A first learning unit that generates the ear parameter estimation model by learning the relationship between the image containing the image of the ear and the ear parameter of the ear
The information processing apparatus according to claim 3, further comprising:
[Claim 5]
The first learning unit
The ear parameter estimation model is generated by learning the relationship between the ear image obtained by rendering the three-dimensional data of the ear generated based on the ear parameter and the ear parameter.
The information processing device according to claim 4.
[Claim 6]
The first learning unit
By learning the relationship between a plurality of ear images obtained by changing the texture of the three-dimensional data of the ear or head, the camera angle in rendering, or the brightness in rendering, and the ear parameter common to the plurality of ear images, generate an ear parameter estimation model
The information processing device according to claim 5.
[Claim 7]
A second learning unit that generates the trained model by learning the relationship between the image containing the image of the ear and the head-related transfer function corresponding to the ear
The information processing apparatus according to claim 4, further comprising:
[Claim 8]
The second learning unit
Acoustic simulation is performed on three-dimensional data obtained by synthesizing the three-dimensional data of the ear and the three-dimensional data of the head generated based on the ear parameters, and the head-related transfer function obtained by the acoustic simulation and the Generating the trained model by learning relationships with ear parameters
The information processing device according to claim 7.
[Claim 9]
The second learning unit
The learned model is generated by compressing the amount of information in the head-related transfer function obtained by the acoustic simulation and learning the relationship between the compressed head-related transfer function and the ear parameter.
The information processing device according to claim 8.
[Claim 10]
The second learning unit
 Set the listening point of the three-dimensional data of the ear generated based on the ear parameter, and perform the acoustic simulation using the set listening point
The information processing device according to claim 8.
[Claim 11]
further comprising a preprocessing unit that identifies an image of the user's ears from a second image that includes an image of the entire head of the user, and detects the identified range as the first image,
The acquisition unit
Acquire the first image detected by the preprocessing unit
The information processing device according to claim 1.
[Claim 12]
The preprocessing unit
The range is specified based on the relationship between the feature points of the user's head included in the second image and the user's posture
The information processing device according to claim 11.
[Claim 13]
The preprocessing unit
If the range cannot be specified based on the relationship between the feature points of the user's head included in the second image and the user's posture, an image different from the second image, the user's head Make a new request to acquire an image that includes the entire video
The information processing device according to claim 12.
[Claim 14]
The preprocessing unit
By correcting the rotation of the second image based on the correction information included in the second image, the image of the user's ears is specified, and the specified range is detected as the first image.
The information processing device according to claim 11.
[Claim 15]
"The computer
 Acquire the first image containing the image of the user's ear,
Corresponding to the user based on the first image acquired using a trained model that is trained to output a head-related transfer function corresponding to the ear when an image including an image of the ear is input. Calculate the head-related transfer function for
Information processing method.
[Claim 16]
the computer,
an acquisition unit that acquires the first image including the image of the user's ear;
Based on the first image acquired by the acquisition unit using a trained model that is trained to output a head-related transfer function corresponding to the ear when an image including an image of the ear is input, a calculation unit that calculates a head-related transfer function corresponding to a user;
Information processing program to function as

Documents

Application Documents

# Name Date
1 202117009272-TRANSLATIOIN OF PRIOIRTY DOCUMENTS ETC. [05-03-2021(online)].pdf 2021-03-05
2 202117009272-STATEMENT OF UNDERTAKING (FORM 3) [05-03-2021(online)].pdf 2021-03-05
3 202117009272-PRIORITY DOCUMENTS [05-03-2021(online)].pdf 2021-03-05
4 202117009272-POWER OF AUTHORITY [05-03-2021(online)].pdf 2021-03-05
5 202117009272-FORM 1 [05-03-2021(online)].pdf 2021-03-05
6 202117009272-DRAWINGS [05-03-2021(online)].pdf 2021-03-05
7 202117009272-DECLARATION OF INVENTORSHIP (FORM 5) [05-03-2021(online)].pdf 2021-03-05
8 202117009272-COMPLETE SPECIFICATION [05-03-2021(online)].pdf 2021-03-05
9 202117009272-Proof of Right [15-04-2021(online)].pdf 2021-04-15
10 202117009272-Proof of Right [24-05-2021(online)].pdf 2021-05-24
11 202117009272-FORM 3 [24-06-2021(online)].pdf 2021-06-24
12 202117009272.pdf 2021-10-19
13 202117009272-FORM 18 [18-08-2022(online)].pdf 2022-08-18
14 202117009272-FER.pdf 2022-11-15
15 202117009272-FORM 3 [13-02-2023(online)].pdf 2023-02-13
16 202117009272-OTHERS [15-05-2023(online)].pdf 2023-05-15
17 202117009272-FER_SER_REPLY [15-05-2023(online)].pdf 2023-05-15
18 202117009272-DRAWING [15-05-2023(online)].pdf 2023-05-15
19 202117009272-CORRESPONDENCE [15-05-2023(online)].pdf 2023-05-15
20 202117009272-COMPLETE SPECIFICATION [15-05-2023(online)].pdf 2023-05-15
21 202117009272-CLAIMS [15-05-2023(online)].pdf 2023-05-15
22 202117009272-ABSTRACT [15-05-2023(online)].pdf 2023-05-15
23 202117009272-US(14)-HearingNotice-(HearingDate-19-11-2024).pdf 2024-10-10
24 202117009272-FORM-26 [18-11-2024(online)].pdf 2024-11-18
25 202117009272-Correspondence to notify the Controller [18-11-2024(online)].pdf 2024-11-18
26 202117009272-Written submissions and relevant documents [04-12-2024(online)].pdf 2024-12-04
27 202117009272-Information under section 8(2) [04-12-2024(online)].pdf 2024-12-04
28 202117009272-PatentCertificate11-06-2025.pdf 2025-06-11
29 202117009272-IntimationOfGrant11-06-2025.pdf 2025-06-11

Search Strategy

1 SearchE_14-11-2022.pdf

ERegister / Renewals