Abstract: The information processing system is equipped with: a recognizing unit that recognizes a first object and a second object on the basis of a signal detected by a plurality of sensors placed in the vicinity of a specific user; an identification unit that identifies the first object and second object recognized by the recognizing unit; an estimating unit that estimates the position of the specific user according to a signal detected by any of the plurality of sensors; and a signal processing unit that processes each signal acquired from sensors in the vicinity of the first and second objects identified by the identifying unit so that the sound when output from a plurality of actuators placed in the vicinity of the specific user is located in the vicinity of the position of the specific user estimated by the estimating unit.
Description
Title of Invention
INFORMATION PROCESSING SYSTEM AND STORAGE MEDIUM
5
Technical Field
[OOOl]
The present invention relates to an information processing system and a
storage medium.
10
Background Art
[0002]
In recent years, various technologies have been proposed in data
communication fields. For example, Patent Literature 1 below proposes technology
15 related to a Machine-to-Machine (M2M) solution. To be specific, the remote
management system written in Patent Literature 1 uses the Internet protocol (IP)
multimedia subsystem (IMS) platform (IS), and through disclosure of presence
information by a device or instant messaging between a user and a device, an
interaction between an authorized user client (UC) and a device client is achieved.
20 [0003]
On the other hand, in acoustic technology fields, various types of array
speakers that can elnit acoustic beams are being developed. For example, Patent
Literature 2 below describes array speakers in which a plurality of speakers forming
a common wave front are attached to a cabinet and which control amounts of delay
25 and levels of the sounds given out from the respective speakers. Further, Patent
Literature 2 below describes that array microphones having the same principle are
being developed. The array microphones can voluntarily set the sound acquisition
point by adjusting the levels and amounts of delay of output signals of the respective
microphones, and thus are capable of acquiring the sound more effectively.
30
Citation List
Patent Literature
[0004]
Patent Literature 1:
Patent Literature 2:
5
Summa~yo f Invention
Technical Problem
[0005]
However, Patent Literature 1 and Patent Literature 2 described above do not
10 mention anything about technology or a communication method that is understood as
means for achieving an augmentation of a user's body by placing many image
sensors, microphones, speakers, and the like over a large area.
[0006]
Accordingly, tlie present disclosure proposes an information processing
15 system and a storage medium which are novel and improved, and which are capable
of providing a feeling of being absorbed in a third space at a time of causing the
space surrounding the user to cooperate with another space.
Solution to Problem
20 [0007]
According to the present disclosure, there is provided an information
processing system including a recognizing unit configured to recognize a first target
and a second target on the basis of signals detected by a plurality of sensors a~~anged
around a specific user, an identifying unit configured to identify the first target and
25 the second target recognized by the recognizing unit, an estimating unit configured to
estimate a position of tlie specific user in accordance with the a signal detected by
any one of the plurality of sensors, and a signal processing unit configured to process
each of signals acquired from sensors around the first target and the second target
identified by the identifying unit in a manner that, when being output from a plurality
30 of actuators arranged around the specific user, the signals are localized near tlie
position of the specific user estimated by the estimating unit.
[0008] --
According to the present disclosure, there is provided an information
processing system including a recognizing unit configured to recognize a first target
and a second target on the basis of signals detected by sensors around a specific user,
5 an identifying unit configured to identify the first target and the second target
recognized by the recognizing unit, and a signal processing unit configured to
generate signals to be output from actuators around the specificuser on the basis of
signals acquired by a plurality of sensors arranged around the first target and the
second target identified by the identifying unit.
10 [0009]
According to the present disclosure, there is provided a storage medium
having a program stored therein, the program being for causing a computer to
function as a recognizing unit configured to recognize a first target and a second
target on the basis of signals detected by a plurality of sensors a~ranged around a
15 specific user, an identifying unit configured to identify the first target and the second
target recognized by the recognizing unit, an estimating unit configured to estimate a
position of the specific user in accordance with the a signal detected by any one of
the plurality of sensors, and a signal processing unit configured to process each of
signals acquired from sensors around the first target and the second target identified
20 by the identifying unit in a manner that, when being output fsom a plurality of
actuators arranged around the specific user, the signals are localized near the position
of the specific user estimated by the estimating unit.
[OOl 01
According to the present disclosure, there is provided a storage medium
25 having a program stored therein, the program being for causing a computer to
function as a recognizing unit configured to recognize a first target and a second
target on the basis of signals detected by sensors around a specific user, an
identifying unit configured to identify the first target and the second target
recognized by the recognizing unit, and a signal processing unit configured to
30 generate signals to be output from actuators around the specific user on the basis of
signals acquired by a plurality of sensors arranged around the first target and the
second target identified by the identifying unit.
Advantageous Effects of Invention
[OOll]
5 According to the present disclosure as described above, a feeling of being
absorbed in a third space can be provided at a time of causing a space surrounding a
user to cooperate with another space.
Brief Description of Drawings
10 [0012]
[FIG. 11 FIG. 1 is a diagram illustrating an outline of an acoustic system according to
an embodiment of the present disclosure.
[FIG. 21 FIG. 2 is a diagram showing a system configuration of an acoustic system
according to an embodiment of the present disclosure.
15 [FIG. 31 FIG. 3 is a block diagram showing a configuration of a signal processing
apparatus according to the present embodiment.
[FIG. 41 FIG. 4 is a diagram illustrating shapes of acoustically closed surfaces
according to the present embodiment.
[FIG. 51 FIG. 5 is a block diagram showing a configuration of a management server
20 according to the present embodiment.
[FIG. 61 FIG. 6 is a flowchart showing a basic process of the acoustic system
according to the present embodiment.
[FIG. 71 FIG. 7 is a flowchart showing a command recognition process according to
the present embodiment.
25 [FIG. 81 FIG. 8 is a flowchart showing a sound acquisition process according to the
present embodiment.
[FIG. 91 FIG. 9 is a diagram illustratirlg constructio~ol f a sound field of a third space
according to the present embodiment.
[FIG. 101 FIG. 10 is a diagram illustrating techlliques of constructing a sound field of
30 a site C.
[FIG. 111 FIG. 11 is a block diagram showing another configuration of the
manxtgement server according to the present embodiment. _
[FIG. 121 FIG. 12 is a diagram illustrating measurement of an acoustic parameter.
[FIG. 131 FIG. 13 is a diagram showing an arrangement of a plurality of microphones
in a measuring environment and an arrangement of a plurality of speakers in a
5 listening environment in comparison with each other.
[FIG. 141 FIG. 14 is a diagram illustrating shapes of closed surfaces in the measuring
environment according to the present embodiment.
[FIG. 151 FIG. 15 is a block diagram showing a configuration of a sound field
reproduction signal processing unit performing sound field construction for
10 providing a feeling of being absorbed in the site C.
[FIG. 16A] FIG. 16A is a diagram illustrating measurement of an impulse response in
the site C.
[FIG. 16B] FIG. 16B is a diagram illustrating an arithmetic operation using an
impulse response group performed by a matrix co~~volutiounn it according to the
15 present embodiment.
[FIG. 171 FIG. 17 is a flowchart showing a sound field reproduction process
according to the present enlbodiment.
[FIG. 18A] FIG. 18A is a diagram illustrating a case where a sound field constructed
in a site B is fixed.
20 [FIG. 18B] FIG. 18B is a diagram illustrating a case where the sound field
constructed in the site B is mobile.
[FIG. 19A] FIG. 19A is a diagram illustrating measurement in a measurement target
space.
[FIG. 19B] FIG. 19B is a diagram illustrating measurement in an anechoic room.
25 [FIG. 19C] FIG. 19C is a diagram illustrating reconstruction in a reproduction target
space.
[FIG. 201 FIG. 20 is a diagram showing another system configuration of the acoustic
system according to the present embodiment.
[FIG. 211 FIG. 21 is a diagram showing an example of a system configuration of an
30 autonomous acoustic system according to the present embodiment.
[FIG. 221 FIG. 22 is a block diagram showing a configuration of a device having the
autonomous acoustic system according to the present embodiment.
[FIG. 231 FIG. 23 is a flowchart showing an operation process of the autonomous
acoustic system according to the present embodiment.
[FIG. 241 FIG. 24 is a diagram illustrating change of operating devices in accordance
5 with movement of a user in the autonomous acoustic system according to the present
embodiment.
[FIG. 251 FIG. 25 is a diagram illustrating a case where a service is provided to a
plurality of users in the autonomous acoustic system according to the present
embodiment.
10
Description of Embodiments
[0013]
Hereinafter, preferred embodiments of the present disclosure will be
described in detail with reference to the appended drawings. Note that, in this
15 specification and the drawings, elements that have substantially the same function
and structure are denoted with the same reference signs, and repeated explanation is
omitted.
[0014]
The description will be given in the following order,
20 1. Outline of acoustic system according to embodiment of present disclosure
2. Basic configuration
2-1 .System configuration
2-2. Signal processing apparatus
2-3. Management server
3. Operation process
3-1. Basic process
3-2. Command recognition process
3-3. Sound acquisition process
4. Construction of sound field of third space
4-1. Configuration of management server
4-2. Configuration of sound field reproduction signal processing
unit
4-3. Sound field reproduction process
5. Supplement
6. Conclusion
5 [0015]
4. Outline of acoustic system according to embodiment of present
disclosur~
First, with reference to FIG. 1, an outline of an acoustic system (information
processing system) according to an embodiment of the present disclosure will be
10 described. FIG. 1 is a diagram illustrating an outline of an acoustic system
according to an embodiment of the present disclosure. As shown in FIG. 1, in the
acoustic system according to the present embodiment, let us assume the situation in
which a large number of sensors and actuators such as microphones 10, image
sensors (not shown), and speakers 20 are arranged everywhere in the world such as
15 rooms, houses, buildings, outdoor sites, regions, and countsies.
[0016]
In the exanlple shown in FIG. 1, on a road or the like in an outdoor area "site
A" at which a user A is currently located, a plurality of microphones IOA are
arranged as examples of the plurality of sensors and a plurality of speakers 20A are
20 arranged as examples of the plurality of actuators. Further, in an indoor area "site
B" at which a user B is currently located, a plurality of microphones 10B and a
plurality of speakers 20B are arranged on the walls, the floor, the ceiling, and the like.
Note that, in the sites A and B, motion sensors and image sensors (which are not
shown) may further be arranged as examples of the sensors.
25 [0017]
Here, the site A and the site B are connectable to each other though a
network, and the signals output fsom and input to the respective microphones and the
respective speakers of the site A and the signals output from and input to the
respective microphones and the respective speakers of the site B are transmitted and
30 received between the sites A and B.
[OOl S]
In this way, the acoustic system according to the present embodiment
reproduces in real time a voice or an image corresponding to a given target (person,
place, building, or the like) though a plurality of speakers and a plurality of displays
assanged around the user. Further, the acoustic system according to the present
5 embodiment can reproduce around the user in real time the voice of the user that has
been acquired by a plurality of microphones arranged around the user. In this way,
the acoustic system according to the present embodiment can cause a space
surrounding a user to cooperate with another space.
[0019]
10 Further, using the microphones 10, the speakers 20, the image sensors, and
the like arranged everywhere, indoor sites and outdoor sites, it becomes possible to
substantially augment over a large area the body such as the mouth, eyes, ears of the
user, and to achieve a new communication method.
[0020]
15 In addition, since microphones and image sensors are arranged everywhere
in the acoustic system according to the present embodiment, the user does not have
to carry a smartphone or a mobile phone terminal. The user specifies a given target
using a voice or a gesture, and can establish connection with a space surrounding the
given target. Hereinaftel; there will be briefly described the application of the
20 acoustic system according to the present embodiment in the case where the user A
located at the site A wants to have a conversation with the user B located at the site B.
[0021]
(Data collection process)
At the site A, a data collection process is continuously performed though
25 the plurality of microphones 10A, the plurality image sensors (not shown), the
plurality of hutnan sensors (not shown), and the like. Specifically, the acoustic
system according to the present embodiment collects voices acquired by the
microphones IOA, captured images obtained by the image sensors, or detection
results of the human sensors, and estimates the user's position on the basis of the
30 collected informatioa.
[0022]
Further, the acoustic system according to the present embodimentmay
select a microphone group arranged at the position at which the user's voice can be
sufficiently acquired on the basis of position information of the plurality of
microphones 10A which are registered in advance and the user's estimated position.
5 Further, the acoustic system according to the present embodiment performs a
microphone array process of a stream group of audio signals acquired by the selected
microphones. In particular, the acoustic system according to the present
embodiment may perform a delay-and-sum array in which a sound acquisition point
is focused on the user A's mouth and can form super directivity of an array
10 microphone. Thus, faint vocalizations such as the user A's muttering can be also
acquired.
[0023]
Further, the acoustic system according to the present embodiment
recognizes a command on the basis of the user A's acquired voice, and executes an
15 operation process according to the command. For example, when the user A located
at the site A says "I'd like to speak with B," the "call origination request to the user
B" is recognized as a command. In this case, the acoustic system according to the
present embodiment identifies the current position of tlie user B, and causes the site
B at which the user B is currently located to be connected with the site A at which the
20 user A is currently located. Through this operation, the user A can speak on the
telephone with the user B.
[0024]
(Object decompositio~p~ro cess)
An object decomposition process such as sound source separation
25 (separation of a noise component around the user A, a conversation of a person
around the user A, and the like), dereverberation, and a noiselecho process is
performed on audio signals (stream data) acquired by the plurality of microphones at
the site A during a telephone call. Through this process, stream data in which an
S/N ratio is high and a reverberant feeling is suppressed is transmitted to the site B.
30 [0025]
Considering a case in which the user A speaks while moving, the acoustic
system according_to the present embodiment can cope with this case bycontinuously
performing the data collection. Specifically, the acoustic system according to the
present embodiment continuously performs data collection on the basis of the
plurality of micropl~ones, the plurality of image sensors, the plurality of human
5 sensors, and the like, and detects a moving path of the user A or a direction in which
the user A is heading. Then, the acoustic system according to the present
embodiment continuously updates selection of an appropriate microphone group
arranged around the moving user A, and continuously performs the array microphone
process so that the sound acquisition point is constantly focused on the moving user
10 A's mouth. Through this operation, the acoustic system according to the present
embodiment can cope with a case in which the user A speaks while moving.
[0026]
Further, separately from stream data of a voice, a moving direction and the
direction of the user A or the like is converted into metadata and transmitted to the
15 site B together with the stream data.
[0027]
(Object synthesis)
Further, the stream data transmitted to the site B is reproduced through the
speakers arranged around the user B located at the site B. At this time, the acoustic
20 system according to the present embodiment performs data collection at the site B
through the plurality of microphones, the plurality of image sensors, and the plurality
of human sensors, estimates the user B's position on the basis of the collected data,
and selects an appropriate speaker group surrounding the user B through an
acoustically closed surface. The stream data transmitted to the site B is reproduced
25 through the selected speaker group, and an area inside the acoustically closed surface
is controlled as an appropriate sound field. In this disclosure, a surface formed such
that positions of a plurality of adjacent speakers or a plurality of adjacent
microphones, are connected to surround an object (the usel; for example) is referred
to conceptually as an "acoustically closed surface." Further, the "acoustically
30 closed surface" does not necessarily configure a perfect closed surface, and is
preferably configured to approximately surround the target object (the user, for
exampla ~ -~
[0028]
Further, the sound field may be appropriately selected by the user B. For
example, in the case where the user B designates the site A as the sound field, the
5 acoustic system according to the present embodiment reconstructs the environment
of the site A in the site B. Specifically, for example, the environment of the site A is
reconstructed in the site B on the basis of sound information as an ambience acquired
in real time and meta information related to the site A that has been acquired in
advance.
10 [0029]
Further, the acoustic system according to the present embodiment may
control the user A's audio image using the plurality of speakers 20B arranged around
the user B at the site B. In other words, the acoustic system according to the present
embodiment may reconstruct the user A's voice (audio image) in the user B's ear or
15 outside the acoustically closed surface by forming an array speaker (beam forming).
Further, the acoustic system according to the present embodiment inay cause the user
A's audio image to move around the user B according to the user A's actual
movement at the site B using nletadata of the moving path or the direction of the user
A.
20 [0030]
The outline of voice communication from the site A to the site B has been
described above in connection with respective steps of the data collection process,
the object decomposition process, and the object synthesis process, but of course, a
similar process is performed in voice communication fiom the site B to the site A.
25 Thus, two-way voice communication can be perfollned between the site A and the
site B.
[0031]
The outline of the acoustic system (information processing system)
according to an embodiment of the present disclosure has been described above.
30 Next, a configuration of the acoustic system according to the present embodiment
will be described in detail with reference to FIGS. 2 to 5.
100321
Q. Basic configuration>
[2-1.System configuration]
FIG. 2 is a diagram illustrating an overall configuration of the acoustic
5 system according to the present embodiment. As shown in FIG. 2, the acoustic
system includes a signal processing apparatus IA, a signal processing apparatus lB,
and a management server 3.
[0033]
The signal processing apparatus 1A and the signal processing apparatus IB
10 are connected to a network 5 in a wired/wireless mannel; and can transmit or receive
data to or from one another via the network 5. The management server 3 is
connected to the network 5, and the signal processing apparatus IA and the signal
processing apparatus 1B can transmit or receive data to or from the management
server 3.
15 [0034]
The signal processing apparatus IA processes signals input or output by the
plurality of microphones 1 OA and the plurality of speakers 20A arranged at the site A.
The signal processing apparatus IB processes signals input or output by the plurality
of microphones 10B and the plurality of speakers 20B arranged at the site B.
20 Further, when it is unnecessary to distinguish the signal processing apparatuses 1A
and 1B fiom one another, the signal processing apparatuses 1A and 1B are referred to
collectively as a "signal processing apparatus 1."
[0035]
The management server 3 has a function of performing a user authentication
25 process and managing a user's absolute position (current position). Further, the
management server 3 may also manage information (for exaniple, IP address)
representing a position of a place or a building.
[0036]
Thus, the signal processing apparatus 1 can send a query for access
30 destination inforn~ation (for example, IP address) of a given target (person, place,
building, or the like) designated by the user to the management server 3 and can
acquire the access destination information.
[0037]
12-2. Signal processing apparatus]
Next, a configuration of the signal processing apparatus I according to the
5 present embodiment will be described in detail. FIG. 3 is a block diagram showing
a configuration of the signal processing apparatus 1 according to the present
embodiment. As shown in FIG. 3, the signal processing apparatus 1 according to
the present embodiment includes a plurality of microphones 10 (array microphone),
an amplifying/analog-to-digital converter (ADC) unit 11, a signal processing unit 13,
10 a microphone position information database (DB) 15, a user position estimating unit
16, a recognizing unit 17, an identifying unit 18, a communication interface (IE) 19,
a speaker position information DB 21, an amplifyingldigital-to-analog converter
(DAC) unit 23, and a plurality of speakers 20 (array speaker). The components will
be described below.
15 100381
(Array microphone)
The plurality of microphones 10 are arranged throughout a certain area (site)
as described above. For example, the plurality of microphones 10 are arranged at
outdoor sites such as roads, electric poles, street lamps, houses, and outer walls of
20 buildings and indoor sites such as floors, walls, and ceilings. The plurality of
microphones 10 acquire ambient sounds, and output the acquired ambient sounds to
the amplifying/ADC unit 11.
100391
( ~ m p l i f ~ i n gu/n~itD) ~
25 The amplifyinglADC unit 11 has a function (amplifier) of amplifying
acoustic waves output from the plurality of microphones 10 and a function (ADC) of
converting an acoustic wave (analog data) into an audio signal (digital data). The
amplifying/ADC unit 11 outputs the converted audio signals to the signal processing
unit 13.
30 [0040]
(Signal processing unit)
The signal processing unit 13 has a function of processing the audio signals
acquiied by the microphones 10 and transmitted through the amplifying/ADC unit 11
and the audio signals reproduced by the speakers 20 through the DACIamplifying
unit 23. Furthel; the signal processing unit 13 according to the present embodiment
5 functions as a microphone array processing unit 131, a high S/N processing unit 133,
and a sound field reproduction signal processing unit 135.
[0041]
-Microphone array processing unit
The microphone array processing unit 131 performs directivity control such
10 that the user's voice is focused 011 (a sound acquisition position is focused on the
user's mouth) in the microphone array process for a plurality of audio signals output
from the amplifying/ADC unit 11.
100421
At this time, the microphone array processing unit 131 may select a
15 microphone group forming the acoustically closed surface surrounding the user
which is optimal for acquisition of the user's voice, on the basis of the user's position
estimated by the user position estimating unit 16 or the positions of the microphones
10 registered to the microphone position information DB 15. Then, the microphone
array processing unit 13 1 performs directivity control on the audio signals acquired
20 by the selected microphone group. Further, the microphone array processing unit
131 may form super directivity of the array microphone through a delay-and-sum
array process and a null generation process.
[0043]
-High S/N processing unit
25 The high S/N processing unit 133 has a function of processing a plurality of
audio signals output from the amplifying/ADC unit 11 to form a monaural signal
having high articulation and a high S/N ratio. Specifically, the high S/N processing
unit 133 perfor~nss ound source separation, and performs dereverberation and noise
reduction.
30 [0044]
Further, the high S/N processing unit 133 may be disposed at a stage
subsequent to the microphon~a rray processing unit 131 . Further, the audio signals
(stream data) processed by the high S/N processing unit 133 are used for voice
recognition performed by the recognizing unit 17 and are transmitted to an outside
though a communication I/F 19.
5 [0045]
-Sound field reproduction signal processing unit
The sound field reproduction signal processing unit 135 performs signal
processing on the audio signals to be reproduced through the plurality of speakers 20,
and performs control such that a sound field is localized around the user's position.
10 Specifically, for example, the sound field reproduction signal processing unit 135
selects an optimal speaker group for forming the acoustically closed surface
sussounding the user on the basis of the user's position estimated by the user position
estimating unit 16 or the positions of the speakers 20 registered to the speaker
position information DB 21. Then, the sound field reproduction signal processing
15 unit 135 writes the audio signals which have been subjected to signal processing in
output buffers of a plurality of cliannels corresponding to the selected speaker group.
[0046]
Further, the sound field reproduction signal processing unit 135 controls an
area inside the acoustically closed surface as an appropriate sound field. As a
20 method of controlling the sound field, for example, the Helmholtz-Kirchhoff intepal
theorem and the Rayleigh integral theorem are lu~own, and wave field synthesis
(WFS) based on the theorems is generally known. Fuither, the sound field
reproduction signal processing unit 135 may apply signal processing techniques
disclosed in JP 46745053 and JP 47351083.
25 [0047]
Note that the shape of the acoustically closed surface formed by the
microphones or the speakers is not particularly limited as long as it is a theedimensional
shape surrounding the user, and, as shown in FIG. 4, examples of the
shape may include an acoustically closed surface 40-1 having an oval shape, an
30 acoustically closed surface 40-2 having a columnar shape, and an acoustically closed
surface 40-3 having a polygonal shape. The examples illustrated in FIG. 4 show as
examples the shapes of the acoustically closed surfaces formed by a plurality of
speakers 20B-1 to 208-12 arranged around the user B in the site B. The examples
also apply to the shapes of the acoustically closed surfaces formed by the plurality of
microphones 10.
5 [0048]
(Microphone position information DB)
The microphone position information DB 15 is a storage unit that stores
position information of the plurality of lnicrophones 10 arranged at the site. The
position information of the plurality of microphones 10 may be registered in advance.
10 [0049]
(User position estimating unit)
The user position estimating unit 16 has a function of estimating the user's
position. Specifically, the user position estimating unit 16 estimates the user's
relative position to the plurality of microphones 10 or the plurality of speakers 20 on
15 the basis of the analysis result of the sounds acquired by the plurality of microphones
10, the analysis result of the captured images obtained by the image sensors, or the
detection result obtained by the human sensors. The user position estimating unit
16 may acquire Global Positioning System (GPS) information and may estimate the
user's absolute position (current position information).
20 [0050]
(Recoguizing unit)
The recognizing unit 17 analyzes the user's voice on the basis of the audio
signals which are acquired by the plurality of microphones 10 and then processed by
the signal processing unit 13, and recognizes a command. For example, the
25 recognizing unit 17 performs morpliological analysis on the voice of the user "I'd
like to speak with B," and recognizes a call origination request command on the basis
of the given target "B" that is designated by the user and the request "I'd like to
speak with."
[005 11
30 (Identifying unit)
The identifying unit 18 has a function of identifying the given target
recognized by the recognizing unit 17. Specifically, for examp&, the identifying
unit 18 may decide the access destination information for acquiring a voice and an
image corresponding to the given target. For example, the identifying unit 18 may
transmit information representing the given target to the management server 3
5 through the communication L/F 19, and acquire the access destination information
(for example, IP address) corresponding to the given target from the management
server 3.
[0052]
(Communication I@)
10 The communication UF 19 is a communication module for transmitting or
receiving data to or from another signal processing apparatus or the management
sewer 3 via the network 5. For example, the communication I/F 19 according to the
present embodiment sends a queiy for access destination information corresponding
to the given target to the management server 3, and transmits the audio signal which
15 is acquired by the microphone 10 and then processed by the signal processing unit 13
to another signal processing apparatus which is an access destination.
[0053]
(Speaker position information DB)
The speaker position information DB 21 is a storage unit that stores position
20 information of the plurality of speakers 20 arranged at the site. The position
information of the plurality of speakers 20 may be registered in advance.
[0054]
(DACIamplifying unit)
The DACIamplifying unit 23 has a function (DAC) of converting the audio
25 signals (digital data), which are witten in the output buffers of the channels, to be
respectively reproduced through the plurality of speakers 20 into acoustic waves
(analog data).
[0055]
In addition, the DACJamplifying unit 23 amplifies the converted acoustic
30 waves, and reproduces (outputs) the acoustic waves through the plurality of speakers
20.
[0056]
(Array speaker)
The plurality of speakers 20 are arranged throughout a cestain area (site) as
described above. For example, the plurality of speakers 20 are arranged at outdoor
5 sites such as roads, electric poles, street lamps, houses, and outer walls of buildings
and indoor sites such as floors, walls, and ceilings. Further, the plurality of
speakers 20 reproduce the acoustic waves (voices) output &om the DACIamplifying
unit 23.
[0057]
10 Heretofore, the configuration of the signal processing apparatus 1 according
to the present embodiment has been described in detail. Next, with reference to FIG.
5, the configuration of tlie management server 3 according to the present
embodiment will be described.
[OOSS]
15 [2-3. Management server]
FIG. 5 is a block diagram showing a configuration of the management server
3 according to the present embodiment. As shown in FIG. 5, the management
server 3 includes a managing unit 32, a searching unit 33, a user position inforniation
DB 35, and a communication I/F 39. The above-mentioned components will be
20 described below.
[0059]
(Managing unit)
The managing unit 32 manages information associated with a place (site) at
which the user is currently located on the basis of a user identification (ID)
25 transmitted fsom the signal processing apparatus 1. For example, the managing unit
32 identifies the user on the basis of the user ID, and stores an IP address of the
signal processing apparatus 1 of a transmission source in the user position
information DB 35 in association with a name of the identified user or the like as the
access destination information. The user ID may include a name, a personal
30 identification number, or biological information. Further, the managing unit 32
may perfornl the user authentication process on the basis of the transmitted user ID.
[OOGO]
(User position information DB)
The user position information DB 35 is a storage unit that stores information
associated with a place at which the user is currently located according to
5 management by the managing unit 32. Specifically, the user position information
DB 35 stores the user ID and the access destination information (for example, an IP
address of a signal processing apparatus corresponding to a site at which the user is
located) in association with each other. Further, current position information of
each user may be constantly updated.
10 [0061]
(Searching unit)
The searching unit 33 searches for the access destination information with
reference to the user position information DB 35 according to the access destination
(call origination destination) query from the signal processing apparatus 1.
15 Specifically, the searching unit 33 searches for the associated access destination
information and extracts the access destination information from the user position
information DB 35 on the basis of, for example, a name of a target user included in
the access destination quely.
[0062]
20 (Communication VF)
The communication VF 39 is a communication module that transmits or
receives data to or from the signal processing apparatus 1 via the network 5. For
example, the communication VF 39 according to the present embodiment receives
the user ID and the access destination quely from the signal processing apparatus 1.
25 Further, the communication VF 39 transmits the access destination information of the
target user in response to the access destination query.
[0063]
Heretofore, the conlponents of the acoustic system according to an
embodiment of the present disclosure have been described in detail. Next, with
30 reference to FIGS. 6 to 9, an operation process of the acoustic system according to
the present embodiment will be described in detail.
[0064]
<3. Operation process>
[3-1. Basic process]
FIG. 6 is a flowchart showing a basic process of the acoustic system
5 according to the present embodiment. As shown in FIG. 6, first of all, in step S103,
the signal processing apparatus 1A transmits an ID of the user A located at the site A
to the management server 3. The signal processing apparatus 1A may acquire an ID
of the user A from a tag such as a radio frequency identification (RFID) tag
possessed by the user A or from the user A's voice. Further, the signal processing
10 apparatus 1A may read biological information from the user A (a face, an eye, a hand,
or the like), and acquire the biological information as an ID.
[0065]
Meanwhile, in step S106, the signal processing apparatus 1B similarly
transmits an ID of the user B located at the site B to the management server 3.
15 100661
Next, in step S109, the management sewer 3 identifies the user on the basis
of the user ID transmitted from each signal processing apparatus 1, and registers, for
example, an IP address of the signal processing apparatus 1 of the transmission
source as the access destination information in association with, for example, the
20 identified user's name.
[0067]
Next, in step S112, the signal processing apparatus 1B estimates the position
of the user B located at the site B. Specifically, the signal processing apparatus 1B
estimates the user B's relative position to the plurality of microphones arranged at the
25 site B.
[0068]
Next, in step S115, the signal processing apparatus 1B performs the
microphone array process on the audio signals acquired by the plurality of
microphones arranged at the site B on the basis of the user B's estimated relative
30 position so that the sound acquisition position is focused on the user B's mouth. As
described above, the signal processing apparatus 1B prepares for the user B to utter
something.
[0069]
On the other hand, in step S118, the signal processing apparatus IA
similarly performs the microphone array process on the audio signals acquired by the
5 plurality of micropliones arranged at the site A so that the sound acquisition position
is focused on the nser A's mouth, and prepares for the user A to utter something.
Then, the signal processing apparatus 1A recognizes a command on the basis of the
nser A's voice (utterance). Here, the description will continue with an example in
which the nser A utters "I'd like to speak with B," and the signal processing
10 apparatus 1A recognizes the utterance as a command of the "call origination request
to the user B." A command recognition process according to the present
embodiment will be described in detail in [3-2. Command recognition process]
which will be described later,
[0070]
15 Next, in step S121, the signal processing apparatus 1A sends the access
destination query to the management server 3. When the command is the "call
origination request to the user B" as described above, the signal processing apparatus
1A queries the access destination information of the user B.
[0071]
20 Next, in step S125, the management server 3 searches for the access
destination information of the user B in response to the access destination query from
the signal processing apparatus lA, and then, in step S126 that follows, transmits the
search result to the signal processing apparatus 1A.
[0072]
25 Next, in step S127, the signal processing apparatus 1A identifies
(determines) an access destination on the basis of the access destination infor~nation
of the user B received from the management server 3.
[0073]
Next, in step S128, the signal processing apparatus 1A performs the process
30 of originating a call to the signal processing apparatus 1B on the basis of the access
destination information of the identified user B, for example, an IP address of the
signal processing apparatus 1B corresponding to the site B at which the user B is
currently located.
[0074]
Next, in step S131, the signal processing apparatus 1B outputs a message
5 asking the user B whether to answer a call from the user A or not (call notification).
Specifically, for example, the signal processing apparatus 1B may reproduce a
corresponding message through the speakers arranged around the user B. Furthel;
the signal processing apparatus 1B recognizes the user B's response to the call
notification on the basis of the user B's voice acquired through the plurality of
10 microphones arranged around the user B.
[0075]
Next, in step S134, the signal processing apparatus 1B transmits the
response of the user B to the signal processing apparatus IA. Here, the user B gives
an OK response, and thus, two-way communication starts between the user A (signal
15 processing apparatus 1A side) and the user B (signal processing apparatus 1B side).
[0076]
Specifically, in step S137, in order to start communication with the signal
processing apparatus IB, tlie signal processing apparatus 1A performs a sound
acquisition process of acquiring the user A's voice at the site A and transmitting an
20 audio stream (audio signals) to the site B (signal processing apparatus 1B side).
The sound acquisition process according to the presetit embodiment will be described
in detail in [3-3. Sound acquisitio~p~ro cess] which will be described later.
[0077]
Then, in step S140, the signal processing apparatus 1B forms the
25 acoustically closed surface surrounding the user B through the plurality of speakers
arranged around the user B, and performs a sound field reproduction process on the
basis of the audio stream transmitted from the signal processing apparatus IA.
Note that the sound field reproduction process according to the present embodiment
can further construct a sound field of a third space (site C), and can provide tlie user
30 speaking on the telephone with another user located at another space with a feeling
of being absorbed in the third space. The sound field reproduction process will be
described in detail in "4. Construction of sound field of third space" which will be
described later.
[0078]
In steps S137 to S140 described above, one-way communication has been
5 described as an example, but in the present embodiment, two-way communication
can be performed. Accordingly, unlike steps S137 to S140 described above, the
signal processing apparatus 1B may perform the sound acquisition process, and the
signal processing apparatus 1A ]nay perform the sound field reproduction process.
[0079]
10 Heretofore, the basic process of the acoustic system according to the present
embodiment has been described. Through the above-described process, the user A
can speak on the telephone with the user B located at a different place by uttering
"I'd like to speak with B" without catlying a mobile phone terminal, a smart phone,
or the like, by using the plurality of microphones and the plurality of speakers
15 arranged around the user A. Next, the command recognition process performed in
step S118 will be described in detail with reference to FIG. 7.
[OOSO]
[3-2. Command recognition process]
FIG. 7 is a flowchart showing the command recognition process according
20 to the present embodiment. As shown in FIG 7, first of all, in step S203, the user
position estimating unit 16 of the signal processing apparatus 1 estimates the user's
position. For example, the user position estimating unit 16 may estimate the
relative position and direction of the user to each microphone, and the position of the
user's mouth on the basis of sounds acquired through the plurality of microphones 10,
25 captured images obtained by the image sensors, an arrangement of the microphones
stored in the microphone position infornlation DB 15, or the like.
[OOSl]
Next, in step S206, the signal processing unit 13 selects the microphone
group forming the acoustically closed surface surrounding the user according to the
30 user's relative position and direction, and the position of the user's n~outhth at have
been estimated.
[0082] --
Next, in step S209, the microphone array processing unit 131 of the signal
processing unit 13 performs the microphone array process on the audio signals
acquired through the selected microphone group, and controls directivity of the
5 lnicrophones to be focused on the user's mouth. Through this process, the signal
processing apparatus 1 can prepare for the user to utter something.
[0083]
Next, in step S212, the high S/N processing unit 133 performs a process
such as dereverberation or noise reduction on the audio signal processed by the
10 microphone array processing unit 131 to improve the S/N ratio.
[0084]
Next, in step S215, the recognizing unit 17 performs voice recognition
(voice analysis) on the basis of the audio signal output fiom the high S/N processing
unit 133.
15 [0085]
Then, in step S218, the recognizing unit 17 performs the command
recognition process on the basis of the recognized voice (audio signal). There is no
particular restriction to concrete content of the command recognition process, but for
example, the recognizing unit 17 may recognize a command by comparing a
20 previously registered (learned) request pattern with the recognized voice.
[0086]
When a command is not recognized in step S218 (No in S218), the signal
processing apparatus 1 repeatedly perforins the process perfonned in steps S203 to
S215. At this time, since steps S203 and S206 are also repeated, the signal
25 processing unit 13 can update the microphone group forming the acoustically closed
surface surrounding the user according to the user's movement.
[0087]
[3-3. Sound acquisition process]
Next, the sound acquisition process performed in step S137 of FIG. 6 will be
30 described in detail with reference to FIG. 8. FIG, 8 is a flowchart showing the
sound acquisition process according to the present embodiment. As shown in FIG.
8, first of all, in step S308, the microphone array processing unit 131 of the signal
processing unit 13 perfornls the microphone array process on the audio signals
acquired through the selectedlupdated microphones, and controls directivity of the
microphones to be focused on the user's mouth.
5 [0088]
Next, in step S312, the high S/N processing unit 133 performs the process
such as dereverberation or noise reduction on the audio signal processed by the
microphone array processing unit 131 to improve the S/N ratio.
[0089]
10 Then, in step S315, the communication I/F 19 transmits the audio signal
output from the high S/N processing unit 133 to the access destination (for example,
signal processing apparatus IB) represented by the access destination information of
the target user identified in step S126 (see FIG. 6). Through this process, a voice
uttered by the user A at the site A is acquired by the plurality of microphones
15 arranged around the user A and then transmitted to the site B.
[0090]
Heretofore, the command recognition process and the sound acquisition
process according to the present embodiment have been described. Next, the sound
field reproduction process according to the present embodiment will be described in
20 detail.
[0091]
<4. Construction of sound field of third space>
As described above, the sound field reproduction process (step S140 in FIG.
6) according to the present embodiment can constmct a sound field of a third space
25 (site C), and can provide the user speaking on the telephone with another user located
at another space with a feeling of being absorbed in the third space. Hereinafter,
with reference to FIG. 9, an outline of sound field construction for providing the user
with the feeling of being absorbed in the third space will be described.
[0092]
30 FIG. 9 is a diagram illustrating construction of the sound field of the third
space according to the present etnbodirnent. As shown in FIG. 9, in the case where
the user A located at the site A speaks on the telephone with the user B located at the
site B, the acoustic system according to the present embodiment constructs a sound
field 42 of the site C which is the third space in each of the sites A and B. Here, as
an example, let us assume that the site A, the site B, and the site C are located away
5 from one another (remote places). In this case, for example, the user B who is in
Tokyo (site B) can speak on the telephone with the user A who is in the USA (site A)
while being provided with the feeling of being absorbed in the space of Italy (site C)
which is a place where the user B is to travel with the user A.
[0093]
10 Specifically, the acoustic system according to the present embodiment may
construct the sound field 42 of the site C using an acoustic information parameter
(characteristics of a parameter such as an impulse response) measured in the site C in
advance or acoustic content (environmental sound) acquired in the site C. Note that
the acoustic information parameter and the acoustic content of the third space may be
15 acquired in the third space in advance and may be accumulated in a management
server.
LO0941
(Technique of constructing sound field of site C)
Here, with reference to FIG. 10, there will be described techniques of
20 const~uctingt he sound field of the site C in the respective sites A and B in the case
where the user A located at the site A speaks on the telephone with the user B located
at the site B. FIG. 10 is a diagram illustrating techniques of constructing the sound
field of the site C. In the example shown in FIG. 10, there will be described the
case where, in the site B at which the user B who is speaking on the telephone with
25 the user A is located, the sound field of the site C is constructed (the feeling of being
absorbed in the site C is provided).
[0095]
As shown in FIG. 10, a technique 1 is as follows: an audio image is
localized so that the voice of the user A is present outside an acoustically closed
30 surface 40B formed by the plurality of speakers 20B surrounding the user B, and is
further modified using an acoustic infolmation parameter in a manner that the voice
of the user A can be heard as if the ~voi-ce were echoed in the site C.
[0096]
Here, as shown in FIG. 10, in the case where the audio image of the user A
is localized outside the acoustically closed surface 40, a wave front at the time at
5 which the voice emitted by the user A located outside the acoustically closed surface
40B crosses the acoustically closed surface 40B is estimated. Then, the
reproduction from the plurality of speakers 20 is performed in a manner that the
estimated wave front is created inside the acoustically closed surface 40B, to thereby
localize the audio image.
10 [0097]
Further, in the case of estimating that the user A emits a voice in the site C,
there may be a case where the voice of the user A may reach the acoustically closed
surface 40B with the inclusion of a reflected sound (reflected sound differing
depending on materials and stmctures) owing to a consttuction or an obstacle of the
15 site C. Accordingly, the acoustic system according to the present embodiment
constructs, in the site B, the sound field 42 in which the voice of the user A can be
heard as if the voice were echoed in the site C, by modifying the voice of the user A
using the acoustic information paranleter (impulse response) that has been measured
in the site C in advance. In this way, the user B can be provided with the feeling of
20 being absorbed in the site C further affluently.
[0098]
A technique 2 is as follows: the voice of the user B located inside the
acoustically closed surface 40 is acquired, the voice is modified using the acoustic
information paranleter of the site C, and the processed voice is reproduced from the
25 plurality of speakers 20B forming the acoustically closed surface 40. That is, in
order that the user B located inside the acoustically closed surface 40 feels the sound
field of the site C as if the user B were actually feeling the sound field, is provided
with the feeling of being absorbed in the site C further affluently, and feels a large
size of the space of the site C, it is also important to feel the change in the voice that
30 the user B himselVherself has emitted (echolocation) in addition to the modification
(technique 1) of the voice of the partner with whom the user B speaks on the
telephone. Accordingly,&the technique 2, tlie sound field 42 is const~ucted in the
site B in a manner that the voice that the user B has emitted can be heard as if the
voice were echoed in the site C. In this way, the user B can further affluently
experience the ambience of the site C and the feeling of being absorbed in the site C.
5 Note that a specific method for achieving the technique 2 will be described later with
reference to FIG. 16A and FIG. 16B.
[0099]
A technique 3 is as follows: by reproducing acoustic content such as
murmuring voices in the site C and an environmental sound from the plurality of
10 speakers 20B forming the acoustically closed surface 40 surrounding the user B, the
ambience of the site C and the feeling of being absorbed in the site C are increased.
The acoustic content of the site C may be recorded in advance or may be acquired in
real time.
[Ol 001
15 Heretofore, there have been described with reference to FIG 10 three
techniques for constructing the sound field for providing the feeling of being
absorbed in the site C. In the acoustic system according to the present embodiment,
the sound field may be constiucted using one technique out of the three techniques,
or the sound field may be const~uctedu sing two or more techniques in combination.
20 [OIOl]
(Designation of site C)
Furthel; in the present embodiment, the third space (site C) may be
appropriately designated by the user or may be the place which is set in advance.
For example, when the user A located at the site A utters "I'd like to speak with tlie
25 user B (first target) at the site C (second target)," the plurality of niicrophoiies 10A
(see FIG. 1) arranged around the user A acquires the utterance, and the utterance is
recognized as a command by the signal processing apparatus 1A.
[O 1021
Next, the signal processing apparatus 1A sends a request for access
30 destination information for speaking with the "user B" on the telephone and data for
constructing the sound field of the designated place to the management seilrer.
Then, the management sellrer transmits the access destination information (here, an
IP address or the like of the signal processing apparatus 1B of the site B at which the
user B is located) and the data for constiucting the sound field (here, the acoustic
information parameter and the acoustic content of the site C) to the signal processing
5 apparatus IA.
[0103]
Further, in the case where communication between the signal processing
apparatus 1A and the signal processing apparatus 1B is started (in the case where the
user B gives an OK response to the call origination from the user A), the data for
10 constructing the sound field is also transmitted to the signal processing apparatus 1B.
In this way, the sound field of the site C is constructed in the site A and the site B,
and the user A and the user B located at the different sites can share the feeling of
being absorbed in the same place.
[0104]
15 Heretofore, the outline of sound field construction for providing the user
with the feeling of being absorbed in the third space has been described.
Subsequently, with reference to FIG 11, a configuration of the management server
which accumulates the acoustic information parameter and the acoustic content of
the third space will be described.
20 [0105]
[4-1. Configuration of management server]
FIG. 11 is a block diagram showing another configuration of the
management server according to the present embodiment. As shown in FIG. 11, a
management server 3' includes a managing unit 32, a searching unit 34, a user
25 position information DB 35, a conununication I/F 39, an acoustic information
parameter DB 36, and an acoustic content DB 37. Since the managing unit 32, the
user position information DB 35, and the conlmunication I/F 39 are the same as those
described with reference to FIG. 5, the description thereof will be omitted here.
[0 1061
30 (Searching unit)
First, the searching unit 34 searches for, in the same manner as the searching
unit 33 described above, the access destination information withreference to the user
position informatioi~ DB 35 according to the access destination (call origination
destination) query from the signal processing apparatus 1. Specifically, the
searching unit 33 searches for the associated access destination it~formation and
5 extracts the access destination information from the user position information DB 35
on the basis of, for example, a name of a target user included in the access
destination query.
[0107]
Further, in response to a request for the data for constructing the sound field
10 from the signal processing apparatus 1, the searching unit 34 searches for the
acoustic information parameter of the designated site and extracts the acoustic
information parameter from the acoustic information parameter DB 36. In addition,
in response to the request for the data for constructing the sound field from the signal
processing apparatus 1, the searching unit 34 searches for the acoustic content of the
15 designated site and extracts the acoustic content from the acoustic content DB 37.
[0108]
(Acoustic information parameter)
The acoustic information parameter DB 36 is a storage unit which stores
acoustic information parameters that have been measured in the respective sites in
20 advance. The acoustic parameter may be obtained by measuring an impulse
response fiom any one point or a plurality of points (position(s) at which audio
image(s) is(/are) attempted to be localized) in each site. Further, the SIN ratio
itnproves by using the time streched pulse (TSP) response, the swept-sine method,
the M-sequence response, and the like in measuring the impulse response.
25 [0109]
Here, with reference to FIG. 12, the measurement of the acoustic
i~lforination parameter will be described. A measurement 1 shown in FIG. 12
describes the measurement of the acoustic information parameter (first acoustic
information parameter) which is used when modifying the voice of the speaking
30 partner localized at any position outside the acoustically closed surface 40 in the
technique 1 described with reference to FIG. 10. As shown in FIG. 12, a plurality of
mjcrophones 10C each having directivity turning outward arranged at the site C
measure how the transmission is performed (impulse response) fiom a sound source
(speaker 20C) to each microphones IOC, the speaker 20C being installed at any
position outside the closed surface 43 formed by the plurality of microphones 10C.
5 [OllO]
In the example shown in FIG. 12, although one speaker 20C is disposed in
the measurement 1, the example is not limited thereto, and a plurality of speakers
20C may be arranged outside the closed surface 43 and the transmission fiom each of
the speakers 20 to each of the microphones 10C may be measured. In this way, the
10 number of locations at which the sound sources of the user A can be localized can be
increased in the technique 1.
[Olll]
Further, a measurement 2 shown in FIG. 12 describes the measurement of
the acoustic information parameter (second acoustic information parameter) which is
15 used when modifying the voice of the user himselfierself located inside the
acoustically closed surface 40 in the technique 2 described with reference to FIG. 10.
As shown in FIG. 12, a plurality of microphones 10C each having directivity turning
outward arranged at the site C measure how the sound (signal for measurement)
output from the sound source (speaker 20C) is transmitted to each microphones 10C
20 with the influence of reflectiodecho in the site C, the speaker 20C being installed
inside the closed surface 43 formed by the plurality of microphones 10C. In the
example shown in FIG. 12, although one speaker 20C is disposed in the ~neasurement
2 as an example, the present embodiment is not limited thereto, and a plurality of
speakers 20C may be arranged inside the closed surface 43 and the transmission kom
25 each of the speakers 20 to each of the n~icrophones 10C may be measured.
[0112]
(Acoustic content)
The acoustic content DB 37 is a storage unit which stores acoustic content
acquired fsom each site. Examples of the acoustic content include ambient sounds
30 (environmental sounds, rnurmnu~ring voices, and the like) that have been recorded
(nieasured) in each site.
[0113] -~
The measurement of the acoustic content is performed as shown in a
measurement 3 of FIG. 12, for example, a plurality of microphones 10C each having
directivity turning outward asranged at the site C measure (record) ambient sounds.
5 The measurement of the ambient sounds may be performed for each hour, weekday,
and holiday. In this way, the acoustic system according to the present embodiment
can construct a sound field of the site C for each hour, weekday, and holiday. For
example, the acoustic content close to the current hour can be reproduced at the site
B which is the reproduction environment.
10 [0114]
Note that the size of the closed surface 43 formed by the plurality of
microphones 10C shown in FIG. 12 may be larger than the acoustically closed
surface of the listening environment (reproduction environment). Hereinafter,
description will be given with reference to FIG. 13. FIG. 13 is a diagram showing
15 an arrangement of the plurality of microphones 10C in a measuring environment (site
C here) and an anange~nent of the plurality of speakers 20B in a listening
environment (site B here) in comparison with each other.
[0115]
As shown in FIG. 13, with respect to the acoustically closed surface 40
20 formed by the plurality of speakers 20B provided so as to surround the user B, the
plurality of microphones 10C used for the measurement at the site C are arranged so
as to form the closed surface 43, the size of which is larger than the size of the
acoustically closed surface 40.
[0116]
25 Fulthel; as described above with reference to FIG. 4, at the site B of the
listening environment (reproduction environment), the three-dimensional
acoustically closed surfaces 40-1, 40-2, and 40-3 are each formed with the plurality
of speakers 20B-1 to 20B-12. Accordingly, as shown in FIG. 14, also at the site C
which is the measuring environment, a plurality of micropliones 10C-1 to 10C-12
30 each having directivity turning outward may form three dimensional closed surfaces
43-1,43-2, and 40-3.
[0117]
Heretofore, each conlponent of the configuration of the management server
3' according to the present embodiment has been described in detail. Subsequently,
there will be described control performed on the site B side of the listening
5 environment (reproduction environment) constructing the sound field of the site C
using the techniques 1 to 3 (see FIG. 12). At the site B side, an optimal sound field
is fonned by the sound field reproduction signal processing unit 135 (see FIG. 3) of
the signal processing apparatus 1B. Hereinafter, with reference to FIG. 15, there
will be described specifically the configuration of the sound field reproduction signal
10 processing unit 135 that consttucts the sound field by achieving the techniques 1 to 3.
[0118]
[4-2. Configuration of sound field reproduction signal processing unit]
FIG. 15 is a block diagram illustrating a configuration of the sound field
reproduction signal processing unit 135 perfornling sound field const~uction for
15 providing a feeling of being absorbed in the site C. Further, FIG. 15 shows main
components of the signal processing apparatus 1B that are related to the description
here, and the other components are omitted.
[0119]
As shown in FIG. 15, the sound field reproduction signal processing unit
20 135 functions as a convolution unit 136, howling reduction units 137 and 139, and a
matrix convolution unit 138.
[O 1201
(Convolution unit)
The convolution unit 136 has a function of achieving the above-mentioned
25 technique 1 (localization of the audio image of the user A, ~nodificationo f the echoed
voice of the user A). Specifically, through the communication I/F 19, the
convolution unit 136 perfotms rendering on an audio signal b (voice of the user A)
acquired (received) frotn the signal processing apparatus 1A of the site A for each
output speaker, using an acoustic information parameter c (first acoustic information
30 parameter) of the site C. Further, in this case, taking into consideration a parameter
a of the position at which the audio itnage of the user A is localized, the convolution
unit 136 may use an acoustic&fosmation parameter c (impulse response) of the site
C corresponding to the localization position. Note that the parameter a of the
position at which the audio image of the user A is localized may be transmitted fiom
the signal processing apparatus 1A or the management server 3' through the
5 communication I/F 19, or may be calculated by the signal processing apparatus 1B
on the basis of the instruction given by the user B. Further, the convolution unit
136 may acquire the acoustic information parameter c (impulse response) of the site
C fiom the management server 3' through the communication I/F 19.
[0121]
10 Then, as shown in FIG. 15, the convolution unit 136 writes the audio signal
subjected to the signal process to the output buffer of each of the output speakers
(plurality of speakers 20B forming the acoustically closed surface 40B surrounding
the user B).
[O 1221
15 (Howling reduction units)
The howling reduction units 137 and 139 are provided as necessary in order
to avoid howling caused by feedback, at the latter part of the amplifyingJADC unit
11 of the microphones and the former part of the DAC/amplifying unit 23 of the
speakers as shown in FIG. 15, respectively, and are capable of operating in
20 cooperation. As described above, in the technique 2, rendering is performed on the
sound acquired from the plurality of microphones 10B arranged around the user B
using the acoustic information parameter (impulse response), and the sound is
reproduced from the plurality of speakers 20B arranged around the user B. In this
case, since the positions of the microphones are close to the positions of the speakers,
25 there is a risk that excessive oscillation may occur due to operations of the
microphones and the speakers. Accordingly, in the example shown in FIG. 15, the
howling reduction units 137 and 139 are provided and the howling reduction process
is executed. Note that, in order to prevent the excessive oscillation, the sound field
reproduction signal processing unit 135 may also include an echo canceler in
30 addition to the howling reduction units 137 and 139.
[0123]
(Matrix convolution -- unit)
The matrix convolution unit 138 has a function of achieving the abovementioned
technique 2 (modification of echoed voice of the user B). Specifically,
the matrix convolution unit 138 performs rendering on an audio signal (sound
5 generated in the acoustically closed surface 40B) acquired by the plurality of
microphones 10B alranged at the site B for each output speaker, using an acoustic
information parameter c (second acoustic information parameter; impulse response
group) of the site C. In this way, the audio signal for constructing, at the site B, the
sound field in which the sound generated in the acoustically closed surface 40B at
10 the site B such as the voice of the user B himselfierself is felt to be echoed at the
site C.
[0 1241
Here, with reference to FIG. 16A and FIG. 16B, the way of achieving the
technique 2 according to the present embodiment will be described specifically.
15 FIG. 16A is a diagram illustrating measurement of an impulse response in the site C.
As shown in FIG. 16A, first, an impulse response from each of the speakers 20C to
each of the microphones 10C is measured, the speakers 20C being arranged at the
site C and facing outside of the closed surface 43, the n~icrophones 10C being also
arranged at the site C and facing outside of the closed surface 43.
20 [0125]
Specifically, an impulse response from a single speaker on the closed
surface 43 to a group including a plurality of microphones on the same closed
surface 43 is measured. From the viewpoint of the frequency axis, the impulse
response can be also considered as a transfer function which is influenced by space
25 acoustics of the constn~ction/obstacleo f the site C.
[0126]
Here, in the example shown in FIG. 16A, the positions of the respective
microphones/speakers on the closed surface 43 are represented by R1, R2, ..., RN.
Then, as shown in FIG. 16A, the respective transfer functions fiom the speaker (SP)
30 disposed at R1 to the microphone disposed at R1, to the microphone disposed at
R2, ..., and to the nlicrophone disposed at RN are measured. Subsequently, the
respective transfer functions from the speaker disposed at R2 to the microphone
disposed at R1, to the microphone disposed at R2, ..., and to the microphone
disposed at RN are measured.
[0127]
5 Next, when the transfer function from the speaker positioned at R1 to the
microphone positioned at R1 is represented by R11, and the transfer function from
the speaker positioned at R1 to the microphone positioned at R2 is represented by
R12, a matrix using the tsansfer functions R can be expressed as Expression (1)
shown in FIG. 16A.
10 [0128]
The matrix data is accumulated in the management server 3' and the like as
an acoustic information parametel; and is used for const~uctingth e sound field of the
site C at the site B. Subsequently, with reference to FIG. 16B, the case where the
sound field of the site C is constructed using the matrix data at the site B will be
15 described.
[0129]
FIG. 16B is a diagram illustrating an arithmetic operation using an impulse
response group performed by the matrix convolution unit 138. In the example
shown in FIG. 16B, let us assume a closed surface in the site B (reproduction
20 environment) side, the closed surface having substantially the same size and the
shape at the time of the measurement at the site C. Further, let us assume the case
where the number of the plurality of microphones 10B and the number of the
plurality of speakers 20B arranged at the site B are also the same as at the time of the
measurement at the site C, and the arrangement positions thereof are the same RI,
25 R2, ..., RN as at the time of the measurement at the site C. However, as shown in
FIG. 16B, the plurality of microphones 10B and the plurality of speakers 20B are
facing inside of the acoustically closed surface 40B.
[0130]
Furthel; as shown in FIG. 16B, frequency axis-based expressions acquired
30 by the respective microphones at the positions of RI, R2, ..., RN at the site B are
represented by V1, V2, ..., VN. Further, output signals (audio signals) output
reproduced) from the respective speakers at the positions-of R1, R2, ..., RN ( - at the
site B are represented by W1, W2, ..., WN.
[0131]
In this case, a wave front of sound (voice of the user A or noise) generated
5 inside the acoustically closed surface 40B of the site B reaches the acoustically
closed surface 40B and is acquired by the microphones 10B facing tlie inside
positioned at R1, R2, ..., RN, and the respective micropliones 10B acquires sound
acquisition signals of V1, V2, ..., VN.
[0132]
10 Then, the matrix convolution unit 138 executes Expression (2) shown in FIG
16B using the signal group (microphone inputs) of V1, V2, ..., VN and the matrix
(Expression (1)) of the transfer function group described with reference to FIG. 16A,
and calculates W1, W2, ..., WN output from the respective speakers 20B.
[0133]
15 As described above, the matrix convolution unit 138 performs the signal
process on the audio signals (Vl, V2, ..., VN) acquired by the plurality of
microphones IOB, using the acoustic information parameters (transfer function
group) of the site C. Further, as shown in FIG. 15, the matrix convolution unit 138
adds the audio signals (Wl, W2, ..., WN) subjected to the signal process to the output
20 buffers of the respective output speakers.
[0134]
(Addition of acoustic content)
As shown in FIG. 15, the sound field reproduction signal processing unit
135 adds acoustic content d of the site C received fiom the management server 3'
25 through the communication I& 19 to each of the output buffers of the respective
output speakers, and thus, the technique 3 described above is achieved.
[0135]
Heretofore, the configuration of the sound field reproduction signal
processing unit 135 of the signal processing apparatus 1B according to the present
30 embodinlent has been described in detail. Next, with reference to FIG. 17, there
will be described specifically a sound field reproduction process at the time of
-- constructing the sound field of the site C at the site B.
[0136]
[4-3. Sound field reproduction process]
FIG. 17 is a flowchart showing a sound field reproduction process according
5 to the present embodiment. As shown in FIG. 17, first, in step S403, tlie user
position estimating unit 16 (see FIG. 3) of the signal processing apparatus 1B
estimates the position of the user B. For example, the user position estimating unit
16 may estimate the relative position, direction, position of the tnouth, and position
of the ear of the user B with respect to each speaker 20B on the basis of sound
10 acquired from the plurality of microphones 10B, captured images obtained by the
image sensors, and arrangement of the speakers stored in the speaker position
information DB 2 1.
[0137]
Next, in step S406, the signal processing unit 13 selects a
15 microphonelspeaker group forming the acoustically closed surface surrounding the
user on the basis of the estimated relative position, direction, position of the tnouth,
and position of tlie ear of the user B.
[0138]
Next, in step S407, tlie sound field reproduction signal processing unit 135
20 of the signal processing unit 13 performs, by the convolution unit 136, the process of
the technique 1 shown in FIG. 10 on the received audio signal b (voice of the user A
acquired at the site A). Specifically, as shown in FIG. 15, the convolution unit 136
perfornis rendering on the audio signal b received from the signal processing
apparatus 1A of the site A for each selected output speaker, using the acoustic
25 information parameter c (first acoustic information parameter) of the site C. Then,
the convolution unit 136 writes the audio signal subjected to the process of the
technique 1 to the output buffer of each selected output speaker.
[0139]
Next, in step S409, the sound field reproduction signal processing unit 135
30 performs, by the matrix convolution unit 138, the process of the technique 2 shown
in FIG. 10 on the voice of tlie user B acquired at the site B using the selected
microphone group. Specifically, the matrix convolution unit 138 performs
rendering on the audio signal acquired by the microphone group (plurality of
micropholles 10B) forming the acoustically closed surface surrounding the user B for
each output speaker, using the acoustic information parameter c (second acoustic
5 information parameter) of the site C. Then, the matrix convolution unit 138 adds
the audio signal subjected to the process of the technique 2 to the output buffer of
each selected output speaker.
[0 1401
Next, in step S411, as the process of the technique 3 shown in FIG. 10, the
10 sound field reproduction signal processing unit 135 adds the acoustic content d of the
site C to the output buffer of each selected output speaker.
[0141]
Then, in step S415, the signal processing apparatus 1B outputs contents of
each output buffer from the speaker group selected in the above step S406 through
15 the DACIamplifying unit 23.
[0 1421
As described above, in the acoustic system according to the present
embodiment, the voice of the user A acquired at the site A is subjected to rendering
using the first acoustic information parameter measured at the site C, and is
20 reproduced from the plurality of speakers 20B of the site B with the echo in the site
C. Ful-thel; the voice of the user B himselfil~erselfa cquired at the site B is subjected
to rendering using the second acoustic infonnation parameter measured at the site C,
and is reproduced from the plurality of speakers 20B of the site B with the echo in
the site C. In addition, the acoustic content acquired at the site C is reproduced
25 from the plurality of speakers 20B of the site B.
[0 1431
In this way, the acoustic system according to the present embodiment can
provide, when one site (site B in this case) cooperates (speaks on the telephone) with
another site (site A in this case), the feeling of being absorbed in the third space (site
30 C in this case). The user B can obtain a sound field feeling as if the user B were
located at the site C with the user A, and can feel more affluent reality.
Further, the sound field reproduction signal processing unit 135 can control
an audio image of the received audio signal (user A's voice) using the speaker group
arranged around the user B. For example, as the at-ray speaker (beam forming) is
5 formed by the plurality of speakers, the sound field reproduction signal processing
unit 135 can reconstruct the user A's voice in the user B's eal; and can reconstruct the
user A's audio image outside the acoustically closed surface surrounding the user B.
[0145]
Note that the continuous execution of steps S403 and S406 enables the
10 signal processing unit 13 to update the speaker group forming the acoustically closed
surface surrounding the user B in accordance with the movement of the user B.
Hereinafter, specific description will be given with reference to FIG. 18A and FIG.
18B.
[0146]
15 FIG. 18A is a diagram illustrating the case where the sound field 42
constructed in the site B is fixed. As shown in FIG. 18A, let us assume the case
where, first, the plurality of speakers 20B are selected (steps S403 and S406) for
forming the acoustically closed surface 40 surrounding the user B, and the sound
field 42 for providing the user B with the feeling of being absorbed in the site C is
20 constmcted. In this case, when the user B moves inside the room or goes out of the
room so that the user B moves out of the acoustically closed surface 40, the user B is
not in the sound field 42, and thus, the user B cannot obtain the feeling of being
absorbed in the site C.
[0 1471
25 Accordingly, as described above, steps S403 and S406 are executed
continuously, and the speaker group forming the acoustically closed surface
surrounding the user B is updated in accordance with the movement of the user B.
FIG. 18B is a diagram illustrating the case where the sound field 42 constructed in
the site B is mobile.
30 [0148]
As shown in FIG. 18B, a speaker group (speaker 20B') forming an
acoustically closed surface 40' surrounding the user B is newly selected (updated) in
accordance with the movenient of the user B, and a sound field 42' is newly
constructed by the updated plurality of speakers 20B'.
[0149]
5 Heretofore, each operation process of the acoustic system according to the
present embodiment has been described in detail. Next, a supplement of the present .
embodiment will be described.
[01 SO]
<5. Supplement>
10 [5-1. Modified example of command input]
In the embodiment above, a command is input by a voice, but the method of
inputting a command in the acoustic system according to the present disclosure is not
limited to the audio input and may be another input method. For example, the
signal processing apparatus 1 according to the present embodiment may detect a user
15 operation on each switch (which is an example of an operation input part) arranged
around a user, and may recognize a command such as a call origination request or the
like. Further, in this case, the signal processing apparatus 1 is capable of accepting
the designation (name or the like of the target user) of the call origination destination
and the designation of the place (name of the place) that the user is to be absorbed in,
20 using a touch panel or the like (an example of operation input part) disposed near the
user.
[0151]
Further, the recognizing unit 17 of the signal processing apparatus 1 may
analyze a gesture of the user on the basis of an image obtained by an imaging unit
25 disposed near the user or a detection result acquired by an IR thermal sensor, and
may recognize the gesture as a command. For example, in the case where the user
performs a gesture of making a telephone call, the recognizing unit 17 recognizes the
gesture as a call origination request command. Further, in this case, the signal
processing apparatus 1 may accept the designation (name or the like of the target
30 user) of the call origination destination and the designation of the place (name of the
place) that the user is to be absorbed in from a touch panel or the like disposed near
the user, or maydetermine the designations on the basis of voice analysis.
[0152]
Further, in the case where a user feels that it is difficult to hear the sound
while speaking on the telephone with another user who is located at another site (in
5 the case where the voice of the other user is reproduced from the plurality of
microphones 10 arranged around the user), the user may send a request for control of
reproducing sound by a gesture. Specifically, for example, the recognizing unit 17
may recognize a gesture of bringing the user's open hand close to the ear or a gesture
of bringing both hands close to the top of the head to imitate ears of a rabbit as a
10 command to increase the volume.
[0153]
As described above, the method of inputting a command in the acoustic
system according to the present disclosure is not limited to the audio input, and may
be the method using the switch operation or the gesture input.
15 [0154]
[5-2. Example of another command]
In the embodiment above, there has been described the case where a person
is designated as a given target and a call origination request (call request) is
recognized as a command, but the command of the acoustic system according to the
20 present disclosure is not limited to the call origination request (call request), and may
be another command. For example, the recognizing unit 17 of the signal processing
apparatus 1 may recognize a command in which a place, a building, a program, a
music piece, or the like which has been designated as a given target is reconstructed
in the space at which the user is located.
25 [0155]
Further, the acoustic system according to the present embodiment may
reconstruct, in the space at which the user is located, another space in real time, or a
past space of a designated place, building, or the like (for example, an excellent
concert which was performed in the past in a famous theater).
30 [0156]
15-3. Conversion of large space into small space]
m e , the embodiment described above assumes a case where the closed
surface of the site B (reproduction environment) side and the closed surface of the
site C (measuring environment) side have substantially the same sizes and shapes
with each other, but the present embodiment is not limited thereto. For example,
5 also in the case where the size of the closed surface of the reproduction environment
side is smaller than the size of the closed surface of the measuring environment side,
the present embodiment is capable of reconst~ucting the sound field (expanse of
space) of the measuring environment in the reproduction environment.
[0157]
10 Such a conversion process from the large space into the small space may be
performed, before executing the techniques 1 and 3 described above with reference
to FIG. 10, on the received audio signal (audio signal of the user A or acoustic
content) in the signal processing apparatus 1. Further, by performing such a
conversion process in real time, the acoustic system according to the present
15 embodiment can solve the problem of inconsistency in a correspondence relationship
of the positions of speakers and microphones between the measuring environment
side and the reproduction environment side.
[Ol58]
Specifically, for example, the sound field reproduction signal processing
20 unit 135 may use the signal process using the transfer function disclosed in JP
4775487B. In JP 4775487B, a transfer function (measurement data of impulse
response) is determined in a sound field of a measuring environment, an audio signal
subjected to an arithmetic process based on the transfer function is reproduced in a
reconstruction environment, and the sound field (reverberation, localization of an
25 audio image, or the like) of the measuring environment is reconstructed in the
reconstruction enviromnent. Hereinafter, with reference to FIGS. 19A to 19C, there
will be described a signal process using a transfer function (measurement data of
impulse response).
[OI 591
30 FIG. 19A is a diagram illustrating measurement in a measurement target
space. First, as shown in FIG. 19A, in a measurement target space (large space), M
microphones forming a large closed surface P are a~anged, and measurement for
Mch (M-speakers-output channel) is performed. Further, the positions of the
respective M microphones are represented by PI, P2, ..., PM. Then, a measurement
signal is output fiom each of speakers (SP) arranged outside the closed surface P, and
5 an impulse response from each of the speakers to each of the microphones arranged
at P1, P2, ..., PM is measured. The thus measured impulse response (transfer
function) is shown in Expression (3) in FIG. 19A.
[0160]
Next, with reference to FIG. 198, measurement in an anechoic room will be
10 described. As shown in FIG. 198, in the anechoic room, M speakers forming a
large closed surface P are ananged, N microphones forming a small closed surface Q
are arranged inside the closed surface P, and measurement for Nch (N-speakersoutput
channel) is performed. Here, the positions of the respective M speakers are
represented by P1, P2, ..., PM, which are the same positions as those of FIG. 19A.
15 Further, the positions of the respective N microphones are represented by Q1, Q2, ...,
QN.
[0161]
Then, a sound (measurement signal) acquired by the microphone disposed at
P1 of FIG. 19A is output from the speaker disposed at P1, and an impulse response to
20 each of the microphones disposed at Q1, 42, ..., QN is measured. Next, a sound
(measurement signal) acquired by the microphone disposed at P2 of FIG. 19A is
output fiom the speaker disposed at P2, and an impulse response to each of the
microphones disposed at Q1, 42, ..., QN is measured. In this way, all of impulse
responses from M speakers to microphones disposed at Q1, 42, ..., QN are measured.
25 [0162]
The thus measured M vertical vectors are put into MxN matrix operation,
and thus are converted into N outputs. That is, the thus measured impulse
responses (transfer functions) are put into a matrix (generating a matrix of transfer
function group) as shown in Expression (4) of FIG. 19B, and in this way, the
30 conversion from the large space (coefficient for Mcli) into the small space
(coefficient for Nch) is achieved.
[0 1631
Next, with reference to FIG. 19C, reconstruction in a reproduction target
space (small space) will be described. As sl~ownin FIG. 19C, N speakers forming a
small closed surface Q surrounding a user B are arranged in the reproduction target
5 space. Here, the positions of the respective N speakers are represented by Q1,
42, ..., QN, which are the same positions as those of FIG. 19B.
[0 1641
In this case, when the respective speakers disposed at Q1, Q2, ..., QN output
the received audio signals (for example, the voice of the user A: audio signal S), the
10 output of each speaker is determined by Expression (5) shown in FIG. 19C.
Expression (5) is an arithmetic operation using the impulse responses (transfer
functions) shown in Expression (3) and Expression (4).
[0165]
In this way, in the case where the audio image of the user A is localized
15 outside the closed surface Q, for example, a wave front at the time at which the voice
emitted by the user A located outside the closed surface Q crosses the closed surface
Q as shown in FIG. 19C is estimated, and the estimated wave front is created inside
the closed surface Q. In this case, the inconsistency between the number of
microphones in the measurement target space and the number of speakers in the
20 reproduction target space is converted by Expression (5), and in this way, the
acoustic system according to the present embodiment can reconstruct the sound field
of the large closed surface P in the small closed surface Q.
[0 1661
[5-4. Video construction]
25 In addition, the provision of the feeling of being absorbed in the third space
is achieved by the sound field construction (sound field reproduction process) in the
embodiment described above, but the acoustic system according to the present
disclosure is not limited thereto, and may also use video construction.
[O 1671
30 For example, when the signal processing apparatus 1 receives a video
obtained by a plurality of image sensors arranged at the third space (site C) from a
given server and reproduces at thc~site B a voice acquired by a plurality of
microphones arranged at the site A, the signal processing apparatus 1 may reproduce
the video of the site C and may reconstruct the space of the site C.
[0168]
5 The reproduction of the video may be space projection using hologram
reproduction, and may be reproduction using a television in a room, a display, or a
head mounted display. In this way, by performing video construction together with
the sound field construction, the user can be provided with a feeling of being
absorbed in the third space, and can feel more affluent reality.
10 [0169]
[5-5. Another system configuration example]
In the system configuration of the acoustic system according to the
embodime~~det scribed with reference to FIG. 1 and FIG. 2, both the call origination
side (site A) and the call destination side (site B) have the plurality of microphones
15 and speakers around the user, and the signal processing apparatuses 1A and 1B
perform the signal process. However, the system configuration of the acoustic
system according to the present embodiment is not limited to the configuration
shown in FIG. 1 and FIG. 2, and may be the configuration as shown in FIG. 20, for
example.
20 [0170]
FIG. 20 is a diagram showing another system configuration of the acoustic
system according to the present embodiment. As shown in FIG. 13, in the acoustic
system according to the present embodiment, a signal processing apparatus 1, a
coinmunication terminal 7, and a management selves 3 are connected to each other
25 tlwough a network 5.
[0171]
The communication terminal 7 includes a mobile phone terminal or a
smartphone including a normal single microphone and a normal single speaker,
which is a legacy interface compared to an advanced interface space according to the
30 present embodiment in which a plurality of microphones and a plurality of speakers
are arranged.
[0 1721 --
The signal processing apparatus 1 according to the present embodiment is
connected to the nornlal communication terminal 7, and can reproduce a voice
received from the communication terminal 7 from the plurality of speakers arranged
5 around the user. Further, the signal processing apparatus 1 according to the present
embodiment can transmit the voice of the user acquired by the plurality of
microphones arranged around the user to the communication terminal 7.
[0173]
As described above, according to the acoustic system according to the
10 present embodiment, a first user located at the space in which the plurality of
microphones and the plurality of speakers are arranged nearby can speak on the
telephone with a second user carrying the normal communication terminal 7. That
is, the configtuation of the acoustic system according to the present embodiment may
be that one of the call origination side and the call destination side is the advanced
15 interface space according to the present embodiment in which the plurality of
microphones and the plurality of speakers are arranged.
[0 1 741
[5-6. Autonomous acoustic system]
As described with reference to FIGS. 1 to 3, in the embodiment described
20 above, the signal processing apparatus 1 controls input and output of the plurality of
microphones 10 and the plurality of speakers 20 arranged around the usel; but the
configuration of the acoustic system according to the present disclosure is not limited
thereto. For example, a plurality of autonomous n~icrophone and speaker devices
may be arranged around the user, the devices may conlmunicate with each other, an
25 acoustically closed surface surrounding the user may be formed by determination of
each of the devices, and thus, the construction of the sound field described above
may be achieved. Hereinafter, such an autonomous acoustic system may be
described specifically with reference to FIGS. 21 to 24. Note that, in the
autononlous acoustic system described here, a case will be described as an example
30 where a plurality of devices 100 each including one n~icrophone 10 and one speaker
20 are arranged around the user.
[0175] -
(System configuration)
FIG 21 is a diagsam showing an example of a system configuration of an
autonomous acoustic system according to the present embodiment. As shown in
5 FIG. 21, the autonomous acoustic system according to the present disclosure includes
a plurality of devices 100 (100-1 to 100-4), a management server 3, a user ID DB 6,
a selvice log DB 8, and a user individual DB 9. Further, as shown in FIG. 21, the
management server 3, the user ID DB 6, the service log DB 8, and the user individual
DB 9 are connected to each other through a network 5.
10 [0176]
-Device
The plurality of devices 100 (100-1 to 100-4) are arranged everywhere in
the world such as rooms, houses, buildings, outdoor sites, regions, and countries. In
the example shown in FIG. 21, a case is shown where the plurality of devices 100 are
15 arranged on the walls and a floor of a public place such as a department store or a
station. Further, the plurality of devices 100 (100-1 to 100-4) are communicable
with each other via wire or radio, and are reporting each other their capabilities.
Further, at least one device 100 (for example, device 100-1) out of the plurality of
devices 100 (100-1 to 100-4) is accessible to the network 5. Further, the devices
20 100 each include a microphone 10 and a speaker 20. Note that the configuration of
the device according to the present embodiment will be described later with
reference to FIG. 22.
[0177]
-Reporting of information
25 As described above, the plurality of devices 100 (100-1 to 100-4) report
each other their capabilities (characteristic information). The characteristic
information to be reposted includes a device ID, a service that the device can provide,
an owner ID of the device, a device lending attribute, and the like. Here, the owner
ID represents the ID of the owner of (person who has installed) the device 100, and it
30 is estimated that each of the devices 100 (100-1 to 100-4) shown in FIG. 21 is
installed by an individual or a corporation. Further, the device lending attribute
represents information indicating an attribute of a service of which lending (use) is
permitted in advance by an individual or a corporation wholwhich has installed the
device.
[0178]
5 Further, the reporting of the information is autonomously and distributedly
performed regularly or on demand. Further, the method of reporting information
according to the present embodiment may use procedure generally known as a
method of constructing a mesh network (beaconing in IEEE802.11~).
[0179]
10 Further, the device 100 may also include a plurality of types of
con~munication IF'S. In Illis case, each device 100 regularly checks which
communication IIF is communicable with which device, and the communication VF
that is capable of co~nmunicating directly with the largest number of devices is
activated with high priority.
15 [0180]
Further, each device 100 may fonvard information reported fiom a device
nearby to a device located several hops away using a wireless I/F, or may transmit
the information to another device through the network 5.
[0181]
20 - Management sewer
As described with reference to FIG. 2 and FIG. 5, tlie management server 3
manages an absolute position (current position) of each user. Alternatively, the
management server 3 may also be the management server 3' which accumulates the
acoustic information parameter of the third space or the like described with reference
25 to FIG. 11.
[0182]
-Service log DB
The service log DB 8 is a storage unit which stores contents of a service, a
device 100 that contributed to the provision of the service, and a user who provides
30 the service in association with eacli other. In this way, the following can be grasped
using the service log DB 8: which device is used for what kind of service provision,
what service is provided to which user, and the like.
[0183]
Further, a service log stored in the service log DB 8 may be used later as
charging information to a user who has used a service, or may be used as kickback
5 information to a person (individuatlcorporation) who has installed the device 100
contributed to the service provision. Here, the kickback infornlation is information
used when providing the owner of (person who has installed) the device 100 with a
part of usage charge depending on a rate of contribution (frequency) or the like to the
service provision of the device 100. Further, the service log stored in the service
10 log DB 8 may be transmitted to the user individual DB 9 as metadata of a user's
action.
[0184]
-User individual DB
The user individual DB 9 stores metadata of user's actions transmitted from
15 the service log DB 8 as data held by the user. The data stored in the user individual
DB 9 may be used in various individualized services and the like.
[0185]
-User ID DB
The user ID DB 6 is a storage unit which stores a registered user's ID (name,
20 identification number, or biological information) and a service the provision of which
is permitted to the user in association with each other. The user ID DB 6 is used
when the device 100 performs user authentication.
[OI 861
Heretofore, the system configuration of the autonomous acoustic system
25 according to the present embodiment has been described with reference to FIG. 21.
Subsequently, with reference to FIG. 22, a configuration of the device 100 (signal
processing apparatus) according to the present embodiment will be described.
[OI 871
(Configuration of device)
30 FIG. 22 is a block diagram showing a configuration of the device 100
according to the present embodiment. As shown in FIG. 22, the device 100 includes
a microphone 10, an amplifying/ADC unit 11, a signal processing unit 200, a
recognizing unit 17, an identifying unit 18, a communication I/F 19, a user
authentication unit 25, a user position estimating unit 16, a DAC/amplifying unit 23,
and a speaker 20. Since the descriptions of the microphone 10, the
5 amplifying/ADC unit 11, the recognizing unit 17, the identifying unit 18, the
communication UF 19, the user position estimating unit 16, the DACIamplifying unit
23, and the speaker 20 have been given with reference to FIG. 3, the descriptions
thereof will be omitted here.
[01 881
10 -Signal processing unit
The signal processing unit 200 includes a high S/N processing unit 210 and
a sound field reproduction signal processing unit 220. In the same manner as the
high S/N processing unit 133 shown in FIG. 3, the high S/N processing unit 210 has a
function of processing an audio signal output from the amplifying/ADC unit 11 to
15 form a monaural signal having high articulation and a high S/N ratio. Specifically,
the high S/N processing unit 210 performs sound source separation, and performs
dereverberation and noise reduction. The audio signal processed by the high S/N
processing unit 210 is output to the recognizing unit 17 and subjected to voice
analysis for command recognition, and transmitted to an external device through the
20 communication VF 19.
[0189]
The sound field reproduction signal processing unit 220 performs a signal
processing on the audio signal reproduced from the speaker 20, and performs control
such that a sound field is localized near the user's position. Further, the sound field
25 reproduction signal processing unit 220 cooperates with another adjacent device 100,
and controls the contents (audio signal) output from the speaker 20 so as to form the
acoustically closed surface surrounding to the uses.
[0190]
-User authentication unit
30 The user authentication unit 25 perfornls user authentication by sending a
query to the user ID DB 6 on the network 5 though the communication I/F 19 on the
basis of a user ID acquired from a tagsuch as an RFID possessed by a user. For
example, in the case where the acquired user ID corresponds to an ID registered in
the user ID DB 6 in advance, the user authentication unit 25 authenticates the user as
a user to whom the provision of the service is permitted.
5 [0191]
Heretofore, the configuration of the device 100 according to the present
embodiment has been described in detail. Subsequently, with reference to FIG. 23,
an operation process of the autonomous acoustic system according to the present
embodiment will be described.
10 [0192]
(Operation process)
FIG. 23 is a flowchart showing an operation process of the autonomous
acoustic system according to the present embodiment. As shown in FIG. 23, first,
in step S503, the device 100 performs an advance preparation process. Specifically,
15 the device 100 and another device 100 report the above-mentioned characteristic
information with each other, and the device 100 confirms a device which can
cooperate with the device 100 (which is worthy of trust).
[0193]
For example, the device 100-1 may confirm whether the adjacent device
20 100-2 is worthy of trust on the basis of an owner ID, a lending attribute, or the like
included in the characteristic information received from the adjacent device 100-2.
The devices which have confirmed that are worthy of trust with each other start their
actuators and capture the output results by the sensors of the adjacent devices, for
example. In this way, the devices can grasp what sort of cooperation can be
25 performed by combining the characteristics of the devices. Such a confirmation
procedure may be performed regularly. Further, tlwongh the confirmation
procedure, each device 100 is capable of loosely grasping what sort of service can be
provided in the space in which the plurality of adjacent devices 100 (100-1 to 100-4)
are airanged.
30 [0194]
Next, in step S506, in the case where a user enters the space in which the
plurality of devices 100 (100-1 to 100-4) are arranged, the devices 100 perform user
authentication. For example, as shown in FIG. 21, in the case where the user
possesses a tag 60 such as an WID, the devices 100-1 to 100-4 arranged nearby inay
receive characteristic information reported from the tag 60 and may detect the
5 entrance of the user. Then, when detecting the entrance of the user, each device 100
authenticates whether the user is a user to whom a service can be provided by
sending a request to the user ID DB 6 on the network 5, on the basis of a user ID
included in the characteristic information reported fiom the tag 60.
[0195]
10 Note that, in the case where the user does not possess the tag 60, the devices
100-1 to 100-4 may detect the entrance of the user using sensors (microphone,
camera, human sensor, thermal sensor, and the like). Further, the devices 100-1 to
100-4 may extract the ID (biological information or the like) of the user by analyzing
the detection results of the sensors.
15 [0196]
Furthel; in the example shown in FIG. 21, the device 100-1 among the
devices 100-1 to 100-4 holds an access route to the user ID DB 6. In this case, the
user authentication may be performed as follows: the device 100-2, 100-3, or 100-4
which has acquired a user ID transmits the user ID to the device 100-1, and the
20 device 100-1 sends a request to the user ID DB 6. In this way, it is not necessary
that all of the plurality of devices 100-1 to 100-4 each hold an access to the user ID
DB 6.
[0 1971
Further, a result of user authentication performed by one device 100 among
25 the devices 100-1 to 100-4 is shared between other devices 100 disposed nearby, and
the devices 100-1 to 100-4 can grasp that a service can be provided to the user.
[0198]
Next, in step S509, the device 100 recognizes a command (service request)
from the user. Here, the device 100 may notify the tag 60 of information of the
30 services that can be provided to the authenticated user. The tag 60 can notify the
user of what sort of service can be received at the place using output means (not
shown) such as a speaker or a display unit. Further, the tag 60 also specifies a
service that the user cwently desires from the user's command input (by microphone,
gyro, key touch, or the like), and notifies the devices 100-1 to 100-4 disposed nearby
of the service.
5 [0199]
Note that, in the case where the user does not possess the tag 60, the devices
100-1 to 100-4 may analyze the voice or the gesture of the user using sensors
(microphone, camera, human sensor, thermal sensor, and the like), and may
recognize the service that the user desires.
10 [0200]
The command recognized by the device 100 here may be, in addition to the
above-mentioned call origination request (call request), a command that requests
reconstmction of a place, a building, a program, a music piece, or the like which has
been designated as a given target.
15 [0201]
Next, in step S512, in the case where the requested service is a service
permitted to the user, the devices 100-1 to 100-4 start providing the service.
Specifically, for example, the devices 100-1 to 100-4 start operations of a sensor (for
example, microphone 10) and an actuator (for example, speaker 20), and make
20 communication paths between the devices be in operation states. Further, the
devices 100-1 to 100-4 may cooperate with each other and determine their own
operations on the basis of the type of the sesvice to provide and the amount of
available con~municationre sources.
[0202]
25 Further, the in the case where the device 100 has a plurality of types of
communication IP's, a communication IIF used for trafic transmission may be
operated according to the needs on the basis of the amount of information to be
provided. Further, the device 100 may increase an operation duty cycle as needed,
and an electric power saving mode may be partially cancelled. In addition, the
30 devices 100 may set each other's traasmissio~l/receptio~tiim e period used for
comnlunication, to thereby transit to the state where the band can be supplied stably
(starting up of resewation access to IEEE802.1 Is). --
[0203]
Next, in step S515, in the case where the completion of the service is
designated by the user, the device 100 co~npletes the provision of the service.
5 Specifically, for example, the device 100 finishes the operations of a sensor (for
example, microphone 10) and an actuator (for example, speaker 20), and make
communication paths between the devices be in stop states.
[0204]
Next, in step S518, the device 100 notifies the sewice log DB 8 of the
10 contents to which the device 100 has contributed in the latest service provision.
Further, the device 100 may also notify the service log DB 8 of information of the
user (authenticated user) to whom the service has been provided.
[0205]
Heretofore, the operation process of the autonomous acoustic system
15 according to the present embodiment has been described specifically with reference
to FIG. 23. Hereinafter, an additional description of the autonomous acoustic
system according to the present embodiment will be described.
[0206]
(Service continuation)
20 The autonomous acoustic system according to the present embodiment can
continuously provide the user with the service even in the case where the
authenticated user walks and moves to another place, by changing the device 100
that provides the user with the service (device 100 that operates). Such a changing
process is performed on the basis of radio wave intensity from the tag 60 possessed
25 by the user and input signals froom sensors (microphone, camera, human sensor, and
the like) possessed by each device, for example. Hereinaftel; description will be
given with reference to FIG. 24.
[0207]
FIG. 24 is a diagram illustrating change of operating devices in accordance
30 with movement of a user in the autonomous acoustic system according to the present
embodiment. As shown in FIG. 24, here, let us assume the case where the user
movecaway from the device 100-1 and the device 100-2 which are operating for
providing the user with a service and comes close to the devices 100-5 and 100-6
which are not performing the operation for the service provision.
[0208]
5 I11 this case, the devices 100-5 and 100-6 detect that the user has approached
on the basis of radio wave intensity from the tag 60 possessed by the user or input
signals fiom sensors included in the devices 100-5 and 100-6. Then, the devices
100-5 and 100-6 receive the ID of the user and information of a service that may be
provided to the user from the device 100-2 and the like which is placed adjacently.
10 [0209]
Then, the devices 100-5 and 100-6 start service provision to the user on the
basis of the received information. On the other hand, the devices 100-1 and 100-2
wlkh have been performing the service provision finish the service provision when
determining that the user moves out fiom the area in which the sensors and the
15 actuators of the devices 100-1 and 100-2 can provide the service, and bring the
operation of the devices and the communication path down.
[0210]
In this way, if the user moves while receiving the service provision, the
device 100 disposed near the destination can takes over the user ID and the contents
20 of the service, and can continuously provide the user with the service.
[0211]
(Access route to network 5)
In the autonomous acoustic system according to the present embodiment
described with reference to FIG. 21, at least one device 100 (here, device 100-1) out
25 of the devices 100-1 to 100-4 has an access route to the network 5. However, the
configuration of the autonomous acoustic system according to the present disclosure
is not limited to the example shown in FIG. 21, and there may be a case where the
devices 100-1 to 100-4 are in a closed network and do not have the access route to
the outside (network 5).
30 [0212]
In sucli a case, the devices 100-1 to 100-4 may use the tag 60 possessed by
the user, for example, as an access gateway to the outside. That is, under the state
in which the devices 100-1 to 100-4 report each other specific information, at the
time point at which the tag 60 appears, the devices 100-1 to 100-4 send a query to the
user ID DB 6 on the network 5 through the tag 60 and perform user authentication.
5 [0213]
(Service provision to plurality of users)
Next, with reference to FIG. 25, there will be described service provision in
the case where a plurality of users enter the space in which the devices 100-1 to 100-
4 are arranged.
10 [0214]
FIG. 25 is a diagram illustrating a case where a service is provided to a
plurality of users in the autonomous acoustic system according to the present
embodiment. As shown in FIG. 25, in the case where the plurality of users enter the
space in which the devices 100-1 to 100-4 are arranged and each perform a service
15 request, each of the devices 100-1 to 100-4 provides a plurality of services.
[02 1 51
In this case, although the operations of the devices 100-1 to 100-4 with
respect to each user is as described with reference to FIGS. 21 to 24, a tag 65
possessed by a user 2 is regarded as one of the devices arranged nearby for a tag 60
20 possessed by a user I. In addition, the tag 60 possessed by the user 1 is also
regarded as one of the devices arranged nearby for the tag 65 possessed by the user 2.
[02 161
Accordingly, the devices 100-1 to 100-4 perform reporting of the
characteristic information also with the tag 60 or the tag 65, and confirms whether
25 the device is worthy of trust. In this way, the characteristics of the tag 60 or the tag
65 may be used for the service provision.
[0217]
For example, in the case where the devices 100-1 to 100-4 are in a closed
network, the devices 100-1 to 100-4 grasp that the tag 65 holds an access route to the
30 outside network 5 fsom the reporting of the characteristic information with the tag 65.
Then, in providing the user 1 possessing the tag 60 with the service, the devices 100-
1 to 100-4 use the tag 65 possessed by t h e ~ ~2e ars one of the devices, and thereby
being able to connect to the outside network 5.
[0218]
In this way, not being limited to the devices 100-1 to 100-4 arranged around
5 the user 1, but also the tag 65 possessed by the user 2 located nearby may provide the
user 1 possessing the tag 60 with an access to the outside network.
[0219]
Note that, in such a case, the contents of the sellrice provided by the tag 65
may be written in the sellrice log DB 8, and may be used when later performing, with
10 respect to the user 2 possessing the tag 65, kickback based on the fact of being
contributed to the service provision to the user 1.
[0220]
16. Conclusion>
As described above, in the acoustic system according to the present
15 embodiment, it becomes possible to provide a feeling of being absorbed in a third
space at a time of causing the space surrounding the user to cooperate with another
space. Specifically, the acoustic system according to the present embodiment can
reproduce a voice and an image corresponding to a first given target (person, place,
building, or the like) though a plurality of speakers and displays arranged around the
20 user. Further, in this case, the acoustic system according to the present embodiment
can reproduce a space of a second given target (place or the like), and can provide
the feeling of being absorbed in the second given target and the ambience. In this
manner, using the microphones 10, tlie speakers 20, the image sensors, and the like
ananged everywhere, indoor sites and outdoor sites, it becomes possible to
25 substantially augment over a large area the body such as the mouth, eyes, ears of the
user, and to achieve a new comniunication method.
[0221]
In addition, since microphones and image sensors are arranged everywhere
in the acoustic system according to the present embodiment, the user does not have
30 to cany a smartphone or a mobile phone terniinal. The user specifies a given target
using a voice or a gesture, and can establish connection with a space surrounding tlie
given target.
[0222]
Further, the configuration of the acoustic system for achieving such a new
commt~nicationm ethod may be achieved by the signal processing apparatus which
5 controls a plurality of microphones and a plurality of speakers. Further, the
acoustic system according to the present embodiment may be achieved as follows:
the devices such as each autonomous microphone and each autonomous speaker
cooperate with another adjacent device.
[0223]
10 The preferred embodiments of the present disclosure have been described
above with reference to the accompanying drawings, whilst the present invention is
not limited to the above examples, of course. A person skilled in the art may find
various alterations and modifications within the scope of the appended claims, and it
should be understood that they will naturally come under the technical scope of the
15 present invention.
[0224]
For example, the configuration of the signal processing apparatus 1 is not
limited to the configuration shown in FIG 3, and the configuration may be that the
recognizing unit 17 and the identifying unit 18 shown in FIG. 3 are not provided to
20 the signal processing apparatus 1 but are provided on the server side which is
connected thereto through a network. In this case, the signal processing apparatus 1
transmits an audio signal output from the signal processing unit 13 to the sewer
through the communication I/F 19. Further, the sewer performs the command
recognition and the process of identifying a given target (person, place, building,
25 program, music piece, or the like) on the basis of the received audio signal, and
transmits the recognition results and the access destination information
corresponding to the identified given target to the signal processing apparatus 1.
[0225]
Additionally, the present technology may also be configured as below.
30 (1)
An information processing system including:
a recognizing unit ~~p configured to recognize a first target and a second target -
on the basis of signals detected by a plurality of sensors arranged around a specific
user;
an identifying unit configured to identify the first target and the second
5 target recognized by the recognizing unit;
an estimating unit configured to estimate a position of the specific user in
accordance with the a signal detected by any one of the plurality of sensors; and
a signal processing unit configured to process each of signals acquired from
sensors around the first target and the second target identified by the identifying unit
10 in a manner that, when being output from a plurality of actuators arranged around the
specific user, the signals are localized near the position of the specific user estimated
by the estimating unit.
(2)
The information processing system according to (I),
15 wherein the first target is a given person and the second target is a given
place, and
wherein the signal processing unit processes signals acquired by sensors
around the given person and signals acquired by sensors around the given place.
(3
20 The information processing system according to (I),
wherein the first target is a given person and the second target is a given
place, and
wherein the signal processing unit processes signals acquired in real time by
sensors around the given person and signals that have been acquired in advance by
25 sensors around the given place and have been accutnulated.
(4)
The information processing system according to any one of (1) to (3),
wherein sensors around the first target and sensors around the second target
are arranged dispersedly at remote places.
30 (5)
The information processing system according to any one of (1) to (4),
wherein -- the plurality of sensors arranged around the specific user are
micropliones, and
wherein the recognizing unit recognizes the first target and the second target
on the basis of audio signals detected by the microphones.
5 (6)
The information processing system according to any one of (1) to (4),
wherein the plurality of sensors arranged around the specific user are image
sensors, and
wherein the recognizing unit recognizes the first target and the second target
10 on the basis of captured images obtained by the image sensors.
(7)
The information processing system according to any one of (1) to (6),
wherein sensors around the first target are of a different type from sensors
around the second target.
16 (8)
The information processing system according to any one of (1) to (7),
wherein the signal processing unit performs processes of modifying signals
acquired by sensors around the first target on the basis of a characteristic of a
parameter corresponding to the second target, and executing addition to signals
20 acquired by sensors around the second target.
(9)
The information processing system according to any one of (1) to (8),
wherein the signal processing unit processes each of the signals acquired
from sensors around the first target and the second target in a manner that each of the
25 signals are localized near a sense organ of the specific uses.
(10)
The information processing system according to any one of (1) to (9),
wherein the sensors around the first target and the second target are
microphones,
30 wherein the plurality of actuators arranged around the specific user are a
plmality of speakers, and
wherein the signal processing unit processes respective audio signals
acquired by the microphones around the first target and the second target in a manner
that a sound field is formed near a position of the specific user when output from the
plurality of speakers, on the basis of respective positions of the plurality of speakers
5 and the estimated position of the user.
(11)
The information processing system according to (lo),
wherein the estimating unit continuously estimates a position of tlie specific
user, and
10 wherein the signal processing unit processes the respective audio signals in
a manner that a sound field is formed near the position of the specific user in
accordance with a change in the position of the specific user.
(12)
An information processing system including:
15 a recognizing unit configured to recognize a first target and a second target
on the basis of signals detected by sensors around a specific user;
an identifying unit configured to identify the first target and the second
target recognized by the recognizing unit; and
a signal processing unit configured to generate signals to be output from
20 actuators around the specific user on the basis of signals acquired by a plurality of
sensors arranged around the first target and the second target identified by the
identifying unit.
(13)
The inforniation processing system according to (12),
25 wherein the first target is a given person and the second target is a given
place, and
wherein the signal processing unit processes signals acquired by a plurality
of sensors arranged around tlie given person and signals acquired by a plurality of
sensors arranged around the given place.
30 (14)
The information processing system according to (12),
wherein the first target is a given person and the second target is a given
place, and
wherein the signal processing unit processes signals acquired in real time by
a plurality of sensors arranged around the given person and signals that have been
5 acquired in advance by sensors around tlie given place and have been accumulated.
(1 5)
A storage medium having a program stored therein, the program being for
causing a computer to function as:
a recognizing unit configured to recognize a first target and a second target
10 on the basis of signals detected by a plurality of sensors arranged around a specific
user;
an identifying unit configured to identify the first target and the second
target recognized by the recognizing unit;
an estimating unit configured to estimate a position of the specific user in
15 accordance with the a signal detected by any one of the plurality of sensors; and
a signal processing unit configured to process each of signals acquired from
sensors around the first target and the second target identified by the identifying unit
in a manner that, when being output from a plurality of actuators arranged around the
specific user, the signals are localized near the position of the specific user estimated
20 by the estimating unit.
(16)
A storage medium having a program stored therein, the program being for
causing a computer to function as:
a recognizing unit configured to recognize a first target and a second target
25 on the basis of signals detected by sensors around a specific user;
an identifying unit configured to identify the first target and the second
target recognized by the recognizing unit; and
a signal processing unit configured to generate signals to be output from
actuators around the specific user on the basis of signals acquired by a plurality of
30 sensors arranged around the first target and the second target identified by tlie
identifying unit.
Reference Signs List
[0226]
1, lA, 1B signal processing apparatus
5 3,3 ' management server
5 network
6 user ID DB
7 communication terminal
8 sellrice log DB
10 9 user individual DB
10, IOA, IOB, IOC microphone
11 amplifying/analog-to-digital converter (ADC) unit
13,200 signal processing unit
15 microphone position information database (DB)
15 16 user position estimating unit
17 recognizing unit
18 identifying unit
19 conlmunication interface (VF)
20,20A, 20B, 20C speaker
20 23 digital-to-analog converter @AC)/amplifying unit
25 user authentication unit
32 managing unit
33,34 searching unit
35 user position information DB
25 36 acoustic information parameter DB
37 acoustic content DB
40,40-1,40-2,40-3 acoustically closed surface
42 sound field
43,43-1,43-2,43-3 closed surface
30 60,65 tag
100, 100-1 to 100-4 device
13 1 microphone array processingunit
133,210high S/N processing unit
135,220sound field reproduction signal processing unit
136 convolution unit
5 137, 139howling reduction unit
138 matrix convolution unit
CLAIMS
Claim 1
An information processing system comprising:
a recognizing unit configured to recognize a first target and a second target
5 on the basis of signals detected by a plurality of sensors arranged around a specific
user;
an identifying unit configured to identify the first target and the second
target recognized by the recognizing unit;
an estimating unit configured to estimate a position of the specific user in
10 accordance with the a signal detected by any one of the plurality of sensors; and
a signal processing unit configured to process each of signals acquired from
sensors around the first target and the second target identified by the identifying unit
in a manner that, when being output from a plurality of actuators arranged around the
specific user, the signals are localized near the position of the specific user estimated
15 by the estimating unit.
Claim 2
The information processing system according to claim 1,
wherein the first target is a given person and the second target is a given
20 place, and
wherein the signal processing unit processes signals acquired by sensors
around the given person and signals acquired by sensors around the given place.
Claim 3
25 The information processing system according to claim 1,
wherein the first target is a given person and the second target is a given
place, and
wherein the signal processing unit processes signals acquired in real time by
sensors around the given person and signals that have been acquired in advance by
30 sensors around the given place and have been accumulated.
Claim 4
The information processing system according to claim 1,
wherein sensors around the first target and sensors around the second target
are arranged dispersedly at remote places.
5
Claim 5
The information processing system according to claim 1,
wherein the plurality of sensors arranged around the specific user are
microphones, and
10 wherein the recognizing unit recognizes the first target and the second target
on the basis of audio signals detected by the microphones.
Claim 6
The information processing system according to claim 1,
15 wherein the plurality of sensors arranged around the specific user are image
sensors, and
wherein the recognizing unit recognizes the first target and the second target
on the basis of captured images obtained by the image sensors.
20 Claim 7
The information processing system according to claim 1,
wherein sensors around the first target are of a different type from sensors
around the second target.
25 Claim 8
The information processing system according to claim 1,
wherein the signal processing unit performs processes of modifying signals
acquired by sensors around the first target on the basis of a characteristic of a
parameter corresponding to the second target, and executing addition to signals
30 acquired by sensors around the second target.
Claim 9
The information processing system according to claim 1,
wherein the signal processing unit processes each of the signals acquired
from sensors around the first target and the second target in a manner that each of the
5 signals are localized near a sense organ of the specific user.
Claim 10
The information processing system according to claim 1,
wherein the sensors around the first target and the second target are
10 microphones,
wherein the plurality of actuators arranged around the specific user are a
plurality of speakers, and
wherein the signal processing unit processes respective audio signals
acquired by the microphones around the first target and the second target in a manner
15 that a sound field is formed near a position of the specific user when output from the
plurality of speakers, on the basis of respective positions of the plurality of speakers
and the estimated position of the user.
Claim 11
20 The information processing system according to claim 10,
wherein the estimating unit continuously estimates a position of the specific
user, and
wherein the signal processing unit processes the respective audio signals in
a manner that a sound field is formed near the position of the specific user in
25 accordance with a change in the position of the specific user.
Claim 12
An information processing system comprising:
a recognizing unit configured to recognize a first target and a second target
30 on the basis of signals detected by sensors around a specific user;
an identifying unit configured to identify the first target and the second
target recognized by the recognizing unit; and
a signal processing unit configured to generate signals to be output from
actuators around the specific user on the basis of signals acquired by a plurality of
sensors arranged around the first target and the second target identified by the
5 identifying unit.
Claim 13
The information processing system according to claim 12,
wherein the first target is a given person and the second target is a given
10 place, and
wherein the signal processing unit processes signals acquired by a plurality
of sensors arranged around the given person and signals acquired by a plurality of
sensors arranged around the given place.
15 Claim 14
The information processing system according to claim 12,
wherein the first target is a given person and the second target is a given
place, and
wherein the signal processing unit processes signals acquired in real time by
20 a plurality of sensors arranged around the given person and signals that have been
acquired in advance by sensors around the given place and have been accumulated.
Claitn 15
A storage medium having a program stored therein, the program being for
25 causing a computer to function as:
a recognizing unit configured to recognize a first target and a second target
on the basis of signals detected by a plurality of sensors arranged around a specific
user;
an identifying unit configured to identify the first target and the second
30 target recognized by the recognizing unit;
an estimating unit configured to estimate a position of the specific user in
accordance with the a signal detected by any one ofthe plurality of sensors; and
a signal processing unit configured to process each of signals acquired from
sensors asound the first target and the second target identified by the identifying unit
in a manner that, when being output from a plurality of actuators arranged around the
5 specific user, the signals are localized near the position of the specific user estimated
- - . .
by the estimating unit.
Claim 16
A storage medium having a program stored therein, the program being for
10 causing a computer to function as:
a recognizing unit configured to recognize a first target and a second target
on the basis of signals detected by sensors around a specific user;
\ ,
an identifying unit configured to identify tlie first target and the second
target recognized by the recognizing unit; and
15 a signal processing unit configured to generate signals to be output from
actuators around the specific user on the basis of signals acquired by a plurality of
sensors asranged around the first target and the second target identified by the
identifying unit.