Abstract: A calibration system for media content comprises a memory, a media device, a plurality of different types of sensors, and a control circuitry. The memory is configured to store a media item and expected-emotions-tagging metadata for the media item. The control circuitry is configured to compute a reaction delay for a first emotional response of the audience captured by the plurality of different types of sensors for a first scene of the media item with respect to a position of the first marker. The first scene of the media item spans a first time slot that corresponds to the first scene in the media item. The control circuitry is configured to calibrate the position of the first marker to associate a first set of frames of the first scene with the first emotional response, based on at least the computed reaction delay.
This application claims priority to U.S. Provisional Patent Application Ser. No. 62/560,460 filed on September 19, 2017, the entire content of which is incorporated herein by reference.
FIELD
[0002] Various embodiments of the disclosure relate to multimedia and audience data based calibration technologies. More specifically, various embodiments of the disclosure relate to a calibration system and method for audience response capture and analysis of media content.
BACKGROUND
[0003] Recent advancements in the field of multimedia technology have led to development of various methods and techniques to encode various types of metadata (such as subtitles, commentary, citations, or emotion tags) in media content (such as a movie, a game, or other audio-video content). In a conventional system, a user may tag media content with emotion tags at various timestamps of the media content. The user may manually tag the media content at various timestamps based on input received from a production team (e.g., directors, cinematographers, composers, creators, and editors) of the media content. The production team may identify certain portions of the media content as interesting points using the manually tagged emotional tags in the media content. The production team may use the interesting points of the media content to generate a preview of media content (such as a movie trailer or a game trailer of the media content). In certain scenarios, the user may tag the media item at erroneous timestamps. In such scenarios, the interesting points of the media content, identified by the production team, may be erroneous. Further, interesting points that are assumed to elicit a certain emotional response may differ from the actual response received from an audience.
[0004] In certain scenarios, gauging an audience response in real time may find application in media or gaming industry, where the success of the media content is driven by how much the audience felt connected or engaged with the consumed media content. Currently, methods used to gauge an audience response are survey based, where a certain audience may be selected and their response may be collected manually or using certain devices, such as a camera. However, there are many challenges associated with manual or automated gauging of the audience response. For example, the audience m ay continue to respond (or elicit a same emotion) to a scene after the scene is already viewed. This may cause inaccurate estimation of an audience response for a subsequent scene.
[0005] Further limitations and disadvantages of conventional and traditional approaches will become apparent to one of skill in the art, through comparison of described systems with some aspects of the present disclosure, as set forth in the remainder of the present application and with reference to the drawings.
SUMMARY
[0006] A calibration system and method for audience response capture and analysis of media content is provided substantially as shown in, and/or described in connection with, at least one of the figures, as set forth more completely in the claims.
[0007] These and other features and advantages of the present disclosure may be appreciated from a review of the following detailed description of the present disclosure, along with the accompanying figures in which like reference numerals refer to like parts throughout.
BRIEF DESCRIPTION OF THE DRAWINGS
[0008] FIG. 1 is a block diagram that illustrates an exemplary network environment for a calibration system for audience response capture and analysis of media content, in accordance with an embodiment of the disclosure.
[0009] FIG. 2 is a block diagram that illustrates a calibration system for media content, in accordance with an embodiment of the disclosure.
[0010] FIG. 3 illustrates a timeline of a media item associated with expected-emotions-tagging-metadata, in accordance with an embodiment of the disclosure.
[0011 ] FIG. 4A illustrates a timeline of a media item inserted with a plurality of markers, in accordance with an embodiment of the disclosure.
[0012] FIG. 4B illustrates a timeline of a media item with the plurality of markers calibrated from previous positions of the plurality of markers in FIG. 4A, in accordance with an embodiment of the disclosure.
[0013] FIG. 5 illustrates a collection of graphs that show sequential transformation of different audio signals captured in a test environment, in accordance with an embodiment of the disclosure.
[0014] FIG. 6A illustrates a set of graphs that illustrate emotional response level of a first user and a second user, in accordance with an embodiment of the disclosure.
[0015] FIG. 6B illustrates a set of graphs that illustrate emotional response levels of the first user and the second user, in accordance with an embodiment of the disclosure.
[0016] FIG. 7 illustrates an exemplary first user interface that displays reaction delays, transmission delays and device weights associated with a plurality of different types of signals, in accordance with an embodiment of the disclosure.
[0017] FIG. 8 illustrates an exemplary second user interface which displays results of a comparison between one or more media items, in accordance with an embodiment of the disclosure.
[0018] FIG. 9 illustrates an exemplary third user interface, which illustrates peak points in a video signal and an audio signal associated with a media item, in accordance with an embodiment of the disclosure.
[0019] FIG. 10 illustrates an exemplary a fourth user interface, which illustrates an amalgamated audience response signal, in accordance with an embodiment of the disclosure.
[0020] FIG. 1 1 illustrates an exemplary fifth user interface, which illustrates a plurality of highlight points of a media item, in accordance with an embodiment of the disclosure.
[0021] FIGs. 12A to 12E collectively depict a flow chart that illustrates an exemplary method for calibration for audience response capture and analysis of media content, in accordance with an embodiment of the disclosure.
DETAILED DESCRIPTION
[0022] The following described implementations may be found in a disclosed calibration system for media content. Exemplary aspects of the disclosure may include the calibration system, which may further include a memory configured to store a media
item and expected-emotions-tagging metadata for the media item. The expected-emotions-tagging metadata may indicate an associative relationship between a set of time slots and a set of specified emotional states that may be expected from an audience at the set of time slots at playback of the media item. The calibration system may further include a media device configured to execute the playback of the media item to the audience and a plurality of different types of sensors configured to capture emotional response data from the audience at the playback of the media item. Further, the calibration system may include a control circuitry that may be configured to embed a first marker in the media item played by the media device. The first marker may be embedded in the media item based on the expected-emotions-tagging metadata. The control circuitry of the calibration system may be further configured to compute a reaction delay for a first emotional response of the audience captured by the plurality of different types of sensors. Such reaction delay may be computed for a first scene of the media item with respect to a position of the first marker that may span a first time slot. Such first time slot may correspond to the first scene in the media item. The control circuitry may be further configured to calibrate the position of the first marker to associate a first set of frames of the first scene with the first emotional response. The association of the first set of frames with the first scene may be done based on at least the computed reaction delay.
[0023] In accordance with an embodiment, the plurality of different types of sensors may include a set of audio sensors, a set of image sensors, and a set of biometric sensors. The control circuitry may be configured to assign a weight to each sensor of the plurality of different types of sensors prior to the playback of the media item by the media device. Thereafter, the control circuitry may be further configured to pre-calibrate the plurality of different types of sensors before the playback of the media item by the media device. The pre-calibration may be done to capture the emotional response data from the audience at the playback of the media item.
[0024] In accordance with an embodiment, the control circuitry may be further configured to pre-calibrate the set of audio sensors before the playback of the media item by the media device. Such pre-calibration may be done based on a playback of a test media item by the media device in a test environment. The playback of the test media item may be done in the presence of a test audience and an absence of the test audience.
[0025] In accordance with an embodiment, the control circuitry may be further configured to generate a noise signal based on at least a difference between a first audio signal and a second audio signal. The first audio signals may be captured by the set of audio sensors in the test environment in the presence of the test audience and the second audio signal may be captured by the set of audio sensors in the test environment in the absence of the test audience. Thereafter, the control circuitry may be further configured to assign a weight to each of the set of audio sensors for the pre-calibration of the set of audio sensors, based on the generated noise signal.
[0026] In accordance with an embodiment, the control circuitry may be further configured to pre-calibrate the set of image sensors of the plurality of different types of sensors before the playback of the media item by the media device. Such calibration may be done based on a playback of a test media item by the media device in a test environment in presence of a test audience. The control circuitry may be further configured to compute a number of faces that may be missed at a detection of number of faces between at least two image frames at playback of the test media item. An error rate may be further estimated for each image sensor of the set of image sensors, based on a change in the number of faces detected for the at least two image frames of the test media item. The control circuitry may assign a weight to each of the set of image sensors for the pre-calibration of the set of image sensors, based on the generated error rate.
[0027] In accordance with an embodiment, the control circuitry may be further configured to pre-calibrate the set of biometric sensors of the plurality of different types of sensors before the playback of the media item by the media device. Such pre-calibration of the set of biometric sensors may be done based on a measurement of biometric data in a test environment in presence of a test audience at playback of a test media. The control circuitry may further determine a standard deviation in the measured biometric data of the test audience at the playback of the test media. A weight may be further assigned by the control circuitry to each biometric sensor of the set of biometric sensors for the pre-calibration of the set of biometric sensors. Such weight may be assigned to each biometric sensor of the set of biometric sensors based on a difference in the measured biometric data of each user of the test audience with respect to the determined standard deviation.
[0028] In accordance with an embodiment, the control circuitry may be further configured to receive a plurality of different types of input signals from the plurality of different types of sensors. Such plurality of different types of input signals may be received for the audience at the playback of the media item. The plurality of different types of input signals may correspond to the emotional response data of the audience. The control circuitry may be further configured to determine a peak emotional response level for each emotional response for each user of the audience based on the received plurality of different types of input signals.
[0029] In accordance with an embodiment, the control circuitry may be further configured to normalize the received plurality of different types of input signals. Such normalization may be done based on the determined peak emotional response level for each emotional response by each user at the playback of the media item. The received plurality of different types of input signals may be normalized further based on a geographical region of the audience. The control circuitry may further synchronize and overlay the normalized plurality of different input signals in a timeline, which may be same as a playback timeline of the media item.
[0030] In accordance with an embodiment, the control circuitry may be further configured to generate an amalgamated audience response signal for the media item. Such amalgamated audience response signals may be generated based on the synchronized plurality of different input signals and a plurality of weights assigned to the plurality of different types of sensors. The control circuitry may further identify a set of common positive peaks and a set of common negative peaks in each of the plurality of different types of input signals based on the overlay of the plurality of different types of input signals. A plurality of highlight points and a plurality of lowlight points may be further calculated by the control circuitry for a plurality of scenes of the media item, based on the identified set of common positive peaks and the set of common negative peaks.
[0031] In accordance with an embodiment, the control circuitry may be further configured to compute a transmission delay for the first emotional response at the playback of the first scene of the media item. The transmission delay may be associated
with the capture and a transmission of the emotional response data from the plurality of different types of sensors. The transmission delay may be computed with respect to the position of the first marker of a plurality of markers that may be embedded in the media item. The control circuitry may further calibrate the position of the first marker to associate the first set of frames of the first scene with the first emotional response, based on the computed transmission delay. The control circuitry may shift a first timestamp of a first lowlight point of the plurality of lowlight points to a second timestamp, based on the computed reaction delay and the transmission delay.
[0032] In accordance with an embodiment, the reaction delay for the first emotional response may be computed further based on a geographical region of the audience, a type of reaction, a type of scene, and a type of expected response from the audience. The control circuitry may further shift a first timestamp of a first highlight point of the plurality of highlight points to a second timestamp, based on the computed reaction delay and the transmission delay.
[0033] In accordance with an embodiment, the control circuitry may be further configured to predict a set of new highlight points and a set of new lowlight points in the media item. Such prediction may be done, based on a change in control parameters using a simulation engine. The control parameters may include a genre of the media item and a geographical region, a race, an age group, and a gender of the audience. The control circuitry may be further configured to compare a first media item with a second media item using the simulation engine, based on a video score, an audio score, and a distribution score. The video score, the audio score, and the distribution score may be
generated using the simulation engine based on an accuracy of the capture of emotion response data.
[0034] FIG. 1 is a block diagram that illustrates an exemplary network environment for a calibration system for audience response capture and analysis of media content, in accordance with an embodiment of the disclosure. With reference to FIG. 1 , there is shown a network environment 100 that includes a calibration system 102. The calibration system 102 may be configured to process a media item 104. The media item 104 may include a plurality of markers at a plurality of positions within the media item 104. For example, as shown, the media item 104 may include a first marker 106 that spans a first timeslot 106A to 106B and a second marker 108 that spans a second timeslot 108A to 108B. The network environment 100 may further include a media device 1 10. An audience 1 12 may engage with the media item 104 via the media device 1 10. The network environment 100 may further comprise a plurality of different types of sensors 1 14. The plurality of different types of sensors 1 14 may include a set of audio sensors 1 14A, a set of image sensors 1 14B, and a set of biometric sensors 1 14C. There is also shown a test audience 1 16.
[0035] In some embodiments, the media device 1 10, the audience 1 12, the plurality of different types of sensors 1 14, and the test audience 1 16 may be present in a closed media presentation environment (such as an auditorium 1 18, a movie theater, an indoor exhibition, and the like). The media device 1 10 may be a display screen suitably positioned to be viewed by the audience 1 12 in the auditorium 1 18. In some embodiments, the media device 1 10, the audience 1 12, the plurality of different types of sensors 1 14, and the test audience 1 16 may be present in an open media presentation
environment (such as an outdoor concert area or an outdoor exhibition). The media device 1 10 and the plurality of different types of sensors 1 14 may be communicatively coupled to the calibration system 102, through a communication network 120. A set of media items, such as the media item 104 or a test media item, may be stored locally on at least one of the calibration system 102 or the media device 1 10.
[0036] The calibration system 102 may comprise suitable logic, circuitry, and interfaces that may be configured to capture an audience response at playback of the media item 104. The calibration system 102 may be further configured to analyze the audience response to calibrate the plurality of markers in the media item 104 of the set of media items. Alternatively stated, the calibration system 102 may embed one or more markers in the media item 104 and calibrate a position of the plurality of markers in the media item 104 based on an audience response at specific durations of playback of the media item 104. In some embodiments, the calibration system 102 may be implemented as an internet of things (IOT) enabled system that operates locally in the network environment 100 or remotely, through the communication network 120. In some embodiments, the calibration system 102 may be implemented as a standalone device that may be integrated within the media device 1 10. Examples of implementation of the calibration system 102 may include a projector, a smart television (TV), a personal computer, a special-purpose device, a media receiver, such as a set top box (STB), a digital media player, a micro-console, a game console, an (High Definition Multimedia Interface) HDMI compliant source device, a smartphone, a tablet computer, a personal computer, a laptop computer, a media processing system, or a calibration device.
[0037] The media device 1 10 may comprise suitable logic, circuitry, and interfaces that may be configured to execute playback of the set of media items, for example, the media item 104. The media device 1 10 may or may not include a display screen or a projection means, an audio device, and a set of input/output (I/O) devices. The media device 1 10 may be placed in a closed environment such that the playback of the set of media items through the media device 1 10 may lie in a field of audio-visual (FOA-V) reception of the audience 1 12. The media device 1 10 may comprise at least a first speaker to output an audio signal of the media item 104. Examples of the first speaker may include, but is not limited to a wireless speaker or a wired speaker.
[0038] In some embodiments, the media device 1 10 may be implemented as a distributed set of audio devices, display screens, and media playback devices. For example, a media device in a multiscreen theater may have different display screens connected to different speakers placed at different selectively determined locations in the multiscreen theater for a surround sound effect, with different speakers and different screens connected to a single media player device. Examples of the media device 1 10 may include, but are not limited to, a liquid crystal display (LCD) television, a Light Emitting Diode (LED) television, and an LCD projector. Examples of the different audio devices may include, but are not limited to, a tweeter speaker, a subwoofer speaker, a woofer speaker, and a treble speaker. Examples of an I/O device may include, but are not limited to, a joystick, a mouse, a keyboard, and a motion sensor, a gesture sensor, a Kinect-based sensor.
[0039] The plurality of different types of sensors 1 14 may comprise suitable logic, circuitry, and interface that may be configured to capture a plurality of different types of input signals from an audience (e.g., the test audience 1 16 or the audience 1 12), at the playback of a media item (e.g., the media item 104). The captured plurality of different types of input signals may correspond to emotional response data associated with the audience 1 12 at the playback of the media item 104. The plurality of different types of sensors 1 14 may include the set of audio sensors 1 14A, the set of image sensors 1 14B, and the set of biometric sensors 1 14C.
[0040] The set of audio sensors 1 14A may be a set of microphones, which may positioned at two different locations within the closed environment, to capture audio signals from the speaker or the audience 1 12 at a pre-calibration stage and a calibration stage. The set of image sensors 1 14B may be a set of image capturing devices, such as a set of cameras. The set of image sensors 1 14B may be positioned at different locations in the closed or open media screening environment, such as the auditorium 1 18, to capture a number of faces from the audience 1 12 at the playback of the media item 104 by the media device 1 10.
[0041] The set of biometric sensors 1 14C may be a plurality of sensors that may be configured to capture biometric data of the audience 1 12 at the playback of the media item 104 of the set of media items. Examples of a biometric sensor may include but is not limited to, a pulse rate sensor, a breath rate sensor, a body temperature sensor, or a skin conductance sensor, or other specialized sensors to measure different emotions aroused in the audience 1 12 at the playback of the set of media items. In some embodiments, the set of biometric sensors 1 14C may be non-invasively attached to the body of each member of the audience 1 12. In other embodiments, the set of biometric sensors 1 14C may be invasively implanted in the body of each member of the audience 1 12, for
example, an ingestible capsule with nano-biometric sensors. In some implementations, instead of a set of biosensors, a single standalone sensor device may be installed in the closed environment, to collectively detect a heart signature and further monitor the biometric data of all the members of the audience 1 12.
[0042] The communication network 120 may include one or more mediums through which the calibration system 102 may communicate with the media device 1 10 and the plurality of different types of sensors 1 14. Examples of the communication network 120 may include, but are not limited to the Internet, a public switched telephone network (PSTN), a third generation (3G), 4G, or 5G mobile network, a radio communication network, a cloud network, or a Wireless Fidelity (Wi-Fi) network and/or a Bluetooth wireless network
[0043] In operation, a control signal may be received by the calibration system 102 from a user or a server. The control signal may utilized to initialize a multistage operation to pre-calibrate the plurality of different types of sensors 1 14 in the auditorium 1 18 and calibrate a position of the plurality of markers in a set of media items (for example, the media item 104) based on analysis of an audience response. The multistage operation may include a pre-calibration stage, a calibration stage, and a post-calibration stage.
[0044] In the pre-calibration stage, the calibration system 102 may pre-calibrate the plurality of different types of sensors 1 14 in the closed environment for the playback of the set of media items. The pre-calibration stage may further include an identification of noise components in the closed environment, a cancellation of the identified noise components, and an estimation of weight for different sensors of the plurality of different types of sensors 1 14 prior to a playback of the set of media items. The pre-calibration may be done to precisely capture emotional response data from the audience 1 12 at the playback of the media item 104, with a minimal effect of noise signals or measurement errors of different sensors on the captured emotional response data.
[0045] The calibration system 102 may pre-calibrate the set of audio sensors 1 14A. Such pre-calibration of the set of audio sensors 1 14A may be done based on a playback of a test media item by the media device 1 10 (i.e., before the playback of the media item 104 before an actual audience (such as the audience 1 12) for capture of audience response data). The test media item may be first played in a test environment in a presence of the test audience 1 16, and further in an absence of the test audience 1 16. Thereafter, the calibration system 102 may generate a noise signal based on analysis of the different audio signals captured by different audio sensors of the set of audio sensors 1 14A in the test environment. For example, a noise signal that may model a babble noise generated by different members of the audience 1 12 based on the analysis of different audio signals received from the set of audio sensors 1 14A placed in the closed environment. Thereafter, the calibration system 102 may be further configured to pre-calibrate the set of image sensors 1 14B of the plurality of different types of sensors 1 14. The set of image sensors 1 14B may be calibrated based on playback of the test media item by the media device 1 10 in the test environment in presence of the test audience 1 16. An error rate may be determined for each image sensor of the set of image sensors 1 14B. Such error rate may be determined based on a change in a number of faces detected for the at least two image frames of the test media item.
[0046] The calibration system 102 may be further configured to pre-calibrate the set of biometric sensors 1 14C of the plurality of different types of sensors 1 14 before the playback of the media item 104 by the media device 1 10. Such pre-calibration of the set of biometric sensors 1 14C may be done based on a measurement of biometric data in the test environment in presence of the test audience 1 16 at the playback of the test media. A standard deviation in the measured biometric data of the test audience 1 16 may be determined at the playback of the test media. The calibration system 102 may be further configured to assign a weight to each sensor of the plurality of different types of sensors 1 14 based on different estimations done for different types of sensors. For example, the set of audio sensors 1 14A may be weighted based on different levels of the captured noise components. The set of image sensors 1 14B may be weighted based on error rate in detection of number of faces. The set of biometric sensors 1 14C may be weighted based on deviations in the measured biometric data of the test audience 1 16. The detailed operation of pre-calibration has been discussed in detail in FIG. 2. The pre-calibration of the plurality of different types of sensors 1 14 enables obtaining of credible data from the plurality of different types of sensors 1 14.
[0047] In accordance with an embodiment, the calibration system 102 may be configured to recognize a new sensor and dynamically add the recognized sensor in a network of sensors, such as the plurality of different types of sensors 1 14, during the pre-calibration stage. Based on the recognition, the new sensor may be automatically categorized as a specific new type of sensor or one of the plurality of different types of sensors 1 14 without prior knowledge about the new sensor or the new type of sensor. The calibration system 102 may be configured to recognize input signals captured by the new sensor from the audience 1 12 or the test audience 1 16. The calibration system 102 may be further configured to correlate various values of the input signals to various
emotional response from the audience 1 12 or the test audience 1 16. The calibration system 102 may be further configured to identify a noise signal in the various input signals captured by the newly added sensor. The newly added sensor may be further weighted based on the identified noise signal in the various input signals.
[0048] After the pre-calibration stage, the calibration system 102 may be configured to embed a set of markers in the media item 104 which may be played by the media device 1 10. The set of markers may be embedded based on the expected-emotions-tagging metadata stored for the media item 104 in a memory of the calibration system 102. The set of markers may be at least one of audio markers, video markers, or frame markers. As an example, an audio marker may be a high frequency human inaudible beacon or an ultrasonic beacon. The expected-emotions-tagging metadata may include an associative relationship between a set of time slots (or a set of scenes) and a set of specified emotional states that may be expected from the audience 1 12 at playback of the media item 104. The first marker 106 may be embedded at a first timestamp (or at the first timeslot 106A to 106B) at a first scene marked within the media item 104. The first timestamp (or time) may be mapped with a first emotional state, such as a happy state, a sad state, an angry state, and a neutral state, of a set of defined emotional states.
[0049] The calibration system 102 may store the expected-emotions-tagging metadata as a pre-stored data in a database managed by the calibration system 102. In some embodiments, the calibration system 102 may fetch the expected-emotions-tagging metadata from a production media server (or dedicated content delivery networks) and store the expected-emotions-tagging metadata in the memory of the calibration system 102. In certain scenarios, the calibration system 102 may be configured to store the
expected-emotions-tagging metadata in the media item 104. The expected-emotions-tagging metadata may be a data structure with a defined schema, such as a scene index number (I), a start timestamp (To), an end timestamp (T1 ), and an expected emotion (EM). In certain scenarios, the expected-emotions-tagging metadata may comprise an expected emotion type and an expected sub-emotion type. An example of the stored expected-emotions-tagging metadata associated with the media item 104 (for example, a movie with a set of scenes), is given below, for example, in Table 1 .
Table 1 : A portion of expected-emotions-tagging metadata for the media item 104.
[0050] The playback of a media item, such as the media item 104, of the set of media items may be initialized at the media device 1 10. The media device 1 10 may be configured to play the media item 104 in the field of audio visual reception (FOAV) of the audience 1 12. With the initialization of the playback of the media item 104 at the media device 1 10, the plurality of different types of sensors 1 14 may be activated to capture a set of emotional responses continuously at a playback duration of the media item 104.
[0051 ] The calibration system 102 may be configured to receive the plurality of different types of input signals from the plurality of different types of sensors 1 14. The plurality of different types of input signals may include audio signals from the set of audio sensors 1 14A, video signals that includes face data of the audience 1 12 from the set of image sensors 1 14B, and biometric signals from the set of biometric sensors 1 14C. Examples of audio signals may include, but are not limited to, audio data of claps, laughter, chatter, and whistles of the audience 1 12. Example of the face data may include but is not limited to a sequence of facial images (or a video feed) of the audience members. The sequence of facial images of the audience members may comprise specific expressions of the audience members. The face data may further comprise eye-tracking, motion-tracking, and gesture tracking information associated with the audience members. Examples of the biometric signals may include, but are not limited to, pulse rate data, heart rate data, body heat data, and body temperature data of each member of the audience 1 12. Such plurality of different types of input signals may correspond to a set of emotional responses of the audience 1 12 at the playback of the media item 1 04 at the media device 1 10. Thereafter, the calibration system 102 may be configured to generate an amalgamated audience response signal based on the synchronized plurality of different input signals and a plurality of weights assigned to the plurality of different types of sensors 1 14 that captures the plurality of different types of input signals.
[0052] Different members of the audience 1 12 may exhibit a different degree of expressive ability, which may be captured in the plurality of different types of input signals in real time or near-real time from the plurality of different types of sensors 1 14. For example, at a certain comic scene at the playback of the media item 104, different members of the audience 1 12 may exhibit different types of smiles or laughter. In order to robustly extract an aggregate emotional response for the comic scene, a normalization of the plurality of different types of input signals may be done. Therefore, the calibration system 102 may be configured to normalize the plurality of different types of input signals for each scene of the media item 104 based on different parameters of the audience 1 12.
[0053] In certain embodiments, the plurality of different types of input signals for different audiences from different geographical locations, for example, India, US, China, Germany, and Japan, may be analyzed. Users from different geographical locations may express emotions at different scale and such behavior may be captured in the plurality of different type of input signals. For example, an American audience may be more expressive as compared to an Asian audience. A different normalization for the plurality of different types of input signals may be applied for different geographical locations to facilitate analysis of the emotional response data from different geographical locations at the same scale.
[0054] In accordance with an embodiment, the calibration system 102 may be configured to compute a set of root mean square (RMS) values for the plurality of different types of input signals. The set of RMS values may be computed based on a geographical region of the audience 1 12, a type of reaction, a type of scene, and a type of expected response from the audience 1 12. For example, magnitude of the set of RMS values
associated with the audience 1 12 may be greater or lesser based on a geographical location of the audience 1 12. The calibration system 102 may be configured to normalize the plurality of different types of input signals based on the geographical region of the audience 1 12. Alternatively stated, the calibration system 102 may be configured to shift the computed set of RMS values of the plurality of different types of input signals based on the geographical region of the audience 1 12.
[0055] In accordance with an embodiment, the calibration system 102 may be configured to compute a reaction delay for a first emotional response of the audience 1 12 captured by the plurality of different types of sensors 1 14 for a first scene of the media item 104 with respect to a position of the first marker 106. The first marker 106 may span a first time slot of the set of time slots in the media item 104. The first time slot may correspond to the first scene in the media item 104. Similar to computation of the reaction delay for the first scene, the calibration system 102 may be configured to compute a set of reaction delays for a set of emotional responses of the audience 1 12. Such set of reaction delays may be captured by the plurality of different types of sensors 1 14 for a set of scenes in the media item 104 with respect to a position of the set of markers. Each marker of the set of markers may span one time slot of the set of time slots in the media item 104.
[0056] The set of reaction delays may be computed further based on a geographical region of the audience 1 12, a type of reaction, a type of scene, and a type of expected response from the audience 1 12. For example, the audience 1 12 may continue generating the emotional response for the first scene even after playback of the first scene elapses and playback of a new scene begins. In such scenarios, the calibration system
102 may be configured to compute reaction delay and compare the computed reaction delay of the audience 1 12 with a numerical threshold. For example, the numerical threshold may be a time interval between "1 " second and "10" seconds. The numerical threshold may be set in the calibration system 102 based on a set of parameters, such as the geographical location of the audience 1 12, a type of media item, the reaction type, and the expected-emotion-tagging metadata associated with the media item 104. In cases where the first scene is associated with sadness, then the first scene may elicit a prolonged emotion response from the audience 1 12 in comparison with a second scene which may be associated with a joke. The computed reaction delay is advantageous for calibration as it ensures that only the frames in the first scene that are responsible to elicit the sadness response are associated with the emotional state of sadness. Other frames in the first scene (i.e., a certain segment of the media item 104) that were initially associated with sadness or other expected emotional responses are disassociated with the sadness emotion.
[0057] The calibration system 102 may be further configured to compute a transmission delay for the set of emotional responses at the playback of the set of scenes of the media item 104. Such transmission delay may be associated with the capture and transmission of the emotional response data from the plurality of different types of sensors 1 14. The transmission delay may be computed with respect to the position of the plurality of markers that may be embedded in the media item 104. For example, for a sad scene in a movie, a reaction delay of "1 .56 seconds" may be computed with a transmission delay of "0.5 seconds". Thus, an overall delay offset of "1 .56+0.5 seconds", i.e. "2.06 seconds" may be computed for the given sad scene. Thereafter, the calibration system 102 may be
configured to calibrate the position of the each marker of the plurality of markers to associate a set of frames of each scene with an emotional response. Such calibration of the position of each of the set of markers may be done based on at least the computed reaction delay and the computed transmission delay for a corresponding scene of the media item 104.
[0058] In accordance with an embodiment, the normalized plurality of different type of input signals may be utilized to determine a peak emotional response level for each emotional response from each user of the audience 1 12. Thereafter, the normalized plurality of different types of input signals may be synchronized and overlaid in a timeline that may be same as a playback timeline of the media item 104. A set of common positive peaks and a set of common negative peaks in each of the plurality of different types of input signals may be further identified based on the overlay of the plurality of different types of input signals. An example of the set of common positive peaks and the set of common negative peaks is shown in FIG. 6B and FIG. 1 1 . The calibration system 102 may be configured to calculate a plurality of highlight points and a plurality of lowlight points for a plurality of scenes of the media item 104. Such plurality of highlight points or the plurality of lowlight points may be calculated based on the identified set of common positive peaks and the set of common negative peaks. In certain scenarios, the calibration system 102 may be configured to identify the set of common positive points and the set of common negative points without using the expected-emotion-tagging metadata associated with the media item 104. The set of common positive points and the set of common negative points may be identified based on the set of emotional responses of the audience 1 12 at the playback of the set of scenes of the media item 104.
[0059] For example, for a "1 .5 hour long movie", the normalized biometric data, the normalized audio signal, and the normalized expression data is overlaid in a timeline. A set of "6 common positive peaks" and a set of "4 common negative peaks" are further identified in the normalized biometric data, the normalized audio signal, and the normalized expression data. The "6 common positive peaks" may correspond to a comic emotion, a sad emotion, and a contempt emotion of the audience 1 12. Similarly, the set of "4 common negative peaks" may correspond to a neutral expression of the audience 1 12. The set of "4 negative peaks" may be mapped with a marker for a specific expected emotions data. A neutral emotion with respect to a specific expected emotion may be identified as a lowlight point.
[0060] In accordance with an embodiment, the calibration system 102 may be configured to predict a set of new highlight points and a set of new lowlight points in the media item 104. Such prediction of the set of new highlight points and the set of new lowlight points may be done using a simulation engine based on application of different techniques and a change in control parameters. Examples of the control parameters may include, but is not limited to a genre of the media item 104, a geographical region, a race, an age group, and a gender of an audience.
[0061] In the post calibration stage, the calibrated set of markers may be utilized to precisely map the emotions expressed by the audience 1 12 in real time or near-real time. A measure of an impact of different marked scenes may be analyzed based on the set of common peaks in the plurality of different types of input signals. Such impact measure may be utilized to derive a rating for each marked scene of the media item 104 and a cumulative rating for the entire media item. Further, calibration system 102 may be
configured to generate, using the simulation engine, a video score, an audio score, and a distribution score for the media item 104 based on an accuracy of the capture of emotion response data from the audience 1 12. Such scores (such as the video score, the audio score, and the distribution score) may be utilized to compare a first media item (e.g., a first movie), with a second media item (e.g. a second movie) of the set of media items. A final rating or impact of different media items may be identified based on the comparison of scores of different media items with each other.
[0062] FIG. 2 is a block diagram that illustrates a calibration system for audience response capture and analysis of media content, in accordance with an embodiment of the disclosure. FIG. 2 is explained in conjunction with elements from FIG. 1 . With reference to FIG. 2, there is shown the calibration system 102. The calibration system 102 may include a control circuitry 202. The control circuitry 202 may include a processor 204, a transform calculator 206, a high point analyzer 208, a face recognizer 210, and a video frame clipper 212. The calibration system 102 may further comprise a transceiver 214 and a memory 216. The memory 216 may comprise a local database. In accordance with an embodiment, the calibration system 102 may communicate with via the communication network 120, to a cloud database 218, using the transceiver 214.
[0063] The processor 204 may comprise suitable logic, circuitry, and interfaces that may be configured to execute a set of instructions stored in the memory 216. The set of instruction may be associated with execution of different operations, such as high point analysis, transform estimation, data normalization, weight calculation, noise estimation, noise cancellation, and the like. Examples of the processor 204 may include, but are not limited to, an x86-based processor, an x64-based processor, a Reduced Instruction Set Computing (RISC) processor, an Application-Specific Integrated Circuit (ASIC) processor, and a Complex Instruction Set Computing (CISC) processor.
[0064] The transform calculator 206 may comprise suitable logic, circuitry, and interfaces that may be configured to execute a signal transformation of the plurality of different types of signals, to identify different noise components in the plurality of different types of input signals. For example, a babble noise created by the audience 1 12 or high frequency noise created by speakers in the closed environment may be identified. The transform calculator 206 may be implemented based on the processor, such as one of a programmable logic controller (PLC), a microcontroller, an X86-based processor, a Reduced Instruction Set Computing (RISC) processor, an Application-Specific Integrated Circuit (ASIC) processor, a Complex Instruction Set Computing (CISC) processor, and/or other hardware processors.
[0065] The high point analyzer 208 may comprise suitable logic, circuitry, and interfaces that may be configured to identify a set of highlight points and a set of lowlight points in a media item, such as the media item 104. The identification of the set of highlight points and the set of lowlight points may be continuously mapped at the playback duration of the media item 104. The high point analyzer 208 may be implemented as a software application or a hardware circuit, such as an Application-Specific Integrated Circuit (ASIC) processor.
[0066] The face recognizer 210 may comprise suitable logic, circuitry, and interfaces that may be configured to detect a change in facial expression of different users in the audience 1 12 by the set of image sensors 1 14B of the plurality of different types of sensors 1 14. The face recognizer 210 may be further configured to detect static facial
expressions of different users in the audience 1 12 by the set of image sensors 1 14B of the plurality of different types of sensors 1 14. The face recognizer 210 may be implemented as a software application or a hardware circuit, such as an Application-Specific Integrated Circuit (ASIC) processor.
[0067] The video frame clipper 212 may comprise suitable logic, circuitry, and interfaces that may be configured to generate a plurality of highpoint media segments based on the identified set of highlight points and the set of lowlight points. The video frame clipper 212 may be implemented based on the processor 204, such as one of a programmable logic controller (PLC), a microcontroller, an X86-based processor, a Reduced Instruction Set Computing (RISC) processor, an Application-Specific Integrated Circuit (ASIC) processor, a Complex Instruction Set Computing (CISC) processor, and/or other hardware processors.
[0068] The transceiver 214 may comprise suitable logic, circuitry, and interfaces that may be configured to communicate with other electronic devices, via the communication network 120. The transceiver 214 may implement known technologies to support wireless communication. The transceiver 214 may include, but are not limited to an antenna, a radio frequency (RF) transceiver, one or more amplifiers, a tuner, one or more oscillators, a digital signal processor, a coder-decoder (CODEC) chipset, a subscriber identity module (SIM) card, and/or a local buffer circuitry. The transceiver 214 may communicate via offline and online wireless communication with networks, such as the Internet, an Intranet, and/or a wireless network, such as a cellular telephone network, a wireless local area network (WLAN), personal area network, and/or a metropolitan area network (MAN). The wireless communication may use any of a plurality of communication standards,
protocols and technologies, such as Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), wideband code division multiple access (W-CDMA), code division multiple access (CDMA), LTE, time division multiple access (TDMA), Bluetooth, Wireless Fidelity (Wi-Fi) (such as IEEE 802.1 1 , IEEE 802.1 1 b, IEEE 802.1 1 g, IEEE 802.1 1 η, and/or any other IEEE 802.1 1 protocol), voice over Internet Protocol (VoIP), Wi-MAX, Internet-of-Things (loT) technology, Machine-Type-Communication (MTC) technology, a protocol for email, instant messaging, and/or Short Message Service (SMS).
[0069] The memory 216 may comprise suitable logic, circuitry, and interfaces that may be configured to store a set of instructions executable by the processor 204, the transform calculator 206, the high point analyzer 208, the face recognizer 210, and the video frame clipper 212. The memory 216 may be configured to further store the set of media items, such as the media item 104 and associated expected-emotions-tagging metadata. Examples of implementation of the memory 216 may include, but are not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Hard Disk Drive (HDD), and/or a flash drive.
[0070] The cloud database 218 may comprise suitable logic, circuitry, and/or interfaces that may be configured to store expected-emotions-tagging metadata that may include an associative relationship between a set of time slots and a set of specified emotional states that are expected from the audience 1 12 in the set of time slots during playback of a media item, for example, the media item 104. The cloud database 218 may be configured to manage a curated repository of a set of media items, which may be retrieved for playback at the media device 1 10.
[0071] In operation, control signals may be received at the calibration system 102 to initialize acquisition of emotional response data based on different sensor-based data and audience 1 12-based data. Such emotional response data may be utilized to precisely update durations of different high points/low points in a media item (for example, the media item 104), which may be scheduled for selected/scheduled for playback at the media device 1 10. Such control signals may be routed to a processor 204, such as the processor 204 of the calibration system 102. Thereafter, the processor 204 may be configured to initially pre-calibrate the plurality of different types of sensors 1 14 prior to the playback of the media item (such as the media item 104), for which emotional response data is to be captured from the audience 1 12 The pre-calibration of the set of audio sensors 1 14A, the set of image sensors 1 14B and the set of biometric sensors 1 14C of the plurality of different types of sensors 1 14 may ensure that credible data from the plurality of different types of sensors 1 14 is captured at next processing stage.
Audio Sensor Pre-Calibration:
[0072] In accordance with an embodiment, the set of audio sensors 1 14A may include two audio sensors, such as a first audio sensor and a second audio sensor. The first audio sensor may be placed near the audio output source (such as the first speaker) of the media device 1 10. The first audio sensor may capture a plurality of audio signals when placed near the audio output source. The plurality of audio signals may comprise an audio signal of the media item 104. In addition to the audio signal of the media item 104, the plurality of audio signals may further comprise other noise signals in the background, for example, environmental noises, such as waves, traffic noise, alarms, babble, bio acoustic noise from animals, birds, and insects, and mechanical noise from devices, such as air
conditioners, refrigerators, power supplies, electric motors, and the like. A second audio sensor may be further placed at a position such that a maximum of audio response can be captured from either of the test audience 1 16 or the actual audience (e.g., at a location where the test audience 1 16 is supposed to be seated). The second audio sensor may also capture noise signals in the background for example, environmental noises, such as waves, traffic noise, alarms, babble, bio acoustic noise from animals, birds, and insects, and mechanical noise from devices, such as air conditioners, refrigerators, power supplies, electric motors, and the like.
[0073] First, a test media item may be recorded in a test environment in absence of the test audience 1 16, such as an empty auditorium. Such recorded test media item may be a first audio signal. Thereafter, the processor 204 may be configured to play the test media item in the test environment in presence of the test audience 1 16. The set of audio sensors 1 14A may be configured to capture a second audio signal from the test environment at playback of the test media item. The second audio signal may comprise sounds, such as claps, whistles, laughs, chatter, and the like, from the test audience 1 16, and also sound of the audio output related to the test media item. The sounds, such as claps, whistles, laughs, chatter, and the like, from the test audience 1 16 may be the emotional response data of the test audience 1 16, which is of interest. The sound of the audio output related to the test media item may be considered a noise signal, which is removed using the first audio signal recorded in the absence of the test audience 1 16, such as an empty auditorium, for reference. The transform calculator 206 may be configured to determine a difference between the second audio signal and the first audio signal.
[0074] In accordance with an embodiment, the transform calculator 206 may be configured to determine the difference based on one of an accord audio framework method, a fast Fourier transformation method, and an independent component analysis method. The transform calculator 206 may be configured to generate an audience noise signal based on the difference between the first audio signal and the second audio signal, which may be estimated based on transform methods, such as Short Time Fourier Transform (STFT). The audience noise signal may include audience noise without any original sound of the audio output related to the test media item. The transform calculator 206 may be further configured to subtract the second audio signal with the audience noise signal to re-generate a final audio signal that may resemble the first audio signal captured in the empty auditorium. Thereafter, the final audio signal and the first audio signal may be compared to estimate a root mean square error (RMS) using known mathematical techniques. The processor 204 may be configured to determine an error rate of each of the set of audio sensors 1 14A based on the comparison.
[0075] In accordance with an embodiment, the processor 204 may be further configured to generate a specific weight for each of the set of audio sensors 1 14A based on the determined error rate associated with the respective audio sensor of the set of audio sensors 1 14A. The processor 204 may be configured to assign a plurality of weights to the set of audio sensors 1 14A. The processor 204 may be configured to play the test media item in the test environment, and execute multiple iterations of noise cancellation to calibrate the set of audio sensors 1 14A based on the assigned weights (or coefficients).
Image Sensor pre-calibration:
[0076] In accordance with an embodiment, the processor 204 may be configured to play the media item 104 in the test environment (such as the auditorium 1 18) using the media device 1 10. The test environment may comprise the test audience 1 16. In some embodiments, the processor 204 may be configured to detect a first set of faces from the test audience 1 16 by each image sensor when a first image frame of the test media item is played by the media device 1 10. The set of image sensors 1 14B may be further configured detect a second set of faces by each image sensor from the test audience 1 16 when a second image frame of the test media item is played by the media device 1 10. Alternatively stated, the processor 204 may be further configured to detect and track a set of faces from the test audience 1 16 by each image sensor of the set of image sensors 1 14B, for a defined playback time period of the test media item. Each face may be assigned a unique identifier (or a UID). Such UIDs may be utilized to identify new faces or find faces missed to be detected in subsequent image frames by a single image sensor. In some embodiments, the set of faces in successive frames may be detected at a cloud server, via an application programming interface (API) for face detection.
[0077] The processor 204 may be configured to determine an error rate for each image sensor of the set of image sensors 1 14B based on a difference in number of faces in the first set of faces and the second set of faces. The processor 204 may be configured to assign a specific weight for each of the set of image sensors 1 14B based on the determined error rate for the respective image sensor for the pre-calibration of the set of image sensors 1 14B.
[0078] For example, for a set of "5 frames" captured a first image sensor of the set of image sensors 1 14B may be stored in an array, as given below, for example, in Table 2.
Table 2
A deviation in a detected number of faces may be computed between every adjacent pair of array element, such as a deviation in the detected number of faces between "Frame 1 " and "Frame 2". A rate of change of detected number of faces may be further computed, as given below, for example, in Table 3.
The processor 204 may calculate an error rate (or error percentage) for the first image sensor based on addition of computed rate of change of the detected number of faces. If the standard error for the first image sensor is "3 faces" and the computed error is "1 + 1 + 2 + 1 faces", i.e. "5 faces", then the error rate for the first image sensor may be "5-3/5", i.e. "0.4" (or "40%"). The weight value for the first image sensor may be computed as "1 -0.4", i.e. "0.6".
[0079] Biometric Sensor Pre-Calibration:
In accordance with an embodiment, the processor 204 may be further configured to play the test media item in the test environment (such as the auditorium 1 18) using the media device 1 10. The test environment may include a test audience (such as the test audience 1 16). In certain scenarios, the set of biometric sensors 1 14C may be configured to capture a plurality of pulse rates from the test audience 1 16. For example, each of the set of biometric sensors 1 14C may be configured to capture a pulse rate of a different user of the test audience 1 16. The processor 204 may be configured to determine a standard deviation of the captured plurality of pulse rates of the test audience 1 16. The processor 204 may be configured to determine a difference of a pulse rate captured by each biometric sensor of the set of biometric sensors 1 14C with respect to the determined standard deviation of the plurality of pulse rates captured by the set of biometric sensors 1 14C.
[0080] The processor 204 may be configured to determine an error rate of each biometric sensor of the set of biometric sensors 1 14C based on the determined difference of the pulse rate captured by the respective biometric sensor with respect to the determined standard deviation. The processor 204 may be configured to assign a specific weight for each of the set of biometric sensors 1 14C based on the determined error rate for the respective biometric sensor for the pre-calibration of the set of biometric sensors 1 14C. The processor 204 may be configured to execute the pre-calibration of the plurality of different types of sensors 1 14 to assign a plurality of weights to the plurality of different types of sensors 1 14. Alternatively stated, the processor 204 may be configured to assign the plurality of weights to the plurality of different types of input signals captured by the plurality of different types of sensors 1 14.
[0081 ] After the pre-calibration, the processor 204 may be configured to embed a set of markers in the media item 104 which may be played by the media device 1 10. The set of markers may be embedded based on the expected-emotions-tagging metadata stored for the media item 104 in a memory of the calibration system 102 (as described in
FIG. 1 A). The set of markers may be at least one of audio markers, video markers, or frame markers. Thereafter, the media device 1 10 may play a scheduled/user-selected media item (for example, the media item 104) in the closed environment occupied by a target audience (or the audience 1 12), for which the emotional response data is intended to be captured in real time.
Data Acquisition at Media Playback:
[0082] At the media playback stage, for different granular durations of the playback (for example, 60 seconds) of the media item 104, the processor 204 may be configured to receive the plurality of different types of input signals from the plurality of different types of sensors 1 14. The plurality of different types of input signals may include audio signals from the set of audio sensors 1 14A, video signals that includes face data of the audience 1 12 from the set of image sensors 1 14B, and biometric signals from the set of biometric sensors 1 14C. Such plurality of different types of input signals may be stored in the memory 216 as an array of a specific size in accordance with a specific duration of the playback, for example, a signal array with "60" values may include values for "60 seconds" of input signal, where "60" denotes a maximum index value of the signal array that stores a value for a type of input signal at every second at an index position within the signal array.
[0083] In accordance with an embodiment, the plurality of different types of input signals may be processed for emotional response estimation from each type of input signal followed by an estimation of a reaction delay of the audience 1 12 for a specific scene of the media item 104. In accordance with another embodiment, the plurality of different types of input signals may be processed for an emotional response estimation
from each type of input signal followed by an estimation of periods of highlight points and low light points in the media item 104.
Emotional Response Estimation:
[0084] The processor 204 may be configured normalize different types of input signals of different users in the audience 1 12, captured by the plurality of different types of sensors 1 14. Such normalization may be done based on a degree of expressive ability of different users of the audience 1 12. The processor 204 may be further configured to detect an emotional response from each user in the audience 1 12 based on the plurality of input video signals. Examples of the emotional response may include, but are not limited to, happy, sad, surprise, anger, fear, disgust, contempt, and neutral emotional response. The processor 204 may be further configured to measure, using the face recognizer 210, an emotional response level of the detected emotional response of different users of the audience 1 12 in the first time interval. The emotional response level of each user of the audience 1 12 may indicate a degree by which a user in the audience 1 12 may express an emotional response in the first time interval.
[0085] For example, the processor 204 may be configured to measure, using the face recognizer 210, the emotional response level of a smile as a response from a first user of the audience 1 12. The processor 204 may be further configured to assign, using the face recognizer 210, the emotional response level on a scale from "0" to "1 " for the smile response from the first user. Similarly, in some instances, when a user shows a grin face, the processor 204 may assign, using the face recognizer 210, an average value for the emotional response level, i.e. a value of "0.5" for the user that shows a grin face. In other instances, when another user laughs, the processor 204 may assign, using the face
recognizer 210, a maximum value of "1 " for the user that laughs. The face recognizer 210 may be configured to determine the emotional response level of each of the audience 1 12 at a plurality of time instants within the first time interval.
[0086] In accordance with an embodiment, the processor 204 may be further configured to determine, using the face recognizer 210, a peak emotional response level associated with each user in the audience 1 12 at the first time interval. The peak emotional response level of each user may correspond to a time duration within which that user in the audience 1 12 provided a maximal emotional response. Alternatively stated, the peak emotional response of each user is the maximum emotional response level that may be measured for the respective user in the first time interval. The processor 204 may be configured to normalize the emotional response level of each user in the audience 1 12 with respect to emotional response levels of other users in the audience 1 12. The processor 204 may be configured to normalize the emotional response level of each user of the audience 1 12 based on the peak emotional response level associated with the respective user.
[0087] In accordance with an embodiment, the set of biometric sensors 1 14C of the plurality of different types of sensors 1 14 may be configured to capture the plurality of pulse rate signals of each user of the audience 1 12. Each biometric sensor of the set of biometric sensors 1 14C may be configured to detect a pulse rate of each user of the audience 1 12 at different time instances of the first time interval at playback of the media item 104. The processor 204 may be configured to determine a peak pulse rate associated with each user of the audience 1 12. The peak pulse rate associated with each user may correspond to a maximum pulse rate generated by the respective user at the
first time interval. The processor 204 may be configured to normalize the detected pulse rate of each user of the audience 1 12 in the first time interval, based on the determined peak pulse rate associated with the respective user.
Emotional Response Estimation for Different Geographies:
[0088] In some embodiments, the emotional response data may be obtained for different geographical regions that may be associated with different demographic factors. Such emotional response data may be obtained from the plurality of different types of input signals recorded from different screening environments (for example, auditorium 1 18 or theatres) spread across different geographical regions.
[0089] A centralized server (such as the calibration system 102) may collect the plurality of different types of input signals as audience response data from different geographical locations. In such data acquisition scenario, the audience 1 12 for a specific geographical location may be marked prior to the playback of the media item 104, with an ethnicity for the geographical location, such as an Indian audience, a Japanese audience, an American audience, and a European audience. Thereafter, each type of input signal from the collected plurality of different types of input signals may be utilized to derive an average value for a specific emotional response data from a specific audience of the geographical locations, through a computation technique, for example, machine learning-based estimation, k-Nearest-Neighbors (k-NN) technique, and decision trees. Such computational technique may facilitate derivation of weight values for different geographical regions and for each emotional response dynamically. Such derived average values may be further utilized to calculate emotional response levels for audiences from different geographical regions, engaged with a specific genre of the
media item 104. In some instances, the derived average values may be stored in an array, which may be utilized to compute weight values for a specific emotional response from different geographical locations.
[0090] In some embodiment, the emotional response data for the audience 1 12 from one geographical location may be utilized to simulate (or generate predicted or expected emotional response data) for other audiences from different geographical locations. Such simulated estimation of the emotional response data for different geographical regions may be done based on previously estimated emotional response data for previously screened media items.
[0091] For example, an emotional response data from an American audience for a specific scene of a media item (A) may be computed and stored within an array of "size 8" as given below, for example, in Table 4.
Table 4
Where a sum of different computed values for different emotional response for a given scene in the media item (A) is equal to 1 . Similarly, for a different media item (B), a previously stored array may include emotional response data (for "Happy" emotion) of audiences from different geographical locations, as given below, for example, in Table 5.
Geographical Region Emotional Response Emotional Response
Media item (A) Media item (B)
USA 0.6 0.7
India Not Available 0.5
Japan Not Available 0.6
Table 5
[0092] The processor 204 may be configured to generate emotional response data for
Indian audience and the Japanese audience for the media item (A) based on the
emotional response data for the media item (B). Such generation of the emotional
response data may be a predicted emotional response data prior to the playback of the
media item (A) for audiences in India and Japan who have not yet engaged with the media
item (A). If the emotional response value for media item (A) may be represented as EMA,
and the emotional response value for media item (B) may be represented as EMB, then
the emotional response value from Indian audience for the media item (A) may be
estimated as:
EM B (India)
BY (India) IN = A (America) ^ x ^ ~
A J A J EMB (America)
0.5
BY (India) = x- = 0.6 12:42
A 0.7
The value "0.42" is a predicted emotional response value for a "Happy" emotional
response from Indian audience for the media item (A).
Reaction Delay Estimation:
[0093] The processor 204 may be further configured to estimate a reaction delay for a first emotional response of the audience 1 12 captured by the plurality of different types of sensors 1 14 for a first scene of the media item 104. The first emotional response may be estimated with respect to a position of the first marker 106 that spans a first time slot that may correspond to the first scene in the media item 104. Similarly, the processor 204 may be configured to estimate a plurality of reaction delays for different types of emotional responses from the audience 1 12 for different marked scenes in the media item 104.
[0094] In some embodiments, the reaction delay may be estimated based on statistical analysis of the continuous reaction data (i.e. different durations of reaction from different users in the audience 1 12), which may be associated a specific reaction pattern for the audience 1 12. For example, a duration of emotional response from a first user (A) in the audience 1 12 may be greater than "1 second" or "5 seconds" and a duration of emotional response from a second user (B) in the audience 1 12 may be greater "1 minute". Therefore, a threshold duration of an emotional response may be set based on analysis of a normalized reaction from the audience 1 12, for example, threshold duration between "1 second" and "10 seconds". The calibration system 102 may continuously track a duration of emotional response from the different members of the audience 1 12 post the playback of the scene in the media item 104 and may further account for the extended duration of emotional response based on the threshold duration. Such threshold duration for emotion response may further vary in accordance with different demographic factors (for example, race, gender, age). For example, a sad scene may arouse a prolonged sad emotional response from the audience 1 12 as compared to a happy scene that may
arouse a short emotional response from the audience 1 12. Such threshold duration for emotion response may vary in accordance with a geographical location of the audience 1 12. For example, in cases where the audience 1 12 is from a first geographical area, the sad scene may arouse a prolonged sad emotional response from the audience 1 12. In other cases where the audience 1 12 is from a second geographical area, the sad scene may arouse a short sad emotional response from the audience 1 12.
[0095] The processor 204 may be further configured to synchronize and overlay the normalized plurality of different input signals in a timeline that may be same as the playback timeline of the media item 104. An amalgamated audience response signal may be further generated based on the synchronized plurality of different input signals and a plurality of weight values assigned to the plurality of different types of sensors 1 14 that may capture the plurality of different types of input signals. A plot of the amalgamated audience response signal may include a plurality of tall peaks, with each tall peak having internal shorter peaks, which may form a Gaussian surface along the tall peak. The processor 204 may further detect and remove such internal shorter peaks to obtain a uniform Gaussian surface that may span for at least the threshold duration of the emotional response from the audience 1 12.
[0096] The processor 204 may further compute the reaction delay based on a difference between a marked duration for an expected emotional response from the audience 1 12 and a duration measured from the peaks observed in the plot of amalgamated audience response signal. For example, if the expected timestamp for an emotional response as per an inserted marker may span from "00:55:22" to "00:55:40" in the media item 104 and the measured peak in the amalgamated emotional response
signal may span from "00:55:33" to "00:56:00", then processor 204 may compute the reaction delay as "00:55:33" - "00:55:22", i.e. "1 1 seconds". Additionally, the processor 204 may further compute the reaction delay for the emotional response based on a geographical region of the audience 1 12, a type of reaction, a type of scene, and a type of expected emotional response from the audience 1 12.
[0097] In practical situations, however, the reaction delay for each emotional response from the audience 1 12 may be further attributed to a set of constraints associated the plurality of different types of sensors 1 14 that captures the plurality of different types of input signals from the audience 1 12. The set of constraints may include, but are not limited to, location of audio sensors and associated phase shifts of received input signal, a sensor sensitivity to the input signal, and attenuation of the input signal in the ambient environment. In one such instance, the processor 204 may further compute a transmission delay associated with the capture and a transmission of the emotional response data from the plurality of different types of sensors 1 14 for the emotional response during the playback of different scenes of the media item 104. Such transmission delay may be computed with respect to a position of an associated marker of the plurality of markers that may be embedded in the media item 104. For example, a first microphone (M1 ) may be positioned near the audience 1 12 seats and a second microphone (M2) may be positioned near the first speaker of the media device 1 10. As the first microphone (M1 ) is close to the source of the input signal with respect to the second microphone (M2), thus the second microphone (M2) may receive the same input signal after a transmission delay of "0.3 milliseconds" as compared to the first microphone (M1 ).
Marker Calibration in Media item:
[0098] In general, the computed reaction delay and the transmission delay provide a measure of deviation of duration or timestamps for an actual emotional response from an expected duration or timestamps of the emotional response from the audience 1 12. Therefore, such deviation may be compensated for the media item 104 with a calibration of position of the plurality of markers by a delay equivalent to a sum of the computed reaction delay from the audience 1 12 and the computed transmission delay associated with the capture and reception of the plurality of different types of input signals at the plurality of different types of sensors 1 14. Thus, the processor 204 may be configured to calibrate the position of each marker to associate a first set of frames of a scene with an emotional response, based on at least the computed reaction delay and the computed transmission delay. The processor 204 may be further configured to perform a plurality of iterations to calibrate the plurality of markers inserted in the media item 104, in accordance with a plurality of different groups of audiences. Each iteration of the plurality of iterations may be executed by the processor 204 for a different group of the plurality of different groups of audiences.
[0099] In some embodiments, the calibration of marker may be done based on computation of the reaction delay for each type of input signal, which may be different for different types of sensors. For example, an increase in a pulse rate may occur in a relatively shorter time interval (for example, "2 seconds" after the scene is screened) with respect to a timestamp for a marker as compared to a duration for generation and capture of audio signal (for example, "5 seconds" after the scene is screened) for the same scene of the media item 104. Therefore, a common peak of the pulse rate and the audio signal
may be identified to obtain a precise value of reaction delay, which may be "4 seconds" based on the reaction delay for the pulse rate and the audio signal. In some embodiments, the processor 204 may be further configured to output a portion of the media item 104 that spans a duration of a calibrated marker in the media item 104. Such outputted portion of the media item 104 may be utilized to analyze an effectiveness of an intended scene of the media item 104 for audiences that may be associated with different geographical locations and different demographics.
Highlight Point and Lowliqht Point Analysis:
[0100] With numerous iterations for media calibration, the durations and timestamps for the expected emotional response may be obtained precisely with reference to precise placement of the plurality of markers in the media item 104. Such calibration may further be utilized to identify a set of highlight points and a set of lowlight points in the media item 104, using the high point analyzer 208. The set of highlight points and the set of lowlight points in the media item 104 may be identified based on analysis of processed plurality of different types of input signals that may be collected over a period.
WE CLAIMS
What is claimed is:
1 . A calibration system for media content, the calibration system comprising :
a memory configured to store a media item and expected-emotions-tagging metadata for the media item, wherein the expected-emotions-tagging metadata indicates an associative relationship between a set of time slots and a set of specified emotional states that are expected from an audience at the set of time slots at playback of the media item;
a media device configured to execute the playback of the media item to the audience;
a plurality of different types of sensors configured to capture emotional response data from the audience at the playback of the media item; and
control circuitry configured to:
embed a first marker in the media item played by the media device based on the expected-emotions-tagging metadata;
compute a reaction delay for a first emotional response of the audience captured by the plurality of different types of sensors for a first scene of the media item with respect to a position of the first marker that spans a first time slot that corresponds to the first scene in the media item; and
calibrate the position of the first marker to associate a first set of frames of the first scene with the first emotional response, based on at least the computed reaction delay.
2. The calibration system of claim 1 , wherein the plurality of different types of sensors includes a set of audio sensors, a set of image sensors, and a set of biometric sensors, and wherein the control circuitry is further configured to assign a weight to each sensor of the plurality of different types of sensors prior to the playback of the media item by the media device.
3. The calibration system of claim 2, wherein the control circuitry is further configured to pre-calibrate the plurality of different types of sensors before the playback of the media item by the media device, for capture of the emotional response data from the audience at the playback of the media item.
4. The calibration system of claim 2, wherein the control circuitry is further configured to pre-calibrate the set of audio sensors before the playback of the media item by the media device, based on playback of a test media item by the media device in a test environment in a presence of a test audience and an absence of the test audience.
5. The calibration system of claim 4, wherein the control circuitry is further configured to generate a noise signal based on at least a difference between a first audio signal captured by the set of audio sensors in the test environment in the presence of the test audience and a second audio signal captured by the set of audio sensors in the test environment in the absence of the test audience, and wherein the control circuitry is configured to assign a weight to each of the set of audio sensors for the pre-calibration of the set of audio sensors, based on the generated noise signal.
6. The calibration system of claim 2, wherein the control circuitry is further configured to pre-calibrate the set of image sensors of the plurality of different types of sensors before the playback of the media item by the media device, based on playback of a test media item by the media device in a test environment in presence of a test audience.
7. The calibration system of claim 6, wherein the control circuitry is further configured to compute a number of faces that are missed to be detected between at least two image frames at playback of the test media item.
8. The calibration system of claim 7, wherein the control circuitry is further configured to determine an error rate for each image sensor of the set of image sensors, based on a change in the number of faces detected for the at least two image frames of the test media item, wherein the control circuitry is configured to assign a weight to each of the set of image sensors for the pre-calibration of the set of image sensors, based on the determined error rate.
9. The calibration system of claim 2, wherein the control circuitry is further configured to pre-calibrate the set of biometric sensors of the plurality of different types of sensors before the playback of the media item by the media device, based on a
measurement of biometric data in a test environment in presence of a test audience at playback of a test media.
10. The calibration system of claim 9, wherein the control circuitry is further configured to determine a standard deviation in the measured biometric data of the test audience at the playback of the test media, and wherein the control circuitry is further configured to assign a weight to each biometric sensor of the set of biometric sensors for the pre-calibration of the set of biometric sensors, based on a difference in the measured biometric data of each user of the test audience with respect to the determined standard deviation.
1 1 . The calibration system of claim 1 , wherein the control circuitry is further configured to receive a plurality of different types of input signals from the plurality of different types of sensors for the audience at the playback of the media item, wherein the plurality of different types of input signals corresponds to the emotional response data of the audience.
12. The calibration system of claim 1 1 , wherein the control circuitry is further configured to determine a peak emotional response level for each emotional response for each user of the audience based on the received plurality of different types of input signals.
13. The calibration system of claim 12, wherein the control circuitry is further configured to normalize the received plurality of different types of input signals based on the determined peak emotional response level for each emotional response for each user at the playback of the media item, and wherein the received plurality of different types of input signals are normalized further based on a geographical region of the audience.
14. The calibration system of claim 13, wherein the control circuitry is further configured to synchronize and overlay the normalized plurality of different types of input signals in a timeline that is same as a playback timeline of the media item.
15. The calibration system of claim 14, wherein the control circuitry is further configured to generate an amalgamated audience response signal based on the synchronized plurality of different types of input signals and a plurality of weights assigned to the plurality of different types of sensors that captures the plurality of different types of input signals.
16. The calibration system of claim 14, wherein the control circuitry is further configured to identify a set of common positive peaks and a set of common negative peaks in each of the plurality of different types of input signals based on the overlay of the plurality of different types of input signals.
17. The calibration system of claim 16, wherein the control circuitry is further configured to calculate a plurality of highlight points and a plurality of lowlight points for a plurality of scenes of the media item, based on the identified set of common positive peaks and the set of common negative peaks.
18. The calibration system of claim 17, wherein the reaction delay for the first emotional response is computed further based on a geographical region of the audience, a type of reaction, a type of scene, and a type of expected response from the audience.
19. The calibration system of claim 17, wherein the control circuitry is further configured to compute a transmission delay associated with the capture and a transmission of the emotional response data from the plurality of different types of sensors for the first emotional response at the playback of the first scene of the media item, and wherein the transmission delay is computed with respect to the position of the first marker of a plurality of markers that are embedded in the media item.
20. The calibration system of claim 19, wherein the control circuitry is further configured to calibrate the position of the first marker to associate the first set of frames of the first scene with the first emotional response, further based on the computed transmission delay.
21 . The calibration system of claim 19, wherein the control circuitry is further configured to shift a first timestamp of a first highlight point of the plurality of highlight points to a second timestamp, based on the computed reaction delay and the transmission delay.
22. The calibration system of claim 19, wherein the control circuitry is further configured to shift a first timestamp of a first lowlight point of the plurality of lowlight points to a second timestamp, based on the computed reaction delay and the transmission delay.
23. The calibration system of claim 19, wherein the control circuitry is further configured to predict a set of new highlight points and a set of new lowlight points in the media item, based on a change in control parameters using a simulation engine, and wherein the control parameters include a genre of the media item and a geographical region, a race, an age group, and a gender of the audience.
24. The calibration system of claim 19, wherein the control circuitry is further configured to compare a first media item a second media item using the simulation engine, based on a video score, an audio score, and a distribution score, and wherein the video score, the audio score, and the distribution score are generated using the simulation engine based on the capture of emotion response data, and further based on a plurality of accuracy scores of the plurality of different types of sensors.
25. A calibration method for media content, comprising:
in a calibration system comprising a memory, a media device, a plurality of different types of sensors, and control circuitry:
embedding, by the control circuitry, a first marker in a media item played by the media device based on expected-emotions-tagging metadata, wherein the expected-emotions-tagging metadata stored in the memory indicates an associative relationship between a set of time slots and a set of specified emotional states that are expected from an audience at the set of time slots at playback of the media item;
computing, by the control circuitry, a reaction delay for a first emotional response of the audience captured by the plurality of different types of sensors for a first scene of the media item with respect to a position of the first marker that spans a first time slot that corresponds to the first scene in the media item; and
calibrating, by the control circuitry, the position of the first marker to associate a first set of frames of the first scene with the first emotional response, based on at least the computed reaction delay.
| # | Name | Date |
|---|---|---|
| 1 | 202017006103-STATEMENT OF UNDERTAKING (FORM 3) [12-02-2020(online)].pdf | 2020-02-12 |
| 2 | 202017006103-REQUEST FOR EXAMINATION (FORM-18) [12-02-2020(online)].pdf | 2020-02-12 |
| 3 | 202017006103-PROOF OF RIGHT [12-02-2020(online)].pdf | 2020-02-12 |
| 4 | 202017006103-PRIORITY DOCUMENTS [12-02-2020(online)].pdf | 2020-02-12 |
| 5 | 202017006103-FORM 18 [12-02-2020(online)].pdf | 2020-02-12 |
| 6 | 202017006103-FORM 1 [12-02-2020(online)].pdf | 2020-02-12 |
| 7 | 202017006103-DRAWINGS [12-02-2020(online)].pdf | 2020-02-12 |
| 8 | 202017006103-DECLARATION OF INVENTORSHIP (FORM 5) [12-02-2020(online)].pdf | 2020-02-12 |
| 9 | 202017006103-COMPLETE SPECIFICATION [12-02-2020(online)].pdf | 2020-02-12 |
| 10 | 202017006103-FORM-26 [18-02-2020(online)].pdf | 2020-02-18 |
| 11 | 202017006103-MARKED COPIES OF AMENDEMENTS [21-02-2020(online)].pdf | 2020-02-21 |
| 12 | 202017006103-FORM 13 [21-02-2020(online)].pdf | 2020-02-21 |
| 13 | 202017006103-AMMENDED DOCUMENTS [21-02-2020(online)].pdf | 2020-02-21 |
| 14 | 202017006103-Information under section 8(2) [05-08-2020(online)].pdf | 2020-08-05 |
| 15 | 202017006103-FORM 3 [05-08-2020(online)].pdf | 2020-08-05 |
| 16 | 202017006103-Information under section 8(2) [24-09-2021(online)].pdf | 2021-09-24 |
| 17 | 202017006103-FORM 3 [24-09-2021(online)].pdf | 2021-09-24 |
| 18 | 202017006103-OTHERS [30-09-2021(online)].pdf | 2021-09-30 |
| 19 | 202017006103-FER_SER_REPLY [30-09-2021(online)].pdf | 2021-09-30 |
| 20 | 202017006103-DRAWING [30-09-2021(online)].pdf | 2021-09-30 |
| 21 | 202017006103-CORRESPONDENCE [30-09-2021(online)].pdf | 2021-09-30 |
| 22 | 202017006103-CLAIMS [30-09-2021(online)].pdf | 2021-09-30 |
| 23 | abstract.jpg | 2021-10-19 |
| 24 | 202017006103.pdf | 2021-10-19 |
| 25 | 202017006103-Power of Attorney-250220.pdf | 2021-10-19 |
| 26 | 202017006103-OTHERS-180220.pdf | 2021-10-19 |
| 27 | 202017006103-OTHERS-180220-.pdf | 2021-10-19 |
| 28 | 202017006103-FER.pdf | 2021-10-19 |
| 29 | 202017006103-Correspondence-250220.pdf | 2021-10-19 |
| 30 | 202017006103-Correspondence-180220.pdf | 2021-10-19 |
| 31 | 202017006103-RELEVANT DOCUMENTS [21-04-2022(online)].pdf | 2022-04-21 |
| 32 | 202017006103-POA [21-04-2022(online)].pdf | 2022-04-21 |
| 33 | 202017006103-FORM 13 [21-04-2022(online)].pdf | 2022-04-21 |
| 34 | 202017006103-US(14)-HearingNotice-(HearingDate-07-06-2023).pdf | 2023-04-23 |
| 35 | 202017006103-FORM-26 [05-06-2023(online)].pdf | 2023-06-05 |
| 36 | 202017006103-Correspondence to notify the Controller [05-06-2023(online)].pdf | 2023-06-05 |
| 37 | 202017006103-Written submissions and relevant documents [23-06-2023(online)].pdf | 2023-06-23 |
| 38 | 202017006103-PatentCertificate11-08-2023.pdf | 2023-08-11 |
| 39 | 202017006103-IntimationOfGrant11-08-2023.pdf | 2023-08-11 |
| 1 | searchstrategyE_30-03-2021.pdf |