Sign In to Follow Application
View All Documents & Correspondence

A Method, Apparatus And Handset

Abstract: A method of annotating, on a display, a plurality of objects in an image of a scene captured by a camera, the method comprising: receiving i) metadata representing the different annotations to be applied to each of the objects, and ii) position information identifying the real-world position of each object in the scene to which the annotations in the image are to be applied; determining the focal length of the camera and the tilt applied to the camera; determining the position of the camera with respect to the scene being captured; and applying the annotation to the image captured by the camera in accordance with the position information. [Figure 21] .

Get Free WhatsApp Updates!
Notices, Deadlines & Correspondence

Patent Information

Application #
Filing Date
23 March 2012
Publication Number
25/2013
Publication Type
INA
Invention Field
COMPUTER SCIENCE
Status
Email
Parent Application

Applicants

SONY CORPORATION
1-7-1 KONAN, MINATO-KU 108-0075 TOKYO

Inventors

1. ROBERT MARK STEFAN PORTER
23 EGBERT ROAD, WINCHESTER, SO23 7EB

Specification

The present invention relates to a method, apparatus and handset. Description of the Prior Art When viewing an event or scene, it may be useful to obtain further details of the event of scene. This is sometimes referred to as augmented reality. One problem associated with augmented reality is the speed and accuracy with which annotations are overlaid on the real-life image. It is an aim of embodiments of the present invention to address these issues. SUMMARY OF THE INVENTION According to a first aspect, there is provided a method of annotating, on a display, a plurality of objects in an image of a scene captured by a camera, the method comprising: receiving i) metadata representing the different annotations to be applied to each of the objects, and ii) position information identifying the real-world position of each object in the scene to which the annotations in the image are to be applied; determining the focal length of the camera and the tilt applied to the camera; determining the position of the camera with respect to the scene being captured; and applying the annotation to the image captured by the camera in accordance with the position information. The method may further comprise obtaining the position information from an image capture device having a field of view of the scene that is different to the camera, and wherein the position information is determined from an image of the scene captured by the image capture device. The method may further comprise identifying at least one of the objects in the image captured by the camera in accordance with the received real-world position information of the object, the position information of the camera, the focal length of the camera and the tilt applied to the camera; and applying the annotation to the image in accordance with the identified object. The step of identifying the object may comprise detecting the object in the image. The method may further comprise identifying the object in accordance with a stored lens distortion characteristic of the camera. The position information may be global positioning system location information The one object may be either a static object located in the scene or is a unique object in the scene. The method may further comprise storing the metadata and displaying the annotation from the stored metadata. According to another aspect, there is provided an apparatus comprising a display and a camera, the display operative in use to display an image of a scene having a plurality of objects, the image being captured by the camera, the apparatus further comprising: a receiving device operable to receive i) metadata representing the different annotations to be applied to each of the objects, and ii) position information identifying the real-world position of each object in the scene to which the annotations in the image are to be applied; a determining device operable to: determine the focal length of the camera and the tilt applied to the camera; and determine the position of the camera with respect to the scene being captured; and the apparatus comprising a controller operable to apply the annotation to the image captured by the camera in accordance with the position information. The apparatus may further comprise an obtaining device operable to obtain the position information from an image capture device having a field of view of the scene that is different to the camera, and wherein the position information is determined from an image of the scene captured by the image capture device. The controller may be further operable to identify at least one of the objects in the image captured by the camera in accordance with the received real-world position information of the object, the position information of the camera, the focal length of the camera and the tilt applied to the camera; and applying the annotation to the image in accordance with the identified object. The controller may be operable to detect the object in the image. The controller may be further operable to identify the object in accordance with a stored lens distortion characteristic of the camera. The position information may be global positioning system location information The one object may be either a static object located in the scene or is a unique object in the scene. The apparatus may further comprise a storage device operable to store the metadata and the controller is operable to display the annotation from the stored metadata. According to another aspect, there is provided a mobile handset comprising a transceiver for connection to a network and an apparatus according to any one of the above embodiments. BRIEF DESCRIPTION OF THE DRAWINGS The above and other objects, features and advantages of the invention will be apparent from the following detailed description of illustrative embodiments which is to be read in connection with the accompanying drawings, in which: Figure 1 shows a system according to a first embodiment of the present invention; Figure 2 shows a client device in the system of the first embodiment; Figure 3 shows a system according to a second embodiment of the present invention; Figure 4A shows a server of the first embodiment of the present invention; Figure 4B shows a server of the second embodiment of the present invention; Figure 5 shows a flow chart explaining the registration process of the client device to the Server according to either the first or second embodiment; Figure 6 shows a flowchart of a method of object tracking in accordance with examples of the present invention applicable to both the first and second embodiments; Figure 7A shows the creation of object keys in accordance with both the first and second embodiment of the present invention; Figure 7B shows the addition of directional indication to a 3D model of the pitch according to both the first and second embodiment of the present invention; Figure 8 shows a plurality of players and their associated bounding boxes according to both the first and second embodiments of the present invention; Figure 9 shows a flow diagram of a method of object tracking and occlusion detection in accordance with both the first and second embodiments of the present invention; Figures 10A and 10B show some examples of object tracking and occlusion detection in accordance with the first and second embodiment of the present invention; Figure 11 shows a reformatting device located within the server according to the first embodiment of the present invention; Figure 12 shows a reformatting device located within the server according to the second embodiment of the present invention; Figure 13 is a schematic diagram of a system for determining the distance between a position of the camera and objects within a field of view of the camera in accordance with both the first and second embodiments of the present invention; Figure 14 is a schematic diagram of a system for determining the distance between a camera and objects within a field of view of the camera in accordance with the first and second embodiments of the present invention; Figure 15A shows the client device according to the first embodiment of the present invention; Figure 15B shows the client device according to the second embodiment of the present invention; Figure 16A shows a client processing device located in the client device of Figure 15A; Figure 16B shows a client processing device located in the client device of Figure 15B; Figure 17 shows a networked system according to another embodiment of the present invention; Figure 18 shows a client device according to either the first or second embodiment located in the networked system of Figure 17 used for generating a highlight package; Figures 19A and 19B show a client device according to either the first or the second embodiment located in the networked system of Figure 17 used for viewing a highlight package; Figure 20 shows a plan view of a stadium in which augmented reality may be implemented on a portable device according to another embodiment of the present invention; Figure 21 shows a block diagram of a portable device according to Figure 20; Figure 22 shows the display of the portable device of Figure 20 and 21 when augmented reality is activated; and Figure 23 shows a flow diagram explaining the augmented reality embodiment of the present invention. DESCRIPTION OF THE PREFERRED EMBODIMENTS Embodiments of the invention will now be described with reference to the accompanying drawings, throughout which like parts are referred to by like references, and in which: A system 100 is shown in Figure 1. In this system 100, images of a scene are captured by the camera arrangement 130. In embodiments, the scene is of a sports event, such as a soccer match, although the invention is not so limited. In this camera arrangement 130, three high definition cameras are located on a rig (not shown). The arrangement 130 enables a stitched image to be generated. The arrangement 130 therefore has each camera capturing a different part of the same scene with a small overlap in the field of view between each camera. The three images are each high definition images, which, when stitched together, result in a super¬high definition image. The three high definition images captured by the three cameras in the 5 camera arrangement 130 are fed into an image processor 135 which performs editing of the images such as colour enhancement. Also, the image processor 135 receives metadata from the cameras in the camera arrangement 130 relating to camera parameters such as focal length, zoom factor and the like. The enhanced images and the metadata are fed into a server 110 of a first embodiment which will be explained later with reference to Figure 4 A or server 110' of a second embodiment which will be explained with reference to Figure 4B. In embodiments, the actual image stitching is carried out in the user devices 200A-N. However, in order to reduce the computational expense within the user devices 200A-N, the parameters required to perform the stitching are calculated within a server 110 to which the image processing device 135 is connected. The server 110 may be wired or wirelessly connected to the image processor 135 directly or via a network, such as a local area network, wide area network, or the Internet. The method of calculating the parameters, and actually performing the stitching, is described in GB 2444566A. Further disclosed in GB 2444566A is a suitable type of camera arrangement 130. The contents of GB 2444566A relating to the calculation of the parameters, the stitching method and the camera arrangement is incorporated herein. As noted in GB 2444566A the camera parameters for each camera in the camera arrangement 130 are determined. These parameters include the focal length and relative yaw, pitch and roll for each camera as well as parameters that correct for lens distortion, barrel distortion and the like and are determined on the server 110. Also, other parameters such as chromatic aberration correction parameters, colourimetry and exposure correction parameters required for stitching the image may also be calculated in the server 110. Moreover, as the skilled person will appreciate, there may be other values calculated in the server 110 which are required in the image stitching process. These values are explained in GB 2444566A and so, for brevity, will not be explained hereinafter. These values calculated in the server 110 are sent to each user device 200A-N as will be explained later. In addition to the image stitching parameters being calculated within the server 110, other calculations take place. For example, object detection and segmentation takes place identifying and extracting objects in the images to which a three dimensional effect may be applied. Positional information identifying the location of each detected object within the image is also determined within the server 110. Moreover, a depth map is generated within the server 110. The depth map allocates each pixel in the image captured by a camera with a corresponding distance from the camera in the captured scene. In other words, once the depth map is complete for a captured image, it is possible to determine the distance between the point in the scene corresponding to the pixel and the camera capturing the image. Also maintained within the server 110 is a background model which is periodically updated. The background model is updated such that different parts of the background image are updated at different rates. Specifically, the background model is updated in dependence on whether the part of the image was detected as a player in the previous frame. Alternatively, the server 110 may have two background models. In this case, within the server 110 a long term background model and a short term background model is maintained. The long term background model defines a background in the image over a longer period of time such as 5 minutes, whereas the short term model defines a background over a shorter period such as 1 second. The use of a short and long term background model enable short term events such as lighting changes to be taken into account. The depth map which is calculated within the server 110 is sent to each user device 200A-N. In embodiments, each camera within the camera arrangement 130 is fixed. This means that the depth map does not change over time. However, the depth map for each camera is sent to each user device 200A-N upon a trigger to allow for new user devices to be connected to the server 110. For example, the depth map may be sent out when the new user device registers with the server 110 or periodically in time. As would be appreciated, if the field of view of the cameras moved, the depth map would need to be recalculated and sent to the user devices 200A-N more frequently. However, it is also envisaged that the depth map be sent continually to each user device 200A-N. The manner in which the depth map and background models are generated will be explained later. Further, the manner in which the object detection and object segmentation is performed will be explained later. Also connected to the server 110 is a plurality of user devices 200A-N. These user devices 200A-N are connected to the server 110, in embodiments, over the Internet 120. However, it is understood that the invention is not so limited and that the user devices 200A-N could be connected to the server 110 over any type of network such as a Local Area Network (LAN), or may be wired to the server 110 or wirelessly connected to the server 110. Also attached to each user device is a corresponding display 205A-N. The display 205A-N may be a television, or monitor or any kind of display capable of displaying images that can be perceived by a user as being a three dimensional image. In embodiments of the invention, the user device 200A-N is a PlayStation

Documents

Application Documents

# Name Date
1 1084-CHE-2012 POWER OF ATTORNEY 23-03-2012.pdf 2012-03-23
2 1084-CHE-2012 FORM-5 23-03-2012.pdf 2012-03-23
3 1084-CHE-2012 FORM-3 23-03-2012.pdf 2012-03-23
4 1084-CHE-2012 FORM-2 23-03-2012.pdf 2012-03-23
5 1084-CHE-2012 FORM-1 23-03-2012.pdf 2012-03-23
6 1084-CHE-2012 DRAWINGS 23-03-2012.pdf 2012-03-23
7 1084-CHE-2012 DESCRIPTION (COMPLETE) 23-03-2012.pdf 2012-03-23
8 1084-CHE-2012 CORRESPONDENCE OTHERS 23-03-2012.pdf 2012-03-23
9 1084-CHE-2012 CLAIMS 23-03-2012.pdf 2012-03-23
10 1084-CHE-2012 ABSTRACT 23-03-2012.pdf 2012-03-23
11 1084-CHE-2012 OTHERS 23-03-2012.pdf 2012-03-23
12 abstract1084-CHE-2012.jpg 2013-04-10
13 1084-CHE-2012 FORM-3 16-05-2013.pdf 2013-05-16
14 1084-CHE-2012 CORRESPONDENCE OTHERS 16-05-2013.pdf 2013-05-16
15 1084-CHE-2012 FORM-3 30-09-2013.pdf 2013-09-30
16 1084-CHE-2012 CORRESPONDENCE OTHERS 30-09-2013.pdf 2013-09-30
17 1084-CHE-2012 FORM-3 07-11-2014.pdf 2014-11-07
18 1084-CHE-2012 CORRESPONDENCE OTHERS 07-11-2014.pdf 2014-11-07