Abstract: An image processing device (10) according to the present disclosure comprises: a skeleton detection unit (11) which detects two-dimensional skeleton structures of a plurality of persons on the basis of an acquired two-dimensional image; a feature value computation unit (12) which computes feature values of the plurality of two-dimensional skeleton structures detected by the skeleton detection unit (11); and a recognition unit (13) which carries out a recognition process of the state of the plurality of persons on the basis of the similarity of the plurality of feature values computed by the feature value computation unit (12).
Title of the Invention: A non-temporary computer-readable medium in which an image processing apparatus, an image processing method, and an image processing program are stored.
Technical field
[0001]
The present invention relates to a non-temporary computer-readable medium in which an image processing device, an image processing method, and an image processing program are stored.
Background technology
[0002]
In recent years, in surveillance systems and the like, techniques for detecting and searching for states such as the posture and behavior of a person from images of surveillance cameras have been used. As related techniques, for example, Patent Documents 1 and 2 are known. Patent Document 1 discloses a technique for searching the posture of a similar person based on key joints such as the head and limbs of the person included in the depth image. Patent Document 2 discloses a technique for searching for a similar image by using posture information such as inclination added to the image, although it is not related to the posture of the person. In addition, Non-Patent Document 1 is known as a technique related to human skeleton estimation.
prior art documents
patent literature
[0003]
Patent Document 1: Japanese Patent Application Laid-Open No. 2014-52035
Patent Document 2: Japanese Unexamined Patent Publication No. 2006-260405
Non-patent literature
[0004]
Non-Patent Document 1: Zhe Cao, Tomas Simon, Shih-En Wei, Yaser Sheikh, "Realtime Multi-Person 2D Pose Optimization using Part Affinity Fields", The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, P. 7291-7299
Outline of the invention
Problems to be Solved by the Invention
[0005]
In the related technique such as Patent Document 1, the posture of a similar person is searched based on the head, limbs, etc. of the person, but it is premised that a depth image (depth image) is acquired. Therefore, it is not possible to search for a state such as a posture from a two-dimensional image. Also, only searches from a specific point of view are considered. Therefore, in the related technique, there is a problem that it is difficult to flexibly perform state recognition processing such as search and classification of the state of a person based on a two-dimensional image.
[0006]
In view of such a problem, the present disclosure is a non-temporary computer in which an image processing device, an image processing method, and an image processing program capable of flexibly performing state recognition processing of a person based on a two-dimensional image are stored. The purpose is to provide a readable medium.
Means to solve problems
[0007]
The image processing apparatus according to the present disclosure calculates a skeleton detecting means for detecting a two-dimensional skeleton structure of a plurality of persons and a feature amount of the detected plurality of two-dimensional skeleton structures based on the acquired two-dimensional image. The feature amount calculation means is provided, and a recognition means for recognizing the state of the plurality of persons based on the similarity of the calculated feature amounts.
[0008]
The image processing method according to the present disclosure detects the two-dimensional skeleton structures of a plurality of persons based on the acquired two-dimensional image, calculates the feature amounts of the detected two-dimensional skeleton structures, and calculates the above. Based on the degree of similarity of the plurality of feature quantities, the state recognition process of the plurality of persons is performed.
[0009]
The non-temporary computer-readable medium in which the image processing program according to the present disclosure is stored detects the two-dimensional skeleton structures of a plurality of persons based on the acquired two-dimensional image, and the detected two-dimensional image. An image processing program for causing a computer to perform a process of recognizing the state of the plurality of persons based on the calculated feature amounts of the skeletal structure and the similarity of the calculated feature amounts is stored. It is a non-temporary computer-readable medium.
Effect of the invention
[0010]
According to the present disclosure, there is provided a non-temporary computer-readable medium in which an image processing device, an image processing method, and an image processing program capable of flexibly performing state recognition processing of a person based on a two-dimensional image are stored. be able to.
A brief description of the drawing
[0011]
FIG. 1 is a configuration diagram showing an outline of an image processing apparatus according to an embodiment.
FIG. 2 is a configuration diagram showing a configuration of an image processing apparatus according to the first embodiment.
FIG. 3 is a flowchart showing an image processing method according to the first embodiment.
FIG. 4 is a flowchart showing a classification method according to the first embodiment.
FIG. 5 is a flowchart showing a search method according to the first embodiment.
FIG. 6 is a diagram showing a detection example of a skeleton structure according to the first embodiment.
FIG. 7 is a diagram showing a human body model according to the first embodiment.
FIG. 8 is a diagram showing a detection example of a skeleton structure according to the first embodiment.
FIG. 9 is a diagram showing a detection example of a skeleton structure according to the first embodiment.
FIG. 10 is a diagram showing a detection example of a skeleton structure according to the first embodiment.
FIG. 11 is a graph showing a specific example of the classification method according to the first embodiment.
FIG. 12 is a diagram showing a display example of a classification result according to the first embodiment.
FIG. 13 is a diagram for explaining a search method according to the first embodiment.
FIG. 14 is a diagram for explaining a search method according to the first embodiment.
FIG. 15 is a diagram for explaining a search method according to the first embodiment.
FIG. 16 is a diagram for explaining a search method according to the first embodiment.
FIG. 17 is a diagram showing a display example of a search result according to the first embodiment.
FIG. 18 is a configuration diagram showing a configuration of an image processing apparatus according to a second embodiment.
FIG. 19 is a flowchart showing an image processing method according to the second embodiment.
FIG. 20 is a flowchart showing a specific example 1 of a height pixel number calculation method according to a second embodiment.
FIG. 21 is a flowchart showing a specific example 2 of the height pixel number calculation method according to the second embodiment.
FIG. 22 is a flowchart showing a specific example 2 of the height pixel number calculation method according to the second embodiment.
FIG. 23 is a flowchart showing a normalization method according to the second embodiment.
FIG. 24 is a diagram showing a human body model according to the second embodiment.
FIG. 25 is a diagram showing a detection example of a skeleton structure according to a second embodiment.
FIG. 26 is a diagram showing a detection example of a skeletal structure according to a second embodiment.
FIG. 27 is a diagram showing a detection example of a skeletal structure according to a second embodiment.
FIG. 28 is a diagram showing a human body model according to the second embodiment.
FIG. 29 is a diagram showing a detection example of a skeletal structure according to a second embodiment.
FIG. 30 is a histogram for explaining a method of calculating the number of height pixels according to the second embodiment.
FIG. 31 is a diagram showing a detection example of a skeleton structure according to a second embodiment.
FIG. 32 is a diagram showing a three-dimensional human body model according to the second embodiment.
FIG. 33 is a diagram for explaining a method for calculating the number of height pixels according to the second embodiment.
FIG. 34 is a diagram for explaining a method for calculating the number of height pixels according to the second embodiment.
FIG. 35 is a diagram for explaining a method for calculating the number of height pixels according to the second embodiment.
FIG. 36 is a diagram for explaining a normalization method according to the second embodiment.
FIG. 37 is a diagram for explaining a normalization method according to the second embodiment.
FIG. 38 is a diagram for explaining a normalization method according to the second embodiment.
FIG. 39 is a configuration diagram showing an outline of computer hardware according to an embodiment.
MODE FOR CARRYING OUT THE INVENTION
[0012]
Hereinafter, embodiments will be described with reference to the drawings. In each drawing, the same elements are designated by the same reference numerals, and duplicate explanations are omitted as necessary.
[0013]
(Examination leading to the embodiment)
In recent years, image recognition technology utilizing machine learning such as deep learning has been applied to various systems. For example, it is being applied to a surveillance system that monitors images from a surveillance camera. By utilizing machine learning for the monitoring system, it is becoming possible to grasp the state such as the posture and behavior of a person from an image to some extent.
[0014]
However, with such related technologies, it may not always be possible to grasp the state of the person desired by the user on demand. For example, there are cases where the state of a person that the user wants to search and grasp can be specified in advance, and there are cases where the state cannot be specifically specified such as an unknown state. Then, in some cases, it is not possible to specify in detail the state of the person that the user wants to search. In addition, if a part of the person's body is hidden, it is not possible to perform a search or the like. With related techniques, it is difficult to flexibly search and classify a desired person's state because the person's state can be searched only from specific search conditions.
[0015]
Therefore, the inventors examined a method of using a skeleton estimation technique such as Non-Patent Document 1 in order to recognize the state of a person desired by a user from an image on demand. In a related skeleton estimation technique such as OpenPose disclosed in Non-Patent Document 1, the skeleton of a person is estimated by learning various patterns of correctly answered image data. In the following embodiments, it is possible to flexibly recognize the state of a person by utilizing such a skeleton estimation technique.
[0016]
The skeletal structure estimated by a skeletal estimation technique such as OpenPose is composed of "key points" which are characteristic points of joints and the like and "bones (bone links)" which indicate links between key points. .. Therefore, in the following embodiments, the skeletal structure will be described using the terms "key point" and "bone", but unless otherwise specified, the "key point" corresponds to the "joint" of a person and is described as "key point". "Bone" corresponds to the "bone" of a person.
[0017]
(Outline of Embodiment)
FIG. 1 shows an outline of the image processing device 10 according to the embodiment. As shown in FIG. 1, the image processing device 10 includes a skeleton detection unit 11, a feature amount calculation unit 12, and a recognition unit 13. The skeleton detection unit 11 detects the two-dimensional skeleton structures of a plurality of people based on the two-dimensional image acquired from a camera or the like. The feature amount calculation unit 12 calculates the feature amount of the plurality of two-dimensional skeleton structures detected by the skeleton detection unit 11. The recognition unit 13 performs the recognition processing of the states of a plurality of persons based on the similarity of the plurality of feature amounts calculated by the feature amount calculation unit 12. The recognition process is a classification process of a person's state, a search process, or the like.
[0018]
As described above, in the embodiment, the two-dimensional skeleton structure of the person is detected from the two-dimensional image, and the recognition process such as classification and examination of the state of the person is performed based on the feature amount calculated from the two-dimensional skeleton structure. This makes it possible to flexibly recognize the state of a desired person.
[0019]
(Embodiment 1)
Hereinafter, the first embodiment will be described with reference to the drawings. FIG. 2 shows the configuration of the image processing device 100 according to the present embodiment. The image processing device 100 constitutes the image processing system 1 together with the camera 200 and the database (DB) 110. The image processing system 1 including the image processing device 100 is a system for classifying and searching states such as the posture and behavior of a person based on the skeleton structure of the person estimated from the image.
[0020]
The camera 200 is an imaging unit such as a surveillance camera that generates a two-dimensional image. The camera 200 is installed at a predetermined location and images a person or the like in the imaging region from the installation location. The camera 200 is directly connected so that the captured image (video) can be output to the image processing device 100, or is connected via a network or the like. The camera 200 may be provided inside the image processing device 100.
[0021]
The database 110 is a database that stores information (data), processing results, and the like necessary for processing of the image processing apparatus 100. The database 110 includes an image acquired by the image acquisition unit 101, a detection result of the skeletal structure detection unit 102, data for machine learning, a feature amount calculated by the feature amount calculation unit 103, a classification result of the classification unit 104, and a search unit 105. The search results etc. of are memorized. The database 110 is directly connected to the image processing device 100 so that data can be input and output as needed, or is connected via a network or the like. The image processing device 100 uses the database 110 as a non-volatile memory such as a flash memory, a hard disk device, or the like. It may be provided inside.
[0022]
As shown in FIG. 2, the image processing apparatus 100 includes an image acquisition unit 101, a skeleton structure detection unit 102, a feature amount calculation unit 103, a classification unit 104, a search unit 105, an input unit 106, and a display unit 107. The configuration of each unit (block) is an example, and may be configured by other units as long as the method (operation) described later is possible. Further, the image processing apparatus 100 is realized by a computer device such as a personal computer or a server that executes a program, for example. good. For example, the input unit 106, the display unit 107, and the like may be external devices. Moreover, both the classification unit 104 and the search unit 105 may be provided, or only one of them may be provided. Both or one of the classification unit 104 and the retrieval unit 105 is a recognition unit that performs recognition processing of the person's state.
[0023]
The image acquisition unit 101 acquires a two-dimensional image including a person captured by the camera 200 . The image acquisition unit 101 acquires, for example, an image including a person (video including a plurality of images) captured by the camera 200 during a predetermined monitoring period. It should be noted that an image including a person prepared in advance may be acquired from the database 110 or the like instead of being acquired from the camera 200 .
[0024]
The skeletal structure detection unit 102 detects the 2D skeletal structure of the person in the image based on the acquired 2D image. The skeletal structure detection unit 102 detects skeletal structures of all persons recognized in the acquired image. The skeletal structure detection unit 102 detects the skeletal structure of a person based on recognized features such as the joints of the person, using a skeletal structure estimation technique using machine learning. The skeleton structure detection unit 102 uses, for example, a skeleton estimation technique such as OpenPose described in Non-Patent Document 1.
[0025]
The feature quantity calculation unit 103 calculates the feature quantity of the detected two-dimensional skeletal structure and stores the calculated feature quantity in the database 110 . The feature amount of the skeletal structure indicates the features of the skeletal structure of the person, and serves as an element for classifying and retrieving the state of the person based on the skeletal structure of the person. The feature amount may be the feature amount of the entire skeleton structure, the feature amount of a part of the skeleton structure, or may include a plurality of feature amounts like each part of the skeleton structure. Any method such as machine learning or normalization may be used as the method for calculating the feature amount, and the minimum value or the maximum value may be obtained as the normalization. As an example, the feature amount is a feature amount obtained by machine-learning the skeletal structure, the size of the skeletal structure from the head to the foot on the image, and the like. The size of the skeletal structure is the vertical height, area, etc. of the skeletal region containing the skeletal structure on the image. The vertical direction (height direction or vertical direction) is the vertical direction (Y-axis direction) in the image, for example, the direction perpendicular to the ground (reference plane). The left-right direction (horizontal direction) is the left-right direction (X-axis direction) in the image, for example, the direction parallel to the ground.
[0026]
In addition, in order to perform the classification and search desired by the user, it is preferable to use features that are robust to classification and search processing. For example, if the user desires classification or retrieval that does not depend on a person's orientation or body shape, a feature quantity that is robust to the person's orientation or body shape may be used. By learning the skeletons of people facing various directions in the same posture and the skeletons of various body types in the same posture, and by extracting features only in the vertical direction of the skeleton, It is possible to obtain features that do not
[0027]
The classification unit 104 classifies (clusters) a plurality of skeletal structures stored in the database 110 based on the degree of similarity of feature amounts of the skeletal structures. It can be said that the classification unit 104 classifies the states of a plurality of persons based on the feature amount of the skeletal structure as the process of recognizing the states of the persons. The degree of similarity is the distance between features of the skeleton structure. The classification unit 104 may classify the skeletal structure according to the similarity of the feature amount of the entire skeletal structure, or may classify the skeletal structure according to the similarity of the feature amount of part of the skeletal structure. Both hands) and the second part (both feet, for example) may be classified according to the similarity of feature amounts. The posture of the person may be classified based on the feature amount of the skeletal structure of the person in each image, or the behavior of the person may be classified based on the change in the feature amount of the skeletal structure of the person in a plurality of consecutive images in time series. can be classified. That is, the classification unit 104 can classify the state of the person, including the posture and behavior of the person, based on the feature amount of the skeletal structure.
For example, the classification unit 104 classifies a plurality of skeletal structures in a plurality of images captured during a predetermined monitoring period. The classification unit 104 obtains the degree of similarity between the feature quantities to be classified, and classifies the skeletal structures with a high degree of similarity into the same cluster (group of similar postures). It should be noted that the user may be allowed to specify the classification condition as in the search. The classification unit 104 stores the classification result of the skeletal structure in the database 110 and displays it on the display unit 107 .
[0028]
The search unit 105 searches a plurality of skeleton structures stored in the database 110 for a skeleton structure with a high degree of similarity to the feature quantity of the search query (query state). It can be said that the search unit 105 searches for a person's state corresponding to a search condition (query state) from among a plurality of persons' states based on the feature amount of the skeletal structure as the recognition processing of the person's state. Similar to classification, similarity is the distance between skeletal structure features. The search unit 105 may search based on the similarity of the feature amount of the entire skeletal structure, or may search based on the similarity of the feature amount of a part of the skeletal structure. Both hands) and the second part (both feet, for example) may be retrieved based on the similarity of feature amounts. The posture of a person may be retrieved based on the feature amount of the skeletal structure of the person in each image, or the behavior of the person may be searched based on changes in the feature amount of the skeletal structure of the person in a plurality of images that are consecutive in time series. can be searched. That is, the search unit 105 can search for a person's state, including the person's posture and behavior, based on the feature amount of the skeletal structure.
For example, the search unit 105 searches feature amounts of a plurality of skeletal structures in a plurality of images captured during a predetermined monitoring period, similar to the classification target. Also, the skeleton structure (posture) specified by the user from among the classification results displayed by the classification unit 104 is used as a search query (search key). It should be noted that the search query may be selected from among a plurality of unclassified skeletal structures, and the user may input a skeletal structure to be the search query. The search unit 105 searches for a feature amount having a high degree of similarity with the feature amount of the skeleton structure of the search query from among the feature amounts to be searched. The search unit 105 stores the search result of the feature amount in the database 110 and displays it on the display unit 107 .
[0029]
The input unit 106 is an input interface that acquires information input by the user who operates the image processing apparatus 100 . For example, the user is a surveillance person who monitors a person in a suspicious state from an image of a surveillance camera. The input unit 106 is, for example, a GUI (Graphical User Interface), and receives information according to user operations from an input device such as a keyboard, mouse, or touch panel. For example, the input unit 106 receives, as a search query, the skeletal structure of a specified person from among the skeletal structures (postures) classified by the classifying unit 104 .
[0030]
The display unit 107 is a display unit that displays the result of the operation (processing) of the image processing apparatus 100, and is, for example, a display device such as a liquid crystal display or an organic EL (Electro Luminescence) display. A display unit 107 displays the classification result of the classification unit 104 and the search result of the search unit 105 on a GUI according to the degree of similarity or the like.
[0031]
3 to 5 show the operation of the image processing apparatus 100 according to this embodiment. 3 shows the flow from image acquisition to search processing in the image processing apparatus 100, FIG. 4 shows the flow of the classification processing (S104) in FIG. 3, and FIG. 5 shows the search processing (S105) in FIG. showing the flow.
[0032]
As shown in FIG. 3, the image processing device 100 acquires an image from the camera 200 (S101). The image acquisition unit 101 acquires an image of a person in order to classify and search from the skeletal structure, and stores the acquired image in the database 110 . The image acquisition unit 101 acquires, for example, a plurality of images captured during a predetermined monitoring period, and performs subsequent processing on all persons included in the plurality of images.
[0033]
Subsequently, the image processing apparatus 100 detects the skeletal structure of the person based on the acquired image of the person (S102). FIG. 6 shows an example of skeletal structure detection. As shown in FIG. 6, an image acquired from a monitoring camera or the like includes a plurality of persons, and the skeletal structure of each person included in the image is detected.
[0034]
FIG. 7 shows the skeletal structure of the human body model 300 detected at this time, and FIGS. 8 to 10 show detection examples of the skeletal structure. A skeleton structure detection unit 102 detects the skeleton structure of a human body model (two-dimensional skeleton model) 300 as shown in FIG. The human body model 300 is a two-dimensional model composed of key points such as human joints and bones connecting the key points.
[0035]
For example, the skeletal structure detection unit 102 extracts feature points that can be keypoints from the image, refers to information obtained by machine learning the image of the keypoints, and detects each keypoint of the person. In the example of FIG. 7, the key points of the person are head A1, neck A2, right shoulder A31, left shoulder A32, right elbow A41, left elbow A42, right hand A51, left hand A52, right hip A61, left hip A62, right knee A71. , left knee A72, right foot A81, and left foot A82. Further, as human bones connecting these key points, bone B1 connecting head A1 and neck A2, bone B21 and bone B22 connecting neck A2 and right shoulder A31 and left shoulder A32, right shoulder A31 and left shoulder A32 and right shoulder A31 and right shoulder A32, respectively. Bone B31 and B32 connecting elbow A41 and left elbow A42 respectively, bone B41 and bone B42 connecting right elbow A41 and left elbow A42 with right hand A51 and left hand A52 respectively, neck A2 connecting right hip A61 and left hip A62 respectively Bone B51 and B52, Bone B61 and B62 connecting right hip A61 and left hip A62 and right knee A71 and left knee A72, respectively, bone B71 and bone connecting right knee A71 and left knee A72, right leg A81 and left leg A82, respectively B72 is detected. The skeletal structure detection unit 102 stores the detected skeletal structure of the person in the database 110 .
[0036]
FIG. 8 is an example of detecting a person standing upright. In FIG. 8, an upright person is imaged from the front, and bones B1, B51 and B52, B61 and B62, and B71 and B72 viewed from the front are detected without overlapping each other. The bones B61 and B71 are slightly more bent than the left leg bones B62 and B72.
[0037]
Fig. 9 is an example of detecting a person who is crouching. In FIG. 9, a crouching person is imaged from the right side, bone B1, bone B51 and bone B52, bone B61 and bone B62, bone B71 and bone B72 seen from the right side are detected, and bone B61 of the right leg is detected. And the bone B71 and the bones B62 and B72 of the left leg are greatly bent and overlapped.
[0038]
FIG. 10 is an example of detecting a sleeping person. In FIG. 10, a person lying down is imaged obliquely from the front left, and bones B1, B51 and B52, B61 and B62, B71 and B72 are detected from the oblique front left, and bones B71 and B72 are detected. The bones B61 and B71 of the left leg and the bones B62 and B72 of the left leg are bent and overlapped.
[0039]
Subsequently, as shown in FIG. 3, the image processing apparatus 100 calculates the feature quantity of the detected skeletal structure (S103). For example, when the height or area of the skeleton region is used as the feature quantity, the feature quantity calculation unit 103 calculates the region including the skeleton structure. Is extracted, and the height (number of pixels) and area (pixel area) of the area are obtained. The height and area of the skeletal region can be obtained from the coordinates of the end of the extracted skeleton region and the coordinates of the key points at the ends. The feature amount calculation unit 103 stores the obtained feature amount of the skeletal structure in the database 110.
[0040]
In the example of FIG. 8, the skeletal region including all bones is extracted from the skeletal structure of an upright person. In this case, the upper end of the skeletal area is the key point A1 of the head, the lower end of the skeletal area is the key point A82 of the left foot, the left end of the skeletal area is the key point A41 of the right elbow, and the right end of the skeletal area is the key point A52 of the left hand. .. Therefore, the height of the skeleton region is obtained from the difference between the Y coordinates of the key point A1 and the key point A82. Further, the width of the skeleton region is obtained from the difference between the X coordinates of the key point A41 and the key point A52, and the area is obtained from the height and width of the skeleton region.
[0041]
In the example of FIG. 9, a skeletal region including all bones is extracted from the skeletal structure of a crouched person. In this case, the upper end of the skeletal area is the key point A1 of the head, the lower end of the skeletal area is the key point A81 of the right foot, the left end of the skeletal area is the key point A61 of the right hip, and the right end of the skeletal area is the key point A51 of the right hand. .. Therefore, the height of the skeleton region is obtained from the difference between the Y coordinates of the key point A1 and the key point A81. Further, the width of the skeleton region is obtained from the difference between the X coordinates of the key point A61 and the key point A51, and the area is obtained from the height and width of the skeleton region.
[0042]
In the example of FIG. 10, a skeletal region including all bones is extracted from the skeletal structure of a person who has fallen in the left-right direction of the image. In this case, the upper end of the skeletal region is the left shoulder key point A32, the lower end of the skeletal region is the left hand key point A52, the left end of the skeletal region is the right hand key point A51, and the right end of the skeletal region is the left foot key point A82. Therefore, the height of the skeleton region is obtained from the difference between the Y coordinates of the key point A32 and the key point A52. Further, the width of the skeleton region is obtained from the difference between the X coordinates of the key point A51 and the key point A82, and the area is obtained from the height and width of the skeleton region.
[0043]
Subsequently, as shown in FIG. 3, the image processing apparatus 100 performs classification processing (S104). In the classification process, as shown in FIG. 4, the classification unit 104 calculates the similarity of the calculated features of the skeletal structure (S111), and classifies the skeletal structure based on the calculated features (S112). .. The classification unit 104 obtains the similarity of the features between all the skeletal structures stored in the database 110 to be classified, and classifies (clusters) the skeletal structures (postures) having the highest similarity into the same cluster. .. Further, the similarity between the classified clusters is obtained and classified, and the classification is repeated until a predetermined number of clusters are obtained. FIG. 11 shows an image of the classification result of the feature amount of the skeletal structure. FIG. 11 is an image of cluster analysis by a two-dimensional classification element, and the two classification elements are, for example, the height of the skeleton region and the area of the skeleton region. In FIG. 11, as a result of classification, the features of the plurality of skeletal structures are classified into three clusters C1 to C3. The clusters C1 to C3 correspond to each posture such as a standing posture, a sitting posture, and a sleeping posture, and the skeletal structure (person) is classified according to the similar posture.
[0044]
In the present embodiment, various classification methods can be used by classifying based on the feature amount of the skeletal structure of a person. The classification method may be preset or may be arbitrarily set by the user. Further, the classification may be performed by the same method as the search method described later. That is, it may be classified according to the same classification conditions as the search conditions. For example, the classification unit 104 classifies by the following classification method. Any classification method may be used, or an arbitrarily selected classification method may be combined.
[0045]
(Classification method 1) Classification by multiple layers
Classification by the skeletal structure of the whole body, classification by the skeletal structure of the upper and lower body, classification by the skeletal structure of the arms and legs, etc. are hierarchically combined and classified. That is, the classification may be performed based on the feature amounts of the first portion and the second portion of the skeletal structure, and further, the feature amounts of the first portion and the second portion may be weighted and classified.
[0046]
(Classification method 2) Classification by multiple images in chronological order
Classification is based on the features of the skeletal structure in multiple images that are continuous in time series. For example, the features may be stacked in the time series direction and classified based on the cumulative value. Further, the classification may be performed based on the change (change amount) of the feature amount of the skeletal structure in a plurality of consecutive images.
[0047]
(Classification method 3) Classification that ignores the left and right sides of the skeletal structure
The skeletal structures with opposite right and left sides of the person are classified as the same skeletal structure.
[0048]
Further, the classification unit 104 displays the classification result of the skeletal structure (S113). The classification unit 104 acquires images of necessary skeletal structures and people from the database 110, and displays the skeletal structure and people on the display unit 107 for each posture (cluster) similar as a classification result. FIG. 12 shows a display example when the postures are classified into three. For example, as shown in FIG. 12, the posture areas WA1 to WA3 for each posture are displayed in the display window W1, and the skeletal structure and the person (image) of the posture corresponding to each of the posture areas WA1 to WA3 are displayed. The posture area WA1 is, for example, a display area for a standing posture, and displays a skeletal structure and a person similar to the standing posture, which are classified into cluster C1. The posture area WA2 is, for example, a sitting posture display area, and displays a skeletal structure and a person similar to the sitting posture classified into cluster C2. The posture area WA3 is, for example, a display area for a sleeping posture, and displays a skeletal structure and a person similar to the sleeping posture classified into cluster C2.
[0049]
Subsequently, as shown in FIG. 3, the image processing device 100 performs a search process (S105). In the search process, as shown in FIG. 5, the search unit 105 accepts the input of the search condition (S121) and searches the skeleton structure based on the search condition (S122). The search unit 105 receives input of a search query, which is a search condition, from the input unit 106 according to a user operation. When inputting a search query from the classification result, for example, in the display example of FIG. 12, the user specifies (selects) the skeleton structure of the posture to be searched from the posture areas WA1 to WA3 displayed in the display window W1. .. Then, the search unit 105 uses the skeleton structure specified by the user as a search query to search for a skeleton structure having a high degree of similarity in the feature amount from all the skeleton structures stored in the database 110 to be searched. The search unit 105 calculates the similarity between the feature amount of the skeleton structure of the search query and the feature amount of the skeleton structure to be searched, and extracts the skeleton structure whose calculated similarity is higher than a predetermined threshold value. As the feature amount of the skeleton structure of the search query, the feature amount calculated in advance may be used, or the feature amount obtained at the time of the search may be used. The search query may be input by moving each part of the skeleton structure according to the operation of the user, or the posture demonstrated by the user in front of the camera may be used as the search query.
[0050]
In the present embodiment, as in the classification method, various search methods can be used by searching based on the feature amount of the skeletal structure of the person. The search method may be preset or may be arbitrarily set by the user. For example, the search unit 105 searches by the following search method. Either search method may be used, or an arbitrarily selected search method may be combined. A plurality of search methods (search conditions) may be combined and searched by a logical expression (for example, AND (logical product), OR (logical sum), NOT (negative)). For example, the search condition may be searched as "(posture in which the right hand is raised) AND (posture in which the left foot is raised)".
[0051]
(Search method 1) Search only by features in the height direction
By searching using only the features in the height direction of the person, the influence of the change in the lateral direction of the person can be suppressed, and the robustness is improved with respect to the change in the direction of the person and the body shape of the person. For example, as in the skeletal structures 501 to 503 of FIG. 13, even if the orientation and body shape of the person are different, the feature amount in the height direction does not change significantly. Therefore, it can be determined that the skeletal structures 501 to 503 have the same posture at the time of searching (at the time of classification).
[0052]
(Search method 2) Partial search
If a part of the person's body is hidden in the image, search using only the information of the recognizable part. For example, as in the skeletal structures 511 and 512 of FIG. 14, even if the key points of the left foot cannot be detected due to the hiding of the left foot, the features of the other key points that have been detected can be used for the search. Therefore, in the skeletal structures 511 and 512, it can be determined that the postures are the same at the time of searching (at the time of classification). That is, it is possible to perform classification and search using the features of some key points instead of all the key points. In the examples of the skeletal structures 521 and 522 in FIG. 15, although the directions of both feet are different, the feature quantities of the key points of the upper body (A1, A2, A31, A32, A41, A42, A51, A52) are used as the search query. Therefore, it can be judged that the posture is the same. Further, the portion (feature point) to be searched may be weighted and searched, or the threshold value for determining the similarity may be changed. When a part of the body is hidden, the hidden part may be ignored and the search may be performed, or the hidden part may be added to the search. By searching including the hidden part, it is possible to search for a posture in which the same part is hidden.
[0053]
(Search method 3) Search ignoring the left and right sides of the skeletal structure
Search for the skeletal structure with the opposite right and left sides of the person as the same skeletal structure. For example, as shown in the skeletal structures 531 and 532 of FIG. 16, the posture in which the right hand is raised and the posture in which the left hand is raised can be searched (classified) as the same posture. In the example of FIG. 16, the skeletal structure 531 and the skeletal structure 532 have different positions of the right hand key point A51, the right elbow key point A41, the left hand key point A52, and the left elbow key point A42, but other key points. The position of is the same. Of the right hand key point A51 and right elbow key point A41 of the skeletal structure 531 and the left hand key point A52 and the left elbow key point A42 of the skeletal structure 532, when one of the key points of the skeletal structure is flipped left and right, the other It becomes the same position as the key point of the skeletal structure of, and one of the key point A52 of the left hand and the key point A42 of the left elbow of the skeletal structure 531 and the key point A51 of the right hand and the key point A41 of the right elbow of the skeletal structure 532. When the key point of the skeletal structure of is inverted left and right, it becomes the same position as the key point of the other skeletal structure, so it is judged that the posture is the same.
[0054]
(Search method 4) Search by features in the vertical and horizontal directions
After searching only with the features in the vertical direction (Y-axis direction) of the person, the obtained results are further searched using the features in the horizontal direction (X-axis direction) of the person.
[0055]
(Search method 5) Search by multiple images in chronological order
Search based on the features of the skeletal structure in multiple images that are continuous in time series. For example, features may be stacked in the time series direction and searched based on the cumulative value. Further, the search may be performed based on the change (change amount) of the feature amount of the skeletal structure in a plurality of consecutive images.
[0056]
Further, the search unit 105 displays the search result of the skeletal structure (S123). The search unit 105 acquires the necessary images of the skeleton structure and the person from the database 110, and displays the skeleton structure and the person obtained as the search result on the display unit 107. For example, when multiple search queries (search conditions) are specified, the search results are displayed for each search query. FIG. 17 shows a display example when a search is performed using three search queries (postures). For example, as shown in FIG. 17, in the display window W2, the skeleton structure and the person of the search queries Q10, Q20, and Q30 specified at the left end are displayed, and each search query is displayed on the right side of the search queries Q10, Q20, and Q30. The skeletal structure and people of the search results Q11, Q21, and Q31 are displayed side by side.To.
[0057]
The order in which the search results are displayed next to the search query may be the order in which the relevant skeletal structure was found, or the order in which the similarity is high. When the parts (feature points) of the partial search are weighted for retrieval, they may be displayed in the order of similarity calculated by weighting. It may be displayed in the order of similarity calculated only from the portions (feature points) selected by the user. Further, images (frames) before and after the time series may be cut out for a certain period of time and displayed, centering on the image (frame) of the search result.
[0058]
As described above, in the present embodiment, it is possible to detect the skeletal structure of a person from a two-dimensional image and perform classification and retrieval based on the feature amount of the detected skeletal structure. As a result, it is possible to classify the postures having a high degree of similarity, and to search for similar postures having a high degree of similarity with the search query (search key). By classifying and displaying similar postures from the image, the posture of the person in the image can be grasped without the user specifying the posture or the like. Since the user can specify the posture of the search query from among the classification results, even if the user does not know the details of the posture to be searched in advance, it is possible to search for the desired posture. For example, classification and retrieval can be performed using the whole or part of a person's skeletal structure as a condition, enabling flexible classification and retrieval.
[0059]
(Embodiment 2)
Embodiment 2 will be described below with reference to the drawings. In this embodiment, a specific example of feature amount calculation in the first embodiment will be described. In this embodiment, the feature amount is obtained by normalization using the height of the person. Others are the same as those of the first embodiment.
[0060]
FIG. 18 shows the configuration of the image processing apparatus 100 according to this embodiment. As shown in FIG. 18, the image processing apparatus 100 further includes a height calculator 108 in addition to the configuration of the first embodiment. Note that the feature amount calculation unit 103 and the height calculation unit 108 may be integrated into one processing unit.
[0061]
A height calculation unit (height estimation unit) 108 calculates the height of a person in a two-dimensional image when standing upright (height in pixels) based on the two-dimensional skeletal structure detected by the skeletal structure detection unit 102 ( presume. It can also be said that the number of height pixels is the height of the person in the two-dimensional image (the length of the whole body of the person in the two-dimensional image space). The height calculation unit 108 obtains the number of height pixels (the number of pixels) from the length of each bone of the detected skeletal structure (the length in the two-dimensional image space).
[0062]
In the following examples, specific examples 1 to 3 are used as the method for obtaining the height pixel count. Any one of the methods of Examples 1 to 3 may be used, or a plurality of arbitrarily selected methods may be used in combination. In specific example 1, the number of height pixels is obtained by totaling the length of the bones from the head to the feet among the bones of the skeletal structure. If the skeletal structure detection unit 102 (skeletal structure estimation technology) does not output the top of the head and the feet, it can be corrected by multiplying by a constant as necessary. In specific example 2, the number of height pixels is calculated using a human body model that indicates the relationship between the length of each bone and the length of the whole body (height in a two-dimensional image space). In specific example 3, the number of height pixels is calculated by fitting a three-dimensional human body model to a two-dimensional skeletal structure.
[0063]
The feature amount calculation unit 103 of the present embodiment is a normalization unit that normalizes the skeletal structure (skeletal information) of the person based on the calculated number of pixels of the height of the person. The feature amount calculation unit 103 stores the normalized feature amount (normalized value) of the skeletal structure in the database 110 . The feature amount calculation unit 103 normalizes the height on the image of each key point (feature point) included in the skeletal structure by the number of height pixels. In this embodiment, for example, the height direction is the vertical direction (Y-axis direction) in the two-dimensional coordinate (XY coordinate) space of the image. In this case, the height of the keypoint can be obtained from the Y coordinate value (the number of pixels) of the keypoint.
Alternatively, the height direction may be the direction of the vertical projection axis (vertical projection direction) obtained by projecting the direction of the vertical axis perpendicular to the ground (reference plane) in the three-dimensional coordinate space of the real world onto the two-dimensional coordinate space. In this case, the height of the keypoint is obtained by calculating the vertical projection axis by projecting the axis perpendicular to the ground in the real world onto the two-dimensional coordinate space based on the camera parameters, and calculating the value along this vertical projection axis (the number of pixels ) can be obtained from Note that the camera parameters are imaging parameters of an image. For example, the camera parameters are the attitude, position, imaging angle, focal length, etc. of the camera 200 . With the camera 200, an object whose length and position are known in advance can be imaged, and camera parameters can be obtained from the image. Distortion occurs at both ends of the captured image, and the vertical direction of the real world may not match the vertical direction of the image. On the other hand, by using the parameters of the camera that captured the image, it is possible to know how much the vertical direction in the real world is tilted in the image. Therefore, by normalizing the values of the keypoints along the vertical projection axis projected into the image based on the camera parameters by the height, it is possible to convert the keypoints into features considering the deviation between the real world and the image. can. The left-right direction (horizontal direction) is the left-right direction (X-axis direction) in the two-dimensional coordinate (XY coordinate) space of the image, or the direction parallel to the ground in the three-dimensional coordinate space of the real world. is projected onto the two-dimensional coordinate space.
[0064]
19 to 23 show the operation of the image processing apparatus 100 according to this embodiment. FIG. 19 shows the flow from image acquisition to search processing in the image processing apparatus 100, and FIGS. 23 shows the flow of the normalization process (S202) in FIG.
[0065]
As shown in FIG. 19, in the present embodiment, height pixel number calculation processing (S201) and normalization processing (S202) are performed as the feature amount calculation processing (S103) in Embodiment 1. Others are the same as those of the first embodiment.
[0066]
After image acquisition (S101) and skeletal structure detection (S102), the image processing apparatus 100 performs height pixel count calculation processing based on the detected skeletal structure (S201). In this example, as shown in FIG. 24, the height of the skeletal structure of the person in the image when standing upright is the number of height pixels (h), and the height of each keypoint of the skeletal structure in the state of the person in the image is the keypoint. Let the height be (yi). Specific examples 1 to 3 of the height pixel number calculation process will be described below.
[0067]
In Specific Example 1, the length of the bone from the head to the foot is used to determine the number of pixels in the height. In specific example 1, as shown in FIG. 20, the height calculation unit 108 acquires the length of each bone (S211), and totals the acquired lengths of each bone (S212).
[0068]
The height calculation unit 108 obtains the length of the bones on the two-dimensional image from the head to the feet of the person and obtains the number of pixels of the height. 24, bone B1 (length L1), bone B51 (length L21), bone B61 (length L31) and bone B71 (length L41), or , bone B1 (length L1), bone B52 (length L22), bone B62 (length L32), and bone B72 (length L42). The length of each bone can be obtained from the coordinates of each keypoint in the two-dimensional image. The height pixel number (h) is calculated by multiplying L1+L21+L31+L41 or L1+L22+L32+L42 by a correction constant. If both values can be calculated, for example, the longer value is used as the number of height pixels. That is, each bone has the longest length in the image when the image is taken from the front, and is displayed to be short when the bone is tilted in the depth direction with respect to the camera. Therefore, the longer bones are more likely to be imaged from the front, and are considered to be closer to the true values. Therefore, it is preferable to choose the longer value.
[0069]
In the example of FIG. 25, bone B1, bone B51 and bone B52, bone B61 and bone B62, bone B71 and bone B72 are detected without overlapping each other. The sums of these bones, L1 + L21 + L31 + L41 and L1 + L22 + L32 + L42, are calculated, and the value obtained by multiplying L1 + L22 + L32 + L42 on the left leg side where the length of the detected bone is longer by a correction constant is taken as the height pixel number.
[0070]
In the example of FIG. 26, bone B1, bone B51 and bone B52, bone B61 and bone B62, bone B71 and bone B72 are respectively detected, and bone B61 and bone B71 of the right leg and bone B62 and bone B72 of the left leg are overlapped. . L1 + L21 + L31 + L41 and L1 + L22 + L32 + L42, which are the sums of these bones, are obtained. For example, the value obtained by multiplying L1 + L21 + L31 + L41 on the right leg side where the length of the detected bone is long by a correction constant is taken as the height pixel number.
[0071]
In the example of FIG. 27, bone B1, bone B51 and bone B52, bone B61 and bone B62, bone B71 and bone B72 are respectively detected, and bone B61 and bone B71 of the right leg and bone B62 and bone B72 of the left leg are overlapped. . The sums of these bones, L1 + L21 + L31 + L41 and L1 + L22 + L32 + L42, are calculated, and the value obtained by multiplying L1 + L22 + L32 + L42 on the left leg side where the length of the detected bone is longer by a correction constant is taken as the height pixel number.
[0072]
In Specific Example 1, since the height can be obtained by totaling the length of the bones from the head to the feet, the number of pixels of the height can be obtained by a simple method. In addition, since it is sufficient to detect at least the skeleton from the head to the feet by the skeleton estimation technology using machine learning, the height pixel count can be accurately calculated even when the whole person is not shown in the image, such as when the person is squatting. can be estimated.
[0073]
In specific example 2, the number of height pixels is obtained using a two-dimensional skeleton model that shows the relationship between the length of the bones included in the two-dimensional skeleton structure and the length of the whole body of a person in the two-dimensional image space.
[0074]
FIG. 28 is a human body model (two-dimensional skeleton model) 301 used in Specific Example 2, showing the relationship between the length of each bone on the two-dimensional image space and the length of the whole body on the two-dimensional image space. As shown in FIG. 28, the relationship between the length of each bone of an average person and the length of the whole body (ratio of the length of each bone to the length of the whole body) is associated with each bone of the human body model 301 . For example, the length of the head bone B1 is the length of the whole body x 0.2 (20%), the length of the right hand bone B41 is the length of the whole body x 0.15 (15%), and the length of the right leg is The length of bone B71 is the length of the whole body×0.25 (25%). By storing such information of the human body model 301 in the database 110, the average length of the whole body can be obtained from the length of each bone. In addition to the human body model of an average person, a human body model may be prepared for each person's attributes such as age, sex, and nationality. As a result, the length of the whole body (height) can be obtained appropriately according to the attributes of the person.
[0075]
In specific example 2, as shown in FIG. 21, the height calculation unit 108 acquires the length of each bone (S221). The height calculator 108 acquires the lengths of all bones (lengths in the two-dimensional image space) in the detected skeletal structure. FIG. 29 shows an example in which a skeletal structure is detected by capturing an image of a squatting person from the right rear oblique direction. In this example, since the face and left side of the person are not shown, the bones of the head, left arm, and left hand cannot be detected. Therefore, the lengths of the detected bones B21, B22, B31, B41, B51, B52, B61, B62, B71, and B72 are obtained.
[0076]
Subsequently, the height calculation unit 108 calculates the number of height pixels from the length of each bone based on the human body model, as shown in FIG. 21 (S222). The height calculator 108 refers to a human body model 301 showing the relationship between each bone and the length of the whole body as shown in FIG. 28, and obtains the number of height pixels from the length of each bone. example For example, since the length of the bone B41 on the right hand is the length of the whole body × 0.15, the number of height pixels based on the bone B41 is obtained from the length of the bone B41 / 0.15. Further, since the length of the bone B71 of the right foot is the length of the whole body × 0.25, the number of height pixels based on the bone B71 is obtained from the length of the bone B71 / 0.25.
[0077]
The human body model referred to at this time is, for example, a human body model of an average person, but a human body model may be selected according to the attributes of the person such as age, gender, and nationality. For example, when a person's face is shown in the captured image, the attribute of the person is identified based on the face, and the human body model corresponding to the identified attribute is referred to. It is possible to recognize a person's attributes from the facial features of the image by referring to the information obtained by machine learning the face for each attribute. Further, when the attribute of the person cannot be identified from the image, the human body model of the average person may be used.
[0078]
Further, the number of height pixels calculated from the length of the bone may be corrected by the camera parameter. For example, when the camera is taken at a high position and looking down at a person, the horizontal length of the shoulder-width bones, etc. is not affected by the depression angle of the camera in the two-dimensional skeletal structure, but the vertical length of the neck-waist bones, etc. The length decreases as the depression angle of the camera increases. Then, the number of height pixels calculated from the horizontal length of the shoulder-width bones and the like tends to be larger than the actual number. Therefore, by utilizing the camera parameters, it is possible to know the angle at which the person is looking down at the camera, and the information on the depression angle can be used to correct the two-dimensional skeleton structure as if it was taken from the front. As a result, the number of height pixels can be calculated more accurately.
[0079]
Subsequently, the height calculation unit 108 calculates the optimum value of the number of height pixels as shown in FIG. 21 (S223). The height calculation unit 108 calculates the optimum value of the number of height pixels from the number of height pixels obtained for each bone. For example, as shown in FIG. 30, a histogram of the number of height pixels obtained for each bone is generated, and a large number of height pixels is selected from the histogram. That is, the number of height pixels longer than the others is selected from the plurality of height pixels obtained based on the plurality of bones. For example, the top 30% is set as a valid value, and in FIG. 30, the number of height pixels by bones B71, B61, and B51 is selected. The average number of selected height pixels may be obtained as the optimum value, or the largest number of height pixels may be used as the optimum value. Since the height is calculated from the length of the bone in the two-dimensional image, the length of the bone is imaged from the front when the bone is not formed from the front, that is, when the bone is tilted in the depth direction when viewed from the camera. It will be shorter than the case. Then, a value having a large number of height pixels is more likely to be imaged from the front than a value having a small number of height pixels, and is a more plausible value. Therefore, a larger value is set as the optimum value.
[0080]
In Specific Example 2, the number of height pixels is calculated based on the detected bones of the skeleton structure using a human body model showing the relationship between the bones in the two-dimensional image space and the length of the whole body, so that all the skeletons from the head to the feet are obtained. Even if is not obtained, the number of height pixels can be obtained from some bones. In particular, the number of height pixels can be estimated accurately by adopting a larger value among the values obtained from a plurality of bones.
[0081]
In Specific Example 3, the two-dimensional skeletal structure is fitted to the three-dimensional human body model (three-dimensional skeletal model), and the skeletal vector of the whole body is obtained using the number of height pixels of the fitted three-dimensional human body model.
[0082]
In Specific Example 3, as shown in FIG. 22, the height calculation unit 108 first calculates the camera parameters based on the image captured by the camera 200 (S231). The height calculation unit 108 extracts an object whose length is known in advance from a plurality of images captured by the camera 200, and obtains a camera parameter from the size (number of pixels) of the extracted object. The camera parameters may be obtained in advance, and the obtained camera parameters may be acquired as needed.
[0083]
Subsequently, the height calculation unit 108 adjusts the arrangement and height of the three-dimensional human body model (S232). The height calculation unit 108 prepares a three-dimensional human body model for calculating the number of height pixels for the detected two-dimensional skeleton structure, and arranges the three-dimensional human body model in the same two-dimensional image based on the camera parameters. Specifically, the "relative positional relationship between the camera and the person in the real world" is specified from the camera parameters and the two-dimensional skeleton structure. For example, assuming that the position of the camera is the coordinates (0, 0, 0), the coordinates (x, y, z) of the position where the person is standing (or sitting) are specified. Then, by assuming an image when the three-dimensional human body model is placed at the same position (x, y, z) as the specified person and captured, the two-dimensional skeleton structure and the three-dimensional human body model are superimposed.
[0084]
FIG. 31 is an example in which a crouching person is imaged diagonally from the front left and the two-dimensional skeleton structure 401 is detected. The two-dimensional skeleton structure 401 has two-dimensional coordinate information. It is preferable that all bones are detected, but some bones may not be detected. For this two-dimensional skeleton structure 401, a three-dimensional human body model 402 as shown in FIG. 32 is prepared. The three-dimensional human body model (three-dimensional skeleton model) 402 is a model of a skeleton having three-dimensional coordinate information and having the same shape as the two-dimensional skeleton structure 401. Then, as shown in FIG. 33, the prepared three-dimensional human body model 402 is arranged and superimposed on the detected two-dimensional skeleton structure 401. In addition, the height of the three-dimensional human body model 402 is adjusted so as to match the two-dimensional skeleton structure 401.
[0085]
The three-dimensional human body model 402 prepared at this time may be a model in a state close to the posture of the two-dimensional skeleton structure 401 as shown in FIG. 33, or may be a model in an upright state. For example, a three-dimensional human body model 402 of the estimated posture may be generated by using a technique of estimating the posture of the three-dimensional space from the two-dimensional image using machine learning. By learning the information of the joints in the two-dimensional image and the joints in the three-dimensional space, the three-dimensional posture can be estimated from the two-dimensional image.
[0086]
Subsequently, the height calculation unit 108 fits the three-dimensional human body model into the two-dimensional skeletal structure as shown in FIG. 22 (S233). As shown in FIG. 34, the height calculation unit 108 superimposes the three-dimensional human body model 402 on the two-dimensional skeleton structure 401 so that the postures of the three-dimensional human body model 402 and the two-dimensional skeleton structure 401 match. The dimensional human body model 402 is transformed. That is, the height, body orientation, and joint angle of the three-dimensional human body model 402 are adjusted and optimized so that there is no difference from the two-dimensional skeletal structure 401. For example, the joints of the three-dimensional human body model 402 are rotated within the movable range of the person, the entire three-dimensional human body model 402 is rotated, and the overall size is adjusted. The fitting of the three-dimensional human body model and the two-dimensional skeleton structure is performed in the two-dimensional space (two-dimensional coordinates). That is, a three-dimensional human body model is mapped in a two-dimensional space, and the three-dimensional human body model is transformed into a two-dimensional skeleton structure in consideration of how the deformed three-dimensional human body model changes in the two-dimensional space (image). Optimize.
[0087]
Subsequently, the height calculation unit 108 calculates the number of height pixels of the fitted three-dimensional human body model as shown in FIG. 22 (S234). As shown in FIG. 35, the height calculation unit 108 obtains the number of height pixels of the three-dimensional human body model 402 in that state when the difference between the three-dimensional human body model 402 and the two-dimensional skeleton structure 401 disappears and the postures match. With the optimized three-dimensional human body model 402 upright, the length of the whole body in the two-dimensional space is obtained based on the camera parameters. For example, the number of height pixels is calculated from the length (number of pixels) of the bones from the head to the foot when the three-dimensional human body model 402 is upright. Similar to Specific Example 1, the lengths of the bones from the head to the foot of the three-dimensional human body model 402 may be totaled.
[0088]
In Specific Example 3, when a three-dimensional human body model is fitted to a two-dimensional skeleton structure based on camera parameters and the number of height pixels is obtained based on the three-dimensional human body model, all bones are not shown in the front. That is, since all the bones are projected diagonally, the number of height pixels can be estimated accurately even when the error is large.
[0089]
As shown in FIG. 19, the image processing device 100 performs a normalization process (S202) following the height pixel number calculation process. In the normalization process, as shown in FIG. 23, the feature amount calculation unit 103 calculates the key point height (S241). The feature amount calculation unit 103 calculates the key point height (number of pixels) of all the key points included in the detected skeleton structure. The key point height is the length (number of pixels) in the height direction from the lowest end of the skeletal structure (for example, the key point of any foot) to the key point. Here, as an example, the height of the key point is obtained from the Y coordinate of the key point in the image. As described above, the key point height may be obtained from the length in the direction along the vertical projection axis based on the camera parameters. For example, in the example of FIG. 24, the height (y i) of the key point A2 of the neck is a value obtained by subtracting the Y coordinate of the key point A81 of the right foot or the key point A82 of the left foot from the Y coordinate of the key point A2.
[0090]
Subsequently, the feature amount calculation unit 103 specifies a reference point for normalization (S242). The reference point is a reference point for expressing the relative height of the key point. The reference point may be preset or may be selectable by the user. The reference point is preferably the center or higher than the center of the skeletal structure (upper and lower in the vertical direction of the image), and for example, the coordinates of the key point of the neck are used as the reference point. The coordinates of the head and other key points, not limited to the neck, may be used as the reference point. Not limited to the key point, any coordinate (for example, the center coordinate of the skeleton structure) may be used as a reference point.
[0091]
Subsequently, the feature amount calculation unit 103 normalizes the key point height (y i) by the number of height pixels (S243). The feature amount calculation unit 103 normalizes each key point by using the key point height, the reference point, and the number of height pixels of each key point. Specifically, the feature amount calculation unit 103 normalizes the relative height of the key point with respect to the reference point by the number of height pixels. Here, as an example focusing only on the height direction, only the Y coordinate is extracted, and the reference point is used as the key point of the neck for normalization. Specifically, the feature amount (normalized value) is obtained by using the following equation (1) with the Y coordinate of the reference point (key point of the neck) as (y c). When using the vertical projection axis based on the camera parameters, (y i) and (y c) are converted into values in the direction along the vertical projection axis.
[Number 1]
[0092]
For example, when there are 18 key points, the coordinates (x 0, y 0), (x 1, y 1), ... (X 17, y 17) of the 18 points of each key point are converted into the above equation (1). ) Is used to convert to 18-dimensional features as follows.
[Number 2]
[0093]
FIG. 36 shows an example of the feature amount of each key point obtained by the feature amount calculation unit 103. In this example, since the key point A2 of the neck is used as the reference point, the feature amount of the key point A2 is 0.0, and the feature amount of the key point A31 on the right shoulder and the key point A32 on the left shoulder at the same height as the neck are also. It is 0.0. The feature amount of the key point A1 of the head higher than the neck is -0.2. The feature amount of the right hand key point A51 and the left hand key point A52 lower than the neck is 0.4, and the feature amount of the right foot key point A81 and the left foot key point A82 is 0.9. When the person raises his / her left hand from this state, the left hand becomes higher than the reference point as shown in FIG. 37, so that the feature amount of the left hand key point A52 is −0.4. On the other hand, since the normalization is performed using only the coordinates of the Y-axis, the feature amount does not change even if the width of the skeleton structure changes as compared with FIG. 36, as shown in FIG. 38. That is, the feature amount (normalized value) of the present embodiment shows the feature of the skeleton structure (key point) in the height direction (Y direction).and is not affected by changes in the lateral direction (X-direction) of the skeletal structure.
[0094]
As described above, in the present embodiment, the skeletal structure of a person is detected from a two-dimensional image, and the number of height pixels (height when standing upright in the two-dimensional image space) obtained from the detected skeletal structure is used to Normalize each keypoint of the skeleton structure. By using this normalized feature amount, it is possible to improve robustness when performing classification, search, and the like. That is, since the feature amount of the present embodiment is not affected by changes in the horizontal direction of the person as described above, it is highly robust against changes in the orientation of the person and the body shape of the person.
[0095]
Furthermore, in the present embodiment, it can be realized by detecting a person's skeleton structure using a skeleton estimation technique such as OpenPose, so there is no need to prepare learning data for learning a person's posture and the like. In addition, by normalizing the key points of the skeletal structure and storing them in a database, it is possible to classify and search for postures of people, etc., so that unknown postures can also be classified and searched. In addition, by normalizing the key points of the skeletal structure, it is possible to obtain clear and easy-to-understand feature quantities, so unlike black-box algorithms such as machine learning, users are highly satisfied with the processing results.
[0096]
Each configuration in the above-described embodiments is configured by hardware or software, or both, and may be configured by one piece of hardware or software, or may be configured by multiple pieces of hardware or software. The functions (processing) of the image processing apparatuses 10 and 100 may be realized by a computer 20 having a processor 21 such as a CPU (Central Processing Unit) and a memory 22 as a storage device, as shown in FIG. For example, a program (image processing program) for performing the method in the embodiment may be stored in the memory 22 and each function may be realized by executing the program stored in the memory 22 with the processor 21 .
[0097]
These programs can be stored and supplied to computers using various types of non-transitory computer readable media. Non-transitory computer-readable media include various types of tangible storage media. Examples of non-transitory computer-readable media include magnetic recording media (eg, flexible discs, magnetic tapes, hard disk drives), magneto-optical recording media (eg, magneto-optical discs), CD-ROMs (Read Only Memory), CD-Rs, CD-R/W, semiconductor memory (eg, mask ROM, PROM (Programmable ROM), EPROM (Erasable PROM), flash ROM, RAM (random access memory)). The program may also be delivered to the computer on various types of transitory computer readable medium. Examples of transitory computer-readable media include electrical signals, optical signals, and electromagnetic waves. Transitory computer-readable media can deliver the program to the computer via wired channels, such as wires and optical fibers, or wireless channels.
[0098]
In addition, the present disclosure is not limited to the above embodiments, and can be modified as appropriate without departing from the scope. For example, although the skeletal structure of a person is detected in the above description, animals other than humans (mammals, reptiles, birds, amphibians, fish, etc.) having a skeletal structure may also be detected.
[0099]
Although the present disclosure has been described above with reference to the embodiments, the present disclosure is not limited to the above embodiments. Various changes that can be understood by those skilled in the art can be made to the configuration and details of the present disclosure within the scope of the present disclosure.
[0100]
Some or all of the above embodiments can also be described as the following additional remarks, but are not limited to the following.
(Appendix 1)
a skeleton detection means for detecting the two-dimensional skeleton structure of a plurality of persons based on the acquired two-dimensional image;
a feature quantity calculation means for calculating the feature quantities of the plurality of detected two-dimensional skeletal structures;
a recognition means for recognizing the states of the plurality of persons based on the calculated similarities of the plurality of feature quantities;
An image processing device equipped with
(Appendix 2)
The recognition means classifies the states of the plurality of persons as the recognition processing,
The image processing device described in appendix 1.
(Appendix 3)
The recognition means classifies the states of the plurality of persons based on the feature amount of the whole or a part of the two-dimensional skeletal structure,
The image processing device described in appendix 2.
(Appendix 4)
The recognition means classifies the states of the plurality of persons based on the feature quantity of the first portion of the two-dimensional skeleton structure and the feature quantity of the second portion of the two-dimensional skeleton structure,
The image processing device described in appendix 2.
(Appendix 5)
The skeleton detection means detects the two-dimensional skeleton structure of the plurality of persons from a plurality of two-dimensional images that are continuous in time series,
The recognition means classifies the states of the plurality of persons based on the feature amount of the 2D skeletal structure detected from the plurality of 2D images,
The image processing device according to any one of Appendices 2 to 4.
(Appendix 6)
The recognition means classifies the states of the plurality of persons based on changes in the feature amount of the 2D skeletal structure in the plurality of 2D images,
The image processing device described in appendix 5.
(Appendix 7)
The recognition means classifies the two-dimensional skeletal structure of the person whose right and left sides are opposite as the same state,
The image processing device according to any one of Appendices 2 to 6.
(Appendix 8)
The recognition means displays the results of classification of the states of the plurality of persons according to the degree of similarity,
The image processing device according to any one of Appendices 2 to 7.
(Appendix 9)
The recognition means searches for the state of the person selected from the displayed classification results from the states of the plurality of persons,
The image processing device described in appendix 8.
(Appendix 10)
The recognition means retrieves the query state from the states of the plurality of persons as the recognition processing,
The image processing device described in appendix 1.
(Appendix 11)
The recognition means retrieves the query state based on the feature amount of the whole or a part of the two-dimensional skeleton structure,
The image processing device according to appendix 10.
(Appendix 12)
The recognition means retrieves the query state based on the feature quantity of the first part of the two-dimensional skeleton structure and the feature quantity of the second part of the two-dimensional skeleton structure,
The image processing device according to appendix 10.
(Appendix 13)
The recognition means retrieves the query state based on the weight of the feature amount of the first part and the feature amount of the second part,
The image processing device according to appendix 12.
(Appendix 14)
The skeleton detection means detects the two-dimensional skeleton structure of the plurality of persons from a plurality of two-dimensional images that are continuous in time series,
The recognition means retrieves the query state based on the feature amount of the two-dimensional skeleton structure detected from the plurality of two-dimensional images,
The image processing device according to any one of Appendices 10 to 13.
(Appendix 15)
The recognition means retrieves the query state based on changes in the feature amount of the two-dimensional skeleton structure in the plurality of two-dimensional images,
The image processing device described in appendix 14.
(Appendix 16)
The recognition means retrieves the two-dimensional skeletal structure in which the right side and the left side of the person are in the same state,
The image processing device according to any one of Appendices 10 to 15.
(Appendix 17)
The recognition means displays search results of the states of the plurality of persons according to the degree of similarity,
The image processing device according to any one of Appendices 10 to 16.
(Appendix 18)
The feature amount is a feature amount only in the height direction of the person,
The image processing device according to any one of Appendices 1 to 17.
(Appendix 19)
The feature amount is a feature amount obtained by normalizing the two-dimensional skeletal structure based on the height of the person standing upright in the two-dimensional image space,
The image processing device described in appendix 18.
(Appendix 20)
Based on the acquired 2D images, the 2D skeletal structures of multiple people are detected,
Calculating the feature values of the detected multiple 2D skeletal structures,
According to the calculated similarities of the plurality of feature quantities, perform recognition processing of the states of the plurality of persons,
Image processing method.
(Appendix 21)
As the recognition process, the states of the plurality of people are classified,
The image processing method described in appendix 20.
(Appendix 22)
As the recognition process, the query state is searched from the states of the plurality of persons,
The image processing method described in appendix 20.
(Appendix 23)
Based on the acquired 2D images, the 2D skeletal structures of multiple people are detected,
Calculating the feature values of the detected multiple 2D skeletal structures,
According to the calculated similarities of the plurality of feature quantities, perform recognition processing of the states of the plurality of persons,
An image processing program that allows a computer to execute processing.
(Appendix 24)
As the recognition process, the states of the plurality of people are classified,
The image processing program described in appendix 23.
(Appendix 25)
As the recognition process, the query state is searched from the states of the plurality of persons,
The image processing program described in appendix 23.
Code description
[0101]
1 image processing system
10 image processing device
11 skeleton detection unit
12 Feature value calculation unit
13 Recognition part
20 computer
21 Processor
22 memory
100 Image processing device
101 image acquisition unit
102 Skeletal structure detection unit
103 feature quantity calculation unit
104 Classification section
105 search unit
106 input unit
107 display unit
108 height calculation unit
110 database
200 camera
300, 301 human body model
401 Two-dimensional skeleton structure
402 Three-dimensional human body model
The scope of the claims
[Claim 1]
a skeleton detection means for detecting the two-dimensional skeleton structure of a plurality of persons based on the acquired two-dimensional image;
a feature quantity calculation means for calculating the feature quantities of the plurality of detected two-dimensional skeletal structures;
a recognition means for recognizing the states of the plurality of persons based on the calculated similarities of the plurality of feature quantities;
An image processing device equipped with
[Claim 2]
The recognition means classifies the states of the plurality of persons as the recognition processing,
The image processing device according to claim 1.
[Claim 3]
The recognition means classifies the states of the plurality of persons based on the feature amount of the whole or a part of the two-dimensional skeletal structure,
The image processing device according to claim 2.
[Claim 4]
The recognition means classifies the states of the plurality of persons based on the feature quantity of the first portion of the two-dimensional skeleton structure and the feature quantity of the second portion of the two-dimensional skeleton structure,
The image processing device according to claim 2.
[Claim 5]
The skeleton detection means detects the two-dimensional skeleton structure of the plurality of persons from a plurality of two-dimensional images that are continuous in time series,
The recognition means classifies the states of the plurality of persons based on the feature amount of the 2D skeletal structure detected from the plurality of 2D images,
The image processing device according to any one of claims 2 to 4.
[Claim 6]
The recognition means classifies the states of the plurality of persons based on changes in the feature amount of the 2D skeletal structure in the plurality of 2D images,
The image processing device according to claim 5.
[Claim 7]
The recognition means classifies the two-dimensional skeletal structure of the person whose right and left sides are opposite as the same state,
The image processing device according to any one of claims 2 to 6.
[Claim 8]
The recognition means displays the results of classification of the states of the plurality of persons according to the degree of similarity,
The image processing apparatus according to any one of claims 2 to 7.
[Claim 9]
The recognition means searches for the state of the person selected from the displayed classification results from the states of the plurality of persons,
The image processing device according to claim 8.
[Claim 10]
The recognition means retrieves the query state from the states of the plurality of persons as the recognition processing,
The image processing device according to claim 1.
[Claim Item 11]
The recognition means retrieves the query state based on the feature amount of the whole or a part of the two-dimensional skeleton structure,
The image processing device according to claim 10.
[Claim 12]
The recognition means retrieves the query state based on the feature quantity of the first part of the two-dimensional skeleton structure and the feature quantity of the second part of the two-dimensional skeleton structure,
The image processing device according to claim 10.
[Claim 13]
The recognition means retrieves the query state based on the weight of the feature amount of the first part and the feature amount of the second part,
The image processing device according to claim 12.
[Claim 14]
The skeleton detection means detects the two-dimensional skeleton structure of the plurality of persons from a plurality of two-dimensional images that are continuous in time series,
The recognition means retrieves the query state based on the feature amount of the two-dimensional skeleton structure detected from the plurality of two-dimensional images,
The image processing apparatus according to any one of claims 10 to 13.
[Claim 15]
The recognition means retrieves the query state based on changes in the feature amount of the two-dimensional skeleton structure in the plurality of two-dimensional images,
The image processing device according to claim 14.
[Claim 16]
The recognition means retrieves the two-dimensional skeletal structure in which the right side and the left side of the person are in the same state,
The image processing apparatus according to any one of claims 10 to 15.
[Claim 17]
The recognition means displays search results of the states of the plurality of persons according to the degree of similarity,
The image processing apparatus according to any one of claims 10 to 16.
[Claim 18]
The feature amount is a feature amount only in the height direction of the person,
The image processing apparatus according to any one of claims 1 to 17.
[Claim 19]
The feature amount is a feature amount obtained by normalizing the two-dimensional skeletal structure based on the height of the person standing upright in the two-dimensional image space,
The image processing device according to claim 18.
[Claim 20]
Based on the acquired 2D images, the 2D skeletal structures of multiple people are detected,
Calculating the feature values of the detected multiple 2D skeletal structures,
According to the calculated similarities of the plurality of feature quantities, perform recognition processing of the states of the plurality of persons,
Image processing method.
[Claim 21]
As the recognition process, the states of the plurality of people are classified,
The image processing method according to claim 20.
[Claim 22]
As the recognition process, the query state is searched from the states of the plurality of persons,
The image processing method according to claim 20.
[Claim 23]
Based on the acquired 2D images, the 2D skeletal structures of multiple people are detected,
Calculating the feature values of the detected multiple 2D skeletal structures,
According to the calculated similarities of the plurality of feature quantities, perform recognition processing of the states of the plurality of persons,
A non-transitory computer-readable medium that stores an image processing program for causing a computer to execute processing.
[Claim 24]
As the recognition process, the states of the plurality of people are classified,
The non-transitory computer-readable medium according to claim 23.
[Claim 25]
As the recognition process, the query state is searched from the states of the plurality of persons,
The non-transitory computer-readable medium according to claim 23.
| # | Name | Date |
|---|---|---|
| 1 | 202217025223.pdf | 2022-04-29 |
| 2 | 202217025223-TRANSLATIOIN OF PRIOIRTY DOCUMENTS ETC. [29-04-2022(online)].pdf | 2022-04-29 |
| 3 | 202217025223-STATEMENT OF UNDERTAKING (FORM 3) [29-04-2022(online)].pdf | 2022-04-29 |
| 4 | 202217025223-REQUEST FOR EXAMINATION (FORM-18) [29-04-2022(online)].pdf | 2022-04-29 |
| 5 | 202217025223-POWER OF AUTHORITY [29-04-2022(online)].pdf | 2022-04-29 |
| 6 | 202217025223-FORM 18 [29-04-2022(online)].pdf | 2022-04-29 |
| 7 | 202217025223-FORM 1 [29-04-2022(online)].pdf | 2022-04-29 |
| 8 | 202217025223-DRAWINGS [29-04-2022(online)].pdf | 2022-04-29 |
| 9 | 202217025223-DECLARATION OF INVENTORSHIP (FORM 5) [29-04-2022(online)].pdf | 2022-04-29 |
| 10 | 202217025223-COMPLETE SPECIFICATION [29-04-2022(online)].pdf | 2022-04-29 |
| 11 | 202217025223-FER.pdf | 2022-09-08 |
| 12 | 202217025223-Proof of Right [09-09-2022(online)].pdf | 2022-09-09 |
| 13 | 202217025223-FORM 3 [20-09-2022(online)].pdf | 2022-09-20 |
| 14 | 202217025223-FORM 3 [21-02-2023(online)].pdf | 2023-02-21 |
| 15 | 202217025223-FORM-26 [06-03-2023(online)].pdf | 2023-03-06 |
| 16 | 202217025223-FER_SER_REPLY [06-03-2023(online)].pdf | 2023-03-06 |
| 17 | 202217025223-DRAWING [06-03-2023(online)].pdf | 2023-03-06 |
| 18 | 202217025223-CORRESPONDENCE [06-03-2023(online)].pdf | 2023-03-06 |
| 19 | 202217025223-COMPLETE SPECIFICATION [06-03-2023(online)].pdf | 2023-03-06 |
| 20 | 202217025223-CLAIMS [06-03-2023(online)].pdf | 2023-03-06 |
| 21 | 202217025223-ABSTRACT [06-03-2023(online)].pdf | 2023-03-06 |
| 22 | 202217025223-Others-060323.pdf | 2023-03-13 |
| 23 | 202217025223-GPA-060323.pdf | 2023-03-13 |
| 24 | 202217025223-Form-5-060323.pdf | 2023-03-13 |
| 25 | 202217025223-Correspondence-060323.pdf | 2023-03-13 |
| 1 | skeltondetectionmE_08-09-2022.pdf |