Abstract: This image-capturing device is provided with: an image-capturing part (10) which includes a pixel area in which a plurality of pixels are arrayed, reads out pixel signals from the pixels included in the pixel area, and outputs the pixel signals; a readout unit control part (123) which controls a readout unit which is set as a part of the pixel area; a recognition part (12) which learns teacher data for each readout unit; a first output part (124) which outputs a recognition result recognized for each readout unit by the recognition part; a second output part (143) which outputs the pixel signals to a rear stage; and a trigger generation part (16a) which generates a trigger signal for controlling a first timing at which the first output part outputs the recognition result and a second timing at which the second output part outputs the pixel signals.
Title of the invention: Imaging apparatus, imaging system, imaging method and imaging program
Technical field
[0001]
The present disclosure relates to an imaging apparatus, an imaging system, an imaging method and an imaging program.
Background technology
[0002]
In recent years, with the increasing performance of image pickup devices such as digital still cameras, digital video cameras, and small cameras mounted on multifunctional mobile phones (smartphones), an image recognition function that recognizes a predetermined object included in a captured image. An imaging device equipped with the above has been developed.
Prior art literature
Patent documents
[0003]
Patent Document 1: Japanese Unexamined Patent Publication No. 2017-11249
Outline of the invention
Problems to be solved by the invention
[0004]
In general, an image suitable for recognition processing in the image recognition function is different from an image suitable for human visual recognition. Therefore, when trying to improve the recognition accuracy in an image pickup device equipped with an image recognition function, it may be difficult for the image captured for the recognition process to sufficiently provide information as a visual image. There is.
[0005]
An object of the present disclosure is to provide an imaging device, an imaging system, an imaging method, and an imaging program capable of achieving both imaging for recognition processing and imaging for visual recognition.
Means to solve problems
[0006]
The imaging apparatus according to the present disclosure has a pixel region in which a plurality of pixels are arranged, an imaging unit that reads and outputs a pixel signal from the pixels included in the pixel region, and a reading unit set as a part of the pixel region. A read unit control unit that controls the above, a recognition unit that learns teacher data for each read unit, a first output unit that outputs the recognition result recognized by the recognition unit for each read unit, and a pixel signal is output to the subsequent stage. A trigger signal is generated to control the second output unit, the first timing at which the first output unit outputs the recognition result, and the second timing at which the second output unit outputs the pixel signal. A trigger generation unit and a trigger generation unit are provided.
A brief description of the drawing
[0007]
FIG. 1 is a block diagram showing a configuration of an example of an imaging device applicable to each embodiment of the present disclosure.
FIG. 2A is a schematic diagram showing an example of a hardware configuration of an imaging device according to each embodiment.
FIG. 2B is a schematic diagram showing an example of a hardware configuration of an imaging device according to each embodiment.
FIG. 3A is a diagram showing an example in which an imaging device according to each embodiment is formed by a two-layer structure laminated CIS.
FIG. 3B is a diagram showing an example in which an imaging device according to each embodiment is formed by a laminated CIS having a three-layer structure.
FIG. 4 is a block diagram showing a configuration of an example of a sensor unit applicable to each embodiment.
FIG. 5A is a schematic diagram for explaining a rolling shutter method.
[Fig. 5B] Fig. 5B is a schematic diagram for explaining a rolling shutter method.
[Fig. 5C] Fig. 5C is a schematic diagram for explaining a rolling shutter method.
FIG. 6A is a schematic diagram for explaining line thinning in the rolling shutter method.
[Fig. 6B] Fig. 6B is a schematic diagram for explaining line thinning in the rolling shutter method.
[Fig. 6C] Fig. 6C is a schematic diagram for explaining line thinning in the rolling shutter method.
FIG. 7A is a diagram schematically showing an example of another imaging method in the rolling shutter method.
FIG. 7B is a diagram schematically showing an example of another imaging method in the rolling shutter method.
[Fig. 8A] Fig. 8A is a schematic diagram for explaining a global shutter method.
[Fig. 8B] Fig. 8B is a schematic diagram for explaining a global shutter method.
[Fig. 8C] Fig. 8C is a schematic diagram for explaining a global shutter method.
[Fig. 9A] Fig. 9A is a diagram schematically showing an example of a sampling pattern that can be realized in the global shutter method.
[Fig. 9B] Fig. 9B is a diagram schematically showing an example of a sampling pattern that can be realized in the global shutter method.
FIG. 10 is a diagram for schematically explaining an image recognition process by CNN.
[Fig. 11] Fig. 11 is a diagram for schematically explaining an image recognition process for obtaining a recognition result from a part of an image to be recognized.
FIG. 12A is a diagram schematically showing an example of identification processing by DNN when time series information is not used.
FIG. 12B is a diagram schematically showing an example of identification processing by DNN when time series information is not used.
FIG. 13A is a diagram schematically showing a first example of identification processing by DNN when time series information is used.
FIG. 13B is a diagram schematically showing a first example of identification processing by DNN when time series information is used.
FIG. 14A is a diagram schematically showing a second example of identification processing by DNN when time series information is used.
FIG. 14B is a diagram schematically showing a second example of identification processing by DNN when time series information is used.
[Fig. 15A] Fig. 15A is a diagram for explaining a relationship between a frame driving speed and a pixel signal reading amount.
[Fig. 15B] Fig. 15B is a diagram for explaining a relationship between a frame driving speed and a pixel signal reading amount.
FIG. 16 is a schematic diagram for schematically explaining the recognition process according to each embodiment of the present disclosure.
FIG. 17 is a flowchart showing an example of recognition processing by the recognition processing unit according to the first embodiment.
FIG. 18 is a diagram showing an example of image data for one frame.
FIG. 19 is a diagram for explaining a flow of machine learning processing executed by the recognition processing unit according to the first embodiment.
FIG. 20A is a schematic diagram for explaining an application example of the first embodiment.
FIG. 20B is a schematic diagram for explaining an application example of the first embodiment.
FIG. 21 is a functional block diagram of an example for explaining the function of the image pickup apparatus according to the second embodiment.
FIG. 22 is a schematic diagram showing in more detail an example of processing in the recognition processing unit according to the second embodiment.
FIG. 23 is a functional block diagram of an example for explaining a function according to a second embodiment.
FIG. 24 is a schematic diagram for explaining a frame reading process according to the second embodiment.
FIG. 25 is a schematic diagram schematically showing a recognition process according to a second embodiment.
[Fig. 26] Fig. 26 is a diagram for explaining an example in which the recognition process is terminated in the middle of reading a frame.
[Fig. 27] Fig. 27 is a diagram for explaining an example in which the recognition process is terminated in the middle of reading a frame.
FIG. 28 is an example flowchart showing a recognition process according to a second embodiment.
FIG. 29A is an example time chart showing an example of control of reading and recognition processing according to the second embodiment.
FIG. 29B is an example time chart showing an example of control of reading and recognition processing according to the second embodiment.
FIG. 30 is an example time chart showing another example of control of reading and recognition processing according to the second embodiment.
FIG. 31 is a flowchart showing an example of control according to a third embodiment.
FIG. 32 is a schematic diagram schematically showing an example of output control processing according to the third embodiment.
FIG. 33A is a functional block diagram showing a function of an example on the recognition processing unit side of the image pickup apparatus according to the third embodiment.
FIG. 33B is a functional block diagram showing a function of an example on the visual processing unit side of the image pickup apparatus according to the third embodiment.
FIG. 34 is a flowchart of an example showing a process when a trigger signal is output according to a time according to a third embodiment.
FIG. 35 is a schematic diagram schematically showing an example of output control processing according to the first modification of the third embodiment.
FIG. 36A is a functional block diagram showing a function of an example on the recognition processing unit side of the image pickup apparatus according to the first modification of the third embodiment.
FIG. 36B is a functional block diagram showing a function of an example on the visual processing unit side of an image pickup apparatus according to a first modification of the third embodiment.
FIG. 37 is a flowchart of an example showing a process according to a first modification of the third embodiment.
FIG. 38 is a schematic diagram schematically showing an example of output control processing according to a second modification of the third embodiment.
FIG. 39A is a functional block diagram showing a function of an example on the recognition processing unit side of the image pickup apparatus according to the second modification of the third embodiment.
FIG. 39B is a functional block diagram showing a function of an example on the visual processing unit side of an image pickup apparatus according to a second modification of the third embodiment.
FIG. 40 is a flowchart of an example showing a process according to a second modification of the third embodiment.
FIG. 41A is a functional block diagram showing a function of an example on the recognition processing unit side of an image pickup apparatus according to a third modification of the third embodiment.
FIG. 41B is a functional block diagram showing a function of an example on the visual processing unit side of an image pickup apparatus according to a third modification of the third embodiment.
FIG. 42 is a flowchart of an example showing a process according to a third modification of the third embodiment.
FIG. 43 is a schematic diagram schematically showing an example of output control processing according to the fourth embodiment.
FIG. 44 is a functional block diagram showing a function of an example of an imaging device according to a fourth embodiment.
FIG. 45 is an example flowchart showing a process according to a fourth embodiment.
FIG. 46 is a functional block diagram showing a function of an example of an imaging device according to a first modification of the fourth embodiment.
FIG. 47 is a flowchart of an example showing a process according to a first modification of the fourth embodiment.
FIG. 48 is a schematic diagram schematically showing an example of output control processing according to a second modification of the fourth embodiment.
FIG. 49 is a functional block diagram showing a function of an example of the image pickup apparatus 1 according to a second modification of the fourth embodiment.
FIG. 50 is a flowchart of an example showing a process according to a first modification of the fourth embodiment.
FIG. 51 is a flowchart of an example for explaining an outline of the arbitration process according to the fifth embodiment.
FIG. 52 is a functional block diagram of an example for explaining the function of the image pickup apparatus 1 applicable to the fifth embodiment.
FIG. 53 is a schematic diagram for explaining the arbitration process according to the fifth embodiment.
FIG. 54 is an example flowchart showing an arbitration process according to a fifth embodiment.
FIG. 55 is a functional block diagram of an example for explaining the function of the image pickup apparatus applicable to the first modification of the fifth embodiment.
FIG. 56 is a schematic diagram for explaining a first example of arbitration processing according to a first modification of the fifth embodiment.
FIG. 57 is a schematic diagram for explaining a second example of the arbitration process according to the first modification of the fifth embodiment.
FIG. 58 is a flowchart of an example showing an arbitration process according to a first modification of the fifth embodiment.
FIG. 59 is a functional block diagram of an example for explaining the function of the image pickup apparatus applicable to the second modification of the fifth embodiment.
FIG. 60 is a schematic diagram for explaining the arbitration process according to the second modification of the fifth embodiment.
FIG. 61 is a flowchart of an example showing an arbitration process according to a second modification of the fifth embodiment.
FIG. 62 is a functional block diagram of an example for explaining the function of the image pickup apparatus 1 applicable to the third modification of the fifth embodiment.
FIG. 63 is a flowchart of an example showing an arbitration process according to a third modification of the fifth embodiment.
[Fig. 64] Fig. 64 is a diagram illustrating a usage example of an image pickup apparatus to which the technique of the present disclosure is applied.
FIG. 65 is a block diagram showing an example of a schematic configuration of a vehicle control system.
[Fig. 66] Fig. 66 is an explanatory view showing an example of installation positions of an external information detection unit and an imaging unit.
Mode for carrying out the invention
[0008]
Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. In the following embodiments, the same parts are designated by the same reference numerals, so that duplicate description will be omitted.
[0009]
Hereinafter, embodiments of the present disclosure will be described in the following order.
1. 1. Configuration example according to each embodiment of the present disclosure
2. Examples of existing technologies applicable to the present disclosure
2-1. Outline of rolling shutter
2-2. Overview of global shutter
2-3. About DNN (Deep Neural Network)
2-3-1. Outline of CNN (Convolutional Neural Network)
2-3-2. Outline of RNN (Recurrent Neural Network)
2-4. Drive speed
3. Outline of the present disclosure
4. First Embodiment
4-1. Operation example by the recognition processing unit
4-2. Specific example of operation by the recognition processing unit
4-3. Application example of the first embodiment
5. Second Embodiment
5-0-1. Configuration Example
5-0-2. Example of processing in the recognition processing unit according to the second embodiment
5-0-3. Details of recognition processing according to the second embodiment
5-0-4. 2. Example of control of reading and recognition processing according to the second embodiment
. Third Embodiment
6-0. Outline of the third embodiment
6-0-1. Example of trigger signal output by time
6-1. First modification of the third embodiment
6-2. Second modification of the third embodiment
6-3. A third modification of the third embodiment
7. Fourth Embodiment
7-1. First modification of the fourth embodiment
7-2. 2. Second modification of the fourth embodiment
8. Fifth Embodiment
8-0-1. Specific examples of arbitration processing
8-1. First modification of the fifth embodiment
8-2. Second modification of the fifth embodiment
8-3. A third modification of the fifth embodiment
9. Sixth Embodiment
[0010]
[1. Configuration Examples of Each Embodiment of the
Disclosure ] The configuration of the imaging apparatus according to the present disclosure will be schematically described. FIG. 1 is a block diagram showing a configuration of an example of an imaging device applicable to each embodiment of the present disclosure. In FIG. 1, the image pickup apparatus 1 includes a sensor unit 10, a sensor control unit 11, a recognition processing unit 12, a memory 13, a visual recognition processing unit 14, and an output control unit 15, and each of these units is CMOS ( It is a CMOS image sensor (CIS) integrally formed by using Complementary Metal Oxide Semiconductor. The image pickup device 1 is not limited to this example, and may be another type of optical sensor such as an infrared light sensor that performs imaging with infrared light.
[0011]
The sensor unit 10 outputs a pixel signal corresponding to the light emitted to the light receiving surface via the optical unit 30. More specifically, the sensor unit 10 has a pixel array in which pixels including at least one photoelectric conversion element are arranged in a matrix. A light receiving surface is formed by each pixel arranged in a matrix in the pixel array. Further, the sensor unit 10 further performs a drive circuit for driving each pixel included in the pixel array and a signal that performs predetermined signal processing on the signal read from each pixel and outputs the signal as a pixel signal of each pixel. Includes processing circuits. The sensor unit 10 outputs the pixel signal of each pixel included in the pixel area as digital image data.
[0012]
Hereinafter, in the pixel array included in the sensor unit 10, a region in which effective pixels for generating a pixel signal are arranged is referred to as a frame. Frame image data is formed by pixel data based on each pixel signal output from each pixel included in the frame. Further, each line in the pixel array of the sensor unit 10 is called a line, and line image data is formed by pixel data based on pixel signals output from each pixel included in the line. Further, the operation in which the sensor unit 10 outputs a pixel signal corresponding to the light applied to the light receiving surface is called imaging. The sensor unit 10 controls the exposure at the time of imaging and the gain (analog gain) with respect to the pixel signal according to the imaging control signal supplied from the sensor control unit 11 described later.
[0013]
The sensor control unit 11 is configured by, for example, a microprocessor, controls the reading of pixel data from the sensor unit 10, and outputs pixel data based on each pixel signal read from each pixel included in the frame. The pixel data output from the sensor control unit 11 is passed to the recognition processing unit 12 and the visual recognition processing unit 14.
[0014]
Further, the sensor control unit 11 generates an image pickup control signal for controlling the image pickup in the sensor unit 10. The sensor control unit 11 generates an imaging control signal, for example, in accordance with instructions from the recognition processing unit 12 and the visual recognition processing unit 14, which will be described later. The image pickup control signal includes the above-mentioned information indicating the exposure and analog gain at the time of image pickup in the sensor unit 10. The image pickup control signal further includes a control signal (vertical synchronization signal, horizontal synchronization signal, etc.) used by the sensor unit 10 to perform an image pickup operation. The sensor control unit 11 supplies the generated image pickup control signal to the sensor unit 10.
[0015]
The optical unit 30 is for irradiating the light receiving surface of the sensor unit 10 with light from the subject, and is arranged at a position corresponding to, for example, the sensor unit 10. The optical unit 30 includes, for example, a plurality of lenses, a diaphragm mechanism for adjusting the size of the opening with respect to the incident light, and a focus mechanism for adjusting the focus of the light applied to the light receiving surface. The optical unit 30 may further include a shutter mechanism (mechanical shutter) that adjusts the time for irradiating the light receiving surface with light. The aperture mechanism, focus mechanism, and shutter mechanism of the optical unit 30 can be controlled by, for example, the sensor control unit 11. Not limited to this, the aperture and focus in the optical unit 30 can be controlled from the outside of the image pickup apparatus 1. It is also possible to integrally configure the optical unit 30 with the image pickup apparatus 1.
[0016]
The recognition processing unit 12 performs recognition processing of an object included in the image based on the pixel data based on the pixel data passed from the sensor control unit 11. In the present disclosure, for example, a DSP (Digital Signal Processor) reads and executes a program that is pre-learned from teacher data and stored as a learning model in the memory 13, and thereby performs recognition processing using DNN (Deep Neural Network). The recognition processing unit 12 as a machine learning unit is configured. The recognition processing unit 12 can instruct the sensor control unit 11 to read the pixel data required for the recognition process from the sensor unit 10. The recognition result by the recognition processing unit 12 is passed to the output control unit 15.
[0017]
The visual recognition processing unit 14 executes processing for obtaining an image suitable for human recognition with respect to the pixel data passed from the sensor control unit 11, and outputs, for example, image data composed of a set of pixel data. To do. For example, the visual processing unit 14 is configured by the ISP (Image Signal Processor) reading and executing a program stored in advance in a memory (not shown).
[0018]
For example, when the visual recognition processing unit 14 is provided with a color filter for each pixel included in the sensor unit 10 and the pixel data has R (red), G (green), and B (blue) color information, demosaic. Processing, white balance processing, etc. can be executed. Further, the visual recognition processing unit 14 can instruct the sensor control unit 11 to read the pixel data required for the visual recognition processing from the sensor unit 10. The image data whose pixel data has been image-processed by the visual recognition processing unit 14 is passed to the output control unit 15.
[0019]
The output control unit 15 is composed of, for example, a microprocessor, and captures one or both of the recognition result passed from the recognition processing unit 12 and the image data passed as the visual recognition processing result from the visual recognition processing unit 14. Output to the outside of 1. The output control unit 15 can output image data to, for example, a display unit 31 having a display device. As a result, the user can visually recognize the image data displayed by the display unit 31. The display unit 31 may be built in the image pickup device 1 or may have an external configuration of the image pickup device 1.
[0020]
2A and 2B are schematic views showing an example of the hardware configuration of the image pickup apparatus 1 according to each embodiment. In FIG. 2A, the sensor unit 10, the sensor control unit 11, the recognition processing unit 12, the memory 13, the visual recognition processing unit 14, and the output control unit 15 are mounted on one chip 2 in the configuration shown in FIG. This is an example. In FIG. 2A, the memory 13 and the output control unit 15 are omitted in order to avoid complication.
[0021]
In the configuration shown in FIG. 2A, the recognition result by the recognition processing unit 12 is output to the outside of the chip 2 via an output control unit 15 (not shown). Further, in the configuration of FIG. 2A, the recognition processing unit 12 can acquire pixel data for use in recognition from the sensor control unit 11 via the internal interface of the chip 2.
[0022]
In FIG. 2B, the sensor unit 10, the sensor control unit 11, the visual recognition processing unit 14, and the output control unit 15 are mounted on one chip 2 in the configuration shown in FIG. 1, and the recognition processing unit 12 and the memory 13 ( (Not shown) is an example of being placed outside the chip 2. Also in FIG. 2B, similarly to FIG. 2A described above, the memory 13 and the output control unit 15 are omitted in order to avoid complication.
[0023]
In the configuration of FIG. 2B, the recognition processing unit 12 acquires pixel data to be used for recognition via an interface for communicating between chips. Further, in FIG. 2B, the recognition result by the recognition processing unit 12 is shown to be directly output to the outside from the recognition processing unit 12, but this is not limited to this example. That is, in the configuration of FIG. 2B, the recognition processing unit 12 may return the recognition result to the chip 2 and output it from the output control unit 15 (not shown) mounted on the chip 2.
[0024]
In the configuration shown in FIG. 2A, the recognition processing unit 12 is mounted on the chip 2 together with the sensor control unit 11, and communication between the recognition processing unit 12 and the sensor control unit 11 can be executed at high speed by the internal interface of the chip 2. .. On the other hand, in the configuration shown in FIG. 2A, the recognition processing unit 12 cannot be replaced, and it is difficult to change the recognition processing. On the other hand, in the configuration shown in FIG. 2B, since the recognition processing unit 12 is provided outside the chip 2, communication between the recognition processing unit 12 and the sensor control unit 11 is performed via the interface between the chips. There is a need. Therefore, the communication between the recognition processing unit 12 and the sensor control unit 11 is slower than the configuration of FIG. 2A, and there is a possibility that a delay may occur in the control. On the other hand, the recognition processing unit 12 can be easily replaced, and various recognition processes can be realized.
[0025]
Hereinafter, unless otherwise specified, the image pickup apparatus 1 has a sensor unit 10, a sensor control unit 11, a recognition processing unit 12, a memory 13, a visual recognition processing unit 14, and an output control unit 15 mounted on one chip 2 in FIG. 2A. The configuration to be used shall be adopted.
[0026]
In the configuration shown in FIG. 2A described above, the image pickup apparatus 1 can be formed on one substrate. Not limited to this, the image pickup apparatus 1 may be a laminated CIS in which a plurality of semiconductor chips are laminated and integrally formed.
[0027]
As an example, the image pickup apparatus 1 can be formed by a two-layer structure in which semiconductor chips are laminated in two layers. FIG. 3A is a diagram showing an example in which the image pickup apparatus 1 according to each embodiment is formed by a two-layer structure laminated CIS. In the structure of FIG. 3A, the pixel portion 20a is formed on the semiconductor chip of the first layer, and the memory + logic portion 20b is formed on the semiconductor chip of the second layer. The pixel unit 20a includes at least the pixel array in the sensor unit 10. The memory + logic unit 20b includes, for example, a sensor control unit 11, a recognition processing unit 12, a memory 13, a visual recognition processing unit 14, and an output control unit 15, and an interface for communicating between the image pickup device 1 and the outside. .. The memory + logic unit 20b further includes a part or all of the drive circuit for driving the pixel array in the sensor unit 10. Further, although not shown, the memory + logic unit 20b can further include, for example, a memory used by the visual recognition processing unit 14 for processing image data.
[0028]
As shown on the right side of FIG. 3A, the image pickup device 1 is configured as one solid-state image pickup device by bonding the semiconductor chip of the first layer and the semiconductor chip of the second layer while electrically contacting each other.
[0029]
As another example, the image pickup apparatus 1 can be formed by a three-layer structure in which semiconductor chips are laminated in three layers. FIG. 3B is a diagram showing an example in which the image pickup apparatus 1 according to each embodiment is formed by a laminated CIS having a three-layer structure. In the structure of FIG. 3B, the pixel portion 20a is formed on the semiconductor chip of the first layer, the memory portion 20c is formed on the semiconductor chip of the second layer, and the logic portion 20b'is formed on the semiconductor chip of the third layer. .. In this case, the logic unit 20b'includes, for example, a sensor control unit 11, a recognition processing unit 12, a visual recognition processing unit 14, and an output control unit 15, and an interface for communicating between the image pickup device 1 and the outside. Further, the memory unit 20c can include a memory 13 and a memory used by, for example, the visual recognition processing unit 14 for processing image data. The memory 13 may be included in the logic unit 20b'.
[0030]
As shown on the right side of FIG. 3B, the image sensor 1 is attached to the image sensor 1 by electrically contacting the semiconductor chip of the first layer, the semiconductor chip of the second layer, and the semiconductor chip of the third layer. It is configured as one solid-state image sensor.
[0031]
FIG. 4 is a block diagram showing a configuration of an example of the sensor unit 10 applicable to each embodiment. In FIG. 4, the sensor unit 10 includes a pixel array unit 101, a vertical scanning unit 102, an AD (Analog to Digital) conversion unit 103, a pixel signal line 106, a vertical signal line VSL, a control unit 1100, and a signal. The processing unit 1101 and the like are included. In FIG. 4, the control unit 1100 and the signal processing unit 1101 may be included in the sensor control unit 11 shown in FIG. 1, for example.
[0032]
The pixel array unit 101 includes a plurality of pixel circuits 100 including, for example, a photoelectric conversion element using a photodiode and a circuit for reading out charges from the photoelectric conversion element, each of which performs photoelectric conversion on the received light. In the pixel array unit 101, the plurality of pixel circuits 100 are arranged in a matrix in the horizontal direction (row direction) and the vertical direction (column direction). In the pixel array unit 101, the arrangement of the pixel circuits 100 in the row direction is called a line. For example, when an image of one frame is formed by 1920 pixels × 1080 lines, the pixel array unit 101 includes at least 1080 lines including at least 1920 pixel circuits 100. An image (image data) of one frame is formed by the pixel signal read from the pixel circuit 100 included in the frame.
[0033]
Hereinafter, the operation of reading the pixel signal from each pixel circuit 100 included in the frame in the sensor unit 10 will be described as appropriate, such as reading the pixel from the frame. Further, the operation of reading the pixel signal from each pixel circuit 100 of the line included in the frame is described as appropriately reading the line.
[0034]
Further, in the pixel array unit 101, a pixel signal line 106 is connected to each row and column of each pixel circuit 100, and a vertical signal line VSL is connected to each column. The end of the pixel signal line 106 that is not connected to the pixel array unit 101 is connected to the vertical scanning unit 102. The vertical scanning unit 102 transmits a control signal such as a drive pulse when reading a pixel signal from a pixel to the pixel array unit 101 via the pixel signal line 106 in accordance with the control of the control unit 1100 described later. The end portion of the vertical signal line VSL that is not connected to the pixel array unit 101 is connected to the AD conversion unit 103. The pixel signal read from the pixel is transmitted to the AD conversion unit 103 via the vertical signal line VSL.
[0035]
The reading control of the pixel signal from the pixel circuit 100 will be schematically described. The pixel signal is read out from the pixel circuit 100 by transferring the electric charge accumulated in the photoelectric conversion element by exposure to the floating diffusion layer (FD) and converting the electric charge transferred in the floating diffusion layer into a voltage. Do. The voltage at which the charge is converted in the floating diffusion layer is output to the vertical signal line VSL via an amplifier.
[0036]
More specifically, in the pixel circuit 100, during exposure, the space between the photoelectric conversion element and the floating diffusion layer is turned off (open), and the photoelectric conversion element is generated according to the light incidented by the photoelectric conversion. Accumulates electric charge. After the end of exposure, the floating diffusion layer and the vertical signal line VSL are connected according to the selection signal supplied via the pixel signal line 106. Further, the floating diffusion layer is connected to the supply line of the power supply voltage VDD or the black level voltage in a short period of time according to the reset pulse supplied via the pixel signal line 106 to reset the floating diffusion layer. A voltage (assumed to be voltage A) at the reset level of the floating diffusion layer is output to the vertical signal line VSL. After that, the transfer pulse supplied via the pixel signal line 106 turns the photoelectric conversion element and the floating diffusion layer into an on (closed) state, and transfers the electric charge accumulated in the photoelectric conversion element to the floating diffusion layer. A voltage (referred to as voltage B) corresponding to the amount of electric charge of the floating diffusion layer is output to the vertical signal line VSL.
[0037]
The AD conversion unit 103 includes an AD converter 107 provided for each vertical signal line VSL, a reference signal generation unit 104, and a horizontal scanning unit 105. The AD converter 107 is a column AD converter that performs AD conversion processing on each column of the pixel array unit 101. The AD converter 107 performs AD conversion processing on the pixel signal supplied from the pixel circuit 100 via the vertical signal line VSL to reduce noise, and is used for Correlated Double Sampling (CDS) processing. Two digital values (values corresponding to voltage A and voltage B, respectively) are generated.
[0038]
The AD converter 107 supplies the two generated digital values to the signal processing unit 1101. The signal processing unit 1101 performs CDS processing based on the two digital values supplied from the AD converter 107, and generates a pixel signal (pixel data) based on the digital signal. The pixel data generated by the signal processing unit 1101 is output to the outside of the sensor unit 10.
[0039]
The reference signal generation unit 104 generates a lamp signal as a reference signal, which is used by each AD converter 107 to convert the pixel signal into two digital values, based on the control signal input from the control unit 1100. The lamp signal is a signal whose level (voltage value) decreases with a constant slope with respect to time, or a signal whose level decreases stepwise. The reference signal generation unit 104 supplies the generated lamp signal to each AD converter 107. The reference signal generation unit 104 is configured by using, for example, a DAC (Digital to Analog Converter) or the like.
[0040]
When the reference signal generation unit 104 supplies a lamp signal whose voltage drops stepwise according to a predetermined inclination, the counter starts counting according to the clock signal. The comparator compares the voltage of the pixel signal supplied from the vertical signal line VSL with the voltage of the lamp signal, and stops the counting by the counter at the timing when the voltage of the lamp signal straddles the voltage of the pixel signal. The AD converter 107 converts the pixel signal of the analog signal into a digital value by outputting a value corresponding to the count value of the time when the count is stopped.
[0041]
The AD converter 107 supplies the two generated digital values to the signal processing unit 1101. The signal processing unit 1101 performs CDS processing based on the two digital values supplied from the AD converter 107, and generates a pixel signal (pixel data) based on the digital signal. The pixel signal generated by the digital signal generated by the signal processing unit 1101 is output to the outside of the sensor unit 10.
[0042]
Under the control of the control unit 1100, the horizontal scanning unit 105 performs selective scanning in which each AD converter 107 is selected in a predetermined order to temporarily hold each digital value of each AD converter 107. The signal processing unit 1101 is sequentially output. The horizontal scanning unit 105 is configured by using, for example, a shift register or an address decoder.
[0043]
The control unit 1100 performs drive control of the vertical scanning unit 102, the AD conversion unit 103, the reference signal generation unit 104, the horizontal scanning unit 105, and the like according to the imaging control signal supplied from the sensor control unit 11. The control unit 1100 generates various drive signals that serve as a reference for the operation of the vertical scanning unit 102, the AD conversion unit 103, the reference signal generation unit 104, and the horizontal scanning unit 105. The control unit 1100 supplies the vertical scanning unit 102 to each pixel circuit 100 via the pixel signal line 106, for example, based on the vertical synchronization signal or the external trigger signal included in the imaging control signal and the horizontal synchronization signal. Generate a control signal. The control unit 1100 supplies the generated control signal to the vertical scanning unit 102.
[0044]
Further, the control unit 1100 passes, for example, information indicating an analog gain included in the image pickup control signal supplied from the sensor control unit 11 to the AD conversion unit 103. The AD conversion unit 103 controls the gain of the pixel signal input to each AD converter 107 included in the AD conversion unit 103 via the vertical signal line VSL according to the information indicating the analog gain.
[0045]
Based on the control signal supplied from the control unit 1100, the vertical scanning unit 102 transmits various signals including a drive pulse to the pixel signal line 106 of the selected pixel line of the pixel array unit 101 to each pixel circuit 100 for each line. It is supplied, and the pixel signal is output from each pixel circuit 100 to the vertical signal line VSL. The vertical scanning unit 102 is configured by using, for example, a shift register or an address decoder. Further, the vertical scanning unit 102 controls the exposure in each pixel circuit 100 according to the information indicating the exposure supplied from the control unit 1100.
[0046]
The sensor unit 10 configured in this way is a column AD type CMOS (Complementary Metal Oxide Semiconductor) image sensor in which AD converters 107 are arranged in columns.
[0047]
[2. Examples of existing technologies applicable to the present disclosure]
Prior to the description of each embodiment according to the present disclosure, the existing technologies applicable to the present disclosure will be outlined in order to facilitate understanding.
[0048]
(2-1. Outline of Rolling Shutter) A
rolling shutter (RS) method and a global shutter (GS) method are known as an imaging method when an image is taken by the pixel array unit 101. First, the rolling shutter method will be schematically described. 5A, 5B and 5C are schematic views for explaining the rolling shutter method. In the rolling shutter method, as shown in FIG. 5A, imaging is performed in order from line 201 at the upper end of the frame 200, for example, in line units.
[0049]
In the above description, "imaging" has been described as referring to an operation in which the sensor unit 10 outputs a pixel signal according to the light applied to the light receiving surface. More specifically, "imaging" refers to a series of operations from exposing a pixel to transferring a pixel signal based on the charge accumulated by the exposure to the photoelectric conversion element included in the pixel to the sensor control unit 11. And. Further, as described above, the frame refers to an area in the pixel array unit 101 in which a pixel circuit 100 effective for generating a pixel signal is arranged.
[0050]
For example, in the configuration of FIG. 4, exposure is simultaneously executed in each pixel circuit 100 included in one line. After the end of the exposure, the pixel signals based on the charges accumulated by the exposure are simultaneously transferred in each pixel circuit 100 included in the line via each vertical signal line VSL corresponding to each pixel circuit 100. By sequentially executing this operation in line units, imaging with a rolling shutter can be realized.
[0051]
FIG. 5B schematically shows an example of the relationship between imaging and time in the rolling shutter method. In FIG. 5B, the vertical axis represents the line position and the horizontal axis represents the time. In the rolling shutter method, since the exposure in each line is performed in sequence, the timing of exposure in each line shifts in order according to the position of the line, as shown in FIG. 5B. Therefore, for example, when the positional relationship between the image pickup device 1 and the subject in the horizontal direction changes at high speed, the image of the captured frame 200 is distorted as illustrated in FIG. 5C. In the example of FIG. 5C, the image 202 corresponding to the frame 200 is an image tilted at an angle corresponding to the speed and direction of change in the horizontal positional relationship between the image pickup device 1 and the subject.
[0052]
In the rolling shutter method, it is also possible to thin out the lines and take an image. 6A, 6B and 6C are schematic views for explaining line thinning in the rolling shutter system. As shown in FIG. 6A, as in the example of FIG. 5A described above, imaging is performed line by line from the line 201 at the upper end of the frame 200 toward the lower end of the frame 200. At this time, imaging is performed while skipping lines at predetermined numbers.
[0053]
Here, for the sake of explanation, it is assumed that imaging is performed every other line by thinning out one line. That is, after the imaging of the nth line, the imaging of the (n + 2) line is performed. At this time, it is assumed that the time from the imaging of the nth line to the imaging of the (n + 2) line is equal to the time from the imaging of the nth line to the imaging of the (n + 1) line when the thinning is not performed.
[0054]
FIG. 6B schematically shows an example of the relationship between imaging and time when one line is thinned out in the rolling shutter method. In FIG. 6B, the vertical axis represents the line position and the horizontal axis represents the time. In FIG. 6B, the exposure A corresponds to the exposure of FIG. 5B without thinning, and the exposure B shows the exposure when one line is thinned out. As shown in the exposure B, by performing the line thinning, it is possible to shorten the deviation of the exposure timing at the same line position as compared with the case where the line thinning is not performed. Therefore, as illustrated as image 203 in FIG. 6C, the distortion in the tilt direction that occurs in the image of the captured frame 200 is smaller than that in the case where the line thinning shown in FIG. 5C is not performed. On the other hand, when the line thinning is performed, the resolution of the image is lower than that when the line thinning is not performed.
[0055]
In the above description, an example in which image pickup is performed line-sequentially from the upper end to the lower end of the frame 200 in the rolling shutter method has been described, but this is not limited to this example. 7A and 7B are diagrams schematically showing examples of other imaging methods in the rolling shutter system. For example, as shown in FIG. 7A, in the rolling shutter system, line-sequential imaging can be performed from the lower end to the upper end of the frame 200. In this case, the horizontal direction of the distortion of the image 202 is opposite to that in the case where the images are sequentially imaged from the upper end to the lower end of the frame 200.
[0056]
Further, for example, by setting the range of the vertical signal line VSL that transfers the pixel signal, it is possible to selectively read a part of the line. Further, by setting the line for performing imaging and the vertical signal line VSL for transferring pixel signals, it is possible to set the lines for starting and ending imaging other than the upper end and the lower end of the frame 200. FIG. 7B schematically shows an example in which a rectangular region 205 whose width and height are less than the width and height of the frame 200 is set as the imaging range. In the example of FIG. 7B, imaging is performed from the line 204 at the upper end of the region 205 toward the lower end of the region 205 in order of line.
[0057]
(2-2. Outline of Global Shutter)
Next, the global shutter (GS) method will be schematically described as an imaging method when imaging is performed by the pixel array unit 101. 8A, 8B and 8C are schematic views for explaining the global shutter method. In the global shutter method, as shown in FIG. 8A, all pixel circuits 100 included in the frame 200 simultaneously expose.
[0058]
When the global shutter method is realized in the configuration of FIG. 4, as an example, it is conceivable that each pixel circuit 100 is further provided with a capacitor between the photoelectric conversion element and the FD. Then, a first switch is provided between the photoelectric conversion element and the capacitor, and a second switch is provided between the capacitor and the floating diffusion layer, and the opening and closing of each of the first and second switches is performed by pixels. The configuration is controlled by a pulse supplied via the signal line 106.
[0059]
In such a configuration, in the all-pixel circuit 100 included in the frame 200 during the exposure period, the first and second switches are opened, respectively, and at the end of the exposure, the first switch is opened and closed, and the photoelectric conversion element is used as a capacitor. Transfer the charge to. Hereinafter, the capacitor is regarded as a photoelectric conversion element, and the electric charge is read from the capacitor in the same sequence as the reading operation described in the rolling shutter method. As a result, simultaneous exposure is possible in all pixel circuits 100 included in the frame 200.
[0060]
FIG. 8B schematically shows an example of the relationship between imaging and time in the global shutter method. In FIG. 8B, the vertical axis represents the line position and the horizontal axis represents the time. In the global shutter method, since exposure is performed simultaneously in all the pixel circuits 100 included in the frame 200, the exposure timing in each line can be made the same as shown in FIG. 8B. Therefore, for example, even when the horizontal positional relationship between the image pickup device 1 and the subject changes at high speed, as illustrated in FIG. 8C, the image 206 of the captured frame 200 responds to the change. No distortion occurs.
[0061]
In the global shutter method, the simultaneity of exposure timing in the all-pixel circuit 100 included in the frame 200 can be ensured. Therefore, by controlling the timing of each pulse supplied by the pixel signal line 106 of each line and the timing of transfer by each vertical signal line VSL, sampling (reading of the pixel signal) in various patterns can be realized.
[0062]
9A and 9B are diagrams schematically showing an example of a sampling pattern that can be realized in the global shutter system. FIG. 9A is an example in which a sample 208 for reading a pixel signal is extracted in a checkered pattern from each of the pixel circuits 100 arranged in a matrix included in the frame 200. Further, FIG. 9B is an example of extracting a sample 208 for reading a pixel signal from each pixel circuit 100 in a grid pattern. Further, also in the global shutter method, as in the rolling shutter method described above, imaging can be performed in line sequence.
[0063]
(2-3. About
DNN ) Next, the recognition process using DNN (Deep Neural Network) applicable to each embodiment will be schematically described. In each embodiment, the image data is recognized by using CNN (Convolutional Neural Network) and RNN (Recurrent Neural Network) among the DNNs. Hereinafter, the "recognition process for image data" will be appropriately referred to as "image recognition process" or the like.
[0064]
(2-3-1. Outline
of CNN ) First, CNN will be described schematically. In the image recognition process by CNN, for example, the image recognition process is performed based on the image information by the pixels arranged in a matrix. FIG. 10 is a diagram for schematically explaining the image recognition process by CNN. The entire pixel information 51 of the image 50 in which the number "8", which is the object to be recognized, is drawn, is processed by the predeterminedly learned CNN 52. As a result, the number "8" is recognized as the recognition result 53.
[0065]
On the other hand, it is also possible to perform processing by CNN based on the image for each line and obtain the recognition result from a part of the image to be recognized. FIG. 11 is a diagram for schematically explaining an image recognition process for obtaining a recognition result from a part of the image to be recognized. In FIG. 11, the image 50'is a partial acquisition of the number "8", which is the object to be recognized, in line units. For example, the pixel information 54a, 54b, and 54c for each line forming the pixel information 51'of the image 50' are sequentially processed by the predeterminedly learned CNN 52'.
[0066]
For example, it is assumed that the recognition result 53a obtained by the recognition process by the CNN 52'for the pixel information 54a of the first line is not a valid recognition result. Here, the valid recognition result refers to, for example, a recognition result in which the score indicating the reliability of the recognized result is equal to or higher than a predetermined value. The CNN 52'updates the internal state 55 based on the recognition result 53a. Next, the pixel information 54b of the second line is recognized by the CNN 52'where the internal state is updated 55 by the previous recognition result 53a. In FIG. 11, as a result, a recognition result 53b indicating that the number to be recognized is either “8” or “9” is obtained. Further, based on this recognition result 53b, the internal information of CNN 52'is updated 55. Next, the pixel information 54c of the third line is recognized by the CNN 52'where the internal state is updated 55 by the previous recognition result 53b. As a result, in FIG. 11, the number to be recognized is narrowed down to “8” out of “8” or “9”.
[0067]
Here, the recognition process shown in FIG. 11 updates the internal state of the CNN using the result of the previous recognition process, and the CNN whose internal state has been updated is adjacent to the line on which the previous recognition process was performed. The recognition process is performed using the pixel information of the line to be used. That is, the recognition process shown in FIG. 11 is executed while sequentially updating the internal state of the CNN with respect to the image based on the previous recognition result. Therefore, the recognition process shown in FIG. 11 is a process that is recursively executed in line sequence, and can be considered to have a structure corresponding to RNN.
[0068]
(2-3-2. Outline
of RNN ) Next, the RNN will be described schematically. 12A and 12B are diagrams schematically showing an example of identification processing (recognition processing) by DNN when time series information is not used. In this case, one image is input to the DNN as shown in FIG. 12A. In DNN, identification processing is performed on the input image, and the identification result is output.
[0069]
FIG. 12B is a diagram for explaining the process of FIG. 12A in more detail. As shown in FIG. 12B, the DNN performs a feature extraction process and an identification process. In DNN, the feature amount is extracted from the input image by the feature extraction process. Further, in the DNN, the identification process is executed on the extracted feature amount, and the identification result is obtained.
[0070]
13A and 13B are diagrams schematically showing a first example of identification processing by DNN when time series information is used. In the examples of FIGS. 13A and 13B, identification processing by DNN is performed using a fixed number of past information on the time series. In the example of FIG. 13A, the image of the time T [T], the image of the time T-1 before the time T [T-1], and the image of the time T-2 before the time T-1 [T-2]. ] And is input to DNN. In the DNN, the identification process is executed for each of the input images [T], [T-1] and [T-2], and the identification result [T] at the time T is obtained.
[0071]
FIG. 13B is a diagram for explaining the process of FIG. 13A in more detail. As shown in FIG. 13B, in DNN, for each of the input images [T], [T-1] and [T-2], a pair of feature extraction processes described with reference to FIG. 12B described above is performed. 1 is executed, and the feature quantities corresponding to the images [T], [T-1] and [T-2] are extracted. In DNN, each feature amount obtained based on these images [T], [T-1] and [T-2] is integrated, an identification process is executed for the integrated feature amount, and identification at time T is performed. The result [T] is obtained.
[0072]
In the methods of FIGS. 13A and 13B, a plurality of configurations for performing feature extraction are required, and a configuration for performing feature extraction is required according to the number of past images that can be used. There is a risk that the configuration of will be large.
[0073]
14A and 14B are diagrams schematically showing a second example of identification processing by DNN when time series information is used. In the example of FIG. 14A, the image [T] of the time T is input to the DNN whose internal state is updated to the state of the time T-1, and the identification result [T] at the time T is obtained.
[0074]
FIG. 14B is a diagram for explaining the process of FIG. 14A in more detail. As shown in FIG. 14B, in the DNN, the feature extraction process described with reference to FIG. 12B described above is executed on the input time T image [T], and the feature amount corresponding to the image [T] is obtained. Extract. In the DNN, the internal state is updated by the image before the time T, and the feature amount related to the updated internal state is stored. The feature amount related to the stored internal information and the feature amount in the image [T] are integrated, and the identification process is executed for the integrated feature amount.
[0075]
The identification process shown in FIGS. 14A and 14B is executed using, for example, a DNN whose internal state has been updated using the immediately preceding identification result, and is a recursive process. A DNN that performs recursive processing in this way is called an RNN (Recurrent Neural Network). The identification process by RNN is generally used for moving image recognition and the like. For example, it is possible to improve the identification accuracy by sequentially updating the internal state of the DNN with a frame image updated in time series. ..
[0076]
In the present disclosure, RNN is applied to a rolling shutter type structure. That is, in the rolling shutter method, the pixel signal is read out in line sequence. Therefore, the pixel signals read out in this line sequence are applied to the RNN as information on the time series. As a result, the identification process based on a plurality of lines can be executed with a smaller configuration than the case where the CNN is used (see FIG. 13B). Not limited to this, RNN can also be applied to the structure of the global shutter system. In this case, for example, it is conceivable to regard adjacent lines as information on a time series.
[0077]
(2-4. Drive Speed)
Next, the relationship between the frame drive speed and the pixel signal reading amount will be described with reference to FIGS. 15A and 15B. FIG. 15A is a diagram showing an example of reading out all the lines in the image. Here, it is assumed that the resolution of the image to be recognized is 640 pixels horizontally × 480 pixels vertically (480 lines). In this case, by driving at a driving speed of 14400 [lines / sec], it is possible to output at 30 [fps (frame per second)].
[0078]
Next, consider thinning out the lines for imaging. For example, as shown in FIG. 15B, it is assumed that imaging is performed by skipping one line at a time and performing imaging by 1/2 thinning-out readout. As a first example of 1/2 thinning, when driving at a driving speed of 14400 [lines / sec] as described above, the number of lines read from the image is halved, so that the resolution is reduced, but thinning is performed. It is possible to output at 60 [fps], which is twice as fast as when it is not performed, and the frame rate can be improved. As a second example of 1/2 thinning, when the drive speed is set to 7200 [fps], which is half of the first example, the frame rate is 30 [fps] as in the case of no thinning, but power saving. Can be converted.
[0079]
When reading out an image line, whether to not thin out, to thin out to increase the drive speed, or to thin out and make the drive speed the same as when not thinning out is based on, for example, the read pixel signal. It can be selected according to the purpose of the recognition process.
[0080]
[3. Outline of the present disclosure]
Hereinafter, each embodiment of the present disclosure will be described in more detail. First, the processing according to each embodiment of the present disclosure will be schematically described. FIG. 16 is a schematic diagram for schematically explaining the recognition process according to each embodiment of the present disclosure. In FIG. 16, in step S1, the imaging device 1 (see FIG. 1) according to each embodiment starts imaging the target image to be recognized.
[0081]
The target image is, for example, an image in which the number "8" is drawn by hand. Further, in the memory 13, a learning model learned so that numbers can be identified by predetermined teacher data is stored in advance as a program, and the recognition processing unit 12 reads this program from the memory 13 and executes it. It is assumed that the numbers contained in the image can be identified. Further, the image pickup apparatus 1 shall perform imaging by the rolling shutter method. Even when the image pickup apparatus 1 performs imaging by the global shutter method, the following processing can be applied in the same manner as in the case of the rolling shutter method.
[0082]
When the imaging is started, the imaging device 1 sequentially reads out the frames in line units from the upper end side to the lower end side of the frame in step S2.
[0083]
When the line is read up to a certain position, the recognition processing unit 12 identifies the number "8" or "9" from the image of the read line (step S3). For example, since the numbers "8" and "9" include a feature portion common to the upper half portion, when the line is read out in order from the top and the feature portion is recognized, the recognized object is the number "8". It can be identified as either "" or "9".
[0084]
Here, as shown in step S4a, the whole picture of the recognized object appears by reading up to the line at the lower end of the frame or the line near the lower end, and as either the number "8" or "9" in step S2. It is determined that the identified object is the number "8".
[0085]
On the other hand, steps S4b and S4c are processes related to the present disclosure.
[0086]
As shown in step S4b, the line is further read from the line position read in step S3, and the recognized object is identified as the number "8" even while reaching the lower end of the number "8". It is possible. For example, the lower half of the number "8" and the lower half of the number "9" have different characteristics. By reading the line up to the part where the difference in the characteristics becomes clear, it becomes possible to identify which of the numbers "8" and "9" the object recognized in step S3 is. In the example of FIG. 16, in step S4b, the object is determined to be the number "8".
[0087]
Further, as shown in step S4c, by further reading from the line position in step S3 in the state of step S3, it can be determined whether the object identified in step S3 is the number "8" or "9". It is also possible to jump to a line position that is likely to be discernible. By reading out the line of the jump destination, it is possible to determine whether the object identified in step S3 is the number "8" or "9". The line position of the jump destination can be determined based on a learning model learned in advance based on predetermined teacher data.
[0088]
Here, when the object is determined in step S4b or step S4c described above, the image pickup apparatus 1 can end the recognition process. This makes it possible to shorten the recognition process and save power in the image pickup apparatus 1.
[0089]
The teacher data is data that holds a plurality of combinations of input signals and output signals for each read unit. As an example, in the task of identifying numbers described above, data for each read unit (line data, subsampled data, etc.) is applied as an input signal, and data indicating a "correct number" is applied as an output signal. Can be done. As another example, for example, in the task of detecting an object, data for each read unit (line data, subsampled data, etc.) is applied as an input signal, and an object class (human body / vehicle / non-object) or an object class (human body / vehicle / non-object) is applied as an output signal. The coordinates of the object (x, y, h, w) and the like can be applied. Further, the output signal may be generated only from the input signal by using self-supervised learning.
[0090]
[4. First Embodiment]
Next, the first embodiment of the present disclosure will be described.
[0091]
(4-1. Example of Operation by Recognition Processing Unit) In
the image pickup apparatus 1 according to the first embodiment, the recognition processing unit 12 has a memory 13 as a learning model pre-learned based on predetermined teacher data as described above. By reading and executing the program stored in, it functions as a recognizer using DNN.
[0092]
FIG. 17 is a flowchart of an example showing the recognition process by the recognition process unit 12 according to the first embodiment. In FIG. 17, in step S121, the DSP constituting the recognition processing unit 12 in the image pickup apparatus 1 reads the learning model from the memory 13 and executes it. As a result, the DSP functions as the recognition processing unit 12.
[0093]
Next, in step S122, the recognition processing unit 12 in the image pickup apparatus 1 instructs the sensor control unit 11 to start reading a frame from the sensor unit 10. In this frame reading, for example, one frame of image data is sequentially read in line units (also referred to as line units). The recognition processing unit 12 determines whether or not a predetermined number of image data in one frame has been read.
[0094]
When the recognition processing unit 12 determines that the image data of a predetermined number of lines in one frame has been read (step S123, “YES”), the processing shifts to step S124. In step S124, the recognition processing unit 12 executes recognition processing as machine learning processing using CNN on the image data for the predetermined number of read lines. That is, the recognition processing unit 12 executes machine learning processing using the learning model with the image data of a predetermined number of lines as a unit area. In machine learning processing using CNN, for example, face detection, face recognition, line-of-sight detection, facial expression recognition, face direction detection, object detection, object recognition, movement (animal body) detection, pet detection, scene recognition, and state detection. , Avoidance object recognition, and other recognition processing and detection processing are executed.
[0095]
Here, the face detection is a process of detecting the face of a person included in the image data. The face authentication is one of biometric authentication, and is a process of authenticating whether or not the face of a person included in the image data matches the face of a person registered in advance. The line-of-sight detection is a process of detecting the direction of the line of sight of a person included in the image data. Facial expression recognition is a process of recognizing a person's facial expression included in image data. The face direction detection is a process of detecting the vertical direction of a person's face included in the image data. Object detection is a process of detecting an object included in image data. Object recognition is a process of recognizing what an object is included in image data. Motion (animal body) detection is a process of detecting an animal body included in image data. Pet detection is a process for detecting pets such as dogs and cats included in image data. Scene recognition is a process of recognizing a scene (sea, mountain, etc.) being shot. The state detection is a process of detecting the state of a person or the like (whether it is a normal state or an abnormal state, etc.) included in the image data. The avoidance object recognition is a process of recognizing an object to be avoided that exists in front of the moving direction when the object moves. The machine learning process executed by the recognition processing unit 12 is not limited to these examples.
[0096]
In step S125, the recognition processing unit 12 determines whether or not the machine learning process using the CNN in step S124 was successful. When the recognition processing unit 12 determines that the machine learning process using the CNN has succeeded (step S125, “YES”), the process shifts to step S129. On the other hand, when the recognition processing unit 12 determines that the machine learning process using the CNN in step S124 has failed (step S125, “NO”), the process shifts to step S126. In step S126, the recognition processing unit 12 waits for the next predetermined number of lines of image data to be read from the sensor control unit 11 (step S126, “NO”).
[0097]
In this description, successful machine learning processing means that, for example, a certain detection result, recognition result, or authentication has been obtained in face detection, face authentication, or the like as illustrated above. On the other hand, failure of the machine learning process means that, for example, sufficient detection results, recognition results, and authentication could not be obtained in face detection, face authentication, and the like as illustrated above.
[0098]
Next, in step S126, when the next predetermined number of image data (unit area) is read (step S126, “YES”), the recognition processing unit 12 determines the number of read predetermined lines in step S127. Machine learning processing using RNN is executed for the image data. In the machine learning process using the RNN, for example, the result of the machine learning process using the CNN or the RNN that has been executed so far on the image data of the same frame is also used.
[0099]
In step S128, when the recognition processing unit 12 determines that the machine learning process using the RNN in step S127 has succeeded (step S128, “YES”), the process shifts to step S129.
[0100]
In step S129, the recognition processing unit 12 supplies the machine learning result succeeded in step S124 or step S127, for example, from the recognition processing unit 12 to the output control unit 15. The machine learning result output in step S129 is, for example, a valid recognition result by the recognition processing unit 12. The recognition processing unit 12 may store the machine learning result in the memory 13.
[0101]
Further, when the recognition processing unit 12 determines in step S128 that the machine learning process using the RNN in step S127 has failed (step S128, “NO”), the process shifts to step S130. In step S130, the recognition processing unit 12 determines whether or not the reading of the image data for one frame is completed. When the recognition processing unit 12 determines that the reading of the image data for one frame has not been completed (step S130, “NO”), the processing is returned to step S126, and the processing for the next predetermined number of lines of image data is performed. Will be executed.
The scope of the claims
[Claim 1]
A read unit control that controls an imaging unit that has a pixel area in which a plurality of pixels are arranged and reads and outputs a pixel signal from the pixels included in the pixel area, and a read unit set as a part of the pixel area. A unit,
a recognition unit that has learned the teacher data for
each read unit, a first output unit that outputs the recognition result recognized by the recognition unit for each read unit , and a second output unit that outputs the
pixel signal to the subsequent stage. and output portion,
the first and the first timing output section outputs the recognition result, the trigger signal for the second output unit controls and a second timing for outputting the pixel signal An
image pickup device including a trigger generator for generating the above .
[Claim 2]
The imaging device according to claim 1, wherein the trigger generation unit generates the trigger signal that synchronizes the first timing with the second timing.
[Claim 3]
The imaging device according to claim 1, wherein the trigger generation unit outputs the trigger signal at a fixed cycle .
[Claim 4]
The
imaging device according to
claim 1, wherein the trigger generation unit generates the trigger signal in response to external control of the imaging device.
[Claim 5]
The imaging device according to claim 1, wherein the trigger generation unit generates the trigger signal according to the recognition certainty indicating the certainty of the recognition result.
[Claim 6]
The output control unit
further includes an output control unit that controls the output of the recognition result by the first output unit and the output of the pixel signal by the second output unit, and the
output control unit is
the first output.
The imaging device according to claim 1, wherein the image pickup apparatus according to claim 1 corrects a deviation between the recognition result output by the unit and the pixel signal output by the second output unit .
[Claim 7]
The output control unit
includes a holding unit that holds the recognition result output to the first output unit, and the recognition result held by the holding unit is obtained by the second output unit of the pixel signal.
The imaging apparatus according to claim 6 , which outputs in synchronization with the output .
[Claim 8]
The holding unit
further holds the pixel signal output to the second output unit, and the
output control unit further holds the
recognition result held by the holding unit and the amount of spatial deviation between the pixel signal. , And based on the estimated deviation amount, correct at least one of the pixel signal and the recognition result held in the holding unit, and the recognition result held in the holding unit and the holding unit.
The image pickup apparatus according to claim 7, wherein the corrected pixel signal held in the image is output in synchronization with the corrected pixel signal .
[Claim 9]
In the trigger generation unit
, the recognition unit recognizes the reading of the pixel signal from the image pickup unit for output by the second output unit in response to the output of the trigger signal to the first output unit.
The imaging device according to claim 1 , wherein the imaging unit is instructed to read the pixel signal from the imaging unit at a speed higher than the speed at which the pixel signal is read from the imaging unit.
[Claim 10]
In
response to the output of the trigger signal to the second output unit , the trigger generation unit reads the pixel signal for recognition by the recognition unit from the image pickup unit, and the second output unit reads the pixel signal.
The imaging device according to claim 1 , wherein the imaging unit is instructed to read the pixel signal for output from the imaging unit at a speed higher than the speed at which the pixel signal is read out from the imaging unit.
[Claim 11]
A read unit control that controls an image pickup unit that has a pixel area in which a plurality of pixels are arranged and reads and outputs a pixel signal from the pixels included in the pixel area, and a read unit set as a part of the pixel area. The unit,
the first output unit that outputs the recognition result recognized by the recognition unit for each read unit,
the second output unit that outputs the pixel signal to the subsequent stage, and
the first output unit are the recognition results. a first timing for outputting the,, and a trigger generator for generating a trigger signal for the second output unit controls and a second timing for outputting the pixel signal
and the image pickup apparatus comprising,
the An image pickup system including an
information processing device
including the recognition unit that has learned teacher data for each read unit .
[Claim 12]
A read unit control step for controlling a read unit set as a part of a pixel area in which a plurality of pixels are arranged, which is executed by a processor,
and a recognition step for learning teacher data for each read unit. A
first output step for outputting the recognition result recognized by the recognition step for each reading unit, a second output step for outputting
a pixel signal read from the pixels included in the pixel area to the subsequent stage, and the
above. A trigger generation step for generating a trigger signal for controlling a first timing for outputting the recognition result by the first output step and a second timing for outputting the pixel signal by the second output step. And,
including imaging methods.
[Claim 13]
A read unit control step for controlling a read unit set as a part of a pixel area in which a plurality of pixels are arranged, which an imaging unit has,
a recognition step for learning teacher data for
each read unit, and each read unit. By the first output step of outputting the recognition result recognized by the recognition step,
the second output step of outputting the pixel signal read from the pixels included in the pixel area to the subsequent stage, and
the first output step. a first timing for outputting the recognition result, the trigger generation step of generating a trigger signal for controlling the second timing for outputting the pixel signal by the second output step
executing a processor Imaging program to make it.
| # | Name | Date |
|---|---|---|
| 1 | 202117002200-TRANSLATIOIN OF PRIOIRTY DOCUMENTS ETC. [18-01-2021(online)].pdf | 2021-01-18 |
| 2 | 202117002200-STATEMENT OF UNDERTAKING (FORM 3) [18-01-2021(online)].pdf | 2021-01-18 |
| 3 | 202117002200-PRIORITY DOCUMENTS [18-01-2021(online)].pdf | 2021-01-18 |
| 4 | 202117002200-POWER OF AUTHORITY [18-01-2021(online)].pdf | 2021-01-18 |
| 5 | 202117002200-FORM 1 [18-01-2021(online)].pdf | 2021-01-18 |
| 6 | 202117002200-DRAWINGS [18-01-2021(online)].pdf | 2021-01-18 |
| 7 | 202117002200-DECLARATION OF INVENTORSHIP (FORM 5) [18-01-2021(online)].pdf | 2021-01-18 |
| 8 | 202117002200-COMPLETE SPECIFICATION [18-01-2021(online)].pdf | 2021-01-18 |
| 9 | 202117002200-Proof of Right [05-02-2021(online)].pdf | 2021-02-05 |
| 10 | 202117002200-FORM-26 [11-02-2021(online)].pdf | 2021-02-11 |
| 11 | 202117002200-Verified English translation [24-03-2021(online)].pdf | 2021-03-24 |
| 12 | 202117002200-Verified English translation [24-03-2021(online)]-1.pdf | 2021-03-24 |
| 13 | 202117002200-Proof of Right [08-04-2021(online)].pdf | 2021-04-08 |
| 14 | 202117002200-Proof of Right [15-04-2021(online)].pdf | 2021-04-15 |
| 15 | 202117002200-FORM 3 [26-04-2021(online)].pdf | 2021-04-26 |
| 16 | 202117002200.pdf | 2021-10-19 |
| 17 | 202117002200-FORM 18 [07-07-2022(online)].pdf | 2022-07-07 |
| 18 | 202117002200-FER.pdf | 2022-10-10 |
| 19 | 202117002200-FER_SER_REPLY [10-04-2023(online)].pdf | 2023-04-10 |
| 20 | 202117002200-CORRESPONDENCE [10-04-2023(online)].pdf | 2023-04-10 |
| 21 | 202117002200-CLAIMS [10-04-2023(online)].pdf | 2023-04-10 |
| 22 | 202117002200-PatentCertificate24-07-2024.pdf | 2024-07-24 |
| 23 | 202117002200-IntimationOfGrant24-07-2024.pdf | 2024-07-24 |
| 1 | 202117002200E_07-10-2022.pdf |