Sign In to Follow Application
View All Documents & Correspondence

System And Method For Temporal Memory Based Monocular Three Dimensional Object Detection In Video Sequences

Abstract: The present invention discloses a system and method for improving monocular three-dimensional (3D) object detection in video sequences using a recurrent Memory-to-Scene (M2S) architecture. The proposed system processes sequential video frames captured from a monocular camera and extracts spatial feature representations using a deep learning–based encoder. A temporal memory bank is introduced to store scene representations inferred from previous frames. A recurrent memory recall unit selectively retrieves relevant scene knowledge from the stored memory and integrates it with the current frame features to enhance prediction accuracy. The system further incorporates motion-aware and appearance-aware feature encoding mechanisms to improve estimation of object depth, orientation, and three-dimensional bounding boxes. A temporal consistency loss function is introduced to minimize prediction discrepancies between consecutive frames and ensure stable object trajectories in dynamic environments. The proposed invention operates using only monocular video input and eliminates the need for additional sensing devices such as LiDAR or stereo cameras. The system is particularly suitable for applications including autonomous driving, intelligent transportation systems, robotic perception, and real-time video surveillance, where reliable three-dimensional scene understanding is required.

Get Free WhatsApp Updates!
Notices, Deadlines & Correspondence

Patent Information

Application #
Filing Date
12 April 2026
Publication Number
17/2026
Publication Type
INA
Invention Field
COMPUTER SCIENCE
Status
Email
Parent Application

Applicants

SR University
Warangal, Ananthasagar, Hasanparthy (PO), Warangal-506371, Telangana, India

Inventors

1. Ms. Dodda Prasanthi
Research Scholar, SR University, Warangal, Anathsagar, Hanumakonda, Telangana, India, 506 371
2. Dr. Ramesh Babu Akarapu
Assistant Professor, SR University, Warangal, Anathsagar, Hanumakonda, Telangana, India, 506 371
3. Dr. Kurumalla Suresh
Professor, St. Martin’s Engineering College (Autonomous), Dhulapally, Hyderabad, India, 500100

Claims

2. the system extracts spatial features from input video frames and stores scene representations in a temporal memory bank to preserve historical scene information.

3. the system includes a memory recall mechanism that retrieves relevant information from previously stored frames to assist in improving current frame predictions.

4. the system integrates motion-aware and appearance-aware features to improve the estimation of object depth, orientation, and three-dimensional bounding box parameters.

5. the system maintains temporal consistency across frames to reduce prediction fluctuations and ensure stable object detection in dynamic environments.

6. the system improves detection performance in challenging conditions such as occlusion, motion blur, object movement, and camera motion.

7. the proposed system operates using only monocular video input and eliminates the requirement for additional sensing devices such as LiDAR, stereo cameras, or depth sensors. the system can be applied in intelligent vision applications including autonomous driving, intelligent transportation systems, robotic perception, and real-time video surveillance.

Specification

Description:The present invention generally relates to the field of computer vision, artificial intelligence, and autonomous perception systems. More particularly, the invention relates to a deep learning–based system and method for monocular 3D object detection from video sequences using a recurrent Memory-to-Scene (M2S) architecture.
The invention specifically focuses on improving the accuracy and temporal consistency of three-dimensional object detection using monocular camera input by incorporating temporal memory mechanisms that recall scene knowledge from consecutive video frames. The proposed system integrates temporal feature encoding, motion-aware memory representation, and appearance-aware scene understanding to enhance prediction of 3D bounding boxes, object depth, and orientation in dynamic environments.
The invention further relates to techniques for memory-based temporal fusion of scene representations, enabling the system to utilize previously inferred 3D information to support current frame predictions. By introducing a temporal memory bank and recurrent memory recall module, the system is capable of maintaining contextual scene knowledge across multiple frames and improving robustness under challenging conditions such as object occlusion, motion blur, camera movement, and dynamic object interactions.
The disclosed invention finds application in various intelligent vision systems including autonomous driving, intelligent transportation systems, robotic perception, smart surveillance, and video analytics, where reliable understanding of three-dimensional scenes from monocular video input is required.

Background of the Invention
Accurate. Importance of 3D Object Detection
Three-dimensional (3D) object detection has become a critical component in modern computer vision and intelligent perception systems. It plays an important role in applications such as autonomous driving, robotics, intelligent transportation systems, and smart surveillance. Accurate estimation of object depth, spatial position, and orientation is essential for safe navigation and reliable environmental understanding.

2. Existing Sensor-Based Approaches
Traditional 3D object detection methods commonly rely on LiDAR sensors, stereo cameras, or depth sensors to obtain accurate depth measurements. Although these technologies provide precise spatial information, they significantly increase hardware costs, system complexity, and computational requirements, which makes large-scale deployment difficult.

3. Emergence of Monocular 3D Detection
To overcome the limitations of expensive sensor systems, researchers have explored monocular 3D object detection, which estimates three-dimensional information using only a single RGB camera. Monocular approaches provide a cost-effective and scalable alternative for real-world vision systems.

4. Limitations of Existing Monocular Methods
Despite their advantages, monocular detection methods suffer from several challenges. Since a single image does not explicitly contain depth information, these methods often experience depth ambiguity, inaccurate object localization, and unstable orientation estimation. Furthermore, many existing approaches process each image independently, ignoring the temporal relationship between frames in a video sequence.
5. Challenges in Dynamic Environments
In real-world scenarios, objects frequently undergo motion, partial occlusion, illumination variation, and motion blur. These factors can lead to inconsistent detection results across consecutive frames, resulting in unstable object trajectories and unreliable predictions.

6. Need for Temporal Scene Understanding
Recent research has attempted to utilize temporal information from video sequences to improve detection accuracy. However, many existing systems lack the capability to store and recall previously inferred scene knowledge, which limits their ability to maintain consistent predictions across frames.

7. Motivation for the Proposed Invention
To address these limitations, there is a need for an advanced system capable of maintaining temporal scene memory and recalling relevant historical information to enhance monocular 3D object detection performance.

8. Proposed Solution
The present invention introduces a recurrent Memory-to-Scene (M2S) architecture that stores scene representations from previous frames and selectively recalls them to assist in current frame prediction. This approach improves the accuracy, stability, and robustness of monocular 3D object detection in video sequences.
Summary of the invention:
The1. Purpose of the Invention
The present invention provides a system and method for improving monocular three-dimensional (3D) object detection in video sequences by utilizing temporal scene knowledge from consecutive frames. The invention introduces a recurrent Memory-to-Scene (M2S) architecture designed to enhance detection accuracy and stability.
2. Monocular Video Frame Processing
The system processes sequential video frames captured from a monocular camera. Each frame is analyzed using a deep learning–based feature encoder to extract spatial features that represent objects and scene structure.
3. Temporal Memory Bank
A temporal memory bank is introduced to store scene representations inferred from previous frames. This component preserves historical scene information that can be reused to assist predictions in subsequent frames.
4. Memory Recall Mechanism
The invention incorporates a recurrent memory recall unit that selectively retrieves relevant scene information from the stored memory. This mechanism enables the system to utilize previously inferred knowledge to improve the interpretation of the current frame.
5. Feature Fusion Strategy
The recalled memory features are combined with the current frame features through a feature fusion module. This module integrates both motion-aware information and appearance-based visual features to create a richer scene representation.
6. 3D Object Prediction
The fused feature representation is processed by a 3D prediction module, which estimates important object parameters including:
• Three-dimensional bounding box coordinates
• Pseudo-depth estimation
• Object orientation and spatial positioning.

7. Temporal Consistency Optimization
To ensure stable detection across frames, the invention introduces a temporal consistency loss mechanism that minimizes discrepancies between predictions of consecutive frames, thereby maintaining smooth and reliable object trajectories.

8. Advantages of the Proposed System
The proposed invention improves monocular 3D detection by:
• Utilizing historical scene information from previous frames
• Enhancing prediction stability in dynamic environments
• Reducing errors caused by occlusion, motion blur, and object movement
• Providing a cost-effective solution that operates using only monocular video input

9. Application Areas
The invention can be effectively applied in various intelligent vision systems including:
• Autonomous driving systems
• Intelligent transportation systems
• Robotic perception and navigation
• Smart surveillance and video analytics
Brief description of the proposed invention:
The1. Purpose of the Invention
The present invention provides a system and method for improving monocular three-dimensional (3D) object detection in video sequences by utilizing temporal scene knowledge from consecutive frames. The invention introduces a recurrent Memory-to-Scene (M2S) architecture designed to enhance detection accuracy and stability.
2. Monocular Video Frame Processing
The system processes sequential video frames captured from a monocular camera. Each frame is analyzed using a deep learning–based feature encoder to extract spatial features that represent objects and scene structure.
3. Temporal Memory Bank
A temporal memory bank is introduced to store scene representations inferred from previous frames. This component preserves historical scene information that can be reused to assist predictions in subsequent frames.
4. Memory Recall Mechanism
The invention incorporates a recurrent memory recall unit that selectively retrieves relevant scene information from the stored memory. This mechanism enables the system to utilize previously inferred knowledge to improve the interpretation of the current frame.
5. Feature Fusion Strategy
The recalled memory features are combined with the current frame features through a feature fusion module. This module integrates both motion-aware information and appearance-based visual features to create a richer scene representation.
6. 3D Object Prediction
The fused feature representation is processed by a 3D prediction module, which estimates important object parameters including:
• Three-dimensional bounding box coordinates
• Pseudo-depth estimation
• Object orientation and spatial positioning.

7. Temporal Consistency Optimization
To ensure stable detection across frames, the invention introduces a temporal consistency loss mechanism that minimizes discrepancies between predictions of consecutive frames, thereby maintaining smooth and reliable object trajectories.

8. Advantages of the Proposed System
The proposed invention improves monocular 3D detection by:
• Utilizing historical scene information from previous frames
• Enhancing prediction stability in dynamic environments
• Reducing errors caused by occlusion, motion blur, and object movement
• Providing a cost-effective solution that operates using only monocular video input

9. Application Areas
The invention can be effectively applied in various intelligent vision systems including:
• Autonomous driving systems
• Intelligent transportation systems
• Robotic perception and navigation
• Smart surveillance and video analytics
Objectives:

1. The Improve 3D Object Detection Accuracy: Enhance the accuracy of monocular 3D object detection in video sequences by utilizing temporal information from consecutive frames.
2. Utilize Temporal Scene Information: Capture motion patterns and scene dynamics across multiple frames using a temporal feature extraction mechanism.
3. Introduce Temporal Memory Bank: Store previously inferred scene representations in a memory bank to preserve historical information for future predictions.
4. Enable Intelligent Memory Recall: Retrieve relevant scene knowledge from stored memory to support and refine predictions for the current frame.
5. Enhance Feature Fusion: Combine motion-aware and appearance-aware features to improve estimation of object depth, orientation, and 3D bounding box parameters.
6. Ensure Temporal Consistency: Maintain stable object detection across frames by reducing prediction inconsistencies in dynamic environments.
7. Improve Robustness in Challenging Conditions: Handle complex situations such as occlusion, motion blur, dynamic object movement, and camera motion more effectively.
8. Reduce Hardware Dependency: Provide accurate 3D scene understanding using only a monocular camera, eliminating the need for expensive sensors such as LiDAR or stereo cameras.
9. Support Intelligent Vision Applications: Enable deployment of the system in applications such as autonomous driving, intelligent transportation systems, robotic perception, and smart surveillance.

, Claims:a system for monocular 3D object detection in video sequences that utilizes a Memory-to-Scene (M2S) architecture to improve the accuracy and stability of object detection using temporal information from consecutive frames.
2. the system extracts spatial features from input video frames and stores scene representations in a temporal memory bank to preserve historical scene information.
3. the system includes a memory recall mechanism that retrieves relevant information from previously stored frames to assist in improving current frame predictions.
4. the system integrates motion-aware and appearance-aware features to improve the estimation of object depth, orientation, and three-dimensional bounding box parameters.
5. the system maintains temporal consistency across frames to reduce prediction fluctuations and ensure stable object detection in dynamic environments.
6. the system improves detection performance in challenging conditions such as occlusion, motion blur, object movement, and camera motion.
7. the proposed system operates using only monocular video input and eliminates the requirement for additional sensing devices such as LiDAR, stereo cameras, or depth sensors.
the system can be applied in intelligent vision applications including autonomous driving, intelligent transportation systems, robotic perception, and real-time video surveillance.

Documents

Application Documents

# Name Date
1 202641046800-STATEMENT OF UNDERTAKING (FORM 3) [12-04-2026(online)].pdf 2026-04-12
2 202641046800-POWER OF AUTHORITY [12-04-2026(online)].pdf 2026-04-12
3 202641046800-FORM-9 [12-04-2026(online)].pdf 2026-04-12
4 202641046800-FORM FOR SMALL ENTITY(FORM-28) [12-04-2026(online)].pdf 2026-04-12
5 202641046800-FORM 1 [12-04-2026(online)].pdf 2026-04-12
6 202641046800-EVIDENCE FOR REGISTRATION UNDER SSI(FORM-28) [12-04-2026(online)].pdf 2026-04-12
7 202641046800-EVIDENCE FOR REGISTRATION UNDER SSI [12-04-2026(online)].pdf 2026-04-12
8 202641046800-EDUCATIONAL INSTITUTION(S) [12-04-2026(online)].pdf 2026-04-12
9 202641046800-DRAWINGS [12-04-2026(online)].pdf 2026-04-12
10 202641046800-DECLARATION OF INVENTORSHIP (FORM 5) [12-04-2026(online)].pdf 2026-04-12
11 202641046800-COMPLETE SPECIFICATION [12-04-2026(online)].pdf 2026-04-12