Sign In to Follow Application
View All Documents & Correspondence

A System For Assisting A User For Operating Third Party Mobile Applications

Abstract: The present invention relates to a system (100) for assisting a user for operating third-party mobile applications. The system (100) includes an input module (10), an optical character recognition (OCR) engine (20), a natural-language processing engine (30), a database module (40), and a guidance generator (50). The input module (10) receives a screenshot of a third-party mobile application together with a user query in text or voice form. The OCR engine (20) analyzes the screenshot to detect user-interface (UI) elements and generate corresponding bounding-box coordinates. The natural-language processing engine (30) interprets the user query in relation to the detected UI elements to determine user intent and produce context-aware guidance parameters. The database module (40) stores annotated screenshots, created workflows, and instructional sequences to support contextual understanding. The guidance generator (50) produces step-wise instructions to actionable UI elements, overlays visual highlights, and triggers haptic-feedback and brightness-adjustment responses for synchronized user interaction.

Get Free WhatsApp Updates!
Notices, Deadlines & Correspondence

Patent Information

Application #
Filing Date
09 February 2026
Publication Number
12/2026
Publication Type
INA
Invention Field
COMPUTER SCIENCE
Status
Email
Parent Application

Applicants

Maharashtra Knowledge Corporation Ltd
ICC Trade Towers, 5th Floor, A-wing, Senapati Bapat Road, Pune-411016, Maharashtra, India.

Inventors

1. SAWANT, Vivek
A-13 , Nalanda Garden , 265 Baner Road, Opp. Mauli Mangal Karyalaya, Pune – 411045.
2. DESAI, Vikas
Dodke Terrace, Flat No 34, 6Th Floor, Chaitanya Chowk, Warje Malwadi, Pune – 411058.
3. KULKARNI, Pawan
Chintamani Concord Pushpak, B 1103, Porwal Road, Lohegaon, Ajinkya D Y Patil College Road, Pune – 411047.
4. GAIKWAD, Dhanshree
2nd floor, Abhas Building, Balakawade Wada, Near Sainath Mitra Mandal, Katraj, Pune 411046.

Specification

Description:Field of the invention
[0001] The present invention relates to the mobile applications. More specifically, the invention relates to a system for assisting a user for operating third-party mobile applications.

Background of the invention
[0002] Smartphone adoption has grown rapidly worldwide, yet a significant portion of users, particularly first-time or older users, struggle with digital literacy. Despite the large number of tutorials, manuals, and app-learning tools, many of these resources fail to provide guidance that is contextually aware or tailored to the specific interface elements a user encounters in real time. Static instructions, whether in the form of printed manuals or video tutorials, often assume a baseline familiarity with smartphone interaction paradigms, leaving new users confused or frustrated when confronted with unfamiliar applications.

[0003] Existing solutions in the domain of digital assistance typically rely on on-screen prompts, video walkthroughs, or step-by-step instructions that require the user to interpret abstract guidance and manually correlate it with the interface on their device. Such approaches lack interactivity and cannot adapt dynamically to the user’s current context, such as which screen is currently active, which buttons have been pressed, or whether a user has correctly completed a given action. This gap between instruction and action creates barriers for users attempting to learn and navigate new applications independently.

[0004] Furthermore, current tools rarely leverage sensory feedback mechanisms to guide user attention physically. Humans naturally respond to tactile and multi-sensory cues in addition to visual prompts, yet most digital assistance systems ignore these modalities. Without such guidance, users may overlook critical interface elements, press incorrect controls, or become disengaged from the learning process. This limitation is particularly pronounced for users with visual impairments, limited experience with touchscreens, or cognitive challenges, highlighting the need for more inclusive and accessible guidance technologies.

[0005] Therefore, there is a need for a system for assisting a user for operating third-party mobile applications to overcome a few or all drawbacks of the existing technologies.

Objects of the invention
[0006] An object of the present invention is to provide a system for assisting a user for operating third-party mobile applications.

[0007] Another object of the present invention is to provide a system for assisting a user for operating third-party mobile applications that improves the user’s ability to understand and follow operational steps on unfamiliar interfaces.

[0008] Another one object of the present invention is to provide a system for assisting a user for operating third-party mobile applications that reduces confusion and errors typically experienced by first-time or digitally-challenged users.

[0009] Another object of the present invention is to provide a system for assisting a user for operating third-party mobile applications that enhances user confidence by offering clear and timely assistance during application navigation.

[0010] Yet another object of the present invention is to provide a system for assisting a user for operating third-party mobile applications that promotes accessibility and inclusivity for users with limited technological familiarity.

[0011] One more object of the present invention is to provide a system for assisting a user for operating third-party mobile applications that supports a smoother learning experience and reduces reliance on external help or static instruction manuals.

Summary of the Invention
[0012] According to the present invention, there is provided a system for assisting a user for operating third-party mobile applications. The system may include an input module, an optical character recognition (OCR) engine, a natural-language processing engine, a database module, and a guidance generator. Further, the system may include an electronic device having a user interface module configured to allow a user to capture a screenshot of a third-party mobile application and to enter a user query in text or voice form. The input module may be configured to receive the screenshot of the third-party mobile application and the user query in text or voice form.

[0013] Further, the OCR engine may be configured to detect user-interface (UI) elements within the received screenshot and to output corresponding bounding-box coordinates. The OCR engine may identify app-specific visual markers to determine the identity of the third-party mobile application represented in the screenshot.

[0014] Further, the natural-language processing engine may be configured to interpret the user query based on the UI elements detected by the OCR engine. The natural-language processing engine may support multilingual input and output and constrains translations to match OCR-detected on-screen text for semantic consistency with the detected text in the screenshot.

[0015] Furthermore, the database module may be configured to store annotated screenshots, created workflows, and instructional sequences. The stored data may provide pre-created app guides for generating complete and contextually consistent instructional steps.

[0016] The guidance generator may be configured to access the stored data from the database module to generate step-wise instructions based on the detected UI elements. The step-wise instructions are generated on the basis of the interpretation of the user query and the stored information of the database module. The guidance generator may be configured to link each instruction to a corresponding actionable UI element detected by the OCR engine. The guidance generator may provide guidance by highlighting the actionable UI element and activating a haptic feedback and screen-brightness adjustment to provide tactile and visual confirmation during user interaction.

[0017] Further, the guidance generator may provide guidance by dimming non-relevant areas, highlighting the actionable UI element detected by the OCR engine, and activating a device vibration motor to provide haptic feedback when the user interacts with the highlighted actionable UI element.

[0018] Furthermore, the guidance generator may adjust the intensity, duration, or pattern of the haptic-feedback based on a confidence score associated with the actionable UI element linked to a corresponding instruction. The guidance generator may control the device vibration motor to generate different vibration patterns corresponding to different actionable UI elements, thereby enabling the user to distinguish between multiple instructional steps through tactile cues. The guidance generator may detect a mismatch between the received screenshot and a previously generated instruction step and, in response, generates updated instructions or requests an updated screenshot. Further, the guidance generator may link each instructional step textual content of a corresponding actionable UI element and to bounding-box coordinates generated by the OCR engine.

Brief description of drawings
[0019] The advantages and features of the present invention will be understood better with reference to the following detailed description and claims taken in conjunction with the accompanying drawings, wherein like elements are identified with like symbols, and in which:
[0020] FIG. 1 shows a block diagram of a system for assisting a user for operating third-party mobile applications in accordance with the present invention; and
[0021] FIG. 2 shows flowchart for a method for assisting a user for operating third-party mobile applications in accordance with the present invention.

Detailed description of the invention
[0022] An embodiment of this invention, illustrating its features, will now be described in detail. The words "comprising," "having," "containing," and "including," and other forms thereof, are intended to be equivalent in meaning and be open-ended in that an item or items following any one of these words is not meant to be an exhaustive listing of such item or items, or meant to be limited to only the listed item or items.

[0023] The present invention relates to a system for assisting a user for operating third-party mobile applications. The system is configured to support effective and intuitive user interaction with mobile interfaces by assisting users in understanding and navigating application workflows under a variety of usage conditions. The system is designed to provide consistent support across diverse user groups, including first-time smartphone users, individuals with limited digital literacy, elderly users, and users operating unfamiliar or complex applications. The present invention promotes improved usability, enhanced user confidence, and dependable assistance outcomes while reducing user frustration, minimizing operational errors, and enabling smoother interaction experiences across a wide range of mobile applications and usage scenarios.
[0024] The terms “first,” “second,” and the like, herein do not denote any order, quantity, or importance, but rather are used to distinguish one element from another, and the terms “a” and “an” herein do not denote a limitation of quantity, but rather denote the presence of at least one of the referenced item.

[0025] The disclosed embodiments are merely exemplary of the invention, which may be embodied in various forms.

[0026] Referring now to FIG. 1, a system (100) for assisting a user for operating third-party mobile applications in accordance with the present invention is illustrated. Specifically, the system (100) includes an input module (10), an optical character recognition (OCR) engine (20), a natural-language processing engine (30), a database module (40), and a guidance generator (50). The system (100) thereby provides a unified architectural framework for delivering multimodal, screen-specific, and hardware-assisted guidance to a user.

[0027] The system (100) includes an electronic device (60) configured to allow a user to provide a screenshot of a third-party mobile application and to submit a user query in text or voice form. The electronic device (60) serves as the front-end interaction layer through which the screenshot and the user query are collected before being forwarded to the input module (10) for subsequent processing. The electronic device (60) includes a user interface module (60A) configured to present a graphical interface for capturing screenshots and receiving user queries in text or voice form. It may be obvious to a person skilled in the art that the electronic device (60) may be implemented as a graphical interface on a mobile device, a web-based interface, or an integrated software panel within a host application.

[0028] The input module (10) is configured to receive a screenshot of a third-party mobile application and a user query in text or voice form from the electronic device (60). The input module (10) is configured to receive the screenshot provided by the user and ensure that the screenshot is captured in a format suitable for subsequent processing. The input module (10) allows the user to upload the screenshot directly from the user device and supports selection of a previously stored screenshot so that the system (100) can initiate guidance even when the user is not operating the app in real time. The input module (10) ensures that the received screenshot maintains adequate clarity for downstream analysis.

[0029] Further, the input module (10) is configured to receive the user query in text or voice form, with voice input converted into text to maintain consistency in processing. The input module (10) performs basic stabilization of both the screenshot and the user query so that the downstream components operate on clean and uniform inputs. The input module (10) ensures that the user query is aligned with the screenshot so that the system (100) can correctly interpret the user’s request in relation to the displayed elements. The input module (10) provides the screenshot to the OCR engine (20), thereby enabling the detection of user-interface (UI) elements contained within the screenshot. The input module (10) provides an audio-output capability for delivering spoken guidance. The input module (10) includes text-to-speech (TTS) functionality that converts the generated instructional steps into natural-sounding audio output, enabling hands-free or accessibility-oriented operation when visual reading of the instructions is impractical.

[0030] The OCR engine (20) is configured to detect user-interface (UI) elements within the received screenshot and to output corresponding bounding-box coordinates. The OCR engine (20) identifies textual labels, icons, and visually distinct regions so that the system (100) establishes a structural representation of the displayed screen. The OCR engine (20) performs this detection in a manner that preserves spatial relationships between the UI elements, enabling reliable downstream interpretation.

[0031] Further, the OCR engine (20) is configured to process variations in screen layouts arising from different device models, orientations, and display densities. The OCR engine (20) applies normalization operations to ensure that the extracted coordinates are consistent across screenshots having differing resolutions or aspect ratios. Further, the OCR engine (20) ensures that the detected UI elements remain stable even when the screenshot contains overlapping components or partially obscured regions.

[0032] Furthermore, the OCR engine (20) analyzes visual characteristics of the detected UI elements to distinguish between actionable UI elements and non-actionable background content. The OCR engine (20) refines each bounding box by evaluating edge clarity, textual prominence, and iconographic structure, thereby improving the ability of the system (100) to correctly target elements when generating guidance. The OCR engine (20) also supports repeated detection across multiple screenshots when the user uploads a sequence of screens during a multi-step interaction. The OCR engine (20) includes one or more AI (Artificial Intelligence) models, including but not limited to deep-learning-based computer-vision models, convolutional neural networks (CNNs), or transformer-based visual recognition models, to improve detection of UI elements and app-specific markers. It may be obvious to a person skilled in the art to use other AI model for enhanced accuracy, adaptability to diverse screen layouts, and robustness to partial occlusion or low-resolution screenshots.

[0033] Moreover, the OCR engine (20) outputs the detected UI elements and coordinates of the bounding-box to the natural-language processing engine (30), ensuring that the natural-language processing engine (30) can interpret the user query in direct relation to the UI context displayed in the screenshot. The OCR engine (20) thus provides structured visual information that allows the natural-language processing engine (30) to understand the user’s intent in context.

[0034] The natural-language processing engine (30) is configured to interpret the user query based on the UI elements detected by the OCR engine (20). The natural-language processing engine (30) receives the detected UI elements and bounding-box coordinates from the OCR engine (20), to associate the user query directly with specific visual elements displayed on the screenshot. This ensures that the natural-language processing engine (30) understands the user’s intent in the precise context of the current screen layout. The natural-language processing engine (30) includes one or more AI (Artificial Intelligence) models, including but not limited to deep-learning-based language models, transformer-based models, or sequence-to-sequence neural networks, to derive user intent, map queries to actionable UI elements, and generate context-aware, step-wise instructions. It may be obvious to a person skilled in the art to use any alternative AI model or reinforcement-learning techniques to improve semantic understanding, multilingual support, and alignment with OCR-detected UI elements.

[0035] Further, the natural-language processing engine (30) analyzes the user query, whether provided in text or voice form, to extract actionable instructions, commands, or questions. The natural-language processing engine (30) uses the UI element data provided by the OCR engine (20) to determine which parts of the screenshot are relevant to the user query. The natural-language processing engine (30) ensures that the generated guidance is contextually accurate and directly applicable to the interface shown, by aligning semantic information from the user query with the visual layout.

[0036] Furthermore, the natural-language processing engine (30) maps the interpreted user intent to the corresponding actionable UI elements identified by the OCR engine (20). The natural-language processing engine (30) links each instruction to specific bounding-box coordinates and UI element types, preparing a structured representation of the user’s requested actions. This structured data serves as a precise input for the guidance generator (50), enabling step-wise instruction generation. The natural-language processing engine (30) communicates the interpreted instructions and linked UI element data to the database module (40), which stores pre-created app guides. This allows the guidance generator (50) to access relevant database content and generate step-wise instructions that are consistent with both the user query and the known workflows for the detected application. The integration of the natural-language processing engine (30) with the database module (40) ensures that each instruction is accurate, contextually relevant, and enriched with pre-created guidance where applicable.

[0037] The database module (40) is configured to store annotated screenshots, created workflows, and instructional sequences, the stored data serving as a knowledge base that provides pre-created app guides useful for generating complete and contextually consistent instructional steps. The database module (40) maintains pre-created app guides that correspond to frequently used third-party applications, ensuring that relevant instructional content is readily available for generating complete and contextually consistent guidance.

[0038] Further, the database module (40) receives metadata from the natural-language processing engine (30), including interpreted user queries and the bounding-box coordinates of actionable UI elements detected by the OCR engine (20). The database module (40) uses this metadata to identify the most appropriate pre-created app guide, aligning the stored instructions and annotated screenshots with the current user context.

[0039] Furthermore, the database module (40) provides the retrieved app guides to the guidance generator (50), enabling the guidance generator (50) to generate step-wise instructions that are both contextually relevant and consistent with known workflows. The database module (40) ensures that each instruction is linked to specific UI elements detected by the OCR engine (20) and interpreted by the natural-language processing engine (30), enhancing accuracy and usability of the guidance. The database module (40) supports updating and expanding the stored knowledge base by including new annotated screenshots, instructional sequences, or user feedback. The database module (40) continuously updates the pre-created guides, thereby improving the performance of the guidance generator (50) over time and maintaining alignment with evolving third-party application interfaces.

[0040] The guidance generator (50) is configured to access the stored data from the database module (40) to enhance contextual understanding of the UI elements detected by the OCR engine (20) and interpreted by the natural-language processing engine (30). The guidance generator (50) uses the retrieved app guides and the user query interpretation to generate step-wise instructions that are relevant, complete, and contextually aligned with the screenshot provided by the input module (10). The guidance generator (50) links each instruction to a corresponding actionable UI element identified by the OCR engine (20). Each instructional step includes textual content associated with bounding-box coordinates corresponding to the actionable UI element. The guidance generator (50) includes one or more AI (Artificial Intelligence) models, including but not limited to deep-learning, transformer-based, or reinforcement-learning models, to adaptively link instructions to actionable UI elements, prioritize instructional steps, optimize visual highlighting and haptic feedback, and tailor guidance to the user’s interaction context. It may be obvious to a person skilled in the art to use any alternative AI model, or adaptive learning techniques to improve step sequencing, context-awareness, and user-specific guidance efficiency.

[0041] Further, the guidance generator (50) overlays interactive visual highlights on these actionable elements, providing the user with a hardware-assisted visual guidance layer that clearly indicates where the next interaction should occur. The guidance generator (50) is configured to trigger haptic-feedback activation and screen-brightness-adjustment actions in synchronization with the highlighted UI elements. The guidance generator (50) ensures that upon user interaction with a highlighted actionable UI element, the device vibration motor generates tactile confirmation while the actionable region remains visually prominent, and non-relevant areas of the screen are dimmed to focus the user’s attention.

[0042] Furthermore, the guidance generator (50) monitors user interactions and dynamically adjusts instructions based on updated input from the natural-language processing engine (30) and OCR engine (20). The guidance generator (50) ensures that each instructional step remains accurate, contextually consistent, and aligned with both the detected UI elements and the pre-created guides stored in the database module (40), thereby delivering a seamless and intuitive guidance experience.

[0043] Moreover, the guidance generator (50) is configured to compute a confidence score for each actionable UI element based on factors including OCR text-recognition certainty, UI-element shape consistency, historical user-interaction patterns, and correlation strength between the interpreted user query and the detected UI element. The guidance generator (50) selectively adjusts the intensity, duration, or vibration pattern of the haptic-feedback response in proportion to the computed confidence score, such that higher-confidence actionable elements trigger shorter or sharper tactile confirmations, while lower-confidence elements trigger longer or more pronounced vibration patterns. This confidence-adaptive haptic strategy improves clarity and reduces ambiguity for the user during interaction.

[0044] The guidance generator (50) controls the device vibration motor to generate distinct vibration patterns corresponding to different actionable UI elements or different steps within a workflow. The vibration patterns may include short pulses, long pulses, multi-pulse sequences, or periodic vibration bursts, each pattern representing a different UI element type or instructional category. The vibration pattern allows the user to differentiate among multiple guided steps through tactile cues.

[0045] The guidance generator (50) is configured to detect a mismatch between a currently received screenshot and a previously generated instruction step. The mismatch detection process compares the layout, detected UI elements, and bounding-box coordinates of the new screenshot with the expected screen state associated with the previous instruction. When discrepancies exceed a predetermined threshold, such as missing UI elements, unexpected screen transitions, or altered visual structures, the guidance generator (50) automatically regenerates updated instructions or prompts the user to capture and upload a new screenshot. The mismatch-resolution mechanism ensures that the instructional flow remains accurate and relevant to the user’s current interface context.

[0046] Further, the system (100) includes an application-identity inference mechanism configured to determine which third-party mobile application is represented in the received screenshot. The mechanism operates by analyzing the screenshot using a combination of visual, textual, and structural cues. Initially, a visual-classification model extracts global image features relating to layout patterns, color schemes, iconography, and interface structures, and compares these features with stored feature profiles of known applications to generate a preliminary identity score. In parallel, the OCR engine (20) provides on-screen text such as page titles, menu labels, and app-specific terminology, which are matched against app-specific vocabularies using keyword-matching and similarity-scoring techniques. The mechanism further performs logo- and app-bar analysis to detect application icons or brand markers through template matching or lightweight object-detection models. When available, metadata embedded within the screenshot, such as package-name indicators or activity labels, is extracted to provide an additional high-confidence signal. The system (100) combines these multiple signals through a weighted decision model to infer the most probable application identity, enabling the database module (40) to retrieve the corresponding pre-created instructional guide and ensuring that the guidance generator (50) produces app-specific, contextually accurate instructional steps.

[0047] In an embodiment, the OCR engine (20) and associated visual-analysis components are configured to identify app-specific visual markers within the screenshot in order to determine the identity of the third-party mobile application. Such markers may include unique iconography, color themes, brand logos, page-layout structures, or characteristic typography. By identifying these app-specific features, the system (100) determines the target application and enables retrieval of the corresponding pre-created workflows and annotated screenshots from the database module (40).

[0048] In another embodiment, the natural-language processing engine (30) supports multilingual input and output, enabling the system (100) to interpret user queries and generate guidance instructions across multiple languages. The natural-language processing engine (30) includes constrained-translation mechanisms that maintain semantic alignment with OCR-detected on-screen text so that translated output preserves terminology appearing within the screenshot. The constrained-translation mechanisms ensures that translated instructions remain consistent with the user’s interface and avoids generating references to text that does not appear on the displayed screen.

[0049] In an embodiment, the system (100) interacts with external application-store APIs (e.g., Google Play Store™ and Apple App Store™ APIs) to retrieve metadata related to the detected third-party mobile application. The metadata may include the application category, developer information, version number, supported features, and user interface conventions. The retrieved information improves the accuracy of query interpretation and instruction generation by enabling the system to tailor guidance based on known app attributes.

[0050] By way of a non-limiting example, operation of the system (100) begins when the user interface (60) receives a screenshot and user query, and the input module (10) forwards the received screenshot to the OCR engine (20), which identifies the user-interface (UI) elements and generates bounding-box coordinates representing the detected regions of interest. The input module (10) forwards the received screenshot to the OCR engine (20), which identifies the user-interface (UI) elements and generates bounding-box coordinates representing the detected regions of interest. Further, the OCR engine (20) outputs the detected UI elements to the natural-language processing engine (30), enabling the natural-language processing engine (30) to interpret the user query in direct relation to the visual context of the screenshot. The natural-language processing engine (30) then produces a context-aware interpretation of user intent, which is passed to the guidance generator (50) for producing step-wise instructions based on the detected UI layout and the retrieved information from the database module (40).

[0051] Once the natural-language processing engine (30) delivers the interpreted user intent, the guidance generator (50) accesses the database module (40) to retrieve pre-created app guides, annotated screenshots, and instructional sequences corresponding to the identified third-party application. The guidance generator (50) combines the interpreted intent with the stored knowledge base and the bounding-box coordinates generated by the OCR engine (20) to produce a set of structured guidance steps that align with the user’s query and the specific UI context visible within the uploaded screenshot. Further, the guidance generator (50) links each generated instruction to a corresponding actionable UI element, ensuring that each step directly reflects an element detected by the OCR engine (20).

[0052] After generating the structured guidance steps, the guidance generator (50) provides a hardware-assisted visual guidance layer by overlaying interactive highlights on the actionable UI elements referenced in each instruction. The guidance generator (50) renders each highlighted region such that it remains in normal brightness while non-relevant regions of the screenshot are subtly dimmed, thereby improving user focus on the intended UI element. Further, upon user interaction with a highlighted actionable UI element, the guidance generator (50) activates the device’s vibration motor to generate a tactile confirmation pulse, allowing the user to physically sense the completion of each guided action.
[0053] The combined operation of the input module (10), the OCR engine (20), the natural-language processing engine (30), the database module (40), and the guidance generator (50) thereby enables the system (100) to deliver multimodal, context-specific, and hardware-assisted guidance for operating third-party mobile applications. This coordinated process ensures that each user receives clear, step-wise support that is visually highlighted, tactically confirmed, and linguistically adapted to the specific screen and query under consideration.

[0054] Referring now to Figure 2, a flowchart of a method (200) for providing guidance for operating a mobile application is illustrated.

[0055] In an aspect, a method (200) for providing guidance for operating a mobile application is provided in accordance with the present invention. Referring to Figure 2, the method (200) is described in conjunction with the system (100) comprising an input module (10), an optical character recognition (OCR) engine (20), a natural-language processing engine (30), a database module (40), and a guidance generator (50). The method (200) includes the following steps:

[0056] The method (200) starts at step (210).

[0057] At step (220), a screenshot of the third-party mobile application and a user query in text or voice form are provided to the input module (10). The input module (10) obtains the screenshot uploaded by the user, shared through a system-level screenshot tool, or captured directly from the host device. The input module (10) collects the screenshot uploaded or captured by the user and accepts the corresponding user query. The input module (10) performs preprocessing such as format standardization and noise reduction to ensure that the screenshot is suitable for analysis by downstream components.

[0058] At step (230), the screenshot is analyzed using the OCR engine (20) to detect user-interface (UI) elements and generate bounding-box coordinates. The OCR engine (20) extracts textual content, identifies UI structures such as icons, buttons, and input fields, and assigns bounding-box coordinates to each detected UI element. The OCR engine (20) thereby produces the spatial and textual UI representation required for downstream interpretation.

[0059] At step (240), the user query is interpreted based on the detected UI elements using the natural-language processing engine (30). The natural-language processing engine (30) determines user intent and correlates the interpreted intent with the UI context derived from the OCR engine (20). The natural-language processing engine (30) thereby produces context-aware parameters that align the user’s query with the visual content present in the screenshot.

[0060] At step (250), an annotated screenshots, workflows, and instructional sequences stored in the database module (40) is accessed to enhance contextual understanding of the detected UI elements. The database module (40) provides pre-created app guides, annotated screenshots, created workflows, and instructional sequences that contribute to improving the completeness and contextual consistency of the forthcoming instructional steps.

[0061] At step (260), a step-wise instructions linked to actionable UI elements selected from the detected UI elements are generated using the guidance generator (50), the generation being based on both the interpreted user query and the accessed database information. The guidance generator (50) aligns each instruction with a specific actionable UI element, creating a structured sequence appropriate to the user’s intent and the detected UI layout.

[0062] At step (270), visual highlights are overlaid on actionable UI elements associated with the instructional steps. Using the bounding-box coordinates produced by the OCR engine (20), the guidance generator (50) renders visual emphasis on each actionable region while preparing the interface for sensory feedback.

[0063] At step (280), a haptic-feedback actions and brightness-adjustment actions are activated in association with the highlighted actionable UI elements. The guidance generator (50) associates tactile and brightness-control responses with each corresponding actionable UI element to enable synchronized sensory reactions during user interaction.

[0064] At step (290), a tactile and visual confirmation is provided when the user interacts with the highlighted actionable UI element. The guidance generator (50) triggers the device vibration motor to generate haptic feedback and adjusts screen brightness such that the actionable region remains at normal brightness while surrounding areas are dimmed.

[0065] The method (200) ends at step (300).

[0066] The present invention has the advantage of providing a system (100) for assisting a user for operating third-party mobile applications capable of delivering accurate, context-aware, and user-adaptive assistance across a wide range of application environments. The input module (10) ensures accurate capture of the user’s on-screen context, while the OCR engine (20) reliably identifies UI elements needed for precise visual mapping. The natural-language processing engine (30) interprets the user query in relation to the detected UI elements, and the database module (40) supplies pre-created guides that support consistent and relevant instruction generation. The guidance generator (50) combines these inputs to produce step-wise instructions linked to actionable UI elements and delivers hardware-assisted visual and tactile feedback to aid user interaction. The present invention thus provides an integrated guidance solution that improves usability and reduces user effort when operating third-party mobile applications.

[0067] The foregoing descriptions of specific embodiments of the present invention have been presented for purposes of illustration and description. They are not intended to be exhaustive or to limit the present invention to the precise forms disclosed, and obviously, many modifications and variations are possible in light of the above teaching. The embodiments were chosen and described in order to explain the principles of the present invention best and its practical application, to thereby enabling others skilled in the art to best utilize the present invention and various embodiments with various modifications as are suited to the particular use contemplated. It is understood that various omission and substitutions of equivalents are contemplated as circumstance may suggest or render expedient, but such are intended to cover the application or implementation without departing from the scope of the claims of the present invention.
, Claims:We Claim:
1. A system (100) for assisting a user for operating third-party mobile applications, the system (100) comprising:
an electronic device (60) having a user interface module (60A) configured to allow a user to capture a screenshot of a third-party mobile application and enter a user query in text or voice form;
an input module (10) configured to receive the screenshot of the third-party mobile application and the user query in text or voice form;
an optical character recognition (OCR) engine (20) configured to detect user-interface (UI) elements within the received screenshot for generating corresponding bounding-box coordinates;
a natural-language processing engine (30) configured to interpret the user query based on the UI elements detected by the OCR engine (20);
a database module (40) configured to store annotated screenshots, created workflows, and instructional sequences, wherein the stored data provides pre-created app guides for generating complete and contextually consistent instructional steps;
a guidance generator (50) configured to access the stored data from the database module (40) to generate step-wise instructions based on the detected UI elements, wherein, the step-wise instructions are generated on the basis of the interpretation of the user query and the stored information of the database module (40);
wherein the guidance generator (50) is configured to link each instruction to a corresponding actionable UI element detected by the OCR engine (20), and provides guidance by highlighting the actionable UI element and activating a haptic feedback and screen-brightness adjustment to provide tactile and visual confirmation during user interaction;
wherein the OCR engine (20), the natural-language processing engine (30), and the guidance generator (50) include one or more artificial intelligence (AI) models configured to detect UI elements, derive user intent, link each instruction corresponding with actionable UI elements, and provide tactile and visual confirmation during user interaction.

2. The system (100) for assisting a user for operating third-party mobile applications as claimed in claim 1, wherein the natural-language processing engine (30) supports multilingual input and output and matches translated responses with OCR-detected on-screen text for semantic consistency with the detected with the detected UI elements.

3. The system (100) for assisting a user for operating third-party mobile applications as claimed in claim 1, wherein the guidance generator (50) provides guidance by dimming non-relevant areas, highlighting the actionable UI element detected by the OCR engine (20), and activating a device vibration motor to provide haptic feedback when the user interacts with the highlighted actionable UI element.

4. The system (100) for assisting a user for operating third-party mobile applications as claimed in claim 1, wherein the guidance generator (50) adjusts the intensity, duration, or pattern of the haptic-feedback based on a confidence score associated with the actionable UI element linked to a corresponding instruction.

5. The system (100) for assisting a user for operating third-party mobile applications as claimed in claim 1, wherein the guidance generator (50) controls the device vibration motor to generate different vibration patterns corresponding to different actionable UI elements, thereby enabling the user to distinguish between multiple instructional steps through tactile cues.

6. The system (100) for assisting a user for operating third-party mobile applications as claimed in claim 1, wherein the guidance generator (50) detects a mismatch between the received screenshot and a previously generated instruction step and, in response, generates updated instructions or requests an updated screenshot.

7. The system (100) for assisting a user for operating third-party mobile applications as claimed in claim 1, wherein the guidance generator (50) links each instructional step textual content of a corresponding actionable UI element and to bounding-box coordinates generated by the OCR engine (20).

8. A method (200) for assisting a user for operating third-party mobile applications, the method (200) comprising the steps of:
receiving a screenshot of the third-party mobile application and a user query in text or voice form using an input module (10);
analyzing the screenshot to detect user-interface (UI) elements and generate bounding-box coordinates using an optical character recognition OCR engine (20);
interpreting the user query based on the detected UI elements using a natural-language processing engine (30);
accessing annotated screenshots, workflows, and instructional sequences stored in a database module (40) to enhance contextual understanding of the detected UI elements;
generating step-wise instructions linked to actionable UI elements using a guidance generator (50) based on both the interpreted user query and the accessed database information;
highlighting the actionable UI elements associated with the instructional steps;
activating a haptic feedback and screen-brightness adjustment associated with the highlighted actionable UI elements; and
providing tactile and visual confirmation when the user interacts with a highlighted actionable UI element

Documents