Sign In to Follow Application
View All Documents & Correspondence

Method And System For Automated Design Rules Extraction And Integration Using Large Language Models (Llms)

Abstract: A method, system and a non-transitory computer-readable medium of automated design rules extraction and integration using Large Language Models (LLMs) are disclosed. A processor fetches a plurality of design documents from one or more repositories. The processor extracts multimodal data from each of the plurality of design documents. The processor interlinks data of different data modalities within the extracted multimodal data to create a plurality of contextually linked datasets. The processor determines one or more design rules based on the plurality of contextually linked datasets using an LLM. The process generates a set of machine-understandable rules based on the one or more design rules. The processor integrates each of the set of machine-understandable rules into a computing tool for enabling automated design rules enforcement within the computing tool. [Fig. 1]

Get Free WhatsApp Updates!
Notices, Deadlines & Correspondence

Patent Information

Application #
Filing Date
09 March 2026
Publication Number
18/2026
Publication Type
INA
Invention Field
COMPUTER SCIENCE
Status
Email
Parent Application

Applicants

HCL Technologies Limited
806, Siddharth, 96, Nehru Place, New Delhi, 110019, India

Inventors

1. Christine Zuzart
HCL Technologies Ltd, Tower 7, Ground Floor, Level Upper, Magarpatta SEZ, Hadapsar Industrial Estate, Pune, Maharashtra, 411013, India
2. Rahul Rajadhyaksha
703-A, 7th Floor, Reliable Tech Park, Gut 31 Kalwa Industrial Area, Airoli, Navi Mumbai, Maharashtra, 400708, India
3. Suvarna Bhagat
Vishwakarma Institute of Information Technology (VIIT) No. 3/4, Kapil Nagar, Kondhwa (Budruk), Pune, Maharashtra, 411048, India
4. Parikshit N Mahalle
Vishwakarma Institute of Information Technology (VIIT) No. 3/4, Kapil Nagar, Kondhwa (Budruk), Pune, Maharashtra, 411048, India

Claims

1. A method for automated design rules extraction and integration using Large Language Models (LLMs) comprising: fetching, by a processor, a plurality of design documents from one or more repositories; extracting, by the processor, multimodal data from each of the plurality of design documents; interlinking, by the processor, data of different data modalities within the extracted multimodal data by associating portions of the multimodal data corresponding to different modalities based on at least one of positional proximity between the portions, reference indicators linking the portions, or semantic correlation between the portions to create a plurality of contextually linked datasets; determining, by the processor, one or more rules based on the plurality of contextually linked datasets using an LLM; generating, by the processor, a set of machine-understandable rules based on the one or more rules; and integrating, by the processor, each of the set of machine-understandable rules into a computing tool for enabling automated rules enforcement within the computing tool.

2. The method of claim 1, wherein each of the plurality of design documents is in a predefined format, and wherein the predefined format comprises a text file, a Portable Document Format (PDF) file, a spreadsheet file, an image file, a Hyper Text Markup Language (HTML) file, a word processing file, a presentation file, a Computer-Aided Design (CAD) drawing file, and a scanned document.

3. The method of claim 1, wherein the multimodal data comprises images, text, tables, charts, and equations.

4. The method of claim 1, wherein determining the one or more rules comprises: classifying, by the processor, the extracted multimodal data into a predefined category, wherein the predefined category comprises a rule category and a non-rule category.

5. The method of claim 1, wherein generating the set of machine-understandable rules comprising: identifying, by the processor, a set of key components from the one or more rules, wherein the set of key components comprises a function, an object, an attribute, an operator, a value, and a unit.

6. The method of claim 1, further comprising: continuously evolving, by the processor, a process for determining the one or more rules based on feedback from a user.

7. The method of claim 1, further comprising: storing, by the processor, the set of machine-understandable rules in a database with associated provenance metadata; and updating, by the processor, the set of machine-understandable rules stored in the database based on one of a new input document or an update to an existing input document of the plurality of design documents.

8. A system for automated rules extraction and integration using Large Language Models (LLMs) comprising: a processor; and a memory communicably coupled to the processor, wherein the memory stores processor-executable instructions, which, when executed by the processor, cause the processor to: fetch a plurality of design documents from one or more repositories; extract multimodal data from each of the plurality of design documents; interlink data of different data modalities within the extracted multimodal data by associating portions of the multimodal data corresponding to different modalities based on at least one of positional proximity between the portions, reference indicators linking the portions, or semantic correlation between the portions to create a plurality of contextually linked datasets; determine one or more rules based on the plurality of contextually linked datasets using an LLM; generate a set of machine-understandable rules based on the one or more rules; and integrate each of the set of machine-understandable rules into a computing tool for enabling automated rules enforcement within the computing tool.

9. The system of claim 8, wherein each of the plurality of design documents is in a predefined format, and wherein the predefined format comprises a text file, a Portable Document Format (PDF) file, a spreadsheet file, an image file, a Hyper Text Markup Language (HTML) file, a word processing file, a presentation file, a Computer-Aided Design (CAD) drawing file, and a scanned document.

10. The system of claim 8, wherein the multimodal data comprises images, text, tables, charts, and equations.

11. The system of claim 8, wherein to determine the one or more rules, the processor-executable instructions, which, when executed by the processor, cause the processor to: classify the extracted multimodal data into a predefined category, wherein the predefined category comprises a rule category and a non-rule category.

12. The system of claim 8, wherein generating the set of machine-understandable rules causes the processor to: identify a set of key components from the one or more rules, wherein the set of key components comprises a function, an object, an attribute, an operator, a value, and a unit.

13. The system of claim 8, wherein the processor-executable instructions, which, when executed by the processor, cause the processor to: continuously evolve a process for determining the one or more rules based on feedback from a user.

14. The system of claim 8, wherein the processor-executable instructions, which, when executed by the processor, cause the processor to: store the set of machine-understandable rules in a database with associated provenance metadata; and update the set of machine-understandable rules stored in the database based on one of a new input document or an update to an existing input document of the plurality of design documents.

15. A non-transitory computer-readable medium storing computer-executable instructions for automated rules extraction and integration using Large Language Models (LLMs), the computer-executable instructions configured for: fetching a plurality of design documents from one or more repositories; extracting multimodal data from each of the plurality of design documents; interlinking data of different data modalities within the extracted multimodal data by associating portions of the multimodal data corresponding to different modalities based on at least one of positional proximity between the portions, reference indicators linking the portions, or semantic correlation between the portions to create a plurality of contextually linked datasets; determining one or more rules based on the plurality of contextually linked datasets using an LLM; generating a set of machine-understandable rules based on the one or more rules; and integrating each of the set of machine-understandable rules into a computing tool for enabling automated rules enforcement within the computing tool.

16. The non-transitory computer-readable medium of claim 15, wherein each of the plurality of design documents is in a predefined format, and wherein the predefined format comprises a text file, a Portable Document Format (PDF) file, a spreadsheet file, an image file, a Hyper Text Markup Language (HTML) file, a word processing file, a presentation file, a Computer-Aided Design (CAD) drawing file, and a scanned document.

17. The non-transitory computer-readable medium of claim 15, wherein for determining the one or more rules, the computer-executable instructions are further configured for: classifying the extracted multimodal data into a predefined category, wherein the predefined category comprises a rule category and a non-rule category.

18. The non-transitory computer-readable medium of claim 15, wherein generating the set of machine-understandable rules comprising: identifying a set of key components from the one or more rules, wherein the set of key components comprises a function, an object, an attribute, an operator, a value, and a unit.

19. The non-transitory computer-readable medium of claim 15, wherein the computer-executable instructions are further configured for: continuously evolving a process for determining the one or more rules based on feedback from a user.

20. The non-transitory computer-readable medium of claim 15, wherein the computer-executable instructions are further configured for: storing the set of machine-understandable rules in a database with associated provenance metadata; and updating the set of machine-understandable rules stored in the database based on one of a new input document or an update to an existing input document of the plurality of design documents.

Specification

Description:DESCRIPTION
TECHNICAL FIELD
[0001] This disclosure relates generally to document processing, and more particularly to a method and a system for automated design rule extraction and integration using Large Language Models (LLMs).
BACKGROUND
[0002] Design rules play a critical role in ensuring that product designs meet manufacturing, quality, cost, and regulatory requirements. These design rules are developed over time by the experts and are documented in various engineering repositories such as design handbooks, standards, and engineering change records. In many organizations, the information about design rules is distributed across multiple systems and document formats, including text, tables, images, and scanned legacy documents. As the volume and complexity of such documentation increase over time, it becomes difficult to consistently identify, interpret, and apply relevant design rules during the design process. This creates a need for more efficient and automated management of design knowledge.
[0003] Conventionally, many software tools in use incorporate design rules into CAD (Computer Aided Design) environments in order to validate designs instantly. These tools lessen the need for manual referencing and aid in design error prevention. However, these solutions rely on domain experts and developers manually extracting and encoding rules from documents. This process is time-consuming, prone to errors, and frequently unable to keep up with complex and dynamic multi-format documents that include text, images, tables, and handwritten notes. Therefore, there is still difficulty in effectively extracting and updating design rules from documents, even with the availability of current automated tools.
[0004] Therefore, there is a need for an improved methodology to speed up these procedures of extracting and applying the design rules from different repositories.
SUMMARY OF THE INVENTION
[0005] In an embodiment, a method of automated design rules extraction and integration using Large Language Models (LLMs) is disclosed. The method may include fetching, by a processor, a plurality of design documents from one or more repositories. The method may further include extracting, by the processor, multimodal data from each of the plurality of design documents. The method may further include interlinking, by the processor, data of different data modalities within the extracted multimodal data by associating portions of the multimodal data corresponding to different modalities based on at least one of positional proximity between the portions, reference indicators linking the portions, or semantic correlation between the portions to create a plurality of contextually linked datasets. The method may include determining, by the processor, one or more design rules based on the plurality of contextually linked datasets using an LLM. The method may further include generating, by the processor, a set of machine-understandable rules based on the one or more design rules. The method may further include integrating, by the processor, each of the set of machine-understandable rules into a computing tool for enabling automated design rules enforcement within the computing tool.
[0006] In another embodiment, a system for automated design rules extraction and integration using Large Language Models (LLMs) is disclosed. The system may include a processor and a memory communicatively coupled to the processor. The memory stores processor-executable instructions, which, when executed by the processor, cause the processor to fetch a plurality of design documents from one or more repositories. The processor-executable instructions may cause the processor to extract multimodal data from each of the plurality of design documents. The processor-executable instructions may cause the processor to interlink data of different data modalities within the extracted multimodal data by associating portions of the multimodal data corresponding to different modalities based on at least one of positional proximity between the portions, reference indicators linking the portions, or semantic correlation between the portions to create a plurality of contextually linked datasets. The processor-executable instructions may further cause the processor to determine one or more design rules based on the plurality of contextually linked datasets using an LLM. The processor-executable instructions may further cause the processor to generate a set of machine-understandable rules based on one or more design rules. The processor-executable instructions may further cause the processor to integrate each of the set of machine-understandable rules into a computing tool for enabling automated design rules enforcement within the computing tool.
[0007] In yet another embodiment, a non-transitory computer-readable medium storing computer-executable instructions for automated design rules extraction and integration using Large Language Models (LLMs) is disclosed. The computer-executable instructions may be configured for fetching a plurality of design documents from one or more repositories. The computer-executable instructions may be configured for extracting multimodal data from each of the plurality of design documents. The computer-executable instructions may be configured for interlinking data of different data modalities within the extracted multimodal data by associating portions of the multimodal data corresponding to different modalities based on at least one of positional proximity between the portions, reference indicators linking the portions, or semantic correlation between the portions to create a plurality of contextually linked datasets. The computer-executable instructions may be configured for determining one or more design rules based on the plurality of contextually linked datasets using an LLM. The computer-executable instructions may be configured for generating a set of machine-understandable rules based on the one or more design rules. The computer-executable instructions may be configured for integrating each of the set of machine-understandable rules into a computing tool for enabling automated design rules enforcement within the computing tool.
[0008] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention, as claimed.
BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The accompanying drawings, which are incorporated in and constitute a part of this disclosure, illustrate exemplary embodiments and, together with the description, serve to explain the disclosed principles.
[0010] FIG. 1 is a block diagram of a system for automated rules extraction and integration using Large Language Models (LLMs), in accordance with an embodiment of the present disclosure.
[0011] FIG. 2 is a functional block diagram of various modules within a memory of a computing device configured for automated rules extraction and integration using Large Language Models (LLMs), in accordance with an embodiment of the present disclosure.
[0012] FIG. 3 illustrates a functional block diagram of a document input module, in accordance with an embodiment of the present disclosure.
[0013] FIG. 4 illustrates a functional block diagram of a multimodal data extraction module 204, in accordance with an embodiment of the present disclosure.
[0014] FIG. 5 illustrates a functional block diagram describing generation of set of machine-understandable rules, in accordance with an embodiment of the present disclosure.
[0015] FIG. 6 illustrates a flowchart of automated rules extraction and integration using Large Language Models (LLMs), in accordance with an exemplary embodiment of the present disclosure.
[0016] FIG. 7 illustrates a flow diagram of an exemplary process for interlinking data of different modalities, in accordance with an embodiment of the present disclosure.
[0017] FIG. 8 illustrates a first exemplary design document, in accordance with an embodiment of the present disclosure.
[0018] FIG. 9 illustrates a second exemplary design document, in accordance with an embodiment of the present disclosure.
[0019] FIG. 10 illustrates a third exemplary design document, in accordance with an embodiment of the present disclosure.
DETAILED DESCRIPTION OF THE DRAWINGS
[0020] Exemplary embodiments are described with reference to the accompanying drawings. Wherever convenient, the same reference numbers are used throughout the drawings to refer to the same or like parts. While examples and features of disclosed principles are described herein, modifications, adaptations, and other implementations are possible without departing from the scope of the disclosed embodiments. It is intended that the following detailed description be considered as exemplary only, with the true scope being indicated by the following claims. Additional illustrative embodiments are listed.
[0021] Further, the phrases “in some embodiments”, “in accordance with some embodiments”, “in the embodiments shown”, “in other embodiments”, and the like, mean a particular feature, structure, or characteristic following the phrase is included in at least one embodiment of the present disclosure and may be included in more than one embodiment. In addition, such phrases do not necessarily refer to the same or different embodiments. It is intended that the following detailed description be considered exemplary only, with the true scope and spirit being indicated by the following claims.
[0022] As discussed earlier, organizations rely heavily on design rules to ensure that product designs are manufacturable, reliable, cost-effective, and compliant with industry and regulatory requirements. Design rules are derived from accumulated engineering experience and are documented across a wide range of repositories, including design handbooks, standard publications, Engineering Change Orders (ECOs), and internal specifications. Over time, this knowledge base grows significantly and becomes distributed across multiple systems such as Product Lifecycle Management (PLM) systems, ECO management platforms, and the like.
[0023] Despite the importance of design rules, their practical use during product design is limited by the way they are documented and maintained. Design rules are typically stored as unstructured or semi-structured content across documents containing text, tables, images, and diagrams, distributed over multiple enterprise systems. Interpreting these documents and converting the embedded knowledge into design rules is mostly manual, time-consuming, and difficult to scale.
[0024] The disclosed invention overcomes these limitations through a design rule extraction and integration mechanism using Large Language Models (LLMs). The method and system of the disclosed invention will be described in detail hereinafter.
[0025] Referring now to FIG. 1, a block diagram of an exemplary system 100 for automated design rules extraction and integration using Large Language Models (LLMs) is illustrated, in accordance with an embodiment of the current disclosure.
[0026] The system 100 may include a computing device 102, a metadata database 104, and a data server 106, which are communicably coupled to each other through a communication network 108 (which may be wired or wireless).
[0027] The metadata database 104 stores structured and unstructured data associated with design documents and extracted rule information. In some embodiments, the metadata 104 stores provenance metadata indicating source document, page location, and version history.
[0028] The data server 106 hosts one or more repositories 107 storing design documents. The repositories 107 may include enterprise data sources such as product lifecycle management (PLM) systems, engineering change order (ECO) systems, document management systems, or cloud-based storage services. The data server 106 may be implemented as a cloud server, distributed server system, or on-premise server.
[0029] The computing device 102, for example, may include but is not limited to a smartphone, a laptop computer, a desktop computer, a notebook, a workstation, a server, a portable computer, a handheld, or a mobile device. The computing device 102 may include a processor 110, a memory 112, and a display 114 communicably coupled to each other. The display 114 may further include a User Interface (UI) 116.
[0030] The computing device 102 may perform automated rules extraction and integration using Large Language Models (LLMs) using a set of instructions processed by the processor 110. Examples of the processor 110 may include but are not limited to, a complex instruction set computer (CISC), a reduced instruction set computer (RISC), a very long instruction word (VILW), a General-purpose Processors (GPPs), microcontrollers (MCUs), a Graphics Processing Units (GPUs), Application-Specific Integrated Circuits (ASICs), a Ultra-Low-Power (IoT/Embedded), and the like, on a chip processors or other future processors.
[0031] The set of instructions executed by the processor 110 may be stored in the memory 112, which, when executed by the processor 110, causes the processor 110 to perform automated rules extraction and integration using Large Language Models (LLMs), as will be discussed in greater detail below.
[0032] In an embodiment, the memory 112 may be a non-volatile memory or a volatile memory. Examples of non-volatile memory may include, but are not limited to, flash memory, a Read Only Memory (ROM), a Programmable ROM (PROM), Erasable PROM (EPROM), and an Electrically EPROM (EEPROM) memory. Further, examples of volatile memory may include but are not limited to Dynamic Random Access Memory (DRAM) and Static Random-Access Memory (SRAM). The memory 112 may enable the efficient storage, retrieval, and processing of the LLM models used for extracting the design rules and for the likewise.
[0033] The user interface 116 may include a variety of interface(s), for example, interfaces for data input and output devices, and the like. The user interface 116 may facilitate the input of instructions by a user communicating with the computing device 102 to enable automated rules extraction and integration using Large Language Models (LLMs). In some embodiments, the display 114 may be external to the computing device 102 and may be connected to the computing device 102 through wired or wireless network protocols, such as HDMI, Bluetooth®, Wi-Fi, or the like.
[0034] The communication network 108 may be a wired or a wireless network or a combination thereof. The communication network 108 can be implemented as one of the different types of networks, such as, but not limited to, Ethernet IP network, intranet, local area network (LAN), wide area network (WAN), the internet, Wi-Fi, LTE network, CDMA network, 5G, and the like. Further, the communication network 108 can either be a dedicated network or a shared network. The shared network represents an association of the different types of networks that use a variety of protocols, for example, Hypertext Transfer Protocol (HTTP), Transmission Control Protocol/Internet Protocol (TCP/IP), Wireless Application Protocol (WAP), and the like, to communicate with one another. Further, the communication network 108 can include a variety of network devices, including routers, bridges, servers, computing devices, storage devices, and the like.
[0035] The processor 110 may be configured to fetch a plurality of design documents from one or more repositories. The processor 110 may further be configured to extract multimodal data from each of the plurality of design documents. The processor 110 may be configured to interlink data of different data modalities within the extracted multimodal data to create a plurality of contextually linked datasets. The processor 110 may be configured to determine one or more rules based on the plurality of contextually linked datasets using an LLM. Upon determination of one or more rules, the processor 110 may be configured to generate a set of machine-understandable rules based on the one or more rules. The processor 110 may further be configured to integrate each of the set of machine-understandable rules into a computing tool for enabling automated rules enforcement within the computing tool. This is further explained in detail in conjunction with FIG. 2-10.
[0036] As discussed before, the computing device 102 is communicably coupled to the database 104 via the communication network 108. The database 104 may be an On-Device (Embedded) database, a client-server architecture, or the like. The database 104 may be configured to store the data of one or more types, which may include, but are not limited to, structured data (tables with rows and columns, for example, design rules explained for manufacturing building and chemicals with images and tables and the like ), semi-structured data (for example, in JSON (JavaScript Object Notation), XML (Extensible Markup Language), Unstructured data (for example, images, text files, PDFs) and the like.
[0037] The computing device 102 and the database 104 may be communicably coupled to the data server 106 via the communication network 108. The data server 106 may be enabled in a remote cloud server or a co-located server and may store a large dataset of the system 100 to perform automated rules extraction and integration using Large Language Models (LLMs). In an embodiment, the data server 106 may store user inputs received via the user interface 116 (i.e., the plurality of design documents) or output generated by the computing device 102 (e.g., a set of machine-understandable rules).
[0038] FIG. 2 illustrates a functional block diagram 200 of various modules within the memory 112 of the computing device 102 configured for automated rules extraction and integration using Large Language Models (LLMs), in accordance with an embodiment of the present disclosure. The memory 112 may include documents input module 202, a multimodal data extraction module 204, a data interlinking module 206, a rules determination module 208, a rules generation module 210, a rules integration module 212, a rules storing module 214 and a rules updating module 216. Further, the rules determination module 208 may include a classification module 218. Further, the rules generation module 210 may include a key components identification module 220.
[0039] The documents input module 202 may fetch a plurality of design documents from the repositories 107. In an embodiment, the one or more repositories may include, but are not limited to, Product Lifecycle Management (PLM) systems, ECO systems, file repositories, databases, or other enterprise document management systems. The documents input module 202 supports multiple document formats, including but not limited to text files, Portable Document Format (PDF) documents, spreadsheets, image files, scanned documents, and HTML files. The documents may include, for example, product design handbooks describing manufacturability guidelines, ECO notices describing changes to allowable tolerances, tabular material property specifications, or diagrams illustrating geometric constraints such as hole-to-bend distances or rib thickness ratios. The documents may include multimodal information including natural language descriptions of design rules, mathematical expressions (e.g., “H = 2t + r”), tabular data defining parameter ranges (e.g., material-specific wall thickness ranges), and graphical representations such as diagrams depicting dimensional relationships (e.g., minimum distance between a hole and a bend). In some cases, documents may include scanned or image-embedded text, which is also retrieved as part of the multimodal content. By supporting ingestion of these heterogeneous document sources and formats, the documents input module 202 enables comprehensive capture of both standardized design rules and organization-specific best practices stored across distributed engineering repositories.
[0040] Further, upon fetching the plurality of design documents from the repositories 107, the multimodal data extraction module 204 may extract multimodal data from each of the plurality of design documents. The multimodal data may include, but is not limited to, text, images, diagrams, tables, charts, equations, captions, labels, and embedded annotations that collectively describe design guidelines or constraints.
[0041] For example, the plurality of design documents may include design guidelines that define geometric relationships, dimensional constraints, manufacturing recommendations, material-dependent limits, or assembly requirements associated with product components. Such design guidelines may specify rules relating to minimum or maximum dimensions of features, spacing between design elements, ratios between geometric parameters, allowable tolerances, or conditions under which certain design features should be used. In some instances, the design guidelines may be expressed as mathematical expressions, conditional statements, tabular values, or descriptive text accompanied by diagrams or illustrative figures identifying the relevant design parameters. For example, a design guideline may specify that a geometric dimension of a component feature should be greater than or equal to a value derived from another parameter such as material thickness, or that a spacing between two features should satisfy a predefined ratio relative to a structural dimension. In other cases, the design guideline may define material-specific parameter ranges, recommended dimensions for manufacturing processes, or constraints intended to prevent structural deformation, manufacturing defects, or assembly interference. Such design guidelines are typically embedded across textual descriptions, diagrams, and tables within the design documents and collectively represent constraints or rules that may be extracted by the multimodal data extraction module 204.
[0042] In an embodiment, the multimodal data extraction module 204 performs page-wise or section-wise analysis of each document to identify and isolate distinct content regions, including text blocks, figure regions, and tabular regions. In an embodiment, the multimodal data extraction module 204 applies Optical Character Recognition (OCR) techniques to extract text from images or scanned documents. The multimodal data extraction module 204 further detects and parses tabular structures to extract structured data and identifies visual elements relevant to design rules, including geometric features, measurement indicators, callouts, and schematic relationships illustrated in engineering diagrams, and associates such visual elements with corresponding textual descriptions or tabular data present within the document.
[0043] For example, the multimodal data extraction module 204 may extract a textual rule describing a minimum distance between a bend and a hole, a diagram illustrating the same dimensional relationship with labeled parameters, and a table specifying permissible ranges for different materials, thereby capturing all modalities of a single design rule context. By extracting and structuring multimodal information in this manner, the multimodal data extraction module 204 enables downstream processing for contextual interlinking, rule identification, and generation of machine-understandable design rules.
[0044] Further, the data interlinking module 206 may interlink data of different data modalities within the extracted multimodal data by associating portions of the multimodal data corresponding to different modalities based on at least one of positional proximity between the portions, reference indicators linking the portions, or semantic correlation between the portions to create a plurality of contextually linked datasets. As used herein, a ‘portion of multimodal data’ refers to a discrete segment of content extracted from a document, including a text segment, an image or diagram region, or a table or table element, that represents part of a design rule or design guideline. The data interlinking module 206 associates related text, images, tables, and diagrams that collectively describe a rule or design guideline. For example, the data interlinking module 206 may link a textual rule description with a corresponding diagram or table that defines parameters, dimensions, or conditions. The data interlinking module 206 may preserve document hierarchy, page references, and proximity information to ensure that related content is processed together.
[0045] In an embodiment, the data interlinking module 206 may perform the association of portions of the multimodal data based on positional proximity between the portions. Positional proximity may include spatial adjacency or closeness of multimodal elements within a same page, within adjacent pages, or within a same section of a document. For example, a paragraph describing a design rule that is located near a figure or table on the same page may be associated based on their relative positions within the document layout. The data interlinking module 206 may utilize document structure information such as page order, section boundaries, and layout regions to determine positional proximity between portions of multimodal data.
[0046] In some embodiments, the data interlinking module 206 performs the association of portions of the multimodal data based on reference indicators linking the portions. Reference indicators may include figure labels, table numbers, captions, callouts, or textual references within a document that link a text segment to a corresponding image, diagram, or table. For example, a text segment referencing “as shown in FIG. 2” or “see Table 1” may be associated with the corresponding figure or table using the reference indicator. The data interlinking module 206 uses such reference indicators to establish explicit relationships between different modalities describing a same design rule.
[0047] In some embodiments, the data interlinking module 206 performs the association of portions of the multimodal data based on semantic correlation between the portions. Semantic correlation may include similarity of terminology, shared design entities, or related contextual meaning between portions of multimodal data. For example, a text segment referring to a “bend radius” may be associated with a diagram showing a bend feature labelled with the same or similar terminology, or with a table specifying bend radius values. The data interlinking module 206 may utilize language model–based analysis or contextual matching techniques to identify semantic relationships between portions of multimodal data.
[0048] In an embodiment, the data interlinking module 206 combines one or more of positional proximity, reference indicators, and semantic correlation to form robust associations between multimodal portions. The resulting contextually linked dataset may include identifiers of the associated portions of multimodal data, document source identifiers, page references, and contextual grouping information, thereby preserving the relationship among the associated portions for subsequent processing.
[0049] Upon interlinking the data of different data modalities, the interlinked data is then passed to the rules determination module 208. The rules determination module 208 may determine one or more rules based on the plurality of contextually linked datasets using a Large Language Model (LLM). In an embodiment, the rules determination module 208 may analyse each contextually linked dataset, which may include associated text segments, images, tables, diagrams, or equations, to identify design constraints, requirements, or conditions expressed across the modalities. The rules determination module 208 evaluates linguistic patterns, semantic meaning, and contextual indicators present in the contextually linked datasets to determine whether the associated content represents a design rule or guideline. For example, the rules determination module 208 may identify language patterns indicating constraints such as minimum or maximum values, ratios, conditional relationships, dimensional requirements, or material-specific limitations. The rules determination module 208 may further utilize the contextual associations preserved during the interlinking process so that textual descriptions, visual representations, and tabular specifications describing the same design rule are interpreted together when determining the rule.
[0050] In an embodiment, the rules determination module 208 may exclude content that does not correspond to enforceable design rules. Such content may include descriptive explanations, background information, illustrative examples, commentary, or consequences of non-compliance, which may be present in the design documents but do not constitute formal design constraints. By filtering out such non-rule content, the rules determination module 208 ensures that only relevant rule statements are forwarded for subsequent rule generation and integration.
[0051] For determining the one or more rules, the classification module 218 may classify the extracted multimodal data into a predefined category. The predefined category may include a rule category and a non-rule category. The classification module 218 may leverage machine learning models, Large Language Models (LLMs), or rule-based logic to perform classification based on features such as linguistic structure, presence of constraint indicators, numerical expressions, or domain-specific terminology associated with design rules. Portions of the multimodal data classified into the rule category are identified as candidate rule segments and provided to the rules determination module 208 for further processing, while portions classified into the non-rule category are excluded from rule generation.
[0052] In an embodiment, the classification performed by the classification module 218 and the rule identification performed by the rules determination module 208 may operate in combination, such that classification results guide the identification of rule segments and the subsequent determination of one or more design rules from the contextually linked datasets. This combined operation enables accurate identification of rule statements from heterogeneous multimodal documents while maintaining contextual relationships among associated text, images, and tables.
[0053] Further, the rules generation module 210 may generate a set of machine-understandable rules based on one or more rules determined by the rule determination module 208. In an embodiment, the rules generation module 210 transforms the determined rule statements from natural language and associated multimodal context into a structured representation suitable for computational processing and automated enforcement. The generated machine-understandable rules are configured to be interpretable by a computing tool for evaluating compliance of a design against the extracted rules. The rules generation module 210 may process each determined rule individually or in groups corresponding to a same contextually linked dataset so that related conditions and constraints are preserved in the generated rule representation.
[0054] For generating the set of machine-understandable rules, the key components identification module 220 may identify a set of key components from the one or more design rules. The set of key components may include, but is not limited to, a function, an object, an attribute, an operator, a value, a unit, and the like. In an embodiment, the key components identification module 220 parses the determined rule statements to identify domain entities representing design features (objects), corresponding characteristics of the design features (attributes), relational or comparative operators indicating a constraint, numerical or symbolic values associated with the constraint, units of measurement, and functions representing relationships between design features. The identified key components are then assembled into a structured rule expression representing the design rule in a machine-understandable format.
[0055] In an embodiment, the machine-understandable rule generated by the rules generation module 210 may represent the design rule as a logical expression, mathematical expression, or parameterized constraint that can be evaluated by a computing tool. Where a design rule includes multiple conditions or constraints, the rules generation module 210 may generate corresponding multiple expressions or structured elements within the machine-understandable rule to represent the complete rule context. The machine-understandable rules thus generated preserve the meaning of the original design rule while enabling automated interpretation and enforcement within the computing tool.
[0056] The rules integration module 212 may integrate each of the set of machine-understandable rules into a computing tool for enabling automated design rules enforcement within the computing tool. In an embodiment, the computing tools may include, but are not limited to, design software tools, Computer-Aided Design / Computer-Aided Manufacturing (CAD/CAM) tools, automated design validation platforms, and the like.
[0057] In an embodiment, the rules integration module 212 provides the machine-understandable rules to the computing tool in a structured format that is compatible with the rule evaluation mechanisms of the computing tool. The computing tool may then apply the machine-understandable rules to one or more design models, design files, or design parameters to determine whether the design satisfies the extracted rules. In this manner, the computing tool may automatically evaluate design compliance based on the rules generated from the plurality of design documents. The computing tool, upon receiving the machine-understandable rules, may execute automated validation operations to identify whether one or more aspects of a design violate the rules. The computing tool may generate validation outputs indicating compliance or non-compliance with the rules, thereby enabling automated design rules enforcement within the computing tool. Such enforcement may assist design engineers in identifying design issues, ensuring adherence to design guidelines, and improving consistency across designs.
[0058] Further, the rules storing module 214 may store the set of machine-understandable rules in the metadata database 104. The rules are stored in a structured and machine-readable format that supports efficient retrieval, execution, and management. The rules storing module 214 may also store associated provenance metadata. The associated provenance metadata may include rule source references, version information, applicability conditions, traceability links to the original design documents and the like.
[0059] Further, the rules updating module 216 may update the set of machine-understandable rules stored in the metadata database 104 based on one of a new design document or an update to an existing design document of the plurality of design documents. For example, updates may be triggered by changes in design documents, new ECOs, or revised standards of PLMs. The rules updating module 216 identifies affected rules, then modifies rule definitions as needed, and ensures that updated rules remain consistent with the latest design guidelines. The rules updating module 216 may be configured to continuously evolve the process for determining the one or more design rules based on feedback from a user.
[0060] It should be noted that all such aforementioned modules 202-220 may be represented as a single module or a combination of different modules. Further, as will be appreciated by those skilled in the art, each of the modules 202-220 may reside, in whole or in parts, on one device or multiple devices in communication with each other. In some embodiments, each of the modules 202-220 may be implemented as dedicated hardware circuit comprising custom application-specific integrated circuit (ASIC) or gate arrays, off-the-shelf semiconductors such as logic chips, transistors, or other discrete components. Each of the modules 202-220 may also be implemented in a programmable hardware device such as a field programmable gate array (FPGA), programmable array logic, programmable logic device, and so forth. Alternatively, each of the modules 202-220 may be implemented in software for execution by various types of processors (e.g., processor 110). An identified module of executable code may, for instance, include one or more physical or logical blocks of computer instructions, which may, for instance, be organized as an object, procedure, function, or other construct. Nevertheless, the executables of an identified module or component need not be physically located together but may include disparate instructions stored in different locations, which, when joined logically together, include the module and achieve the stated purpose of the module. Indeed, a module of executable code could be a single instruction, or many instructions, and may even be distributed over several different code segments, among different applications, and across several memory devices.
[0061] In particular, as will be appreciated by those of ordinary skill in the art, control logic and/or automated routines for performing the techniques and steps described herein may be implemented by the system 100 and the associated computing device 102 either by hardware, software, or combinations of hardware and software. For example, suitable code may be accessed and executed by the one or more processor(s) 110 on the system 100 to perform some or all of the techniques described herein. Similarly, application-specific integrated circuits (ASICs) configured to perform some or all of the processes described herein may be included in the one or more processor(s) 110 on the system 100.
[0062] Referring now to FIG. 3, a functional block diagram of the document input module 202 is illustrated, in accordance with an embodiment of the present disclosure. FIG. 3 is explained in conjunction with FIGs. 1 and 2.
[0063] The documents input module 202 may obtain the plurality of design documents from the one or more repositories 107. The one or more repositories 107 may include repositories associated with enterprise systems such as Product Lifecycle Management (PLM) systems 302 and Engineering Change Order (ECO) systems 304. The documents input module 202 fetches the plurality of design documents from the repositories 107 for subsequent processing by the system.
[0064] In an embodiment, the plurality of design documents obtained from the repositories 107 correspond to documents in predefined formats including text files, Portable Document Format (PDF) files, spreadsheet files, image files, Hyper Text Markup Language (HTML) files, word processing files, presentation files, Computer-Aided Design (CAD) drawing files, and scanned documents. The plurality of design documents include content describing design rules, design constraints, design guidelines, or design requirements.
The documents input module 202, upon fetching the plurality of design documents, provides the plurality of design documents to the metadata database 104 for storage and subsequent processing. The metadata database 104 stores data corresponding to the plurality of design documents to enable extraction of multimodal data from each of the plurality of design documents, interlinking of data of different data modalities to create a plurality of contextually linked datasets, and determination of one or more rules based on the plurality of contextually linked datasets using the Large Language Model (LLM).
[0065] Referring now to FIG. 4, a functional block diagram of the multimodal data extraction module 204 is illustrated, in accordance with an embodiment of the present disclosure. FIG. 4 is explained in conjunction with FIG. 3.
[0066] The multimodal data extraction module 204 may extract the multimodal data from each of the plurality of design documents fetched by the documents input module 202 from the repositories 107. The multimodal data extraction module 204 extracts multimodal data including images, text, tables, charts, and equations from each of the plurality of design documents.
[0067] In an embodiment, the multimodal data extraction module 204 includes a text extractor 402, an image extractor 404, and a table extractor 406. The text extractor 402 extracts text from each of the plurality of design documents, including text present in text files, PDF files, word processing files, presentation files, HTML files, and scanned documents. The image extractor 404 extracts images from each of the plurality of design documents, including images present in image files, PDF files, and scanned documents. The table extractor 406 extracts tables from each of the plurality of design documents, including tables present in spreadsheet files, PDF files, and word processing files.
[0068] In an embodiment, the multimodal data extraction module 204 extracts the multimodal data from each of the plurality of design documents such that the extracted multimodal data includes images, text, tables, charts, and equations associated with design rules, design constraints, design guidelines, or design requirements contained in the plurality of design documents.
[0069] The multimodal data extraction module 204 provides the extracted multimodal data to the data interlinking module 206 for interlinking data of different data modalities within the extracted multimodal data to create the plurality of contextually linked datasets. Accordingly, FIG. 4 illustrates that the multimodal data extraction module 204, including the text extractor 402, the image extractor 404, and the table extractor 406, extracts multimodal data from each of the plurality of design documents and provides the extracted multimodal data for subsequent interlinking and rule determination in the automated design rules extraction and integration using Large Language Models (LLMs).
[0070] Referring now to FIG. 5, a functional block diagram 500 describing generation of set of machine-understandable rules is illustrated, in accordance with an exemplary embodiment of the present disclosure. FIG. 5 is explained in conjunction with FIGs. 3 and 4.
[0071] The plurality of design documents stored in the metadata database 506 may be used for determining one or more rules based on a plurality of contextually linked datasets using a Large Language Model (LLM). The rule generation engine 502 accesses the data stored in the metadata database 506 to generate one or more rules corresponding to design rules, design constraints, design guidelines, or design requirements contained in the plurality of design documents.
[0072] In an embodiment, the rule generation engine 502 generates one or more rules from the plurality of design documents and provides the one or more rules to a rule database 504. The rule database 504 stores the one or more rules determined from the plurality of design documents for subsequent generation of a set of machine-understandable rules.
[0073] The design rules translator 508 accesses the one or more rules stored in the rule database 504 and generates a set of machine-understandable rules based on the one or more rules. The set of machine-understandable rules corresponds to rules that are interpretable by a computing tool for automated rules enforcement within the computing tool.
[0074] The design rules translator 508 provides the set of machine-understandable rules to a design rules database 510. The design rules database 510 stores the set of machine-understandable rules for integration into a computing tool. In an embodiment, a design rules manager 512 accesses the set of machine-understandable rules stored in the design rules database 510 and provides the set of machine-understandable rules for integration into a computing tool for enabling automated rules enforcement within the computing tool.
[0075] Referring now to FIG. 6, a flowchart of an exemplary method 600 of automated rules extraction and integration using LLMs is illustrated, in accordance with an embodiment of the present disclosure. FIG. 6 is explained in conjunction with FIGs. 1-5.
[0076] Each step of the method 600 may be executed by various modules, similar to the modules of the computing device 102, for automated rules extraction and integration using LLMs. The method 600 outlines the general method for implementing automated rules extraction and integration using LLMs in the system 100.
[0077] At step 602, the processor 110 may fetch a plurality of input documents from one or more repositories. Further, at step 604, the processor 110 may extract multimodal data from each of the plurality of input documents fetched at step 602. Upon extracting the multimodal data from each of the plurality of input documents, at step 606, the processor 110 may interlink data of different data modalities within the extracted multimodal data to create a plurality of contextually linked datasets.
[0078] At step 608, the processor 110 may determine one or more rules based on the plurality of contextually linked datasets using an LLM. Further, at step 610, the processor 110 may generate a set of machine-understandable rules based on the one or more rules determined at step 608. Upon generating the set of machine-understandable rules, the processor 110, at step 612, may integrate each of the set of machine-understandable rules into a computing tool for enabling automated rules enforcement within the computing tool.
[0079] Referring now to FIG. 7, a flow diagram of an exemplary process 700 for interlinking data of different data modalities is illustrated, in accordance with an embodiment of the present disclosure. FIG. 7 is explained in conjunction with FIGs. 1-6.
[0080] Each step of the method 700 may be executed by various modules, similar to the modules of the computing device 102, for automated rules extraction and integration using LLMs. The method 700 outlines the general method for implementing a flow diagram of the exemplary process 700 for interlinking data of different data modalities in the system 100.
[0081] At step 702, the processor 110 may fetch a plurality of input documents from one or more repositories. The fetching operation may be performed periodically, on demand, or in response to a detected update or revision in the repositories.
[0082] Further, the processor 110 may process each page of the plurality of input documents fetched at step 704. In an embodiment, processing each page includes identifying page boundaries, recording page sequence information, and extracting multimodal content present on each page.
[0083] Further, at step 706, the processor 110 may extract multimodal data from each of the plurality of input documents 706. The multimodal data may include, but is not limited to, textual data, images, tables, diagrams, and scanned or handwritten information.
[0084] . Additionally, the step 706 may include extracting the text data from each of the plurality of input documents at step 708. The step 706 may further include extracting the image data from each of the plurality of input documents at step 710. Moreover, the step 706 may further include extracting the table data from each of the plurality of input documents at step 712.
[0085] Once the image data is extracted at step 710, the processor 110 may perform a check to determine whether the extracted image data includes the text data, at step 714. If the extracted image data includes the text data, the processor follows step 708 to extract the text data from the extracted image data. In another case, if the extracted image data does not include the text data, the processor follows step 716. At step 716, a check is performed by the processor 110 to determine whether the extracted image data includes the table data. If the extracted image data includes the table data, the processor follows step 712 to extract the table data from the extracted image data. In another case, if the extracted image data does not include the table data, the processor follows step 710.
[0086] Further, upon extracting the multimodal data, at step 718, the processor 110 may process the extracted multimodal data and break it down into contextual chunks. It should be noted that each of the contextual chunks refers to a similar context. Moreover, upon processing the extracted multimodal data, the processor 110 may interlink data of different data modalities within the extracted multimodal data to create a plurality of contextually linked datasets, at step 720. Upon interlinking the data of different data modalities within the extracted multimodal data, the processor 110 may store the plurality of contextually linked datasets in a predefined metadata database (for example, the metadata database 306), at step 722. It should be noted that in some embodiments, the predefined metadata database is analogous to the database 104.
[0087] Referring now to FIG. 8, a first exemplary design document 800 is illustrated, in accordance with an embodiment of the present disclosure. FIG. 8 is explained in conjunction with FIGs. 1-7.
[0088] The first exemplary design document 800 describes a rule defining the distance between a hole and a bend, or a slot and a bend. It explains the rule's importance, specifies the manufacturing features (hole, slot), outlines the check's constraints, and includes a descriptive diagram.
[0089] The first exemplary design document 800 includes multimodal data comprising textual descriptions, mathematical expressions, and diagrams associated with a design rule defining a minimum distance between a bend and a feature selected from a hole or a slot.
[0090] In the illustrated example, the first exemplary design document 800 includes a textual description specifying that when a bend is formed too close to a hole, the hole may deform, and further specifies that a minimum distance D between the bend and the hole or slot is defined as a function of a material thickness T and a bend radius R. The first exemplary design document 800 includes multiple rule expressions including: (i) for a hole having a diameter less than a predefined threshold, the minimum distance D is defined as D = 2T + R, and (ii) for a slot or hole having a diameter greater than the predefined threshold, the minimum distance D is defined as D = 2.5T + R. The first exemplary design document 800 further includes diagrams depicting the geometric relationship between the bend and the hole or slot, and identifies the variables D, T, and R within the diagrams.
[0091] In an embodiment, the multimodal data extraction module 204 extracts the textual description, the mathematical expressions, and the diagrams from the first exemplary design document 800 as multimodal data, and the data interlinking module interlinks portions of the multimodal data corresponding to the textual description, the mathematical expressions, and the diagrams to create a contextually linked dataset representing the design rule. The rules determination module determines one or more rules from the contextually linked dataset using a Large Language Model (LLM), and the rules generation module generates a set of machine-understandable rules based on the one or more rules.
[0092] In an embodiment, the machine-understandable rule generated from the first exemplary design document 800 may include a structured representation including key components comprising an object, an attribute, an operator, a value, a unit, and a function. An example of a machine-understandable rule corresponding to the design document 800 is illustrated below:
Output
{
Rule Name: Minimum Bend Distance for Hole
Function: Distance
Object 1: Bend
Object 2: Hole
Operator: ≥ Expression: Distance(Bend, Hole) ≥ 2 × Material.Thickness + Bend.Radius
}
[0093] The set of machine-understandable rules generated from the first exemplary design document 800 are provided for integration into a computing tool to enable automated rules enforcement within the computing tool, as described with reference to FIG. 5. Accordingly, FIG. 8 illustrates an example of a design document from which multimodal data is extracted, interlinked, and transformed into machine-understandable rules for automated design rules extraction and integration using Large Language Models (LLMs).
[0094] Referring now to FIG. 9, a second exemplary design document 900 is illustrated, in accordance with an embodiment of the present disclosure. FIG. 9 is explained in conjunction with FIGs. 1-7.
[0095] The second exemplary design document 900 corresponds to one of the plurality of design documents fetched from one or more repositories and processed by the documents input module 202, the multimodal data extraction module 204, and the data interlinking module as described with reference to FIGs. 3-5. The second exemplary design document 900 includes multimodal data comprising textual content and an image representing a design guideline related to rib thickness in a component.
[0096] In the illustrated example, the second exemplary design document 900 includes textual content describing various considerations affecting rib thickness, including material properties, processing conditions, surface texture, and aesthetic considerations. Such textual content includes explanatory information and background information that describe reasons for selecting appropriate rib thickness values. The second exemplary design document 900 also includes rule content specifying a relationship between rib geometry parameters, including a draft value, a base radius value, and a base thickness value. The accompanying image represents the geometric features of the rib and visually illustrates dimensional relationships and parameters associated with the rule.
[0097] In an embodiment, the multimodal data extraction module 204 extracts the textual content and the image from the second exemplary design document 900 as multimodal data, and the data interlinking module interlinks portions of the multimodal data corresponding to the textual content and the image to create a contextually linked dataset. The rules determination module, using a Large Language Model (LLM), evaluates the contextually linked dataset to classify portions of the multimodal data into a rule category and a non-rule category. The explanatory information and background information are classified into the non-rule category, while the portion of the multimodal data representing the dimensional constraints of the rib is classified into the rule category.
[0098] Based on the portions classified into the rule category, the rules determination module determines one or more rules corresponding to the rib thickness design guideline. The rules generation module generates a set of machine-understandable rules based on the determined one or more rules by identifying key components comprising an object, an attribute, an operator, a value, a unit, and a function from the rule content.
[0099] In an embodiment, an example machine-understandable rule generated from the second exemplary design document 900 may include a structured representation as follows:
Output {
Rule Name: Rib Thickness Constraint
Object: Rib
Attributes: BaseThickness, BaseRadius, Draft
Operator: ≥ and =
Expression 1: Rib.Draft ≥ 0.5 degrees
Expression 2: Rib.BaseRadius = 0.125 × Rib.BaseThickness
}
[00100] The set of machine-understandable rules generated from the second exemplary design document 900 are provided for integration into a computing tool to enable automated design rules enforcement within the computing tool, as described with reference to FIG. 5. Accordingly, FIG. 9 illustrates an example of a design document containing both rule and non-rule content, wherein the system identifies, extracts, and converts the rule content into machine-understandable rules for automated design rules extraction and integration using Large Language Models (LLMs).
[00101] Referring now to FIG. 10, a third exemplary design document 1000 is illustrated, in accordance with an embodiment of the present disclosure. FIG. 10 is explained in conjunction with FIGs. 1-7.
[00102] The third design document 1000 corresponds to one of the plurality of design documents fetched from one or more repositories and processed by the documents input module 202, the multimodal data extraction module 204, and the data interlinking module 206 as described with reference to FIGs. 3-5. The third exemplary design document 1000 includes multimodal data comprising textual content and a table associated with design rules defining wall thickness constraints for components formed using different materials.
[00103] In the illustrated example, the third exemplary design document 1000 includes textual content describing general design considerations for wall thickness selection, including material selection, manufacturing effects such as sink marks or warpage, and general recommendations for maintaining uniform wall thickness. Such textual content represents explanatory information and background information that is not directly used for automated design rules enforcement. The third exemplary design document 1000 further includes a table specifying values and ranges corresponding to wall thickness constraints for different materials, including materials such as ABS, Acetal, Acrylic, and Liquid Crystal Polymer. The table includes numerical ranges expressed in units corresponding to wall thickness values for each material.
[00104] In an embodiment, the multimodal data extraction module extracts the textual content and the table from the third exemplary design document 1000 as multimodal data, and the data interlinking module interlinks portions of the multimodal data corresponding to the textual content and the table to create a contextually linked dataset. The rules determination module, using a Large Language Model (LLM), evaluates the contextually linked dataset to classify portions of the multimodal data into a rule category and a non-rule category. The explanatory textual content is classified into the non-rule category, while the tabular data representing material-specific wall thickness ranges is classified into the rule category.
[00105] Based on the portions classified into the rule category, the rules determination module determines one or more rules corresponding to wall thickness constraints for different materials. The rules generation module generates a set of machine-understandable rules based on the determined one or more rules by identifying key components including an object, an attribute, an operator, a value, and a unit from the tabular rule content.
[00106] In an embodiment, example machine-understandable rules generated from the third exemplary design document 1000 may include structured representations as follows:
Output
{
Rule Name : Wall Thickness based on Material Function :
Object1 : Wall
Object1_Attributes : [Thickness]

Constraint : Expression : PartBody.Material , Operator : == , Value : ABS
{
Expression : Wall.Thickness >= 1.0 mm
Operator : Range
Value : [0.045, 0.140]
Units : inches
}
Constraint : Expression : PartBody.Material , Operator : == , Value : Acetal
{
Expression : Wall.Thickness >= 1.0 mm
Operator : Range
Value : [0.030, 0.120]
Units : inches
}

Constraint : Expression : PartBody.Material , Operator : == , Value : Acrylic
{
Expression : Wall.Thickness >= 1.0 mm
Operator : Range
Value : [0.025, 0.500]
Units : inches
}
Constraint : Expression : PartBody.Material , Operator : == , Value : Liquid Crystal Polymer
{
Expression : Wall.Thickness >= 1.0 mm
Operator : Range
Value : [0.030, 0.120]
Units : inches
}
}
[00107] The set of machine-understandable rules generated from the third exemplary design document 1000 are provided for integration into a computing tool to enable automated design rules enforcement within the computing tool, as described with reference to FIG. 5. Accordingly, FIG. 10 illustrates an example of a design document containing both rule and non-rule content, wherein tabular constraint data is identified, extracted, and converted into machine-understandable rules for automated design rules extraction and integration using Large Language Models (LLMs).
[00108] As will be appreciated by those skilled in the art, the systems described in various embodiments discussed above are not routine, or conventional, or well understood in the art. The system 100 discussed above may be capable of offering several advantages.
[00109] Thus, the proposed invention addresses the technical problem of inefficient, error-prone, and time-consuming extraction of design rules from various design documents by automatically processing multimodal data using artificial intelligence techniques. The proposed invention efficiently improves design quality and accelerates product development by eliminating manual interpretation of design handbooks, standards, and changing documents, thereby reducing the time of design iteration cycles.
[00110] The proposed invention also improves operational efficiency by continuously analyzing design document repositories, such as from PLM systems, ECO systems, and other repositories, to identify, categorize, and determine design rules from non-rule content. By selectively prioritizing rule-relevant information while preserving contextual metadata, the system 100 ensures optimal utilization of computational resources and maintains high relevance and accuracy of extracted design rules.
[00111] The proposed invention also introduces intelligent translation of natural-language design rules into structured, machine-understandable format representations by identifying domain objects, attributes, operators, units of measurement, and mathematical relationships. This automated validation and formalization ensures consistency, precision, and reliability of extracted design rules. This enables their direct integration within predefined design tools such as CAD/CAM tools and design software tools without compromising data integrity.
[00112] Furthermore, the proposed invention provides flexibility by interlinking diverse document formats, including text files, PDFs, scanned images, tables, and handwritten legacy data using OCR and multimodal AI techniques.
[00113] Thus, without human intervention, the proposed invention can automatically determine, extract, categorize, and store relevant design rules from changing engineering documents, enabling continuous and automated design validation. This enhances system scalability, design accuracy, and manufacturing efficiency in complex engineering environments. This also reduces manual effort and dependency on domain experts despite increased system intelligence and processing complexity.
[00114] In light of the above-mentioned advantages and the technical advancements provided by the disclosed method 600 and system 100, the claimed steps as discussed above are not routine, conventional, or well understood in the art, as the claimed steps enable the following solutions to the existing problems in conventional technologies. Further, the claimed steps bring an improvement in the functioning of the device itself, as the claimed steps provide a technical solution to a technical problem.
[00115] The specification has described the method and the system 100 for automated design rules extraction and integration using Large Language Models. The illustrated steps are set out to explain the exemplary embodiments shown, and it should be anticipated that ongoing technological development will change the manner in which particular functions are performed. These examples are presented herein for purposes of illustration, and not limitation. Further, the boundaries of the functional building blocks have been arbitrarily defined herein for the convenience of the description. Alternative boundaries can be defined so long as the specified functions and relationships thereof are appropriately performed. Alternatives (including equivalents, extensions, variations, deviations, etc., of those described herein) will be apparent to persons skilled in the relevant art(s) based on the teachings contained herein. Such alternatives fall within the scope and spirit of the disclosed embodiments.
[00116] Furthermore, one or more computer-readable storage media may be utilized in implementing embodiments consistent with the present disclosure. A computer-readable storage medium refers to any type of physical memory on which information or data readable by a processor 110 may be stored. Thus, a computer-readable storage medium may store instructions for execution by one or more processors 110, including instructions for causing the processor(s) 110 to perform steps or stages consistent with the embodiments described herein. The term “computer-readable medium” should be understood to include tangible items and exclude carrier waves and transient signals, i.e., be non-transitory. Examples include random access memory (RAM), read-only memory (ROM), volatile memory, nonvolatile memory, hard drives, CD-ROMS, DVDs, flash drives, disks, and any other known physical storage media.
[00117] It is intended that the disclosure and examples be considered as exemplary only, with a true scope of disclosed embodiments being indicated by the following claims. , Claims:CLAIMS
What is claimed is:
1. A method for automated design rules extraction and integration using Large Language Models (LLMs) comprising:
fetching, by a processor, a plurality of design documents from one or more repositories;
extracting, by the processor, multimodal data from each of the plurality of design documents;
interlinking, by the processor, data of different data modalities within the extracted multimodal data by associating portions of the multimodal data corresponding to different modalities based on at least one of positional proximity between the portions, reference indicators linking the portions, or semantic correlation between the portions to create a plurality of contextually linked datasets;
determining, by the processor, one or more rules based on the plurality of contextually linked datasets using an LLM;
generating, by the processor, a set of machine-understandable rules based on the one or more rules; and
integrating, by the processor, each of the set of machine-understandable rules into a computing tool for enabling automated rules enforcement within the computing tool.

2. The method of claim 1, wherein each of the plurality of design documents is in a predefined format, and wherein the predefined format comprises a text file, a Portable Document Format (PDF) file, a spreadsheet file, an image file, a Hyper Text Markup Language (HTML) file, a word processing file, a presentation file, a Computer-Aided Design (CAD) drawing file, and a scanned document.

3. The method of claim 1, wherein the multimodal data comprises images, text, tables, charts, and equations.

4. The method of claim 1, wherein determining the one or more rules comprises:
classifying, by the processor, the extracted multimodal data into a predefined category, wherein the predefined category comprises a rule category and a non-rule category.

5. The method of claim 1, wherein generating the set of machine-understandable rules comprising:
identifying, by the processor, a set of key components from the one or more rules, wherein the set of key components comprises a function, an object, an attribute, an operator, a value, and a unit.

6. The method of claim 1, further comprising:
continuously evolving, by the processor, a process for determining the one or more rules based on feedback from a user.

7. The method of claim 1, further comprising:
storing, by the processor, the set of machine-understandable rules in a database with associated provenance metadata; and
updating, by the processor, the set of machine-understandable rules stored in the database based on one of a new input document or an update to an existing input document of the plurality of design documents.

8. A system for automated rules extraction and integration using Large Language Models (LLMs) comprising:
a processor; and
a memory communicably coupled to the processor, wherein the memory stores processor-executable instructions, which, when executed by the processor, cause the processor to:
fetch a plurality of design documents from one or more repositories;
extract multimodal data from each of the plurality of design documents;
interlink data of different data modalities within the extracted multimodal data by associating portions of the multimodal data corresponding to different modalities based on at least one of positional proximity between the portions, reference indicators linking the portions, or semantic correlation between the portions to create a plurality of contextually linked datasets;
determine one or more rules based on the plurality of contextually linked datasets using an LLM;
generate a set of machine-understandable rules based on the one or more rules; and
integrate each of the set of machine-understandable rules into a computing tool for enabling automated rules enforcement within the computing tool.

9. The system of claim 8, wherein each of the plurality of design documents is in a predefined format, and wherein the predefined format comprises a text file, a Portable Document Format (PDF) file, a spreadsheet file, an image file, a Hyper Text Markup Language (HTML) file, a word processing file, a presentation file, a Computer-Aided Design (CAD) drawing file, and a scanned document.

10. The system of claim 8, wherein the multimodal data comprises images, text, tables, charts, and equations.

11. The system of claim 8, wherein to determine the one or more rules, the processor-executable instructions, which, when executed by the processor, cause the processor to:
classify the extracted multimodal data into a predefined category, wherein the predefined category comprises a rule category and a non-rule category.

12. The system of claim 8, wherein generating the set of machine-understandable rules causes the processor to:
identify a set of key components from the one or more rules, wherein the set of key components comprises a function, an object, an attribute, an operator, a value, and a unit.

13. The system of claim 8, wherein the processor-executable instructions, which, when executed by the processor, cause the processor to:
continuously evolve a process for determining the one or more rules based on feedback from a user.

14. The system of claim 8, wherein the processor-executable instructions, which, when executed by the processor, cause the processor to:
store the set of machine-understandable rules in a database with associated provenance metadata; and
update the set of machine-understandable rules stored in the database based on one of a new input document or an update to an existing input document of the plurality of design documents.

15. A non-transitory computer-readable medium storing computer-executable instructions for automated rules extraction and integration using Large Language Models (LLMs), the computer-executable instructions configured for:
fetching a plurality of design documents from one or more repositories;
extracting multimodal data from each of the plurality of design documents;
interlinking data of different data modalities within the extracted multimodal data by associating portions of the multimodal data corresponding to different modalities based on at least one of positional proximity between the portions, reference indicators linking the portions, or semantic correlation between the portions to create a plurality of contextually linked datasets;
determining one or more rules based on the plurality of contextually linked datasets using an LLM;
generating a set of machine-understandable rules based on the one or more rules; and
integrating each of the set of machine-understandable rules into a computing tool for enabling automated rules enforcement within the computing tool.

16. The non-transitory computer-readable medium of claim 15, wherein each of the plurality of design documents is in a predefined format, and wherein the predefined format comprises a text file, a Portable Document Format (PDF) file, a spreadsheet file, an image file, a Hyper Text Markup Language (HTML) file, a word processing file, a presentation file, a Computer-Aided Design (CAD) drawing file, and a scanned document.

17. The non-transitory computer-readable medium of claim 15, wherein for determining the one or more rules, the computer-executable instructions are further configured for:
classifying the extracted multimodal data into a predefined category, wherein the predefined category comprises a rule category and a non-rule category.

18. The non-transitory computer-readable medium of claim 15, wherein generating the set of machine-understandable rules comprising:
identifying a set of key components from the one or more rules, wherein the set of key components comprises a function, an object, an attribute, an operator, a value, and a unit.

19. The non-transitory computer-readable medium of claim 15, wherein the computer-executable instructions are further configured for:
continuously evolving a process for determining the one or more rules based on feedback from a user.

20. The non-transitory computer-readable medium of claim 15, wherein the computer-executable instructions are further configured for:
storing the set of machine-understandable rules in a database with associated provenance metadata; and
updating the set of machine-understandable rules stored in the database based on one of a new input document or an update to an existing input document of the plurality of design documents.

Documents

Application Documents

# Name Date
1 202611027482-STATEMENT OF UNDERTAKING (FORM 3) [09-03-2026(online)].pdf 2026-03-09
2 202611027482-PROOF OF RIGHT [09-03-2026(online)].pdf 2026-03-09
3 202611027482-POWER OF AUTHORITY [09-03-2026(online)].pdf 2026-03-09
4 202611027482-FORM-9 [09-03-2026(online)].pdf 2026-03-09
5 202611027482-FORM 18 [09-03-2026(online)].pdf 2026-03-09
6 202611027482-FORM 1 [09-03-2026(online)].pdf 2026-03-09
7 202611027482-FIGURE OF ABSTRACT [09-03-2026(online)].pdf 2026-03-09
8 202611027482-DRAWINGS [09-03-2026(online)].pdf 2026-03-09
9 202611027482-DECLARATION OF INVENTORSHIP (FORM 5) [09-03-2026(online)].pdf 2026-03-09
10 202611027482-COMPLETE SPECIFICATION [09-03-2026(online)].pdf 2026-03-09
11 202611027482-PATENT_APPLICATION_PUBLICATION.pdf 2026-05-04