Abstract: The present disclosure relates to a method and a system for recommending data analytics processes to a user. The method receives input data related to the data analytics processes from the user and identifies the one or more data analytics processes comprising at least one of one or more partial data analytics process and one or more complete data analytics processes from a knowledge repository based on the received input data. Furthermore, the method recommends at least one of the one or more partial data analytics process and the one or more complete data analytics processes to the user based on relevance of each of the one or more identified data analytics processes to the input data. Fig. 2
1. A data analytics process recommendation system for recommending data analytics processes to a user, the data analytics process recommendation system comprising: a processor; and a memory communicatively coupled to the processor, wherein the memory stores processor-executable instructions, which, on execution, causes the processor to: receive input data related to the data analytics processes from the user; identify at least one of one or more partial data analytics process and one or more complete data analytics processes from a knowledge repository associated with the data analytics process recommendation system based on the received input data related to the data analytics processes; and recommend at least one of the one or more partial data analytics process and the one or more complete data analytics processes to the user based on relevance of each of the one or more identified data analytics processes to the input data thus received, wherein each of the one or more partial data analytics process comprises a data analytics operator and each of the one or more complete data analytics processes comprises plurality of data analytics operators.
2. The process recommendation system as claimed in claim 1 further comprises recommending one or more subsequent data analytics operators upon receiving user selection on the data analytics operator of one of the one or more partial data analytics processes thus recommended.
3. The process recommendation system as claimed in claim 2, wherein a machine learning technique comprised in the knowledge repository learns the user selection on the data analytics operator of one of the one or more partial data analytics processes recommended based on the input data for improving the relevance of recommendation of the data analytics processes.
4. The process recommendation system as claimed in claim 1 and 2, wherein a display interface in the data analytics process recommendation system displays at least one of the one or more partial data analytics process and the one or more complete data analytics processes, and the one or more subsequent data analytics operators being recommended to the user.
5. The process recommendation system as claimed in claim 4 further comprises providing a notification to the user, wherein the notification displays one or more information related to the recommended at least one of the one or more partial data analytics process and the one or more complete data analytics processes, and the subsequent data analytics operators.
6. The process recommendation system as claimed in claim 1 further comprises receiving at least one of industry specific information and goal information related to the data analytics processes for recommending at least one of the one or more partial data analytics process and the one or more complete data analytics processes to the user.
7. The process recommendation system as claimed in claim 1, wherein a similarity matrix configured in the knowledge repository stores one or more data analytics processes which are functionally related to data analytics operators.
8. The process recommendation system as claimed in claim 1 further comprises optimizing the recommended at least one of the one or more partial data analytics process and the one or more complete data analytics processes using predefined configuration details.
9. A method of recommending data analytics processes to a user, the method comprising: receiving, by a processor of a data analytics process recommendation system, input data related to the data analytics processes from the user; identifying, by the processor, at least one of one or more partial data analytics process and one or more complete data analytics processes from a knowledge repository associated with the data analytics process recommendation system based on the received input data related to the data analytics processes; and recommending, by the processor, at least one of the one or more partial data analytics process and the one or more complete data analytics processes to the user based on relevance of each of the one or more identified data analytics processes to the input data thus received, wherein each of the one or more partial data analytics process comprises a data analytics operator and each of the one or more complete data analytics processes comprises plurality of data analytics operators.
10. A non-transitory computer readable medium including instructions stored thereon that when processed by at least one processor cause a data analytics process recommendation system to perform operations comprising: receiving, by a processor of a data analytics process recommendation system, input data related to the data analytics processes from the user; identifying, by the processor, at least one of one or more partial data analytics process and one or more complete data analytics processes from a knowledge repository associated with the data analytics process recommendation system based on the received input data related to the data analytics processes; and recommending, by the processor, at least one of the one or more partial data analytics process and the one or more complete data analytics processes to the user based on relevance of each of the one or more identified data analytics processes to the input data thus received, wherein each of the one or more partial data analytics process comprises a data analytics operator and each of the one or more complete data analytics processes comprises plurality of data analytics operators. , Description:FIELD OF THE DISCLOSURE: The present subject matter is related, in general to data analytics, and more particularly, but not exclusively to a method and system for recommending data analytics processes to a user. BACKGROUND: Data analytics is the process of examining large data sets containing a variety of data types to uncover hidden patterns, unknown correlations, market trends, customer preferences and other useful business information. The primary goal of data analytics is to help organizations make more informed business decisions by enabling data scientists, predictive modelers and other analytics professionals to analyze large volumes of transaction data, as well as other forms of data that may be untapped by conventional business intelligence methods. Various forms of unstructured data may include, but are not limited, Web server logs and Internet clickstream data, social media content and social network activity reports, mobile-phone call detail records and machine data captured by sensors connected to the Internet of Things (IoT). The unstructured data may not fit well in traditional data warehouses that are based on relational databases. Furthermore, data warehouses may not be able to handle the processing demands posed by sets of big data that need to be updated frequently or even continually. As a result, many organizations looking to collect, process and analyze big data have turned to a newer class of applications that support the processing of large and diverse data sets across clustered systems. However, these data analytics applications can be viewed as a pipeline or analytics process of operators i.e. algorithms, processes, libraries and hardware exposed via Application Program Interface (API). The analytic data processes includes various operations related to the process of Extract-Transform-Load (ETL) processes, combined with analytic, querying and reporting that consumes a variety of data types. These processes are complex to design, difficult to implement and maintain since they span a variety of systems. Additionally, various quality objectives such as reliability, maintainability, performance, pipeline portioning and parallelization have to be met. Therefore, specific decision support systems are needed to reduce development time and maintenance costs associated with the analytics process. Conventional data analytic processes are built using inputs provided by Domain/ IT expert as per desired functionalities. Even if the data analytic processes are functionally correct, improving the data analytics process is still difficult without the domain specific inputs. Also, conventional techniques doesn’t utilize the already available versions of the process stored in a knowledge repository, rather handles the data analytic processes based on inputs from Domain/ IT expert. Therefore, there is a need for a method and a system that automatically builds data analytic processes and utilizes the already available version of the processes without the involvement of Domain/IT expert, thereby reducing the design cycle time and cost involved and overcoming the limitations of the existing art. SUMMARY: One or more shortcomings of the prior art are overcome and additional advantages are provided through the present disclosure. Additional features and advantages are realized through the techniques of the present disclosure. Other embodiments and aspects of the disclosure are described in detail herein and are considered a part of the claimed disclosure. Accordingly, the present disclosure relates to a method of recommending data analytics processes to a user. The method comprises the steps of receiving input data related to the data analytics processes from the user. The method further comprises the steps of identifying at least one of one or more partial data analytics process and one or more complete data analytics processes from a knowledge repository associated with the data analytics process recommendation system based on the received input data related to the data analytics processes. Upon identifying, the method recommends at least one of the one or more partial data analytics process and the one or more complete data analytics processes to the user based on relevance of each of the one or more identified data analytics processes to the input data thus received. Each of the one or more partial data analytics process comprises a data analytics operator and each of the one or more complete data analytics processes comprises plurality of data analytics operators. Further, the present disclosure relates to a data analytics process recommendation system for recommending data analytics processes to a user. The system comprises a processor and a memory communicatively coupled to the processor, wherein the memory stores processor-executable instructions, which, on execution, cause the processor to receive input data related to the data analytics operator from the user. The processor is further configured to identify at least one of one or more partial data analytics process and one or more complete data analytics processes from a knowledge repository associated with the data analytics process recommendation system based on the received input data related to the data analytics processes. The processor is furthermore configured to recommend at least one of the one or more partial data analytics process and the one or more complete data analytics processes to the user based on relevance of each of the one or more identified data analytics processes to the input data thus received, wherein each of the one or more partial data analytics process comprises a data analytics operator and each of the one or more complete data analytics processes comprises plurality of data analytics operators. Furthermore, the present disclosure relates to a non-transitory computer readable medium including instructions stored thereon that when processed by at least one processor causes a data analytics process recommendation system to perform operations comprising receiving input data related to the data analytics operator from the user. Further, the instructions cause the processor to identify at least one of one or more partial data analytics process and one or more complete data analytics processes from a knowledge repository associated with the data analytics process recommendation system based on the received input data related to the data analytics processes. Furthermore, at least one of the one or more partial data analytics process and the one or more complete data analytics processes to the user based on relevance of each of the one or more identified data analytics processes to the input data thus received, wherein each of the one or more partial data analytics process comprises a data analytics operator and each of the one or more complete data analytics processes comprises plurality of data analytics operators. The foregoing summary is illustrative only and is not intended to be in any way limiting. In addition to the illustrative aspects, embodiments, and features described above, further aspects, embodiments, and features will become apparent by reference to the drawings and the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS: The accompanying drawings, which are incorporated in and constitute a part of this disclosure, illustrate exemplary embodiments and, together with the description, serve to explain the disclosed principles. In the figures, the left-most digit(s) of a reference number identifies the figure in which the reference number first appears. The same numbers are used throughout the figures to reference like features and components. Some embodiments of system and/or methods in accordance with embodiments of the present subject matter are now described, by way of example only, and with reference to the accompanying figures, in which: Fig. 1 illustrates an architecture diagram of an exemplary system for recommending data analytics processes to a user in accordance with some embodiments of the present disclosure; Fig. 2 illustrates an exemplary block diagram of a data analytics process recommendation system of Fig. 1 in accordance with some embodiments of the present disclosure; Figs. 3a and 3b illustrate an exemplary method of representing and recommending data analytics processes on a display interface of the data analytics process recommendation system; Fig. 4a illustrates a detailed process chart of an exemplary method of recommending data analytics processes to a user in accordance with some embodiments of the present disclosure; Fig. 4b illustrates the method of recommending data analytics processes to a user in accordance with an exemplary embodiment of the present disclosure; and Fig. 5 is a block diagram of an exemplary computer system for implementing embodiments consistent with the present disclosure. It should be appreciated by those skilled in the art that any block diagrams herein represent conceptual views of illustrative systems embodying the principles of the present subject matter. Similarly, it will be appreciated that any process charts, process diagrams, state transition diagrams, pseudo code, and the like represent various processes which may be substantially represented in computer readable medium and executed by a computer or processor, whether or not such computer or processor is explicitly shown. DETAILED DESCRIPTION: In the present document, the word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any embodiment or implementation of the present subject matter described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments. While the disclosure is susceptible to various modifications and alternative forms, specific embodiment thereof has been shown by way of example in the drawings and will be described in detail below. It should be understood, however that it is not intended to limit the disclosure to the particular forms disclosed, but on the contrary, the disclosure is to cover all modifications, equivalents, and alternative falling within the scope of the disclosure. The terms “comprises”, “comprising”, or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a setup, device or method that comprises a list of components or steps does not include only those components or steps but may include other components or steps not expressly listed or inherent to such setup or device or method. In other words, one or more elements in a system or apparatus proceeded by “comprises… a” does not, without more constraints, preclude the existence of other elements or additional elements in the system or apparatus. The present disclosure relates to a method and a system for recommending data analytics processes to a user. In one embodiment, the method receives input data related to the data analytics operator from the user. The method further identifies the one or more data analytics processes comprising at least one of one or more partial data analytics process and one or more complete data analytics processes from a knowledge repository associated with the data analytics process recommendation system based on the received input data related to the data analytics operators. Each of the one or more partial data analytics process comprises a data analytics operator. Further, each of the one or more complete data analytics processes comprises plurality of data analytics operators. Furthermore, the method recommends the identified at least one of the one or more partial data analytics process and the one or more complete data analytics processes to the user based on relevance of each of the one or more identified data analytics processes to the input data thus received. In an embodiment of the present disclosure, the method also recommends one or more subsequent data analytics operators upon receiving user selection on the data analytics operator of one of the one or more partial data analytics processes thus recommended. Therefore, the proposed data analytics process recommendation system makes use of the related to the data analytics operators and previously stored data analytics processes to recommend one or more subsequent data analytics operators and data analytics processes with minimum optimization time and cost. In the following detailed description of the embodiments of the disclosure, reference is made to the accompanying drawings that form a part hereof, and in which are shown by way of illustration specific embodiments in which the disclosure may be practiced. These embodiments are described in sufficient detail to enable those skilled in the art to practice the disclosure, and it is to be understood that other embodiments may be utilized and that changes may be made without departing from the scope of the present disclosure. The following description is, therefore, not to be taken in a limiting sense. Fig. 1 illustrates an architecture diagram of an exemplary system for recommending data analytics processes to a user in accordance with some embodiments of the present disclosure. As shown in Fig. 1, the exemplary system 100 comprises one or more components configured to recommend one or more data analytics processes 213 to a user. In one embodiment, the exemplary system 100 comprises a data source 101, and a data analytics process recommendation system 109. In one embodiment, the data source 101 comprises input data 103 related to one or more data analytics processes 213. In an embodiment, the data analytics processes 213 may be one or more partial data analytics processes. Each of the one or more partial data analytics process comprises a single data analytics operator. In another embodiment, the data analytics process 213 may be a complete data analytics process. The complete data analytics process comprises plurality of data analytics operators 211. As an example the data analytics processes 213 received from the user may include a reading operation, aggregation of rows of data in a database, parsing of data in a database and running a simple machine learning algorithm. The data analytics process recommendation system 109 recommends at least one of the one or more partial data analytics process and the one or more complete data analytics processes to the users. In an embodiment, the input data 103 may include, but not limited to, a data file of a specific format which comprises one or more information related to the data analytics processes. As an example, the input file may be in one of the data file formats, such as txt, doc, pdf and xml. In another embodiment, the input data 103 may also contain at least one of one or more data analytics operators 211, data analytics processes 213, goal information 110, industry specific information 111 and data to be analyzed 104. The input data 103 received from the data source 101 may be used by the data analytics process recommendation system 109 for recommending one or more data analytics processes 213 to the user. Fig. 2 illustrates an exemplary block diagram of a data analytics process recommendation system of Fig. 1 in accordance with some embodiments of the present disclosure. In one embodiment, the data analytics process recommendation system 109 comprises a central processing unit (“CPU” or “processor”) 113, a memory 115 and an I/O interface 201. The I/O interface 201 is coupled with the processor 113 and one or more I/O devices. An I/O device, such as a keyboard, is configured to receive inputs via the I/O interface 201. Similarly, the processed outputs for displaying in an I/O device, such as a monitor, are transmitted via the I/O interface 201. In an embodiment, the memory 115 of the data analytics process recommendation system 109 further comprises a knowledge repository 219 and modules 205 comprising, a recommendation module 117 and a knowledge repository interface module 221. In one example, the knowledge repository 219 may include, but not limited to, a similarity matrix 209 and one or more configuration information 105. The configuration information 105 may include, but not limited to, goal information 110 and industry specific information 111 associated with the one or more data analytics processes 213. The goal information 110 related to a data analytics process 213 defines the goal of the data analytics process recommendation. In an embodiment, the goal information 110 may include, but not limited to, text processing, predictive modeling, descriptive modeling, pattern discovery, forecasting, and text analytics. The industry specific information 111 related to a data analytics process 213 comprises at least one of name of the industry, subject matter of the analytics industry and one or more specific rules associated with the analytics industry. As an example, the analytics industry may be Healthcare, Oil and Gas, Mining, Social Media and any other suitable analytics industry. In an embodiment, the goal information 110 and the industry specific information 111 may be stored in the form of one or more metadata files. In an embodiment, the user uses the one or more goal information 110 and the industry specific information 111 stored in the similarity matrix 209 to prune the recommendation of one or more data analytics processes 213. In an embodiment, one or more data analytics operators 211 and one or more data analytics processes 213 may be stored in the similarity matrix 209. As an example, the similarity matrix 209 can be a structured database storing one or more data analytics processes 213 and one or more data analytics operators 211 related to a specific combination of the goal information 110 and the industry specific information 111. The similarity matrix 209 may also store one or more data analytics operators 211 and one or more data analytics processes that are previously recommended to a user. In an embodiment, the one or more data analytics operators 211 and the one or more data analytics processes 213 stored in the similarity matrix 209 may be alternatively generated by industry expert personnel based on the one or more goal information 110 and the industry specific information 111. In yet another embodiment, the data analytics process recommendation system 109 may optimize the one or more recommended data analytics processes 213 based on the one or more optimization parameters. The optimization parameters determine the type of optimization trade-offs needed by the user. The one or more optimized data analytics processes stored in the knowledge repository 219 can be selected based on the optimization parameters. The optimization parameters may include, but not limited to, measures for reliability, maintainability, performance, pipelining and parallelization. In an example, the reliability measures are determined by the total number of recovery points present in a data analytics process 213. The measures for maintainability may include length of the data analytics process 213, complexity in execution of the data analytics process 213, degree of modularity associated with the data analytics process 213 and the extent of coupling between the one or more modules 205 in the data analytics process 213. The measures for performance may include total memory 115 space consumed while performing the data analytics process 213, total processor 113 runtime consumed and throughput associated with the execution of the data analytics process 213. The measures for pipelining and parallelization may include various factors associated with pipelining, such as, partition type, number of data volumes in each pipeline partition, length of data volumes in each pipeline partition and number of processing units under execution. In an exemplary embodiment, the user may want to optimize a process to include maximum fault tolerance. Hence, the recommendation module 117 has to implement the measures for reliability in order to generate a fault tolerant process. Then, the recommendation module 117 has to consider including effective number of recovery points in the optimized data analytics process 213. As an example, for a simple data analytics process with ‘n’ number of operators, there must be ‘n-1’ recovery points to ensure maximum fault tolerance. Further, upon optimizing the one or more data analytics processes 213 to the user, the optimized data analytics processes are stored back into the knowledge repository 219. The stored optimized data analytics processes may be used by the data analytics process recommendation system 109 to quickly recommend one or more data analytics processes 213 to the user upon receiving a relevant data analytics operator 211. In an embodiment, the data 203 may be stored in the memory 115 in the form of various data structures. Additionally, the aforementioned data can be organized using data models, such as relational or hierarchical data models. The other data 219 may be also referred to as reference repository for storing recommended implementation approaches as reference data. In an embodiment, the other data 219 may store on or more configuration data related to the one or more partial data analytics processes and one or more complete data analytics processes. The other data 219 also stores data, including temporary data and temporary files, generated by the modules 205 for performing the various functions of the data analytics process recommendation system 109. The modules 205 may include, for example, the recommendation module 117, a knowledge repository interface module 221 and other modules 223. The other modules 223 perform various miscellaneous functionalities of the data analytics process recommendation system 109. It will be appreciated that such aforementioned modules may be represented as a single module or a combination of different modules. The modules 205 may be implemented in the form of software, hardware and/or firmware. The data analytics process recommendation system 109 receives the input data 103 associated with the data analytics processes 213 from a user. The data analytics processes 213 may include at least one of one or more partial data analytics process and one or more complete data analytics processes stored in a knowledge repository 219 associated with the data analytics process recommendation system 109. In an embodiment, the recommendation module 117 recommends one or more data analytics processes 213 to the users based on the relevance of each of the data analytics processes 213 stored in the knowledge repository 219 to the data to be analyzed 104. In one embodiment, the knowledge repository 219 may be integrated within the data analytics process recommendation system 109. In another embodiment, the knowledge repository 219 may be implemented independent of the data analytics process recommendation system 109. The knowledge repository 219 may be for example, a data warehouse configured to store metadata related to one or more optimized data analytics processes that are previously executed and saved by the user. The recommendation module 117 selects one of the one or more data analytics processes 213, having maximum relevance to the data analytics processes 213 stored in the knowledge repository 219 and recommends it to the user. The recommended processes are then saved to the knowledge repository 219 for their future use. The knowledge repository interface module 221 is used to access the knowledge repository 219 that may be integrated within the data analytics process recommendation system 109. In an embodiment, the knowledge repository 219 can be integrated independent of the data analytics process recommendation system 109. Figs. 3a and 3b illustrate an exemplary method of representing and recommending data analytics processes on a display interface of the data analytics process recommendation system. Fig. 3a illustrates the method of recommending one or more data analytics processes 213 to the user. In an embodiment, the display interface 300 in the data analytics process recommendation system 109 may be further divided into a selection panel 301 and a process recommendation area 303. The data analytics process recommendation system 109 displays one or more information, such as, one or more operators, one or more domains, input data 103 and one or more other configuration data on the selection panel 301 of the display interface 300. The user may select one of the one or more operators, the one or more domains, the input data 103 and the one or more other configuration data, displayed in the selection panel 301, to obtain a most relevant data analytics process 213 as the recommendation from the data analytics process recommendation system 109. In an exemplary embodiment, the one or more information displayed on the selection panel 301 may be enabled with a drag and drop technique, using which, the user may drag one or more required information from the selection pane 301 and drop them on the process recommendation area 303 of the display interface 300. In an exemplary embodiment, the user may input a data to the data analytics process recommendation system 109. In an embodiment, the input data 103 may include one of one or more data analytics operators 211, data analytics processes 213, goal information 110, industry specific information 111 and one or more data 104 to which the user needs to receive one or more recommendations. Accordingly, the data analytics process recommendation system 109 identifies one or more data analytics processes 213, from the knowledge repository 219, that are functionally equivalent to the received input data 103. Further, the data analytics process recommendation system 109 recommends one or more identified data analytics processes X, Y and Z, and one or more partial data analytics processes, namely, Opr 1 and Opr 2 on the process recommendation area 303 based on the relevance of each of the identified data analytics processes to the received input data 103. In an embodiment, the user may select one of the one or more recommended data analytics processes X, Y and Z, and Opr 1 and Opr 2 if the user finds the recommended processes to be in match with the pre-selected goal information 110 and industry specific information 111. In yet another embodiment, as illustrated in Fig. 3b, if none of the recommendations are in compliance with the pre-selected domains, the user may wish to change one or more individual data analytics operators 211 in one of the recommended data analytics processes. As an example, the user may prefer to alter one of the data analytics operator Y1 of the recommended data analytics process Y, if the complete data analytics process Y is not acceptable by the user. Alternatively, the data analytics process recommendation system 109 recommends a plurality of data analytics operators 211 to the selected data analytic operator Y1, such that, the recommended data analytics operators 211 are functionally equivalent to Y1, thus recommending a complete data analytics process 213. Further, the one or more data analytics operators 211 and the data analytics process 213 selected by the user are saved back in the knowledge repository 219 for aiding in the future recommendations. Figure 4a illustrates a detailed process chart of an exemplary method of recommending data analytics processes to a user in accordance with some embodiments of the present disclosure; As illustrated in Figure 4, the method 400 comprises one or more blocks implemented by the processor 113 for recommending data analytics processes 213 to a user. The method 400 may be described in the general context of computer executable instructions. Generally, computer executable instructions can include routines, programs, objects, components, data structures, procedures, modules, and functions, which perform particular functions or implement particular abstract data types. The order in which the method 400 is described is not intended to be construed as a limitation, and any number of the described method blocks can be combined in any order to implement the method 400. Additionally, individual blocks may be deleted from the method 400 without departing from the spirit and scope of the subject matter described herein. Furthermore, the method 400 can be implemented in any suitable hardware, software, firmware, or combination thereof. At block 401, the data analytics process recommendation system 109 receives input data 103 related to the data analytics operator from the user. The input data 103 may contain one or more data analytics operators 211, data analytics processes 213, goal information 110, industry specific information 111 and data to be analyzed 104. The input data 103 received from the data source 101 may be used by the data analytics process recommendation system 109 for recommending one or more data analytics processes 213 to the data to be analyzed 104. At block 403, the data analytics process recommendation system 109 identifies at least one of one or more partial data analytics process and one or more complete data analytics processes from a knowledge repository 219 associated with the data analytics process recommendation system 109 based on the received data related to the data analytics processes 213. At block 405, the process recommendation module 117 recommends at least one of the one or more partial data analytics process and the one or more complete data analytics processes to the user based on relevance of each of the one or more identified data analytics processes 213 to the input data 103, wherein each of the one or more partial data analytics process comprises a data analytics operator and each of the one or more complete data analytics processes comprises plurality of data analytics operators 211. Fig. 4b illustrates the method of recommending data analytics processes to a user in accordance with an exemplary embodiment of the present disclosure. In an embodiment, the data analytics process recommendation system 109 may receive one or more input data 103 from the user. The input data 103 may include one or more information related to the data analytics processes 213 in one of the forms including, but not limited to, a CSV, a file, a tabular representation of the information and data to be 104, having a specific schema, on which the user wants to perform the data analytics operations. Additionally the user may also input the goal information 110 and the industry specific information 111, related to the data analytics process 213, to the data analytics process recommendation system 109. The data analytics process recommendation system 109 may use the received goal information 110 and the industry specific information 111 to improve the recommendations on one of the partial data analytics process and the one or more complete data analytics flows. In an embodiment, the data analytics process recommendation system 109 as shown in Fig. 4b may further use the received goal information 110 and the industry specific information 111 to derive a most relevant dependency tree corresponding to the data to be analyzed 104 and other information. The dependency trees are the data structures that are stored in the knowledge repository 219 of the recommendation system 109 along with the data sets previously used for the data analytics processes 213 and the other info like industry specific information 111. Furthermore, the recommendation module 117 makes use of usage history and one or more configuration details associated with the data analytics process recommendation module 117 to identify one or more additional data analytics processes 213 from the obtained dependency tree. In an embodiment, the data analytics process recommendation system 109 may save one or more samples of each of the data including the goal information 110 and the industry specific information 111 in the knowledge repository 219 for its future use. The data analytics process recommendation system 109 may save the sample data each time the user requests a recommendation. In an embodiment, the recommendation module 117 may also save one or more information to find out a match between the data analytics process 213 and the used data. Similarly, the recommendation module 117 may derive a match between the goal information 110 and the corresponding data analytic operators which can be recommended to find out the recommended combination of goal information 110 and operators. The data stored in the knowledge repository 219 may enable the user of the data analytics process recommendation system 109 to select a most appropriate dependency tree and also to rank the suitable branches among the various branches going from one node in dependency tree. Additionally, the recommendation module 117 may also save one or more feedback information received from the user for improving the future recommendations. Table A depicted below illustrates an exemplary similarity matrix 209 in which the recommendation module 117 receives the one or more data sets and the corresponding goal information 110 to find a most relevant dependency tree comprising one or more data analytics operators 211. Data Set Goal Data type Operators tree WellLocationData_inTableX Text_processing Csv T3_Text Key2 Predictive_Modellin g CSv,Table, File Learners_Predcitors Table A The Table B illustrates an exemplary method of ranking or scoring the one or more data analytics operators 211 comprised in at least one of a partial data analytics flow and one or more complete data analytics flows based on the relevance of each of the operators with the received industry specific information 111. Industry Operator Relevance Oil&Gas Op1 2.5 Mining Op2 0 Table B In an embodiment, as shown in Fig. 4b, the user may select the goal of recommendation to be predictive modeling. Upon receiving the user selection of goal information 110 and the data to be analyzed 104, the recommendation module 117 finds one or more appropriate learner unit from its knowledge base. As an example, the recommendation module 117 may select a specific decision tree comprising one or more predictor units, in order to recommend a most relevant data analytics process 213 to the user. In an embodiment, the recommendation module 117 may also refer to a scorer module for identifying the most relevant data analytics operator for the identified dependency tree. The data analytics operator ranking is based on the one or more goal information 110 and the usage of each of the data analytics operators 211 for the goal information 110 stored in the knowledge repository 219. Further, upon identifying the most relevant data analytics process 213, the recommendation module 117 proceeds to display the identified data analytics processes 213 on a display interface, as shown in Fig. 3a and Fig. 3b. In an embodiment, one of a visualization technique may be used to display the identified data analytics processes 213 on the display interface. As an example, the visualization techniques used by the recommendation module 117 may include, but not limited to, bqplot, scatter plot and any other suitable interface. Figure 5 is a block diagram of an exemplary computer system for implementing embodiments consistent with the present disclosure. Variations of computer system 501 may be used for implementing all the computing systems that may be utilized to implement the features of the present disclosure. Computer system 501 may comprise a central processing unit (“CPU” or “processor”) 502. Processor 502 may comprise at least one data processor for executing program components for executing user or system-generated requests. The processor may include specialized processing units such as integrated system (bus) controllers, memory management control units, floating point units, graphics processing units, digital signal processing units, etc. The processor 502 may include a microprocessor, such as AMD Athlon, Duron or Opteron, ARM’s application, embedded or secure processors, IBM PowerPC, Intel’s Core, Itanium, Xeon, Celeron or other line of processors, etc. The processor 502 may be implemented using mainframe, distributed processor, multi-core, parallel, grid, or other architectures. Some embodiments may utilize embedded technologies like application-specific integrated circuits (ASICs), digital signal processors (DSPs), Field Programmable Gate Arrays (FPGAs), etc. Processor 502 may be disposed in communication with one or more input/output (I/O) devices via I/O interface 503. The I/O interface 503 may employ communication protocols/methods such as, without limitation, audio, analog, digital, monoaural, RCA, stereo, IEEE-1394, serial bus, universal serial bus (USB), infrared, PS/2, BNC, coaxial, component, composite, digital visual interface (DVI), high-definition multimedia interface (HDMI), RF antennas, S-Video, VGA, IEEE 802.n /b/g/n/x, Bluetooth, cellular (e.g., code-division multiple access (CDMA), high-speed packet access (HSPA+), global system for mobile communications (GSM), long-term evolution (LTE), WiMax, or the like), etc. Using the I/O interface 503, the computer system 501 may communicate with one or more I/O devices. For example, the input device 504 may be an antenna, keyboard, mouse, joystick, (infrared) remote control, camera, card reader, fax machine, dongle, biometric reader, microphone, touch screen, touchpad, trackball, sensor (e.g., accelerometer, light sensor, GPS, gyroscope, proximity sensor, or the like), stylus, scanner, storage device, transceiver, video device/source, visors, etc. Output device 505 may be a printer, fax machine, video display (e.g., cathode ray tube (CRT), liquid crystal display (LCD), light-emitting diode (LED), plasma, or the like), audio speaker, etc. In some embodiments, a transceiver 506 may be disposed in connection with the processor 502. The transceiver may facilitate various types of wireless transmission or reception. For example, the transceiver may include an antenna operatively connected to a transceiver chip (e.g., Texas Instruments WiLink WL1283, Broadcom BCM4750IUB8, Infineon Technologies X-Gold 618-PMB9800, or the like), providing IEEE 802.11a/b/g/n, Bluetooth, FM, global positioning system (GPS), 2G/3G HSDPA/HSUPA communications, etc. In some embodiments, the processor 502 may be disposed in communication with a communication network 508 via a network interface 507. The network interface 507 may communicate with the communication network 508. The network interface 507 may employ connection protocols including, without limitation, direct connect, Ethernet (e.g., twisted pair 10/40/400 Base T), transmission control protocol/internet protocol (TCP/IP), token ring, IEEE 802.11a/b/g/n/x, etc. The communication network 508 may include, without limitation, a direct interconnection, local area network (LAN), wide area network (WAN), wireless network (e.g., using Wireless Application Protocol), the Internet, etc. Using the network interface 507 and the communication network 508, the computer system 501 may communicate with devices 509, 510, and 511. These devices may include, without limitation, personal computer(s), server(s), fax machines, printers, scanners, various mobile devices such as cellular telephones, smartphones (e.g., Apple iPhone, Blackberry, Android-based phones, etc.), tablet computers, eBook readers (Amazon Kindle, Nook, etc.), laptop computers, notebooks, gaming consoles (Microsoft Xbox, Nintendo DS, Sony PlayStation, etc.), or the like. In some embodiments, the computer system 501 may itself embody one or more of these devices. In some embodiments, the processor 502 may be disposed in communication with one or more memory devices (e.g., RAM 513, ROM 5Error! Reference source not found.14, etc.) via a storage interface 512. The storage interface may connect to memory devices including, without limitation, memory drives, removable disc drives, etc., employing connection protocols such as serial advanced technology attachment (SATA), integrated drive electronics (IDE), IEEE-1394, universal serial bus (USB), fiber channel, small computer systems interface (SCSI), etc. The memory drives may further include a drum, magnetic disc drive, magneto-optical drive, optical drive, redundant array of independent discs (RAID), solid-state memory devices, solid-state drives, etc. The memory 515 may store a collection of program or database components, including, without limitation, an operating system 5Error! Reference source not found.16, user interface application 5Error! Reference source not found.17, web browser 518, mail server 519, mail client 520, user/application data 521 (e.g., any data variables or data records discussed in this disclosure), etc. The operating system 516 may facilitate resource management and operation of the computer system 501. Examples of operating systems include, without limitation, Apple Macintosh OS X, UNIX, Unix-like system distributions (e.g., Berkeley Software Distribution (BSD), FreeBSD, NetBSD, OpenBSD, etc.), Linux distributions (e.g., Red Hat, Ubuntu, Kubuntu, etc.), IBM OS/2, Microsoft Windows (XP, Vista/7/8, etc.), Apple iOS, Google Android, Blackberry OS, or the like. User interface 517 may facilitate display, execution, interaction, manipulation, or operation of program components through textual or graphical facilities. For example, user interfaces may provide computer interaction interface elements on a display system operatively connected to the computer system 501, such as cursors, icons, check boxes, menus, scrollers, windows, widgets, etc. Graphical user interfaces (GUIs) may be employed, including, without limitation, Apple Macintosh operating systems’ Aqua, IBM OS/2, Microsoft Windows (e.g., Aero, Metro, etc.), Unix X-Windows, web interface libraries (e.g., ActiveX, Java, Javascript, AJAX, HTML, Adobe Flash, etc.), or the like. In some embodiments, the computer system 501 may implement a web browser 518 stored program component. The web browser may be a hypertext viewing application, such as Microsoft Internet Explorer, Google Chrome, Mozilla Firefox, Apple Safari, etc. Secure web browsing may be provided using HTTPS (secure hypertext transport protocol), secure sockets layer (SSL), Transport Layer Security (TLS), etc. Web browsers may utilize facilities such as AJAX, DHTML, Adobe Flash, JavaScript, Java and application programming interfaces (APIs), etc. In some embodiments, the computer system 501 may implement a mail server 519 stored program component. The mail server may be an Internet mail server such as Microsoft Exchange, or the like. The mail server may utilize facilities such as ASP, ActiveX, ANSI C++/C#, Microsoft .NET, CGI scripts, Java, JavaScript, PERL, PHP, Python, WebObjects, etc. The mail server may utilize communication protocols such as internet message access protocol (IMAP), messaging application programming interface (MAPI), Microsoft Exchange, post office protocol (POP), simple mail transfer protocol (SMTP), or the like. In some embodiments, the computer system 501 may implement a mail client 520 stored program component. The mail client may be a mail viewing application, such as Apple Mail, Microsoft Entourage, Microsoft Outlook, Mozilla Thunderbird, etc. In some embodiments, computer system 501 may store user/application data 521, such as the data, variables, records, etc. as described in this disclosure. Such databases may be implemented as fault-tolerant, relational, scalable, secure databases such as Oracle or Sybase. Alternatively, such databases may be implemented using standardized data structures, such as an array, hash, linked list, struct, structured text file (e.g., XML), table, or as object-oriented databases (e.g., using ObjectStore, Poet, Zope, etc.). Such databases may be consolidated or distributed, sometimes among the various computer systems discussed above in this disclosure. It is to be understood that the structure and operation of the any computer or database component may be combined, consolidated, or distributed in any working combination. As described above, the modules 205, amongst other things, include routines, programs, objects, components, and data structures, which perform particular tasks or implement particular abstract data types. The modules 205 may also be implemented as, signal processor(s), state machine(s), logic circuitries, and/or any other device or component that manipulate signals based on operational instructions. Further, the modules 205 can be implemented by one or more hardware components, by computer-readable instructions executed by a processing unit, or by a combination thereof. The illustrated steps are set out to explain the exemplary embodiments shown, and it should be anticipated that ongoing technological development will change the manner in which particular functions are performed. These examples are presented herein for purposes of illustration, and not limitation. Further, the boundaries of the functional building blocks have been arbitrarily defined herein for the convenience of the description. Alternative boundaries can be defined so long as the specified functions and relationships thereof are appropriately performed. Alternatives (including equivalents, extensions, variations, deviations, etc., of those described herein) will be apparent to persons skilled in the relevant art(s) based on the teachings contained herein. Such alternatives fall within the scope and spirit of the disclosed embodiments. Also, the words "comprising," "having," "containing," and "including," and other similar forms are intended to be equivalent in meaning and be open ended in that an item or items following any one of these words is not meant to be an exhaustive listing of such item or items, or meant to be limited to only the listed item or items. It must also be noted that as used herein and in the appended claims, the singular forms “a,” “an,” and “the” include plural references unless the context clearly dictates otherwise. Furthermore, one or more computer-readable storage media may be utilized in implementing embodiments consistent with the present disclosure. A computer-readable storage medium refers to any type of physical memory on which information or data readable by a processor may be stored. Thus, a computer-readable storage medium may store instructions for execution by one or more processors, including instructions for causing the processor(s) to perform steps or stages consistent with the embodiments described herein. The term “computer-readable medium” should be understood to include tangible items and exclude carrier waves and transient signals, i.e., are non-transitory. Examples include random access memory (RAM), read-only memory (ROM), volatile memory, nonvolatile memory, hard drives, CD ROMs, DVDs, flash drives, disks, and any other known physical storage media. It is intended that the disclosure and examples be considered as exemplary only, with a true scope and spirit of disclosed embodiments being indicated by the following claims. Referral Numerals: Reference Number Description 100 Environment 101 Data source 103 Input data 104 Data to be analyzed 105 Configuration information 109 Data analytics process recommendation system 110 Goal information 111 Industry specific information 113 Processor 115 Memory 117 Recommendation module 201 I/O Interface 205 Modules 209 Similarity matrix 219 Knowledge repository 211 Data analytics operators 213 Data analytics processes 215 Other data 221 Knowledge Repository Interface Module 223 Other modules
Claims:1. A data analytics process recommendation system for recommending data analytics processes to a user, the data analytics process recommendation system comprising:
a processor; and
a memory communicatively coupled to the processor, wherein the memory stores processor-executable instructions, which, on execution, causes the processor to:
receive input data related to the data analytics processes from the user;
identify at least one of one or more partial data analytics process and one or more complete data analytics processes from a knowledge repository associated with the data analytics process recommendation system based on the received input data related to the data analytics processes; and
recommend at least one of the one or more partial data analytics process and the one or more complete data analytics processes to the user based on relevance of each of the one or more identified data analytics processes to the input data thus received, wherein each of the one or more partial data analytics process comprises a data analytics operator and each of the one or more complete data analytics processes comprises plurality of data analytics operators.
2. The process recommendation system as claimed in claim 1 further comprises recommending one or more subsequent data analytics operators upon receiving user selection on the data analytics operator of one of the one or more partial data analytics processes thus recommended.
3. The process recommendation system as claimed in claim 2, wherein a machine learning technique comprised in the knowledge repository learns the user selection on the data analytics operator of one of the one or more partial data analytics processes recommended based on the input data for improving the relevance of recommendation of the data analytics processes.
4. The process recommendation system as claimed in claim 1 and 2, wherein a display interface in the data analytics process recommendation system displays at least one of the one or more partial data analytics process and the one or more complete data analytics processes, and the one or more subsequent data analytics operators being recommended to the user.
5. The process recommendation system as claimed in claim 4 further comprises providing a notification to the user, wherein the notification displays one or more information related to the recommended at least one of the one or more partial data analytics process and the one or more complete data analytics processes, and the subsequent data analytics operators.
6. The process recommendation system as claimed in claim 1 further comprises receiving at least one of industry specific information and goal information related to the data analytics processes for recommending at least one of the one or more partial data analytics process and the one or more complete data analytics processes to the user.
7. The process recommendation system as claimed in claim 1, wherein a similarity matrix configured in the knowledge repository stores one or more data analytics processes which are functionally related to data analytics operators.
8. The process recommendation system as claimed in claim 1 further comprises optimizing the recommended at least one of the one or more partial data analytics process and the one or more complete data analytics processes using predefined configuration details.
9. A method of recommending data analytics processes to a user, the method comprising:
receiving, by a processor of a data analytics process recommendation system, input data related to the data analytics processes from the user;
identifying, by the processor, at least one of one or more partial data analytics process and one or more complete data analytics processes from a knowledge repository associated with the data analytics process recommendation system based on the received input data related to the data analytics processes; and
recommending, by the processor, at least one of the one or more partial data analytics process and the one or more complete data analytics processes to the user based on relevance of each of the one or more identified data analytics processes to the input data thus received, wherein each of the one or more partial data analytics process comprises a data analytics operator and each of the one or more complete data analytics processes comprises plurality of data analytics operators.
10. A non-transitory computer readable medium including instructions stored thereon that when processed by at least one processor cause a data analytics process recommendation system to perform operations comprising:
receiving, by a processor of a data analytics process recommendation system, input data related to the data analytics processes from the user;
identifying, by the processor, at least one of one or more partial data analytics process and one or more complete data analytics processes from a knowledge repository associated with the data analytics process recommendation system based on the received input data related to the data analytics processes; and
recommending, by the processor, at least one of the one or more partial data analytics process and the one or more complete data analytics processes to the user based on relevance of each of the one or more identified data analytics processes to the input data thus received, wherein each of the one or more partial data analytics process comprises a data analytics operator and each of the one or more complete data analytics processes comprises plurality of data analytics operators.
, Description:FIELD OF THE DISCLOSURE:
The present subject matter is related, in general to data analytics, and more particularly, but not exclusively to a method and system for recommending data analytics processes to a user.
BACKGROUND:
Data analytics is the process of examining large data sets containing a variety of data types to uncover hidden patterns, unknown correlations, market trends, customer preferences and other useful business information. The primary goal of data analytics is to help organizations make more informed business decisions by enabling data scientists, predictive modelers and other analytics professionals to analyze large volumes of transaction data, as well as other forms of data that may be untapped by conventional business intelligence methods. Various forms of unstructured data may include, but are not limited, Web server logs and Internet clickstream data, social media content and social network activity reports, mobile-phone call detail records and machine data captured by sensors connected to the Internet of Things (IoT).
The unstructured data may not fit well in traditional data warehouses that are based on relational databases. Furthermore, data warehouses may not be able to handle the processing demands posed by sets of big data that need to be updated frequently or even continually. As a result, many organizations looking to collect, process and analyze big data have turned to a newer class of applications that support the processing of large and diverse data sets across clustered systems.
However, these data analytics applications can be viewed as a pipeline or analytics process of operators i.e. algorithms, processes, libraries and hardware exposed via Application Program Interface (API). The analytic data processes includes various operations related to the process of Extract-Transform-Load (ETL) processes, combined with analytic, querying and reporting that consumes a variety of data types. These processes are complex to design, difficult to implement and maintain since they span a variety of systems. Additionally, various quality objectives such as reliability, maintainability, performance, pipeline portioning and parallelization have to be met. Therefore, specific decision support systems are needed to reduce development time and maintenance costs associated with the analytics process.
Conventional data analytic processes are built using inputs provided by Domain/ IT expert as per desired functionalities. Even if the data analytic processes are functionally correct, improving the data analytics process is still difficult without the domain specific inputs. Also, conventional techniques doesn’t utilize the already available versions of the process stored in a knowledge repository, rather handles the data analytic processes based on inputs from Domain/ IT expert.
Therefore, there is a need for a method and a system that automatically builds data analytic processes and utilizes the already available version of the processes without the involvement of Domain/IT expert, thereby reducing the design cycle time and cost involved and overcoming the limitations of the existing art.
SUMMARY:
One or more shortcomings of the prior art are overcome and additional advantages are provided through the present disclosure. Additional features and advantages are realized through the techniques of the present disclosure. Other embodiments and aspects of the disclosure are described in detail herein and are considered a part of the claimed disclosure.
Accordingly, the present disclosure relates to a method of recommending data analytics processes to a user. The method comprises the steps of receiving input data related to the data analytics processes from the user. The method further comprises the steps of identifying at least one of one or more partial data analytics process and one or more complete data analytics processes from a knowledge repository associated with the data analytics process recommendation system based on the received input data related to the data analytics processes. Upon identifying, the method recommends at least one of the one or more partial data analytics process and the one or more complete data analytics processes to the user based on relevance of each of the one or more identified data analytics processes to the input data thus received. Each of the one or more partial data analytics process comprises a data analytics operator and each of the one or more complete data analytics processes comprises plurality of data analytics operators.
Further, the present disclosure relates to a data analytics process recommendation system for recommending data analytics processes to a user. The system comprises a processor and a memory communicatively coupled to the processor, wherein the memory stores processor-executable instructions, which, on execution, cause the processor to receive input data related to the data analytics operator from the user. The processor is further configured to identify at least one of one or more partial data analytics process and one or more complete data analytics processes from a knowledge repository associated with the data analytics process recommendation system based on the received input data related to the data analytics processes. The processor is furthermore configured to recommend at least one of the one or more partial data analytics process and the one or more complete data analytics processes to the user based on relevance of each of the one or more identified data analytics processes to the input data thus received, wherein each of the one or more partial data analytics process comprises a data analytics operator and each of the one or more complete data analytics processes comprises plurality of data analytics operators.
Furthermore, the present disclosure relates to a non-transitory computer readable medium including instructions stored thereon that when processed by at least one processor causes a data analytics process recommendation system to perform operations comprising receiving input data related to the data analytics operator from the user. Further, the instructions cause the processor to identify at least one of one or more partial data analytics process and one or more complete data analytics processes from a knowledge repository associated with the data analytics process recommendation system based on the received input data related to the data analytics processes. Furthermore, at least one of the one or more partial data analytics process and the one or more complete data analytics processes to the user based on relevance of each of the one or more identified data analytics processes to the input data thus received, wherein each of the one or more partial data analytics process comprises a data analytics operator and each of the one or more complete data analytics processes comprises plurality of data analytics operators.
The foregoing summary is illustrative only and is not intended to be in any way limiting. In addition to the illustrative aspects, embodiments, and features described above, further aspects, embodiments, and features will become apparent by reference to the drawings and the following detailed description.
BRIEF DESCRIPTION OF THE DRAWINGS:
The accompanying drawings, which are incorporated in and constitute a part of this disclosure, illustrate exemplary embodiments and, together with the description, serve to explain the disclosed principles. In the figures, the left-most digit(s) of a reference number identifies the figure in which the reference number first appears. The same numbers are used throughout the figures to reference like features and components. Some embodiments of system and/or methods in accordance with embodiments of the present subject matter are now described, by way of example only, and with reference to the accompanying figures, in which:
Fig. 1 illustrates an architecture diagram of an exemplary system for recommending data analytics processes to a user in accordance with some embodiments of the present disclosure;
Fig. 2 illustrates an exemplary block diagram of a data analytics process recommendation system of Fig. 1 in accordance with some embodiments of the present disclosure;
Figs. 3a and 3b illustrate an exemplary method of representing and recommending data analytics processes on a display interface of the data analytics process recommendation system;
Fig. 4a illustrates a detailed process chart of an exemplary method of recommending data analytics processes to a user in accordance with some embodiments of the present disclosure;
Fig. 4b illustrates the method of recommending data analytics processes to a user in accordance with an exemplary embodiment of the present disclosure; and
Fig. 5 is a block diagram of an exemplary computer system for implementing embodiments consistent with the present disclosure.
It should be appreciated by those skilled in the art that any block diagrams herein represent conceptual views of illustrative systems embodying the principles of the present subject matter. Similarly, it will be appreciated that any process charts, process diagrams, state transition diagrams, pseudo code, and the like represent various processes which may be substantially represented in computer readable medium and executed by a computer or processor, whether or not such computer or processor is explicitly shown.
DETAILED DESCRIPTION:
In the present document, the word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any embodiment or implementation of the present subject matter described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments.
While the disclosure is susceptible to various modifications and alternative forms, specific embodiment thereof has been shown by way of example in the drawings and will be described in detail below. It should be understood, however that it is not intended to limit the disclosure to the particular forms disclosed, but on the contrary, the disclosure is to cover all modifications, equivalents, and alternative falling within the scope of the disclosure.
The terms “comprises”, “comprising”, or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a setup, device or method that comprises a list of components or steps does not include only those components or steps but may include other components or steps not expressly listed or inherent to such setup or device or method. In other words, one or more elements in a system or apparatus proceeded by “comprises… a” does not, without more constraints, preclude the existence of other elements or additional elements in the system or apparatus.
The present disclosure relates to a method and a system for recommending data analytics processes to a user. In one embodiment, the method receives input data related to the data analytics operator from the user. The method further identifies the one or more data analytics processes comprising at least one of one or more partial data analytics process and one or more complete data analytics processes from a knowledge repository associated with the data analytics process recommendation system based on the received input data related to the data analytics operators. Each of the one or more partial data analytics process comprises a data analytics operator. Further, each of the one or more complete data analytics processes comprises plurality of data analytics operators. Furthermore, the method recommends the identified at least one of the one or more partial data analytics process and the one or more complete data analytics processes to the user based on relevance of each of the one or more identified data analytics processes to the input data thus received. In an embodiment of the present disclosure, the method also recommends one or more subsequent data analytics operators upon receiving user selection on the data analytics operator of one of the one or more partial data analytics processes thus recommended. Therefore, the proposed data analytics process recommendation system makes use of the related to the data analytics operators and previously stored data analytics processes to recommend one or more subsequent data analytics operators and data analytics processes with minimum optimization time and cost.
In the following detailed description of the embodiments of the disclosure, reference is made to the accompanying drawings that form a part hereof, and in which are shown by way of illustration specific embodiments in which the disclosure may be practiced. These embodiments are described in sufficient detail to enable those skilled in the art to practice the disclosure, and it is to be understood that other embodiments may be utilized and that changes may be made without departing from the scope of the present disclosure. The following description is, therefore, not to be taken in a limiting sense.
Fig. 1 illustrates an architecture diagram of an exemplary system for recommending data analytics processes to a user in accordance with some embodiments of the present disclosure.
As shown in Fig. 1, the exemplary system 100 comprises one or more components configured to recommend one or more data analytics processes 213 to a user. In one embodiment, the exemplary system 100 comprises a data source 101, and a data analytics process recommendation system 109.
In one embodiment, the data source 101 comprises input data 103 related to one or more data analytics processes 213. In an embodiment, the data analytics processes 213 may be one or more partial data analytics processes. Each of the one or more partial data analytics process comprises a single data analytics operator. In another embodiment, the data analytics process 213 may be a complete data analytics process. The complete data analytics process comprises plurality of data analytics operators 211.
As an example the data analytics processes 213 received from the user may include a reading operation, aggregation of rows of data in a database, parsing of data in a database and running a simple machine learning algorithm. The data analytics process recommendation system 109 recommends at least one of the one or more partial data analytics process and the one or more complete data analytics processes to the users. In an embodiment, the input data 103 may include, but not limited to, a data file of a specific format which comprises one or more information related to the data analytics processes. As an example, the input file may be in one of the data file formats, such as txt, doc, pdf and xml. In another embodiment, the input data 103 may also contain at least one of one or more data analytics operators 211, data analytics processes 213, goal information 110, industry specific information 111 and data to be analyzed 104. The input data 103 received from the data source 101 may be used by the data analytics process recommendation system 109 for recommending one or more data analytics processes 213 to the user.
Fig. 2 illustrates an exemplary block diagram of a data analytics process recommendation system of Fig. 1 in accordance with some embodiments of the present disclosure.
In one embodiment, the data analytics process recommendation system 109 comprises a central processing unit (“CPU” or “processor”) 113, a memory 115 and an I/O interface 201. The I/O interface 201 is coupled with the processor 113 and one or more I/O devices. An I/O device, such as a keyboard, is configured to receive inputs via the I/O interface 201. Similarly, the processed outputs for displaying in an I/O device, such as a monitor, are transmitted via the I/O interface 201.
In an embodiment, the memory 115 of the data analytics process recommendation system 109 further comprises a knowledge repository 219 and modules 205 comprising, a recommendation module 117 and a knowledge repository interface module 221. In one example, the knowledge repository 219 may include, but not limited to, a similarity matrix 209 and one or more configuration information 105. The configuration information 105 may include, but not limited to, goal information 110 and industry specific information 111 associated with the one or more data analytics processes 213. The goal information 110 related to a data analytics process 213 defines the goal of the data analytics process recommendation. In an embodiment, the goal information 110 may include, but not limited to, text processing, predictive modeling, descriptive modeling, pattern discovery, forecasting, and text analytics. The industry specific information 111 related to a data analytics process 213 comprises at least one of name of the industry, subject matter of the analytics industry and one or more specific rules associated with the analytics industry. As an example, the analytics industry may be Healthcare, Oil and Gas, Mining, Social Media and any other suitable analytics industry. In an embodiment, the goal information 110 and the industry specific information 111 may be stored in the form of one or more metadata files. In an embodiment, the user uses the one or more goal information 110 and the industry specific information 111 stored in the similarity matrix 209 to prune the recommendation of one or more data analytics processes 213.
In an embodiment, one or more data analytics operators 211 and one or more data analytics processes 213 may be stored in the similarity matrix 209. As an example, the similarity matrix 209 can be a structured database storing one or more data analytics processes 213 and one or more data analytics operators 211 related to a specific combination of the goal information 110 and the industry specific information 111. The similarity matrix 209 may also store one or more data analytics operators 211 and one or more data analytics processes that are previously recommended to a user. In an embodiment, the one or more data analytics operators 211 and the one or more data analytics processes 213 stored in the similarity matrix 209 may be alternatively generated by industry expert personnel based on the one or more goal information 110 and the industry specific information 111.
In yet another embodiment, the data analytics process recommendation system 109 may optimize the one or more recommended data analytics processes 213 based on the one or more optimization parameters. The optimization parameters determine the type of optimization trade-offs needed by the user. The one or more optimized data analytics processes stored in the knowledge repository 219 can be selected based on the optimization parameters. The optimization parameters may include, but not limited to, measures for reliability, maintainability, performance, pipelining and parallelization. In an example, the reliability measures are determined by the total number of recovery points present in a data analytics process 213. The measures for maintainability may include length of the data analytics process 213, complexity in execution of the data analytics process 213, degree of modularity associated with the data analytics process 213 and the extent of coupling between the one or more modules 205 in the data analytics process 213. The measures for performance may include total memory 115 space consumed while performing the data analytics process 213, total processor 113 runtime consumed and throughput associated with the execution of the data analytics process 213. The measures for pipelining and parallelization may include various factors associated with pipelining, such as, partition type, number of data volumes in each pipeline partition, length of data volumes in each pipeline partition and number of processing units under execution.
In an exemplary embodiment, the user may want to optimize a process to include maximum fault tolerance. Hence, the recommendation module 117 has to implement the measures for reliability in order to generate a fault tolerant process. Then, the recommendation module 117 has to consider including effective number of recovery points in the optimized data analytics process 213. As an example, for a simple data analytics process with ‘n’ number of operators, there must be ‘n-1’ recovery points to ensure maximum fault tolerance. Further, upon optimizing the one or more data analytics processes 213 to the user, the optimized data analytics processes are stored back into the knowledge repository 219. The stored optimized data analytics processes may be used by the data analytics process recommendation system 109 to quickly recommend one or more data analytics processes 213 to the user upon receiving a relevant data analytics operator 211.
In an embodiment, the data 203 may be stored in the memory 115 in the form of various data structures. Additionally, the aforementioned data can be organized using data models, such as relational or hierarchical data models. The other data 219 may be also referred to as reference repository for storing recommended implementation approaches as reference data. In an embodiment, the other data 219 may store on or more configuration data related to the one or more partial data analytics processes and one or more complete data analytics processes. The other data 219 also stores data, including temporary data and temporary files, generated by the modules 205 for performing the various functions of the data analytics process recommendation system 109.
The modules 205 may include, for example, the recommendation module 117, a knowledge repository interface module 221 and other modules 223. The other modules 223 perform various miscellaneous functionalities of the data analytics process recommendation system 109. It will be appreciated that such aforementioned modules may be represented as a single module or a combination of different modules. The modules 205 may be implemented in the form of software, hardware and/or firmware.
The data analytics process recommendation system 109 receives the input data 103 associated with the data analytics processes 213 from a user. The data analytics processes 213 may include at least one of one or more partial data analytics process and one or more complete data analytics processes stored in a knowledge repository 219 associated with the data analytics process recommendation system 109.
In an embodiment, the recommendation module 117 recommends one or more data analytics processes 213 to the users based on the relevance of each of the data analytics processes 213 stored in the knowledge repository 219 to the data to be analyzed 104. In one embodiment, the knowledge repository 219 may be integrated within the data analytics process recommendation system 109. In another embodiment, the knowledge repository 219 may be implemented independent of the data analytics process recommendation system 109. The knowledge repository 219 may be for example, a data warehouse configured to store metadata related to one or more optimized data analytics processes that are previously executed and saved by the user.
The recommendation module 117 selects one of the one or more data analytics processes 213, having maximum relevance to the data analytics processes 213 stored in the knowledge repository 219 and recommends it to the user. The recommended processes are then saved to the knowledge repository 219 for their future use. The knowledge repository interface module 221 is used to access the knowledge repository 219 that may be integrated within the data analytics process recommendation system 109. In an embodiment, the knowledge repository 219 can be integrated independent of the data analytics process recommendation system 109.
Figs. 3a and 3b illustrate an exemplary method of representing and recommending data analytics processes on a display interface of the data analytics process recommendation system.
Fig. 3a illustrates the method of recommending one or more data analytics processes 213 to the user. In an embodiment, the display interface 300 in the data analytics process recommendation system 109 may be further divided into a selection panel 301 and a process recommendation area 303. The data analytics process recommendation system 109 displays one or more information, such as, one or more operators, one or more domains, input data 103 and one or more other configuration data on the selection panel 301 of the display interface 300. The user may select one of the one or more operators, the one or more domains, the input data 103 and the one or more other configuration data, displayed in the selection panel 301, to obtain a most relevant data analytics process 213 as the recommendation from the data analytics process recommendation system 109. In an exemplary embodiment, the one or more information displayed on the selection panel 301 may be enabled with a drag and drop technique, using which, the user may drag one or more required information from the selection pane 301 and drop them on the process recommendation area 303 of the display interface 300.
In an exemplary embodiment, the user may input a data to the data analytics process recommendation system 109. In an embodiment, the input data 103 may include one of one or more data analytics operators 211, data analytics processes 213, goal information 110, industry specific information 111 and one or more data 104 to which the user needs to receive one or more recommendations. Accordingly, the data analytics process recommendation system 109 identifies one or more data analytics processes 213, from the knowledge repository 219, that are functionally equivalent to the received input data 103. Further, the data analytics process recommendation system 109 recommends one or more identified data analytics processes X, Y and Z, and one or more partial data analytics processes, namely, Opr 1 and Opr 2 on the process recommendation area 303 based on the relevance of each of the identified data analytics processes to the received input data 103. In an embodiment, the user may select one of the one or more recommended data analytics processes X, Y and Z, and Opr 1 and Opr 2 if the user finds the recommended processes to be in match with the pre-selected goal information 110 and industry specific information 111.
In yet another embodiment, as illustrated in Fig. 3b, if none of the recommendations are in compliance with the pre-selected domains, the user may wish to change one or more individual data analytics operators 211 in one of the recommended data analytics processes. As an example, the user may prefer to alter one of the data analytics operator Y1 of the recommended data analytics process Y, if the complete data analytics process Y is not acceptable by the user. Alternatively, the data analytics process recommendation system 109 recommends a plurality of data analytics operators 211 to the selected data analytic operator Y1, such that, the recommended data analytics operators 211 are functionally equivalent to Y1, thus recommending a complete data analytics process 213. Further, the one or more data analytics operators 211 and the data analytics process 213 selected by the user are saved back in the knowledge repository 219 for aiding in the future recommendations.
Figure 4a illustrates a detailed process chart of an exemplary method of recommending data analytics processes to a user in accordance with some embodiments of the present disclosure;
As illustrated in Figure 4, the method 400 comprises one or more blocks implemented by the processor 113 for recommending data analytics processes 213 to a user. The method 400 may be described in the general context of computer executable instructions. Generally, computer executable instructions can include routines, programs, objects, components, data structures, procedures, modules, and functions, which perform particular functions or implement particular abstract data types.
The order in which the method 400 is described is not intended to be construed as a limitation, and any number of the described method blocks can be combined in any order to implement the method 400. Additionally, individual blocks may be deleted from the method 400 without departing from the spirit and scope of the subject matter described herein. Furthermore, the method 400 can be implemented in any suitable hardware, software, firmware, or combination thereof.
At block 401, the data analytics process recommendation system 109 receives input data 103 related to the data analytics operator from the user. The input data 103 may contain one or more data analytics operators 211, data analytics processes 213, goal information 110, industry specific information 111 and data to be analyzed 104. The input data 103 received from the data source 101 may be used by the data analytics process recommendation system 109 for recommending one or more data analytics processes 213 to the data to be analyzed 104.
At block 403, the data analytics process recommendation system 109 identifies at least one of one or more partial data analytics process and one or more complete data analytics processes from a knowledge repository 219 associated with the data analytics process recommendation system 109 based on the received data related to the data analytics processes 213.
At block 405, the process recommendation module 117 recommends at least one of the one or more partial data analytics process and the one or more complete data analytics processes to the user based on relevance of each of the one or more identified data analytics processes 213 to the input data 103, wherein each of the one or more partial data analytics process comprises a data analytics operator and each of the one or more complete data analytics processes comprises plurality of data analytics operators 211.
Fig. 4b illustrates the method of recommending data analytics processes to a user in accordance with an exemplary embodiment of the present disclosure.
In an embodiment, the data analytics process recommendation system 109 may receive one or more input data 103 from the user. The input data 103 may include one or more information related to the data analytics processes 213 in one of the forms including, but not limited to, a CSV, a file, a tabular representation of the information and data to be 104, having a specific schema, on which the user wants to perform the data analytics operations. Additionally the user may also input the goal information 110 and the industry specific information 111, related to the data analytics process 213, to the data analytics process recommendation system 109. The data analytics process recommendation system 109 may use the received goal information 110 and the industry specific information 111 to improve the recommendations on one of the partial data analytics process and the one or more complete data analytics flows.
In an embodiment, the data analytics process recommendation system 109 as shown in Fig. 4b may further use the received goal information 110 and the industry specific information 111 to derive a most relevant dependency tree corresponding to the data to be analyzed 104 and other information. The dependency trees are the data structures that are stored in the knowledge repository 219 of the recommendation system 109 along with the data sets previously used for the data analytics processes 213 and the other info like industry specific information 111. Furthermore, the recommendation module 117 makes use of usage history and one or more configuration details associated with the data analytics process recommendation module 117 to identify one or more additional data analytics processes 213 from the obtained dependency tree.
In an embodiment, the data analytics process recommendation system 109 may save one or more samples of each of the data including the goal information 110 and the industry specific information 111 in the knowledge repository 219 for its future use. The data analytics process recommendation system 109 may save the sample data each time the user requests a recommendation. In an embodiment, the recommendation module 117 may also save one or more information to find out a match between the data analytics process 213 and the used data. Similarly, the recommendation module 117 may derive a match between the goal information 110 and the corresponding data analytic operators which can be recommended to find out the recommended combination of goal information 110 and operators. The data stored in the knowledge repository 219 may enable the user of the data analytics process recommendation system 109 to select a most appropriate dependency tree and also to rank the suitable branches among the various branches going from one node in dependency tree. Additionally, the recommendation module 117 may also save one or more feedback information received from the user for improving the future recommendations.
Table A depicted below illustrates an exemplary similarity matrix 209 in which the recommendation module 117 receives the one or more data sets and the corresponding goal information 110 to find a most relevant dependency tree comprising one or more data analytics operators 211.
Data Set Goal Data type Operators tree
WellLocationData_inTableX Text_processing Csv T3_Text
Key2 Predictive_Modellin g CSv,Table, File Learners_Predcitors
Table A
The Table B illustrates an exemplary method of ranking or scoring the one or more data analytics operators 211 comprised in at least one of a partial data analytics flow and one or more complete data analytics flows based on the relevance of each of the operators with the received industry specific information 111.
Industry Operator Relevance
Oil&Gas Op1 2.5
Mining Op2 0
Table B
In an embodiment, as shown in Fig. 4b, the user may select the goal of recommendation to be predictive modeling. Upon receiving the user selection of goal information 110 and the data to be analyzed 104, the recommendation module 117 finds one or more appropriate learner unit from its knowledge base. As an example, the recommendation module 117 may select a specific decision tree comprising one or more predictor units, in order to recommend a most relevant data analytics process 213 to the user. In an embodiment, the recommendation module 117 may also refer to a scorer module for identifying the most relevant data analytics operator for the identified dependency tree. The data analytics operator ranking is based on the one or more goal information 110 and the usage of each of the data analytics operators 211 for the goal information 110 stored in the knowledge repository 219. Further, upon identifying the most relevant data analytics process 213, the recommendation module 117 proceeds to display the identified data analytics processes 213 on a display interface, as shown in Fig. 3a and Fig. 3b. In an embodiment, one of a visualization technique may be used to display the identified data analytics processes 213 on the display interface. As an example, the visualization techniques used by the recommendation module 117 may include, but not limited to, bqplot, scatter plot and any other suitable interface.
Figure 5 is a block diagram of an exemplary computer system for implementing embodiments consistent with the present disclosure.
Variations of computer system 501 may be used for implementing all the computing systems that may be utilized to implement the features of the present disclosure. Computer system 501 may comprise a central processing unit (“CPU” or “processor”) 502. Processor 502 may comprise at least one data processor for executing program components for executing user or system-generated requests. The processor may include specialized processing units such as integrated system (bus) controllers, memory management control units, floating point units, graphics processing units, digital signal processing units, etc. The processor 502 may include a microprocessor, such as AMD Athlon, Duron or Opteron, ARM’s application, embedded or secure processors, IBM PowerPC, Intel’s Core, Itanium, Xeon, Celeron or other line of processors, etc. The processor 502 may be implemented using mainframe, distributed processor, multi-core, parallel, grid, or other architectures. Some embodiments may utilize embedded technologies like application-specific integrated circuits (ASICs), digital signal processors (DSPs), Field Programmable Gate Arrays (FPGAs), etc.
Processor 502 may be disposed in communication with one or more input/output (I/O) devices via I/O interface 503. The I/O interface 503 may employ communication protocols/methods such as, without limitation, audio, analog, digital, monoaural, RCA, stereo, IEEE-1394, serial bus, universal serial bus (USB), infrared, PS/2, BNC, coaxial, component, composite, digital visual interface (DVI), high-definition multimedia interface (HDMI), RF antennas, S-Video, VGA, IEEE 802.n /b/g/n/x, Bluetooth, cellular (e.g., code-division multiple access (CDMA), high-speed packet access (HSPA+), global system for mobile communications (GSM), long-term evolution (LTE), WiMax, or the like), etc.
Using the I/O interface 503, the computer system 501 may communicate with one or more I/O devices. For example, the input device 504 may be an antenna, keyboard, mouse, joystick, (infrared) remote control, camera, card reader, fax machine, dongle, biometric reader, microphone, touch screen, touchpad, trackball, sensor (e.g., accelerometer, light sensor, GPS, gyroscope, proximity sensor, or the like), stylus, scanner, storage device, transceiver, video device/source, visors, etc. Output device 505 may be a printer, fax machine, video display (e.g., cathode ray tube (CRT), liquid crystal display (LCD), light-emitting diode (LED), plasma, or the like), audio speaker, etc. In some embodiments, a transceiver 506 may be disposed in connection with the processor 502. The transceiver may facilitate various types of wireless transmission or reception. For example, the transceiver may include an antenna operatively connected to a transceiver chip (e.g., Texas Instruments WiLink WL1283, Broadcom BCM4750IUB8, Infineon Technologies X-Gold 618-PMB9800, or the like), providing IEEE 802.11a/b/g/n, Bluetooth, FM, global positioning system (GPS), 2G/3G HSDPA/HSUPA communications, etc.
In some embodiments, the processor 502 may be disposed in communication with a communication network 508 via a network interface 507. The network interface 507 may communicate with the communication network 508. The network interface 507 may employ connection protocols including, without limitation, direct connect, Ethernet (e.g., twisted pair 10/40/400 Base T), transmission control protocol/internet protocol (TCP/IP), token ring, IEEE 802.11a/b/g/n/x, etc. The communication network 508 may include, without limitation, a direct interconnection, local area network (LAN), wide area network (WAN), wireless network (e.g., using Wireless Application Protocol), the Internet, etc. Using the network interface 507 and the communication network 508, the computer system 501 may communicate with devices 509, 510, and 511. These devices may include, without limitation, personal computer(s), server(s), fax machines, printers, scanners, various mobile devices such as cellular telephones, smartphones (e.g., Apple iPhone, Blackberry, Android-based phones, etc.), tablet computers, eBook readers (Amazon Kindle, Nook, etc.), laptop computers, notebooks, gaming consoles (Microsoft Xbox, Nintendo DS, Sony PlayStation, etc.), or the like. In some embodiments, the computer system 501 may itself embody one or more of these devices.
In some embodiments, the processor 502 may be disposed in communication with one or more memory devices (e.g., RAM 513, ROM 5Error! Reference source not found.14, etc.) via a storage interface 512. The storage interface may connect to memory devices including, without limitation, memory drives, removable disc drives, etc., employing connection protocols such as serial advanced technology attachment (SATA), integrated drive electronics (IDE), IEEE-1394, universal serial bus (USB), fiber channel, small computer systems interface (SCSI), etc. The memory drives may further include a drum, magnetic disc drive, magneto-optical drive, optical drive, redundant array of independent discs (RAID), solid-state memory devices, solid-state drives, etc.
The memory 515 may store a collection of program or database components, including, without limitation, an operating system 5Error! Reference source not found.16, user interface application 5Error! Reference source not found.17, web browser 518, mail server 519, mail client 520, user/application data 521 (e.g., any data variables or data records discussed in this disclosure), etc. The operating system 516 may facilitate resource management and operation of the computer system 501. Examples of operating systems include, without limitation, Apple Macintosh OS X, UNIX, Unix-like system distributions (e.g., Berkeley Software Distribution (BSD), FreeBSD, NetBSD, OpenBSD, etc.), Linux distributions (e.g., Red Hat, Ubuntu, Kubuntu, etc.), IBM OS/2, Microsoft Windows (XP, Vista/7/8, etc.), Apple iOS, Google Android, Blackberry OS, or the like. User interface 517 may facilitate display, execution, interaction, manipulation, or operation of program components through textual or graphical facilities. For example, user interfaces may provide computer interaction interface elements on a display system operatively connected to the computer system 501, such as cursors, icons, check boxes, menus, scrollers, windows, widgets, etc. Graphical user interfaces (GUIs) may be employed, including, without limitation, Apple Macintosh operating systems’ Aqua, IBM OS/2, Microsoft Windows (e.g., Aero, Metro, etc.), Unix X-Windows, web interface libraries (e.g., ActiveX, Java, Javascript, AJAX, HTML, Adobe Flash, etc.), or the like.
In some embodiments, the computer system 501 may implement a web browser 518 stored program component. The web browser may be a hypertext viewing application, such as Microsoft Internet Explorer, Google Chrome, Mozilla Firefox, Apple Safari, etc. Secure web browsing may be provided using HTTPS (secure hypertext transport protocol), secure sockets layer (SSL), Transport Layer Security (TLS), etc. Web browsers may utilize facilities such as AJAX, DHTML, Adobe Flash, JavaScript, Java and application programming interfaces (APIs), etc. In some embodiments, the computer system 501 may implement a mail server 519 stored program component. The mail server may be an Internet mail server such as Microsoft Exchange, or the like. The mail server may utilize facilities such as ASP, ActiveX, ANSI C++/C#, Microsoft .NET, CGI scripts, Java, JavaScript, PERL, PHP, Python, WebObjects, etc. The mail server may utilize communication protocols such as internet message access protocol (IMAP), messaging application programming interface (MAPI), Microsoft Exchange, post office protocol (POP), simple mail transfer protocol (SMTP), or the like. In some embodiments, the computer system 501 may implement a mail client 520 stored program component. The mail client may be a mail viewing application, such as Apple Mail, Microsoft Entourage, Microsoft Outlook, Mozilla Thunderbird, etc.
In some embodiments, computer system 501 may store user/application data 521, such as the data, variables, records, etc. as described in this disclosure. Such databases may be implemented as fault-tolerant, relational, scalable, secure databases such as Oracle or Sybase. Alternatively, such databases may be implemented using standardized data structures, such as an array, hash, linked list, struct, structured text file (e.g., XML), table, or as object-oriented databases (e.g., using ObjectStore, Poet, Zope, etc.). Such databases may be consolidated or distributed, sometimes among the various computer systems discussed above in this disclosure. It is to be understood that the structure and operation of the any computer or database component may be combined, consolidated, or distributed in any working combination.
As described above, the modules 205, amongst other things, include routines, programs, objects, components, and data structures, which perform particular tasks or implement particular abstract data types. The modules 205 may also be implemented as, signal processor(s), state machine(s), logic circuitries, and/or any other device or component that manipulate signals based on operational instructions. Further, the modules 205 can be implemented by one or more hardware components, by computer-readable instructions executed by a processing unit, or by a combination thereof.
The illustrated steps are set out to explain the exemplary embodiments shown, and it should be anticipated that ongoing technological development will change the manner in which particular functions are performed. These examples are presented herein for purposes of illustration, and not limitation. Further, the boundaries of the functional building blocks have been arbitrarily defined herein for the convenience of the description. Alternative boundaries can be defined so long as the specified functions and relationships thereof are appropriately performed. Alternatives (including equivalents, extensions, variations, deviations, etc., of those described herein) will be apparent to persons skilled in the relevant art(s) based on the teachings contained herein. Such alternatives fall within the scope and spirit of the disclosed embodiments. Also, the words "comprising," "having," "containing," and "including," and other similar forms are intended to be equivalent in meaning and be open ended in that an item or items following any one of these words is not meant to be an exhaustive listing of such item or items, or meant to be limited to only the listed item or items. It must also be noted that as used herein and in the appended claims, the singular forms “a,” “an,” and “the” include plural references unless the context clearly dictates otherwise.
Furthermore, one or more computer-readable storage media may be utilized in implementing embodiments consistent with the present disclosure. A computer-readable storage medium refers to any type of physical memory on which information or data readable by a processor may be stored. Thus, a computer-readable storage medium may store instructions for execution by one or more processors, including instructions for causing the processor(s) to perform steps or stages consistent with the embodiments described herein. The term “computer-readable medium” should be understood to include tangible items and exclude carrier waves and transient signals, i.e., are non-transitory. Examples include random access memory (RAM), read-only memory (ROM), volatile memory, nonvolatile memory, hard drives, CD ROMs, DVDs, flash drives, disks, and any other known physical storage media.
It is intended that the disclosure and examples be considered as exemplary only, with a true scope and spirit of disclosed embodiments being indicated by the following claims.
Referral Numerals:
Reference Number Description
100 Environment
101 Data source
103 Input data
104 Data to be analyzed
105 Configuration information
109 Data analytics process recommendation system
110 Goal information
111 Industry specific information
113 Processor
115 Memory
117 Recommendation module
201 I/O Interface
205 Modules
209 Similarity matrix
219 Knowledge repository
211 Data analytics operators
213 Data analytics processes
215 Other data
221 Knowledge Repository Interface Module
223 Other modules
| # | Name | Date |
|---|---|---|
| 1 | Form 5 [27-11-2015(online)].pdf | 2015-11-27 |
| 2 | Form 3 [27-11-2015(online)].pdf | 2015-11-27 |
| 3 | Drawing [27-11-2015(online)].pdf | 2015-11-27 |
| 4 | Description(Complete) [27-11-2015(online)].pdf | 2015-11-27 |
| 5 | 6377-CHE-2015-Form 1-101215.pdf | 2016-06-07 |
| 6 | 6377-CHE-2015-Correspondence-F1-101215.pdf | 2016-06-07 |
| 7 | 6377-CHE-2015-Power of Attorney-060116.pdf | 2016-06-17 |
| 8 | 6377-CHE-2015-Correspondence-PA-060116.pdf | 2016-06-17 |
| 9 | abstract-6377-CHE-2015-jpeg.jpg | 2016-09-16 |
| 10 | 6377-CHE-2015-FER.pdf | 2020-02-13 |
| 11 | 6377-CHE-2015-OTHERS [27-07-2020(online)].pdf | 2020-07-27 |
| 12 | 6377-CHE-2015-FER_SER_REPLY [27-07-2020(online)].pdf | 2020-07-27 |
| 13 | 6377-CHE-2015-DRAWING [27-07-2020(online)].pdf | 2020-07-27 |
| 14 | 6377-CHE-2015-CLAIMS [27-07-2020(online)].pdf | 2020-07-27 |
| 15 | 6377-CHE-2015-PatentCertificate27-07-2022.pdf | 2022-07-27 |
| 16 | 6377-CHE-2015-IntimationOfGrant27-07-2022.pdf | 2022-07-27 |
| 1 | SearchStrategy_amended_stageAE_09-02-2021.pdf |
| 2 | SearchStrategy_15-01-2020.pdf |
| 3 | CopyofSearchStartegy_31-12-2019.pdf |