Abstract: METHOD AND SYSTEM FOR DETECTING THE ROOT CAUSE OF ANOMALIES IN THE DATA SET ABSTRACT The present invention relates to a method of ranking variables for detecting the root cause of anomalies in the data set (101). The method comprises obtaining a plurality of variables and a target variable (102) in the data set (101). Further, one or more anomalies in the variation of the target variable (102) with respect to the plurality of variables is identified based on outliers present in the variation. Furthermore, one or more variables (805) from the plurality of variables causing the detected one or more anomalies is identified based on one or more statistical analysis performed on the plurality of variables. Finally, a rank is assigned to the variation of the target variable (102) with respect to each of the identified one or more variables (805), wherein each of the variation is displayed based on the ranking for detecting the root cause of anomalies in the data set (101). FIGURE 3
1. A method of ranking variables for detecting the root cause of anomalies in the data set (101), the method comprising: obtaining, by a ranking system (200), a plurality of variables from the data set (101) and a target variable (102) in the data set (101); identifying, by the ranking system (200), a variation of the target variable (102) with respect to the plurality of variables to detect one or more anomalies in the variation of the target variable (102) with respect to the plurality of variables based on outliers present in the variation; identifying, by the ranking system (200), one or more variables from the plurality of variables causing the detected one or more anomalies based on one or more statistical analysis performed on the plurality of variables; and ranking, by the ranking system (200), the variation of the target variable (102) with respect to each of the identified one or more variables (805), wherein each of the variation is displayed based on the ranking for detecting the root cause of anomalies in the data set (101).
2. The method as claimed in claim 1, wherein the plurality of variables is aggregated into one or more data types based on meta data (103) of the plurality of variables.
3. The method as claimed in claim 2, wherein the one or more data types comprises at least one of a numerical data type, a categorical data type, a temporal data type and a spatial data type.
4. The method as claimed in claim 2, wherein the meta data (103) comprises at least one of set of rules to identify a data type of the plurality of variables, one or more graphical representations (111) for displaying the variation of the target variable (102) with respect to each of the identified one or more variables (805) and the one or more data types supported for a graphical representation (111), and one or more data analysis method to analyze the variation of the target variable (102) with respect to the plurality of variables and the one or more data types supported for corresponding data analysis method.
5. The method as claimed in claim 4, wherein the one or more data analysis method comprises at least one of a clustering analysis (601), quartile analysis (604) and a time series analysis (606).
6. The method as claimed in claim 1, wherein detecting the one or more anomalies comprises: applying on each variation at least one data analysis method from one or more data analysis method based on a data type of the plurality of variables; and identifying the outliers present in the plurality of variables based on results of the applied at least one data analysis method.
7. The method as claimed in claim 1, wherein the outliers indicative of at least one data point deviating from the plurality of data point in each variation, are used to generate an aggregated outlier data set (106) and an aggregated normal data set (205) of the plurality of variables without the outliers.
8. The method as claimed in claim 1, wherein identifying the one or more variables (805) comprises: applying at least one statistical analysis from the one or more statistical analysis on an aggregated outlier data set (106) and an aggregated normal data set (205) of the plurality of variables without the outliers; and identifying the one or more variables (805) by computing a difference between result of the statistical analysis on the aggregated outlier data set (106) and the aggregated normal data set (205) of the plurality of variables without the outliers.
9. The method as claimed in claim 1, wherein the one or more statistical analysis comprises at least one of a correlation analysis (801), a regression and a learning algorithm.
10. The method as claimed in claim 1, wherein ranking the variation of the target variable (102) with respect to each of the identified one or more variables (805) is based on a computed difference between result of an statistical analysis on an aggregated outlier data set (106) and an aggregated normal data set (205) of the plurality of variables without the outliers, wherein a graphical representation (111) is selected from the one or more graphical representations (111) for displaying the variation.
11. A ranking system (200) comprising: a processor (203); and a memory (202) communicatively coupled to the processor (203), wherein the memory (202) stores the processor (203) instructions, which, on execution, causes the processor (203) to: obtain, a plurality of variables from the data set (101) and a target variable (102) in the data set (101); identify, a variation of the target variable (102) with respect to the plurality of variables to detect one or more anomalies in the variation of the target variable (102) with respect to the plurality of variables based on outliers present in the variation; identify, one or more variables (805) from the plurality of variables causing the detected one or more anomalies based on one or more statistical analysis performed on the plurality of variables; and rank, the variation of the target variable (102) with respect to each of the identified one or more variables (805), wherein each of the variation is displayed based on the ranking for detecting the root cause of anomalies in the data set (101).
12. The ranking system (200) as claimed in claim 11, wherein the processor (203) is configured to aggregate the obtained plurality of variables into one or more data types based on meta data (103) of the plurality of variables.
13. The ranking system (200) as claimed in claim 12, wherein the one or more data types comprises at least one of a numerical data type, a categorical data type, a temporal data type and a spatial data type.
14. The ranking system (200) as claimed in claim 12, wherein the meta data (103) comprises at least one of set of rules to identify a data type of the plurality of variables, one or more graphical representations (111) for displaying the variation of the target variable (102) with respect to each of the identified one or more variables (805) and the one or more data types supported for a graphical representation (111), and one or more data analysis method to analyze the variation of the target variable (102) with respect to the plurality of variables and the one or more data types supported for the corresponding data analysis method.
15. The ranking system (200) as claimed in claim 14, wherein the one or more data analysis method comprises at least one of a clustering analysis (601), quartile analysis (604) and a time series analysis (606).
16. The ranking system (200) as claimed in claim 11, wherein the processor (203) is configured to detect the one or more anomalies comprises: applying on each variation at least one data analysis method from one or more data analysis method based on a data type of the plurality of variables; and identifying the outliers present in the plurality of variables based on results of the applied at least one data analysis method.
17. The ranking system (200) as claimed in claim 11, wherein the processor (203) is configured to generate an aggregated outlier data set (106) and an aggregated normal data set (205) of the plurality of variables without the outliers based on identified outliers in the plurality of variables, further wherein the outliers are indicative of at least one data point deviating from the plurality of data point in each variation.
18. The ranking system (200) as claimed in claim 11, wherein the processor (203) is configured to identify the one or more variables (805) comprises: applying at least one statistical analysis from the one or more statistical analysis on an aggregated outlier data set (106) and an aggregated normal data set (205) of the plurality of variables without the outliers; and identifying the one or more variables (805) by computing a difference between result of the statistical analysis on the aggregated outlier data set (106) and the aggregated normal data set (205) of the plurality of variables without the outliers.
19. The ranking system (200) as claimed in claim 11, wherein the one or more statistical analysis comprises at least one of a correlation analysis (801), a regression and a learning algorithm.
20. The ranking system (200) as claimed in claim 11, wherein the processor (203) is configured to rank the variation of the target variable (102) with respect to each of the identified one or more variables (805) is based on a computed difference between result of an statistical analysis on an aggregated outlier data set (106) and an aggregated normal data set (205) of the plurality of variables without the outliers, wherein a graphical representation (111) is selected from the one or more graphical representation (111) for displaying the variation. Dated this on July 18th, 2019 MADHUSUDAN S.T. IN/PA-1297 K & S PARTNERS AGENT FOR THE APPLICANT , Description:FORM 2 THE PATENTS ACT 1970 [39 OF 1970] & THE PATENTS RULES, 2003 COMPLETE SPECIFICATION [See section 10 and Rule 13] TITLE: “METHOD AND SYSTEM FOR DETECTING THE ROOT CAUSE OF ANOMALIES IN THE DATA SET” Name and Address of the Applicant: Hitachi, Ltd., 6-6, Marunouchi 1-chome, Chiyoda-ku, Tokyo 100-8280, Japan Nationality: Japan The following specification particularly describes the invention and the manner in which it is to be performed. TECHNICAL FIELD The present disclosure relates to the field of data analysis. Particularly, but not exclusively, the present disclosure relates to a method and system for detecting root cause of anomalies in the data set. BACKGROUND Generally, it is a challenging task to identify an anomaly in a large-scale network. The data set collected from the distributed systems in a large-scale network is very huge and may contain noisy data points, so manual checking of the data set to detect one or more anomalies is time consuming and error prone. Therefore, an automated system is used to detect the one or more anomalies in the data set by ranking the graphical representation of the variation of the target variable with respect to the plurality of variables in the dataset. The identified one or more anomalies must be rectified. Hence, we need the root cause of the one or more anomalies. The existing graphical representation systems automatically generate a graphical representation based on pre-assigned mapping rules to represent one or more data variables from a data set. The pre-assigned mapping rules are based on matching data properties with recommended graphical representation charts. In addition, existing graphical representation systems also continuously rank the graphical representations based on a historical user preference and recommend them to the users. Further, the ranking is based on properties of a data field, and many users use exploratory data analysis technique to identify interesting insights from the data relevant to the use case application. So, choosing an appropriate graphical representation when large amounts of data fields are available and finding right graphical representation chart is time consuming task. An issue with the existing techniques is the time consumed to generate the rankings and the usage of adulterated historical preferences. This leads to in-accurate ranking and ultimately resulting in sub-optimal graphical representation chart recommendations. An issue with the existing techniques is redundant recommendations due to user's past preference bias and for new data sets the rankings are not accurate due to the usage of historical data regarding the user past preferences and data from the user profile to assign the rankings. The information disclosed in this background of the disclosure section is only for enhancement of understanding of the general background of the invention and should not be taken as an acknowledgement or any form of suggestion that this information forms the prior art already known to a person skilled in the art. SUMMARY One or more shortcomings of the prior art are overcome, and additional advantages are provided through the provision of method of the present disclosure. Additional features and advantages are realized through the techniques of the present disclosure. Other embodiments and aspects of the disclosure are described in detail herein and are considered a part of the claimed disclosure. Disclosed herein is a method of ranking variables for detecting the root cause of anomalies in the data set. The method includes obtaining a plurality of variables from the data set and a target variable in the data set. Further, the method includes identifying a variation of the target variable with respect to the plurality of variables to detect one or more anomalies in the variation of the target variable with respect to the plurality of variables based on outliers present in the variation. Furthermore, the method includes identifying one or more variables from the plurality of variables causing the detected one or more anomalies based on one or more statistical analysis performed on the plurality of variables. Finally, the method includes ranking the variation of the target variable with respect to each of the identified one or more variables, wherein each of the variation is displayed based on the ranking for detecting the root cause of anomalies in the data set. Further, the present disclosure discloses a ranking system including a processor and a memory communicatively coupled to the processor, wherein the memory stores the processor instructions, which, on execution, causes the processor to obtain, a plurality of variables from the data set and a target variable in the data set. Further, the processor is configured to identify a variation of the target variable with respect to the plurality of variables to detect one or more anomalies in the variation of the target variable with respect to the plurality of variables based on outliers present in the variation. Furthermore, the processor is configured to identify one or more variables from the plurality of variables causing the detected one or more anomalies based on one or more statistical analysis performed on the plurality of variables. Finally, the processor is configured to rank the variation of the target variable with respect to each of the identified one or more variables, wherein each of the variation is displayed based on the ranking for detecting the root cause of anomalies in the data set. The foregoing summary is illustrative only and is not intended to be in any way limiting. In addition to the illustrative aspects, embodiments, and features described above, further aspects, embodiments, and features may become apparent by reference to the drawings and the following detailed description. BRIEF DESCRIPTION OF THE ACCOMPANYING DRAWINGS The novel features and characteristic of the disclosure are set forth in the appended claims. The disclosure itself, however, as well as a preferred mode of use, further objectives and advantages thereof, may best be understood by reference to the following detailed description of an illustrative embodiment when read in conjunction with the accompanying drawings. The accompanying drawings, which are incorporated in and constitute a part of this disclosure, illustrate exemplary embodiments and, together with the description, serve to explain the disclosed principles. In the figures, the left-most digit(s) of a reference number identifies the figure in which the reference number first appears. One or more embodiments are now described, by way of example only, with reference to the accompanying figures wherein like reference numerals represent like elements and in which: FIGURE 1 shows an exemplary system for detecting the root cause of anomalies in the data set, in accordance with some embodiments of the present disclosure; FIGURE 2 shows a detailed block diagram of a ranking system, in accordance with some embodiments of the present disclosure; FIGURE 3 shows a flowchart illustrating method for detecting the root cause of anomalies in the data set, in accordance with some embodiment of the present disclosure; FIGURE 4A shows an exemplary data set, in accordance with some embodiments of the present disclosure; FIGURE 4B shows an exemplary aggregated data set based on the categorical data type of the plurality of variables, in accordance with some embodiments of the present disclosure; FIGURE 4C shows an exemplary aggregated data set based on the temporal data type of the plurality of variables, in accordance with some embodiments of the present disclosure; FIGURE 5A shows an exemplary meta data for identifying the data type of plurality of variables in a data set, in accordance with some embodiments of the present disclosure; FIGURE 5B shows an exemplary meta data for selecting a graphical representation based on the data type of plurality of variables, in accordance with some embodiments of the present disclosure; FIGURE 5C shows an exemplary meta data for selecting a data analysis method based on the data type of plurality of variables, in accordance with some embodiments of the present disclosure; FIGURE 6 shows an exemplary detection of outliers present in the variation, in accordance with some embodiments of the present disclosure; FIGURE 7A shows an exemplary clustering analysis, in accordance with some embodiments of the present disclosure; FIGURE 7B shows an exemplary quartile analysis, in accordance with some embodiments of the present disclosure; FIGURE 7C shows an exemplary time series analysis, in accordance with some embodiments of the present disclosure; FIGURE 8 shows an exemplary identification of one or more variables using a correlation analysis, in accordance with some embodiments of the present disclosure; and FIGURE 9 shows a general-purpose computer system to rank variables for detecting the root cause of anomalies in the data set, in accordance with embodiments of the present disclosure. It should be appreciated by those skilled in the art that any block diagrams herein represent conceptual views of illustrative systems embodying the principles of the present subject matter. Similarly, it may be appreciated that any flow charts, flow diagrams, state transition diagrams, pseudo code, and the like represent various processes which may be substantially represented in computer readable medium and executed by a computer or processor, whether or not such computer or processor is explicitly shown. DETAILED DESCRIPTION In the present document, the word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any embodiment or implementation of the present subject matter described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments. While the disclosure is susceptible to various modifications and alternative forms, specific embodiment thereof has been shown by way of example in the drawings and may be described in detail below. It should be understood, however that it is not intended to limit the disclosure to the particular forms disclosed, but on the contrary, the disclosure is to cover all modifications, equivalents, and alternative falling within the scope of the disclosure. The terms “comprises”, “includes” “comprising”, “including” or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a setup, device or method that comprises a list of components or steps does not include only those components or steps but may include other components or steps not expressly listed or inherent to such setup or device or method. In other words, one or more elements in a system or apparatus proceeded by “comprises… a” or “includes…a” does not, without more constraints, preclude the existence of other elements or additional elements in the system or apparatus. The present disclosure describes a method of ranking variables for detecting the root cause of anomalies in the data set. The method includes obtaining a plurality of variables from the data set and a target variable in the data set. Further, the method includes identifying a variation of the target variable with respect to the plurality of variables to detect one or more anomalies in the variation of the target variable with respect to the plurality of variables based on outliers present in the variation. Furthermore, the method includes identifying one or more variables from the plurality of variables causing the detected one or more anomalies based on one or more statistical analysis performed on the plurality of variables. Finally, the method includes ranking the variation of the target variable with respect to each of the identified one or more variables, wherein each of the variation is displayed based on the ranking for detecting the root cause of anomalies in the data set. In the following detailed description of the embodiments of the disclosure, reference is made to the accompanying drawings that form a part hereof, and in which are shown by way of illustration specific embodiments in which the disclosure may be practiced. These embodiments are described in sufficient detail to enable those skilled in the art to practice the disclosure, and it is to be understood that other embodiments may be utilized and that changes may be made without departing from the scope of the present disclosure. The following description is, therefore, not to be taken in a limiting sense. FIGURE 1 shows an exemplary system for detecting the root cause of anomalies in the data set (101), in accordance with some embodiments of the present disclosure. In an embodiment, a user may provide a data set (101) including a plurality of rows and columns. In another embodiment, the data set may be retrieved from the memory (202) of the ranking system (109). Further, the user selects a target variable (102) from the data set (101) for analysis. The data set (101) is a collection of related and discrete items of data that may be accessed individually or in combination or managed as a whole entity. The columns of the data set (101) constitute plurality of variables and every row of the data set (101) constitutes a value taken by the plurality of variables. The target variable (102) denoted as a dependent variable, is a column of the data set (101) under observation. The plurality of variables in the data set (101) excluding the target variable (102) is denoted as independent variables. The plurality of variables excluding the target variable (102) in the data set (101) is aggregated based on the data type (for example numerical, categorical, temporal, spatial) of the plurality of variables. The meta data (103) of the data set (101) includes a set rules to identify a data type of the plurality of variables. Further, a variation of the target variable (102) with respect to the aggregated plurality of variables corresponding to a data type is identified. Based on the variation corresponding to the data type, one or more anomalies are detected based on outliers present in the variation by the outlier detection unit (105). The outliers are indicative of at least one data point deviating from the plurality of data point in each variation. The detected outliers with respect to the aggregated plurality of variables is stored in the outliers data set (106). Furthermore, one or more variables (805) from the plurality of variables causing the detected one or more anomalies is identified based on one or more statistical analysis performed on the plurality of variables by the variable identification unit (107). The identified one or more variables (805) from the plurality of variables casing the detected one or more anomalies are stored in the identified variables data set (108). The one or more statistical analysis suitable for the plurality of variables is obtained from the meta data (103) associated with the data set (101). Finally, the variation of the target variable (102) with respect to each of the identified one or more variables (805) is ranked by the ranking unit (109). For example, if V1, V2, and V3 are the identified one or more variables, a rank 1 is assigned to the variation of the target variable with the identified one or more variables V1, a rank 2 is assigned to the variation of the target variable with the identified one or more variables V2, and a rank 3 is assigned to the variation of the target variable with the identified one or more variables V3 as shown in Figure 1. For the ranked variation (110) a suitable graphical representation (111) is selected from the meta data (103) associated with the data set (101) based on the data type of the one or more variables (805) in the variation. For example, let the data type of the one or more variables V1 be numerical and the one or more variables V2 be spatial. Therefore, a line chart is selected to display the variation of the one or more variables V1 with respect to the target variable and a Geo chart is selected to display the variation of the one or more variables V2 with respect to the target variable as shown in Figure 1. In one embodiment, the order of providing the graphical representation (111) may be based on the rank associated with the one or more variables. For example, the line chart is displayed at top indicating that the corresponding variable V3 is ranked 1 and likewise, the Geo chart may be displayed below the line chart indicating that the corresponding variable V2 is ranked 2. The ranked variations along with the corresponding graphical representations (111) are displayed to the user for detecting the root cause of anomalies in the data set (101). FIGURE 2 shows a detailed block diagram of a ranking system (200), in accordance with some embodiments of the present disclosure. The ranking system (200) may include Central Processing Unit (“CPU” or “processor”) (203) and a memory (202) storing instructions executable by the processor (203). The processor (203) may include at least one data processor for executing program components for executing user or system-generated requests. The memory (202) may be communicatively coupled to the processor (203). The ranking system (200) further includes an Input/ Output (I/O) interface (201). The I/O interface (201) may be coupled with the processor (203) through which an input signal or/and an output signal may be communicated. In one embodiment, the ranking system (200) may receive the data set (101) and the target variable (102) through the I/O interface (201). In some implementations, the ranking system (200) may include data (204) and modules (207). As an example, the data (204) and modules (207) may be stored in the memory (202) configured in the ranking system (200) as shown in the FIGURE 2. In one embodiment, the data (204) may include, for example, a data set (101), a meta data (103), a normal data set (205), an outlier data set (106), identified variables (108) and other data (206). In the illustrated FIGURE 2, modules (207) are described herein in detail. In an embodiment, the data set (101) is a collection of related and discrete items of data that may be accessed individually or in combination or managed as a whole entity. The data set (101) includes arrangement of related data in a plurality of rows and columns. The columns of the data set (101) constitute the plurality of variables and every row of the data set (101) constitutes a value taken by the plurality of variables. FIGURE 4A shows an exemplary data set (101) arranged in the form of table including a plurality of rows and columns. The columns (for example Date Time, Merchant ID, Customer ID, Transaction Amount, Location, Merchant Category Code (MCC) constitute the plurality of variables. The rows in table of FIGURE 4A constitutes the value taken by the corresponding variable of the column. For example, consider row 1 in the table of FIGURE 4A wherein “M101” is the value taken by the variable “Merchant ID”, “Bangalore” is the value taken by the variable “Location” and the like. In an embodiment, the metadata is data that provides information about the data set (101). The distinct types of metadata are descriptive metadata, structural metadata, administrative metadata, reference metadata and statistical metadata. The descriptive metadata describes a resource for purposes such as discovery and identification. It can include elements such as title, abstract, author, and keywords. The structural metadata is metadata about containers of data and indicates how compound objects are put together, for example, how pages are ordered to form chapters. It describes the types, versions, relationships and other characteristics of digital materials. The administrative metadata provides information to help manage a resource, such as when and how it was created, file type and other technical information, and who can access it. The reference metadata describes the contents and quality of statistical data. The statistical metadata may also describe processes that collect, process, or produce statistical data; such metadata are also called process data. Further, the meta data (103) includes at least one of set of rules to identify a data type of the plurality of variables, one or more graphical representations (111) for displaying the variation of the target variable (102) with respect to each of the identified one or more variables (805) and the one or more data types supported for a graphical representation (111), and one or more data analysis method to analyze the variation of the target variable (102) with respect to the plurality of variables and the one or more data types supported for corresponding data analysis method. As shown in table of FIGURE 5A the meta data (103) includes one or more rules for identifying the data type of the plurality of variables in the data set (101). For example, consider the row 1 in the table of FIGURE 5A, the one or more variables (805) in the data set (101) having a value of an integer type or a double type is identified as numerical data type and the like. As shown in table of FIGURE 5B the meta data (103) includes one or more graphical representations (111) for displaying the variation of the target variable (102) with respect to each of the identified one or more variables (805). For example, consider column 1 in table of Figure 5B, for the data types of numerical and categorical representing the X-axis and numerical data type representing the Y-axis a “Bar chart” is used and the like. As shown in table of FIGURE 5C the meta data (103) includes one or more data analysis method to analyze the variation of the target variable (102) with respect to the plurality of variables and the one or more data types supported for corresponding data analysis method. For example, consider the column 1 in table of the Figure 5C, for a variation of the target variable (102) of the data type temporal or numerical with respect to the plurality of the variables of the data type temporal or numerical a “time series analysis (606)” method is used to analyze the variation and a line chart or a bar chart is used to represent the variation and the like. In an embodiment, the normal data set (205) includes a subset of rows and columns of the data set (101) excluding the rows identified as outliers in the variation of the target variable (102) of a corresponding data type with respect to the plurality of variables. Further, the normal data set (205) may include an aggregate the normal data set (205) obtained by the variation of the target variable (102) corresponding to the one or more data types. In an embodiment, the outlier data set (106) includes a subset of rows and columns of the data set (101) identified as outliers in the variation of the target variable (102) of a corresponding data type with respect to the plurality of variables. The outliers in the variation are identified based on the one or more data analysis method including at least one of a clustering analysis (601), quartile analysis (604) and a time series analysis (606). Further, the outliers are indicative of at least one data point deviating from the plurality of data point in each variation. Furthermore, the outlier data set (106) may include an aggregate the outlier data set (106) obtained by the variation of the target variable (102) corresponding to the one or more data types. In an embodiment, the identified variables (108) include one or more variables (805) from the plurality of variables causing the detected one or more anomalies based on one or more statistical analysis performed on the plurality of variables. The one or more statistical analysis comprises at least one of a correlation analysis (801), a regression and a learning algorithm. For example, consider the table of FIGURE 4A, wherein the “Location” variable may be identified as the one or more variable from the plurality of variables causing the detected one or more anomalies (for example high transaction amount in one month). In an embodiment, the other data (206) may include data regarding the one or more data analysis method including at least one of a clustering analysis (601), quartile analysis (604) and a time series analysis (606) and the one or more statistical analysis including at least one of a correlation analysis (801), a regression and a learning algorithm. Further, the other data (206) may include data regarding the distance measurement (603) between two data points in a variation, computing and comparing the inter quartile distance (605) between the upper and lower quartiles, and computing and comparing the expected bound value (607) between the upper and lower bounds. In some embodiments, data (204) may be stored in the memory (202) in form of various data structures. Additionally, the data (204) may be organized using data models, such as relational or hierarchical data models. The other data (206) may store data, including temporary data and temporary files, generated by the modules (207) for performing the various functions of the ranking system (200). In some embodiments, the data (204) stored in the memory (202) may be processed by the modules (207) of the ranking system (200). The modules (207) may be stored within the memory (202). In an example, the modules (207) communicatively coupled to the processor (203) configured in the ranking system (200), may also be present outside the memory (202) as shown in FIGURE 2 and implemented as hardware. As used herein, the term modules (207) may refer to an application specific integrated circuit (ASIC), a FPGA (Field Programmable Gate Array), an electronic circuit, a processor (203) (shared, dedicated, or group) and memory (202) that execute one or more software or firmware programs, a combinational logic circuit, and/or other suitable components that provide the described functionality. In some other embodiments, the modules (207) may be implemented using at least one of ASICs and FPGAs. In one implementation, the modules (207) may include, for example, an outlier detection unit (105), a variable identification unit (107), an output module (209), an input module (208), a ranking unit (109) and other module (210). It may be appreciated that such aforementioned modules (207) may be represented as a single module or a combination of different modules (207). In an embodiment, the outlier detection unit (105) is used to identify the outliers present in the plurality of variables based on results of the applied at least one data analysis method. The outliers are indicative of at least one data point deviating from the plurality of data point in each variation. The one or more data analysis method includes at least one of a clustering analysis (601), quartile analysis (604) and a time series analysis (606). For example, the detected outliers are indicated in the FIGURE 7A, FIGURE 7B, and FIGURE 7C. In an embodiment, the variable identification unit (107) is used to identify the one or more variables (805) from the plurality of variables causing the detected one or more anomalies based on the one or more statistical analysis performed on the plurality of variables. Initially the variable identification unit (107) applies at least one statistical analysis from the one or more statistical analysis on an aggregated outlier data set (106) and an aggregated normal data set (205) of the plurality of variables without the outliers. Further, the variable identification unit (107) identifies the one or more variables (805) by computing a difference between result of the statistical analysis on the aggregated outlier data set (106) and the aggregated normal data set (205) of the plurality of variables without the outliers. The one or more statistical analysis includes at least one of a correlation analysis (801), a regression and a learning algorithm. For example, consider the table of FIGURE 4A, wherein the “Location” variable may be identified as the one or more variable from the plurality of variables causing the detected one or more anomalies (for example high transaction amount in one month). In an embodiment, the ranking unit (109) is responsible for assigning a rank to the variation of the target variable (102) with respect to each of the identified one or more variables (805). The rank is assigned based on a computed difference between result of a statistical analysis on an aggregated outlier data set (106) and an aggregated normal data set (205) of the plurality of variables without the outliers. Further, a graphical representation (111) is selected from the one or more graphical representations (111) and the variation is displayed to the user using the output module (209). In an embodiment, the output module (209) is responsible for displaying a graphical representation (111) of each of the variation of the target variable (102) with respect to each of the identified one or more variables (805) based on the rank assigned by the ranking unit (109). Further, the ranked graphical representations (111) displayed to the user enables the user to detect the root cause of anomalies in the data set (101). In an embodiment, the input module (208) is responsible for obtaining a plurality of variables from the data set (101) and a target variable (102) in the data set (101) for detecting the root cause of anomalies in the data set (101). In an embodiment, the other module (210) is responsible for measuring the distance between two data points in a variation, measuring the distance between a data point and a group of data points, computing and comparing the inter quartile distance (605) between the upper and lower quartiles, and computing and comparing the expected bound value (607) between the upper and lower bounds. FIGURE 3 shows a flowchart illustrating method of detecting the root cause of anomalies in the data set (101), in accordance with some embodiment of the present disclosure. The order in which the method (300) may be described is not intended to be construed as a limitation, and any number of the described method blocks may be combined in any order to implement the method. Additionally, individual blocks may be deleted from the methods without departing from the spirit and scope of the subject matter described herein. Furthermore, the method may be implemented in any suitable hardware, software, firmware, or combination thereof. At the step (301), the ranking system (200) obtains a plurality of variables from the data set (101) and a target variable (102) in the data set (101) from the user. The columns of the data set (101) constitute the plurality of variables and every row of the data set (101) constitutes a value taken by the plurality of variables. An exemplary data set (101) is shown in table of FIGURE 4A. The target variable (102) obtained from the user corresponds to one variable from the plurality of variables in the dataset. For example, consider “Transaction amount” as the target variable (102) from the data set (101) obtained from the user. In an embodiment, the plurality of variables from the data set (101) is aggregated into one or more data types based on meta data (103) of the plurality of variables. The meta data (103) may contain one or more rules for identifying the one or more data types of the plurality of variables. The one or more data types comprises at least one of a numerical data type, a categorical data type, a temporal data type and a spatial data type. As shown in table of the FIGURE 4A the “data time” and “transaction amount” are identified as numerical data type and “Merchant ID”, “Customer ID”, “Location” and “Merchant Category Code (MCC)” are identified as categorical data type. The data type of the plurality of the variables is identified based on the meta data (103) as shown in table of FIGURE 5A. Further, as shown in table of FIGURE 4B an aggregated categorical data set (104B) corresponding to the categorical data type of the plurality of variables with respect to the “Transaction amount” is obtained and as shown in table of FIGURE 4C an aggregated numerical data set (104A) corresponding to the numerical data type of the plurality of variables with respect to the “Transaction amount” is obtained. At the step (302), the ranking system (200) identifies a variation of the target variable (102) with respect to the plurality of variables to detect one or more anomalies in the variation based on outliers present in the variation. As shown in FIGURE 6 the one or more anomalies are detected by applying on each variation at least one data analysis method from one or more data analysis method based on a data type of the plurality of variables. In an embodiment, for a numerical data set (104A) the clustering analysis (601) is applied, for categorical data set (104B) or spatial data set (104D) a quartile analysis (604) is applied and for temporal data set (104C) a time series analysis (606) is applied. Clustering analysis (601) is the task of grouping a set of data points from a data set (101) into a cluster group (602) such that the data points in the cluster group (602) are similar to other data (206) points in the cluster as compared with data points in the other cluster group (602). The outliers are identified as the data points at a distance greater than a threshold distance from the center of the cluster. In an embodiment, the threshold distance may be set to half the inter cluster distance. A quartile is a statistical term used to divide a set of data points in the data set (101) into four defined intervals based upon the values of the data points. The values of the data points that separate the intervals are called the first, second, and third quartiles. Suppose a data set (101) includes “2n” data points then, the first quartile (Q1) is computed as a median of the n smallest entries in the data set (101), the second quartile Q2 is the computed median of all “2n” entries in the data set (101) and the third quartile (Q3) is computed as a median of the n largest entries in the data set (101). In an embodiment, the third quartile (Q3) may be referred to as the upper quartile and the first quartile (Q1) may be referred to as the lower quartile. Further, an Interquartile (IQR) distance is computed as a difference between the third quartile Q3 and the first quartile Q1. Furthermore, the outliers are identified as the data points that fall below (Q1 - 1.5 * IQR) and / or above (Q3 + 1.5 * IQR). A boxplot is used for graphically depicting groups of data points in the data set (101) through the respective quartiles. As shown in FIGURE 7B, a box plot may also have lines extending vertically from the boxes indicating variability of the data points outside the upper and lower quartiles. The band inside the box represents the second quartile. Further, the outliers may be plotted as individual points as shown in FIGURE 7B. A time series is a series of data points indexed in time order. The time series is a sequence taken at successive equally spaced points in time. Thus, the time series forms a sequence of discrete-time data. The time series analysis (606) comprises methods for analyzing time series data and extract meaningful statistics (for example auto correlation, cross correlation, mean and the like) and other characteristics of the data. For example, an Auto Regressive Integrated Moving Average (ARIMA) model may be used to perform the time series analysis. Further, the outliers present in the plurality of variables is identified based on results of the applied at least one data analysis method. The outliers are indicative of at least one data point deviating from the plurality of data point in each variation. The identified outliers are used to generate an aggregated data set (101) of the outliers and an aggregated normal data set (205) of the plurality of variables without the outliers. In an embodiment a distance measure is used to identify the outliers based on the one or more cluster group (602) generated by the clustering analysis (601), an inter-quartile distance is used to identify the outliers based on the results of quartile analysis (604) and the expected bound values (607) are used to identify the outliers based on the results of time series analysis (606) as shown in FIGURE 6. Further, FIGURE 7A shows an exemplary outlier identified using the distance measure of the data point with respect to the center of the cluster based on the one or more cluster group (602) generated by the clustering analysis (601). FIGURE 7B shows an exemplary outlier identified by comparing the inter quartile distance (605). Figure 7C shows an exemplary outlier identified based on the timer series analysis. The identified outliers are separated from the aggregated data set (101) corresponding to a data type and stored in the outliers data set (101) and the data in the aggregated data set (101) corresponding to a data type excluding the outlier data is stored in the normal data set (205) as shown in FIGURE 6. The outlier data set (106), and the normal data set (205) of the one or more data types are combined by the data aggregation (608) to generate an aggregated outlier data set (106) and an aggregated normal data set (205) of the plurality of variables without the outliers as shown in FIGURE 6. In an embodiment, the aggregated outlier data set corresponds to the detected one or more anomalies. At the step (303), the ranking system (200) identifies one or more variables (805) from the plurality of variables causing the detected one or more anomalies based on one or more statistical analysis performed on the plurality of variables. In an embodiment at least one statistical analysis from the one or more statistical analysis is applied on an aggregated outlier data set (106) and an aggregated normal data set (205) of the plurality of variables without the outliers and identifying the one or more variables (805) by computing a difference between result of the statistical analysis on the aggregated outlier data set (106) and the aggregated normal data set (205) of the plurality of variables without the outliers. The one or more statistical analysis comprises at least one of a correlation analysis (801), a regression and a learning algorithm. FIGURE 8 shows an exemplary correlation analysis (801) to identify one or more variables (805) from the plurality of variables causing the detected one or more anomalies. A correlation analysis (801) identifies a statistical relationship between two or more variables. A correlation between data points in the outlier data set (106) and the target variable (102) is identified and outlier data correlation (802) is generated. A correlation between data points in the normal data set (205) and the target variable (102) is identified and normal data correlation (803) is generated. The result of correlation analysis (801) generates a numerical measure of correlation denoted as correlation coefficient. The value of correlation coefficient ranges from -1 to +1, wherein +1 indicates strong positive correlation (or agreement), 0 indicates no correlation (or agreement) and -1 indicates strong negative correlation (or agreement) of the data points with respect to the target variable (102). As shown in FIGURE 8 the outlier data correlation (802) indicates V1 having a correlation coefficient of +0.9, V2 having a correlation coefficient of +0.2 and V3 having a correlation coefficient of -1, wherein V1, V2 and V3 indicates the plurality of variables in the outlier data set (106). Further, the normal data correlation (803) indicates V1 having a correlation coefficient of +0.7, V2 having a correlation coefficient of +0.3 and V3 having a correlation coefficient of +1, wherein V1, V2 and V3 indicates the plurality of variables in the normal data set (205) as shown in FIGURE 8. Furthermore, an absolute correlation difference (804) between the correlation coefficient of outlier data correlation (802) and the correlation coefficient of normal data correlation (803) is computed. The V1 has a correlation difference (804) of +0.2 (|+0.9 – 0.7|), V2 has a correlation difference (804) of +0.1 (|+0.2 – 0.3|) and V3 has a correlation difference (804) of +2 (|-1 - 1|). The variable V3 having the highest absolute correlation difference (804) is identified as the one or more variables (805) causing the detected one or more anomalies as compared to the plurality of variables V1 and V2 based on correlation analysis (801) performed on the plurality of variables. In an embodiment a regression model fitted on the plurality of variables may be used to identify the one or more variables (805) causing the detected one or more anomalies. The regression model is fitted using a mathematical equation Y = ß0 + ß1 * V1 + ß2 * V2 + ? + ßn* Vn wherein V1, V2, …, Vn indicates the plurality of variables, Y indicates the target variable (102), ß0, ß1, …, ßn indicates the weights of the regression model. The regression model is fitted to the outlier dataset and the normal data set (205). A difference between the weights of the regression model is fitted to the outlier dataset and the normal data set (205) is computed and one or more variables (805) having the difference greater than a threshold is identified as the cause for the detected one or more anomalies based on regression performed on the plurality of variables. In an embodiment a decision tree or a neural network with a learning algorithm may be used to identify the one or more variables (805) causing the detected one or more anomalies based on correlation analysis (801) performed on the plurality of variables. At the step (304), the ranking system (200) ranks the variation of the target variable (102) with respect to each of the identified one or more variables (805). The variation is ranked based on the computed difference between result of the statistical analysis on an aggregated data set (101) of the outliers and an aggregated normal data set (205) of the plurality of variables without the outliers. Further, a graphical representation (111) is selected from the one or more graphical representations (111) using the meta data (103) for displaying the variation. Further, each of the variation is displayed to the user based on the assigned rank. Furthermore, the user may use the ranked variations for detecting the root cause of anomalies in the data set (101). In an embodiment, the variation of the target variable (102) with respect to each of the identified one or more variables (805) at the step (303) is ranked based on the computed difference. As shown in FIGURE 8 V3 is assigned a rank 1 followed by V2 assigned a rank 2 and V1 assigned a rank 3. Further, a graphical representation (111) for the variation of the V3 with respect to the target variable (102), V2 with respect to the target variable (102) and V1 with respect to the target variable (102) from the one or more graphical representation (111) is chosen using the meta data (103) of the data set (101). As shown in Table 2 of the FIGURE 5 based on the data type of the one or more variable representing the X-axis and the data type of the target variable (102) representing the Y-axis a suitable chart or graphical representation (111) is selected. As an example consider column 1 in Table 2 of the FIGURE 5 a Bar chart is selected as the graphical representation (111) for display the variation of the target variable (102) with the corresponding data type as numerical with respect to each of the identified one or more variables (805) with the corresponding data type as categorical or numerical. Further, the graphical representations (111) selected for the variation of the target variable (102) with respect to each of the identified one or more variables (805) is displayed to the user based on the assigned ranks as shown in FIGURE 1. In an embodiment, the user may select the one or more detected anomalies based on the displayed ranked graphical representations (111) to identify the root cause of anomalies using the existing drill down techniques. As an example, consider a network of distributed systems for detecting the root cause of anomalies. The network of distributed systems includes one or more computing nodes interconnected to form a network. Every computing node is a network of distributed systems generate and collect logs for example a CPU usage, amount of free memory, percentage of time CPU blocked for Input/output, number of blocks read per second, number of blocks written per second and the like. The one or more anomalies are detected using the information in the logs. The one or more anomalies include erroneous behavior in the network of distributed systems or unexpected long response times from the network. These anomalies may be caused by hardware problems, network communication congestion or software bugs in distributed system components. The detection of one or more anomalies may be done using a time series analysis. The detected one or more anomalies are ranked and displayed to the user. Further, the root cause of the anomalies may be detected by performing a drill down analysis based on the ranked one or more anomalies. As another example consider a banking or a financial services domain wherein card-based transactions are performed from various sources like points-of-sale, ecommerce sites and mobile apps. The detection of one or more anomalies includes for example suspicious transactions, service outage, sudden decrease in transaction in a location and the like may be performed by the ranking system (109). The one or more anomalies may result in poor quality of service to customers from banks. So, the banks need to continuously monitor the transaction patterns to detect and identify the route cause for the one or more anomalies. The one or more anomalies may be detected using the information recorded in the bank payment network for example transaction logs, customer relationship data and the like. Consider an anomaly for example, service outage in a location. The ranking system (109) may identify the service outage using the time series and the categorical analysis. The one or more variables causing the detected anomaly may be identified by the ranking system (109). The one or more variables may include for example a payment network failure due to huge traffic in the location. The one or more identified variables are ranked and displayed to the user using a suitable graphical representation (111). Further, the root cause of the anomalies may be detected by performing a drill down analysis based on the ranked one or more anomalies. COMPUTER SYSTEM FIGURE 9 illustrates a block diagram of an exemplary computer system (900) for implementing embodiments consistent with the present disclosure. In an embodiment, the computer system (900) may be used to implement the method of ranking variables for detecting the root cause of anomalies in the data set. The computer system (900) may comprise a central processing unit (“CPU” or “processor”) (902). The processor (902) may comprise at least one data processor for executing program components for dynamic resource allocation at run time. The processor (902) may include specialized processing units such as integrated system (bus) controllers, memory (202) management control units, floating point units, graphics processing units, digital signal processing units, etc. The processor (902) may be disposed in communication with one or more input/output (I/O) devices (not shown) via I/O interface (901). The I/O interface (901) may employ communication protocols/methods such as, without limitation, audio, analog, digital, monoaural, RCA, stereo, IEEE-(139)4, serial bus, universal serial bus (USB), infrared, PS/2, BNC, coaxial, component, composite, digital visual interface (DVI), high-definition multimedia interface (HDMI), RF antennas, S-Video, VGA, IEEE (802).n /b/g/n/x, Bluetooth, cellular (e.g., code-division multiple access (CDMA), high-speed packet access (HSPA+), global system for mobile communications (GSM), long-term evolution (LTE), WiMax, or the like), etc. Using the I/O interface (901), the computer system (900) may communicate with one or more I/O devices. For example, the input device (910) may be an antenna, keyboard, mouse, joystick, (infrared) remote control, camera, card reader, fax machine, dongle, biometric reader, microphone, touch screen, touchpad, trackball, stylus, scanner, storage device, transceiver, video device/source, etc. The output device (911) may be a printer, fax machine, video display (e.g., cathode ray tube (CRT), liquid crystal display (LCD), light-emitting diode (LED), plasma, Plasma display panel (PDP), Organic light-emitting diode display (OLED) or the like), audio speaker, etc. In some embodiments, the computer system (900) is connected to the service operator through a communication network (909). The processor (902) may be disposed in communication with the communication network (909) via a network interface (903). The network interface (903) may communicate with the communication network (909). The network interface (903) may employ connection protocols including, without limitation, direct connect, Ethernet (e.g., twisted pair 10/100/1000 Base T), transmission control protocol/Internet protocol (TCP/IP), token ring, IEEE 802.11a/b/g/n/x, etc. The communication network (909) may include, without limitation, a direct interconnection, e-commerce network, a peer to peer (P2P) network, local area network (LAN), wide area network (WAN), wireless network (e.g., using Wireless Application Protocol), the Internet, Wi-Fi, etc. Using the network interface (903) and the communication network (909), the computer system (900) may communicate with the one or more service operators. In some embodiments, the processor (902) may be disposed in communication with a memory (905) (e.g., RAM, ROM, etc. not shown in Figure 9 via a storage interface (904). The storage interface (904) may connect to memory (905) including, without limitation, memory (202) drives, removable disc drives, etc., employing connection protocols such as serial advanced technology attachment (SATA), Integrated Drive Electronics (IDE), IEEE-1394, Universal Serial Bus (USB), fiber channel, Small Computer Systems Interface (SCSI), etc. The memory (202) drives may further include a drum, magnetic disc drive, magneto-optical drive, optical drive, Redundant Array of Independent Discs (RAID), solid-state memory (202) devices, solid-state drives, etc. The memory (905) may store a collection of program or database components, including, without limitation, user interface (906), an operating system (907), web server (908) etc. In some embodiments, computer system (900) may store user/application data (906), such as the data, variables, records, etc. as described in this disclosure. Such databases may be implemented as fault-tolerant, relational, scalable, secure databases such as Oracle or Sybase. The operating system (907) may facilitate resource management and operation of the computer system (900). Examples of operating systems include, without limitation, APPLE® MACINTOSH® OS X®, UNIX®, UNIX-like system distributions (E.G., BERKELEY SOFTWARE DISTRIBUTION® (BSD), FREEBSD®, NETBSD®, OPENBSD, etc.), LINUX® DISTRIBUTIONS (E.G., RED HAT®, UBUNTU®, KUBUNTU®, etc.), IBM®OS/2®, MICROSOFT® WINDOWS® (XP®, VISTA®/7/8, 10 etc.), APPLE® IOS®, GOOGLETM ANDROIDTM, BLACKBERRY® OS, or the like. In some embodiments, the computer system (900) may implement a web browser (908) stored program component. The web browser (908) may be a hypertext viewing application, such as MICROSOFT® INTERNET EXPLORER®, GOOGLETM CHROMETM, MOZILLA® FIREFOX®, APPLE® SAFARI®, etc. Secure web browsing may be provided using Secure Hypertext Transport Protocol (HTTPS), Secure Sockets Layer (SSL), Transport Layer Security (TLS), etc. Web browsers (908) may utilize facilities such as AJAX, DHTML, ADOBE® FLASH®, JAVASCRIPT®, JAVA®, Application Programming Interfaces (APIs), etc. In some embodiments, the computer system (900) may implement a mail server stored program component. The mail server may be an Internet mail server such as Microsoft Exchange, or the like. The mail server may utilize facilities such as Active Server Pages (ASP), ACTIVEX®, ANSI® C++/C#, MICROSOFT®, .NET, CGI SCRIPTS, JAVA®, JAVASCRIPT®, PERL®, PHP, PYTHON®, WEBOBJECTS®, etc. The mail server may utilize communication protocols such as Internet Message Access Protocol (IMAP), Messaging Application Programming Interface (MAPI), MICROSOFT® Exchange, Post Office Protocol (POP), Simple Mail Transfer Protocol (SMTP), or the like. In some embodiments, the computer system (900) may implement a mail client stored program component. The mail client may be a mail viewing application, such as APPLE® MAIL, MICROSOFT® ENTOURAGE®, MICROSOFT® OUTLOOK®, MOZILLA® THUNDERBIRD®, etc. Furthermore, one or more computer-readable storage media may be utilized in implementing embodiments consistent with the present invention. A computer-readable storage medium refers to any type of physical memory (202) on which information or data readable by a processor (203) may be stored. Thus, a computer-readable storage medium may store instructions for execution by one or more processors, including instructions for causing the processors to perform steps or stages consistent with the embodiments described herein. The term “computer-readable medium” should be understood to include tangible items and exclude carrier waves and transient signals, i.e., non-transitory. Examples include Random Access memory (202) (RAM), Read-Only memory (202) (ROM), volatile memory (202), non-volatile memory (202), hard drives, Compact Disc (CD) ROMs, Digital Video Disc (DVDs), flash drives, disks, and any other known physical storage media. The terms "an embodiment", "embodiment", "embodiments", "the embodiment", "the embodiments", "one or more embodiments", "some embodiments", and "one embodiment" mean "one or more (but not all) embodiments of the invention(s)" unless expressly specified otherwise. The terms "including", "comprising", “having” and variations thereof mean "including but not limited to", unless expressly specified otherwise. The enumerated listing of items does not imply that any or all of the items are mutually exclusive, unless expressly specified otherwise. The terms "a", "an" and "the" mean "one or more", unless expressly specified otherwise. A description of an embodiment with several components in communication with each other does not imply that all such components are required. On the contrary, a variety of optional components are described to illustrate the wide variety of possible embodiments of the invention. When a single device or article is described herein, it may be readily apparent that more than one device/article (whether or not they cooperate) may be used in place of a single device/article. Similarly, where more than one device or article is described herein (whether or not they cooperate), it may be readily apparent that a single device/article may be used in place of the more than one device or article or a different number of devices/articles may be used instead of the shown number of devices or programs. The functionality and/or the features of a device may be alternatively embodied by one or more other devices which are not explicitly described as having such functionality/features. Thus, other embodiments of the invention need not include the device itself. The illustrated operations of FIGURE 3 show certain events occurring in a certain order. In alternative embodiments, certain operations may be performed in a different order, modified or removed. Moreover, steps may be added to the above described logic and still conform to the described embodiments. Further, operations described herein may occur sequentially or certain operations may be processed in parallel. Yet further, operations may be performed by a single processing unit or by distributed processing units. Thus, the ranking variables for detecting the root cause of anomalies in the data set (101) recommends one or more graphical representations (111) for displaying the variation of the target variable (102) with respect to each of the identified one or more variables (805) using the meta data (103) associated with the data set (101). Further, the ranked graphical representations (111) provide drill-down insights based on the detected one or more anomalies in data set (101) and a relevant graphical representation (111) explaining one or more variables (805) causing the one or more anomalies. Finally, the language used in the specification has been principally selected for readability and instructional purposes, and it may not have been selected to delineate or circumscribe the inventive subject matter. It is therefore intended that the scope of the invention be limited not by this detailed description, but rather by any claims that issue on an application based here on. Accordingly, the disclosure of the embodiments of the invention is intended to be illustrative, but not limiting, of the scope of the invention, which is set forth in the following claims. While various aspects and embodiments have been disclosed herein, other aspects and embodiments may be apparent to those skilled in the art. The various aspects and embodiments disclosed herein are for purposes of illustration and are not intended to be limiting, with the true scope and spirit being indicated by the following claims. REFERRAL NUMERALS: Reference number Description 101 Data Set 102 Target variable 103 Meta data 104A Numerical data set 104B Categorical data set 104C Temporal data set 104D Spatial data set 105 Outlier Detection Unit 106 Outlier Data set 107 Variable Identification Unit 108 Identified variables data set 109 Ranking Unit 110 Ranked variation 111 Graphical representation 200 Ranking system 201 I/O Interface 202 Memory 203 Processor 204 Data 205 Normal data set 206 Other data 207 Modules 208 Input Module (209) Output Module (210) Other Module 601 Clustering Analysis 602 Cluster group 603 Distance measurement 604 Quartile Analysis 605 Compare with Inter Quartile Distance 606 Time Series Analysis 607 Expected Bound Value 608 Data aggregation 801 Correlation Analysis 802 Outlier Data Correlation 803 Normal Data Correlation 804 Correlation Difference 805 One or more variables 900 Computer System 901 I/O interface 902 Processor 903 Network Interface 904 Storage Interface 905 Memory 906 User Interface 907 Operating System 908 Web Server 909 Communication Network 910 Input Device 911 Output Device 912 Remote Devices
Claims:We Claim:
1. A method of ranking variables for detecting the root cause of anomalies in the data set (101), the method comprising:
obtaining, by a ranking system (200), a plurality of variables from the data set (101) and a target variable (102) in the data set (101);
identifying, by the ranking system (200), a variation of the target variable (102) with respect to the plurality of variables to detect one or more anomalies in the variation of the target variable (102) with respect to the plurality of variables based on outliers present in the variation;
identifying, by the ranking system (200), one or more variables from the plurality of variables causing the detected one or more anomalies based on one or more statistical analysis performed on the plurality of variables; and
ranking, by the ranking system (200), the variation of the target variable (102) with respect to each of the identified one or more variables (805), wherein each of the variation is displayed based on the ranking for detecting the root cause of anomalies in the data set (101).
2. The method as claimed in claim 1, wherein the plurality of variables is aggregated into one or more data types based on meta data (103) of the plurality of variables.
3. The method as claimed in claim 2, wherein the one or more data types comprises at least one of a numerical data type, a categorical data type, a temporal data type and a spatial data type.
4. The method as claimed in claim 2, wherein the meta data (103) comprises at least one of set of rules to identify a data type of the plurality of variables, one or more graphical representations (111) for displaying the variation of the target variable (102) with respect to each of the identified one or more variables (805) and the one or more data types supported for a graphical representation (111), and one or more data analysis method to analyze the variation of the target variable (102) with respect to the plurality of variables and the one or more data types supported for corresponding data analysis method.
5. The method as claimed in claim 4, wherein the one or more data analysis method comprises at least one of a clustering analysis (601), quartile analysis (604) and a time series analysis (606).
6. The method as claimed in claim 1, wherein detecting the one or more anomalies comprises:
applying on each variation at least one data analysis method from one or more data analysis method based on a data type of the plurality of variables; and
identifying the outliers present in the plurality of variables based on results of the applied at least one data analysis method.
7. The method as claimed in claim 1, wherein the outliers indicative of at least one data point deviating from the plurality of data point in each variation, are used to generate an aggregated outlier data set (106) and an aggregated normal data set (205) of the plurality of variables without the outliers.
8. The method as claimed in claim 1, wherein identifying the one or more variables (805) comprises:
applying at least one statistical analysis from the one or more statistical analysis on an aggregated outlier data set (106) and an aggregated normal data set (205) of the plurality of variables without the outliers; and
identifying the one or more variables (805) by computing a difference between result of the statistical analysis on the aggregated outlier data set (106) and the aggregated normal data set (205) of the plurality of variables without the outliers.
9. The method as claimed in claim 1, wherein the one or more statistical analysis comprises at least one of a correlation analysis (801), a regression and a learning algorithm.
10. The method as claimed in claim 1, wherein ranking the variation of the target variable (102) with respect to each of the identified one or more variables (805) is based on a computed difference between result of an statistical analysis on an aggregated outlier data set (106) and an aggregated normal data set (205) of the plurality of variables without the outliers, wherein a graphical representation (111) is selected from the one or more graphical representations (111) for displaying the variation.
11. A ranking system (200) comprising:
a processor (203); and
a memory (202) communicatively coupled to the processor (203), wherein the memory (202) stores the processor (203) instructions, which, on execution, causes the processor (203) to:
obtain, a plurality of variables from the data set (101) and a target variable (102) in the data set (101);
identify, a variation of the target variable (102) with respect to the plurality of variables to detect one or more anomalies in the variation of the target variable (102) with respect to the plurality of variables based on outliers present in the variation;
identify, one or more variables (805) from the plurality of variables causing the detected one or more anomalies based on one or more statistical analysis performed on the plurality of variables; and
rank, the variation of the target variable (102) with respect to each of the identified one or more variables (805), wherein each of the variation is displayed based on the ranking for detecting the root cause of anomalies in the data set (101).
12. The ranking system (200) as claimed in claim 11, wherein the processor (203) is configured to aggregate the obtained plurality of variables into one or more data types based on meta data (103) of the plurality of variables.
13. The ranking system (200) as claimed in claim 12, wherein the one or more data types comprises at least one of a numerical data type, a categorical data type, a temporal data type and a spatial data type.
14. The ranking system (200) as claimed in claim 12, wherein the meta data (103) comprises at least one of set of rules to identify a data type of the plurality of variables, one or more graphical representations (111) for displaying the variation of the target variable (102) with respect to each of the identified one or more variables (805) and the one or more data types supported for a graphical representation (111), and one or more data analysis method to analyze the variation of the target variable (102) with respect to the plurality of variables and the one or more data types supported for the corresponding data analysis method.
15. The ranking system (200) as claimed in claim 14, wherein the one or more data analysis method comprises at least one of a clustering analysis (601), quartile analysis (604) and a time series analysis (606).
16. The ranking system (200) as claimed in claim 11, wherein the processor (203) is configured to detect the one or more anomalies comprises:
applying on each variation at least one data analysis method from one or more data analysis method based on a data type of the plurality of variables; and
identifying the outliers present in the plurality of variables based on results of the applied at least one data analysis method.
17. The ranking system (200) as claimed in claim 11, wherein the processor (203) is configured to generate an aggregated outlier data set (106) and an aggregated normal data set (205) of the plurality of variables without the outliers based on identified outliers in the plurality of variables, further wherein the outliers are indicative of at least one data point deviating from the plurality of data point in each variation.
18. The ranking system (200) as claimed in claim 11, wherein the processor (203) is configured to identify the one or more variables (805) comprises:
applying at least one statistical analysis from the one or more statistical analysis on an aggregated outlier data set (106) and an aggregated normal data set (205) of the plurality of variables without the outliers; and
identifying the one or more variables (805) by computing a difference between result of the statistical analysis on the aggregated outlier data set (106) and the aggregated normal data set (205) of the plurality of variables without the outliers.
19. The ranking system (200) as claimed in claim 11, wherein the one or more statistical analysis comprises at least one of a correlation analysis (801), a regression and a learning algorithm.
20. The ranking system (200) as claimed in claim 11, wherein the processor (203) is configured to rank the variation of the target variable (102) with respect to each of the identified one or more variables (805) is based on a computed difference between result of an statistical analysis on an aggregated outlier data set (106) and an aggregated normal data set (205) of the plurality of variables without the outliers, wherein a graphical representation (111) is selected from the one or more graphical representation (111) for displaying the variation.
Dated this on July 18th, 2019
MADHUSUDAN S.T.
IN/PA-1297
K & S PARTNERS
AGENT FOR THE APPLICANT
, Description:FORM 2
THE PATENTS ACT 1970
[39 OF 1970]
&
THE PATENTS RULES, 2003
COMPLETE SPECIFICATION
[See section 10 and Rule 13]
TITLE: “METHOD AND SYSTEM FOR DETECTING THE ROOT CAUSE OF ANOMALIES IN THE DATA SET”
Name and Address of the Applicant: Hitachi, Ltd., 6-6, Marunouchi 1-chome, Chiyoda-ku, Tokyo 100-8280, Japan
Nationality: Japan
The following specification particularly describes the invention and the manner in which it is to be performed.
TECHNICAL FIELD
The present disclosure relates to the field of data analysis. Particularly, but not exclusively, the present disclosure relates to a method and system for detecting root cause of anomalies in the data set.
BACKGROUND
Generally, it is a challenging task to identify an anomaly in a large-scale network. The data set collected from the distributed systems in a large-scale network is very huge and may contain noisy data points, so manual checking of the data set to detect one or more anomalies is time consuming and error prone. Therefore, an automated system is used to detect the one or more anomalies in the data set by ranking the graphical representation of the variation of the target variable with respect to the plurality of variables in the dataset. The identified one or more anomalies must be rectified. Hence, we need the root cause of the one or more anomalies. The existing graphical representation systems automatically generate a graphical representation based on pre-assigned mapping rules to represent one or more data variables from a data set. The pre-assigned mapping rules are based on matching data properties with recommended graphical representation charts. In addition, existing graphical representation systems also continuously rank the graphical representations based on a historical user preference and recommend them to the users. Further, the ranking is based on properties of a data field, and many users use exploratory data analysis technique to identify interesting insights from the data relevant to the use case application. So, choosing an appropriate graphical representation when large amounts of data fields are available and finding right graphical representation chart is time consuming task.
An issue with the existing techniques is the time consumed to generate the rankings and the usage of adulterated historical preferences. This leads to in-accurate ranking and ultimately resulting in sub-optimal graphical representation chart recommendations.
An issue with the existing techniques is redundant recommendations due to user's past preference bias and for new data sets the rankings are not accurate due to the usage of historical data regarding the user past preferences and data from the user profile to assign the rankings.
The information disclosed in this background of the disclosure section is only for enhancement of understanding of the general background of the invention and should not be taken as an acknowledgement or any form of suggestion that this information forms the prior art already known to a person skilled in the art.
SUMMARY
One or more shortcomings of the prior art are overcome, and additional advantages are provided through the provision of method of the present disclosure.
Additional features and advantages are realized through the techniques of the present disclosure. Other embodiments and aspects of the disclosure are described in detail herein and are considered a part of the claimed disclosure.
Disclosed herein is a method of ranking variables for detecting the root cause of anomalies in the data set. The method includes obtaining a plurality of variables from the data set and a target variable in the data set. Further, the method includes identifying a variation of the target variable with respect to the plurality of variables to detect one or more anomalies in the variation of the target variable with respect to the plurality of variables based on outliers present in the variation. Furthermore, the method includes identifying one or more variables from the plurality of variables causing the detected one or more anomalies based on one or more statistical analysis performed on the plurality of variables. Finally, the method includes ranking the variation of the target variable with respect to each of the identified one or more variables, wherein each of the variation is displayed based on the ranking for detecting the root cause of anomalies in the data set.
Further, the present disclosure discloses a ranking system including a processor and a memory communicatively coupled to the processor, wherein the memory stores the processor instructions, which, on execution, causes the processor to obtain, a plurality of variables from the data set and a target variable in the data set. Further, the processor is configured to identify a variation of the target variable with respect to the plurality of variables to detect one or more anomalies in the variation of the target variable with respect to the plurality of variables based on outliers present in the variation. Furthermore, the processor is configured to identify one or more variables from the plurality of variables causing the detected one or more anomalies based on one or more statistical analysis performed on the plurality of variables. Finally, the processor is configured to rank the variation of the target variable with respect to each of the identified one or more variables, wherein each of the variation is displayed based on the ranking for detecting the root cause of anomalies in the data set.
The foregoing summary is illustrative only and is not intended to be in any way limiting. In addition to the illustrative aspects, embodiments, and features described above, further aspects, embodiments, and features may become apparent by reference to the drawings and the following detailed description.
BRIEF DESCRIPTION OF THE ACCOMPANYING DRAWINGS
The novel features and characteristic of the disclosure are set forth in the appended claims. The disclosure itself, however, as well as a preferred mode of use, further objectives and advantages thereof, may best be understood by reference to the following detailed description of an illustrative embodiment when read in conjunction with the accompanying drawings. The accompanying drawings, which are incorporated in and constitute a part of this disclosure, illustrate exemplary embodiments and, together with the description, serve to explain the disclosed principles. In the figures, the left-most digit(s) of a reference number identifies the figure in which the reference number first appears. One or more embodiments are now described, by way of example only, with reference to the accompanying figures wherein like reference numerals represent like elements and in which:
FIGURE 1 shows an exemplary system for detecting the root cause of anomalies in the data set, in accordance with some embodiments of the present disclosure;
FIGURE 2 shows a detailed block diagram of a ranking system, in accordance with some embodiments of the present disclosure;
FIGURE 3 shows a flowchart illustrating method for detecting the root cause of anomalies in the data set, in accordance with some embodiment of the present disclosure;
FIGURE 4A shows an exemplary data set, in accordance with some embodiments of the present disclosure;
FIGURE 4B shows an exemplary aggregated data set based on the categorical data type of the plurality of variables, in accordance with some embodiments of the present disclosure;
FIGURE 4C shows an exemplary aggregated data set based on the temporal data type of the plurality of variables, in accordance with some embodiments of the present disclosure;
FIGURE 5A shows an exemplary meta data for identifying the data type of plurality of variables in a data set, in accordance with some embodiments of the present disclosure;
FIGURE 5B shows an exemplary meta data for selecting a graphical representation based on the data type of plurality of variables, in accordance with some embodiments of the present disclosure;
FIGURE 5C shows an exemplary meta data for selecting a data analysis method based on the data type of plurality of variables, in accordance with some embodiments of the present disclosure;
FIGURE 6 shows an exemplary detection of outliers present in the variation, in accordance with some embodiments of the present disclosure;
FIGURE 7A shows an exemplary clustering analysis, in accordance with some embodiments of the present disclosure;
FIGURE 7B shows an exemplary quartile analysis, in accordance with some embodiments of the present disclosure;
FIGURE 7C shows an exemplary time series analysis, in accordance with some embodiments of the present disclosure;
FIGURE 8 shows an exemplary identification of one or more variables using a correlation analysis, in accordance with some embodiments of the present disclosure; and
FIGURE 9 shows a general-purpose computer system to rank variables for detecting the root cause of anomalies in the data set, in accordance with embodiments of the present disclosure.
It should be appreciated by those skilled in the art that any block diagrams herein represent conceptual views of illustrative systems embodying the principles of the present subject matter. Similarly, it may be appreciated that any flow charts, flow diagrams, state transition diagrams, pseudo code, and the like represent various processes which may be substantially represented in computer readable medium and executed by a computer or processor, whether or not such computer or processor is explicitly shown.
DETAILED DESCRIPTION
In the present document, the word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any embodiment or implementation of the present subject matter described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments.
While the disclosure is susceptible to various modifications and alternative forms, specific embodiment thereof has been shown by way of example in the drawings and may be described in detail below. It should be understood, however that it is not intended to limit the disclosure to the particular forms disclosed, but on the contrary, the disclosure is to cover all modifications, equivalents, and alternative falling within the scope of the disclosure.
The terms “comprises”, “includes” “comprising”, “including” or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a setup, device or method that comprises a list of components or steps does not include only those components or steps but may include other components or steps not expressly listed or inherent to such setup or device or method. In other words, one or more elements in a system or apparatus proceeded by “comprises… a” or “includes…a” does not, without more constraints, preclude the existence of other elements or additional elements in the system or apparatus.
The present disclosure describes a method of ranking variables for detecting the root cause of anomalies in the data set. The method includes obtaining a plurality of variables from the data set and a target variable in the data set. Further, the method includes identifying a variation of the target variable with respect to the plurality of variables to detect one or more anomalies in the variation of the target variable with respect to the plurality of variables based on outliers present in the variation. Furthermore, the method includes identifying one or more variables from the plurality of variables causing the detected one or more anomalies based on one or more statistical analysis performed on the plurality of variables. Finally, the method includes ranking the variation of the target variable with respect to each of the identified one or more variables, wherein each of the variation is displayed based on the ranking for detecting the root cause of anomalies in the data set.
In the following detailed description of the embodiments of the disclosure, reference is made to the accompanying drawings that form a part hereof, and in which are shown by way of illustration specific embodiments in which the disclosure may be practiced. These embodiments are described in sufficient detail to enable those skilled in the art to practice the disclosure, and it is to be understood that other embodiments may be utilized and that changes may be made without departing from the scope of the present disclosure. The following description is, therefore, not to be taken in a limiting sense.
FIGURE 1 shows an exemplary system for detecting the root cause of anomalies in the data set (101), in accordance with some embodiments of the present disclosure.
In an embodiment, a user may provide a data set (101) including a plurality of rows and columns. In another embodiment, the data set may be retrieved from the memory (202) of the ranking system (109). Further, the user selects a target variable (102) from the data set (101) for analysis. The data set (101) is a collection of related and discrete items of data that may be accessed individually or in combination or managed as a whole entity. The columns of the data set (101) constitute plurality of variables and every row of the data set (101) constitutes a value taken by the plurality of variables. The target variable (102) denoted as a dependent variable, is a column of the data set (101) under observation. The plurality of variables in the data set (101) excluding the target variable (102) is denoted as independent variables. The plurality of variables excluding the target variable (102) in the data set (101) is aggregated based on the data type (for example numerical, categorical, temporal, spatial) of the plurality of variables. The meta data (103) of the data set (101) includes a set rules to identify a data type of the plurality of variables. Further, a variation of the target variable (102) with respect to the aggregated plurality of variables corresponding to a data type is identified. Based on the variation corresponding to the data type, one or more anomalies are detected based on outliers present in the variation by the outlier detection unit (105). The outliers are indicative of at least one data point deviating from the plurality of data point in each variation. The detected outliers with respect to the aggregated plurality of variables is stored in the outliers data set (106). Furthermore, one or more variables (805) from the plurality of variables causing the detected one or more anomalies is identified based on one or more statistical analysis performed on the plurality of variables by the variable identification unit (107). The identified one or more variables (805) from the plurality of variables casing the detected one or more anomalies are stored in the identified variables data set (108). The one or more statistical analysis suitable for the plurality of variables is obtained from the meta data (103) associated with the data set (101). Finally, the variation of the target variable (102) with respect to each of the identified one or more variables (805) is ranked by the ranking unit (109). For example, if V1, V2, and V3 are the identified one or more variables, a rank 1 is assigned to the variation of the target variable with the identified one or more variables V1, a rank 2 is assigned to the variation of the target variable with the identified one or more variables V2, and a rank 3 is assigned to the variation of the target variable with the identified one or more variables V3 as shown in Figure 1. For the ranked variation (110) a suitable graphical representation (111) is selected from the meta data (103) associated with the data set (101) based on the data type of the one or more variables (805) in the variation. For example, let the data type of the one or more variables V1 be numerical and the one or more variables V2 be spatial. Therefore, a line chart is selected to display the variation of the one or more variables V1 with respect to the target variable and a Geo chart is selected to display the variation of the one or more variables V2 with respect to the target variable as shown in Figure 1. In one embodiment, the order of providing the graphical representation (111) may be based on the rank associated with the one or more variables. For example, the line chart is displayed at top indicating that the corresponding variable V3 is ranked 1 and likewise, the Geo chart may be displayed below the line chart indicating that the corresponding variable V2 is ranked 2. The ranked variations along with the corresponding graphical representations (111) are displayed to the user for detecting the root cause of anomalies in the data set (101).
FIGURE 2 shows a detailed block diagram of a ranking system (200), in accordance with some embodiments of the present disclosure.
The ranking system (200) may include Central Processing Unit (“CPU” or “processor”) (203) and a memory (202) storing instructions executable by the processor (203). The processor (203) may include at least one data processor for executing program components for executing user or system-generated requests. The memory (202) may be communicatively coupled to the processor (203). The ranking system (200) further includes an Input/ Output (I/O) interface (201). The I/O interface (201) may be coupled with the processor (203) through which an input signal or/and an output signal may be communicated. In one embodiment, the ranking system (200) may receive the data set (101) and the target variable (102) through the I/O interface (201).
In some implementations, the ranking system (200) may include data (204) and modules (207). As an example, the data (204) and modules (207) may be stored in the memory (202) configured in the ranking system (200) as shown in the FIGURE 2. In one embodiment, the data (204) may include, for example, a data set (101), a meta data (103), a normal data set (205), an outlier data set (106), identified variables (108) and other data (206). In the illustrated FIGURE 2, modules (207) are described herein in detail.
In an embodiment, the data set (101) is a collection of related and discrete items of data that may be accessed individually or in combination or managed as a whole entity. The data set (101) includes arrangement of related data in a plurality of rows and columns. The columns of the data set (101) constitute the plurality of variables and every row of the data set (101) constitutes a value taken by the plurality of variables. FIGURE 4A shows an exemplary data set (101) arranged in the form of table including a plurality of rows and columns. The columns (for example Date Time, Merchant ID, Customer ID, Transaction Amount, Location, Merchant Category Code (MCC) constitute the plurality of variables. The rows in table of FIGURE 4A constitutes the value taken by the corresponding variable of the column. For example, consider row 1 in the table of FIGURE 4A wherein “M101” is the value taken by the variable “Merchant ID”, “Bangalore” is the value taken by the variable “Location” and the like.
In an embodiment, the metadata is data that provides information about the data set (101). The distinct types of metadata are descriptive metadata, structural metadata, administrative metadata, reference metadata and statistical metadata. The descriptive metadata describes a resource for purposes such as discovery and identification. It can include elements such as title, abstract, author, and keywords. The structural metadata is metadata about containers of data and indicates how compound objects are put together, for example, how pages are ordered to form chapters. It describes the types, versions, relationships and other characteristics of digital materials. The administrative metadata provides information to help manage a resource, such as when and how it was created, file type and other technical information, and who can access it. The reference metadata describes the contents and quality of statistical data. The statistical metadata may also describe processes that collect, process, or produce statistical data; such metadata are also called process data. Further, the meta data (103) includes at least one of set of rules to identify a data type of the plurality of variables, one or more graphical representations (111) for displaying the variation of the target variable (102) with respect to each of the identified one or more variables (805) and the one or more data types supported for a graphical representation (111), and one or more data analysis method to analyze the variation of the target variable (102) with respect to the plurality of variables and the one or more data types supported for corresponding data analysis method. As shown in table of FIGURE 5A the meta data (103) includes one or more rules for identifying the data type of the plurality of variables in the data set (101). For example, consider the row 1 in the table of FIGURE 5A, the one or more variables (805) in the data set (101) having a value of an integer type or a double type is identified as numerical data type and the like. As shown in table of FIGURE 5B the meta data (103) includes one or more graphical representations (111) for displaying the variation of the target variable (102) with respect to each of the identified one or more variables (805). For example, consider column 1 in table of Figure 5B, for the data types of numerical and categorical representing the X-axis and numerical data type representing the Y-axis a “Bar chart” is used and the like. As shown in table of FIGURE 5C the meta data (103) includes one or more data analysis method to analyze the variation of the target variable (102) with respect to the plurality of variables and the one or more data types supported for corresponding data analysis method. For example, consider the column 1 in table of the Figure 5C, for a variation of the target variable (102) of the data type temporal or numerical with respect to the plurality of the variables of the data type temporal or numerical a “time series analysis (606)” method is used to analyze the variation and a line chart or a bar chart is used to represent the variation and the like.
In an embodiment, the normal data set (205) includes a subset of rows and columns of the data set (101) excluding the rows identified as outliers in the variation of the target variable (102) of a corresponding data type with respect to the plurality of variables. Further, the normal data set (205) may include an aggregate the normal data set (205) obtained by the variation of the target variable (102) corresponding to the one or more data types.
In an embodiment, the outlier data set (106) includes a subset of rows and columns of the data set (101) identified as outliers in the variation of the target variable (102) of a corresponding data type with respect to the plurality of variables. The outliers in the variation are identified based on the one or more data analysis method including at least one of a clustering analysis (601), quartile analysis (604) and a time series analysis (606). Further, the outliers are indicative of at least one data point deviating from the plurality of data point in each variation. Furthermore, the outlier data set (106) may include an aggregate the outlier data set (106) obtained by the variation of the target variable (102) corresponding to the one or more data types.
In an embodiment, the identified variables (108) include one or more variables (805) from the plurality of variables causing the detected one or more anomalies based on one or more statistical analysis performed on the plurality of variables. The one or more statistical analysis comprises at least one of a correlation analysis (801), a regression and a learning algorithm. For example, consider the table of FIGURE 4A, wherein the “Location” variable may be identified as the one or more variable from the plurality of variables causing the detected one or more anomalies (for example high transaction amount in one month).
In an embodiment, the other data (206) may include data regarding the one or more data analysis method including at least one of a clustering analysis (601), quartile analysis (604) and a time series analysis (606) and the one or more statistical analysis including at least one of a correlation analysis (801), a regression and a learning algorithm. Further, the other data (206) may include data regarding the distance measurement (603) between two data points in a variation, computing and comparing the inter quartile distance (605) between the upper and lower quartiles, and computing and comparing the expected bound value (607) between the upper and lower bounds.
In some embodiments, data (204) may be stored in the memory (202) in form of various data structures. Additionally, the data (204) may be organized using data models, such as relational or hierarchical data models. The other data (206) may store data, including temporary data and temporary files, generated by the modules (207) for performing the various functions of the ranking system (200).
In some embodiments, the data (204) stored in the memory (202) may be processed by the modules (207) of the ranking system (200). The modules (207) may be stored within the memory (202). In an example, the modules (207) communicatively coupled to the processor (203) configured in the ranking system (200), may also be present outside the memory (202) as shown in FIGURE 2 and implemented as hardware. As used herein, the term modules (207) may refer to an application specific integrated circuit (ASIC), a FPGA (Field Programmable Gate Array), an electronic circuit, a processor (203) (shared, dedicated, or group) and memory (202) that execute one or more software or firmware programs, a combinational logic circuit, and/or other suitable components that provide the described functionality. In some other embodiments, the modules (207) may be implemented using at least one of ASICs and FPGAs.
In one implementation, the modules (207) may include, for example, an outlier detection unit (105), a variable identification unit (107), an output module (209), an input module (208), a ranking unit (109) and other module (210). It may be appreciated that such aforementioned modules (207) may be represented as a single module or a combination of different modules (207).
In an embodiment, the outlier detection unit (105) is used to identify the outliers present in the plurality of variables based on results of the applied at least one data analysis method. The outliers are indicative of at least one data point deviating from the plurality of data point in each variation. The one or more data analysis method includes at least one of a clustering analysis (601), quartile analysis (604) and a time series analysis (606). For example, the detected outliers are indicated in the FIGURE 7A, FIGURE 7B, and FIGURE 7C.
In an embodiment, the variable identification unit (107) is used to identify the one or more variables (805) from the plurality of variables causing the detected one or more anomalies based on the one or more statistical analysis performed on the plurality of variables. Initially the variable identification unit (107) applies at least one statistical analysis from the one or more statistical analysis on an aggregated outlier data set (106) and an aggregated normal data set (205) of the plurality of variables without the outliers. Further, the variable identification unit (107) identifies the one or more variables (805) by computing a difference between result of the statistical analysis on the aggregated outlier data set (106) and the aggregated normal data set (205) of the plurality of variables without the outliers. The one or more statistical analysis includes at least one of a correlation analysis (801), a regression and a learning algorithm. For example, consider the table of FIGURE 4A, wherein the “Location” variable may be identified as the one or more variable from the plurality of variables causing the detected one or more anomalies (for example high transaction amount in one month).
In an embodiment, the ranking unit (109) is responsible for assigning a rank to the variation of the target variable (102) with respect to each of the identified one or more variables (805). The rank is assigned based on a computed difference between result of a statistical analysis on an aggregated outlier data set (106) and an aggregated normal data set (205) of the plurality of variables without the outliers. Further, a graphical representation (111) is selected from the one or more graphical representations (111) and the variation is displayed to the user using the output module (209).
In an embodiment, the output module (209) is responsible for displaying a graphical representation (111) of each of the variation of the target variable (102) with respect to each of the identified one or more variables (805) based on the rank assigned by the ranking unit (109). Further, the ranked graphical representations (111) displayed to the user enables the user to detect the root cause of anomalies in the data set (101).
In an embodiment, the input module (208) is responsible for obtaining a plurality of variables from the data set (101) and a target variable (102) in the data set (101) for detecting the root cause of anomalies in the data set (101).
In an embodiment, the other module (210) is responsible for measuring the distance between two data points in a variation, measuring the distance between a data point and a group of data points, computing and comparing the inter quartile distance (605) between the upper and lower quartiles, and computing and comparing the expected bound value (607) between the upper and lower bounds.
FIGURE 3 shows a flowchart illustrating method of detecting the root cause of anomalies in the data set (101), in accordance with some embodiment of the present disclosure.
The order in which the method (300) may be described is not intended to be construed as a limitation, and any number of the described method blocks may be combined in any order to implement the method. Additionally, individual blocks may be deleted from the methods without departing from the spirit and scope of the subject matter described herein. Furthermore, the method may be implemented in any suitable hardware, software, firmware, or combination thereof.
At the step (301), the ranking system (200) obtains a plurality of variables from the data set (101) and a target variable (102) in the data set (101) from the user. The columns of the data set (101) constitute the plurality of variables and every row of the data set (101) constitutes a value taken by the plurality of variables. An exemplary data set (101) is shown in table of FIGURE 4A. The target variable (102) obtained from the user corresponds to one variable from the plurality of variables in the dataset. For example, consider “Transaction amount” as the target variable (102) from the data set (101) obtained from the user.
In an embodiment, the plurality of variables from the data set (101) is aggregated into one or more data types based on meta data (103) of the plurality of variables. The meta data (103) may contain one or more rules for identifying the one or more data types of the plurality of variables. The one or more data types comprises at least one of a numerical data type, a categorical data type, a temporal data type and a spatial data type.
As shown in table of the FIGURE 4A the “data time” and “transaction amount” are identified as numerical data type and “Merchant ID”, “Customer ID”, “Location” and “Merchant Category Code (MCC)” are identified as categorical data type. The data type of the plurality of the variables is identified based on the meta data (103) as shown in table of FIGURE 5A. Further, as shown in table of FIGURE 4B an aggregated categorical data set (104B) corresponding to the categorical data type of the plurality of variables with respect to the “Transaction amount” is obtained and as shown in table of FIGURE 4C an aggregated numerical data set (104A) corresponding to the numerical data type of the plurality of variables with respect to the “Transaction amount” is obtained.
At the step (302), the ranking system (200) identifies a variation of the target variable (102) with respect to the plurality of variables to detect one or more anomalies in the variation based on outliers present in the variation. As shown in FIGURE 6 the one or more anomalies are detected by applying on each variation at least one data analysis method from one or more data analysis method based on a data type of the plurality of variables.
In an embodiment, for a numerical data set (104A) the clustering analysis (601) is applied, for categorical data set (104B) or spatial data set (104D) a quartile analysis (604) is applied and for temporal data set (104C) a time series analysis (606) is applied. Clustering analysis (601) is the task of grouping a set of data points from a data set (101) into a cluster group (602) such that the data points in the cluster group (602) are similar to other data (206) points in the cluster as compared with data points in the other cluster group (602). The outliers are identified as the data points at a distance greater than a threshold distance from the center of the cluster. In an embodiment, the threshold distance may be set to half the inter cluster distance. A quartile is a statistical term used to divide a set of data points in the data set (101) into four defined intervals based upon the values of the data points. The values of the data points that separate the intervals are called the first, second, and third quartiles. Suppose a data set (101) includes “2n” data points then, the first quartile (Q1) is computed as a median of the n smallest entries in the data set (101), the second quartile Q2 is the computed median of all “2n” entries in the data set (101) and the third quartile (Q3) is computed as a median of the n largest entries in the data set (101). In an embodiment, the third quartile (Q3) may be referred to as the upper quartile and the first quartile (Q1) may be referred to as the lower quartile. Further, an Interquartile (IQR) distance is computed as a difference between the third quartile Q3 and the first quartile Q1. Furthermore, the outliers are identified as the data points that fall below (Q1 - 1.5 * IQR) and / or above (Q3 + 1.5 * IQR). A boxplot is used for graphically depicting groups of data points in the data set (101) through the respective quartiles. As shown in FIGURE 7B, a box plot may also have lines extending vertically from the boxes indicating variability of the data points outside the upper and lower quartiles. The band inside the box represents the second quartile. Further, the outliers may be plotted as individual points as shown in FIGURE 7B. A time series is a series of data points indexed in time order. The time series is a sequence taken at successive equally spaced points in time. Thus, the time series forms a sequence of discrete-time data. The time series analysis (606) comprises methods for analyzing time series data and extract meaningful statistics (for example auto correlation, cross correlation, mean and the like) and other characteristics of the data. For example, an Auto Regressive Integrated Moving Average (ARIMA) model may be used to perform the time series analysis.
Further, the outliers present in the plurality of variables is identified based on results of the applied at least one data analysis method. The outliers are indicative of at least one data point deviating from the plurality of data point in each variation. The identified outliers are used to generate an aggregated data set (101) of the outliers and an aggregated normal data set (205) of the plurality of variables without the outliers. In an embodiment a distance measure is used to identify the outliers based on the one or more cluster group (602) generated by the clustering analysis (601), an inter-quartile distance is used to identify the outliers based on the results of quartile analysis (604) and the expected bound values (607) are used to identify the outliers based on the results of time series analysis (606) as shown in FIGURE 6.
Further, FIGURE 7A shows an exemplary outlier identified using the distance measure of the data point with respect to the center of the cluster based on the one or more cluster group (602) generated by the clustering analysis (601). FIGURE 7B shows an exemplary outlier identified by comparing the inter quartile distance (605). Figure 7C shows an exemplary outlier identified based on the timer series analysis.
The identified outliers are separated from the aggregated data set (101) corresponding to a data type and stored in the outliers data set (101) and the data in the aggregated data set (101) corresponding to a data type excluding the outlier data is stored in the normal data set (205) as shown in FIGURE 6. The outlier data set (106), and the normal data set (205) of the one or more data types are combined by the data aggregation (608) to generate an aggregated outlier data set (106) and an aggregated normal data set (205) of the plurality of variables without the outliers as shown in FIGURE 6.
In an embodiment, the aggregated outlier data set corresponds to the detected one or more anomalies.
At the step (303), the ranking system (200) identifies one or more variables (805) from the plurality of variables causing the detected one or more anomalies based on one or more statistical analysis performed on the plurality of variables. In an embodiment at least one statistical analysis from the one or more statistical analysis is applied on an aggregated outlier data set (106) and an aggregated normal data set (205) of the plurality of variables without the outliers and identifying the one or more variables (805) by computing a difference between result of the statistical analysis on the aggregated outlier data set (106) and the aggregated normal data set (205) of the plurality of variables without the outliers. The one or more statistical analysis comprises at least one of a correlation analysis (801), a regression and a learning algorithm.
FIGURE 8 shows an exemplary correlation analysis (801) to identify one or more variables (805) from the plurality of variables causing the detected one or more anomalies. A correlation analysis (801) identifies a statistical relationship between two or more variables. A correlation between data points in the outlier data set (106) and the target variable (102) is identified and outlier data correlation (802) is generated. A correlation between data points in the normal data set (205) and the target variable (102) is identified and normal data correlation (803) is generated. The result of correlation analysis (801) generates a numerical measure of correlation denoted as correlation coefficient. The value of correlation coefficient ranges from -1 to +1, wherein +1 indicates strong positive correlation (or agreement), 0 indicates no correlation (or agreement) and -1 indicates strong negative correlation (or agreement) of the data points with respect to the target variable (102). As shown in FIGURE 8 the outlier data correlation (802) indicates V1 having a correlation coefficient of +0.9, V2 having a correlation coefficient of +0.2 and V3 having a correlation coefficient of -1, wherein V1, V2 and V3 indicates the plurality of variables in the outlier data set (106). Further, the normal data correlation (803) indicates V1 having a correlation coefficient of +0.7, V2 having a correlation coefficient of +0.3 and V3 having a correlation coefficient of +1, wherein V1, V2 and V3 indicates the plurality of variables in the normal data set (205) as shown in FIGURE 8. Furthermore, an absolute correlation difference (804) between the correlation coefficient of outlier data correlation (802) and the correlation coefficient of normal data correlation (803) is computed. The V1 has a correlation difference (804) of +0.2 (|+0.9 – 0.7|), V2 has a correlation difference (804) of +0.1 (|+0.2 – 0.3|) and V3 has a correlation difference (804) of +2 (|-1 - 1|). The variable V3 having the highest absolute correlation difference (804) is identified as the one or more variables (805) causing the detected one or more anomalies as compared to the plurality of variables V1 and V2 based on correlation analysis (801) performed on the plurality of variables.
In an embodiment a regression model fitted on the plurality of variables may be used to identify the one or more variables (805) causing the detected one or more anomalies. The regression model is fitted using a mathematical equation
Y = ß0 + ß1 * V1 + ß2 * V2 + ? + ßn* Vn
wherein V1, V2, …, Vn indicates the plurality of variables, Y indicates the target variable (102), ß0, ß1, …, ßn indicates the weights of the regression model. The regression model is fitted to the outlier dataset and the normal data set (205). A difference between the weights of the regression model is fitted to the outlier dataset and the normal data set (205) is computed and one or more variables (805) having the difference greater than a threshold is identified as the cause for the detected one or more anomalies based on regression performed on the plurality of variables.
In an embodiment a decision tree or a neural network with a learning algorithm may be used to identify the one or more variables (805) causing the detected one or more anomalies based on correlation analysis (801) performed on the plurality of variables.
At the step (304), the ranking system (200) ranks the variation of the target variable (102) with respect to each of the identified one or more variables (805). The variation is ranked based on the computed difference between result of the statistical analysis on an aggregated data set (101) of the outliers and an aggregated normal data set (205) of the plurality of variables without the outliers. Further, a graphical representation (111) is selected from the one or more graphical representations (111) using the meta data (103) for displaying the variation. Further, each of the variation is displayed to the user based on the assigned rank. Furthermore, the user may use the ranked variations for detecting the root cause of anomalies in the data set (101).
In an embodiment, the variation of the target variable (102) with respect to each of the identified one or more variables (805) at the step (303) is ranked based on the computed difference. As shown in FIGURE 8 V3 is assigned a rank 1 followed by V2 assigned a rank 2 and V1 assigned a rank 3. Further, a graphical representation (111) for the variation of the V3 with respect to the target variable (102), V2 with respect to the target variable (102) and V1 with respect to the target variable (102) from the one or more graphical representation (111) is chosen using the meta data (103) of the data set (101). As shown in Table 2 of the FIGURE 5 based on the data type of the one or more variable representing the X-axis and the data type of the target variable (102) representing the Y-axis a suitable chart or graphical representation (111) is selected. As an example consider column 1 in Table 2 of the FIGURE 5 a Bar chart is selected as the graphical representation (111) for display the variation of the target variable (102) with the corresponding data type as numerical with respect to each of the identified one or more variables (805) with the corresponding data type as categorical or numerical. Further, the graphical representations (111) selected for the variation of the target variable (102) with respect to each of the identified one or more variables (805) is displayed to the user based on the assigned ranks as shown in FIGURE 1.
In an embodiment, the user may select the one or more detected anomalies based on the displayed ranked graphical representations (111) to identify the root cause of anomalies using the existing drill down techniques.
As an example, consider a network of distributed systems for detecting the root cause of anomalies. The network of distributed systems includes one or more computing nodes interconnected to form a network. Every computing node is a network of distributed systems generate and collect logs for example a CPU usage, amount of free memory, percentage of time CPU blocked for Input/output, number of blocks read per second, number of blocks written per second and the like. The one or more anomalies are detected using the information in the logs. The one or more anomalies include erroneous behavior in the network of distributed systems or unexpected long response times from the network. These anomalies may be caused by hardware problems, network communication congestion or software bugs in distributed system components. The detection of one or more anomalies may be done using a time series analysis. The detected one or more anomalies are ranked and displayed to the user. Further, the root cause of the anomalies may be detected by performing a drill down analysis based on the ranked one or more anomalies.
As another example consider a banking or a financial services domain wherein card-based transactions are performed from various sources like points-of-sale, ecommerce sites and mobile apps. The detection of one or more anomalies includes for example suspicious transactions, service outage, sudden decrease in transaction in a location and the like may be performed by the ranking system (109). The one or more anomalies may result in poor quality of service to customers from banks. So, the banks need to continuously monitor the transaction patterns to detect and identify the route cause for the one or more anomalies. The one or more anomalies may be detected using the information recorded in the bank payment network for example transaction logs, customer relationship data and the like. Consider an anomaly for example, service outage in a location. The ranking system (109) may identify the service outage using the time series and the categorical analysis. The one or more variables causing the detected anomaly may be identified by the ranking system (109). The one or more variables may include for example a payment network failure due to huge traffic in the location. The one or more identified variables are ranked and displayed to the user using a suitable graphical representation (111). Further, the root cause of the anomalies may be detected by performing a drill down analysis based on the ranked one or more anomalies.
COMPUTER SYSTEM
FIGURE 9 illustrates a block diagram of an exemplary computer system (900) for implementing embodiments consistent with the present disclosure. In an embodiment, the computer system (900) may be used to implement the method of ranking variables for detecting the root cause of anomalies in the data set. The computer system (900) may comprise a central processing unit (“CPU” or “processor”) (902). The processor (902) may comprise at least one data processor for executing program components for dynamic resource allocation at run time. The processor (902) may include specialized processing units such as integrated system (bus) controllers, memory (202) management control units, floating point units, graphics processing units, digital signal processing units, etc.
The processor (902) may be disposed in communication with one or more input/output (I/O) devices (not shown) via I/O interface (901). The I/O interface (901) may employ communication protocols/methods such as, without limitation, audio, analog, digital, monoaural, RCA, stereo, IEEE-(139)4, serial bus, universal serial bus (USB), infrared, PS/2, BNC, coaxial, component, composite, digital visual interface (DVI), high-definition multimedia interface (HDMI), RF antennas, S-Video, VGA, IEEE (802).n /b/g/n/x, Bluetooth, cellular (e.g., code-division multiple access (CDMA), high-speed packet access (HSPA+), global system for mobile communications (GSM), long-term evolution (LTE), WiMax, or the like), etc.
Using the I/O interface (901), the computer system (900) may communicate with one or more I/O devices. For example, the input device (910) may be an antenna, keyboard, mouse, joystick, (infrared) remote control, camera, card reader, fax machine, dongle, biometric reader, microphone, touch screen, touchpad, trackball, stylus, scanner, storage device, transceiver, video device/source, etc. The output device (911) may be a printer, fax machine, video display (e.g., cathode ray tube (CRT), liquid crystal display (LCD), light-emitting diode (LED), plasma, Plasma display panel (PDP), Organic light-emitting diode display (OLED) or the like), audio speaker, etc.
In some embodiments, the computer system (900) is connected to the service operator through a communication network (909). The processor (902) may be disposed in communication with the communication network (909) via a network interface (903). The network interface (903) may communicate with the communication network (909). The network interface (903) may employ connection protocols including, without limitation, direct connect, Ethernet (e.g., twisted pair 10/100/1000 Base T), transmission control protocol/Internet protocol (TCP/IP), token ring, IEEE 802.11a/b/g/n/x, etc. The communication network (909) may include, without limitation, a direct interconnection, e-commerce network, a peer to peer (P2P) network, local area network (LAN), wide area network (WAN), wireless network (e.g., using Wireless Application Protocol), the Internet, Wi-Fi, etc. Using the network interface (903) and the communication network (909), the computer system (900) may communicate with the one or more service operators.
In some embodiments, the processor (902) may be disposed in communication with a memory (905) (e.g., RAM, ROM, etc. not shown in Figure 9 via a storage interface (904). The storage interface (904) may connect to memory (905) including, without limitation, memory (202) drives, removable disc drives, etc., employing connection protocols such as serial advanced technology attachment (SATA), Integrated Drive Electronics (IDE), IEEE-1394, Universal Serial Bus (USB), fiber channel, Small Computer Systems Interface (SCSI), etc. The memory (202) drives may further include a drum, magnetic disc drive, magneto-optical drive, optical drive, Redundant Array of Independent Discs (RAID), solid-state memory (202) devices, solid-state drives, etc.
The memory (905) may store a collection of program or database components, including, without limitation, user interface (906), an operating system (907), web server (908) etc. In some embodiments, computer system (900) may store user/application data (906), such as the data, variables, records, etc. as described in this disclosure. Such databases may be implemented as fault-tolerant, relational, scalable, secure databases such as Oracle or Sybase.
The operating system (907) may facilitate resource management and operation of the computer system (900). Examples of operating systems include, without limitation, APPLE® MACINTOSH® OS X®, UNIX®, UNIX-like system distributions (E.G., BERKELEY SOFTWARE DISTRIBUTION® (BSD), FREEBSD®, NETBSD®, OPENBSD, etc.), LINUX® DISTRIBUTIONS (E.G., RED HAT®, UBUNTU®, KUBUNTU®, etc.), IBM®OS/2®, MICROSOFT® WINDOWS® (XP®, VISTA®/7/8, 10 etc.), APPLE® IOS®, GOOGLETM ANDROIDTM, BLACKBERRY® OS, or the like.
In some embodiments, the computer system (900) may implement a web browser (908) stored program component. The web browser (908) may be a hypertext viewing application, such as MICROSOFT® INTERNET EXPLORER®, GOOGLETM CHROMETM, MOZILLA® FIREFOX®, APPLE® SAFARI®, etc. Secure web browsing may be provided using Secure Hypertext Transport Protocol (HTTPS), Secure Sockets Layer (SSL), Transport Layer Security (TLS), etc. Web browsers (908) may utilize facilities such as AJAX, DHTML, ADOBE® FLASH®, JAVASCRIPT®, JAVA®, Application Programming Interfaces (APIs), etc. In some embodiments, the computer system (900) may implement a mail server stored program component. The mail server may be an Internet mail server such as Microsoft Exchange, or the like. The mail server may utilize facilities such as Active Server Pages (ASP), ACTIVEX®, ANSI® C++/C#, MICROSOFT®, .NET, CGI SCRIPTS, JAVA®, JAVASCRIPT®, PERL®, PHP, PYTHON®, WEBOBJECTS®, etc. The mail server may utilize communication protocols such as Internet Message Access Protocol (IMAP), Messaging Application Programming Interface (MAPI), MICROSOFT® Exchange, Post Office Protocol (POP), Simple Mail Transfer Protocol (SMTP), or the like. In some embodiments, the computer system (900) may implement a mail client stored program component. The mail client may be a mail viewing application, such as APPLE® MAIL, MICROSOFT® ENTOURAGE®, MICROSOFT® OUTLOOK®, MOZILLA® THUNDERBIRD®, etc.
Furthermore, one or more computer-readable storage media may be utilized in implementing embodiments consistent with the present invention. A computer-readable storage medium refers to any type of physical memory (202) on which information or data readable by a processor (203) may be stored. Thus, a computer-readable storage medium may store instructions for execution by one or more processors, including instructions for causing the processors to perform steps or stages consistent with the embodiments described herein. The term “computer-readable medium” should be understood to include tangible items and exclude carrier waves and transient signals, i.e., non-transitory. Examples include Random Access memory (202) (RAM), Read-Only memory (202) (ROM), volatile memory (202), non-volatile memory (202), hard drives, Compact Disc (CD) ROMs, Digital Video Disc (DVDs), flash drives, disks, and any other known physical storage media.
The terms "an embodiment", "embodiment", "embodiments", "the embodiment", "the embodiments", "one or more embodiments", "some embodiments", and "one embodiment" mean "one or more (but not all) embodiments of the invention(s)" unless expressly specified otherwise.
The terms "including", "comprising", “having” and variations thereof mean "including but not limited to", unless expressly specified otherwise.
The enumerated listing of items does not imply that any or all of the items are mutually exclusive, unless expressly specified otherwise. The terms "a", "an" and "the" mean "one or more", unless expressly specified otherwise.
A description of an embodiment with several components in communication with each other does not imply that all such components are required. On the contrary, a variety of optional components are described to illustrate the wide variety of possible embodiments of the invention.
When a single device or article is described herein, it may be readily apparent that more than one device/article (whether or not they cooperate) may be used in place of a single device/article. Similarly, where more than one device or article is described herein (whether or not they cooperate), it may be readily apparent that a single device/article may be used in place of the more than one device or article or a different number of devices/articles may be used instead of the shown number of devices or programs. The functionality and/or the features of a device may be alternatively embodied by one or more other devices which are not explicitly described as having such functionality/features. Thus, other embodiments of the invention need not include the device itself.
The illustrated operations of FIGURE 3 show certain events occurring in a certain order. In alternative embodiments, certain operations may be performed in a different order, modified or removed. Moreover, steps may be added to the above described logic and still conform to the described embodiments. Further, operations described herein may occur sequentially or certain operations may be processed in parallel. Yet further, operations may be performed by a single processing unit or by distributed processing units.
Thus, the ranking variables for detecting the root cause of anomalies in the data set (101) recommends one or more graphical representations (111) for displaying the variation of the target variable (102) with respect to each of the identified one or more variables (805) using the meta data (103) associated with the data set (101). Further, the ranked graphical representations (111) provide drill-down insights based on the detected one or more anomalies in data set (101) and a relevant graphical representation (111) explaining one or more variables (805) causing the one or more anomalies.
Finally, the language used in the specification has been principally selected for readability and instructional purposes, and it may not have been selected to delineate or circumscribe the inventive subject matter. It is therefore intended that the scope of the invention be limited not by this detailed description, but rather by any claims that issue on an application based here on. Accordingly, the disclosure of the embodiments of the invention is intended to be illustrative, but not limiting, of the scope of the invention, which is set forth in the following claims.
While various aspects and embodiments have been disclosed herein, other aspects and embodiments may be apparent to those skilled in the art. The various aspects and embodiments disclosed herein are for purposes of illustration and are not intended to be limiting, with the true scope and spirit being indicated by the following claims.
REFERRAL NUMERALS:
Reference number Description
101 Data Set
102 Target variable
103 Meta data
104A Numerical data set
104B Categorical data set
104C Temporal data set
104D Spatial data set
105 Outlier Detection Unit
106 Outlier Data set
107 Variable Identification Unit
108 Identified variables data set
109 Ranking Unit
110 Ranked variation
111 Graphical representation
200 Ranking system
201 I/O Interface
202 Memory
203 Processor
204 Data
205 Normal data set
206 Other data
207 Modules
208 Input Module
(209) Output Module
(210) Other Module
601 Clustering Analysis
602 Cluster group
603 Distance measurement
604 Quartile Analysis
605 Compare with Inter Quartile Distance
606 Time Series Analysis
607 Expected Bound Value
608 Data aggregation
801 Correlation Analysis
802 Outlier Data Correlation
803 Normal Data Correlation
804 Correlation Difference
805 One or more variables
900 Computer System
901 I/O interface
902 Processor
903 Network Interface
904 Storage Interface
905 Memory
906 User Interface
907 Operating System
908 Web Server
909 Communication Network
910 Input Device
911 Output Device
912 Remote Devices
| # | Name | Date |
|---|---|---|
| 1 | 201941028908-STATEMENT OF UNDERTAKING (FORM 3) [18-07-2019(online)].pdf | 2019-07-18 |
| 2 | 201941028908-REQUEST FOR EXAMINATION (FORM-18) [18-07-2019(online)].pdf | 2019-07-18 |
| 3 | 201941028908-POWER OF AUTHORITY [18-07-2019(online)].pdf | 2019-07-18 |
| 4 | 201941028908-FORM-26 [18-07-2019(online)].pdf | 2019-07-18 |
| 5 | 201941028908-FORM 18 [18-07-2019(online)].pdf | 2019-07-18 |
| 6 | 201941028908-FORM 1 [18-07-2019(online)].pdf | 2019-07-18 |
| 7 | 201941028908-DRAWINGS [18-07-2019(online)].pdf | 2019-07-18 |
| 8 | 201941028908-DECLARATION OF INVENTORSHIP (FORM 5) [18-07-2019(online)].pdf | 2019-07-18 |
| 9 | 201941028908-COMPLETE SPECIFICATION [18-07-2019(online)].pdf | 2019-07-18 |
| 10 | Correspondence by Agent_Power of Attorney_22-07-2019.pdf | 2019-07-22 |
| 11 | Abstract 201941028908.jpg | 2019-07-22 |
| 12 | 201941028908-Proof of Right (MANDATORY) [06-01-2020(online)].pdf | 2020-01-06 |
| 13 | 201941028908-Request Letter-Correspondence [22-05-2020(online)].pdf | 2020-05-22 |
| 14 | 201941028908-Power of Attorney [22-05-2020(online)].pdf | 2020-05-22 |
| 15 | 201941028908-Form 1 (Submitted on date of filing) [22-05-2020(online)].pdf | 2020-05-22 |
| 16 | 201941028908-FER_SER_REPLY [16-09-2021(online)].pdf | 2021-09-16 |
| 17 | 201941028908-FER.pdf | 2021-10-17 |
| 18 | 201941028908-US(14)-HearingNotice-(HearingDate-13-06-2024).pdf | 2024-05-10 |
| 19 | 201941028908-REQUEST FOR ADJOURNMENT OF HEARING UNDER RULE 129A [10-06-2024(online)].pdf | 2024-06-10 |
| 20 | 201941028908-FORM 3 [10-06-2024(online)].pdf | 2024-06-10 |
| 21 | 201941028908-US(14)-ExtendedHearingNotice-(HearingDate-10-07-2024).pdf | 2024-06-12 |
| 22 | 201941028908-Correspondence to notify the Controller [08-07-2024(online)].pdf | 2024-07-08 |
| 23 | 201941028908-FORM-26 [09-07-2024(online)].pdf | 2024-07-09 |
| 24 | 201941028908-Written submissions and relevant documents [25-07-2024(online)].pdf | 2024-07-25 |
| 25 | 201941028908-Written submissions and relevant documents [25-07-2024(online)]-1.pdf | 2024-07-25 |
| 26 | 201941028908-PETITION UNDER RULE 137 [25-07-2024(online)].pdf | 2024-07-25 |
| 27 | 201941028908-PatentCertificate09-08-2024.pdf | 2024-08-09 |
| 28 | 201941028908-IntimationOfGrant09-08-2024.pdf | 2024-08-09 |
| 1 | 2021-03-1815-14-26E_18-03-2021.pdf |