Sign In to Follow Application
View All Documents & Correspondence

System And Method For Multivariate Statistical Analysis Of High Dimensional Data

Abstract: System and Method for Multivariate Statistical Analysis of High-Dimensional Data Abstract The present invention relates to a system for multivariate statistical analysis of high-dimensional data. The system includes a data receiver configured to receive high-dimensional data comprising a plurality of variables, and a processor configured to execute instructions for determining a correlation structure among the variables, performing a principal component analysis to identify principal components of the high-dimensional data, determining a contribution of each principal component to the variability of the high-dimensional data, selecting a subset of the principal components based on the contribution of each principal component to the variability of the data, performing a clustering analysis on the subset of principal components to identify clusters of similar data points, determining statistical properties of each cluster, and providing an output based on the statistical properties of each cluster. The system allows for the analysis of complex datasets with multiple variables, providing valuable insights and strategies for various applications, such as marketing, finance, and healthcare. The system can store the high-dimensional data and output in memory and display the output on a display device. The processor can also receive user input for adjusting the parameters used in the analysis, allowing the system to be customized to suit the specific needs of the user. The invention is highly versatile and can be used to analyze any type of high-dimensional data, making it a valuable tool for data analysis in various fields.

Get Free WhatsApp Updates!
Notices, Deadlines & Correspondence

Patent Information

Application #
Filing Date
09 May 2023
Publication Number
23/2023
Publication Type
INA
Invention Field
COMPUTER SCIENCE
Status
Email
Parent Application

Applicants

BANASTHALI VIDYAPITH
BANASTHALI VIDYAPITH, P.O. BANASTHALI, BANASTHALI, RAJASTHAN, INDIA, 304022 JAIPUR

Inventors

1. DR. SHALINI CHANDRA
BANASTHALI VIDYAPITH, P.O. BANASTHALI, BANASTHALI, RAJASTHAN, INDIA, 304022 JAIPUR

Claims

1. A method for multivariate statistical analysis of high-dimensional data, comprising: receiving high-dimensional data comprising a plurality of variables; determining a correlation structure among the plurality of variables; performing a principal component analysis to identify principal components of the high-dimensional data; determining a contribution of each of the principal components to the variability of the high-dimensional data; selecting a subset of the principal components based on the contribution of each of the principal components to the variability of the high-dimensional data; performing a clustering analysis on the subset of the principal components to identify clusters of similar data points; determining statistical properties of each of the clusters; and providing an output based on the statistical properties of each of the clusters.

2. The method of claim 1, wherein determining a correlation structure comprises computing a covariance matrix or a correlation matrix of the plurality of variables.

3. The method of claim 1, wherein selecting a subset of the principal components comprises selecting several principal components that explain a predetermined percentage of the variability of the high-dimensional data.

4. The method of claim 1, wherein performing a clustering analysis comprises using a clustering algorithm selected from the group consisting of k-means, hierarchical clustering, and mixture modelling.

5. The method of claim 1, wherein determining statistical properties comprises determining one or more of mean, median, mode, variance, skewness, and kurtosis of each of the clusters.

6. A system for multivariate statistical analysis of high-dimensional data, comprising: a data receiver configured to receive high-dimensional data comprising a plurality of variables; a processor configured to execute instructions for: determining a correlation structure among the plurality of variables; performing a principal component analysis to identify principal components of the high-dimensional data; determining a contribution of each of the principal components to the variability of the high-dimensional data; selecting a subset of the principal components based on the contribution of each of the principal components to the variability of the high-dimensional data; performing a clustering analysis on the subset of the principal components to identify clusters of similar data points; determining statistical properties of each of the clusters; and providing an output based on the statistical properties of each of the clusters.

7. The system of claim 6, further comprising a memory for storing the high-dimensional data and the output.

8. The system of claim 6, wherein the processor is further configured to display the output on a display device.

9. The system of claim 6, wherein the processor is further configured to receive user input for adjusting the parameters used in the method. System and Method for Multivariate Statistical Analysis of High-Dimensional Data Abstract The present invention relates to a system for multivariate statistical analysis of high-dimensional data. The system includes a data receiver configured to receive high-dimensional data comprising a plurality of variables, and a processor configured to execute instructions for determining a correlation structure among the variables, performing a principal component analysis to identify principal components of the high-dimensional data, determining a contribution of each principal component to the variability of the high-dimensional data, selecting a subset of the principal components based on the contribution of each principal component to the variability of the data, performing a clustering analysis on the subset of principal components to identify clusters of similar data points, determining statistical properties of each cluster, and providing an output based on the statistical properties of each cluster. The system allows for the analysis of complex datasets with multiple variables, providing valuable insights and strategies for various applications, such as marketing, finance, and healthcare. The system can store the high-dimensional data and output in memory and display the output on a display device. The processor can also receive user input for adjusting the parameters used in the analysis, allowing the system to be customized to suit the specific needs of the user. The invention is highly versatile and can be used to analyze any type of high-dimensional data, making it a valuable tool for data analysis in various fields. , Claims: Claims :

1. A method for multivariate statistical analysis of high-dimensional data, comprising: receiving high-dimensional data comprising a plurality of variables; determining a correlation structure among the plurality of variables; performing a principal component analysis to identify principal components of the high-dimensional data; determining a contribution of each of the principal components to the variability of the high-dimensional data; selecting a subset of the principal components based on the contribution of each of the principal components to the variability of the high-dimensional data; performing a clustering analysis on the subset of the principal components to identify clusters of similar data points; determining statistical properties of each of the clusters; and providing an output based on the statistical properties of each of the clusters.

2. The method of claim 1, wherein determining a correlation structure comprises computing a covariance matrix or a correlation matrix of the plurality of variables.

3. The method of claim 1, wherein selecting a subset of the principal components comprises selecting several principal components that explain a predetermined percentage of the variability of the high-dimensional data.

4. The method of claim 1, wherein performing a clustering analysis comprises using a clustering algorithm selected from the group consisting of k-means, hierarchical clustering, and mixture modelling.

5. The method of claim 1, wherein determining statistical properties comprises determining one or more of mean, median, mode, variance, skewness, and kurtosis of each of the clusters.

6. A system for multivariate statistical analysis of high-dimensional data, comprising: a data receiver configured to receive high-dimensional data comprising a plurality of variables; a processor configured to execute instructions for: determining a correlation structure among the plurality of variables; performing a principal component analysis to identify principal components of the high-dimensional data; determining a contribution of each of the principal components to the variability of the high-dimensional data; selecting a subset of the principal components based on the contribution of each of the principal components to the variability of the high-dimensional data; performing a clustering analysis on the subset of the principal components to identify clusters of similar data points; determining statistical properties of each of the clusters; and providing an output based on the statistical properties of each of the clusters.

7. The system of claim 6, further comprising a memory for storing the high-dimensional data and the output.

8. The system of claim 6, wherein the processor is further configured to display the output on a display device.

9. The system of claim 6, wherein the processor is further configured to receive user input for adjusting the parameters used in the method.

Specification

Description:System and Method for Multivariate Statistical Analysis of High-Dimensional Data
Field of the Invention
[0001] The present invention relates to systems and methods for performing multivariate statistical analysis on high-dimensional data. More specifically, the invention provides a novel approach for analyzing complex datasets that involve multiple variables or factors that are potentially interrelated. The invention employs advanced statistical techniques, including machine learning algorithms and feature selection methods, to identify the most relevant variables and relationships in the data and generate meaningful insights and visualizations.
Background
[0002] The background description includes information that may be useful in understanding the present invention. It is not an admission that any of the information provided herein is prior art or relevant to the presently claimed invention, or that any publication specifically or implicitly referenced is prior art.
[0003] With the advent of new technologies and techniques for data collection, large and complex data sets have become increasingly common in various fields, including finance, biology, engineering, and many others. High-dimensional data sets contain a large number of variables or features, and the analysis of such data presents unique challenges for statistical analysis. Traditional statistical methods that work well with low-dimensional data often fail when applied to high-dimensional data. This is because the increase in the number of variables leads to an exponential increase in the number of possible combinations, which makes it difficult to identify patterns and relationships between variables.
[0004] Multivariate statistical analysis is a powerful tool for analyzing high-dimensional data sets. Multivariate statistical analysis involves analyzing the relationships between multiple variables in a data set. This approach can uncover hidden patterns and structures in high-dimensional data that may not be apparent when analyzing individual variables separately. However, the complexity and size of high-dimensional data sets make multivariate analysis challenging.
[0005] The system and method for multivariate statistical analysis of high-dimensional data is a technique that uses multivariate statistical analysis to analyze high-dimensional data sets. The method 100 involves reducing the dimensionality of the data set to a smaller number of variables by using techniques such as principal component analysis, factor analysis, or clustering. This reduction in dimensionality allows for a more manageable data set that can be analyzed using traditional statistical methods.
[0006] Principal component analysis (PCA) is a popular technique used for dimensionality reduction in multivariate analysis. PCA identifies the most important variables that explain the maximum amount of variation in the data set. The technique finds the linear combination of variables that can explain the maximum variance in the data, and then uses these combinations as the new variables for further analysis. This results in a smaller number of variables that capture the most important information in the data set.
[0007] Factor analysis is another technique used for dimensionality reduction in multivariate analysis. Factor analysis assumes that the observed variables in the data set are influenced by a smaller number of underlying factors. The technique identifies these factors by analyzing the correlations between variables and then groups the variables into these factors. The new factors represent the underlying relationships in the data and are used for further analysis.
[0008] Clustering is a technique used for identifying groups of similar observations or variables in the data set. Clustering algorithms group variables together based on their similarity, and the resulting groups can be used for further analysis. Clustering can help identify patterns and relationships in high-dimensional data sets that may not be apparent through other techniques.
[0009] The system and method for multivariate statistical analysis of high-dimensional data has a wide range of applications. In finance, the technique can be used to analyze stock market data to identify patterns and trends in the market. In biology, the technique can be used to analyze gene expression data to identify genes that are co-regulated or that play a role in a particular biological process. In engineering, the technique can be used to analyze sensor data to identify patterns and anomalies in the data that may indicate a malfunction or other problem.
[00010] Thus, the system and method for multivariate statistical analysis of high-dimensional data is a powerful tool for analyzing complex data sets and has the potential to revolutionize many fields by providing new insights and understanding. The technique enables researchers and analysts to identify relationships and patterns in large and complex data sets that may not be apparent through traditional methods, providing a more complete understanding of the data and its underlying processes.
[00011] All references, including publications, patent applications, and patents, cited herein are hereby incorporated by reference to the same extent as if each reference were individually and specifically indicated to be incorporated by reference and were set forth in its entirety herein.
Summary
[00012] Various objects, features, and advantages of the disclosed subject matter can be more fully appreciated with reference to the following detailed description of the disclosed subject matter when considered in connection with the following drawings, in which like reference numerals identify like elements.
[00013] The present invention relates to systems and methods for performing multivariate statistical analysis on high-dimensional data. More specifically, the invention provides a novel approach for analyzing complex datasets that involve multiple variables or factors that are potentially interrelated. The invention employs advanced statistical techniques, including machine learning algorithms and feature selection methods, to identify the most relevant variables and relationships in the data and generate meaningful insights and visualizations.
[00014] The method for multivariate statistical analysis of high-dimensional data is a powerful tool for analyzing large datasets comprising a plurality of variables. The method involves several steps, including receiving high-dimensional data, determining a correlation structure among the variables, performing a principal component analysis to identify principal components, selecting a subset of the principal components, performing a clustering analysis on the subset of principal components, determining statistical properties of each cluster, and providing an output based on the statistical properties of each cluster.
[00015] To determine the correlation structure among the variables, the method involves computing a covariance matrix or a correlation matrix of the variables. The principal component analysis is then performed to identify the principal components of the high-dimensional data. The contribution of each principal component to the variability of the high-dimensional data is then determined, and a subset of the principal components is selected based on the contribution of each principal component to the variability of the data. This can be done by selecting several principal components that explain a predetermined percentage of the variability of the high-dimensional data.
[00016] Once the subset of principal components is selected, a clustering analysis is performed on this subset to identify clusters of similar data points. This can be done using a clustering algorithm such as k-means, hierarchical clustering, or mixture modelling. The statistical properties of each cluster are then determined, which can include determining the mean, median, mode, variance, skewness, and kurtosis of each cluster.
[00017] Based on the statistical properties of each cluster, the method provides an output that can be used to gain insights into complex data and develop targeted strategies for various applications, such as marketing, finance, and healthcare. The method can be used to analyze any type of high-dimensional data, such as customer purchase history, financial data, and medical data.
[00018] In summary, the method for multivariate statistical analysis of high-dimensional data provides a powerful tool for analyzing large datasets with multiple variables. The method involves several steps, including determining a correlation structure, performing a principal component analysis, selecting a subset of principal components, performing a clustering analysis, determining statistical properties of each cluster, and providing an output based on the statistical properties of each cluster. The method can be used for various applications and can provide valuable insights into complex data.
[00019] The system for multivariate statistical analysis of high-dimensional data is a powerful tool for analyzing large datasets with multiple variables. The system includes a data receiver, a processor, and a memory for storing the high-dimensional data and output.
[00020] The processor is configured to execute instructions for determining a correlation structure among the variables, performing a principal component analysis, selecting a subset of the principal components, performing a clustering analysis, determining statistical properties of each cluster, and providing an output based on the statistical properties of each cluster. The processor can also display the output on a display device and receive user input for adjusting the parameters used in the method.
[00021] To determine the correlation structure among the variables, the processor can compute a covariance matrix or a correlation matrix of the variables. The principal component analysis is then performed to identify the principal components of the high-dimensional data. The contribution of each principal component to the variability of the high-dimensional data is then determined, and a subset of the principal components is selected based on the contribution of each principal component to the variability of the data.
[00022] Once the subset of principal components is selected, the processor performs a clustering analysis on this subset to identify clusters of similar data points. This can be done using a clustering algorithm such as k-means, hierarchical clustering, or mixture modelling. The statistical properties of each cluster are then determined, which can include determining the mean, median, mode, variance, skewness, and kurtosis of each cluster.
[00023] Based on the statistical properties of each cluster, the processor provides an output that can be used to gain insights into complex data and develop targeted strategies for various applications, such as marketing, finance, and healthcare. The output can be stored in the memory device for later reference.
[00024] The system is highly versatile and can be used to analyze any type of high-dimensional data, such as customer purchase history, financial data, and medical data. The system can be customized to suit the specific needs of the user by adjusting the parameters used in the method.
[00025] In summary, the system for multivariate statistical analysis of high-dimensional data provides a powerful tool for analyzing large datasets with multiple variables. The system includes a data receiver, a processor, and a memory for storing the high-dimensional data and output. The processor is configured to perform several statistical analyses, including determining a correlation structure, performing a principal component analysis, selecting a subset of principal components, performing a clustering analysis, determining statistical properties of each cluster, and providing an output based on the statistical properties of each cluster. The system is highly versatile and can be customized to suit the specific needs of the user.
Brief Description of the Drawings
[00026] The features and advantages of the present disclosure would be more clearly understood from the following description taken in conjunction with the accompanying drawings in which:
[00027] FIG. 1 shows an exemplary flowchart that outlines the steps involved in method for performing multivariate statistical analysis on high-dimensional data, according to some embodiments of the present disclosure.
[00028] FIG. 2 is diagram that demonstrate configuration of system for multivariate statistical analysis of high-dimensional data, according to some embodiments of the present disclosure.
Detailed Description
[00029] The following is a detailed description of exemplary embodiments to illustrate the principles of the invention. The embodiments are provided to illustrate aspects of the invention, but the invention is not limited to any embodiment. The scope of the invention encompasses numerous alternatives, modifications and equivalent; it is limited only by the claims.
[00030] In view of the many possible embodiments to which the principles of the present discussion may be applied, it should be recognized that the embodiments described herein with respect to the drawing figures are meant to be illustrative only and should not be taken as limiting the scope of the claims. Therefore, the techniques as described herein contemplate all such embodiments as may come within the scope of the following claims and equivalents thereof.
[00031] The detailed description is described with reference to the accompanying figures. In the figures, the left-most digit(s) of a reference number identifies the figure in which the reference number first appears. The use of the same reference numbers in different instances in the description and the figures may indicate similar or identical items.
[00032] The present invention relates to systems and methods for performing multivariate statistical analysis on high-dimensional data. More specifically, the invention provides a novel approach for analyzing complex datasets that involve multiple variables or factors that are potentially interrelated. The invention employs advanced statistical techniques, including machine learning algorithms and feature selection methods, to identify the most relevant variables and relationships in the data and generate meaningful insights and visualizations.
[00033] Fig 1 illustrates a method 100 for multivariate statistical analysis of high-dimensional data involves (at step 102) receiving a dataset comprising a plurality of variables, each variable representing a different aspect of the data. The method then involves (at step 104) determining the correlation structure among the variables to identify which variables are most strongly correlated with one another. Next, (at step 106) a principal component analysis is performed to identify principal components of the high-dimensional data. The principal components are combinations of the original variables that explain the most variance in the data. The contribution of each principal component to the variability of the high-dimensional data is determined (at step 108). Based on this, a subset of the principal components is selected (at step 110), which represents the most important aspects of the data.
[00034] Referring to the preceding embodiment, next, (at step 112) a clustering analysis is performed on the subset of principal components to identify clusters of similar data points. Each cluster is characterized by its statistical properties (at step 114) , such as mean, variance, and distribution. Finally, an output (at step 116) is provided based on the statistical properties of each cluster, which can be used to gain insights into the data and make informed decisions. For example, this method can be applied in medical research to analyse high-dimensional data, such as gene expression data, to identify clusters of patients with similar genetic profiles and determine their response to treatment. It can also be used in finance to analyse stock prices, exchange rates, and commodity prices to identify clusters of investments with similar risk and return profiles.
[00035] In an embodiment, determining a correlation structure among the plurality of variables typically involves computing a covariance matrix or a correlation matrix of the variables. The covariance matrix represents the covariances between all possible pairs of variables in the dataset. The covariance between two variables measures how much they vary together, while the correlation between two variables measures the strength of their linear relationship, regardless of their scale. For example, in a study on housing prices, the dataset may include variables such as the size of the house, the number of rooms, the location, and the age of the house. To determine the correlation structure among these variables, a covariance matrix or a correlation matrix can be computed. The covariance matrix will represent how much the variables vary together, while the correlation matrix will represent the strength of their linear relationship.
[00036] In an embodiment, after computing the covariance or correlation matrix, the method may involve further analysis, such as performing a factor analysis or a principal component analysis, to identify the underlying factors or components that explain the most variance in the data. This can help to reduce the dimensionality of the data and focus on the most important aspects. Thus, determining the correlation structure of the variables is an important step in multivariate statistical analysis, as it helps to identify the relationships and dependencies among the variables, which can aid in interpreting the results and making informed decisions.
[00037] Referring to the preceding embodiment, selecting a subset of the principal components involves choosing the most important components that explain the majority of the variability in the high-dimensional data. One common approach is to select a predetermined percentage of the total variability in the data, such as 80% or 90%, and choose the principal components that explain this percentage of the variability. For example, in a study of customer behavior, the high-dimensional data may include variables such as purchase history, demographics, and web browsing history. After performing a principal component analysis to identify the most important components of the data, the method may involve selecting the top principal components that explain, for instance, 80% of the variability in the data. This would involve choosing a subset of the principal components that capture the majority of the underlying patterns and relationships in the data.
[00038] Referring to the preceding embodiment, by selecting a subset of the principal components that explain a large percentage of the variability, the method can effectively reduce the dimensionality of the data and focus on the most important aspects. This can make it easier to interpret the results and gain insights into the data. Thus, selecting a subset of the principal components is an important step in multivariate statistical analysis, as it helps to identify the most important components of the data and focus on the most relevant information.
[00039] Fig 2 illustrates a system 200 for multivariate statistical analysis of high-dimensional data is described herein, which can be used to analyse large datasets comprising a plurality of variables. The system 200 comprises a data receiver 202 configured to receive high-dimensional data, a processor 204 configured to execute instructions for performing various statistical analyses, and a memory 206 for storing the high-dimensional data and the output.
[00040] In an embodiment, the data receiver can be any suitable device capable of receiving high-dimensional data, such as a computer or a server. The high-dimensional data can include any number of variables, and can be in any suitable format, such as a spreadsheet or a database.
[00041] In an embodiment, the processor is configured to execute instructions for determining a correlation structure among the plurality of variables in the high-dimensional data. This can be done using any suitable statistical method, such as Pearson correlation or Spearman correlation. Once the correlation structure is determined, the processor can perform a principal component analysis to identify principal components of the high-dimensional data. This can be done using any suitable method, such as the singular value decomposition (SVD) method.
[00042] In an embodiment, the processor determines the contribution of each of the principal components to the variability of the high-dimensional data. This can be done using any suitable method, such as the eigenvalue method. Based on the contribution of each principal component, the processor selects a subset of the principal components to analyse further. This subset can be selected based on any suitable criteria, such as the top n principal components or the principal components with the highest contributions to the variability of the high-dimensional data.
[00043] In an embodiment, the processor then performs a clustering analysis on the subset of the principal components to identify clusters of similar data points. This can be done using any suitable method, such as the k-means clustering method or the hierarchical clustering method. Once the clusters are identified, the processor determines statistical properties of each of the clusters. This can include, for example, the mean and standard deviation of each variable within each cluster. Finally, the processor provides an output based on the statistical properties of each of the clusters. This output can be displayed on a display device, such as a computer monitor or a mobile device. The output can include, for example, graphs or charts showing the statistical properties of each cluster, as well as any other relevant information, such as the number of data points in each cluster.
[00044] In an embodiment, the system can further comprise a memory for storing the high-dimensional data and the output. This can include any suitable memory device, such as a hard drive or a cloud-based storage system. The system can also be configured to receive user input for adjusting the parameters used in the method. For example, the user can adjust the number of principal components selected for analysis or the clustering method used.
[00045] Another embodiment of the system for multivariate statistical analysis of high-dimensional data is described herein. In this embodiment, the system comprises a data receiver, a processor, and a memory for storing the high-dimensional data and the output. The processor is configured to execute instructions for performing various statistical analyses, as described below. The data receiver can receive high-dimensional data from any suitable source, such as a sensor or a database. The high-dimensional data can comprise any number of variables, and can be in any suitable format, such as a spreadsheet or a database.
[00046] In an embodiment, the processor is configured to execute instructions for determining a correlation structure among the plurality of variables in the high-dimensional data. This can be done using any suitable statistical method, such as the Pearson correlation or the Spearman correlation. Once the correlation structure is determined, the processor performs a principal component analysis to identify principal components of the high-dimensional data. This can be done using any suitable method, such as the singular value decomposition (SVD) method.
[00047] In an embodiment, after identifying the principal components, the processor determines the contribution of each of the principal components to the variability of the high-dimensional data. This can be done using any suitable method, such as the eigenvalue method. Based on the contribution of each principal component, the processor selects a subset of the principal components to analyse further. This subset can be selected based on any suitable criteria, such as the top n principal components or the principal components with the highest contributions to the variability of the high-dimensional data.
[00048] In an embodiment, the processor then performs a clustering analysis on the subset of the principal components to identify clusters of similar data points. This can be done using any suitable method, such as the k-means clustering method or the hierarchical clustering method. Once the clusters are identified, the processor determines statistical properties of each of the clusters. This can include, for example, the mean and standard deviation of each variable within each cluster, as well as any other relevant statistical measures. Finally, the processor provides an output based on the statistical properties of each of the clusters. This output can be displayed on a display device, such as a computer monitor or a mobile device. The output can include, for example, graphs or charts showing the statistical properties of each cluster, as well as any other relevant information, such as the number of data points in each cluster.
[00049] In addition to the above, the system can also comprise a memory for storing the high-dimensional data and the output. The memory can be any suitable memory device, such as a hard drive or a cloud-based storage system. The processor can also be configured to receive user input for adjusting the parameters used in the method. For example, the user can adjust the number of principal components selected for analysis or the clustering method used. This allows the user to customize the analysis to their specific needs and requirements.
[00050] In an embodiment, to illustrate the use of the system, consider an example in which the system is used to analyse a dataset of customer purchase history for a retail store. The dataset includes information on customer demographics (age, gender, income), purchase history (items purchased, price, quantity), and time of purchase. Using the system, the correlation structure among the variables is first determined, which allows the identification of the principal components of the dataset. The contribution of each principal component to the variability of the dataset is then determined, and a subset of the principal components is selected for further analysis.
[00051] In an embodiment, using clustering analysis, the system identifies clusters of customers with similar purchase histories. For example, one cluster may consist of customers who frequently purchase high-end electronics, while another cluster may consist of customers who primarily purchase low-cost clothing items. The statistical properties of each cluster are then determined, allowing the retailer to gain insights into the purchasing habits of their customers.
[00052] Based on the output of the system, the retailer can develop targeted marketing strategies for each cluster of customers. For example, the retailer may offer discounts on high-end electronics to the cluster of customers who frequently purchase these items, while offering discounts on clothing to the cluster of customers who primarily purchase low-cost clothing items.
[00053] Thus, the system for multivariate statistical analysis of high-dimensional data described herein provides a powerful tool for analysing large datasets comprising a plurality of variables. The system allows the identification of principal components of the dataset, the clustering of similar data points, and the determination of statistical properties of each cluster. The output of the system can be used to gain insights into complex data and develop targeted strategies for various applications, such as marketing, finance, and healthcare.
[00054] Further, while operations are depicted in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Likewise, while several specific implementation details are contained in the above discussions, these should not be construed as limitations on the scope of the subject matter described herein, but rather as descriptions of features that may be specific to particular embodiments. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable sub-combination. Moreover, although features may be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a sub-combination or variation of a sub-combination.
[00055] The term “memory,” as used herein relates to a volatile or persistent medium, such as a magnetic disk, or optical disk, in which a computer can store data or software for any duration. Optionally, the memory is non-volatile mass storage such as physical storage media. Furthermore, a single memory may encompass and in a scenario wherein computing system is distributed, the processing, memory and/or storage capability may be distributed as well.
[00056] Throughout the present disclosure, the term ‘server’ relates to a structure and/or module that include programmable and/or non-programmable components configured to store, process and/or share information. Optionally, the server includes any arrangement of physical or virtual computational entities capable of enhancing information to perform various computational tasks.
[00057] Throughout the present disclosure, the term “network” relates to an arrangement of interconnected programmable and/or non-programmable components that are configured to facilitate data communication between one or more electronic devices and/or databases, whether available or known at the time of filing or as later developed. Furthermore, the network may include, but is not limited to, one or more peer-to-peer network, a hybrid peer-to-peer network, local area networks (LANs), radio access networks (RANs), metropolitan area networks (MANS), wide area networks (WANs), all or a portion of a public network such as the global computer network known as the Internet, a private network, a cellular network and any other communication system or systems at one or more locations.
[00058] Throughout the present disclosure, the term “process”* relates to any collection or set of instructions executable by a computer or other digital system so as to configure the computer or the digital system to perform a task that is the intent of the process.
[00059] Throughout the present disclosure, the term ‘Artificial intelligence (AI)’ as used herein relates to any mechanism or computationally intelligent system that combines knowledge, techniques, and methodologies for controlling a bot or other element within a computing environment. Furthermore, the artificial intelligence (AI) is configured to apply knowledge and that can adapt it-self and learn to do better in changing environments. Additionally, employing any computationally intelligent technique, the artificial intelligence (AI) is operable to adapt to unknown or changing environment for better performance. The artificial intelligence (AI) includes fuzzy logic engines, decision-making engines, preset targeting accuracy levels, and/or programmatically intelligent software.

Claims
I/We Claim:
1. A method for multivariate statistical analysis of high-dimensional data, comprising:
receiving high-dimensional data comprising a plurality of variables;
determining a correlation structure among the plurality of variables;
performing a principal component analysis to identify principal components of the high-dimensional data;
determining a contribution of each of the principal components to the variability of the high-dimensional data;
selecting a subset of the principal components based on the contribution of each of the principal components to the variability of the high-dimensional data;
performing a clustering analysis on the subset of the principal components to identify clusters of similar data points;
determining statistical properties of each of the clusters; and
providing an output based on the statistical properties of each of the clusters.
2. The method of claim 1, wherein determining a correlation structure comprises computing a covariance matrix or a correlation matrix of the plurality of variables.
3. The method of claim 1, wherein selecting a subset of the principal components comprises selecting several principal components that explain a predetermined percentage of the variability of the high-dimensional data.
4. The method of claim 1, wherein performing a clustering analysis comprises using a clustering algorithm selected from the group consisting of k-means, hierarchical clustering, and mixture modelling.
5. The method of claim 1, wherein determining statistical properties comprises determining one or more of mean, median, mode, variance, skewness, and kurtosis of each of the clusters.
6. A system for multivariate statistical analysis of high-dimensional data, comprising:
a data receiver configured to receive high-dimensional data comprising a plurality of variables;
a processor configured to execute instructions for:
determining a correlation structure among the plurality of variables;
performing a principal component analysis to identify principal components of the high-dimensional data;
determining a contribution of each of the principal components to the variability of the high-dimensional data;
selecting a subset of the principal components based on the contribution of each of the principal components to the variability of the high-dimensional data;
performing a clustering analysis on the subset of the principal components to identify clusters of similar data points;
determining statistical properties of each of the clusters; and
providing an output based on the statistical properties of each of the clusters.
7. The system of claim 6, further comprising a memory for storing the high-dimensional data and the output.
8. The system of claim 6, wherein the processor is further configured to display the output on a display device.
9. The system of claim 6, wherein the processor is further configured to receive user input for adjusting the parameters used in the method.

System and Method for Multivariate Statistical Analysis of High-Dimensional Data
Abstract
The present invention relates to a system for multivariate statistical analysis of high-dimensional data. The system includes a data receiver configured to receive high-dimensional data comprising a plurality of variables, and a processor configured to execute instructions for determining a correlation structure among the variables, performing a principal component analysis to identify principal components of the high-dimensional data, determining a contribution of each principal component to the variability of the high-dimensional data, selecting a subset of the principal components based on the contribution of each principal component to the variability of the data, performing a clustering analysis on the subset of principal components to identify clusters of similar data points, determining statistical properties of each cluster, and providing an output based on the statistical properties of each cluster. The system allows for the analysis of complex datasets with multiple variables, providing valuable insights and strategies for various applications, such as marketing, finance, and healthcare. The system can store the high-dimensional data and output in memory and display the output on a display device. The processor can also receive user input for adjusting the parameters used in the analysis, allowing the system to be customized to suit the specific needs of the user. The invention is highly versatile and can be used to analyze any type of high-dimensional data, making it a valuable tool for data analysis in various fields. , Claims:
Claims
I/We Claim:
1. A method for multivariate statistical analysis of high-dimensional data, comprising:
receiving high-dimensional data comprising a plurality of variables;
determining a correlation structure among the plurality of variables;
performing a principal component analysis to identify principal components of the high-dimensional data;
determining a contribution of each of the principal components to the variability of the high-dimensional data;
selecting a subset of the principal components based on the contribution of each of the principal components to the variability of the high-dimensional data;
performing a clustering analysis on the subset of the principal components to identify clusters of similar data points;
determining statistical properties of each of the clusters; and
providing an output based on the statistical properties of each of the clusters.
2. The method of claim 1, wherein determining a correlation structure comprises computing a covariance matrix or a correlation matrix of the plurality of variables.
3. The method of claim 1, wherein selecting a subset of the principal components comprises selecting several principal components that explain a predetermined percentage of the variability of the high-dimensional data.
4. The method of claim 1, wherein performing a clustering analysis comprises using a clustering algorithm selected from the group consisting of k-means, hierarchical clustering, and mixture modelling.
5. The method of claim 1, wherein determining statistical properties comprises determining one or more of mean, median, mode, variance, skewness, and kurtosis of each of the clusters.
6. A system for multivariate statistical analysis of high-dimensional data, comprising:
a data receiver configured to receive high-dimensional data comprising a plurality of variables;
a processor configured to execute instructions for:
determining a correlation structure among the plurality of variables;
performing a principal component analysis to identify principal components of the high-dimensional data;
determining a contribution of each of the principal components to the variability of the high-dimensional data;
selecting a subset of the principal components based on the contribution of each of the principal components to the variability of the high-dimensional data;
performing a clustering analysis on the subset of the principal components to identify clusters of similar data points;
determining statistical properties of each of the clusters; and
providing an output based on the statistical properties of each of the clusters.
7. The system of claim 6, further comprising a memory for storing the high-dimensional data and the output.
8. The system of claim 6, wherein the processor is further configured to display the output on a display device.
9. The system of claim 6, wherein the processor is further configured to receive user input for adjusting the parameters used in the method.

Documents

Application Documents

# Name Date
1 202311032857-REQUEST FOR EARLY PUBLICATION(FORM-9) [09-05-2023(online)].pdf 2023-05-09
2 202311032857-POWER OF AUTHORITY [09-05-2023(online)].pdf 2023-05-09
3 202311032857-OTHERS [09-05-2023(online)].pdf 2023-05-09
4 202311032857-FORM-9 [09-05-2023(online)].pdf 2023-05-09
5 202311032857-FORM FOR SMALL ENTITY(FORM-28) [09-05-2023(online)].pdf 2023-05-09
6 202311032857-FORM 1 [09-05-2023(online)].pdf 2023-05-09
7 202311032857-EVIDENCE FOR REGISTRATION UNDER SSI(FORM-28) [09-05-2023(online)].pdf 2023-05-09
8 202311032857-EDUCATIONAL INSTITUTION(S) [09-05-2023(online)].pdf 2023-05-09
9 202311032857-DRAWINGS [09-05-2023(online)].pdf 2023-05-09
10 202311032857-DECLARATION OF INVENTORSHIP (FORM 5) [09-05-2023(online)].pdf 2023-05-09
11 202311032857-COMPLETE SPECIFICATION [09-05-2023(online)].pdf 2023-05-09