Sign In to Follow Application
View All Documents & Correspondence

Multivariate Analysis In Big Data Analytics And Machine Learning

Abstract: Multivariate Analysis in Big Data Analytics and Machine Learning Abstract This invention describes a system for multivariate analysis in big data analytics and machine learning. The system comprises a computer processor, memory, a dataset stored in memory, and program instructions stored in memory. The program instructions include code for applying principal component analysis (PCA) to the dataset to reduce its dimensionality, code for applying a machine learning algorithm to the reduced dataset to generate a model, and code for applying the model to one or more additional datasets to make predictions or classifications. The system can be applied in various fields such as finance, healthcare, and customer segmentation. By using PCA to reduce the dimensionality of the dataset and applying machine learning algorithms to generate models, the system can extract meaningful insights and make accurate predictions, revolutionizing the way we analyze and utilize big data.

Get Free WhatsApp Updates!
Notices, Deadlines & Correspondence

Patent Information

Application #
Filing Date
09 May 2023
Publication Number
25/2023
Publication Type
INA
Invention Field
COMPUTER SCIENCE
Status
Email
Parent Application

Applicants

BANASTHALI VIDYAPITH
BANASTHALI VIDYAPITH, P.O. BANASTHALI, BANASTHALI, RAJASTHAN, INDIA, 304022 JAIPUR

Inventors

1. DR. GARGI TYAGI
BANASTHALI VIDYAPITH, P.O. BANASTHALI, BANASTHALI, RAJASTHAN, INDIA, 304022 JAIPUR

Claims

1. A method for multivariate analysis in big data analytics and machine learning, comprising: receiving a dataset comprising a plurality of variables; applying principal component analysis (PCA) to the dataset to reduce the dimensionality of the dataset; applying a machine learning algorithm to the reduced dataset to generate a model; and applying the model to one or more additional datasets to make predictions or classifications.

2. The method of claim 1, wherein the machine learning algorithm is a supervised learning algorithm.

3. The method of claim 1, wherein the machine learning algorithm is an unsupervised learning algorithm.

4. The method of claim 1, wherein the dataset is pre-processed prior to applying PCA, including normalization, scaling, or imputation.

5. The method of claim 1, wherein the PCA is performed using an iterative algorithm.

6. The method of claim 1, wherein the model generated by the machine learning algorithm is a regression model.

7. The method of claim 1, wherein the model generated by the machine learning algorithm is a classification model.

8. A system for multivariate analysis in big data analytics and machine learning, comprising: a computer processor; memory; a dataset stored in memory; program instructions stored in memory, wherein the program instructions comprising: code for applying principal component analysis (PCA) to the dataset to reduce the dimensionality of the dataset; code for applying a machine learning algorithm to the reduced dataset to generate a model; and code for applying the model to one or more additional datasets to make predictions or classifications. Multivariate Analysis in Big Data Analytics and Machine Learning Abstract This invention describes a system for multivariate analysis in big data analytics and machine learning. The system comprises a computer processor, memory, a dataset stored in memory, and program instructions stored in memory. The program instructions include code for applying principal component analysis (PCA) to the dataset to reduce its dimensionality, code for applying a machine learning algorithm to the reduced dataset to generate a model, and code for applying the model to one or more additional datasets to make predictions or classifications. The system can be applied in various fields such as finance, healthcare, and customer segmentation. By using PCA to reduce the dimensionality of the dataset and applying machine learning algorithms to generate models, the system can extract meaningful insights and make accurate predictions, revolutionizing the way we analyze and utilize big data. , Claims:Claims :

1. A method for multivariate analysis in big data analytics and machine learning, comprising: receiving a dataset comprising a plurality of variables; applying principal component analysis (PCA) to the dataset to reduce the dimensionality of the dataset; applying a machine learning algorithm to the reduced dataset to generate a model; and applying the model to one or more additional datasets to make predictions or classifications.

2. The method of claim 1, wherein the machine learning algorithm is a supervised learning algorithm.

3. The method of claim 1, wherein the machine learning algorithm is an unsupervised learning algorithm.

4. The method of claim 1, wherein the dataset is pre-processed prior to applying PCA, including normalization, scaling, or imputation.

5. The method of claim 1, wherein the PCA is performed using an iterative algorithm.

6. The method of claim 1, wherein the model generated by the machine learning algorithm is a regression model.

7. The method of claim 1, wherein the model generated by the machine learning algorithm is a classification model.

8. A system for multivariate analysis in big data analytics and machine learning, comprising: a computer processor; memory; a dataset stored in memory; program instructions stored in memory, wherein the program instructions comprising: code for applying principal component analysis (PCA) to the dataset to reduce the dimensionality of the dataset; code for applying a machine learning algorithm to the reduced dataset to generate a model; and code for applying the model to one or more additional datasets to make predictions or classifications.

Specification

Description:Multivariate Analysis in Big Data Analytics and Machine Learning
Field of the Invention
[0001] The present invention relates to multivariate analysis in big data analytics and machine learning. More specifically, the invention provides a novel approach for analyzing and modeling large, complex data sets using advanced multivariate analysis techniques.
Background
[0002] The background description includes information that may be useful in understanding the present invention. It is not an admission that any of the information provided herein is prior art or relevant to the presently claimed invention, or that any publication specifically or implicitly referenced is prior art.
[0003] Multivariate analysis is a statistical technique that involves analyzing multiple variables simultaneously to uncover relationships and patterns within large datasets. In the context of big data analytics and machine learning, multivariate analysis is an essential tool for discovering hidden patterns and trends in complex datasets with many variables.
[0004] The goal of multivariate analysis is to identify relationships between variables, and to understand how changes in one variable affect the behavior of the others. This can be done using a variety of techniques, including principal component analysis (PCA), factor analysis, cluster analysis, discriminant analysis, and canonical correlation analysis.
[0005] Principal component analysis (PCA) is a popular technique for reducing the dimensionality of large datasets, while retaining as much of the original information as possible. PCA works by identifying the most important variables in the dataset, and projecting the data onto a new set of variables that capture the most important patterns and trends.
[0006] Factor analysis is a technique for identifying underlying factors that explain the correlations between the variables in the dataset. Factor analysis is used to reduce the number of variables in the dataset by identifying the most important factors that explain the variation in the data.
[0007] Cluster analysis is a technique for grouping similar observations together based on their similarity or distance from each other in the dataset. Cluster analysis is useful for identifying patterns and relationships between variables that may not be immediately apparent from the raw data.
[0008] Discriminant analysis is a technique for identifying the variables that best discriminate between different groups in the dataset. Discriminant analysis is often used in machine learning to classify observations into different categories based on their characteristics.
[0009] Canonical correlation analysis is a technique for identifying the linear relationships between two sets of variables. Canonical correlation analysis is useful for identifying the most important relationships between variables in large datasets with many variables.
[00010] Multivariate analysis is an essential tool in big data analytics and machine learning, as it allows analysts to discover hidden patterns and relationships in large and complex datasets. By reducing the dimensionality of the data and identifying the most important variables and factors, multivariate analysis can help analysts to make more informed decisions and predictions based on the data.
[00011] All references, including publications, patent applications, and patents, cited herein are hereby incorporated by reference to the same extent as if each reference were individually and specifically indicated to be incorporated by reference and were set forth in its entirety herein.
Summary
[00012] Various objects, features, and advantages of the disclosed subject matter can be more fully appreciated with reference to the following detailed description of the disclosed subject matter when considered in connection with the following drawings, in which like reference numerals identify like elements.
[00013] The present invention relates to multivariate analysis in big data analytics and machine learning. More specifically, the invention provides a novel approach for analyzing and modeling large, complex data sets using advanced multivariate analysis techniques.
[00014] The method described in this patent application involves using principal component analysis (PCA) in conjunction with machine learning algorithms to analyze large and complex datasets. The method involves receiving a dataset with a large number of variables, and then applying PCA to reduce the dimensionality of the dataset. By reducing the number of variables, PCA simplifies the dataset, making it easier to analyze.
[00015] Once PCA has been applied, a machine learning algorithm is used to generate a model from the reduced dataset. The type of machine learning algorithm used can be either supervised or unsupervised, depending on the nature of the analysis being performed. Supervised learning algorithms are used when the desired output is known, such as in regression analysis or classification. Unsupervised learning algorithms, on the other hand, are used when the desired output is not known, such as in clustering analysis.
[00016] Prior to applying PCA, the dataset may need to be pre-processed. This can involve normalization, scaling, or imputation, depending on the nature of the data. For example, if the data is skewed, normalization may be required to make the data more symmetrical. Similarly, if the data has a wide range of values, scaling may be required to bring the data into a common range.
[00017] PCA is performed using an iterative algorithm, which involves a series of linear transformations that reduce the dimensionality of the data by identifying the most important variables. This process helps to eliminate redundancy and noise, which can improve the accuracy of the analysis. The model generated by the machine learning algorithm can be either a regression model or a classification model, depending on the type of analysis being performed. Regression models are used when the goal is to predict a continuous outcome, such as the price of a stock or the number of sales for a particular product. Classification models, on the other hand, are used when the goal is to categorize data into discrete groups, such as determining whether a customer is likely to purchase a product or not.
[00018] Once the model has been generated, it can be applied to one or more additional datasets to make predictions or classifications. This allows the method to be used in a wide range of applications, such as predicting stock prices, analyzing customer behavior, or identifying patterns in large datasets. Overall, this method provides a powerful tool for analyzing large and complex datasets, allowing researchers and analysts to extract meaningful insights and make accurate predictions.
[00019] This invention describes a system for multivariate analysis in big data analytics and machine learning. The system comprises a computer processor, memory, a dataset stored in memory, and program instructions stored in memory. The program instructions include code for applying principal component analysis (PCA) to the dataset to reduce its dimensionality, code for applying a machine learning algorithm to the reduced dataset to generate a model, and code for applying the model to one or more additional datasets to make predictions or classifications.
[00020] The system is designed to handle large and complex datasets, which can be difficult to analyze using traditional statistical methods. By using PCA to reduce the dimensionality of the dataset, the system simplifies the data and makes it easier to analyze. The machine learning algorithm is then used to generate a model from the reduced dataset. The type of machine learning algorithm used can be either supervised or unsupervised, depending on the nature of the analysis being performed.
[00021] The system also includes code for pre-processing the dataset prior to applying PCA. This can involve normalization, scaling, or imputation, depending on the nature of the data. For example, if the data is skewed, normalization may be required to make the data more symmetrical. Similarly, if the data has a wide range of values, scaling may be required to bring the data into a common range.
[00022] PCA is performed using an iterative algorithm, which involves a series of linear transformations that reduce the dimensionality of the data by identifying the most important variables. This process helps to eliminate redundancy and noise, which can improve the accuracy of the analysis.The model generated by the machine learning algorithm can be either a regression model or a classification model, depending on the type of analysis being performed. Regression models are used when the goal is to predict a continuous outcome, such as the price of a stock or the number of sales for a particular product. Classification models, on the other hand, are used when the goal is to categorize data into discrete groups, such as determining whether a customer is likely to purchase a product or not.
[00023] Finally, the system includes code for applying the model to one or more additional datasets to make predictions or classifications. This allows the system to be used in a wide range of applications, such as predicting stock prices, analyzing customer behavior, or identifying patterns in large datasets. Overall, this system provides a powerful tool for analyzing large and complex datasets, allowing researchers and analysts to extract meaningful insights and make accurate predictions.
Brief Description of the Drawings
[00024] The features and advantages of the present disclosure would be more clearly understood from the following description taken in conjunction with the accompanying drawings in which:
[00025] FIG. 1 shows an exemplary flowchart exemplifying a method for multivariate analysis in big data analytics and machine learning, according to some embodiments of the present disclosure.
[00026] FIG. 2 denotes a representative system for multivariate analysis in big data analytics and machine learning, according to some embodiments of the present disclosure.
Detailed Description
[00027] The following is a detailed description of exemplary embodiments to illustrate the principles of the invention. The embodiments are provided to illustrate aspects of the invention, but the invention is not limited to any embodiment. The scope of the invention encompasses numerous alternatives, modifications and equivalent; it is limited only by the claims.
[00028] In view of the many possible embodiments to which the principles of the present discussion may be applied, it should be recognized that the embodiments described herein with respect to the drawing figures are meant to be illustrative only and should not be taken as limiting the scope of the claims. Therefore, the techniques as described herein contemplate all such embodiments as may come within the scope of the following claims and equivalents thereof.
[00029] The detailed description is described with reference to the accompanying figures. In the figures, the left-most digit(s) of a reference number identifies the figure in which the reference number first appears. The use of the same reference numbers in different instances in the description and the figures may indicate similar or identical items.
[00030] The present invention relates to multivariate analysis in big data analytics and machine learning. More specifically, the invention provides a novel approach for analyzing and modeling large, complex data sets using advanced multivariate analysis techniques.
[00031] FIG. 1 showcases a method 100 for multivariate analysis in big data analytics and machine learning, comprising steps such as receiving (at step 102) a dataset comprising a plurality of variables; applying (at step 104) principal component analysis (PCA) to the dataset to reduce the dimensionality of the dataset; applying (at step 106) a machine learning algorithm to the reduced dataset to generate a model; and applying (at step 108) the model to one or more additional datasets to make predictions or classifications.
[00032] In an embodiment, the method 100 for multivariate analysis could be used for medical data analysis, such as predicting disease outcomes or identifying risk factors for certain conditions. The dataset could include a large number of variables related to a patient's medical history, lifestyle, and genetic makeup. Principal component analysis (PCA) could be applied to the dataset to identify the most important variables and reduce the dimensionality of the dataset. A machine learning algorithm, such as logistic regression or a decision tree, could be applied to the reduced dataset to generate a model that can predict disease outcomes or identify risk factors. The model could then be applied to one or more additional datasets to make predictions or classifications, such as identifying patients at risk for a certain disease or predicting the success of a particular treatment.
[00033] In an embodiment, the method for multivariate analysis could also be used for financial data analysis, such as predicting stock prices or identifying investment opportunities. The dataset could include a large number of variables related to financial metrics, such as stock prices, economic indicators, and company financial statements. PCA could be applied to the dataset to identify the most important variables and reduce the dimensionality of the dataset. A machine learning algorithm, such as a random forest or neural network, could be applied to the reduced dataset to generate a model that can predict stock prices or identify investment opportunities. The model could then be applied to one or more additional datasets to make predictions or classifications, such as predicting the future performance of a particular stock or identifying undervalued companies.
[00034] In an embodiment, the method for multivariate analysis could also be used for customer behaviour analysis, such as predicting customer churn or identifying customer preferences. The dataset could include a large number of variables related to customer behaviour, such as purchase history, demographic information, and website usage patterns. PCA could be applied to the dataset to identify the most important variables and reduce the dimensionality of the dataset. A machine learning algorithm, such as a support vector machine or k-nearest neighbours, could be applied to the reduced dataset to generate a model that can predict customer churn or identify customer preferences. The model could then be applied to one or more additional datasets to make predictions or classifications, such as identifying customers at risk for churn or recommending products based on customer preferences.
[00035] In an embodiment, the method for multivariate analysis could be used for marketing analysis, such as predicting customer responses to marketing campaigns or identifying the most effective marketing channels. The dataset could include a large number of variables related to customer demographics, purchase history, and marketing campaign data. PCA could be applied to the dataset to identify the most important variables and reduce the dimensionality of the dataset. A machine learning algorithm, such as a decision tree or logistic regression, could be applied to the reduced dataset to generate a model that can predict customer responses to marketing campaigns or identify the most effective marketing channels. The model could then be applied to one or more additional datasets to make predictions or classifications, such as optimizing marketing campaigns or identifying new target markets.
[00036] In an embodiment, the method for multivariate analysis could also be used for image recognition, such as identifying objects in images or detecting anomalies in medical images. The dataset could include a large number of variables related to the pixels in the images. PCA could be applied to the dataset to identify the most important variables and reduce the dimensionality of the dataset. A machine learning algorithm, such as a convolutional neural network or decision tree, could be applied to the reduced dataset to generate a model that can identify objects in images or detect anomalies. The model could then be applied to one or more additional datasets to make predictions or classifications, such as identifying objects in real-time or detecting anomalies in medical images.
[00037] In an embodiment, the method for multivariate analysis could be used for fraud detection, such as identifying fraudulent transactions or activity. The dataset could include a large number of variables related to financial transactions, such as transaction amounts, location, and frequency. PCA could be applied to the dataset to identify the most important variables and reduce the dimensionality of the dataset. A machine learning algorithm, such as a support vector machine or decision tree, could be applied to the reduced dataset to generate a model that can identify fraudulent transactions or activity. The model could then be applied to one or more additional datasets to make predictions or classifications, such as identifying fraudulent transactions in real-time or detecting patterns of fraudulent activity.
[00038] FIG 2. describes a system 200 for predicting stock prices using multivariate analysis in big data analytics and machine learning. The system comprises a computer processor 202, memory 204, a dataset 206 stored in memory, and program instructions 208 stored in memory.
[00039] The dataset consists of various financial variables such as company financial reports, market trends, and economic indicators. The program instructions include code for applying PCA to the dataset to reduce its dimensionality and code for applying a supervised machine learning algorithm to generate a regression model.
[00040] The system first pre-processes the dataset by normalizing the variables to make them comparable. Then, PCA is applied to identify the most important variables that contribute to the variation in the dataset. The reduced dataset is then used to generate a regression model using a supervised machine learning algorithm.
[00041] The system can then apply the model to new datasets to predict the stock prices of various companies. For example, if a user inputs a new dataset with company financial reports, market trends, and economic indicators for a particular company, the system can predict its future stock prices based on the regression model generated from the previous dataset.
[00042] This embodiment describes a system for customer segmentation using multivariate analysis in big data analytics and machine learning. The system comprises a computer processor, memory, a dataset stored in memory, and program instructions stored in memory. The dataset consists of various customer data such as demographics, purchase history, and social media activity. The program instructions include code for applying PCA to the dataset to reduce its dimensionality and code for applying an unsupervised machine learning algorithm to generate a clustering model.
[00043] The system first pre-processes the dataset by imputing missing values and scaling the variables. Then, PCA is applied to identify the most important variables that contribute to the variation in the dataset. The reduced dataset is then used to generate a clustering model using an unsupervised machine learning algorithm.
[00044] The system can then apply the model to new datasets to segment customers based on their characteristics. For example, if a user inputs a new dataset with customer data, the system can segment the customers into different groups based on their demographics, purchase history, and social media activity. This information can be used for targeted marketing campaigns or personalized recommendations.
[00045] This embodiment describes a system for fraud detection using multivariate analysis in big data analytics and machine learning. The system comprises a computer processor, memory, a dataset stored in memory, and program instructions stored in memory.
[00046] The dataset consists of various transaction data such as purchase amounts, transaction times, and user locations. The program instructions include code for applying PCA to the dataset to reduce its dimensionality and code for applying a supervised machine learning algorithm to generate a classification model.
[00047] The system first pre-processes the dataset by normalizing the variables and removing outliers. Then, PCA is applied to identify the most important variables that contribute to the variation in the dataset. The reduced dataset is then used to generate a classification model using a supervised machine learning algorithm.
[00048] The system can then apply the model to new datasets to detect fraud in transactions. For example, if a user inputs a new dataset with transaction data, the system can classify the transactions as fraudulent or non-fraudulent based on the classification model generated from the previous dataset. This information can be used to prevent fraudulent transactions and reduce financial losses.
[00049] This embodiment describes a system for medical diagnosis using multivariate analysis in big data analytics and machine learning. The system comprises a computer processor, memory, a dataset stored in memory, and program instructions stored in memory.
[00050] The dataset consists of various medical data such as patient symptoms, laboratory test results, and medical history. The program instructions include code for applying PCA to the dataset to reduce its dimensionality and code for applying a supervised machine learning algorithm to generate a classification model.
[00051] The system first pre-processes the dataset by imputing missing values and scaling the variables. Then, PCA is applied to identify the most important variables that contribute to the variation in the dataset. The reduced dataset is then used to generate a classification model using a supervised machine learning algorithm.
[00052] The system can then apply the model to new datasets to diagnose medical conditions. For example, if a user inputs a new dataset with patient symptoms, laboratory test results, and medical history, the system can classify the patient's condition based on the classification model generated from the previous dataset. This information can be used to assist medical professionals in making accurate diagnoses and providing effective treatments.
[00053] This embodiment describes a system for detecting fraudulent activities in financial services using multivariate analysis in big data analytics and machine learning. The system comprises a computer processor, memory, a dataset stored in memory, and program instructions stored in memory.
[00054] The dataset consists of various financial data such as transaction amounts, transaction times, and user locations. The program instructions include code for applying PCA to the dataset to reduce its dimensionality and code for applying a supervised machine learning algorithm to generate a classification model.
[00055] The system first pre-processes the dataset by normalizing the variables and removing outliers. Then, PCA is applied to identify the most important variables that contribute to the variation in the dataset. The reduced dataset is then used to generate a classification model using a supervised machine learning algorithm.
[00056] The system can then apply the model to new datasets to detect fraudulent activities in financial transactions. For example, if a user inputs a new dataset with transaction data, the system can classify the transactions as fraudulent or non-fraudulent based on the classification model generated from the previous dataset. This information can be used to prevent fraudulent transactions and reduce financial losses for financial services providers.
[00057] In conclusion, the embodiments discussed above demonstrate the versatility and power of the system for multivariate analysis in big data analytics and machine learning described in this patent application. The system is capable of handling large and complex datasets and can be applied in various fields such as finance, healthcare, and customer segmentation. By using PCA to reduce the dimensionality of the dataset and applying machine learning algorithms to generate models, the system can extract meaningful insights and make accurate predictions. The system has the potential to revolutionize the way we analyze and utilize big data, leading to improved decision-making and better outcomes.
[00058] Further, while operations are depicted in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Likewise, while several specific implementation details are contained in the above discussions, these should not be construed as limitations on the scope of the subject matter described herein, but rather as descriptions of features that may be specific to particular embodiments. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable sub-combination. Moreover, although features may be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a sub-combination or variation of a sub-combination.
[00059] The term “memory,” as used herein relates to a volatile or persistent medium, such as a magnetic disk, or optical disk, in which a computer can store data or software for any duration. Optionally, the memory is non-volatile mass storage such as physical storage media. Furthermore, a single memory may encompass and in a scenario wherein computing system is distributed, the processing, memory and/or storage capability may be distributed as well.
[00060] Throughout the present disclosure, the term ‘server’ relates to a structure and/or module that include programmable and/or non-programmable components configured to store, process and/or share information. Optionally, the server includes any arrangement of physical or virtual computational entities capable of enhancing information to perform various computational tasks.
[00061] Throughout the present disclosure, the term “network” relates to an arrangement of interconnected programmable and/or non-programmable components that are configured to facilitate data communication between one or more electronic devices and/or databases, whether available or known at the time of filing or as later developed. Furthermore, the network may include, but is not limited to, one or more peer-to-peer network, a hybrid peer-to-peer network, local area networks (LANs), radio access networks (RANs), metropolitan area networks (MANS), wide area networks (WANs), all or a portion of a public network such as the global computer network known as the Internet, a private network, a cellular network and any other communication system or systems at one or more locations.
[00062] Throughout the present disclosure, the term “process”* relates to any collection or set of instructions executable by a computer or other digital system so as to configure the computer or the digital system to perform a task that is the intent of the process.
[00063] Throughout the present disclosure, the term ‘Artificial intelligence (AI)’ as used herein relates to any mechanism or computationally intelligent system that combines knowledge, techniques, and methodologies for controlling a bot or other element within a computing environment. Furthermore, the artificial intelligence (AI) is configured to apply knowledge and that can adapt it-self and learn to do better in changing environments. Additionally, employing any computationally intelligent technique, the artificial intelligence (AI) is operable to adapt to unknown or changing environment for better performance. The artificial intelligence (AI) includes fuzzy logic engines, decision-making engines, preset targeting accuracy levels, and/or programmatically intelligent software.

Claims
I/We Claim:
1. A method for multivariate analysis in big data analytics and machine learning, comprising: receiving a dataset comprising a plurality of variables; applying principal component analysis (PCA) to the dataset to reduce the dimensionality of the dataset; applying a machine learning algorithm to the reduced dataset to generate a model; and applying the model to one or more additional datasets to make predictions or classifications.
2. The method of claim 1, wherein the machine learning algorithm is a supervised learning algorithm.
3. The method of claim 1, wherein the machine learning algorithm is an unsupervised learning algorithm.
4. The method of claim 1, wherein the dataset is pre-processed prior to applying PCA, including normalization, scaling, or imputation.
5. The method of claim 1, wherein the PCA is performed using an iterative algorithm.
6. The method of claim 1, wherein the model generated by the machine learning algorithm is a regression model.
7. The method of claim 1, wherein the model generated by the machine learning algorithm is a classification model.
8. A system for multivariate analysis in big data analytics and machine learning, comprising: a computer processor; memory; a dataset stored in memory; program instructions stored in memory, wherein the program instructions comprising: code for applying principal component analysis (PCA) to the dataset to reduce the dimensionality of the dataset; code for applying a machine learning algorithm to the reduced dataset to generate a model; and code for applying the model to one or more additional datasets to make predictions or classifications.

Multivariate Analysis in Big Data Analytics and Machine Learning
Abstract
This invention describes a system for multivariate analysis in big data analytics and machine learning. The system comprises a computer processor, memory, a dataset stored in memory, and program instructions stored in memory. The program instructions include code for applying principal component analysis (PCA) to the dataset to reduce its dimensionality, code for applying a machine learning algorithm to the reduced dataset to generate a model, and code for applying the model to one or more additional datasets to make predictions or classifications. The system can be applied in various fields such as finance, healthcare, and customer segmentation. By using PCA to reduce the dimensionality of the dataset and applying machine learning algorithms to generate models, the system can extract meaningful insights and make accurate predictions, revolutionizing the way we analyze and utilize big data. , Claims:Claims
I/We Claim:
1. A method for multivariate analysis in big data analytics and machine learning, comprising: receiving a dataset comprising a plurality of variables; applying principal component analysis (PCA) to the dataset to reduce the dimensionality of the dataset; applying a machine learning algorithm to the reduced dataset to generate a model; and applying the model to one or more additional datasets to make predictions or classifications.
2. The method of claim 1, wherein the machine learning algorithm is a supervised learning algorithm.
3. The method of claim 1, wherein the machine learning algorithm is an unsupervised learning algorithm.
4. The method of claim 1, wherein the dataset is pre-processed prior to applying PCA, including normalization, scaling, or imputation.
5. The method of claim 1, wherein the PCA is performed using an iterative algorithm.
6. The method of claim 1, wherein the model generated by the machine learning algorithm is a regression model.
7. The method of claim 1, wherein the model generated by the machine learning algorithm is a classification model.
8. A system for multivariate analysis in big data analytics and machine learning, comprising: a computer processor; memory; a dataset stored in memory; program instructions stored in memory, wherein the program instructions comprising: code for applying principal component analysis (PCA) to the dataset to reduce the dimensionality of the dataset; code for applying a machine learning algorithm to the reduced dataset to generate a model; and code for applying the model to one or more additional datasets to make predictions or classifications.

Documents

Application Documents

# Name Date
1 202311032877-REQUEST FOR EARLY PUBLICATION(FORM-9) [09-05-2023(online)].pdf 2023-05-09
2 202311032877-POWER OF AUTHORITY [09-05-2023(online)].pdf 2023-05-09
3 202311032877-OTHERS [09-05-2023(online)].pdf 2023-05-09
4 202311032877-FORM-9 [09-05-2023(online)].pdf 2023-05-09
5 202311032877-FORM FOR SMALL ENTITY(FORM-28) [09-05-2023(online)].pdf 2023-05-09
6 202311032877-FORM 1 [09-05-2023(online)].pdf 2023-05-09
7 202311032877-EVIDENCE FOR REGISTRATION UNDER SSI(FORM-28) [09-05-2023(online)].pdf 2023-05-09
8 202311032877-EDUCATIONAL INSTITUTION(S) [09-05-2023(online)].pdf 2023-05-09
9 202311032877-DRAWINGS [09-05-2023(online)].pdf 2023-05-09
10 202311032877-DECLARATION OF INVENTORSHIP (FORM 5) [09-05-2023(online)].pdf 2023-05-09
11 202311032877-COMPLETE SPECIFICATION [09-05-2023(online)].pdf 2023-05-09