Sign In to Follow Application
View All Documents & Correspondence

An Amalgamated Prediction Method For Breast Cancer Detection Using Fuzzy Features

Abstract: The present invention relates to an amalgamated prediction method for breast cancer detection with fuzzy features. Further, the present invention discloses an amalgamation of supervised, unsupervised learning techniques along with a powerful data filtering method. Fuzzy set theory has also been used for achieving effective feature reduction.

Get Free WhatsApp Updates!
Notices, Deadlines & Correspondence

Patent Information

Application #
Filing Date
14 December 2022
Publication Number
52/2022
Publication Type
INA
Invention Field
COMPUTER SCIENCE
Status
Email
pooja@innoveintellects.com
Parent Application

Applicants

Banasthali Vidyapith
Banasthali Vidyapith, P.O. Banasthali Banasthali Rajasthan India 304022
Dr. Seema Verma
Banasthali Vidyapith, P.O. Banasthali Banasthali Rajasthan India 304022
Dr. Rajesh Kumar
Malaviya National Institute of Technology, Jaipur Rajasthan India
Dr. Smita Jhajharia
Banasthali Vidyapith, P.O. Banasthali Banasthali Rajasthan India 304022
Dr. Manisha Jailia
Banasthali Vidyapith, P.O. Banasthali Banasthali Rajasthan India 304022
Dr. Manisha Agarwal
Banasthali Vidyapith, P.O. Banasthali Banasthali Rajasthan India 304022

Inventors

1. Dr. Seema Verma
Banasthali Vidyapith, P.O. Banasthali Banasthali Rajasthan India 304022
2. Dr. Rajesh Kumar
Malaviya National Institute of Technology, Jaipur Rajasthan India
3. Dr. Smita Jhajharia
Banasthali Vidyapith, P.O. Banasthali Banasthali Rajasthan India 304022
4. Dr. Manisha Jailia
Banasthali Vidyapith, P.O. Banasthali Banasthali Rajasthan India 304022
5. Dr. Manisha Agarwal
Banasthali Vidyapith, P.O. Banasthali Banasthali Rajasthan India 304022

Claims

1. An amalgamated prediction method for breast cancer detection using fuzzy features wherein the method comprises the steps of: a. initializing dataset and separate for each point; b. calculating apriori state as probability of point; c. updating posteriori probability of each point in the cluster using fuzzy likelihood; d. calculating new centroid for each point and for each cluster and updating kalman gain; e. checking if the centroid movement is greater than the threshold; f. if yes, then updating covariance and apriori state for each selected cluster; g. if the cluster quality is good, then selecting features and loading sum to initialize hyperplane; h. calculating distance of hyperplane with each point and for maximum width of street update weights and classifying as benign of malignant.

2. The method as claimed in claim 1, wherein the extended Kalman filter (EKF) and fuzzy K-means clustering algorithms have been combined into a hybrid algorithm with improved functionality

3. The method as claimed in claim 1, wherein the fuzzy K-means comprises: a. starting the k clusters until united for every pair; b. computing probability of point belonging to the cluster; c. assigning the values of probability membership of the points to the clusters; d. re-computing the centres of the clusters.

4. The method as claimed in claim 1, wherein the method comprises placing them as far as possible from each other is a better choice.

5. The method as claimed in claim 1, wherein the first step is completed when no point is awaiting and then the early group age is done.

6. The method as claimed in claim 1, wherein the between the new centre that is nearest and the same points of dataset,

7. The method as claimed in claim 1, wherein a new binding is to be done after obtaining these new k centroids.

8. The method as claimed in claim 1, wherein the generation of a loop takes place. Due to this loop, we can notice that there is no movement in the centres or in other words, until no further changes are done, the location is changed by the k centres step by step.

9. The method as claimed in claim 1, wherein the wherein the hybrid algorithm comprises the steps of: a. initialising dataset and separate for each point i do; b. calculating apriori state as probability of point; c. updating apriori probability as posteriori probability of each point in the cluster using fuzzy likelihood; d. calculating each cluster for each point for new centroid and update Kalman gain and check If centroid movement > threshold; e. updating covariance and apriori state for each selected cluster do; f. checking if cluster quality is good, IF no then reject feature, ELSE select feature and load sum to initialise hyperplane; g. calculating distance of hyperplane with each point and for maximum width of street update weights; and h. classifying as benign or malignant.

Specification

FIELD OF INVENTION:
The present invention relates to an amalgamated prediction method for breast cancer detection with fuzzy features.

BACKGROUND OF THE INVENTION:
Detecting cancer within the appropriate time-frame is critical given the usual nature of late diagnosis associated with most of the cancer types. Quite pertinently, the predictive detection problem (Kaur and Bawa, 2015) assumes an immediately utmost importance. Research being carried across the world has not yet been able to comprehensively infer the prognostic approach since it is challenged by a multitude of concerns. Accurately appraising the risk of cancer for individuals with high degree of accuracy and thus enabling cancer prognosis for patients are important to control the morbidity and deaths caused due to cancer. Predictive cancer models (Jhajharia et al., 2016a) additionally provide modern advancements in determining risk and identifying high-risk individuals
(Ramachandran et al., 2014) thus facilitating the planning and design of clinical cancer trials, promoting the development and analysis of risk-benefit indices, and also estimating the cancer burden and cost associated with this condition. Models may also help in the evaluation of the treatments and management.
Cancer, also called as malignant tumour or malignant neoplasm, involves unusual cell growth with metastasis. However, all tumours need not be cancerous; benign tumours are immobile. The signs and symptoms that are characteristics of different types of cancers include: lump formation, abnormal bleeds, persistence of cough for a long period, sudden weight loss (> 10%) and bowel movements irregularity, etc. Humans are affected by more than 100 different types of cancers. In a cancerous condition, the orderly process of cell growth and production goes wrong and cells begin to show an uncontrollable growth. Data has shown tobacco to have caused about 22% of all cancer related deaths. 10% deaths may be attributed to an unbalanced diet, obesity, sedentary lifestyle and consumption of alcohol (Cancer Fact Sheet, 2015). Among the less common factors are things like infections, ionising radiation and environmental pollutants. The same WHO
report (Cancer Fact Sheet, 2015) brings out the fact that in the developing countries, about 20% of cancers are caused due to infections like Hepatitis B, Hepatitis C, and Human Papilloma Virus (HPV). These factors may be responsible for partially mutating the genes in a cell, however a lot of genetic changes may be required for the development of cancer. Approximately 5%–10% of cancers are known to be due to genetic defects inherited from one or both parents. Signs and symptoms may show the presence of cancer but certain tests like medical imaging and biopsy may be necessary to confirm its presence. Unfortunately, even after diagnosis and several stages of treatment, there remains a silent risk of regeneration of the cancerous form, which needs to be appropriately detected and predicted in appropriate time. Having considered all these facts, the learning and intelligence-based predictive models developed on the foundations of machine learning theory will prove to be highly beneficial.
Even when the same data mining approach is applied to the same dataset, the results may be different since different researchers use different feature extraction and selection methods. It is important that the data is pre-processed before data mining is applied so that redundant information can be eliminated or the unstructured data can be quantified by data transformation. Theoretical guidelines for choosing appropriate patterns and features vary for various problems and methodologies. Indeed, the data collection and pattern generation processes are usually not controllable directly. Therefore, utilising feature extraction and selection is the key to simplifying the training part of the data mining process and improving the performance without changing the main body of data mining algorithms (Zheng et al., 2013). Feature extraction, also called data transformation, is the process of transforming the feature data into a quantitative data structure for training convenience.

Given the above, there arises a need to develop an improved process for the diagnosis of cancer. The present invention provides a better solution for the same.
The proposed approach is an amalgamation of supervised, unsupervised learning
techniques along with a powerful data filtering method. Fuzzy set theory has also been used for achieving effective feature reduction. The various constituent algorithms have been discussed in this section.
The amalgamated algorithm was used for prognosis of breast cancer by identifying a sample to be belonging to either ‘benign’ or ‘malignant’ class.

OBJECTIVE OF THE INVENTION:
An objective of the present invention is to provide an amalgamated prediction model for breast cancer detection using fuzzy features.
An objective of the present invention is to provide an amalgamated prediction method for breast cancer detection using fuzzy features.

SUMMARY OF THE INVENTION:
Accordingly, the present invention provides an amalgamated prediction method for breast cancer detection using fuzzy features wherein the method comprises the steps of initializing dataset and separate for each point; calculating apriori state as probability of point; updating posteriori probability of each point in the cluster using fuzzy likelihood; calculating new centroid for each point and for each cluster and updating kalman gain; checking if the centroid movement is greater than the threshold; if yes, then updating covariance and apriori state for each selected cluster; if the cluster quality is good, then selecting features and loading sum to initialize hyperplane; calculating distance of hyperplane with each point and for maximum width of street update weights and classifying as benign of malignant.
In an embodiment, the present invention provides that the extended Kalman filter (EKF) and fuzzy K-means clustering algorithms have been combined into a hybrid algorithm with improved functionality
In an embodiment, the present invention provides that fuzzy K-means comprises:
a. starting the k clusters until united for every pair;
b. computing probability of point belonging to the cluster;
c. assigning the values of probability membership of the points to the clusters;
d. re-computing the centres of the clusters.
In an embodiment, the present invention provides that the method comprises placing them as far as possible from each other is a better choice.
In an embodiment, the present invention provides that the first step is completed when no point is awaiting and then the early group age is done.
In an embodiment, the present invention provides that the between the new centre that is nearest and the same points of dataset,
In an embodiment, the present invention provides that a new binding is to be done after obtaining these new k centroids.
In an embodiment, the present invention provides that the generation of a loop takes place. Due to this loop, we can notice that there is no movement in the centres or in other words, until no further changes are done, the location is changed by the k centres step by step.
In an embodiment, the present invention provides that the wherein the hybrid algorithm comprises the steps of:
a. initialising dataset and separate for each point i do;
b. calculating apriori state as probability of point;
c. updating apriori probability as posteriori probability of each point in the cluster using fuzzy likelihood;
d. calculating each cluster for each point for new centroid and update Kalman gain and check If centroid movement > threshold;
e. updating covariance and apriori state for each selected cluster do;
f. checking if cluster quality is good, IF no then reject feature, ELSE select feature and load sum to initialise hyperplane;
g. calculating distance of hyperplane with each point and for maximum width of street update weights; and
h. classifying as benign or malignant.
BRIEF DESCRIPTION OF DRAWINGS:
These and other features, aspects, and advantages of the present invention will become better understood when the following detailed description is read concerning the accompanying drawings in which like characters represent like parts throughout the drawings, wherein:
Fig. 1: Applicability of an SVM-based technique for breast cancer prediction;
Fig. 2: Hybrid algorithm.
Further, skilled artisans will appreciate that elements in the drawings are illustrated for simplicity and may not have been necessarily been drawn to scale. Furthermore, the drawings may show only those specific details that are pertinent to understanding the embodiments of the present invention so as not to obscure the drawings with details that will be readily apparent to those of ordinary skill in the art having the benefit of the description herein.

DETAILED DESCRIPTION:
To promote an understanding of the principles of the invention, reference will now be made to the embodiment illustrated in the drawings and specific language will be used to describe the same. It will nevertheless be understood that no limitation of the scope of the invention is thereby intended, such alterations and further modifications in the illustrated system, and such further applications of the principles of the invention as illustrated therein being contemplated as would normally occur to one skilled in the art to which the invention relates.

It will be understood by those skilled in the art that the foregoing general description and the following detailed description are exemplary and explanatory of the invention and are not intended to be restrictive thereof.

Reference throughout this specification to “an aspect”, “another aspect” or similar language means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present invention. Thus, appearances of the phrase “in an embodiment”, “in another embodiment”, and similar language throughout this specification may, but do not necessarily, all refer to the same embodiment.

The terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process or method that comprises a list of steps does not include only those steps but may include other steps not expressly listed or inherent to such process or method. Similarly, one or more devices or sub-systems or elements or structures or components proceeded by "comprises...a" does not, without more constraints, preclude the existence of other devices or other sub-systems or other elements or other structures or other components or additional devices or additional sub-systems or additional elements or additional structures or additional components.

Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. The system, methods, and examples provided herein are illustrative only and not intended to be limiting.
Embodiments of the present invention will be described below in detail concerning the accompanying drawings.

The present invention discloses an amalgamation of supervised, unsupervised learning techniques along with a powerful data filtering method. Fuzzy set theory has also been used for achieving effective feature reduction.
The hybrid algorithm of the present invention called fuzzy K-means support vector machine (SVM), reduces support vectors by combining the fuzzy K-means clustering technique and SVM. Since the K-means clustering technique can almost preserve the underlying structure and distribution of the original data, the testing accuracy of such hybridised classifiers can be put under control to some degree even though reducing support vectors could incur a degradation of testing accuracy. In the present invention, the number of clusters is added into the training process as the input parameter except the kernel parameters and the penalty factor in SVM. In unsupervised learning (clustering), usually the number of clusters is subjectively determined by users by applying domain knowledge. However, when the K-means clustering technique is combined with SVM to solve the problems in supervised learning, i.e., classification, some objective criteria independent of applications can be adopted to determine these input parameters.
A hybrid algorithm combining fuzzy feature processing and enhanced Kalman filter with a SVM-based classifier. The amalgamated algorithm is used for prognosis of breast cancer by identifying a sample to be belonging to either ‘benign’ or ‘malignant’ class. Three types of datasets were used for testing the algorithm. Apart from the publicly available datasets of the UCI and WDBC, efforts were put in collecting data from actual patients in the state of Rajasthan, India. The satisfactory results obtained with the proposed hybrid algorithm provide encouragement for the adoption of machine learning-based approaches for cancer prognosis.
Extended Kalman filter
Kalman filter is a discrete time, linear time varying and finite dimensional system by which the mean-square error minimisation state is estimated. Kalman filters are a form of predictor-corrector algorithm used extensively in the control systems engineering for the estimation of a processes unmeasured states. For the control law design, the estimated states might be used then as a part of the required strategy (Chen, 2012). In order to get an approximate filter, the extended Kalman filter (EKF) is obtained as an extension of the Kalman filter to the nonlinear system model domain. This is done by applying the Kalman filter to the error system that is linear (Wan et al., 2000). In the sense of mean squared error, no optimality guarantee is offered by the EKF since it is obtained with the use of a nonlinear systems linear approximation. However, for getting the good estimates of accuracy rates from the system, the method of EKF has proven to be useful for a number of systems. This could be understood in other terms like the process where the previous value is recurred back so as to compute the value at present.
Fuzzy K-means:
An extension of the famous technique of clustering, K-means is fuzzy K-means that is also termed as fuzzy C-means (Dehariya et al., 2010). While hard clusters are discovered by K-means, a method that is formalised more statistically is fuzzy K-means by which soft clusters are discovered where with certain probability, to more than one cluster, a fuzzy point can belong to (Chang et al., 2011). Fuzzy K-means, like K-means, works on the objects whose representation is possible in the vector space that is n-dimensional and has a defined distance measure. The algorithm is described in Algorithm 1.
Algorithm 1 Fuzzy K-means

Implementation of the design
One of the simplest learning algorithms that are unsupervised is the k-means using which a clustering problem that is well known and whose solutions are available. In order to classify the given set of data, an easy and simple way is followed by the procedure through definite number of clusters set apriori. To define the k centres is the main idea, one for every cluster. Since different results are obtained for different locations therefore in a sly way, these centres must be positioned (Laskar et al., 2013). Placing them as far as possible from each other is a better choice. Each point that belongs to a given set of data is taken in the next step and then is associated to the nearest centre. First step is completed when no point is awaiting and then the early group age is done. From the previous step, the resulted clusters centroid are required to be re-calculated at this point as k new centroids. Between the new centre that is nearest and the same points of dataset, a new binding is to be done after obtaining these new k centroids. Generation of a loop takes place. Due to this loop, we can notice that there is no movement in the centres or in other words, until no further changes are done, the location is changed by the k centres step by step. Similar to K-means mentioned above, an input file is accepted by the fuzzy K-means that contains the vector points. The input directories here, similar to K-means, are not modified by the program. The output of the cluster for every iteration is stored in the directory that is cluster-N. In the cluster-N directory, the files are created. The mapper/driver/reducer/combiner are used by the code as follows: fuzzy K-means driver. This is comparable to K-means driver until it is converged or for the number of iterations specified, it iterates over the cluster points and input points. A new directory cluster-i is created during every iteration containing cluster centres that are modified and obtained during the iteration of fuzzy K-means. In the next iteration, this is fed as the input clusters (Apache Mahout, 2016). Fuzzy K-means mapper: during the method configure, it reads the input cluster and then of a point of each of the cluster, the probability of cluster membership is computed. There is an inverse relationship between distance and cluster membership. By making use of the distance measure supplied by the user, the distance is computed. The encoded clustered is the output key (Apache Mahout, 2016). The cluster observations that contain the observation statistics are the output values. Fuzzy K-means combiner: from the mapper, all the pairs of key value are received by this and then the cluster membership probabilities partial sums times each clusters input vectors are produced. The encoded cluster identifier is the output key. The cluster observations having the statistics of the observation ate the output values. Fuzzy K-means reducer: certain keys are received by multiple receivers along with the values that are associated with the keys. For the cluster, a new centroid is produced by the reducer by summing up the values and this is the output. The encoded cluster identifier is the output key. The clusters that are not converged are encoded by the reducer with the cluster Id ‘Cn and with cluster Id ‘Vn, the converged clusters are encoded.

Support vector machine
Classification techniques like SVMs (Waddell et al., 2005; Tseng et al., 2013) have been applied recently for breast cancer research. Since this is a powerful technique based on the optimal separating hyperplane theory, an SVM-based classifier model can accurately divide the data related to cancer (obtained either in digits or as images) into suitable classes to detect cancer in advance or estimate the survivability of a patient possessing an advanced stage. The use of the separating hyperplane theory for a breast cancer prediction problem, as regards to the applicability of an SVM based on maximum margin classifier approach, has been shown in Figure 1.

Data sources
Collection of data: real and digital sources
Cancer-related information sources, available from reputed institutes like the American Cancer Society, National Cancer Institute, WHO, Breast Cancer India, NCBI, Regional Cancer Centre, Bikaner (India) were referred to for this study. Efforts were also made to infer knowledge from real field experiences. Accordingly, patients suspected or diagnosed with cancer, including those attending the Out Patients Department (OPD) and those admitted with diagnosis of breast cancer, inside the wards of the Regional Cancer Centre Bikaner, Rajasthan were interacted with to gain deeper insights. Another set of patients who were suspected or diagnosed with cancer and attending the OPD at the Rajiv
Gandhi Cancer Institute, New Delhi were also spoken to for valuable information.
Relevant data and fairly recent technical content was searched from online portals like IEEE Xplore Digital Library, ACM digital library, Elsevier and PubMed. Some other journals with good impact factors were thoroughly referred to and some noteworthy points were also taken from relevant papers published in highly acclaimed conference proceedings.
Real dataset
This data is from diagnosis of breast cancer, inside the wards of the Regional Cancer Centre Bikaner, Rajasthan. Number of instances: 569, class distribution: 357 benign, 212 malignant, number of attributes: 32 (ID, diagnosis, 30 real-valued input features).
Attribute information are (ID number, diagnosis (M = malignant, B = benign) ten
real-valued features are computed for each cell nucleus:
1 radius (mean of distances from centre to points on the perimeter);
2 texture (standard deviation of gray-scale values);
3 perimeter;
4 area;
5 smoothness (local variation in radius lengths);
6 compactness;
7 concavity;
8 concave points;
9 symmetry; and
10 fractal dimension.
Out of a total 569 samples, 357 are benign and 212 are malignant cases. There are no missing attribute value in dataset.
UCI repository
The original Wisconsin breast cancer dataset was utilised in this study and it was
downloaded from the University of California, UCI machine learning repository.
Recordings are made from biopsies of real patients in the hospitals of Wisconsin, USA.
The dataset has 699 instances or samples characterised by nine attributes or features although there are 11 columns in total. The 1st column is for identifying each instance with an ID while column 11 stands for the class label for describing the severity of the tumour, the labels being benign = 2 or malignant = 4. The nine attributes of the data (excluding the class label) represent bio-physical characteristics of the tumour biopsy and are expressed as integral values. The information of the nine attributes can be reproduced from the relevant web page of the UCI repository website as:
1 clump thickness
2 uniformity of cell size
3 uniformity of cell shape
4 marginal adhesion
5 single epithelial cell size
6 bare nuclei
7 bland chromatin
8 normal nucleoli
9 mitoses.
Out of a total 699 samples, 241 are malignant and 485 are benign cases (Jhajharia et al., 2016b). The next section discusses how the training dataset obtained after random sampling of the 699 observations has been used for performing feature processing using hybrid algorithm.

Hybrid algorithm description
Since different selection methods and feature extraction are used by different researchers therefore different results might be there even when on the same set of data, same approach of data mining is applied. Before the mining is applied, it is necessary that pre-processed data is there so that by data transformation, quantification could be done of the unstructured data or elimination of the redundant information could be done. For the different methodologies and different problems, the theoretical guidelines for the selection of proper features and patterns differ. In fact, the processes of the pattern generation and collection of data are not directly handy often. To simplify the data mining processes training part and improve the performance without changing the data mining algorithms main body, the key is utilisation of the feature selection and extraction. For the convenience of training, the process of transformation of feature data
into the quantitative data is termed as Feature extraction that is also known as data transformation.

The approach is to implement an EKF optimised hybrid fuzzy K-means algorithm with SVMs for breast cancer detection. The fuzzy membership functions will be calculated in order to create a relationship between the attributes and the labels. The likelihood of each point in the clusters is calculated and is considered as state vectors. The apriori state is updated using the Kalman gain and the relationship is utilised by K-means and it will create a new modified set of attributes, which will be given to the SVM classifier. Although SVM can build classifiers with high testing accuracy, the response time of SVM classifiers still needs to improve when applied into real-time BI systems. Two elements affecting the response time of SVM classifiers are the number of input variables and that of the support vectors. The traditional k-means clustering algorithm suffers from serious drawbacks like difficulty in finding the correct method for the cluster initialisation, making a correct choice of number of clusters (k). Moreover k-means is not efficient for overlapped dataset. There have been many methods and techniques proposed to address these drawbacks of k-means. Fuzzy K-means is one of the algorithms, which provide better result than k-means for overlapped dataset. The number of features is very important to the accuracy and prediction time of the algorithm. Some features might be redundant having the same information while others might also deviate the classifier. To address this problem, proposed novel scheme of features selection is implemented, which would extract the meaningful features out of the large dimensional dataset. Although K-means is a well-known technique for clustering data without any labels, hence it is called unsupervised learning. Traditional K-means generates crisp partitions and assignment of each object is done only to one cluster. But this fails when the objects are located between clusters. The fuzzy K-means clustering algorithm partitions data points into k clusters Sl(l = 1, 2, …, k) and clusters Sl are associated with representatives (cluster centre) Cl. The relationship between a data point and cluster representative is fuzzy. That is, a membership ui,j [0, 1] is used to represent the degree of belongingness of data point Xi and cluster centre Cj. Fuzzy clustering becomes more relevant in such situations for extracting meaningful information out of these data structures. Also uncertainty handling is another challenge to crisp clustering techniques. And thus only partial knowledge is extracted during the crisp clustering approach. This paper utilises fuzzy clusters so that any object is represented by membership in a cluster rather than a crisp value. Although the developed EKF optimised fuzzy K-means algorithm is a clustering algorithm, it is utilised in this paper as an aid to the classification of breast cancer. Only few features out of the input support vectors in a particular cluster are selected. The hybrid algorithm reduces support vectors by combining the fuzzy K-means clustering technique and SVM. Since the K-means clustering technique can almost preserve the underlying structure and distribution of the original data, the testing accuracy of hybrid classifiers can be under control to some degree even though reducing support vectors could incur a degradation of testing accuracy. In the hybrid algorithm, the number of clusters is added into the training process as the input parameter except the kernel parameters and the penalty factor in SVM. The approach has been shown as a schematic in Figure 2. A learning method that is supervised is SVM. For the data classification and analysis, it is a good tool. Even for the data that is large, the SVM classifier has a quick learning speed. For the two or more problems of class classification, SVM is used. SVM is based on the notion of the optimal decision planes. The one that separates among the items set that have a different membership of class is the decision plane. With the use of the technique of the SVM, the detection and classification of the brain tumour was done. With the support vector algorithm, a hybrid fuzzy K-means algorithm has been attempted to be implemented. In order to create a relationship among the labels and attributes, calculation of the functions of the fuzzy membership will be done. By K-means, the relationship would be utilised and a new modified set will be created by this of the attributes that are then given to the SVM classifier. By combining the SVM and fuzzy K-means clustering technique, the support vectors are reduced by the proposed algorithm. Since the original data distribution and the underlying structures maintenance is almost done by the technique of K-means clustering, to some extent, the hybrid classifiers testing accuracy can be under control even though the degradation of testing accuracy could be incurred by reducing the support vectors. In the training process, the number of clusters is added in the proposed algorithm as the parameter of input except the kernel parameters and in SVM, the penalty factor. With the help of domain knowledge, the number of clusters is usually determined subjectively in unsupervised learning. However, in supervised learning, e.g., classification, to solve the problem, when the technique of K-means clustering is combined with the SVM, for the determination of these input parameters, some objective criteria could be adopted that is independent of the applications.
The following examples define the invention by way of illustration which does not limit the scope of the invention.

Example 1:
The proposed algorithm was tested for two different datasets: the UCI (Wisconsin) dataset, which is publicly available and a dataset compiled from real patients in Rajasthan, India (referred to here as the ‘real’ dataset).
Since the possible classes are ‘benign’ and ‘malignant’ only, the classification output was either ‘0’ or ‘1’ respectively. Table 1 shows the confusion matrix found for classification test done on the UCI dataset:

When used on the ‘real’ dataset, the confusion matrix was found to be as shown in Table 2.

The results obtained for individual datasets have been summarised in Table 3, in terms of classification accuracy.



We Claim:

1. An amalgamated prediction method for breast cancer detection using fuzzy features wherein the method comprises the steps of:
a. initializing dataset and separate for each point;
b. calculating apriori state as probability of point;
c. updating posteriori probability of each point in the cluster using fuzzy likelihood;
d. calculating new centroid for each point and for each cluster and updating kalman gain;
e. checking if the centroid movement is greater than the threshold;
f. if yes, then updating covariance and apriori state for each selected cluster;
g. if the cluster quality is good, then selecting features and loading sum to initialize hyperplane;
h. calculating distance of hyperplane with each point and for maximum width of street update weights and classifying as benign of malignant.
2. The method as claimed in claim 1, wherein the extended Kalman filter (EKF) and fuzzy K-means clustering algorithms have been combined into a hybrid algorithm with improved functionality
3. The method as claimed in claim 1, wherein the fuzzy K-means comprises:
a. starting the k clusters until united for every pair;
b. computing probability of point belonging to the cluster;
c. assigning the values of probability membership of the points to the clusters;
d. re-computing the centres of the clusters.
4. The method as claimed in claim 1, wherein the method comprises placing them as far as possible from each other is a better choice.
5. The method as claimed in claim 1, wherein the first step is completed when no point is awaiting and then the early group age is done.
6. The method as claimed in claim 1, wherein the between the new centre that is nearest and the same points of dataset,
7. The method as claimed in claim 1, wherein a new binding is to be done after obtaining these new k centroids.
8. The method as claimed in claim 1, wherein the generation of a loop takes place. Due to this loop, we can notice that there is no movement in the centres or in other words, until no further changes are done, the location is changed by the k centres step by step.
9. The method as claimed in claim 1, wherein the wherein the hybrid algorithm comprises the steps of:
a. initialising dataset and separate for each point i do;
b. calculating apriori state as probability of point;
c. updating apriori probability as posteriori probability of each point in the cluster using fuzzy likelihood;
d. calculating each cluster for each point for new centroid and update Kalman gain and check If centroid movement > threshold;
e. updating covariance and apriori state for each selected cluster do;
f. checking if cluster quality is good, IF no then reject feature, ELSE select feature and load sum to initialise hyperplane;
g. calculating distance of hyperplane with each point and for maximum width of street update weights; and
h. classifying as benign or malignant.

Documents

Application Documents

# Name Date
1 202211072105-STATEMENT OF UNDERTAKING (FORM 3) [14-12-2022(online)].pdf 2022-12-14
2 202211072105-REQUEST FOR EARLY PUBLICATION(FORM-9) [14-12-2022(online)].pdf 2022-12-14
3 202211072105-POWER OF AUTHORITY [14-12-2022(online)].pdf 2022-12-14
4 202211072105-FORM-9 [14-12-2022(online)].pdf 2022-12-14
5 202211072105-FORM FOR SMALL ENTITY(FORM-28) [14-12-2022(online)].pdf 2022-12-14
6 202211072105-FORM 1 [14-12-2022(online)].pdf 2022-12-14
7 202211072105-EVIDENCE FOR REGISTRATION UNDER SSI(FORM-28) [14-12-2022(online)].pdf 2022-12-14
8 202211072105-EVIDENCE FOR REGISTRATION UNDER SSI [14-12-2022(online)].pdf 2022-12-14
9 202211072105-EDUCATIONAL INSTITUTION(S) [14-12-2022(online)].pdf 2022-12-14
10 202211072105-DRAWINGS [14-12-2022(online)].pdf 2022-12-14
11 202211072105-DECLARATION OF INVENTORSHIP (FORM 5) [14-12-2022(online)].pdf 2022-12-14
12 202211072105-COMPLETE SPECIFICATION [14-12-2022(online)].pdf 2022-12-14
13 202211072105-FORM 18 [30-01-2023(online)].pdf 2023-01-30
14 202211072105-FER.pdf 2024-01-04
15 202211072105-FER_SER_REPLY [04-07-2024(online)].pdf 2024-07-04

Search Strategy

1 SearchHistory(2)E_29-12-2023.pdf