Sign In to Follow Application
View All Documents & Correspondence

Method For Prediction Of Bug For Optimizing Parameters And Improving Accuracy Using Machine Learning

Abstract: “METHOD FOR PREDICTION OF BUG FOR OPTIMIZING PARAMETERS AND IMPROVING ACCURACY USING MACHINE LEARNING” Accordingly, embodiments herein disclose a method for prediction of bug for optimizing parameters and improving accuracy using a machine learning (ML). The method involves acquiring data from a reputed promise repository; and obtaining the data preprocessing by applying some feature selection techniques. The data modelling is provided with data metrics, and the data is to be visualized. The feature selection forms the foundation for machine learning (ML) which contributes to feature measure or assessment criterion in the data model. After the feature selection process, the training and testing are done by a random sampling chniquest and taking ratio of bugged and not bugged instances, training an ML algorithm to predict labels from characteristics, tweaking it for the business need, and verifying it on outlier data are all part of the modelling process. The training and testing ratio of 80:20 has been taken into consideration and enhancing the learning procedure. Figure to be published with Abstract: Figure 1 Dated this 27th day of April, 2023 POOJA AGENT FOR THE APPLICANT IN/PA/1838

Get Free WhatsApp Updates!
Notices, Deadlines & Correspondence

Patent Information

Application #
Filing Date
05 May 2023
Publication Number
25/2023
Publication Type
INA
Invention Field
COMPUTER SCIENCE
Status
Email
Parent Application

Applicants

Banasthali Vidyapith
Banasthali Vidyapith, P.O. Banasthali, Banasthali, Rajasthan, India, 304022
Dr. Manisha Agarwal
Banasthali Vidyapith, P.O. Banasthali, Banasthali, Rajasthan, India, 304022
Mrs. Nidhi Srivastava
Banasthali Vidyapith, P.O. Banasthali, Banasthali, Rajasthan, India, 304022

Inventors

1. Dr. Manisha Agarwal
Banasthali Vidyapith, P.O. Banasthali, Banasthali, Rajasthan, India, 304022
2. Mrs. Nidhi Srivastava
Banasthali Vidyapith, P.O. Banasthali, Banasthali, Rajasthan, India, 304022

Claims

1. A method for prediction of bug for optimizing parameters and improving accuracy using a machine learning (ML) model, comprising the steps of: acquiring data from a reputed promise repository; obtaining the data preprocessing by applying some feature selection techniques; providing the data modelling with data metrics; and visualizing the data, wherein the feature selection forms the foundation for machine learning (ML) which contributes to feature measure or assessment criterion in the data model.

2. The method as claimed in the claim 1, wherein the dataset matrices can help in selecting and finding the important variables.

3. The method as claimed in the claim 1, further comprising after the feature selection process, the training and testing are done by a random sampling chniquest and taking ratio of bugged and not bugged instances, training an ML algorithm to predict labels from characteristics, tweaking it for the business need, and verifying it on outlier data are all part of the modelling process.

4. The method as claimed in the claim 1, wherein the training and testing ratio of 80:20 has been taken into consideration and enhancing the learning procedure. Dated this 27th day of April, 2023 POOJA AGENT FOR THE APPLICANT IN/PA/1838

Specification

Description:FIELD OF INVENTION
[0001] The present disclosure relates to the field of machine learning (ML) technology and more particularly, to a method for prediction of bug for optimizing parameters and improving accuracy using a machine learning (ML) model.
BACKGROUND OF INVENTION
[0002] Bug prediction can be accomplished with the aid of machine learning and predictive analysis. Developers can make improvements as they create code by integrating the prediction models into their development environments. Unreliable predictions, however, are inevitable. These erroneous predictions fall into two categories: those that mistakenly label clean code as buggy and those that inappropriately label buggy code as clean.
[0003] Typically, the problem statement contains a significant gap that must be filled in for an ideal model to balance out the inaccurate prediction and inspire developers to believe in the model. Even with a careful software application, problem definition centered on delivering bug-free software is a difficult task. The models have been studied in terms of their level of accuracy and complexity despite the lack of common benchmarks for model comparison. The accuracy of the model is greatly influenced by the selected metrics and this becomes the most important step in the bug prediction. The method becomes more difficult as the number of metrics in the model increases. The inclusion of pointless measurements can significantly reduce accuracy.
[0004] Software development challenges represent a learning process that varies depending on the conditions and the stages of development in which we find and can easily detect the problem. Through Fig. 1, it has been shown how the data development process is carried out on three levels i.e. Level 1, Level 2, and Level 3. At Level 1, data filtration and extraction have been performed, and then all of the extracted data is dissected into Training, Testing, and Validation on the data in Level 2, and Level 3, which provides the actual notable data for an analyst to work on and compare to the entire system. It is always recommended to perform resource-intensive, time-consuming, and expensive sedation activities.
[0005] Therefore, the proposed invention is to determine the best bug detection algorithm using machine learning, evaluate the accuracy of the entire algorithm, and compare them. The optimal algorithm will make it simple for the user to evaluate the findings.
[0006] Thus, it is desired to address the above-mentioned disadvantages or other shortcomings or at least provide a useful alternative.
OBJECTIVE OF INVENTION
[0007] The principal object of the embodiments herein is to determine the best bug detection algorithm using machine learning.
[0008] Another object of the embodiments herein is to evaluate the accuracy of the entire algorithm.
[0009] Another object of the embodiments herein is simple for the user to evaluate the findings
[0010] Still, another object of the embodiments herein is user friendly.
SUMMARYOF INVENTION
[0011] The following presents a simplified summary in order to provide a basic understanding of some aspects of the disclosed invention. This summary is not an extensive overview, and it is not intended to identify key/critical elements or to delineate the scope thereof. Its sole purpose is to present some concepts in a simplified form as a prelude to the more detailed description that is presented later.
[0012] Accordingly, there is provided to a method for prediction of bug for optimizing parameters and improving accuracy using a machine learning (ML) model. The proposed method is to demonstrate the applicability of machine learning algorithms namely Neural Network, SVM, Decision Tree and Cubist in using different performance metrics i.e. R, R square, Root Mean Square Error, Accuracy and obtaining the optimal outcome-based algorithm for a Bug report on diversion dataset from PROMISE repository.
[0013] In accordance with an embodiment of the present subject matter relates to a method for prediction of bug for optimizing parameters and improving accuracy using a machine learning (ML). The method involves acquiring data from a reputed promise repository; and obtaining the data preprocessing by applying some feature selection techniques. The data modelling is provided with data metrics, and the data is to be visualized. The feature selection forms the foundation for machine learning (ML) which contributes to feature measure or assessment criterion in the data model. After the feature selection process, the training and testing are done by a random sampling chniquest and taking ratio of bugged and not bugged instances, training an ML algorithm to predict labels from characteristics, tweaking it for the business need, and verifying it on outlier data are all part of the modelling process. The training and testing ratio of 80:20 has been taken into consideration and enhancing the learning procedure.
[0014] Accordingly, the major contribution of the suggested approach is the evaluation of ML classifiers, i.e. in software bug prediction, the use of NN, SVM, DT, and Cubist. Seven diverse datasets, including Ant, Ivcene, Camel 1.6, Poi3, Synapse, Tomcat, and velocity, are subjected to the discussed ML classifiers to achieve bug-free software.
[0015] The proposed method is highlighting the accuracy obtained by the current approaches that are significant for research scholars and solution providers.
[0016] These hand other aspects of the embodiments herein will be better appreciated and understood when considered in conjunction with the following description and the accompanying drawings. It should be understood, however, that the following descriptions, while indicating preferred embodiments and numerous specific details thereof, are given by way of illustration and not of limitation. Many changes and modifications may be made within the scope of the embodiments herein without departing from the scope thereof, and the embodiments herein include all such modifications.
BRIEF DESCRIPTION OF FIGURES
[0017] These and other features, aspects and advantages of the present invention will be better understood if the following detailed description is read with reference to the accompanying drawings, in which like characters represent like parts in the drawings, where:
[0018] Figure 1 illustrates flow diagram of a method for prediction of bug for optimizing parameters and improving accuracy using a machine learning (ML) model, according to the present invention.
[0019] It will be understood by those skilled in the art that the elements in the drawings are shown for convenience and have not necessarily been drawn to scale. For example, the flowcharts illustrate the process using key steps to enhance understanding of aspects of the present invention. In addition, one or more components of the device may be represented in the drawings by conventional symbols, and the drawings may show only the specific details relevant to an understanding of embodiments of the present invention, so as not to clutter the drawings with details that are readily apparent to those skilled in the art who are familiar with the present description.
DETAILED DESCRIPTION OF INVENTION
[0020] The embodiments herein and the various features and advantageous details thereof are explained more fully with reference to the non-limiting embodiments that are illustrated in the accompanying drawings and detailed in the following description. Descriptions of well-known components and processing techniques are omitted so as to not unnecessarily obscure the embodiments herein. Also, the various embodiments described herein are not necessarily mutually exclusive, as some embodiments can be combined with one or more other embodiments to form new embodiments. The term “or” as used herein, refers to a non-exclusive or, unless otherwise indicated. The examples used herein are intended merely to facilitate an understanding of ways in which the embodiments herein can be practiced and to further enable those skilled in the art to practice the embodiments herein. Accordingly, the examples should not be construed as limiting the scope of the embodiments herein.
[0021] The accompanying drawings are used to help easily understand various technical features and it should be understood that the embodiments presented herein are not limited by the accompanying drawings. As such, the present disclosure should be construed to extend to any alterations, equivalents and substitutes in addition to those which are particularly set out in the accompanying drawings. Although the terms first, second, etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are generally only used to distinguish one element from another.
[0022] Accordingly, embodiments here are provided to a method for prediction of bug for optimizing parameters and improving accuracy using a machine learning (ML) model. The proposed method is to demonstrate the applicability of machine learning algorithms namely Neural Network, SVM, Decision Tree and Cubist in using different performance metrics i.e. R, R square, Root Mean Square Error, Accuracy and obtaining the optimal outcome-based algorithm for a Bug report on diversion dataset from PROMISE repository.
[0023] The major contribution of the suggested approach is the evaluation of ML classifiers, i.e. in software bug prediction, the use of NN, SVM, DT, and Cubist. Seven diverse datasets, including Ant, Ivcene, Camel 1.6, Poi3, Synapse, Tomcat, and velocity, are subjected to the discussed ML classifiers to achieve bug-free software.
[0024] Referring Figure 1 illustrates flow diagram of method for prediction of bug for optimizing parameters and improving accuracy using a machine learning (ML) model. Regression is one of the ML techniques for determining the relationship between relevant variables; specifically, regression allows for the selection of the curve that best fits the available data. Many regression techniques are available for resolving the engineering problem. The goal of the regression is to reduce the total squared errors (least squares). In the model entire process is divided into four different stages which are shown in Figure 1.
[0025] In an embodiment, the method involves acquiring data from a reputed promise repository; and obtaining the data preprocessing by applying some feature selection techniques. The data modelling is provided with data metrics, and the data is to be visualized. The feature selection forms the foundation for machine learning (ML) which contributes to feature measure or assessment criterion in the data model. After the feature selection process, the training and testing are done by a random sampling chniquest and taking ratio of bugged and not bugged instances, training an ML algorithm to predict labels from characteristics, tweaking it for the business need, and verifying it on outlier data are all part of the modelling process. The training and testing ratio of 80:20 has been taken into consideration and enhancing the learning procedure.
[0026] Accordingly, for the purpose of research, Bug dataset is used, where 20 metrics i.e.WMC,MFA,DIT,CAM,NOC,IC,CBO,CBM,RFC,AMC,LCOM,Ca,LCOM3,Ce,NPM,Max_cc,DAM,Avg_CC,MOA,LOC are used as the features (i.e., independent variables) and the metric “bug” is used as the response or dependent variable a detailed illustration of the same is given in TABLE 1 . The feature variables (i.e., independent variables) are 20 metrics, and the response (or dependent) parameter is the number of bugs.
S No. Metrics Description
1. WMC Weighted methods per class
2. MFA A measure of Functional Abstraction
3. DIT Depth of Inheritance Tree
4. CAM Cohesion Among Methods of
Class
5. NOC Number of Children IC Inheritance Coupling
6. IC Inheritance Coupling
7. CBO Coupling between object classes
8. CBM Coupling Between Methods
9. RFC Response to a Class
10. AMC Average Method Complexity
11. LCOM Lack of cohesion in methods
12. Ca Afferent couplings
13. LCOM3 Lack of cohesion in methods
14. Ce Efferent couplings
15. NPM Number of Public Methods
16. Max_CC Maximum of McCabe's cyclomatic complexity
17. DAM Data Access Metric
18. Avg_CC Average of McCabe's cyclomatic complexity
19. MOA Measure of Aggregation
20. LOC Lines of Code
21. Bug Bug
TABLE 1: The Java Metrics in the bug prediction dataset
[0027] The generalization performance of the ML algorithm is frequently influenced by the data preprocessing. One of the most challenging inductive ML challenges is the removal of noise instances. Another frequently addressed concern in data preprocessing is missing data handling. It is generally best to determine well-known data preprocessing methods, like data normalization, feature selection, and training & Testing of data. Feature selection forms the foundation for ML, it contributes to feature measure or assessment criterion in data model. Boruta deals with the issue of increasing the system's randomization. The basic concept is pretty straightforward: simply duplicate the system using randomization, combine it with the original, and then develop a classifier for this expanded system. Then contrast it with that of the randomized variables to determine the variable's significance in the original system. Only variables are considered important if their importance exceeds that of the randomized variables. After implementing Boruta in the proposed method on the dataset matrices that help in selecting and finding important variables.
[0028] Regression algorithms are being used in the research to get the expected outcomes from the existing data. Different regression algorithms used in this paper are Neural Network (NN), Support Vector Machine (SVM), Decision Tree (DT) and Cubist with the help of all the above algorithms results are evaluated.
[0029] Neural Network (NN): Neural Network can also be defined as the number of hidden nodes included in the model or the number of inputs and outputs present in each node.
[0030] Support Vector Machine (SVM): Each data point in the SVM algorithm is plotted as a point in n-dimensional space (where n is no. of features in a dataset). The classification is then carried out by locating the hyper-plane that best distinguishes the two classes.
[0031] Decision Tree (DT): The decision Tree algorithm begins at the root node and takes a step up to predict the class of a given dataset. This algorithm follows the branch and jumps to the mentioned node based on a comparison of the values of the data.
[0032] Cubist: Cubist is constructed using the predictors from earlier splits. Additionally, there are intermediate linear models at every stage of the tree, and by combining them one can create a tree with various rates of growth.
[0033] Accordingly, the different performance indicators are used for ML tasks in regression. There are numerous metrics for the problems, and the model performance can be determined by Root Mean square error (RMSE, Eq.(2)) is used to calculate how close the regression line is to a set of points, coefficient of determination are used to assess the model's correctness during training and validation R squared (R2, Eq. (1)) and Accuracy (Eq. 3) is the proportion of correct predictions over total predictions are employed together for the performance of the models. Mathematical expressions of several measures are given below:
(1)
Where:
= actual y value – predicted y value
= predicted value
(2)
Where:
t = number of instances.
A = Actual observed value
P= Predicted value

(3)
[0034] The proposed method is selecting the best model from a group of effective models can be reduced to simply identifying the portions of the model that provide the highest accuracy or lowest loss while ensuring that the model doesn't overfit. The model visualization generates actions by applying previously learned information to new input. In TABLE 2 the average size of bugs which were seen earlier is helping in data visualization.

System #version No. of classes Avg no. of Bugs
Ant 338 19.58
Camel 696 18.87
Ivcene2.4 235 24.92
Poi3 345 49.82
Synapse 212 23.6
Tomcat 858 8.97
Velocity 213 58.47
TABLE 2 Java dataset used in OPABP with average no. of bugs
[0035] Typically, there are seven datasets (Ant, Camel, Iucene 2.4, Poi3, Synapse, Tomcat, and Velocity) in OPABP detailed information is already discussed and on every dataset, four different regression algorithms i.e. Neural Network, SVM, Decision Tree and cubist were being implemented, and with the help of the Mean Square Error, R square and accuracy is being calculated on each dataset. Also, it is being found out the algorithm that is showing the best accuracy after judging their values.
Ant Dataset
R R2 RMSE Accuracy
Neural Network 0.02 0 0.76 88.2
Decision Tree 0.52 0.27 0.5 88
SVM 0.93 0.86 0.1 100
Cubist -0.05 0 0.75 83.65
Camel 1.6 Dataset
R R2 RMSE Accuracy
Neural Network -0.05 0 0.8 89.69
Decision Tree 0.02 0 0.79 85.57
SVM 0.38 0.14 0.61 90.7
Cubist 0.55 0.3 0.7 90.86

Ivcene Dataset
R R2 RMSE Accuracy
Neural Network 0.11 0.01 1.8 33.33
Decision Tree 0.64 0.41 1.4 47.83
SVM 0.64 0.41 1.24 56.2
Cubist 0.66 0.44 1.15 58.48
Poi3 Dataset
R R2 RMSE Accuracy
Neural Network 0.1 0.01 0.91 75.28
Decision Tree 0.6 0.36 0.82 87.64
SVM 0.4 0.16 0.84 80.23
Cubist 0.51 0.26 0.82 81.53
Synapse Dataset
R R2 RMSE Accuracy
Neural Network 0.31 0.1 0.57 84.62
Decision Tree 0.51 0.26 0.63 82.69
SVM 0.38 0.14 0.53 87.38
Cubist 0.58 0.34 0.54 87.6
Tomcat Dataset
R R2 RMSE Accuracy
Neural Network 0.32 0.1 0.16 99.42
Decision Tree 0.43 0.18 0.23 93.6
SVM 0.41 0.17 0.18 96.51
Cubist 0.42 0.18 0.16 97.21
Velocity Dataset
R R2 RMSE Accuracy
Neural Network 0.25 0.06 1 58.7
Decision Tree 0.36 0.13 1.06 73.91
SVM 0.45 0.2 0.79 77.17
Cubist 0.45 0.2 0.72 79.13
TABLE 3 Result of Different performance measures
[0036] In the TABLE 3 can be seen that there are different performance measures which are evaluated on 7 different datasets and with the implementation of four different algorithms. Feature selection is done on every dataset and for important features, the only performance measure is calculated.
[0037] Further, the proposed method includes a comparative examination of ML algorithms and demonstrates the performance accuracy and capabilities in software bug prediction. The SVM is found to be the most précised model with the highest value of accuracy. Whereas, in some places, Decision Tree and Neural Network also come up with better values and for some datasets, Cubist is also producing significant values as shown in TABLE 4 and accuracy values used for the different ML algorithms.
Ant Camel 1.6 Ivcene Poi3 Synapse Tomcat Velocity
NN 88.20 89.69 33.33 75.28 84.62 99.42 58.70
DT 88.00 85.57 47.83 87.64 82.69 93.60 73.91
SVM 100.00 90.70 56.20 80.23 87.38 96.51 77.17
Cubist 83.65 90.86 58.48 81.53 87.60 97.21 79.13
TABLE 4 Comparing the Data Model in respect of Accuracy
[0038] The foregoing description of the specific embodiments will so fully reveal the general nature of the embodiments herein that others can, by applying current knowledge, readily modify and/or adapt for various applications such specific embodiments without departing from the generic concept, and, therefore, such adaptations and modifications should and are intended to be comprehended within the meaning and range of equivalents of the disclosed embodiments. It is to be understood that the phraseology or terminology employed herein is for the purpose of description and not of limitation. Therefore, while the embodiments herein have been described in terms of preferred embodiments, those skilled in the art will recognize that the embodiments herein can be practiced with modification within the spirit and scope of the embodiments as described herein.
, Claims:We claim:
1. A method for prediction of bug for optimizing parameters and improving accuracy using a machine learning (ML) model, comprising the steps of:
acquiring data from a reputed promise repository;
obtaining the data preprocessing by applying some feature selection techniques;
providing the data modelling with data metrics; and
visualizing the data,
wherein the feature selection forms the foundation for machine learning (ML) which contributes to feature measure or assessment criterion in the data model.
2. The method as claimed in the claim 1, wherein the dataset matrices can help in selecting and finding the important variables.
3. The method as claimed in the claim 1, further comprising after the feature selection process, the training and testing are done by a random sampling chniquest and taking ratio of bugged and not bugged instances, training an ML algorithm to predict labels from characteristics, tweaking it for the business need, and verifying it on outlier data are all part of the modelling process.
4. The method as claimed in the claim 1, wherein the training and testing ratio of 80:20 has been taken into consideration and enhancing the learning procedure.
Dated this 27th day of April, 2023

POOJA
AGENT FOR THE APPLICANT
IN/PA/1838

Documents

Application Documents

# Name Date
1 202311031990-STATEMENT OF UNDERTAKING (FORM 3) [05-05-2023(online)].pdf 2023-05-05
2 202311031990-REQUEST FOR EARLY PUBLICATION(FORM-9) [05-05-2023(online)].pdf 2023-05-05
3 202311031990-POWER OF AUTHORITY [05-05-2023(online)].pdf 2023-05-05
4 202311031990-FORM-9 [05-05-2023(online)].pdf 2023-05-05
5 202311031990-FORM FOR SMALL ENTITY(FORM-28) [05-05-2023(online)].pdf 2023-05-05
6 202311031990-FORM FOR SMALL ENTITY [05-05-2023(online)].pdf 2023-05-05
7 202311031990-FORM 1 [05-05-2023(online)].pdf 2023-05-05
8 202311031990-FIGURE OF ABSTRACT [05-05-2023(online)].pdf 2023-05-05
9 202311031990-EVIDENCE FOR REGISTRATION UNDER SSI(FORM-28) [05-05-2023(online)].pdf 2023-05-05
10 202311031990-EVIDENCE FOR REGISTRATION UNDER SSI [05-05-2023(online)].pdf 2023-05-05
11 202311031990-DRAWINGS [05-05-2023(online)].pdf 2023-05-05
12 202311031990-DECLARATION OF INVENTORSHIP (FORM 5) [05-05-2023(online)].pdf 2023-05-05
13 202311031990-COMPLETE SPECIFICATION [05-05-2023(online)].pdf 2023-05-05
14 202311031990-FORM 18 [10-10-2023(online)].pdf 2023-10-10
15 202311031990-FER.pdf 2025-03-22
16 202311031990-FER_SER_REPLY [22-09-2025(online)].pdf 2025-09-22
17 202311031990-DRAWING [22-09-2025(online)].pdf 2025-09-22
18 202311031990-CLAIMS [22-09-2025(online)].pdf 2025-09-22

Search Strategy

1 202311031990E_30-08-2024.pdf