Proceedings Article10.1109/ICMLA.2013.13
Improving Software Quality Estimation by Combining Boosting and Feature Selection
Kehan Gao,Taghi M. Khoshgoftaar,Amri Napolitano +2 more
- 04 Dec 2013
- Vol. 1, pp 27-33
3
TL;DR: The experimental results demonstrate that with the exception of one learner, feature selection combined with boosting provides better classification performance than when either is applied alone or when neither are applied.
read more
Abstract: The predictive accuracy of a classification modelis often affected by the quality of training data. However, there are two problems which may affect the quality of the training data: high dimensionality (too many independent attributes in a dataset) and class imbalance (many more instances of one class than the other class in a binary-classification problem). In this study, we present an iterative feature selection approach working with an ensemble learning method to solve both of these problems. The iterative feature selection approach samples the dataset k times and applies feature ranking to each sampled dataset, the k different rankings are then aggregated to create a single feature ranking. The ensemble learning method used is RUSBoost, in which random under sampling(RUS) is integrated into a boosting algorithm. The main purpose of this paper is to investigate the impact of feature selection as well as the RUSBoost approach on the classification performance in the context of software quality prediction. In the experiment, we explore six rankers, each used along with RUS in the iterative feature selection process. Following feature selection, models are built either using a plain learner or byusing the RUSBoost algorithm. We also examine the case of no feature selection and use this as the baseline for comparisons. The experimental results demonstrate that with the exception of one learner, feature selection combined with boosting provides better classification performance than when either is applied alone or when neither are applied.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
•Proceedings Article
Choosing the Best Classification Performance Metric for Wrapper-based Software Metric Selection for Defect Prediction.
Huanjing Wang,Taghi M. Khoshgoftaar,Amri Napolitano +2 more
- 01 Jan 2014
TL;DR: Five wrapper-based feature selection methods to remove irrelevant and redundant features are used and it is demonstrated that BAM is the best performance metric used within the wrapper.
2
Patent
Exception prediction before an actual exception during debugging
Vikas Chandra,Sarika Sinha +1 more
- 26 Jan 2016
TL;DR: In this article, an approach is provided for predicting an exception during debugging of software code before the debugging encounters the exception, based on the execution of the upcoming lines, a prediction is determined that the exception will be encountered at line number M, which is within a range of line numbers L+1 and X.
2
References
Borderline-SMOTE: a new over-sampling method in imbalanced data sets learning
Hui Han,Wenyuan Wang,Binghuan Mao +2 more
- 23 Aug 2005
TL;DR: Two new minority over-sampling methods are presented, borderline- SMOTE1 and borderline-SMOTE2, in which only the minority examples near the borderline are over- Sampling, which achieve better TP rate and F-value than SMOTE and random over-Sampling methods.
A Practical Approach to Feature Selection
Kenji Kira,Larry A. Rendell +1 more
- 07 Jul 1992
TL;DR: Comparison with other feature selection algorithms shows Relief's advantages in terms of learning time and the accuracy of the learned concept, suggesting Relief's practicality.
3.3K
•Journal Article
An extensive empirical study of feature selection metrics for text classification
TL;DR: An empirical comparison of twelve feature selection methods evaluated on a benchmark of 229 text classification problem instances, revealing that a new feature selection metric, called 'Bi-Normal Separation' (BNS), outperformed the others by a substantial margin in most situations and was the top single choice for all goals except precision.
SMOTEBoost: Improving Prediction of the Minority Class in Boosting
Nitesh V. Chawla,Aleksandar Lazarevic,Lawrence O. Hall,Kevin W. Bowyer +3 more
- 22 Sep 2003
TL;DR: This paper presents a novel approach for learning from imbalanced data sets, based on a combination of the SMOTE algorithm and the boosting procedure, which shows improvement in prediction performance on the minority class and overall improved F-values.