Proceedings Article10.1109/ICMLA.2013.13
Improving Software Quality Estimation by Combining Boosting and Feature Selection
Kehan Gao,Taghi M. Khoshgoftaar,Amri Napolitano +2 more
- 04 Dec 2013
- Vol. 1, pp 27-33
3
TL;DR: The experimental results demonstrate that with the exception of one learner, feature selection combined with boosting provides better classification performance than when either is applied alone or when neither are applied.
read more
Abstract: The predictive accuracy of a classification modelis often affected by the quality of training data. However, there are two problems which may affect the quality of the training data: high dimensionality (too many independent attributes in a dataset) and class imbalance (many more instances of one class than the other class in a binary-classification problem). In this study, we present an iterative feature selection approach working with an ensemble learning method to solve both of these problems. The iterative feature selection approach samples the dataset k times and applies feature ranking to each sampled dataset, the k different rankings are then aggregated to create a single feature ranking. The ensemble learning method used is RUSBoost, in which random under sampling(RUS) is integrated into a boosting algorithm. The main purpose of this paper is to investigate the impact of feature selection as well as the RUSBoost approach on the classification performance in the context of software quality prediction. In the experiment, we explore six rankers, each used along with RUS in the iterative feature selection process. Following feature selection, models are built either using a plain learner or byusing the RUSBoost algorithm. We also examine the case of no feature selection and use this as the baseline for comparisons. The experimental results demonstrate that with the exception of one learner, feature selection combined with boosting provides better classification performance than when either is applied alone or when neither are applied.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
•Proceedings Article
Choosing the Best Classification Performance Metric for Wrapper-based Software Metric Selection for Defect Prediction.
Huanjing Wang,Taghi M. Khoshgoftaar,Amri Napolitano +2 more
- 01 Jan 2014
TL;DR: Five wrapper-based feature selection methods to remove irrelevant and redundant features are used and it is demonstrated that BAM is the best performance metric used within the wrapper.
2
Patent
Exception prediction before an actual exception during debugging
Vikas Chandra,Sarika Sinha +1 more
- 26 Jan 2016
TL;DR: In this article, an approach is provided for predicting an exception during debugging of software code before the debugging encounters the exception, based on the execution of the upcoming lines, a prediction is determined that the exception will be encountered at line number M, which is within a range of line numbers L+1 and X.
2
References
RUSBoost: A Hybrid Approach to Alleviating Class Imbalance
C. Seiffert,Taghi M. Khoshgoftaar,J. Van Hulse,Amri Napolitano +3 more
- 01 Jan 2010
TL;DR: This paper presents a new hybrid sampling/boosting algorithm, called RUSBoost, for learning from skewed training data, which provides a simpler and faster alternative to SMOTEBoost, which is another algorithm that combines boosting and data sampling.
Feature Selection: An Ever Evolving Frontier in Data Mining
Huan Liu,Hiroshi Motoda,Rudy Setiono,Zheng Zhao +3 more
- 26 May 2010
TL;DR: The key components of feature selection are introduced, and its developments with the growth of data mining are reviewed, and some potential lines of research that require multidisciplinary research are identified.
A General Software Defect-Proneness Prediction Framework
TL;DR: The results show that the proposed framework for software defect prediction is more effective and less prone to bias than previous approaches and that small details in conducting how evaluations are conducted can completely reverse findings.
Support Vector Machines.
Nello Cristianini,Elisa Ricci +1 more
- 01 Jan 2008
TL;DR: A new command is introduced, svmachines, which is a thin wrapper for the widely deployed libsvm and can be applied to continuous, binary, and categorical outcomes analogous to Gaussian, logistic, and multinomial regression.
333