Proceedings Article10.1109/ICMLA.2013.13
Improving Software Quality Estimation by Combining Boosting and Feature Selection
Kehan Gao,Taghi M. Khoshgoftaar,Amri Napolitano +2 more
- 04 Dec 2013
- Vol. 1, pp 27-33
3
TL;DR: The experimental results demonstrate that with the exception of one learner, feature selection combined with boosting provides better classification performance than when either is applied alone or when neither are applied.
read more
Abstract: The predictive accuracy of a classification modelis often affected by the quality of training data. However, there are two problems which may affect the quality of the training data: high dimensionality (too many independent attributes in a dataset) and class imbalance (many more instances of one class than the other class in a binary-classification problem). In this study, we present an iterative feature selection approach working with an ensemble learning method to solve both of these problems. The iterative feature selection approach samples the dataset k times and applies feature ranking to each sampled dataset, the k different rankings are then aggregated to create a single feature ranking. The ensemble learning method used is RUSBoost, in which random under sampling(RUS) is integrated into a boosting algorithm. The main purpose of this paper is to investigate the impact of feature selection as well as the RUSBoost approach on the classification performance in the context of software quality prediction. In the experiment, we explore six rankers, each used along with RUS in the iterative feature selection process. Following feature selection, models are built either using a plain learner or byusing the RUSBoost algorithm. We also examine the case of no feature selection and use this as the baseline for comparisons. The experimental results demonstrate that with the exception of one learner, feature selection combined with boosting provides better classification performance than when either is applied alone or when neither are applied.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
•Proceedings Article
Choosing the Best Classification Performance Metric for Wrapper-based Software Metric Selection for Defect Prediction.
Huanjing Wang,Taghi M. Khoshgoftaar,Amri Napolitano +2 more
- 01 Jan 2014
TL;DR: Five wrapper-based feature selection methods to remove irrelevant and redundant features are used and it is demonstrated that BAM is the best performance metric used within the wrapper.
2
Patent
Exception prediction before an actual exception during debugging
Vikas Chandra,Sarika Sinha +1 more
- 26 Jan 2016
TL;DR: In this article, an approach is provided for predicting an exception during debugging of software code before the debugging encounters the exception, based on the execution of the upcoming lines, a prediction is determined that the exception will be encountered at line number M, which is within a range of line numbers L+1 and X.
2
References
The Imbalanced Training Sample Problem: Under or over Sampling?
TL;DR: This paper presents a study concerning the relative merits of several re-sizing techniques for handling the imbalance issue and the convenience of combining some of these techniques.
Stable Gene Selection from Microarray Data via Sample Weighting
Lei Yu,Yue Han,Michael E. Berens +2 more
TL;DR: A general framework of sample weighting is proposed to improve the stability of feature selection methods under sample variations and leads to more stable gene signatures than the state-of-the-art ensemble method, particularly for small signature sizes.
105
•Proceedings Article
A study on feature selection and classification techniques for automatic genre classification of traditional Malay music
Shyamala Doraisamy,Shahram Golzari,Noris Mohd Norowi,Md. Nasir Sulaiman,Nur Izura Udzir +4 more
- 01 Jan 2008
TL;DR: This study performs a more comprehensive investigation on improving the classification of Traditional Malay Music (TMM), identifying potentially useful classifiers and showing the impact of adding a feature selection phase for TMM genre classification.
A comparative study of iterative and non-iterative feature selection techniques for software defect prediction
TL;DR: The proposed iterative feature selection approach outperforms the non-iterative approach and is designed to find a ranked feature list which is particularly effective on the more balanced dataset resulting from sampling while minimizing the risk of losing data through the sampling step and missing important features.
84
Feature selection in proteomic pattern data with support vector machines
Kees H. de Jong,Elena Marchiori,Michèle Sebag,A. van der Vaart +3 more
- 07 Oct 2004
TL;DR: This work introduces novel methods for feature selection (FS) based on support vector machines (SVM) that combine feature subsets produced by a variant of SVM-RFE, a popular feature ranking/selection algorithm based on SVM.