Journal Article10.1007/S12065-020-00536-Z
Rank-based univariate feature selection methods on machine learning classifiers for code smell detection
Shivani Jain,Anju Saha +1 more
27
TL;DR: In this paper, the authors have implemented 32 machine learning algorithms after performing feature selection through six variations of the filter method, including Mutual Information, fisher score, and univariate ROC-AUC feature selection techniques with brute force and random forest correlation strategies.
read more
Abstract: Detecting code smells and treating them with refactoring are trivial part of maintaining vast and sophisticated software. There is an urgent need for automatic system to treat code smells. Tools provide variable results, based on threshold values and subjective interpretation of smells. Machine learning is one of the best approaches that provides effective solution to this problem. Practitioners do not need expert knowledge on smell’s characteristics for detection, which makes this approach accessible. In this paper, we have implemented 32 machine learning algorithms after performing feature selection through six variations of the filter method. We have used multiple correlation methodologies to discard similar features. Mutual information, fisher score, and univariate ROC–AUC feature selection techniques were used with brute force and random forest correlation strategies. Feature selection eliminates dimensionality curse and improves performance measures drastically. It is the selection of relevant feature subset based on the relation between dependent and independent variables. We have compared performance of classifiers implemented with and without performing feature selection. Results show that accuracy of machine learning models has increased up to 26.5%, f-measure by 70.9%, area under ROC curve has surged up to 26.74%, and average training time has reduced up to 62 s as compared to performance measures of machine learning models executed without feature selection. Mutual information feature selection strategy with random forest correlation methodology has the highest impact on performance measures among all the filter methods. Among 32 classifiers, boosted decision trees (J48) and Naive Bayes algorithms gave best performance after dimensionality reduction.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
Improving performance with hybrid feature selection and ensemble machine learning techniques for code smell detection
Shivani Jain,Anju Saha +1 more
TL;DR: In this article, three hybrid feature selection techniques with ensemble machine learning algorithms are employed to improve the performance in detecting code smells, including boosting designs, stacking methods, and bagging.
45
Key control variables affecting interior visual comfort for automated louver control in open-plan office –– a study using machine learning
TL;DR: In this article, the importance of control-related variables available for machine learning-assisted automated louvers in open-plan offices is explored, and a validated simulation was performed to generate data samples that were used for feature analysis.
25
FeatureEnVi: Visual Analytics for Feature Engineering Using Stepwise Selection and Semi-Automatic Extraction Approaches
TL;DR: FeatureEnVi as discussed by the authors is a visual analytics system specifically designed to assist with the feature engineering process, which helps users to choose the most important feature, to transform the original features into powerful alternatives, and to experiment with different feature generation combinations.
Explaining xgboost predictions with shap value: a comprehensive guide to interpreting decision tree-based models
11 Apr 2023
TL;DR: In this paper , the SHAP (SHapley Additive explanations) index is used to evaluate the influence of each feature on the forecasts made by the model and it is demonstrated that the contribution of features to model learning may be precisely estimated when utilizing SHAP values with decision tree-based models, which are frequently used to represent tabular data.
A decomposition-based multi-objective immune algorithm for feature selection in learning to rank
TL;DR: In this article, a decomposition-based multi-objective immune algorithm for feature selection in learning-to-rank (L2R) is proposed, which associates each solution with a scalar subproblem based on Tchebycheff decomposition approach, which makes the optimization process more efficient.
12
References
•Journal Article
Scikit-learn: Machine Learning in Python
Fabian Pedregosa,Gaël Varoquaux,Alexandre Gramfort,Vincent Michel,Bertrand Thirion,Olivier Grisel,Mathieu Blondel,Peter Prettenhofer,Ron Weiss,Vincent Dubourg,Jake Vanderplas,Alexandre Passos,David Cournapeau,Matthieu Brucher,Matthieu Perrot,Edouard Duchesnay +15 more
TL;DR: Scikit-learn is a Python module integrating a wide range of state-of-the-art machine learning algorithms for medium-scale supervised and unsupervised problems, focusing on bringing machine learning to non-specialists using a general-purpose high-level language.
•Book
Data Mining: Concepts and Techniques
Jiawei Han,Micheline Kamber,Jian Pei +2 more
- 08 Sep 2000
TL;DR: This book presents dozens of algorithms and implementation examples, all in pseudo-code and suitable for use in real-world, large-scale data mining projects, and provides a comprehensive, practical look at the concepts and techniques you need to get the most out of real business data.
The WEKA data mining software: an update
TL;DR: This paper provides an introduction to the WEKA workbench, reviews the history of the project, and, in light of the recent 3.6 stable release, briefly discusses what has been added since the last stable version (Weka 3.4) released in 2003.
Classification and Regression by randomForest
Andy Liaw,Matthew C. Wiener +1 more
- 01 Jan 2007
TL;DR: random forests are proposed, which add an additional layer of randomness to bagging and are robust against overfitting, and the randomForest package provides an R interface to the Fortran programs by Breiman and Cutler.
An introduction to variable and feature selection
Isabelle Guyon,André Elisseeff +1 more
TL;DR: The contributions of this special issue cover a wide range of aspects of variable selection: providing a better definition of the objective function, feature construction, feature ranking, multivariate feature selection, efficient search methods, and feature validity assessment methods.