查看文章

mdpi.com 中的 [HTML]

Impact of feature selection methods on the predictive performance of software defect prediction models: an extensive empirical study

作者

Abdullateef O Balogun, Shuib Basri, Saipunidzam Mahamad, Said J Abdulkadir, Malek A Almomani, Victor E Adeyemo, Qasem Al-Tashi, Hammed A Mojeed, Abdullahi A Imam, Amos O Bajeh

发表日期

2020/7/9

期刊

Symmetry

卷号

期号

页码范围

1147

出版商

MDPI

简介

Feature selection (FS) is a feasible solution for mitigating high dimensionality problem, and many FS methods have been proposed in the context of software defect prediction (SDP). Moreover, many empirical studies on the impact and effectiveness of FS methods on SDP models often lead to contradictory experimental results and inconsistent findings. These contradictions can be attributed to relative study limitations such as small datasets, limited FS search methods, and unsuitable prediction models in the respective scope of studies. It is hence critical to conduct an extensive empirical study to address these contradictions to guide researchers and buttress the scientific tenacity of experimental conclusions. In this study, we investigated the impact of 46 FS methods using Naïve Bayes and Decision Tree classifiers over 25 software defect datasets from 4 software repositories (NASA, PROMISE, ReLink, and AEEEM). The ensuing prediction models were evaluated based on accuracy and AUC values. Scott–KnottESD and the novel Double Scott–KnottESD rank statistical methods were used for statistical ranking of the studied FS methods. The experimental results showed that there is no one best FS method as their respective performances depends on the choice of classifiers, performance evaluation metrics, and dataset. However, we recommend the use of statistical-based, probability-based, and classifier-based filter feature ranking (FFR) methods, respectively, in SDP. For filter subset selection (FSS) methods, correlation-based feature selection (CFS) with metaheuristic search methods is recommended. For wrapper feature selection (WFS …

引用总数

被引用次数：58

202020212022202320243 23 13 11 8

学术搜索中的文章

Impact of feature selection methods on the predictive performance of software defect prediction models: an extensive empirical study

AO Balogun, S Basri, S Mahamad, SJ Abdulkadir… - Symmetry, 2020

被引用次数：58 相关文章所有 11 个版本