Résumé
Software defect prediction helps practitioners prioritize their quality assurance efforts by identifying defective components in their software projects. Recent studies proposed various machine learning and deep learning-based approaches to identify whether a given artifact is defective or non-defective. However, these approaches tend to focus on model performance, which typically leads to inherently complex models that lack transparency, in turn, limiting their adoption by developers in practice. Although various model-agnostic approaches (e.g.,LIME and BreakDown) have been proposed to provide local explanations for such complex models, recent studies show that these approaches are inconsistent and tend to disagree in several cases. In this study, we investigate whether it is possible to learn simple and interpretable software defect prediction models while maintaining prominent predictive performance. To this end, we propose a novel approach, namely CASDP (Complexity Aware software defect prediction), by formulating software defect prediction as a multi-objective search-based problem to find the best trade-off between the model’s performance and complexity. To evaluate our approach, we conducted an empirical study to investigate the trade-off between model performance and complexity. We compare our proposed approach with different state-of-the-art machine learning, deep learning and search-based approaches. The obtained results indicate that CASDP achieves competitive performance scores that reach the state-of-the-art ensemble models (e.g.,Random Forest and Deep Forest) performance while learning models as compact as Fast and Frugal trees models. In conclusion, our findings demonstrate that we can learn simple software defect prediction models with reduced complexity and limited degradation of performance.
| langue originale | Anglais |
|---|---|
| Numéro d'article | 5 |
| journal | Empirical Software Engineering |
| Volume | 32 |
| Numéro de publication | 1 |
| Les DOIs | |
| état | Publié - févr. 2027 |
Empreinte digitale
Voici les principaux termes ou expressions associés à « Less is more: balancing models performance and complexity for software defects prediction ». Ces libellés thématiques sont générés à partir du titre et du résumé de la publication. Ensemble, ils forment une empreinte digitale unique.Citer cette ressorce
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS