Skip to main navigation Skip to search Skip to main content

Less is more: balancing models performance and complexity for software defects prediction

  • Moataz Chouchen
  • , Ali Ouni
  • , Gopi Krishnan Rajbahadur
  • , Ahmed E. Hassan
  • Concordia University
  • Huawei Technologies Co., Ltd.
  • Queen's University Kingston

Research output: Contribution to journalJournal Articlepeer-review

Abstract

Software defect prediction helps practitioners prioritize their quality assurance efforts by identifying defective components in their software projects. Recent studies proposed various machine learning and deep learning-based approaches to identify whether a given artifact is defective or non-defective. However, these approaches tend to focus on model performance, which typically leads to inherently complex models that lack transparency, in turn, limiting their adoption by developers in practice. Although various model-agnostic approaches (e.g.,LIME and BreakDown) have been proposed to provide local explanations for such complex models, recent studies show that these approaches are inconsistent and tend to disagree in several cases. In this study, we investigate whether it is possible to learn simple and interpretable software defect prediction models while maintaining prominent predictive performance. To this end, we propose a novel approach, namely CASDP (Complexity Aware software defect prediction), by formulating software defect prediction as a multi-objective search-based problem to find the best trade-off between the model’s performance and complexity. To evaluate our approach, we conducted an empirical study to investigate the trade-off between model performance and complexity. We compare our proposed approach with different state-of-the-art machine learning, deep learning and search-based approaches. The obtained results indicate that CASDP achieves competitive performance scores that reach the state-of-the-art ensemble models (e.g.,Random Forest and Deep Forest) performance while learning models as compact as Fast and Frugal trees models. In conclusion, our findings demonstrate that we can learn simple software defect prediction models with reduced complexity and limited degradation of performance.

Original languageEnglish
Article number5
JournalEmpirical Software Engineering
Volume32
Issue number1
DOIs
Publication statusPublished - Feb 2027

!!!Keywords

  • Explainable software analytics
  • Genetic programming
  • Software quality assurance

Fingerprint

Dive into the research topics of 'Less is more: balancing models performance and complexity for software defects prediction'. These topics are generated from the title and abstract of the publication. Together, they form a unique fingerprint.

Cite this