TY - GEN
T1 - Predicting Intermittent Job Failure Categories for Diagnosis Using Few-Shot Fine-Tuned Language Models
AU - Aïdasso, Henri
AU - Bordeleau, Francis
AU - Tizghadam, Ali
N1 - Publisher Copyright:
© 2026 Copyright held by the owner/author(s).
PY - 2026/7/17
Y1 - 2026/7/17
N2 - In principle, failures in Continuous Integration (CI) pipelines provide valuable feedback to developers about code-related errors. In practice, however, pipeline jobs often fail intermittently due to non-deterministic tests, network outages, infrastructure failures, resource exhaustion, and other reliability issues. These intermittent (flaky) job failures lead to substantial inefficiencies: wasted computational resources from repeated reruns and significant diagnosis time that distracts developers from core activities and often requires intervention from specialized teams. Prior studies have proposed machine learning techniques to detect intermittent failures, but the subsequent diagnosis remains underexplored. To fill this gap, we introduce FlaXifyer, a few-shot learning approach for predicting intermittent job failure categories using pre-trained language models. FlaXifyer requires only job execution logs and achieves 84.3% Macro F1 and 92.0% Top-2 accuracy with just 12 labeled examples per category. We also propose LogSift, an interpretability technique that identifies influential log statements in under one second, reducing review effort by 74.4% while surfacing relevant failure information in 87% of cases. Evaluated on 2,458 job failures from TELUS, FlaXifyer and LogSift enable practitioners to automate triage and accelerate the diagnosis of intermittent job failures.
AB - In principle, failures in Continuous Integration (CI) pipelines provide valuable feedback to developers about code-related errors. In practice, however, pipeline jobs often fail intermittently due to non-deterministic tests, network outages, infrastructure failures, resource exhaustion, and other reliability issues. These intermittent (flaky) job failures lead to substantial inefficiencies: wasted computational resources from repeated reruns and significant diagnosis time that distracts developers from core activities and often requires intervention from specialized teams. Prior studies have proposed machine learning techniques to detect intermittent failures, but the subsequent diagnosis remains underexplored. To fill this gap, we introduce FlaXifyer, a few-shot learning approach for predicting intermittent job failure categories using pre-trained language models. FlaXifyer requires only job execution logs and achieves 84.3% Macro F1 and 92.0% Top-2 accuracy with just 12 labeled examples per category. We also propose LogSift, an interpretability technique that identifies influential log statements in under one second, reducing review effort by 74.4% while surfacing relevant failure information in 87% of cases. Evaluated on 2,458 job failures from TELUS, FlaXifyer and LogSift enable practitioners to automate triage and accelerate the diagnosis of intermittent job failures.
KW - CI
KW - classification
KW - failure diagnosis
KW - few-shot learning
KW - intermittent job failures
KW - interpretability
KW - language models
KW - logs
UR - https://www.scopus.com/pages/publications/105045841876
U2 - 10.1145/3803437.3805224
DO - 10.1145/3803437.3805224
M3 - Contribution to conference proceedings
AN - SCOPUS:105045841876
T3 - FSE Companion 2026 - Proceedings of the 34th ACM International Conference on the Foundations of Software Engineering
SP - 513
EP - 523
BT - FSE Companion 2026 - Proceedings of the 34th ACM International Conference on the Foundations of Software Engineering
A2 - Tan, Shin Hwei
A2 - Khomh, Foutse
PB - Association for Computing Machinery, Inc
T2 - ACM International Conference on the Foundations of Software Engineering, FSE 2026
Y2 - 5 July 2026 through 9 July 2026
ER -