TY - GEN
T1 - Pretraining Helps When Capacity Allows
T2 - 2026 IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2026
AU - Muralidharan, Srikanth
AU - Medeiros, Heitor R.
AU - Aminbeidokhti, Masih
AU - Granger, Eric
AU - Pedersoli, Marco
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026
Y1 - 2026
N2 - Robust visual recognition on embedded platforms requires models that both generalize out-of-distribution (OOD) and fit into tiny compute/memory budgets. While pre-training is a standard route to robustness for mid/large backbones, its value in the ultra-small regime remains unclear. We present a capacity-aware study of pre-training for two efficient ConvNet families (EfficientNet and MobileNetV3) scaled from "small"to "ultra-small"via a simple, reproducible recipe. We compare three initializations - ImageNet→COCO pretraining, ImageNet classification pretraining, and training from scratch - on two axes of distribution shift: (i) cross-dataset RGB→RGB transfer between LLVIP and FLIR (ii) cross-modality detection where models are fine-tuned on RGB and evaluated on infrared (IR). A complementary classification study on DomainNet probes whether the trends extend beyond detection. Across settings, we find that pretraining's benefit is conditional on both backbone capacity and shift difficulty. Task-aligned Imagenet→COCO pretraining is the most reliable starting point at moderate sizes and for the easier transfer direction. In the low-capacity regimes, differences are typically within run-to-run variation, and training from scratch can match or surpass pre-training. Classification mirrors this capacity gating. Our results test the premise "pretraining always helps"and instead quantify when task-aligned pretraining pays off for ultra-small backbones and when it likely does not 1.
AB - Robust visual recognition on embedded platforms requires models that both generalize out-of-distribution (OOD) and fit into tiny compute/memory budgets. While pre-training is a standard route to robustness for mid/large backbones, its value in the ultra-small regime remains unclear. We present a capacity-aware study of pre-training for two efficient ConvNet families (EfficientNet and MobileNetV3) scaled from "small"to "ultra-small"via a simple, reproducible recipe. We compare three initializations - ImageNet→COCO pretraining, ImageNet classification pretraining, and training from scratch - on two axes of distribution shift: (i) cross-dataset RGB→RGB transfer between LLVIP and FLIR (ii) cross-modality detection where models are fine-tuned on RGB and evaluated on infrared (IR). A complementary classification study on DomainNet probes whether the trends extend beyond detection. Across settings, we find that pretraining's benefit is conditional on both backbone capacity and shift difficulty. Task-aligned Imagenet→COCO pretraining is the most reliable starting point at moderate sizes and for the easier transfer direction. In the low-capacity regimes, differences are typically within run-to-run variation, and training from scratch can match or surpass pre-training. Classification mirrors this capacity gating. Our results test the premise "pretraining always helps"and instead quantify when task-aligned pretraining pays off for ultra-small backbones and when it likely does not 1.
KW - domain adaptation
KW - object detection
KW - out of distribution robustness
UR - https://www.scopus.com/pages/publications/105041306005
U2 - 10.1109/WACV61042.2026.00804
DO - 10.1109/WACV61042.2026.00804
M3 - Contribution to conference proceedings
AN - SCOPUS:105041306005
T3 - Proceedings - 2026 IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2026
SP - 8333
EP - 8342
BT - Proceedings - 2026 IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2026
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 6 March 2026 through 10 March 2026
ER -