Passer à la navigation principale Passer à la recherche Passer au contenu principal

High-Rate Mixout: Revisiting Mixout for Robust Domain Generalization

  • École de technologie supérieure

Résultats de recherche: Chapitre dans un livre, rapport, actes de conférenceParticipation à un ouvrage collectif lié à un colloque ou une conférenceRevue par des pairs

Résumé

Ensembling fine-tuned models initialized from powerful pre-trained weights is a common strategy to improve robustness under distribution shifts, but it comes with substantial computational costs due to the need to train and store multiple models. Dropout offers a lightweight alternative by simulating ensembles through random neuron deactivation; however, when applied to pre-trained models, it tends to over-regularize and disrupt critical representations necessary for generalization. In this work, we investigate Mixout, a stochastic regularization technique that provides an alternative to Dropout for domain generalization. Rather than deactivating neurons, Mixout mitigates overfitting by probabilistically swapping a subset of fine-tuned weights with their pre-trained counterparts during training, thereby maintaining a balance between adaptation and retention of prior knowledge. Our study reveals that achieving strong performance with Mixout on domain generalization benchmarks requires a notably high masking probability of 0.9 for ViTs and 0.8 for ResNets. While this may seem like a simple adjustment, it yields two key advantages for domain generalization: (1) higher masking rates more strongly penalize deviations from the pre-trained parameters, promoting better generalization to unseen domains; and (2) high-rate masking substantially reduces computational overhead, cutting gradient computation by up to 45% and gradient memory usage by up to 90%. Experiments across five domain generalization benchmarks - PACS, VLCS, OfficeHome, TerraIncognita, and DomainNet - using ResNet and ViT architectures show that our approach, High-rate Mixout, achieves out-of-domain accuracy comparable to ensemble-based methods while significantly reducing training costs.

langue originaleAnglais
titreProceedings - 2026 IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2026
EditeurInstitute of Electrical and Electronics Engineers Inc.
Pages3803-3812
Nombre de pages10
ISBN (Electronique)9798331555115
Les DOIs
étatPublié - 2026
Evénement2026 IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2026 - Tucson, Etats-Unis
Durée: 6 mars 202610 mars 2026

Série de publications

NomProceedings - 2026 IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2026

Conférence

Conférence2026 IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2026
Pays/TerritoireEtats-Unis
La villeTucson
période6/03/2610/03/26

Empreinte digitale

Voici les principaux termes ou expressions associés à « High-Rate Mixout: Revisiting Mixout for Robust Domain Generalization ». Ces libellés thématiques sont générés à partir du titre et du résumé de la publication. Ensemble, ils forment une empreinte digitale unique.

Citer cette ressorce