Passer à la navigation principale Passer à la recherche Passer au contenu principal

A Closer Look at the Few-Shot Adaptation of Large Vision-Language Models

  • École de technologie supérieure

Résultats de recherche: Chapitre dans un livre, rapport, actes de conférenceParticipation à un ouvrage collectif lié à un colloque ou une conférenceRevue par des pairs

48 Citations (Scopus)

Résumé

Efficient transfer learning (ETL) is receiving increasing attention to adapt large pre-trained language-vision models on downstream tasks with a few labeled samples. While significant progress has been made, we reveal that state-of-the-art ETL approaches exhibit strong performance only in narrowly-defined experimental setups, and with a careful adjustment of hyperparameters based on a large corpus of labeled samples. In particular, we make two interesting, and surprising empirical observations. First, to out-perform a simple Linear Probing baseline, these methods require to optimize their hyper-parameters on each target task. And second, they typically underperform -sometimes dramatically- standard zero-shot predictions in the presence of distributional drifts. Motivated by the unrealistic assumptions made in the existing literature, i.e., access to a large validation set and case-specific grid-search for optimal hyperparameters, we propose a novel approach that meets the requirements of real-world scenarios. More concretely, we introduce a CLass-Adaptive linear Probe (CLAP) objective, whose balancing term is optimized via an adaptation of the general Augmented Lagrangian method tailored to this context. We comprehensively evaluate CLAP on a broad span of datasets and scenarios, demonstrating that it consistently outperforms SoTA approaches, while yet being a much more efficient alternative. Code available at https://github.com/jusiro/CLAP.

langue originaleAnglais
titreProceedings - 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024
EditeurIEEE Computer Society
Pages23681-23690
Nombre de pages10
ISBN (Electronique)9798350353006
Les DOIs
étatPublié - 2024
Evénement2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024 - Seattle, Etats-Unis
Durée: 16 juin 202422 juin 2024

Série de publications

NomProceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition
ISSN (imprimé)1063-6919

Conférence

Conférence2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024
Pays/TerritoireEtats-Unis
La villeSeattle
période16/06/2422/06/24

Empreinte digitale

Voici les principaux termes ou expressions associés à « A Closer Look at the Few-Shot Adaptation of Large Vision-Language Models ». Ces libellés thématiques sont générés à partir du titre et du résumé de la publication. Ensemble, ils forment une empreinte digitale unique.

Citer cette ressorce