Passer à la navigation principale Passer à la recherche Passer au contenu principal

CLIP-IT: CLIP-based Pairing of Histology Images with Privileged Textual Information

  • École de technologie supérieure
  • University of Cagliari
  • McGill University

Résultats de recherche: Chapitre dans un livre, rapport, actes de conférenceParticipation à un ouvrage collectif lié à un colloque ou une conférenceRevue par des pairs

Résumé

Multimodal learning has shown promise in medical imaging, combining complementary modalities like images and text. Vision-language models (VLMs) capture rich diagnostic cues but often require large paired datasets and promptor text-based inference. Their practicality is therefore limited due to annotation cost, privacy, and compute demands. Unpaired external text, like pathology reports, can still provide complementary diagnostic cues if semantically relevant content is retrievable per image. To address this, we introduce CLIP-IT, a novel framework that relies on rich unpaired text reports. Specifically, CLIP-IT uses a CLIP model pre-trained on histology image-text pairs from a separate dataset to retrieve the most relevant unpaired textual report for each image in the downstream unimodal dataset. These reports, sourced from the same disease domain and tissue type, form pseudo-pairs that reflect shared clinical semantics rather than exact alignment. Knowledge from these texts is distilled into the vision model during training, while LoRA-based adaptation mitigates the semantic gap between unaligned modalities. At inference, only the vision model is used, maintaining low overhead while still benefiting from multimodal training without requiring paired data in the downstream dataset. Experiments1 show that CLIP-IT consistently improves classification accuracy over both unimodal and multimodal CLIP-based baselines in most cases, without requiring paired annotations per dataset or incurring additional inference-time complexity.

langue originaleAnglais
titreProceedings - 2026 IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2026
EditeurInstitute of Electrical and Electronics Engineers Inc.
Pages3700-3709
Nombre de pages10
ISBN (Electronique)9798331555115
Les DOIs
étatPublié - 2026
Evénement2026 IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2026 - Tucson, Etats-Unis
Durée: 6 mars 202610 mars 2026

Série de publications

NomProceedings - 2026 IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2026

Conférence

Conférence2026 IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2026
Pays/TerritoireEtats-Unis
La villeTucson
période6/03/2610/03/26

Empreinte digitale

Voici les principaux termes ou expressions associés à « CLIP-IT: CLIP-based Pairing of Histology Images with Privileged Textual Information ». Ces libellés thématiques sont générés à partir du titre et du résumé de la publication. Ensemble, ils forment une empreinte digitale unique.

Citer cette ressorce