TY - GEN
T1 - CLIP-IT
T2 - 2026 IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2026
AU - Karimian, Banafsheh
AU - Avanzato, Giulia
AU - Belharbi, Soufiane
AU - Guichemerre, Alexis
AU - McCaffrey, Luke
AU - Shateri, Mohammadhadi
AU - Granger, Eric
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026
Y1 - 2026
N2 - Multimodal learning has shown promise in medical imaging, combining complementary modalities like images and text. Vision-language models (VLMs) capture rich diagnostic cues but often require large paired datasets and promptor text-based inference. Their practicality is therefore limited due to annotation cost, privacy, and compute demands. Unpaired external text, like pathology reports, can still provide complementary diagnostic cues if semantically relevant content is retrievable per image. To address this, we introduce CLIP-IT, a novel framework that relies on rich unpaired text reports. Specifically, CLIP-IT uses a CLIP model pre-trained on histology image-text pairs from a separate dataset to retrieve the most relevant unpaired textual report for each image in the downstream unimodal dataset. These reports, sourced from the same disease domain and tissue type, form pseudo-pairs that reflect shared clinical semantics rather than exact alignment. Knowledge from these texts is distilled into the vision model during training, while LoRA-based adaptation mitigates the semantic gap between unaligned modalities. At inference, only the vision model is used, maintaining low overhead while still benefiting from multimodal training without requiring paired data in the downstream dataset. Experiments1 show that CLIP-IT consistently improves classification accuracy over both unimodal and multimodal CLIP-based baselines in most cases, without requiring paired annotations per dataset or incurring additional inference-time complexity.
AB - Multimodal learning has shown promise in medical imaging, combining complementary modalities like images and text. Vision-language models (VLMs) capture rich diagnostic cues but often require large paired datasets and promptor text-based inference. Their practicality is therefore limited due to annotation cost, privacy, and compute demands. Unpaired external text, like pathology reports, can still provide complementary diagnostic cues if semantically relevant content is retrievable per image. To address this, we introduce CLIP-IT, a novel framework that relies on rich unpaired text reports. Specifically, CLIP-IT uses a CLIP model pre-trained on histology image-text pairs from a separate dataset to retrieve the most relevant unpaired textual report for each image in the downstream unimodal dataset. These reports, sourced from the same disease domain and tissue type, form pseudo-pairs that reflect shared clinical semantics rather than exact alignment. Knowledge from these texts is distilled into the vision model during training, while LoRA-based adaptation mitigates the semantic gap between unaligned modalities. At inference, only the vision model is used, maintaining low overhead while still benefiting from multimodal training without requiring paired data in the downstream dataset. Experiments1 show that CLIP-IT consistently improves classification accuracy over both unimodal and multimodal CLIP-based baselines in most cases, without requiring paired annotations per dataset or incurring additional inference-time complexity.
KW - histopathology classification
KW - knowledge distillation
KW - multimodal learning
UR - https://www.scopus.com/pages/publications/105041248813
U2 - 10.1109/WACV61042.2026.00361
DO - 10.1109/WACV61042.2026.00361
M3 - Contribution to conference proceedings
AN - SCOPUS:105041248813
T3 - Proceedings - 2026 IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2026
SP - 3700
EP - 3709
BT - Proceedings - 2026 IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2026
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 6 March 2026 through 10 March 2026
ER -