Skip to main navigation Skip to search Skip to main content

Dealing with domain shift in deep learning: from training-time generalization to test-time adaptation

  • Mehrdad Noori

Student thesis: Doctoral thesisDoctorate in Engineering: Engineering

Abstract

Deep learning models have achieved remarkable progress across a wide range of computer vision tasks, from image classification and object detection to segmentation. Despite these advances, most models are developed under the simplifying assumption that the data observed during deployment will resemble that seen during training. In real-world scenarios, this assumption is rarely valid. Real-world conditions, such as variations in illumination, imaging sensors, viewpoints, or textures, can substantially alter the data distribution and lead to a significant degradation in model performance. This vulnerability raises critical concerns about the robustness and reliability of vision systems deployed in the wild. To address this limitation, this thesis explores methods that move beyond the traditional training centered paradigm toward models capable of maintaining performance under novel and unseen conditions. Specifically, it investigates two directions: Domain Generalization (DG), which aims to learn domain-invariant representations during training, and Test-Time Adaptation (TTA), which adjusts models dynamically during inference using only unlabeled test data. Together, these approaches address the growing need for reliable and adaptive models in both traditional vision architectures and modern foundation models such as vision–language systems. In the first part, we study DG and propose two novel methods. (1) TFS-ViT (Token-Level Feature Stylization for Vision Transformers) introduces the first token-level feature stylization framework for Vision Transformers, mixing normalization statistics across samples to enforce structure rather than texture-dependent representations. An attention-aware variant further exploits class-token saliency to guide stylization toward semantically relevant regions, achieving state of-the-art generalization across standard DG benchmarks. (2) FDS (Feedback-Guided Domain Synthesis) presents a diffusion-based framework that trains a single multi-source conditional model to generate pseudo-domains spanning inter-domain gaps. A feedback-driven filtering mechanism selects challenging synthetic samples that explicitly encourage domain-invariant feature learning, yielding substantial robustness gains while incurring no inference-time cost. With the advent of large-scale foundation models, which are pretrained once and reused across diverse tasks, it becomes crucial to develop mechanisms that enable adaptation at test time without access to source data. This motivates the second part of this thesis, which explores fully test-time adaptation strategies for Vision–Language Models (VLMs). (3) MLMP is the f irst TTA framework for Open-Vocabulary Semantic Segmentation (OVSS), combining adaptive multi-layer fusion with multi-prompt optimization to exploit VLMs’ inherent prompt sensitivity as a stable adaptation signal. Furthermore, we establish the first comprehensive OVSS-TTA benchmark covering nine datasets and over eighty test scenarios, providing a standardized protocol for future research in this field. (4) Histopath-C introduces the first benchmark for evaluating VLM TTA in digital histopathology, simulating realistic clinical domain shifts such as stain variation, blur, and contamination, thereby providing a valuable foundation for studying model robustness under clinically relevant domain shifts. Building upon this benchmark, we also propose LATTE (Low-rank Adaptation with Transductive Template Ensembling), a simple yet powerful adaptation strategy that combines multiple textual templates with low-rank updates to enhance model stability and robustness. This histopathology-specific method achieves significant performance improvements across diverse datasets and represents one of the most realistic and practically important applications of TTA.
Date4 Mar 2026
Original languageAmerican English
Awarding Institution
  • École de technologie supérieure
SupervisorChristian Desrosiers (Supervisor) & Ismail Ben Ayed (Co-supervisor)

Cite this

'