Visible-infrared person re-identification (V-I ReID) has emerged as a crucial task for modern surveillance systems, requiring the accurate matching of individuals captured across different modalities—visible (V) and infrared (I) cameras. However, significant challenges arise from the substantial domain gap between modalities, resulting in pronounced discrepancies in appearance, texture, and illumination. This gap severely hampers the generalization ability of deep learning models, leading to degraded cross-modal retrieval performance. Furthermore, conventional feature extraction approaches often lack discriminability due to their reliance on global representations, overlooking fine-grained, identity-specific attributes critical for robust person matching. Existing evaluation protocols exacerbate these limitations by assuming a unimodal gallery during testing, which is unrealistic for real-world 24-hour surveillance systems that naturally operate under mixed-modality conditions.
This thesis, titled Gradual domain generalization for visible-infrared person re-identification, systematically addresses these challenges through a gradual and structured domain generalization framework. Three major contributions are proposed: (1) Adaptive generation of privileged intermediate information (AGPI2), introducing dynamic intermediate modalities to guide the learning of modality-invariant features; (2) Bidirectional multi-step domain generalization (BMDG), progressively refining part-based feature alignments to bridge domain gaps more effectively; and (3) Mixed-modality re identification (MixER), disentangling modality-specific and identity-specific features to enhance performance under realistic mixed-gallery conditions. Extensive experiments conducted on benchmark datasets such as SYSU-MM01, RegDB, and LLCM demonstrate that the proposed approaches achieve state-of-the-art results across a variety of challenging scenarios, including occlusion, extreme lighting variations, and mixed-modal settings.
Beyond cross-modal person re-identification, this thesis also investigates a complementary and timely problem: enhancing the interpretability and interoperability of deep vision models. Inspired by the prototype discovery mechanisms introduced for ReID, a novel framework for interpretable classification is developed, enabling the automatic discovery of part prototypical primitive concepts without requiring manual annotations. This approach not only improves the transparency and faithfulness of deep models but also ensures robust and stable explanations even under challenging conditions. By bridging the gap between performance and interpretability, this thesis contributes toward the development of more trustworthy and practically deployable deep learning systems for real world applications.
Alehdaghi, M. (Author),
Granger (Supervisor) &
Menelau Oliveira Cruz (Co-supervisor),
12 Jan 2026Student thesis: Doctoral thesis › Doctorate in Engineering: Engineering