Person re-identification (ReID) is a crucial video surveillance task, allowing one to match images of individuals captured by non-overlapping cameras. This task poses significant challenges due to factors such as varying camera positions, capture conditions (e.g., illumination, weather, background), complex body shapes, and diverse clothing styles. These factors result in a wide range of potential capture conditions leading to datasets that only cover a small fraction of the potential scenarios, and consequently to models’ learning and evaluation data unsuited to the design of a robust framework. Under these constraints, a cost-effective model must be built to capture the complex and discriminative personal features while allowing to perform real-time data processing.
Among the previous aspects, the visible modality commonly used in traditional ReID frameworks is highly dependent on the prevailing luminosity. Low-light conditions can severely affect the quality of captured scenes, resulting in inaccurate ReID. This introduces an additional obstacle in the person ReID task, compounded by the potential presence of noisy or blurry captures. Infrared cameras can mitigate the issues caused by lighting conditions because they do not rely on light information for scene encoding, but do not capture color information and are similarly affected by sensor encoding issues. Therefore, visible and infrared sensors are versatile and relevant sensors in the context of person ReID, but relying solely on one of these modalities compromises the effectiveness of the framework in outdoor conditions or under complex capture conditions.
In this thesis, the multimodal setting and especially the fusion of visible and infrared (V-I) modalities are considered to address these challenges. By having a distinct encoding process while capturing videos from the same individuals, V-I cameras allow for correlated captures while limiting the effect of eventual encoding issues.
Chapter 1 provides some background on deep learning models and techniques for person ReID. Then, a review of multimodal fusion techniques and real-world data for person ReID is provided in Chapter 2, allowing us to assess the key challenges in the area. As such, multiple aspects must be considered. First, multimodal fusion models are seen to neglect modality-specific features, mostly focusing on shared knowledge instead, and consequently missing an important part of discriminant information. In addition, recent approaches emphasize the need to deepen evaluation protocols by artificially corrupting datasets but also to implement specific learning strategies in this regard. However, while such approaches have been explored and shown to be effective in the unimodal setting, this has yet to be done for fusion algorithms, similarly affected by real-world unpredictability.
This work presents a novel multimodal model architecture Chapter 3 to fully exploit modality knowledge and handle real-world data. The model is composed of three backbones, two concentrate on extracting modality-specific features, while the third leverage shared knowledge from a fused modality representation. Furthermore, attention-based approaches are investigated to enable dynamic feature selection, which is likely suitable for multimodal feature fusion under challenging operational conditions. These conditions are reproduced through the proposed V-I corrupted datasets that replicate realistic and highly challenging conditions for both co-located and not co-located camera scenarios, allowing an in-depth model evaluation. For co-located cameras, eventual corruptions correlations are considered, not expected, and consequently not applied for not co-located cameras since each V and I cameras are at distinct locations. Finally, a multimodal data augmentation that enhances the multimodal model’s capacity for generalization is proposed and works at promoting collaboration among modalities and priming the model to face modality-specific local or global corruptions.
The use of three datasets and two corrupted evaluation scenarios through twenty V and I corruptions allowed us to show that the multimodal ReID strategy can improve ReID accuracy while conserving moderate system complexity. Specifically, with the appropriate learning approach, the proposed multimodal model can outperform related state-of-the-art systems under ideal and challenging noisy real-world conditions.
Josi, A. (Author),
Granger (Supervisor) &
Menelau Oliveira Cruz (Co-supervisor),
16 Aug 2023Student thesis: Master's thesis › Master in Engineering: Automated Manufacturing Engineering