Skip to main navigation Skip to search Skip to main content

Adaptation of deep object detectors for new modalities

  • Heitor Rapela Medeiros

Student thesis: Doctoral thesisDoctorate in Engineering: Engineering

Abstract

The performance of deep object detectors significantly deteriorates when deployed across different sensing modalities, such as RGB, infrared, and depth. This degradation arises from the shift between modalities, which drastically affects the performance of the models. Existing adaptation methods, such as pixel-level translation and feature-space alignment, are often limited to specific modalities or require extensive retraining, leading to increased computational cost and potential loss of previously acquired knowledge. In contrast to the dominant approach of adapting the model, here we investigated the potential of adapting the input, or making minor changes, preserving as much prior knowledge while incorporating new modality knowledge. In this context, this thesis studies modality adaptation strategies to bridge the gap between modalities while preserving detection performance on the source pre-trained RGB model. In this thesis, we first introduce our search and contributions. Then, in the first chapter, we provided a general background to understand the current different strategies presented in this thesis with different mechanisms used to adapt object detectors, ranging from input-level (image modification) to middle-level (mechanisms in backbones) to output-level (adaptation of boxes or pseudo-level modifications). In the second chapter, we study how to incorporate knowledge of two different modalities in a single modality-agnostic shared encoder for detectors in an efficient and powerful way. Then, in the third chapter, we explore progressive modality adaptation, first adapting the detector knowledge from the RGB source data (e.g., COCO dataset) to the target RGB dataset (e.g., LLVIP RGB) and then adapting the input from IR data (e.g., LLVIP IR) to a pseudo-RGB representation with this detector feedback. In the fourth chapter, we focused on input modality adaptation, preserving the knowledge of the source pre-trained model (e.g, COCO dataset) and adapting directly to the IR dataset (e.g., LLVIP IR), without the intermediate step of the prior chapter, and focusing on maximizing the detection performance while keeping the source zero-shot knowledge. In the fifth chapter, we explored how to incorporate language in the input modality adaptation for visual-language object detectors; therefore, our goal was still to preserve zero-shot knowledge of the detector, but also to understand how to incorporate powerful visual modality adaptation techniques, along with prompt adaptation. Our main contributions include: for the second chapter, we introduced MiPa, a mixed-patch training strategy for transformer-based object detectors that enables a single shared encoder to be modality-agnostic to RGB and infrared inputs. MiPa stochastically samples and combines complementary RGB/IR patches during training, effectively capturing cross-modal information without requiring both modalities at inference. In the fourth chapter, we introduced ModTr, a modality translation framework for adapting pre-trained RGB object detectors to new modalities, such as infrared (IR), without changing the detector’s parameters. ModTr preserves the detector’s original knowledge, enabling a single model to serve multiple modalities through dedicated translators, reducing memory and computation costs. ModTr introduces simple yet effective fusion strategies, such as the Hadamard product–based gating, to blend the translated and original inputs. In the fifth chapter, we introduced ModPrompt, a visual prompt–based framework for adapting open-vocabulary object detectors (OV-ODs) to new visual modalities, such as infrared, depth, and LiDAR, without compromising their zero-shot capabilities. Unlike pixel-level prompt strategies used in classification, ModPrompt employs an encoder–decoder visual prompt module that generates modality-specific prompts tailored to each input image. It further proposes Modality Prompt Decoupled Residuals (MPDR), which enhance adaptation by introducing lightweight, inference-friendly residual parameters, enabling modality alignment without losing pre-trained language knowledge. Finally, in the last part of this thesis, we provided an overall conclusion of our thesis and how we can leverage pre-trained RGB knowledge of detectors while we adapt to new modalities and recommendations for future work.
Date18 Nov 2025
Original languageAmerican English
Awarding Institution
  • École de technologie supérieure
SupervisorMarco Pedersoli (Supervisor) & Éric Granger (Co-supervisor)

Cite this

'