Diffusion models achieve state-of-the-art performance in multimodal generation. Yet, their practical adoption remains limited by (i) the high computational cost of iterative sampling and (ii) the difficulty of adapting a pretrained diffusion model to new domains from few references. This thesis investigates the use of diffusion models for adaptation, personalized multimodal generation, and editing in image and sound-effect generation under low-data regimes, with a focus on efficient adaptation, identity preservation, and controllable variation generation for production workflows.
First, this thesis presents contributions stemming from a collaborative work on the Uni DAD paper. It introduces a unified training framework for simultaneous distillation and adaptation of diffusion models, enabling few-shot and few-step image generation. The work contrasts classical two-stage pipelines of adapt-then-distill or distill-then-adapt, which entail complex designs, overfitting, and limited diversity. The contribution presented here focuses on subject-driven personalization (SDP), enhanced by integrating textual conditioning. Evaluated on the DreamBooth benchmark, Uni-DAD SDP achieves competitive image quality compared to reference adaptation methods while requiring only a single sampling step. The end result is a few-step generator capable of adapting and producing diverse, high-quality images in novel domains under few-shot conditions. Ultimately, this checkpoint-agnostic method facilitates the deployment of diffusion models in personalized, real-time applications.
Second, this work contributes to applied research on audio generation for sound effect (SFX) production. In particular, it investigates the ability of modern generative models to produce diverse variations from a reference clip while preserving the sound event’s identity. Experiments analyze the models’ capacity to reproduce or edit key signal characteristics, such as temporal structure and energy curve, under production-oriented constraints (e.g., duration control, style transfer, and alignment cues). Furthermore, this analysis helps structure an overview and enables fair comparisons between existing methods for SFX variation and editing, thereby bridging recent advances in audio generation with practical industry requirements. This work also establishes a foundation for future applied research on efficient, controllable variation generation in production settings.
By bridging the efficiency of diffusion-based generation and the expressiveness of reference-guided modeling across modalities, these contributions enable user-friendly, controllable adaptation and editing of both images and SFX clips.
| Date | 11 May 2026 |
|---|
| Original language | French |
|---|
| Awarding Institution | - École de technologie supérieure
|
|---|
| Supervisor | Mohammadhadi Shateri (Supervisor) & Éric Granger (Co-supervisor) |
|---|
Desbos, M. (Author),
Shateri (Supervisor) &
Granger (Co-supervisor),
11 May 2026Student thesis: Master's thesis › Master in Engineering: Automated Manufacturing Engineering