Skip to main navigation Skip to search Skip to main content

Vital signs estimation using remote photoplethysmography rPPG

  • Mohamed Khalil Ben Salah

Student thesis: Doctoral thesisDoctorate in Engineering: Engineering

Abstract

Vital sign monitoring in Pediatric Intensive Care Units (PICUs) is critical for managing vulnerable pediatric patients. Conventional approaches, such as electrocardiography, rely on physical contact and are often invasive, costly, and unsuitable for newborns or patients with contagious conditions. Remote photoplethysmography (rPPG) offers a non-contact alternative by capturing subtle variations in skin color caused by pulsatile blood flow. In pediatric intensive care, it provides a safer solution than adhesive sensors, which can cause irritation and increase the risk of infection. However, deploying rPPG in real clinical environments remains challenging due to frequent occlusions from medical equipment, patient motion, illumination variability, and a domain gap between controlled laboratory data and PICU recordings. These limitations are compounded by the scarcity of annotated clinical datasets. Addressing these constraints requires models that are physiologically interpretable, computationally efficient, and resilient to domain shifts. This thesis introduces a unified framework that integrates efficient spatiotemporal feature learning, anatomically consistent region detection, and curriculum-based self-supervised pretraining to achieve accurate and real-time estimation of heart rate in complex clinical environments. To extract reliable rPPG signals from unconstrained facial videos, a hybrid architecture is proposed that combines 3D convolutional blocks with temporal difference kernels (3DCDC-T) and multi-head self-attention from vision transformers. The model captures local spatiotemporal gradients indicative of blood volume changes while modeling longer-range dependencies required to resolve complete cardiac cycles. Attention mechanisms further refine feature focus on physiologically informative facial regions, and the feed-forward design ensures computational efficiency by limiting the transformer’s input to compact feature embeddings. Evaluated on public datasets, the model achieves an MAE of 0.79 bpm and RMSE of 0.80 bpm, with a Pearson correlation of 0.99, improving over existing methods both in accuracy and inference cost. Accurate rPPG estimation requires stable anatomical tracking of face and thoracoabdominal regions, particularly in videos affected by rotation, bed tilt, or caregiver occlusion. A dedicated detection module is developed using the Divided Space–Time Mamba (DST-Mamba) model. This architecture decouples spatial and temporal processing through Selective State Space Models (SSMs), enabling linear-time complexity and low-latency inference across longer video sequences. The model predicts oriented bounding boxes (OBBs) to preserve rotation alignment under non-standard camera angles and integrates RGB-D inputs to improve robustness against visual occlusions. DST-Mamba achieves 0.96 [email protected] and 0.95 rotated IoU on a clinical dataset, maintaining temporal stability while operating at 23 FPS on standard hardware. To mitigate the scarcity of labeled PICU data, a curriculum-based self-supervised learning strategy is introduced. A Mamba-based adaptive masking controller assigns spatiotemporal importance scores to input patches and applies strategic masking using differentiable Gumbel sampling. This adversarial masking forces the model to reconstruct physiological signals from degraded inputs, encouraging robustness to clinical occlusions and distractions. The learning process follows a structured curriculum: initial training on public datasets, simulation of occlusion patterns observed in PICU recordings, and domain adaptation on 500 unlabeled clinical videos. A lightweight teacher–student distillation module transfers physiological priors from expert models. This pipeline reduces supervised data requirements by 80%, achieving an MAE of 3.2 bpm using only 160 labeled patients, compared to 18.2 bpm with direct supervised training. The framework is validated on an extensive dataset collected at CHU Sainte-Justine, demonstrat ing generalization across ages, skin tones, and occlusion conditions. The system maintains MAE under 7.2 bpm with over 70% facial occlusion, achieving 3.8 bpm for neonates and 3.5 bpm for mechanically ventilated patients. It operates in real-time within clinical constraints, consuming 169.7 GFLOPs and 6.1 GB memory at 30 FPS throughput. Together, these contributions address key barriers to clinical rPPG deployment, including domain adaptation, anatomical tracking, and data efficiency, moving non-contact physiological monitoring toward practical use in pediatric intensive care.
Date1 Dec 2025
Original languageAmerican English
Awarding Institution
  • École de technologie supérieure
SupervisorRita Noumeir (Supervisor) & Philippe Jouvet (Co-supervisor)

Cite this

'