Skip to main navigation Skip to search Skip to main content

Assessment of the acute respiratory distress using a depth camera

  • Wajahat Nawaz

Student thesis: Doctoral thesisDoctorate in Engineering: Engineering

Abstract

Acute respiratory distress is an early phase of respiratory failure marked by severely impaired gas exchange, causing inadequate oxygenation of arterial blood and/or insufficient carbon dioxide removal, which can rapidly progress to respiratory failure if not treated promptly. This clinical emergency manifests through various observable signs that reflect the body’s compensatory mechanisms attempting to maintain adequate oxygenation. Common clinical manifestations include tachypnea, nasal flaring, expiratory grunting, thoracoabdominal asynchrony and chest retractions. Among these clinical indicators, signs of chest retraction serve as highly specific and sensitive markers, as they directly indicate increased work of breathing. The presence of these signs constitutes a medical emergency necessitating prompt clinical intervention. The timely and accurate detection of these critical signs is paramount for initiating appropriate therapeutic interventions to prevent respiratory failure. Current clinical practice relies primarily on visual assessment, a process where healthcare professionals physically observe patients at the bedside to score the severity of respiratory distress through identification of retraction signs. Visual assessment offers several advantages, including being non-invasive, providing immediate results, and requiring no specialized equipment. However, this manual, intermittent monitoring approach suffers from inter-observer variability and is resource-intensive, requiring continuous expert supervision. These limitations are particularly pronounced in resource-constrained settings and pandemic scenarios where clinical resources are strained. This thesis presents an artificial intelligence-based contactless acute respiratory distress (ARD) detection system that mitigates the deficiencies of visual examination by automating the assessment process. The proposed system leverages RGB-D (color and depth) camera technology to capture visual and temporal information of the patient’s chest wall in a non-invasive and continuous manner. The system further utilizes deep learning models to accurately localize chest wall regions and segment clinically meaningful temporal windows while effectively removing the motion artifacts. Advanced video analysis algorithms subsequently extract discriminative spatiotemporal features from the refined multi-modal data streams for automated respiratory distress identification. This thesis makes three key contributions through interconnected studies. First, we evaluate various deep learning-based video analysis architectures for ARD detection in case of limited clinical data settings. Our evaluation reveals that real-world clinical datasets exhibit inherent spatial biases. To address this challenge, we propose a spatial-temporal selection framework. Systematic evaluation demonstrates that clinically relevant regions and appropriate temporal window length are critical for accurate and computationally efficient detection. Further analysis reveals that models performing temporal downsampling alongside spatial feature extraction demonstrate superior performance compared to architectures that retain full temporal information. The proposed ARD system, leveraging the CSN-R101 model, attains an accuracy of an accuracy of 82%, precision of 80%, recall of 89%, and F1 score of 84. Second, we investigate multi-modal data fusion for enhanced detection accuracy. We first establish that depth information alone is insufficient for robust ARD detection. Subsequently, we demonstrate that late feature fusion of RGB and depth modalities substantially outperforms single-modality approaches, achieving 85.2% accuracy, 86.7% precision, 85.2% recall, and 85.8% F1 score, significantly improving upon the RGB-only (82.2% accuracy, 87.2% precision, 77.7% recall, 82.1% F1 score). These findings demonstrate that while depth alone is inadequate, it provides essential complementary features that significantly improve detection when combined with RGB data. Third, we address critical deployment challenges by developing a real-time, computationally efficient system for automated region-of-interest (ROI) detection and filtering of clinically irrelevant movements. We employed an oriented bounding box-based detection network that precisely localizes the thoracoabdominal region, achieving an 84% mean Average Precision (mAP) at IoU thresholds 0.5 to 0.95. This oriented approach outperforms traditional axis aligned methods by reducing false activations caused by surrounding medical equipment and environmental artifacts. Additionally, we propose an optical flow-based, region-aware clinically irrelevant movement detector that attains a 93% F1 score in identifying video segments where retraction symptoms are difficult to observable, ensuring the system focuses exclusively on diagnostically relevant periods. This thesis presents a comprehensive methodology for automated acute respiratory ARD detection system. The proposed system validates the feasibility of objective, continuous respiratory monitoring with the potential to reduce clinician workload, improve diagnostic reliability, and enable monitoring in resource-constrained healthcare settings.
Date2 Dec 2025
Original languageAmerican English
Awarding Institution
  • École de technologie supérieure
SupervisorRita Noumeir (Supervisor) & Philippe Jouvet (Co-supervisor)

Cite this

'