Skip to main navigation Skip to search Skip to main content

Deep regression models for spatio-temporal expression recognition in videos

  • Gnana Praveen Rajasekhar

Student thesis: Doctoral thesisDoctorate in Engineering: Engineering

Abstract

Automatic expression recognition (ER) is a challenging problem in the field of affective computing, playing an important role in human behavior understanding in, e.g., human-computer interaction, sociable robots, and driver assistance. ER can be formulated as the problem of classification or regression of expressions. Though regression of expressions plays a crucial role in many healthcare applications, such as estimating pain and fatigue levels, it remains relatively less explored compared to the classification of expressions. Fatigue detection is widely used in applications such as autonomous driving and employee engagement. Similarly, automatic pain assessment has an important potential diagnostic value for infants, young children, and people with communicative or neurological impairments. Fatigue is synchronous with pain, where high fatigue is associated with high pain, which can be found with the correlation of Visual Analog Scores (VASs) of fatigue and pain. Often pain expressions happen over a shorter period of time, while fatigue happens over a longer duration. Some of the major challenges in dealing with regression of expressions are subtle variations across individuals, ambiguity across the contiguous frames pertinent to the intensities of expressions, identity bias, and sensor capture conditions. Moreover, most deep learning (DL) models demand a huge amount of data with annotations, which requires a lot of human support with domain expertise. Therefore, leveraging DL models for the regression of expressions with limited annotations remains to be a major bottleneck. Although audio-visual fusion is expected to outperform the unimodal performance, failing to efficiently leverage the complementary relationship across the audio and visual modalities often results in poor performance. This Thesis focus on the development of DL models for two problems: (1) weakly supervised domain adaptation (WSDA) for estimating the levels of pain and fatigue and (2) audio-visual (A-V) fusion for dimensional emotion recognition. As a first contribution, a detailed review of weakly supervised learning (WSL) approaches is presented for facial behavior analysis. To provide a comprehensive review, action units (AUs), which is defined by the fundamental actions of individual facial movements or a group of facial movements, are also included along with expressions for both classification and regression. In particular, a taxonomy of methods in the literature for different WSL scenarios has been provided, along with their respective strengths and limitations. A review of widely used public datasets, experimental protocols, and experimental results is also provided for the evaluation of these state-of-the-art methods. Finally, our critical analysis of these methods provides insight into the potential research directions to leverage weakly-labeled data for facial behavior analysis. This review concludes that although WSL methods are promising in handling the weak labels of facial expressions in real-world scenarios, they are not effectively explored in the literature, and there is much room for advancing the state-of-the-art facial ER performance given data with weak annotations. As a second contribution, a novel DL model for WSDA with ordinal regression (WSDA-OR) is proposed to estimate the levels of pain and fatigue from videos. DA has been widely explored to alleviate the problem of domain shifts that typically occur between video data captured across various source (laboratory) and target (operational) domains. In this work, WSDA is leveraged to adapt a DL model to different persons and capture conditions when the videos are weakly annotated. Contrary to prior state-of-the-art WSL models for estimating pain intensity in videos, the proposed model enforces the ordinal relationship among the pain intensity levels of the target sequences along with the temporal coherence of multiple consecutive frames. In particular, it learns discriminant and domain-invariant feature representations by integrating multiple-instance learning with deep adversarial DA, where soft Gaussian labels are used to efficiently represent weak ordinal sequence-level labels from the target domain. Experimental results on UNBC-McMaster, BIOVID, and Fatigue (private) datasets indicate that our proposed approach can significantly improve performance over state-of-the-art models, allowing us to achieve a greater pain localization accuracy. As a third contribution, a joint cross-attention model is proposed for A-V fusion in dimensional ER based on facial and vocal modalities. Most state-of-the-art methods for A-V fusion rely on recurrent networks or conventional attention mechanisms that do not effectively leverage the complementary nature of A-V modalities. In this work, the complementary relationship across A-V modalities is effectively explored to extract the salient features, allowing for accurate prediction of continuous values of valence and arousal. Experimental results on RECOLA and Affwild2 indicate that our joint cross-attentional A-V fusion model provides a cost-effective solution that can outperform state-of-the-art approaches. The work described in this Thesis indicates that efficiently adapting DL models with weakly labeled videos shows significant improvement over prior state-of-the-art methods for estimating pain and fatigue levels. This work shows that there is much room to further improve the proposed WSDA model to leverage the potential of DL models for unsupervised domain adaptation for the regression of expressions. This work has further shown that leveraging the complementary relationship across A and V modalities is a promising research direction for effective AV fusion. The proposed joint cross-attentional approach can also be further improved using gating mechanisms for effective modeling of intra and intermodal relationships as well as to handle corrupted modalities.
Date2 Jun 2023
Original languageAmerican English
Awarding Institution
  • École de technologie supérieure
SupervisorÉric Granger (Supervisor) & Patrick Cardinal (Co-supervisor)

Cite this

'