Skip to main navigation Skip to search Skip to main content

Reconnaissance multi-dimensionnelle de l’émotion par apprentissage profond de caractéristiques spatio-temporelles sur séquences vidéo

Translated title of the thesis: Multi-dimensional emotion recognition with deep learning of spatio-temporal features on video sequences
  • Thomas Teixeira

Student thesis: Master's thesisMaster in Engineering: Electrical Engineering

Abstract

Affect computing and emotion recognition have shown an increased interest in several research areas for the past decades. Notably, facial expressions are one of the most powerful ways for depicting specific patterns in human behavior and describing human emotional state. Nevertheless, even for human, identifying facial expressions is difficult, and automatic facial expression recognition (FER) systems based on images have often suffered from a lack of various and cross-cultures training data. With the slight shift to video sequences with in-the-wild settings and more complex emotion representation such as dimensional models, deep FER systems has the ability to learn more accurate and discriminative features. Furthermore, most models, based on Convolutional Neural Networks (CNNs) and combined with Recurrent Neural Networks (RNNs), have been proposed for recognizing emotions but often lied on short video sequences for categorical model predictions. And still, few studies are interested in 3D-CNN models for recognizing emotion and based on multi-dimensional representation. Moreover, few pre-trained 3D-CNN models are currently available for FER tasks. Which make the development of 3D-CNN more complex, regarding the amount of available data. Furthermore, most models, based on Convolutional Neural Networks (CNNs) and combined with Recurrent Neural Networks (RNNs), have been proposed for recognizing emotions but often lied on short video sequences for categorical model predictions. And still, few studies are interested in 3D-CNN models for recognizing emotion and based on multi-dimensional representation. Moreover, few pre-trained 3D-CNN models are currently available for FER tasks. Which make the development of 3D-CNN more complex, regarding the amount of available data. Firstly, our study describe the different main stages for the design of deep FER models (preprocessing, transfer learning, post-processing). Then, we detail, the development steps of each architecture, and the related variables for our approach. The design of i3D models showed particularly flexible regarding the initialization of model parameters and allowed us to develop a new fine tuning method of deep architecture. Thanks to the weight inflation method, it is possible to make distinction between initial 2D weights and extended weights, thus differenciating weights associated to the spatial domain from weights associated to the temporal domain. Finally, the last part, details the experimental results of our different approaches validating several assumptions from the litterature regarding convolutional models.
Date10 Sept 2020
Original languageFrench
Awarding Institution
  • École de technologie supérieure
SupervisorAlessandro Lameiras Koerich (Supervisor) & Éric Granger (Co-supervisor)

Cite this

'