In this thesis, focus is set on spatiotemporal 3D convolutional neural networks (3D CNNs) for facial expression recognition (FER) in videos. Over the last decade, deep learning has emerged as a state-of-the-art paradigm for FER and spatiotemporal recognition. The transition of research focus toward deep learning was also a regression from spatiotemporal to spatial methods. Models are increasingly dependent on the quantity of training data, favouring 2D-image datasets over videos which are more difficult to collect, label and process.
Different approaches to spatiotemporal FER are evaluated, leveraging pretrained deep-learning models. The 3D-convolution paradigm for video classification is analyzed, questioning the relevance of considering spatial and temporal dimensions of video data as forming a unified 3D volume. To cope with the computational requirements and scarcity of data, clip sampling is commonly adopted for training 3D CNNs. To increase performance within this framework, a new method is developed. The proposed temporal stochastic softmax is based on a weighted clip-sampling mechanism. This method allows the model to focus on the most relevant clips, for efficient training and accurate recognition. Experiments are carried out on several video classification tasks, focusing on facial expression recognition, and discussions are provided concerning the relevance of such weighted temporal sampling and pooling mechanisms in addressing common issues such as occlusion, inaccurate trimming and coarse annotation of videos, and uneven distribution of discriminant cues across time. In addition, the study explores visual attention mechanisms, as a way to implement more complex weighting behaviors for masking or highlighting regions of the input videos. The different attentional behaviors that can be developed with 3D CNNs are analyzed, and their relevance is discussed in the context of spatiotemporal recognition and FER specifically. Experiments notably demonstrate the benefits of guiding attention with contextual representations, summarizing the global information of the video. The study discusses the relative importance of temporal frames in a video, to address the heterogeneous distribution of relevant cues in time. The proposed methods increase the performance of 3D CNNs on all benchmarks, providing better ways to learn from data. Such methods for efficient spatiotemporal recognition should become more and more important as larger datasets become available in the future, allowing richer training of 3D CNNs.
| Date | 10 Jun 2021 |
|---|
| Original language | American English |
|---|
| Awarding Institution | - École de technologie supérieure
|
|---|
| Supervisor | Éric Granger (Supervisor), Marco Pedersoli (Co-supervisor) & Simon Bacon (Co-supervisor) |
|---|
Ayral, T. (Author),
Granger (Supervisor),
Pedersoli (Co-supervisor) & Bacon (Co-supervisor),
10 Jun 2021Student thesis: Master's thesis › Master in Engineering: Automated Manufacturing Engineering