Automatic emotions recognition (AER) from speech is a challenging task especially when dealing with real-life affective expressions. Spontaneous emotions are often subtle, sometimes mixed, of short periods, with large intra-class variability, in addition to have a skewed class distribution. It is in this context that our objective to propose a methodology capable of improving the performance of current AER systems is inscribed.
The proposed methodology is motivated by prior knowledge on theoretical models of emotion in psychology. The idea is to integrate the concepts of dimensional emotion model in the design of discrete emotions classifiers. Two concepts were identified from the dimensional model: the existence of a dimensional space in which categorical emotions can be projected and the existence of a similarity relationship between these categories of emotion with respect to each of these dimensions. The first concept leads to the extraction of high-level features that are intended to play a role similar to that played by the dimensions of the theoretical model. The second concept has motivated the adoption of a similarity-based approach for emotions representation and classification. We have shown that the likelihood scores generated by GMM models are powerful similarity-based features for AER task and responds well to the issue of the short duration length of utterances.
We have proposed a first method of classification, entitled weighted ordered class-nearest neighbors. This method is built around a new feature vector describing an utterance by its pattern of neighboring emotion classes. The classes inside the pattern are ordered according to their proximities and estimated on the likelihood scores basis. Unlike the Bayes decision rule, the ranks of all scores influence the classification decision. Two types of models have been proposed and tested: linear and nonlinear.
We also proposed anchor models as emotion classification method but which can be also used as a tool for emotional content analysis in psychology studies. The utterances are projected in a continuous space where each dimension is spanned by an emotion class model that measures the similarity level of an utterance with respect to this class. We have shown that it is also possible to successfully apply anchor models for multi-class problem context as for a binary classification one by expanding the anchor space with new external models. We analyzed and compared Euclidean- and cosine-based anchor models performances based on geometric properties of their decision boundaries. Furthermore, we showed that the Anchor models can also be used as a powerful method of combining classifiers subject to scores normalization more suited for the fusion context. Their performances and their properties (e.g., insensitivity to skewed class distribution) make of anchor models very suitable solutions for AER task compared to more complex systems.
Finally, in terms of acoustic descriptors, new and more discriminative features have been proposed. The results achieved by fusion of these features using the anchor models outperformed the state-of-the-art when tested on FAU Emotions AIBO, a benchmark spontaneous emotion corpus for the AER research community.
| Date | 30 Nov 2015 |
|---|
| Original language | French |
|---|
| Awarding Institution | - École de technologie supérieure
|
|---|
| Supervisor | Pierre Dumouchel (Supervisor) |
|---|
Attabi, Y. (Author),
Dumouchel (Supervisor),
30 Nov 2015Student thesis: Doctoral thesis › Doctorate in Engineering: Engineering