The study of human gestures in general and the movements of the head and hands in particular is an important part of non-verbal communication. For this reason the number of studies related to gestures analysis follows an increasing curve. Among these studies are those which concern the videos analysis in order to recognize and interpret specific information in a conversation through gestures annotation. Noting head and hand movements of elderly people can be a difficult task, since motor and sensory skills decline with age. Currently, the annotation of gestures in videos is done manually. This has several limitations. It is tedious and imprecise since it is subject to the variability of the annotators. Likewise, this type of annotation is expensive since it must be carried out only by experts. In addition, the conclusions drawn are not robust since manual annotation is often applied to small samples. In this research work, we have proposed two automatic approaches to annotate hand movements and head movements. These approaches are based on deep learning techniques namely, CNN and RNN. We use ground truth carried out by experts in order to validate the proposed approaches.
First, we proposed an approach to annotate head movements. The developed method is inspired from standards developed by experts in the linguistic field. The annotation problem has been modeled as a classification task. Each simple gesture, consisting of a single movement, constitutes a class while complex gestures, consisting of more than one movement, are classified in a single class. To do this, we implemented a method based on the MTCNN technique, this technique was used to detect the face. Then the LSTM is applied to predict the class of each movement.
Secondly, we proposed an approach to automate the gestural phases of the hands that has existed in the literature since the Eighties. Our work is based on previous studies used in gestural phases classification. The strategy used to automatically annotate the gestural phases is similar to that proposed to annotate the movements of hands. Indeed, the annotation problem is considered a classification problem. Technically, we used MobileNet to detect hands and LSTM to predict the phase in question.
The proposed process focuses on the automatic annotation of head and hands, reducing the cost and the time of the annotation process and decreasing the impact of subjectivity caused by the imprecision generated by a visual analysis of gestures on a video. The results suggest two research focuses for further exploration. First, the automatic annotation process for nonverbal behaviors in videos is quite feasible; second, machine learning algorithms have the capacity to identify features that are not totally in sync with what humans are used to. These characteristics clearly open up a new dialog between artificial intelligence and researchers in linguistics, and researchers in fields related to communication and aging in terms of interpreting results and features in videos
| Date | 29 Sept 2021 |
|---|
| Original language | French |
|---|
| Awarding Institution | - École de technologie supérieure
|
|---|
| Supervisor | Sylvie Ratté (Supervisor) & Luc Duong (Co-supervisor) |
|---|
Garraoui, H. (Author),
Ratté (Supervisor) &
Duong (Co-supervisor),
29 Sept 2021Student thesis: Doctoral thesis › Doctorate in Engineering: Engineering