Skip to main navigation Skip to search Skip to main content

Automatic evaluation of Alzheimer’s disease, a multimodal analysis of spontaneous conversations

  • Arlen Pérez Arana

Student thesis: Doctoral thesisDoctorate in Engineering: Engineering

Abstract

Alzheimer’s disease (AD) patients present verbal and nonverbal communication difficulties, which has led to growing interest in the role nonverbal communication plays in the lives of people with dementia (Rousseaux et al., 2010). It is estimated that 55-97% of the message communicated in adult interaction consists of nonverbal behavior ((Gross, 1990), (Hargie et al., 1981)), which includes body movement, facial expressions (FE), touch, physical appearance, personal space, and vocal communication features such as pitch, intonation, and speech rate. As a result of the above, there are some studies related to Alzheimer’s Disease (AD), where verbal and nonverbal communication has been studied, some examples are eye movements, facial expressions, speech rate, vocal communication, and sentiment analysis during performing some tasks. According to these studies, facial expressions and acoustic features of AD patients could suggest certain characteristics in early stages of AD that can be automatically analyzed. In this thesis, we introduce a method to automatically analyze and evaluate the correlation between verbal and nonverbal and AD during video recorded natural conversations. Our objective is to automatically classify AD subjects or Healthy Controls (HC) through facial expressions features. We analyze 23 conversations, with an average duration of 16 minutes. For the purpose of the facial analysis, we tracked 3 groups of features: eye gaze and landmarks, face landmarks, and Facial Action Units (FAU). Additionally, for the purpose of the verbal, analysis we obtained 2 groups of features: silences and phonetic features (13 Mel-Frequency Cepstral Coefficients (MFCC) features). In general, we used four classifiers to discern between AD and HC: Random Forest Classifier (RFC), K Nearest Neighbor (KNN), Support Vector Machines (SVM), and Naïve Bayes (NB). Regarding the analysis with the silences and phonetic features, the best performance obtained was 81% accuracy and 91% of sensitivity-specificity rate ‘Receiver Operating Characteristic’ (ROC) curve. Likewise, the multimodal analysis showed 90% accuracy and 93% ROC curve. These results were obtained with the KNN classifier trained with all the features (verbal and nonverbal). Notably, the RFC showed the best performance in all the experiments performed in this study, training the algorithm exclusively using facial expressions and gaze features, we obtained a 93% accuracy and 98% ROC. The classification accuracy for the KNN classifier was 91%, the sensitivity-specificity was 97.7%, trained with facial landmarks as features. These results present a better performance trained with facial landmarks in comparison with training the classifier with all the features. In conclusion, facial expressions and phonetic features while speaking could provide signs of AD in early stages. In this work, we presented a methodology for discriminating between AD and HC. The principal objective while using this method is to use a non-invasive way to analyze and do classification through recording natural conversations. Consequently, we can provide clinicians a non-invasive and automatic tool for the early detection of signs of AD.
Date14 Mar 2022
Original languageAmerican English
Awarding Institution
  • École de technologie supérieure
SupervisorSylvie Ratté (Supervisor) & Luc Duong (Co-supervisor)

Cite this

'