Skip to main navigation Skip to search Skip to main content

Utilisation des caractéristiques prosodiques pour optimiser un système de compréhension du langage naturel

Translated title of the thesis: Use of prosodic features to optimize a natural language understanding system
  • Simon Boutin

Student thesis: Master's thesisMaster in Engineering: Information Technology Engineering

Abstract

In general, companies working in the human-machine dialog systems industry offer several computer applications to their customers, including automatic natural language understanding. Current human-machine dialog systems are composed of three weakly coupled components: -The automatic speech recognition system (ASR); -The natural language understanding system (NLU) -The dialog system or conversational agent (CA). In this architecture, the outputs of the preceding components are the inputs of the following. The characteristics of the acoustic signal are not included in the output of the first component. But it is possible that additional information in the original signal can directly help the NLU system to perform its task. This so called "prosodic" information concerns intonation, intensity and duration of sound, which is obviously absent from the written text. For example, the identification of free text in a voice command is particularly difficult for the current NLU system. The literature does not directly address the identification of free texts. The most similar concept is the identification of quotations. By focusing on free texts, the originality of this study is that the author of the quotation and its narrator correspond to the same entity. The first objective of this thesis was to determine whether there is a correlation between prosodic information of an acoustic signal and the presence or absence of free texts. Three types of prosodic features were extracted from a large set of voice commands, and their sample distributions were examined. Distributions for free texts were compared to those of other concepts using the two-sample Kolmogorov-Smirnov test (K-S test). The results showed that there was indeed a correlation. The second objective was to verify whether it is possible to improve the performance of an NLU system through the integration of prosodic information. Given a minimal NLU system, the performance gains of models based on lexical features alone were compared against models augmented with prosodic features. The McNemar's test was used to verify whether the gains were significant. Prosodic information has indeed improved this system’s performance.
Date27 Jun 2016
Original languageFrench
Awarding Institution
  • École de technologie supérieure
SupervisorPierre Dumouchel (Supervisor), Réal Tremblay (Co-supervisor) & Patrick Cardinal (Co-supervisor)

Cite this

'