Skip to main navigation Skip to search Skip to main content

Classification of nonverbal human-produced audio events

  • Philippe Chabot

Student thesis: Master's thesisMaster in Engineering: Electrical Engineering

Abstract

Noise Induced Hearing Loss due to excessive noise exposure in the workplace is affecting an increasing number of workers and can be reduced by properly using Hearing Protection Devices (HPD). These protection devices are often found in the form of intra-aural (plugs) or circumaural (earmuffs) devices. When intra-aural devices are properly inserted, they create an occlusion effect that amplifies physiological noise and makes it easier to be perceived by the wearer. These amplified physiological noises can be captured using a microphone placed inside the occluded ear canal and with a detection algorithm, these nonverbal audio events could be used for many applications, such as: the user could interact with an audio wearable device in a discreet manner by clicking his teeth or tongue; the user’s health could be monitored by detecting coughing or clearing of the throat events and the worker’s noise dose exposure calculation could be more accurately measured by the removal of the wearers’ own noises, which are harmless to his hearing. The objective of this project are threefold: 1) build a classification algorithm of nonverbal human-produced audio events 2) build a nonverbal audio event detection algorithm and 3) validate the performances and ensure that these algorithms can run in real-time on a low computational power device. Ten nonverbal audio events were selected from an existing database featuring nonverbal events recorded using an in-ear microphone. In total, 3037 samples, each 400 ms in length were extracted from the database in order to train the classification algorithm. To create the audio event classifier, several state-of-the-art machine learning algorithms found in the literature such as the Support Vector Machine, Gaussian Mixture Model, Multilayer Perceptron, Convolutional Neural Network and Bag-of-Audio-Words (BoAW) were successively implemented and tested. To provide input for these algorithms, three types of features found in the literature were tested: the Mel-Frequency Ceptral Coeffients (MFCC), mainly used for speech recognition, the Auditory-inspired Amplitude Modulation Features typically used for speaker verification on whispered speech and Per-Channel Energy Normalization (PCEN) features often used for far-field keyword spotting. Optimal performance of the classification algorithm was found using the BOAW classifier coupled with the MFCC and PCEN features with a sensitivity of 81.5% and a precision of 83%. The real-time detector was tested in a noisy environment on 10 new test subjects. It showed a sensitivity of 69.9% and a precision of 78.9% in a quiet environment. This makes it a promising algorithm to be implemented in new lines of smart digital HPD capable of protecting worker’s hearing while opening the way for a new range of health monitoring and human-machine interfaces.
Date8 Oct 2020
Original languageAmerican English
Awarding Institution
  • École de technologie supérieure
SupervisorJérémie Voix (Supervisor) & Patrick Cardinal (Co-supervisor)

Cite this

'