Swallowing is one of the most frequently performed motor functions in the body and constitutes a major clinical marker of neurological and behavioural health. Both its frequency and temporal dynamics are altered by dysphagia, neurodegenerative diseases and ageing, and are also modulated by emotional state, cognitive load and food intake. Reference clinical tools provide brief observations of swallowing in controlled environments and do not allow continuous monitoring in daily life. Neck-worn sensors, although promising, remain limited due to limited social acceptability for prolonged use.
Hearables, wearable devices designed for the ear, constitute a socially acceptable platform for non-invasive physiological monitoring. In particular, in-ear microphones, placed in the occluded ear canal, exploit the occlusion effect to amplify sounds transmitted via bone and tissue conduction, including swallowing sounds. However, the captured signal is a heterogeneous bioacoustic mixture (breathing, cardiovascular sounds, speech, chewing, physical movements) within which swallowing constitutes a minority event class, brief (average duration ∼860 ms) and highly variable between subjects. Added to this challenge is the absence of public databases of in-ear audio synchronised with swallowing annotations validated by reference measurements.
This thesis proposes a complete framework for swallowing detection from bilateral in-ear audio signals, structured around two complementary contributions. The first is the design and validation of a multimodal database collected from 34 healthy adults (14 women, 20 men, aged 20 to 29 years), comprising synchronised recordings of in-ear and outer-ear audio signals as well as multimodal recordings from a thoracic belt (electrocardiogram, respiratory signals and accelerometry), and tongue ultrasound imaging serving as annotation reference. The protocol covers verbal and non-verbal tasks under three acoustic conditions (silent, babble noise, factory noise) and several bolus types (dry, water, yoghurt). As a proof of concept, the first contribution presents a fully connected network trained on Yet Another Mobile Network (YAMNet) vector representations supplemented by the zero-crossing rate, which achieves an F1 score of 0.875 ± 0.013 for binary swallowing classification.
The second contribution is an event detection system based on a bilateral Convolutional Recurrent Neural Network (CRNN) architecture operating on Per Channel Energy Normalization (PCEN) spectrograms, with confidence-based adaptive fusion of both ears and hysteresis post-processing. Evaluated on the 34 participants by five-fold subject disjoint cross-validation, the system achieves a window-level classification F1 score of 0.892 ± 0.006 and an event-level detection F1 score of 0.835 ± 0.012, with a precision of 0.779 ± 0.014 and a recall of 0.900 ± 0.020, recall consistently exceeding 0.870 across all noise conditions. These results demonstrate the feasibility of robust swallowing detection from in-ear audio in realistic acoustic conditions and pave the way for longitudinal monitoring of dysphagia and sialorrhoea using hearables.
Ben Cheikh, E. (Author),
Bouserhal (Supervisor) &
Laporte (Co-supervisor),
24 Jun 2026Student thesis: Master's thesis › Master in Engineering: Engineering