Ultrasound (US) imaging has emerged as a valuable tool in speech sciences, offering a non invasive, real-time window into tongue movements during articulation. For articulatory analysis, and particularly in visual biofeedback systems for second language (L2) learning or clinical intervention, the hard palate trace can provide a significant added value. It serves not only as a passive target for many consonants but also as a stable frame of reference to normalize articulatory measurements, compare productions, and guide learners toward more targeted gestures. However, the use of the palate trace in biofeedback is limited by a major challenge : its invisibility during speech, caused by the air-tissue interface that blocks the ultrasound waves. Existing methods, which are few, are therefore limited to static reconstruction of the contour, often acquired through a swallowing task.
This thesis proposes two complementary contributions. First, we establish best practices for reliable palate tracing. Our comparative analysis (51 swallowing videos, 17 participants, 3 tasks, 3 methods ) demonstrates that the dry swallow task combined with the Cumulative Echo Skeleton (CES) method yields the best inter-rater agreement (mean error : 2.87 mm) and validates the viability of the automatic CES approach (mean error : 2.63 mm).
Second, this thesis introduces an automatic palate tracking method. The approach relies on inferring palatal motion from a consistently visible anatomical landmark : the genioglossus tendon. A hybrid system, combining a YOLOv8 detector with a particle filter, ensures robust tendon tracking, allowing the palate’s position to be inferred via a rigid transformation model. Evaluated on 71 videos (swallowing and free speech), the method achieves mean errors from 1.34 to 2.68 mm and remains reliable even when palate visibility drops below 5%. The hypothesis of anatomical inference is validated by a significant correlation (r = 0.64, p = 0.001) between tendon and palate tracking accuracy.
These advances lay the foundation for improved visual biofeedback systems for speech therapy and L2 learning. The first study establishes a best practice (dry swallow and CES) for obtaining a reliable reference trace and validates the automatic CES method, while noting its limitations with artifacts caused by the presence of liquids. The second contribution builds on this to propose continuous tracking despite invisibility. The potential of this tracking is demonstrated by the ReaPT prototype, developed as part of this project and praised by users. Current limitations, such as the rigid transformation hypothesis being affected by jaw movement, open clear future directions : non-rigid models, re-initialization, and clinical validation.
| Date | 19 Dec 2025 |
|---|
| Original language | French |
|---|
| Awarding Institution | - École de technologie supérieure
|
|---|
| Supervisor | Catherine Laporte (Supervisor), Lucie Ménard (Co-supervisor) & Walcir Cardoso (Co-supervisor) |
|---|
Ben Asker, H. (Author),
Laporte (Supervisor), Ménard (Co-supervisor) & Cardoso (Co-supervisor),
19 Dec 2025Student thesis: Master's thesis › Master in Engineering: Information Technology Engineering