Ultrasound imaging is a helpful tool to observe tongue movements while minimally interfering with natural speech. There exists a variety of models to quantify tongue shape based on contours extracted from ultrasound images. However, these can be affected by poor image quality, e.g., when parts of the tongue are missing from the images due to imaging artifacts. In this study, we investigate the effects of various contour extraction errors on the accuracy and consistency of different shape measures.
We developed exponential and polynomial contour perturbation models, then simulated missing tongue tip and root, and investigated the impact of these perturbations on shape measures based on the discrete Fourier transform (DFT), modified curvature index (MCI), and triangular fitting. This was applied to a set of CV utterances collected from healthy speakers and speakers who were diagnosed with speech deficits. Results demonstrate the effectiveness of DFT, MCI and triangular fitting in clustering different phonemes despite the added noise. There is also a trade-off between the robustness of the model and sensitivity to minor actual differences in tongue shape. Sometimes, these slight differences help group tongue shapes that differ, e.g., due to coarticulation effects. Therefore, we have attempted to improve the precision of the DFT model by adding palatal contact information. Our experiment shows that the new shape model is robust to the noise and can successfully classify our target CV utterances and increase the classification score by 23% for speakers with speech deficits.
| Date | 11 Jul 2022 |
|---|
| Original language | American English |
|---|
| Awarding Institution | - École de technologie supérieure
|
|---|
| Supervisor | Catherine Laporte (Supervisor) & Lucie Ménard (Co-supervisor) |
|---|
Changizi, S. (Author),
Laporte (Supervisor) & Ménard (Co-supervisor),
11 Jul 2022Student thesis: Master's thesis › Master in Engineering: Engineering