Speech is shaped not only by internal cognitive and motor processes but also by the external sensory conditions and technological tools through which it is produced and recorded. Contemporary developments such as increasingly immersive communication environments and widespread use of in-ear wearable devices (hearable) have introduced new challenges and opportunities for understanding speech in contexts that more closely reflect everyday life. The overarching objective of this thesis is to investigate how auditory–visual integration, altered listening conditions, and novel recording methods influence speech production and its analysis, with the goal of advancing theoretical models and informing applications in communication and health monitoring.
The first study examines the multisensory basis of speech control by investigating how visual and auditory characteristics of a room jointly affect speech level regulation. Using immersive virtual reality (VR) environments with varying acoustics and visuals, it is shown that both modalities shape vocal output dynamically, with auditory information exerting a stronger influence but visual information modulating speech earlier in time.
The second study addresses the combined effects of noise, ear occlusion, and hearing impairment on speech production. A new bilingual speech corpus including the use of hearable devices was developed, incorporating recordings across systematically varied listening conditions and a wide range of hearing thresholds. Analyses revealed complex individual differences in speech level regulation, including reduced reactivity to noise in participants with greater hearing impairment under high-occlusion conditions. These findings highlight the need for individualized rather than group-level modeling approaches.
The third study evaluates how novel in-ear microphones (IEMs) and outer-ear microphones (OEMs) in hearable devices capture acoustic measures of voice quality as compared to standard laboratory microphones. Results indicate systematic discrepancies across recording methods, highlighting the importance of developing new standards for voice evaluation with hearables.
Together, these studies extend our understanding of how speech production is regulated under varied environmental, sensory, and technological constraints. By situating speech within the multisensory and technological conditions of modern communication, the thesis contributes to theoretical models of speech motor control and provides empirical insights for applications in virtual communication, occupational safety, and wearable technologies.
| Date | 28 Nov 2025 |
|---|
| Original language | American English |
|---|
| Awarding Institution | - École de technologie supérieure
|
|---|
| Supervisor | Rachel Bouserhal (Supervisor) & Ingrid Verduyckt (Co-supervisor) |
|---|
Zhang, X. (Author),
Bouserhal (Supervisor) & Verduyckt (Co-supervisor),
28 Nov 2025Student thesis: Doctoral thesis › Doctorate in Engineering: Engineering