Automatic speech recognition (ASR) is a technology widely used in daily life, but not completely solved. ASR systems are still prone to errors, especially when confronted with nonstandard conditions, different from those used to train them. This technology is especially challenged when used with speech from non-English speakers and aged voices. In certain domains, such as the study of neurodegenerative diseases, it is known that language impairments appear in early stages of the disease, and that the analysis of patients' narrative discourse helps to obtain a timely diagnosis. Manual analysis, as it is done so far, is costly in terms of time and resources, so automatic speech recognition could make the process more efficient. However, the high error rates in these systems prevent them from being widely used in science and research.
In this paper, we propose a new post-editing method of error detection and correction for an ASR system that generates automatic transcriptions of the speech of French-speaking adults and older adults describing an image.
By means of natural language processing techniques, we extract the most common vocabulary from correct manual transcriptions to build a phonemicized correction dictionary; then we extract out-of-context sentences from the automatic transcriptions, which are then compared through a fuzzy phonetic search with the correction dictionary to find and apply the best corrections. Experimental results show an error detection accuracy of 80% and our best system yields an average WER improvement of 1.9%, with values ranging from 0.6% to 6.4%.
| Date | 26 Jan 2022 |
|---|
| Original language | French |
|---|
| Awarding Institution | - École de technologie supérieure
|
|---|
| Supervisor | Sylvie Ratté (Supervisor) |
|---|
García Cano Castillo, E. U. (Author),
Ratté (Supervisor),
26 Jan 2022Student thesis: Master's thesis › Master in Engineering: Information Technology Engineering