The induction of multi-media documents arouses currently a great interest on the experimental as well as theoretical level. Particularly, the détection of key words in sound files is a sector in fuU progress. However, despite the progress realized in the field of voice indexation, much to be donc remains and in particular for the search of key words in the spontaneous speech.
Our work presented in this manuscript registers within the Framework of the voice annotations indexation in a context of documentary management. First of ail we will present some automatic speech recognition Systems. Based on selection criterion, we hâve identified two speech recognition engines. They are the subject of our experiments.
Then, we will propose a keyword détection system in the voice annotations. This latter will be based on the two automatic speed récognition engines which we have chosen; namely the engines of Dragon Naturally Speaking and the one of Microsoft. In order to test the performance of the two Systems, we have built a corpus of voice annotations. The evaluation of the transcription performances was realized through being based upon the percentage of correct words and that of precision. On the other hand, the evaluation of the indexation performances was realized through being based on ROC curves and the recall and precision rates.
The best results were observed with the Microsoft recognition engine for the profile without training. While for the trained profile, the Dragon engine presents the best performances. So in order to improve the performances, we propose to involve the model language with a great corpus of written annotations text.
| Date | 8 Jul 2010 |
|---|
| Original language | French |
|---|
| Awarding Institution | - École de technologie supérieure
|
|---|
| Supervisor | Pierre Dumouchel (Supervisor) |
|---|
Ouali, C. (Author),
Dumouchel (Supervisor),
8 Jul 2010Student thesis: Master's thesis › Master in Engineering: Engineering