This thesis concerns the Arabic dialect identification task which is still unresolved due to the high similarity between the Arabic dialects. We conduct the experiments on a small and a large databases. The main novelty of the thesis is the use of dynamic selection which has never been used for language identification to the best of our knowledge. Dynamic selection allows to choose one or many classifiers for each observation to be classified.We want to know its potential for the present task. This thesis also proposes to improve the results in two ways. First, we use transfer learning to improve the results of the small database by using the large one. Then, we use the deep metric learning by changing the traditional softmax cost function by triplets because we think that this will have a positive impact on the dynamic selection. The results showed that dynamic selection has potential for this task even if it has been under-exploited because we obtained results comparable to those of the best teams who worked on these databases. Future works will need to find a distance metric suited to the problem and make classifiers more diverse. We have also noted that transfer learning has greatly contributed to the improvement of the small database while we obtained mixed results with the deep metric learning.
Thibault, P.-M. (Author),
Cardinal (Supervisor) &
Menelau Oliveira Cruz (Co-supervisor),
20 Jan 2022Student thesis: Master's thesis › Master in Engineering: Information Technology Engineering