Today, many domains and communication mediums such as telecommunications, speech recognition and audio-visual systems use speech enhancement as a way of improving the quality of speech signals, typically by reducing the level of background noise. Speech signals often contain underlying noise, originating either from the acquisition process or the transmission channel. In recent years, there has been significant research in the field of speech enhancement using machine learning techniques. These techniques have been used in many speech processing tasks since they have provided very satisfactory results. Accordingly, in this thesis, the main objective of our project has been to improve speech signals, using a framework based on machine learning using neural networks.
Speech signals are composed of speech and non-speech segments and in speech enhancement, classifying speech and non-speech segments of a speech signal is an important task as it helps for targeted enhancement of speech signals. Our model uses an algorithm with a windowing process which improves its accuracy compared to other methods. It involves dividing the input signal into short-time frames or windows and analyzing each window separately to determine if it contains a speech or non-speech signal.
Our neural network-based framework has been implemented in order to fulfill two fundamental tasks. The first task is to classify speech and non-speech signals and the second task consists of enhancing the speech and non-speech signals in order to obtain an improved speech signal as a result. To achieve the classification, we have used the NOIZEUS dataset for training our models. We successfully developed a comprehensive framework that classifies speech and non-speech segments of the noisy speech signals. The enhancement relies on a criterion that is based on the type of each window. This approach allows us to apply specific enhancement methods to different segments, resulting in a fully enhanced and denoised final signal. Our results have shown that the classification scheme of speech and non-speech signals together with our cleaning strategy on signals corrupted by additive noise have been very effective, obtaining a really improved speech signal in terms of SNR and listening quality.
| Date | 8 Apr 2024 |
|---|
| Original language | American English |
|---|
| Awarding Institution | - École de technologie supérieure
|
|---|
| Supervisor | Gheorghe Marcel Gabrea (Supervisor) |
|---|
Darabpour, A. (Author),
Gabrea (Supervisor),
8 Apr 2024Student thesis: Master's thesis › Master in Engineering: Electrical Engineering