Skip to main navigation Skip to search Skip to main content

End-to-end deep learning for audio classification: from waveforms to a security perspective

  • Sajjad Abdoli

Student thesis: Doctoral thesisDoctorate in Engineering: Engineering

Abstract

Audio processing is one of the challenging problems in machine learning. Deep learning models have recently had a significant impact on sound processing problems such as audio classification, speech recognition, speaker identification, etc. Most of the methods based on deep models usually rely on feeding the models by 2D representations like spectrograms. Inspired by deep learning models designed for computer vision, these 2D representations are treated as images. Recently, end-to-end audio processing models have been developed for various tasks. In this case, as a 1D vector, the raw audio signal is used as the input to deep models to benefit from the full potential of such models and eliminate signal processing modules for generating the handcrafted features from audio signals. This study explores how deep models can be used for audio processing, notably, audio classification problems. The general methodologies for audio processing based on handcrafted representations and raw audio signals as inputs to the models are reviewed. A novel end-to-end architecture based on a convolutional neural network for Environmental Sound Classification (ESC) is proposed. The proposed model eliminates the necessary signal processing modules for generating handcrafted features and outperforms most state-of-the-art approaches that use handcrafted features as input. The experimental results demonstrate the power and the potential of such an end-to-end model for the classification problem. Moreover, recent studies demonstrate the vulnerability of deep learning models to a range of adversarial attacks, which threaten the functionality and credibility of such models. Such threats cover a wide range of goals such as evasion, data, model poisoning, model extraction, etc. This study also explores such attacks focusing on end-to-end audio processing models by mentioning their strengths and weaknesses. Adversarial perturbation is one of the evasion attacks, which is a severe threat against such models. This attack injects a quasi-imperceptible perturbation to the input sample to deceive the model into misclassifying it. In this context, universal adversarial perturbation is a single adversarial perturbation that can fool the classifier for most input samples once added to the input samples. This study proposes two methodologies for crafting universal adversarial audio perturbations. One method is inspired by an iterative greedy algorithm, which is well-known in computer vision, and the other is based on a novel penalty formulation. The experimental results indicate the effectiveness of such perturbation for fooling a family of end-to-end ESC and speech recognition models. In this regard, suitable defensive mechanisms and mitigation strategies must be considered for designing and implementing such end-to-end models.
Date21 Dec 2021
Original languageAmerican English
Awarding Institution
  • École de technologie supérieure
SupervisorAlessandro Lameiras Koerich (Supervisor) & Patrick Cardinal (Co-supervisor)

Cite this

'