Environmental sound classification (ESC) and automatic speech recognition (ASR) have always attracted increasing interest from industry and academia due to their extensive range of practical applications in real-life. For instance, multimedia sensor networks, surveillance systems, and voice assistance applications embedded into our smartphones unanimously employ ESC and ASR models. Regarding the significant progress made over the last few decades, the recognition accuracy of the cutting-edge classifiers introduced in these domains has competitively reached to human-level of understanding. However, these state-of-the-art data-driven models are intensely vulnerable against adversarial signals, which are carefully crafted to fool the classifiers toward any incorrect output phrases. Technically, an adversarial signal carries a slight perturbation achievable through an optimization formulation, and it forces the recognition model to predict incorrect outputs as predefined by an adversary. This poses a major security concern since adversarial signals are not detectable by subjective evaluations either. Moreover, these malicious signals are bijectively transferable to both 1D (i.e., Mel-frequency cepstral coefficient - MFCC) and 2D representations (2D spectrograms) such as short-time Fourier and discrete wavelet transforms. Since the majority of the advanced ESC and ASR models are trained on representations, hence such adversarial spectrograms can effectively debase the recognition accuracy of these models. Unfortunately, there is a limited number of investigations on defending classifiers against various targeted and non-targeted adversarial attacks. Additionally, these approaches might not be reliable enough to secure models against strong white and black-box attacks.
Since there is no standard definition for the reliability of an adversarial defense algorithm, we define our implications from reliability and impose three main conditions. Firstly, a reliable defense algorithm should avoid any filtration operations resulting in obfuscating gradient information or shattering the Jacobian matrix. Secondly, it should make a reasonable tradeoff among recognition accuracy, robustness against adversarial attack (fooling rate), and the algorithm’s computational complexity to work in real-time. Thirdly, it should be designed to yield an inherently strong classifier to maximize the cost of attack (e.g., the total number of required gradient computations) for the adversary. Moreover, complying with each of these conditions should not conflict with another. This thesis develops reliable defense and attack algorithms for the advanced end-to-end and representation-level ESC and ASR systems organized into four chapters and five appendices.
Our first contribution is developing an ESC classifier mainly in regard to our third defense reliability conditions. More specifically, we design an ensemble-based classifier in the frontend since it is more robust against adversarial attacks. Furthermore, we exploit a generative adversarial network (GAN) with optimized architectures for both the generator and discriminator networks in the back-end for spectrogram augmentation purposes. We demonstrate that this classification framework outperforms other conventional (e.g., support vector machines) and deep learning-based architectures on benchmarking ESC datasets.
As a second contribution, we develop a robust approach for securing ESC models from a wide range of white and black-box adversarial attacks. This algorithm complies with all the defense reliability conditions mentioned above, and it makes a reasonable trade-off between recognition accuracy and attack fooling rate. Moreover, we study the adversarial transferability ratio between conventional and neural network-based classifiers. According to these findings, we reconfigured our back-end configuration to fill the gap between robustness against attacks and the performance of the front-end classifier. For instance, we employed highboost filtering, dimensionality reduction operation, various logarithmic spectrogram visualizations, and convolutional denoising autoencoder. Our conducted experiments on four challenging datasets corroborate the superior performance of our defense approach compared to other algorithms.
Our third contribution is experimentally characterizing the inverse relation between the recognition accuracy and robustness of the victim classifier against targeted and non-targeted adversarial attacks. Additionally, we identified a few spectrogram settings that maximize the adversary’s cost of attack. This is completely in line with our third reliability condition which obliges us to develop an inherently strong recognition classifier. These settings should be applied before spectrogram production; therefore they do not negatively affect the Jacobian matrix’s distribution either during training or runtime.
As a fourth contribution, we develop an upscale defense approach for end-to-end ASR systems, particularly speech-to-text transcription models. This algorithm is based on synthesizing a new signal using the adjusted chordal distance, and it entirely meets our predefined defense reliability conditions. We employ a multi-discriminator GAN with novel residual-convolutional architectures for the generator and discriminator networks. Then, we train this generative model in the Sobolev space since it is closely related to coefficients of Fourier series, such as Mel-frequency cepstral coefficients (MFCC). Furthermore, we propose a new constraining technique for the generator network to improve its stability and generalizability during training and real-time execution, respectively. Finally, we run our experiments against white and black-box adversarial attacks benchmarked on the advanced DeepSpeech, Kaldi, and Lingvo transcription systems. These experiments indicate that our proposed defense algorithm outperforms other approaches both in terms of word error rate and sentence-level accuracy.
The rest of our contributions published in the flagship signal processing conference and journal letters are organized into appendices. They include four defense and one adversarial attack algorithm developed for both ESC and ASR systems. Our main motivation for developing an adversarial attack algorithm is introducing a fast and robust attack for exploiting in the reliable defense frameworks such as adversarially training.
Esmaeilpour, M. (Author),
Cardinal (Supervisor) &
Lameiras Koerich (Co-supervisor),
20 Dec 2021Student thesis: Doctoral thesis › Doctorate in Engineering: Engineering