Representations aim to capture significant, high-level information from raw data, most commonly as low dimensional vectors. When considered as input features for a downstream classification task, they reduce classifier complexity, and help in transfer learning and domain adaptation. An interpretable representation captures underlying meaningful factors, and can be used for understanding data, or to solve tasks that need access to these factors. In natural language processing (NLP), representations such as word or sentence embeddings have recently become important components of most natural language understanding models. They are trained without supervision on very large, unannotated corpora, allowing powerful models that capture semantic relations important in many NLP tasks. In speech processing, deep network-based representations such as bottlenecks and x-vectors have had some success, but are limited to supervised or partly supervised settings where annotations are available and are not optimized to separate underlying factors.
An unsupervised representation for speech, i.e. one that could be trained directly with large amounts of unlabelled speech recordings, would have a major impact on many speech processing tasks. Annotating speech data requires expensive manual transcription and is often a limiting factor, especially for low-resource languages. Disentangling speaker and phonetic variability in the representation would eliminate major nuisance factors for downstream tasks in speech or speaker recognition. But despite this potential, unsupervised representation has received less attention than its supervised counterpart.
In this thesis, we propose a non-supervised generative model that can learn interpretable speech representations. More specifically, we propose several extensions to the variational autoencoder (VAE) model, a unified probabilistic framework which combines generative modelling and deep neural networks. To induce the model to capture and disentangle meaningful underlying factors, we impose biases inspired by articulatory and acoustic theories of speech production.
We first propose time filtering as a bias to induce representations at a different time scale for each latent variable. It allows the model to separate several latent variables along a continuous range of time scale properties, as opposed to binary oppositions or hierarchical factorization that have been previously proposed.
We also show how to impose a multimodal prior to induce discrete latent variables, and present two new tractable VAE loss functions that apply to discrete variables, using expectation-maximization reestimation with matched divergence, and divergence sampling.
In addition, we propose self-attention to add sequence modelling capacity to the VAE model, to our knowledge the first time self-attention is used for learning in an unsupervised speech task.
We use simulated data to confirm that the proposed model can accurately recover phonetic and speaker underlying factors. We find that, given only a realistic high dimensional log filterbank signal, the model is able to accurately recover the generating factors, and that both frame and sequence level variables are essential for accurate reconstruction and well-disentangled representation.
On TIMIT, a corpus of read English speech, the proposed biases yield representations that separate phonetic and speaker information, as evidenced by unsupervised results on downstream phoneme and speaker classification tasks using a simple k-means classifier. Jointly optimizing for multiple latent variables, with a distinct bias for each one, makes it possible to disentangle underlying factors that a single latent variable is not able to capture simultaneously.
We explored some of the underlying factors potentially useful for applications where annotated data is scarce or non-existent. The approach proposed in this thesis, which induces a generative model to learn disentangled and interpretable representations, opens the way for exploration of new factors and inductive biases.
| Date | 27 Aug 2020 |
|---|
| Original language | American English |
|---|
| Awarding Institution | - École de technologie supérieure
|
|---|
| Supervisor | Pierre Dumouchel (Supervisor) |
|---|
Boulianne, G. (Author),
Dumouchel (Supervisor),
27 Aug 2020Student thesis: Doctoral thesis › Doctorate in Engineering: Engineering