Skip to main navigation Skip to search Skip to main content

Automatic audio anonymization

  • Guillaume Baril

Student thesis: Master's thesisMaster in Engineering: Information Technology Engineering

Abstract

Data anonymization is often a task carried out by humans. Automating it would reduce the cost and time required to complete this task. This work shows that the anonymization of audio data in French can be automated. We propose a pipeline, which takes audio files with their transcriptions and removes the named entities present in the audio. Our pipeline is made up of two components. The first component is the aligner which will align the words in the transcript with the audio and the second component is the template which performs named entity recognition. Then, we replace the audio corresponding to the named entities with a silence. We compared several aligners and several models to find the best ones for our scenario. We evaluated our pipeline on a small hand-annotated dataset, and it achieved a F1 score of 76.9%. This result proves that automating this task is feasible. However, by having a more extensive dataset it would be possible to get better results, train the model to recognize new named entities, and train an end-to-end model that would likely perform better than a pipeline with components trained separately.
Date17 Dec 2021
Original languageAmerican English
Awarding Institution
  • École de technologie supérieure
SupervisorPatrick Cardinal (Supervisor) & Alessandro Lameiras Koerich (Co-supervisor)

Cite this

'