Skip to main navigation Skip to search Skip to main content

Multimedia copy detection using audio and video fingerprints

  • Chahid Ouali

Student thesis: Doctoral thesisDoctorate in Engineering: Engineering

Abstract

According to a study by the International Data Corporation (IDC), the digital universe is doubling in size every two years to rich 44 trillion gigabytes by 2020. A large part of this big universe consists of audio and videos (e.g. music, TV shows and films), which are distributed over the Internet in an effortless way. This has increased the need for powerful tools to handle this data in terms of identification, filtering and retrieval. In this context, multimedia copy detection, which consists of identifying duplicate (or near duplicate) multimedia content, has become an emerging and active research area due to its broad applications. Multimedia copy detection can be used in a wide variety of applications such as broadcast monitoring, music identification, copyright control, law enforcement investigation and music library organization. Content-Based Copy Detection (CBCD) has been recently introduced as a solution to the problem of multimedia copy detection. This approach extracts fingerprints from a candidate copy and then compares them against fingerprints of the original content. However, audio and video signals are subjected to various kinds of transformations that make robust fingerprint extraction challenging. Thus, fingerprints should be robust to a variety of audio and video transformations and also discriminate against imposter fingerprints. In addition, the search of a candidate copy against a large dataset of fingerprints should be very fast. In this thesis, we propose an efficient multimedia copy detection system that is highly robust to a variety of audio and video transformations. We first describe a new audio feature extraction schema that allows the generation of three kinds of audio fingerprints. We then address the problem of video copy detection and we describe two video fingerprint extraction algorithms. In addition, we propose a fusion technique that combines the results achieved separately from the audio and the video parts to tackle the problem of audio+video copy detection. In the last part of this thesis, we address the problem of fingerprint retrieval, and we propose two solutions to improve the speed of the search algorithm. In the first solution we propose to parallelize the similarity search algorithm by using a Graphics Processing Unit (GPU), whereas the second solution is based on a clustering technique. We evaluate the proposed systems on the TRECVID 2009 and 2010 datasets, and we evaluate our approaches in terms of detection performance, localization accuracy and run time. In addition, we demonstrate the effectiveness of our methods by comparing them to several state-of-the-art audio and video copy detection systems.
Date25 Oct 2016
Original languageAmerican English
Awarding Institution
  • École de technologie supérieure
SupervisorPierre Dumouchel (Supervisor) & Vishwa Gupta (Co-supervisor)

Cite this

'