Face recognition for static face images has been well explored and is generally very successful, but video face images taken in unconstrained environments pose more difficult challenges as the image samples suffer from more issues such as pose variation, blur, illumination variation, lower resolution and lower quality. This thesis, addresses face recognition for video-based applications. First, we explore face re-identification for video surveillance applications, attempting pairwise face matching to identify a person in a database using a deep Siamese network. Next, we explore video description, attempting to capture the distribution of the identity samples from a movie using the same Siamese network with clustering techniques. To address this problem, other researchers have labeled a large amount of data in order to enhance their model, which is problematic as it is time- and resource-intensive. The objective is to a adapt deep Siamese network trained on public datasets of static face images to the unconstrained video domain in an unsupervised manner, removing the need to label data manually. To this end, we use triplet loss to learn and adapt discriminative face features in a practical manner for real-world video applications.
Recentwork in video surveillance has used supervised adaptation to close the domain gap between static images and videos. Other researchers have used weakly-supervised or unsupervised domain adaptation, but there are very few works based on a deep Siamese network. These require the target domain to be either a closed-set problem or have a very large amount of unlabeled data, both of which impractical. In this thesis, we introduce an unsupervised domain adaptation named Dual-Triplet learning which is a variant of triplet learning. It simultaneously uses triplets from source and target domains to adapt robust static representation to newly installed video sources using only a few unlabeled samples. The methodology is validated with the COX-S2V dataset whith which we are able to get 3% to 7% gain in classification accuracy.
In regards to video description, we intend to use a deep Siamese network with tracklet information to group face samples of the same identity. With such a tool, it will be possible to describe faces on any video automatically. To this end, robust static face CNN backbones are adapted to a movie using unlabeled data from the movie itself. By using spatio-temporal information (tracklets) of video samples, it is possible to produce positive and negative pairs for triplet loss training. With this, we attempt self-supervised learning with a deep Siamese network, using the first episode of the television series The Big Bang Theory, to learn robust and discriminative features of face samples from the movie. We show that self-supervised learning can enhance the clustering V measure by 15%. We also show that with a sufficient number of samples, the tracklets can be used as a single representation to perform faster and more accurate clustering.
| Date | 5 Oct 2021 |
|---|
| Original language | American English |
|---|
| Awarding Institution | - École de technologie supérieure
|
|---|
| Supervisor | Éric Granger (Supervisor) & Mohamed Dahmane (Co-supervisor) |
|---|
Lemoine St-André, H. (Author),
Granger (Supervisor) & Dahmane (Co-supervisor),
5 Oct 2021Student thesis: Master's thesis › Master in Engineering: Systems Engineering