Face recognition (FR) has attracted a considerable amount of interest from both academia and industry due to the wide range of applications as found in surveillance and security. Despite the recent progress in computer vision and machine learning, designing a robust system for video-based FR in real-world surveillance applications has been a long-standing challenge. One key issue is the visual domain shift between faces from source domain, where high-quality reference faces are captured under controlled conditions from still cameras, and those from the target domain, where video frames are captured with video cameras under uncontrolled conditions with variations in pose, illumination, expression, etc. The appearance of the faces captured in the videos corresponds to multiple non-stationary data distributions can differ considerably from faces captured during enrollment. Another challenge in video-based FR is the limited number of reference stills that are available per target individual to design facial models. This is a common scenario in security and surveillance applications, as found in, e.g., biometric authentication and watch-list screening. The performance of video-based FR systems can decline significantly due to the limited information available to represent the intra-class variations seen in video frames. This thesis proposes 3 data augmentation techniques based on face synthesis to overcome the challenges of such visual domain shift and limited training set. The main advantage of the proposed approaches is the ability to provide a compact set that can accurately represent the original reference face with relevant intra-class variations corresponding to the capture conditions in the target domain. In particular, this thesis presents new systems for domain-invariant still-to-video FR that are based on augmenting the reference gallery set synthetically which are described with more details in the following.
As a first contribution, a face synthesis approach is proposed that exploits the representative intra-class variational information available from the generic set in target domain. The proposed approach, called domain-specific face synthesis, generates a set of synthetic faces that resemble individuals of interest under the capture conditions relevant to the target domain. In a particular implementation based on sparse representation, the generated synthetic faces are employed to form a cross-domain dictionary that accounts for structured sparsity where the dictionary blocks combine the original and synthetic faces of each individual. Experimental results obtained with videos from the Chokepoint and COX-S2V datasets reveal that augmenting the reference gallery set of still-to-video FR systems using the proposed face synthesizing approach can provide a significantly higher level of accuracy compared to state-of-the-art approaches.
As a second contribution, a paired sparse representation model is proposed allowing for joint use of generic variational information and synthetic face images. The proposed model, called synthetic plus variational model, reconstructs a probe image by jointly using (1) a variational dictionary designed with generic set and (2) a gallery dictionary augmented with a set of synthetic images generated over a wide diversity of pose angles. The augmented gallery dictionary is then encouraged to share the same sparsity pattern with the variational dictionary for similar pose angles by solving a simultaneous sparsity-based optimization problem. Experimental results obtained on Chokepoint and COX-S2V datasets, indicate that the proposed approach can outperform state-of-the-art methods for still-to-video FR with a single sample per person.
As a third contribution, a deep Siamese network, referred as SiamSRC, is proposed where performs face matching using sparse coding. The proposed approach extends the gallery using a set of synthetic face images and exploits sparse representation with a block structure for pairwise face matching that finds the representation of a probe image that requires the minimum number of blocks from the gallery. Experimental results obtained using the Chokepoint and COX-S2V datasets suggest that the proposed SiamSRC network allows for efficient representation of intra-class variations with only a moderate increase in time complexity. Results show that the performance of still-to-video FR systems based on SiamSRC can improve through face synthesis, with no need to collect a large amount of training data.
Results indicate that our proposed techniques which are the integration of face synthesis and generic learning can effectively resolve the challenges of the visual domain shift and limited number of reference stills and provide a higher level of accuracy compared to state-of-the-art approaches under unconstrained surveillance conditions.
| Date | 24 Jan 2020 |
|---|
| Original language | American English |
|---|
| Awarding Institution | - École de technologie supérieure
|
|---|
| Supervisor | Éric Granger (Supervisor) |
|---|
Mokhayyeri, F. (Author),
Granger, É. (Supervisor),
24 Jan 2020Student thesis: Doctoral thesis › Doctorate in Engineering: Engineering