In the past decades, face recognition (FR) has received a growing attention in security applications such as intelligent video surveillance (VS). Embedded in decision support tools, FR allows to detect the presence of individuals of interest in video streams in a discrete and nonintrusive way, which is of a particular interest for applications such as watchlist screening, search and retrieval or face re-identification. However, recognizing faces corresponding to target individuals remains a challenging problem in VS. FR systems are usually presented with videos exhibiting a wide range of variations caused by uncontrolled observation conditions, most notably in illumination condition, image resolution, motion blur, facial pose and expression. To perform recognition, facial models of target individuals are typically designed with a limited number of reference stills or videos captured during an enrollment process, and these variations contribute to a growing divergence between these models and the underlying data distribution. Although facial models can be adapted when new reference videos that may become available over time, incremental learning with faces captured under different conditions remains challenging, as it may lead to knowledge corruption. Furthermore, only a subset of this knowledge may be relevant to classify a given facial capture, and relying on information related to different capture conditions may even deteriorate system performance.
In this thesis, a new framework is proposed for the automatic detection of individuals of interest for VS applications. A human-centric scenario is considered, where a FR system is Embedded in a decision support tool that alerts an analyst to the presence of individuals of interest in multiple video feeds. Individuals can be added or removed from the system by the analyst, and their facial models can be refined over time with new reference sequences. In this framework, the use of concept change detection is proposed to guide an ensemble learning strategy. Each enrolled individual is modeled by a dedicated ensemble of two-class classifiers, each one specialized in a different conditions detected in reference sequences. In addition, this Framework allows for a dynamic adaptation of its behavior to changing capture conditions during operations. A dynamic ensemble fusion rule is proposed, relying on concept models to estimate the relevance of each classifier w.r.t. each operational input. Finally, system decisions are accumulated over tracks following faces across consecutive frames, to provide robust spatio-temporal recognition.
In Chapter 2, concept change detection is first investigated to reduce the growth in complexity of a self-updating template-matching system for FR in video. A context-sensitive approach is proposed for self-updating, where galleries of reference images are only updated with highlyconfident captures exhibiting significant changes in capture conditions. Proof of concept experiments have been conducted with a standard template matching system detecting changes in illumination conditions, using thee publicly-available face databases. Simulation results indicate that the proposed approach allows to maintain system performance while mitigating the growth in system complexity. It exhibits the level of performance than a regular self-updating template matching system, with gallery sizes reduced by half.
In Chapter 3, a new framework for an adaptive multi-classifier system is proposed for FR in VS. It is comprised of an ensemble of incremental learning classifiers per enrolled individual, and relies on concept change detection to refine facial models with new reference data available over time while mitigating knowledge corruption. An hybrid strategy is proposed, where individual-specific ensembles are only augmented with new classifiers when an abrupt change is detected in reference data. When a gradual change is detected, knowledge about corresponding concepts is refined through incremental update of corresponding classifiers. For proof of concept experiments, a particular implementation is proposed, using ensembles of probabilistic Fuzzy-ARTMAP classifiers generated and updated with dynamic Particle Swarm Optimization, and the Hellinger Drift Detection Method for change detection. Experimental results with the FIA video surveillance database indicate that the proposed framework allows to maintain system performance over time, effectively mitigating the effects of knowledge corruption. It exhibits higher classification performance than a similar passive system, and reference probabilistic kNN and TCM-kNN systems.
In Chapter 4, an evolution of the framework presented in Chapter 3 is presented, that allows to adapt system behavior to changing operating conditions. A new dynamic weighting fusing rule is proposed for ensembles of classifiers, where each classifier is weighted by its competence to classify each operational input. Furthermore, to provide a lightweight competence estimation that doesn’t interfere with live operations, classifier competence is estimated from the concept models used for change detection during training. An evolution of the particular implementation presented in Chapter 3 is proposed, where concept models are estimated with the Fuzzy C-Means clustering algorithm, and ensemble fusion is performed through dynamic weighted score-average. Experimental simulations with the FIA and ChokePoint videosurveillance datasets shows that the proposed dynamic fusion method provides a higher classification performance than the DS-OLA dynamic selection method, for a significantly lower computational complexity. In addition, the proposed system exhibits higher performance than reference probabilistic kNN, TCM-kNN and Adaptive Sparse Coding systems.
| Date | 8 oct. 2015 |
|---|
| langue originale | Anglais américain |
|---|
| Établissement diplômant | - École de technologie supérieure
|
|---|
| Superviseur | Éric Granger (Directeur(-trice)) & Robert Sabourin (Codirecteur(-trice)) |
|---|