Skip to main navigation Skip to search Skip to main content

Transductive few-shot learning

  • Malik Boudiaf

Student thesis: Doctoral thesisDoctorate in Engineering: Engineering

Abstract

Deep learning models have achieved unprecedented success, approaching human-level performances when trained on large-scale labeled data. However, the generalization of such models might be seriously challenged when dealing with new (unseen) classes, with only a few labeled instances per class. Humans, however, can learn new tasks rapidly from a handful of instances, by leveraging context and prior knowledge. To bridge this gap, the few-shot learning community has relied on meta-training strategies, in an attempt to provide the model with intrinsic generalization abilities. In this thesis, we see the few-shot problem in a different light. Noticing the opportunities emerging from foundation models, those large pre-trained models training once on billion-scaled datasets, we shift from the usual training-centered paradigm to an inference-centered one. Throughout this thesis, we aim to develop modular inference procedures that can efficiently adapt any model, regardless of its architecture or how it was trained, to few-shot tasks. To achieve that challenging task, we explore the benefits and limitations of transduction as an inference principle, demonstrating promising results on few-shot classification and few-shot segmentation tasks. As a first contribution, we tackle the most popular problem of few-shot image classification. We develop a highly modular, transductive inference procedure based on the maximization of the mutual information between extracted features and label predictions. We observe very promising results, in both standard few-shot settings, and with domain shift between labeled and unlabeled samples. As a second contribution, we explore the impact on transductive methods of introducing class imbalance in the unlabeled test data of each task. Our findings demonstrate strong adverse effects for all transductive methods, leading some to underperform inductive baselines. To cope with that setting, we diagnose and extend the mutual information-based inference procedure previously described with a-divergences, whose gradients allow more deviation from the uniform prior encoded in the mutual information. Empirically, we observe substantial gains in the class-imbalanced scenario. As a third contribution, we continue to explore potential adverse properties of the unlabeled data on transductive methods. In particular, we investigate the few-shot open-set problem, in which distracting classes can be introduced in the unlabeled data. Motivated by the observation that existing transductive methods perform poorly in open-set scenarios, we propose a generalization of the maximum likelihood principle, in which latent scores down-weighing the influence of potential outliers are introduced alongside the usual parametric model. We show that this method surpasses existing inductive and transductive methods on both aspects of open-set recognition, namely closed-set classification and outlier detection. As a final contribution, we examine the challenging setting of few-shot segmentation, which exhibits both adverse effects mentioned above: class imbalance and openness. We present the first method to completely forego meta-learning and custom architectures. Instead, it uses a standard backbone, trained with standard cross-entropy, and focuses on formulating a per-image transductive inference for each new task. Beyond simplicity, we find this new approach exhibits strong advantages, including a much-improved capacity to leverage an increasing amount of supervision, surpassing by 6 % mIoU previous state-of-the-art in the 10-shot scenario, on the most popular few-shot benchmark.
Date26 Apr 2023
Original languageAmerican English
Awarding Institution
  • École de technologie supérieure
SupervisorIsmail Ben Ayed (Supervisor) & Pablo Piantanida (Co-supervisor)

Cite this

'