Skip to main navigation Skip to search Skip to main content

Optimization problems for deep neural networks

  • Jérôme Rony

Student thesis: Doctoral thesisDoctorate in Engineering: Engineering

Abstract

Deep learning methods heavily rely on gradient descent to solve a variety of pattern recognition problems. Given the known limitations of this optimization method, it is of paramount importance to carefully craft objectives to accurately and efficiently solve learning and verification tasks arising in this domain. In particular, researchers have mostly relied on penalty methods to handle constraints and introduced many hyperparameters to ease the optimization process. In this thesis, we propose to revisit some of these problems, analyze them through the lens of optimization, and leverage well-known tools from the optimization literature to more accurately and efficiently solve them. Our first contribution is to take a step back from the deep metric learning literature, and notice that most pairwise methods proposed in recent years have similar objectives. In fact, they all correspond to maximizing the same quantity: the mutual information. Additionally, minimizing the well-known cross-entropy loss can also be viewed as maximizing the mutual information. This suggests that using the cross-entropy to learn the parameters of a deep metric learning model is a viable solution. This is confirmed experimentally, where the simplicity of the cross-entropy yields state-of-the-art results on all commonly used datasets. As a second contribution, we investigate several problems related to adversarial robustness, and adversarial attacks in particular. These problems can be formulated as the minimization of a discrepancy measure under one (or several) misclassification constraint(s), with additional input space constraints. We develop a first simple algorithm to generate minimal ℓ2-norms adversarial perturbations for classifications models. Like several of the later published adversarial attacks, this method is efficient, but lacks generality as it is customized to one particular distance. Therefore, we develop a second adversarial attack for classification models based on the augmented Lagrangian framework. This attack enjoys the generality of penalty based approaches, as it can handle many smooth discrepancy measures, and the computational efficiency of distance-specific algorithms. Our goal is to provide a general framework that can serve as a starting point to future researchers when designing adversarial attacks for new measures. Finally, we investigate attacks in the context of a dense prediction task: semantic segmentation. Adversarial attacks in this context can be formulated as optimization problems with millions of misclassification constraints. Therefore, we leverage our augmented Lagrangian based method to handle such large numbers of constraints, and combine it with a proximal splitting to minimize the non-smooth ℓ∞-norm. This attack is, to the best of our knowledge, the first to accurately solve the minimal adversarial perturbation problem for semantic segmentation. Our third contribution focuses on calibration of deep neural networks in classification tasks. Following recent work that showed the advantage of using constraints on the output of a model to improve calibration, we generalize the approach in an augmented Lagrangian framework. In particular, we tackle the constraints with adaptive class-wise penalties. This results in a scalable method that can be applied to classification, as well as segmentation, and obtains state-of-the-art classification and calibration performances.
Date24 Apr 2023
Original languageAmerican English
Awarding Institution
  • École de technologie supérieure
SupervisorIsmail Ben Ayed (Supervisor) & Éric Granger (Co-supervisor)

Cite this

'