Deep learning has achieved remarkable success in the field of computer vision with the advancement of Convolutional Neural Networks (CNN). However, the design of the optimal CNN models (architectures) for a new application is a time-consuming task, requiring expert knowledge and trial and error. Neural Architecture Search (NAS) has emerged in the past decade as a solution to automate the process of finding optimal architecture design. One-shot methods are one of the main NAS approaches that reduce the computational cost of NAS by training a single supernet that contains all possible architectures in the search space and directly inheriting these weights for architecture performance evaluation. However, a fundamental issue in one-shot NAS is the degradation of the quality of architecture performance estimation based on the supernet due to conflict and co-adaptation of weights during supernet training. This issue can be addressed by reducing the weight sharing by using multiple supernets for various parts of the search space or by focusing only on promising parts of the search space during training by using a sampling method.
This thesis focuses on improving one-shot NAS methods from various aspects. We first provide an introduction to our research and contributions. We then provide general background about CNNs, various architecture designs, Monte-Carlo Tree Search (MCTS), and NAS. For our first contribution, we focus on optimizing the downsampling configuration of CNN as a NAS problem. We propose a balanced mixture of supernets to partition the search space and reduce weight sharing by utilizing distinct supernets for each partition. We propose to learn the partitioning and association of architectures to each partition in a balanced manner to ensure fairness in training multiple supernets. Next, we propose two approaches for learning the hierarchical structure (search tree) of the NAS search space for MCTS simultaneously with supernet training. Our first approach is to learn the hierarchy of the search space in an unsupervised manner. We propose to use the functional similarity of architectures based on their output vector to construct the hierarchy. Unlike previous works, our method does not use a default hierarchical design or inaccurate performance predictions from the supernet. The second approach is an iterative method to refine the hierarchy based on increasingly better performance estimations from the supernet. Both approaches facilitate one-shot NAS by providing a better exploration-exploitation trade-off, improving the final performance with reduced NAS cost. Finally, as our last contribution, we investigate the gradient conflict and cooperation of sampled architectures during supernet training. Since gradient conflict is not uniform in the search space, we focus on the effective training each architecture receives in an optimization step. We propose a gradient density metric that estimates the effective training received by an architecture by measuring how aligned its gradient is with the rest of the search space. We propose a density-aware sampling method to reduce the bias in effective training received during supernet optimization. Finally, we provide a conclusion and future directions for improving one-shot NAS.
| Date | 26 May 2026 |
|---|
| Original language | American English |
|---|
| Awarding Institution | - École de technologie supérieure
|
|---|
| Supervisor | Marco Pedersoli (Supervisor) & Matthew Toews (Co-supervisor) |
|---|
Javan Roshtkhari, M. (Author),
Pedersoli (Supervisor) &
Toews (Co-supervisor),
26 May 2026Student thesis: Doctoral thesis › Doctorate in Engineering: Engineering