Skip to main navigation Skip to search Skip to main content

Structural pruning of convolutional neural networks for efficient and sustainable model optimization

  • Sadegh Tofigh

Student thesis: Doctoral thesisDoctorate in Engineering: Engineering

Abstract

The proliferation of Deep Convolutional Neural Networks (CNNs) has fundamentally transformed the landscape of computer vision and autonomous systems. However, the state-of-the-art performance of these models is largely predicated on extreme over-parameterization, leading to a "Red AI" paradigm characterized by massive computational overhead and significant environmental impact. While structural pruning has emerged as a primary solution for deploying these models on resource-constrained edge devices, current methodologies suffer from three critical bottlenecks: a heavy reliance on original training data for saliency analysis, an expensive "search tax" incurred by iterative greedy algorithms, and a lack of theoretical grounding in the functional interdependencies between neural layers. This dissertation addresses these challenges by establishing a mathematically grounded, data-free, and sustainable framework for structural model compression. The core theoretical contribution of this research is the move away from empirical, magnitude-based heuristics toward a principled investigation of signal propagation. We derive analytical upper bounds for the Average Absolute Error (AAE) propagated from a pruned layer to its successor. By modeling the network as a sequence of interdependent manifolds rather than isolated layers, this framework characterizes the "cascading impact" of structural modifications. We further prove the γ-weak property of our importance function, which provides a rigorous mathematical guarantee for the stability and convergence of the selection process. This theoretical foundation allows the framework to identify redundancy using only the intrinsic algebraic properties of weight tensors, facilitating a completely data-free pruning process that preserves the network’s representational fidelity without requiring access to sensitive training distributions. To resolve the computational inefficiencies of traditional pruning, this thesis introduces an oblivious (single-pass) selection paradigm. Conventional greedy algorithms, while locally optimal, impose complexity tax due to their iterative re-inference requirements. Our research identifies a "complexity-performance gap," demonstrating that a well-informed oblivious algorithm can achieve predictive accuracy parity with greedy variants at a fraction of the temporal cost. This efficiency ensures that the optimization process itself is sustainable, preventing the carbon footprint of the pruning phase from outweighing the energy savings gained during model inference. A significant structural innovation presented in this work is the transition from binary "pruning-as-elimination" to a more granular strategy of non-zero filter replacement. In many high-sensitivity architectural regions, the total removal of a filter (replacement with a "zero filter") leads to a catastrophic collapse of the signal manifold. We propose an optimization-based framework that substitutes redundant filters with mathematically optimized, lower dimensional alternatives. This replacement strategy preserves the "knowledge flow" of the network, providing a more stable architectural initialization. This approach is particularly effective in data-constrained environments or scenarios where post-pruning fine-tuning is restricted, offering a superior trade-off between structural sparsity and accuracy. Beyond algorithmic performance, this thesis operationalizes the concept of environmental sustainability through the introduction of the Resource Efficiency (RE) metric. The RE metric provides a standardized protocol to quantify the total computational lifecycle of model optimization, accounting for saliency analysis, selection, and recovery. We further propose a hardware agnostic carbon footprint model to estimate the CO2 equivalents (CO2e) of the pruning lifecycle. This shifts the evaluation of compression algorithms from static inference-speed benchmarks to a comprehensive lifecycle analysis, ensuring that the development of efficient AI is itself an efficient process. The proposed methodologies were validated across a spectrum of standard benchmarks, including VGG, ResNet, and MobileNet architectures, evaluated on the CIFAR-10 and ImageNet datasets. The results demonstrate that the proposed data-free framework consistently achieves state-of-the-art (SOTA) performance, maintaining high accuracy even at aggressive pruning scales. The findings verify that oblivious selection and non zero replacement provide a robust, fast, and simple pipeline for model optimization. The impact of this research is evidenced by its dissemination in premier academic venues. The theoretical derivations and data-free paradigms have been peer-reviewed and published in IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), IEEE Transactions on Neural Networks and Learning Systems (TNNLS), and IEEE Signal Processing Letters. These publications establish the framework as a fundamental contribution to the field of sustainable artificial intelligence. This dissertation provides a comprehensive solution to the inefficiencies of modern neural network compression. By bridging the gap between mathematical optimization theory and practical engineering, the research establishes a new standard for data-blind, interdependency-aware, and environmentally responsible model optimization. The framework serves as a vital tool for the next generation of high-performance intelligence, enabling the deployment of complex neural architectures on the edge while adhering to the global imperatives of "Green AI."
Date22 Jul 2026
Original languageAmerican English
Awarding Institution
  • École de technologie supérieure
SupervisorKim Khoa Nguyen (Supervisor)

Cite this

'