Deep learning has revolutionized 3D perception, enabling accurate recognition, segmentation, and reconstruction of point-cloud data across robotics, autonomous navigation, and augmented reality. However, two major challenges remain. First, large-scale 3D models rely heavily on labeled datasets, which are costly and time-consuming to acquire. This motivates the development of self-supervised learning methods that can pretrain models on unlabeled data and learn transferable geometric representations for downstream tasks with limited annotations. Second, models trained under fixed conditions often fail to generalize when exposed to real-world distribution shifts caused by noise, sensor variation, or environmental changes. This challenge motivates the development of learning frameworks that are both robust to distribution shifts and adaptive to new environments without requiring labeled data.
This thesis advances the robustness and generalizability of 3D deep learning through a unified exploration of self-supervised representation learning and test-time learning (TTL). The first part investigates how to construct geometric priors that enable models to learn meaningful and transferable 3D representations. (1) GeoMask3D introduces a geometry-aware masked modeling strategy that explicitly aligns masked pretraining with structural cues of 3D shapes, improving the interpretability and invariance of learned features. 2) Spectral-Informed Mamba extends state-space models to point clouds by leveraging the Laplacian spectrum of the underlying graph manifold, producing an isometry-invariant traversal order that strengthens the quality of self-supervised geometric representations while maintaining linear computational complexity compared to quadratic Transformer designs.
The second part addresses adaptation under unseen test conditions, proposing efficient and reliable methods for test-time training (TTT) and test-time adaptation (TTA). (3) Sampling Variation Weight Averaging (SVWA) presents the first fully TTA strategy for point clouds, combining sampling variation and weight averaging to achieve robust adaptation through f lat-minima optimization. (4) SMART-PC introduces a skeleton-based TTT framework that learns compact geometric abstractions during pretraining, enabling real-time adaptation without backpropagation by updating only BatchNorm statistics.
Together, these contributions establish a cohesive framework for robust and generalizable deep geometric representation learning in 3D. By unifying self-supervised pretraining with efficient test-time learning, this thesis advances toward 3D perception systems that are stable, adaptive, and resilient to real-world distribution shifts while preserving interpretability and computational efficiency.
| Date | 14 Mar 2026 |
|---|
| Original language | American English |
|---|
| Awarding Institution | - École de technologie supérieure
|
|---|
| Supervisor | Christian Desrosiers (Supervisor) & Ismail Ben Ayed (Co-supervisor) |
|---|
Bahri, A. (Author),
Desrosiers (Supervisor) &
Ben Ayed (Co-supervisor),
14 Mar 2026Student thesis: Doctoral thesis › Doctorate in Engineering: Engineering