Digitized archives are increasingly using multispectral (MS) imaging to reveal weak content (faded inks, palimpsests, annotations, etc.) and separate text from background. However, spectral decomposition remains difficult : conventional approaches (e.g., fixed-rank NMF, PCA or GMM) require ad hoc settings, cumbersome pre/post-processing and generalize poorly to the diversity of substrates, inks and spectral acquisition conditions.
To address these challenges, we first introduce an end-to-end learning framework for multispectral decomposition that combines a convolutional auto-encoder, coupled with a constrained unmixing head (non-negativity, interpretability, orthogonality), enriched with layout priors (attention block), to preserve glyph structure while modeling the spectro-spatial context. This hybrid approach integrates NMF principles into an auto-encoder architecture, exploiting the complementary advantages of both approaches. Secondly, in response to the open problem of manual rank selection, we propose a mechanism for its automatic selection via pruning guided by minimum description length (MDL), learned jointly. Uninformative components are then progressively removed to simultaneously minimize reconstruction error and model complexity. Finally, in a third step, we show that this framework, named PRISM, holds for different MS image configurations, for both overdetermined (i.e., more bands than sources) and underdetermined (i.e., fewer bands, e.g. RGB) cases, and generalizes beyond multispectral documents.
Evaluated on MSBin and MStex, two varied document datasets (e.g., letters, forms, manuscripts) from different periods and states, PRISM consistently improves ink/background separation by +29.5 F-score points against Howe’s binarization and outperforms ACE v2 by +1.32 points (state-of-the-art). Furthermore, for unsupervised MS image decomposition, PRISM remains up to 7.4× faster than VBONMF, the best competing NMF approach. Tests on reference hyperspectral scenes, Jasper Ridge and Urban, as well as on RGB images, confirm good transferability beyond the documentary domain. Ablation studies validate the contribution of the MDL pruning and the various priors. These results show that combining physical constraints and spatial context enables interpretable and adaptive decompositions, useful for transcription and restoration. PRISM code, weights and hyperparameters are available on Github and accompany this thesis, whose contributions have been integrated into an ICCV 2025 VisionDocs workshop publication.
| Date | 18 Nov 2025 |
|---|
| Original language | French |
|---|
| Awarding Institution | - École de technologie supérieure
|
|---|
| Supervisor | Mohamed Cheriet (Supervisor) |
|---|