Résumé
Paragraph-level handwritten text recognition presents unique challenges distinct from line-or word-level recognition: the need to simultaneously capture fine-grained stroke details and long-range semantic dependencies across multiple lines without explicit segmentation boundaries. Existing methods often rely on computationally expensive recurrent architectures, external lexicons, or line-level annotations, which limit their practical applicability. This work proposes a novel multi-scale architecture that addresses these challenges through three tightly integrated components: (1) an encoder integrating lightweight CNNs for local feature extraction and Vision Transformers for global context modeling, unified through a learnable spatial attention mechanism that adaptively weights their contributions at each spatial location; (2) an autoregressive Transformer decoder that generates character sequences directly from the fused representation, eliminating recurrence while maintaining strong sequence modeling capabilities; and (3) a multilevel feature consistency loss that enforces alignment between CNN and ViT representations, improving optimization stability and cross-scale coherence. The fully end-toend framework achieves state-of-the-art segmentation-free, lexicon-independent performance on IAM (WER 13.23%, CER 4.50%) and RIMES (WER 4.84%, CER 2.18%), while converging about 20% faster than recurrent baselines. These results indicate that adaptive multi-scale fusion with recurrence-free decoding is an effective and scalable solution for real-world paragraph-level HTR.
| langue originale | Anglais |
|---|---|
| titre | Proceedings - 2026 IEEE/CVF Winter Conference on Applications of Computer Vision Workshops, WACVW 2026 |
| Editeur | Institute of Electrical and Electronics Engineers Inc. |
| Pages | 710-718 |
| Nombre de pages | 9 |
| ISBN (Electronique) | 9798331591496 |
| Les DOIs | |
| état | Publié - 2026 |
| Evénement | 2026 IEEE/CVF Winter Conference on Applications of Computer Vision Workshops, WACVW 2026 - Tucson, Etats-Unis Durée: 6 mars 2026 → 10 mars 2026 |
Série de publications
| Nom | Proceedings - 2026 IEEE/CVF Winter Conference on Applications of Computer Vision Workshops, WACVW 2026 |
|---|
Conférence
| Conférence | 2026 IEEE/CVF Winter Conference on Applications of Computer Vision Workshops, WACVW 2026 |
|---|---|
| Pays/Territoire | Etats-Unis |
| La ville | Tucson |
| période | 6/03/26 → 10/03/26 |
SDG des Nations Unies
Ce résultat contribue à ou aux Objectifs de développement durable suivants
-
SDG 3 – Bonne santé et bien-être
Empreinte digitale
Voici les principaux termes ou expressions associés à « Adaptive Multi-Scale Feature Fusion for Paragraph-Level Handwritten Text Recognition via Spatial Attention ». Ces libellés thématiques sont générés à partir du titre et du résumé de la publication. Ensemble, ils forment une empreinte digitale unique.Citer cette ressource
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver