Abstract
Paragraph-level handwritten text recognition presents unique challenges distinct from line-or word-level recognition: the need to simultaneously capture fine-grained stroke details and long-range semantic dependencies across multiple lines without explicit segmentation boundaries. Existing methods often rely on computationally expensive recurrent architectures, external lexicons, or line-level annotations, which limit their practical applicability. This work proposes a novel multi-scale architecture that addresses these challenges through three tightly integrated components: (1) an encoder integrating lightweight CNNs for local feature extraction and Vision Transformers for global context modeling, unified through a learnable spatial attention mechanism that adaptively weights their contributions at each spatial location; (2) an autoregressive Transformer decoder that generates character sequences directly from the fused representation, eliminating recurrence while maintaining strong sequence modeling capabilities; and (3) a multilevel feature consistency loss that enforces alignment between CNN and ViT representations, improving optimization stability and cross-scale coherence. The fully end-toend framework achieves state-of-the-art segmentation-free, lexicon-independent performance on IAM (WER 13.23%, CER 4.50%) and RIMES (WER 4.84%, CER 2.18%), while converging about 20% faster than recurrent baselines. These results indicate that adaptive multi-scale fusion with recurrence-free decoding is an effective and scalable solution for real-world paragraph-level HTR.
| Original language | English |
|---|---|
| Title of host publication | Proceedings - 2026 IEEE/CVF Winter Conference on Applications of Computer Vision Workshops, WACVW 2026 |
| Publisher | Institute of Electrical and Electronics Engineers Inc. |
| Pages | 710-718 |
| Number of pages | 9 |
| ISBN (Electronic) | 9798331591496 |
| DOIs | |
| Publication status | Published - 2026 |
| Event | 2026 IEEE/CVF Winter Conference on Applications of Computer Vision Workshops, WACVW 2026 - Tucson, United States Duration: 6 Mar 2026 → 10 Mar 2026 |
Publication series
| Name | Proceedings - 2026 IEEE/CVF Winter Conference on Applications of Computer Vision Workshops, WACVW 2026 |
|---|
Conference
| Conference | 2026 IEEE/CVF Winter Conference on Applications of Computer Vision Workshops, WACVW 2026 |
|---|---|
| Country/Territory | United States |
| City | Tucson |
| Period | 6/03/26 → 10/03/26 |
UN SDGs
This output contributes to the following UN Sustainable Development Goals (SDGs)
-
SDG 3 Good Health and Well-being
!!!Keywords
- Paragraph-level HTR multi-scale fusion Vision Transformer spatial attention segmentation-free recognition
Fingerprint
Dive into the research topics of 'Adaptive Multi-Scale Feature Fusion for Paragraph-Level Handwritten Text Recognition via Spatial Attention'. These topics are generated from the title and abstract of the publication. Together, they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver