Skip to main navigation Skip to search Skip to main content

Adaptive Multi-Scale Feature Fusion for Paragraph-Level Handwritten Text Recognition via Spatial Attention

  • École de technologie supérieure

Research output: Contribution to Book/Report typesContribution to conference proceedingspeer-review

Abstract

Paragraph-level handwritten text recognition presents unique challenges distinct from line-or word-level recognition: the need to simultaneously capture fine-grained stroke details and long-range semantic dependencies across multiple lines without explicit segmentation boundaries. Existing methods often rely on computationally expensive recurrent architectures, external lexicons, or line-level annotations, which limit their practical applicability. This work proposes a novel multi-scale architecture that addresses these challenges through three tightly integrated components: (1) an encoder integrating lightweight CNNs for local feature extraction and Vision Transformers for global context modeling, unified through a learnable spatial attention mechanism that adaptively weights their contributions at each spatial location; (2) an autoregressive Transformer decoder that generates character sequences directly from the fused representation, eliminating recurrence while maintaining strong sequence modeling capabilities; and (3) a multilevel feature consistency loss that enforces alignment between CNN and ViT representations, improving optimization stability and cross-scale coherence. The fully end-toend framework achieves state-of-the-art segmentation-free, lexicon-independent performance on IAM (WER 13.23%, CER 4.50%) and RIMES (WER 4.84%, CER 2.18%), while converging about 20% faster than recurrent baselines. These results indicate that adaptive multi-scale fusion with recurrence-free decoding is an effective and scalable solution for real-world paragraph-level HTR.

Original languageEnglish
Title of host publicationProceedings - 2026 IEEE/CVF Winter Conference on Applications of Computer Vision Workshops, WACVW 2026
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages710-718
Number of pages9
ISBN (Electronic)9798331591496
DOIs
Publication statusPublished - 2026
Event2026 IEEE/CVF Winter Conference on Applications of Computer Vision Workshops, WACVW 2026 - Tucson, United States
Duration: 6 Mar 202610 Mar 2026

Publication series

NameProceedings - 2026 IEEE/CVF Winter Conference on Applications of Computer Vision Workshops, WACVW 2026

Conference

Conference2026 IEEE/CVF Winter Conference on Applications of Computer Vision Workshops, WACVW 2026
Country/TerritoryUnited States
CityTucson
Period6/03/2610/03/26

UN SDGs

This output contributes to the following UN Sustainable Development Goals (SDGs)

  1. SDG 3 - Good Health and Well-being
    SDG 3 Good Health and Well-being

!!!Keywords

  • Paragraph-level HTR multi-scale fusion Vision Transformer spatial attention segmentation-free recognition

Fingerprint

Dive into the research topics of 'Adaptive Multi-Scale Feature Fusion for Paragraph-Level Handwritten Text Recognition via Spatial Attention'. These topics are generated from the title and abstract of the publication. Together, they form a unique fingerprint.

Cite this