hand-read2016-double-page

Part of HAND: Unified Text–Layout Decoding for Handwritten Document Recognition (Hamdan, Rahiche, Cheriet). Code: github.com/DocumentRecognitionModels/HAND-Decoding

The page model after continued training on READ 2016 double-page images (311 epochs; selected at epoch 220).

  • Architecture: fully convolutional encoder and eight-layer transformer decoder; one output stream of characters and layout tokens (page, page number, section, annotation, body).
  • Parameters: 7,033,700 plus 365,968 in the speculative draft heads (m = 5).
  • Input: an RGB document image, resampled to 150 dpi.
  • Result reported in the paper: READ 2016 double-page test (24 images): CER 3.60 %, WER 13.27 %, LOER 0.0474, mAP-CER 0.925.
  • Scope: Reads two pages. On a single page it opens further page elements after the page; on a triple page it stops after two pages.

Verified before release: decoding the stored example image with these weights reproduces the prediction recorded in the repository token for token; with the draft heads, speculative decoding reproduces greedy decoding.

# git clone https://github.com/DocumentRecognitionModels/HAND-Decoding && cd HAND-Decoding && pip install -r requirements.txt
import sys; sys.path.insert(0, "release")
from hand_release.inference import HANDRecognizer

model = HANDRecognizer.from_pretrained("MHamdan/hand-read2016-double-page", kv_cache=True, speculative=True)
page = model.read("page.jpg", source_dpi=300)   # 300 for a raw scan, 150 if already resampled
print(page.text)      # transcription; page.raw keeps the layout tokens inline

Licence and attribution

These weights are released under the Creative Commons Attribution 4.0 International licence (CC BY 4.0, https://creativecommons.org/licenses/by/4.0/). They are Adapted Material of:

  • Pretrained Document Attention Network for Handwritten Text Recognition, Denis Coquenet, Zenodo, DOI 10.5281/zenodo.7244382, CC BY 4.0, file fcn_read_2016_line_syn.pt (sha256 557d6c349131b113c1e316effdb028eff9b37bf39a48d0689125a3783b86a7c3). Its encoder weights initialised this model, which was then trained further on READ 2016 page images and synthetic pages. The model was trained on READ 2016 (Sánchez et al., Zenodo 10.5281/zenodo.1297399, CC BY 4.0).

The weights are provided as-is, without warranty. The code that loads them is MIT-licensed, with DAN-derived files under CeCILL-C; see release/NOTICE.md in https://github.com/DocumentRecognitionModels/HAND-Decoding.

Citation

The paper is under review. Until it is published, please cite the preprint, which appeared under an earlier title:

@article{hamdan2024hand,
  title   = {{HAND}: Hierarchical Attention Network for Multi-Scale Handwritten
             Document Recognition and Layout Analysis},
  author  = {Hamdan, Mohammed and Rahiche, Abderrahmane and Cheriet, Mohamed},
  journal = {arXiv preprint arXiv:2412.18981},
  year    = {2024},
  url     = {https://arxiv.org/abs/2412.18981}
}
Downloads last month
7
Safetensors
Model size
7.03M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including MHamdan/hand-read2016-double-page

Paper for MHamdan/hand-read2016-double-page