hand-read2016-page
Part of HAND: Unified Text–Layout Decoding for Handwritten Document Recognition (Hamdan, Rahiche, Cheriet). Code: github.com/DocumentRecognitionModels/HAND-Decoding
The READ 2016 single-page model reported in the paper (1,258,600 training samples).
- Architecture: fully convolutional encoder and eight-layer transformer decoder; one output stream of characters and layout tokens (page, page number, section, annotation, body).
- Parameters: 7,033,700 plus 365,968 in the speculative draft heads (m = 5).
- Input: an RGB document image, resampled to 150 dpi.
- Result reported in the paper: READ 2016 single-page test (50 pages): CER 3.55 %, WER 13.31 %, LOER 0.0529, mAP-CER 0.9264.
- Scope: Reads one page. On a double- or triple-page image it reads the first page and stops.
Verified before release: decoding the stored example image with these weights reproduces the prediction recorded in the repository token for token; with the draft heads, speculative decoding reproduces greedy decoding.
# git clone https://github.com/DocumentRecognitionModels/HAND-Decoding && cd HAND-Decoding && pip install -r requirements.txt
import sys; sys.path.insert(0, "release")
from hand_release.inference import HANDRecognizer
model = HANDRecognizer.from_pretrained("MHamdan/hand-read2016-page", kv_cache=True, speculative=True)
page = model.read("page.jpg", source_dpi=300) # 300 for a raw scan, 150 if already resampled
print(page.text) # transcription; page.raw keeps the layout tokens inline
Licence and attribution
These weights are released under the Creative Commons Attribution 4.0 International licence (CC BY 4.0, https://creativecommons.org/licenses/by/4.0/). They are Adapted Material of:
- Pretrained Document Attention Network for Handwritten Text Recognition, Denis Coquenet,
Zenodo, DOI 10.5281/zenodo.7244382, CC BY 4.0, file
fcn_read_2016_line_syn.pt(sha256 557d6c349131b113c1e316effdb028eff9b37bf39a48d0689125a3783b86a7c3). Its encoder weights initialised this model, which was then trained further on READ 2016 page images and synthetic pages. The model was trained on READ 2016 (Sánchez et al., Zenodo 10.5281/zenodo.1297399, CC BY 4.0).
The weights are provided as-is, without warranty. The code that loads them is MIT-licensed, with
DAN-derived files under CeCILL-C; see release/NOTICE.md in
https://github.com/DocumentRecognitionModels/HAND-Decoding.
Citation
The paper is under review. Until it is published, please cite the preprint, which appeared under an earlier title:
@article{hamdan2024hand,
title = {{HAND}: Hierarchical Attention Network for Multi-Scale Handwritten
Document Recognition and Layout Analysis},
author = {Hamdan, Mohammed and Rahiche, Abderrahmane and Cheriet, Mohamed},
journal = {arXiv preprint arXiv:2412.18981},
year = {2024},
url = {https://arxiv.org/abs/2412.18981}
}
- Downloads last month
- 9