MHamdan commited on
Commit
eacef65
·
verified ·
1 Parent(s): a43911f

Release weights (HAND-Decoding v1.1.0)

Browse files
README.md ADDED
@@ -0,0 +1,70 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: cc-by-4.0
3
+ library_name: pytorch
4
+ pipeline_tag: image-to-text
5
+ tags:
6
+ - handwritten-text-recognition
7
+ - document-layout-analysis
8
+ - htr
9
+ - ocr
10
+ - historical-documents
11
+ - read-2016
12
+ ---
13
+
14
+ # hand-read2016-page
15
+
16
+ Part of **HAND: Unified Text–Layout Decoding for Handwritten Document Recognition** (Hamdan, Rahiche, Cheriet).
17
+ Code: [github.com/DocumentRecognitionModels/HAND-Decoding](https://github.com/DocumentRecognitionModels/HAND-Decoding) · Demo: [HAND-Decoding demo](https://huggingface.co/spaces/MHamdan/HAND-Decoding-demo)
18
+
19
+ The READ 2016 single-page model reported in the paper (1,258,600 training samples).
20
+
21
+ - **Architecture:** fully convolutional encoder and eight-layer transformer decoder; one output
22
+ stream of characters and layout tokens (page, page number, section, annotation, body).
23
+ - **Parameters:** 7,033,700 plus 365,968 in the speculative draft heads (m = 5).
24
+ - **Input:** an RGB document image, resampled to 150 dpi.
25
+ - **Result reported in the paper:** READ 2016 single-page test (50 pages): CER 3.55 %, WER 13.31 %, LOER 0.0529, mAP-CER 0.9264.
26
+ - **Scope:** Reads one page. On a double- or triple-page image it reads the first page and stops.
27
+
28
+ Verified before release: decoding the stored example image with these weights reproduces the
29
+ prediction recorded in the repository token for token; with the draft heads, speculative decoding reproduces greedy decoding.
30
+
31
+ ```python
32
+ # git clone https://github.com/DocumentRecognitionModels/HAND-Decoding && cd HAND-Decoding && pip install -r requirements.txt
33
+ import sys; sys.path.insert(0, "release")
34
+ from hand_release.inference import HANDRecognizer
35
+
36
+ model = HANDRecognizer.from_pretrained("MHamdan/hand-read2016-page", kv_cache=True, speculative=True)
37
+ page = model.read("page.jpg", source_dpi=300) # 300 for a raw scan, 150 if already resampled
38
+ print(page.text) # transcription; page.raw keeps the layout tokens inline
39
+ ```
40
+
41
+ ## Licence and attribution
42
+
43
+ These weights are released under the Creative Commons Attribution 4.0 International licence
44
+ (CC BY 4.0, https://creativecommons.org/licenses/by/4.0/). They are Adapted Material of:
45
+
46
+ - **Pretrained Document Attention Network for Handwritten Text Recognition**, Denis Coquenet,
47
+ Zenodo, DOI 10.5281/zenodo.7244382, CC BY 4.0, file `fcn_read_2016_line_syn.pt`
48
+ (sha256 557d6c349131b113c1e316effdb028eff9b37bf39a48d0689125a3783b86a7c3). Its encoder weights
49
+ initialised this model, which was then trained further on READ 2016 page images and synthetic
50
+ pages. The model was trained on READ 2016 (Sánchez et al., Zenodo 10.5281/zenodo.1297399,
51
+ CC BY 4.0).
52
+
53
+ The weights are provided as-is, without warranty. The code that loads them is MIT-licensed, with
54
+ DAN-derived files under CeCILL-C; see `release/NOTICE.md` in
55
+ https://github.com/DocumentRecognitionModels/HAND-Decoding.
56
+
57
+ ## Citation
58
+
59
+ The paper is under review. Until it is published, please cite the preprint, which appeared under an earlier title:
60
+
61
+ ```bibtex
62
+ @article{hamdan2024hand,
63
+ title = {{HAND}: Hierarchical Attention Network for Multi-Scale Handwritten
64
+ Document Recognition and Layout Analysis},
65
+ author = {Hamdan, Mohammed and Rahiche, Abderrahmane and Cheriet, Mohamed},
66
+ journal = {arXiv preprint arXiv:2412.18981},
67
+ year = {2024},
68
+ url = {https://arxiv.org/abs/2412.18981}
69
+ }
70
+ ```
charset.json ADDED
@@ -0,0 +1,111 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "charset": [
3
+ "\n",
4
+ " ",
5
+ "(",
6
+ ")",
7
+ "+",
8
+ ",",
9
+ "-",
10
+ ".",
11
+ "/",
12
+ "0",
13
+ "1",
14
+ "2",
15
+ "3",
16
+ "4",
17
+ "5",
18
+ "6",
19
+ "7",
20
+ "8",
21
+ "9",
22
+ ":",
23
+ "A",
24
+ "B",
25
+ "C",
26
+ "D",
27
+ "E",
28
+ "F",
29
+ "G",
30
+ "H",
31
+ "I",
32
+ "J",
33
+ "K",
34
+ "L",
35
+ "M",
36
+ "N",
37
+ "O",
38
+ "P",
39
+ "Q",
40
+ "R",
41
+ "S",
42
+ "T",
43
+ "U",
44
+ "V",
45
+ "W",
46
+ "Y",
47
+ "Z",
48
+ "[",
49
+ "]",
50
+ "a",
51
+ "b",
52
+ "c",
53
+ "d",
54
+ "e",
55
+ "f",
56
+ "g",
57
+ "h",
58
+ "i",
59
+ "j",
60
+ "k",
61
+ "l",
62
+ "m",
63
+ "n",
64
+ "o",
65
+ "p",
66
+ "q",
67
+ "r",
68
+ "s",
69
+ "t",
70
+ "u",
71
+ "v",
72
+ "w",
73
+ "x",
74
+ "y",
75
+ "z",
76
+ "¬",
77
+ "¾",
78
+ "Ö",
79
+ "ß",
80
+ "ä",
81
+ "ö",
82
+ "ü",
83
+ "ÿ",
84
+ "ā",
85
+ "ē",
86
+ "ō",
87
+ "ū",
88
+ "ȳ",
89
+ "̄",
90
+ "̈",
91
+ "—",
92
+ "Ⓐ",
93
+ "Ⓑ",
94
+ "Ⓝ",
95
+ "Ⓟ",
96
+ "Ⓢ",
97
+ "ⓐ",
98
+ "ⓑ",
99
+ "ⓝ",
100
+ "ⓟ",
101
+ "ⓢ"
102
+ ],
103
+ "vocab_size": 99,
104
+ "additional_tokens": 1,
105
+ "special_tokens": {
106
+ "end": 99,
107
+ "start": 100,
108
+ "pad": 101
109
+ },
110
+ "note": "index convention from hand/OCR/ocr_dataset_manager.py:74-99 with charset_mode='seq2seq'; the output layer covers 0..vocab_size, i.e. the charset plus <end>"
111
+ }
config.json ADDED
@@ -0,0 +1,49 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "model_type": "hand-dan-page",
3
+ "input_channels": 3,
4
+ "dropout": 0.5,
5
+ "enc_dim": 256,
6
+ "nb_layers": 5,
7
+ "pe_h_max": 500,
8
+ "pe_w_max": 1000,
9
+ "l_max": 15000,
10
+ "dec_num_layers": 8,
11
+ "dec_num_heads": 4,
12
+ "dec_res_dropout": 0.1,
13
+ "dec_pred_dropout": 0.1,
14
+ "dec_att_dropout": 0.1,
15
+ "dec_dim_feedforward": 256,
16
+ "attention_win": 100,
17
+ "use_tokens_from_all_lines": true,
18
+ "use_first_pass_tokens": true,
19
+ "two_step_pos_enc_mode": "cat",
20
+ "use_line_indices": false,
21
+ "vocab_size": 99,
22
+ "additional_tokens": 1,
23
+ "max_char_prediction": 3000,
24
+ "max_line_pred": 100,
25
+ "max_pred_per_line": 150,
26
+ "working_dpi": 150,
27
+ "image_mean": [
28
+ 202.70203512453844,
29
+ 190.85819291654516,
30
+ 136.08993740340918
31
+ ],
32
+ "image_std": [
33
+ 71.95683449787683,
34
+ 72.00679752035164,
35
+ 57.33656372156373
36
+ ],
37
+ "expected_parameters": 7033700,
38
+ "parameters_encoder": 1706240,
39
+ "parameters_decoder": 5327460,
40
+ "source_checkpoint": {
41
+ "path": "outputs/e14_budget_1p26M_s0/checkpoints/best_3580.pt",
42
+ "sha256": "d94af1419554a057ae671733af2fe4cf3f63d1f8b3652f9de17277b432632f5f",
43
+ "bytes": 84821823,
44
+ "epoch": 3580,
45
+ "step": 1253350,
46
+ "recorded_best_valid_cer": 0.0409,
47
+ "use_line_indices_source": "checkpoint carries no 'use_line_indices' key; inferred from additional_tokens == 1, i.e. (additional_tokens == 3) -> False"
48
+ }
49
+ }
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:b842d1e1e1d9568825253b79d01b91702635a665d19227c9a2748de0ec49515b
3
+ size 28169776
preprocessor_config.json ADDED
@@ -0,0 +1,20 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "do_rgb": true,
3
+ "working_dpi": 150,
4
+ "resample": "PIL.Image.BILINEAR",
5
+ "image_mean": [
6
+ 202.70203512453844,
7
+ 190.85819291654516,
8
+ 136.08993740340918
9
+ ],
10
+ "image_std": [
11
+ 71.95683449787683,
12
+ 72.00679752035164,
13
+ 57.33656372156373
14
+ ],
15
+ "do_resize_to_fixed_size": false,
16
+ "do_pad": false,
17
+ "encoder_reduction_h": 32,
18
+ "encoder_reduction_w": 8,
19
+ "note": "Training-set channel statistics for READ_2016_page_sem_dan, read from outputs/e14_budget_1p26M_s0/results/params.txt:202-221. Scale is 0-255, NOT 0-1, and these are NOT ImageNet statistics. The page must reach the model at 150 dpi: raw READ 2016 scans are 300 dpi and must be halved (PIL BILINEAR); the pages under formatted/ already are 150 dpi and must NOT be resized again."
20
+ }
spec_heads_m5.json ADDED
@@ -0,0 +1,13 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "m": 5,
3
+ "hidden": 256,
4
+ "vocab_out": 100,
5
+ "parameters": 365968,
6
+ "d_model": 256,
7
+ "base_checkpoint_sha256": "d94af1419554a057ae671733af2fe4cf3f63d1f8b3652f9de17277b432632f5f",
8
+ "source": {
9
+ "path": "outputs/spec_heads_m5/heads.pt",
10
+ "sha256": "3598151fd92a50fd722b3ac0931310ad0ca2776b62dd84339ad51ade2dd71976",
11
+ "bytes": 1469127
12
+ }
13
+ }
spec_heads_m5.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:fc73d191871da0241bcd82ea68d2844ffea2385635712c9951802ecfe9bbbf2a
3
+ size 1465160