tetrak_hy — Armenian text recognition for EasyOCR
An Armenian text recogniser packaged as an EasyOCR custom model, trained by tetrak-hy-trainer for Tetrak, an OCR pipeline for community archives. The architecture is EasyOCR's own generation2 recognition network (VGG feature extractor, two BiLSTM layers, CTC head), so the model drops into a stock EasyOCR install.
Status: v6, alpha
v6 is v5 fine-tuned on twice as many real crops: 99,521 cut from scanned pages of 16 works, not rendered ones. It reads eight kinds of Armenian print: an encyclopedia, a medical encyclopedia, a bilingual dictionary, a scholarly history, and literary editions in Eastern and Western Armenian.
With the companion package's post-processing (below), v6 leads every Armenian OCR engine we have measured on word recall on three of those eight registers, and is within 0.001 on a fourth. Those engines include Calfa's hye-calfa-n and hye-paddle, which are CC BY-NC. This model is Apache 2.0.
| Register (held-out pages) | v6 | v6 + fold + word list | best other engine |
|---|---|---|---|
| Otyan, Works (Western Armenian) | 0.934 | 0.944 | 0.934 hye-paddle |
| Totovents, Works | 0.929 | 0.945 | 0.941 hye-calfa-n |
| Baronian, Works vol. 10 | 0.911 | 0.920 | 0.918 hye-paddle |
| Tumanyan, academic edition vol. 5 | 0.895 | 0.905 | 0.913 hye-paddle |
| Faustus of Byzantium (1968) | 0.873 | 0.897 | 0.922 hye-paddle |
| Medical encyclopedia | 0.926 | 0.929 | 0.933 hye-paddle |
| Armenian Soviet Encyclopedia vol. 2 | 0.837 | 0.839 | 0.865 hye-paddle |
| Armenian–English dictionary | 0.667 | 0.677 | 0.678 marker |
Word recall: the share of the transcript's words found in the output. Ten pages per register (five for the dictionary), never trained on, proofread on Armenian Wikisource. Lines are joined in detector order, as EasyOCR returns them. Pages that print an evaluation page's text are kept out of v6's new training data, but not out of what v6 inherits from v5 (below).
Character similarity measures reading order here, not recognition. In detector order, v6 scores 0.91–0.95 on the single-column registers, and 0.30–0.35 on the multi-column encyclopedias, where EasyOCR's output is not in column order. In Tetrak's pipeline, which orders columns before joining, the same weights score 0.970 on the encyclopedia and 0.965 on the medical encyclopedia.
These figures are not comparable with those published for v5 and earlier. The metric was corrected for v6. Character similarity used difflib with
autojunkon, which on a page of text ignores most of the alphabet. The ASCII colon and the Armenian full stop, and.and the abbreviation dot․, were also scored as different characters, though transcribers type one for the other. Under the corrected metric, v6 is ahead of v5 on every register (mean word recall 0.872 against 0.851).
Use v3 or later. v0 and v1 carry two defects that v2 fixed: 21% of their training labels were wrapped in quotation marks the images do not show, and their charset has no U+2024 ONE DOT LEADER, so 5.8% of the evaluation words were unwinnable. Both are recorded in those versions'
provenance.json.
The fold and the word list
The recognition head has no language model, so two kinds of slip remain that a post-process can fix without retraining. Both ship in tetrak-easyocr-armenian:
fold_scriptturns a Latin twin inside an Armenian word back into the Armenian letter (h→հ,:→։). It also readsՉandշas the digit 2 inside a number, which bold italic page numbers confuse.tetrak_hy.lexiconlooks at each word whose reading is not in a word list. It takes the most probable listed reading from the model's own alternatives for that word, if it is nearly as probable. On the held-out pages it fixed 465 words and broke 13: proper nouns, and classical or edition spellings the list does not know.
wordlist.tsv.gz in this repository is that list. It holds 1.13 million word forms, counted from proofread Armenian Wikisource transcripts with the evaluation pages excluded. It adds the Nayiri Armenian Lexicon (© Serouj Ourishian, CC BY 4.0). Because it is derived from those sources, the list is licensed CC BY-SA 4.0, not Apache 2.0 like the weights.
import tetrak_hy
from tetrak_hy import lexicon
reader = tetrak_hy.reader()
lexicon.use_lexicon(reader, lexicon.load_wordlist("wordlist.tsv.gz"))
results = [
(box, tetrak_hy.fold_script(text), confidence)
for box, text, confidence in reader.readtext("page.png", decoder="beamsearch")
]
decoder="beamsearch" is how the model's probabilities reach the word list. Without it, readtext decodes greedily as before.
What is still lost, and why
- Faustus of Byzantium: the index and notes are set in a bold italic whose digits v6 confuses (
2read as8or7). - Tumanyan: the academic edition's apparatus quotes Russian, and the charset has no Cyrillic.
Known defects
- Inherited evaluation-text exposure. 25 harvested pages print an evaluation page's text: variants and reprints in Tumanyan's academic edition and Baronian's collected works vol. 10, two pages of Faustus of Byzantium and one of the Armenian Soviet Encyclopedia. They are excluded from v6's new real crops, its all-caps set and the word list. But v6 starts from v5's weights and reuses v5's synthetic crops, both made before the exclusion existed, so the exposure is inherited. v6's figures on those four registers may be flattered by it. Removing it needs a new pre-train.
- Synthetic validation trained on. v6 also trained on the 750 crops meant to validate its all-caps set. No figure above is measured on them.
Charset
174 characters plus the CTC blank, 175 classes, unchanged since v5. A charset change is a new model by construction, because CTC class indices are positional. Always take the .yaml and the .pth from the same revision.
Validation
95.2% word accuracy (0.9933 normalised edit distance) on 10,597 real crops from pages held out of the fine-tune, split by page. Held-out crop accuracy is a poor proxy for page-level recall: it plateaus early and then measures overfitting. The page figures above are the ones to trust.
Files
tetrak_hy.pth— the weights exactly as the trainer saved them (keys carry themodule.prefix EasyOCR's loader expects to handle). This is the file EasyOCR loads.model.safetensors— the same tensors with themodule.prefix stripped, for anything that isn't EasyOCR.tetrak_hy.yaml— charset, language list and network parameters.tetrak_hy.py— the architecture module EasyOCR imports by name.wordlist.tsv.gz— the word list fortetrak_hy.lexicon(v6 on).provenance.json— training recipe, dataset revision, charset and checksums for this release.
Use with EasyOCR
Download the three EasyOCR files and place them where EasyOCR looks for custom models:
from huggingface_hub import hf_hub_download
for filename in ("tetrak_hy.pth", "tetrak_hy.py", "tetrak_hy.yaml"):
hf_hub_download("tetrak/easyocr-armenian", filename, revision="v6")
tetrak_hy.yamlandtetrak_hy.pygo in the user network directory (by default~/.EasyOCR/user_network/).tetrak_hy.pthgoes in the model directory (by default~/.EasyOCR/model/).
Then:
import easyocr
reader = easyocr.Reader(["en"], recog_network="tetrak_hy")
results = reader.readtext("page.png")
Note the ["en"]: with a custom recog_network, the language list
selects EasyOCR's dictionaries rather than the model — the recogniser
itself is chosen by recog_network, and this model's charset covers
Armenian plus basic Latin, digits and punctuation.
Pin revision= when downloading: each weights release is tagged, and
provenance.json records the exact dataset revision it was trained
from.
Training data
v6 is a fine-tune of v5, which was pre-trained on 351,000 synthetic line crops. Those crops were rendered in 15 Armenian faces from about 7,400 proofread Armenian Wikisource pages (CC BY-SA), then fine-tuned on real crops.
v6's real crops (99,521 for training) were cut from 862 scanned pages of 16 works by detection-assisted alignment against their transcripts. They are mixed in every batch with v5's synthetic crops and with all-caps lines for encyclopedia headwords. Transcript colons inside Armenian words are labelled as the Armenian full stop, which is what the page prints.
The real crops are not published. They are reproducible from the transcripts with the trainer's harvester. provenance.json records the recipe, the sources, the charset and the checksums. No font file is redistributed.
Licence
The weights, like the trainer, are Apache 2.0. The training text is CC BY-SA; we publish the text itself, share-alike, in the dataset repository above, and take the position — shared by most of the ecosystem, though not legally settled — that trained weights are not a redistribution or adaptation of the training text.
wordlist.tsv.gz is different: it is counted directly from CC BY-SA 3.0
Wikisource text and includes the CC BY 4.0 Nayiri Armenian Lexicon, so it
is licensed CC BY-SA 4.0, with the attribution recorded in
provenance.json.
Related
- tetrak-hy-trainer — synthesis, training and packaging (Apache 2.0).
- tetrak/armenian-ocr-crops — the training data (CC BY-SA 4.0).
- Tetrak — the OCR pipeline this model ships in.