docTR PARSeq — Russian + Latin text recognition (ONNX)
Word-level text recognition model for printed Russian documents (Cyrillic + Latin + digits + punctuation),
fine-tuned from the docTR parseq checkpoint and exported to ONNX.
It is a drop-in recognition model for docTR / OnnxTR OCR pipelines:
pair it with any docTR text detector (e.g. db_resnet50) to read full pages.
Кратко по-русски. Модель распознавания слов для печатных документов на русском языке (кириллица + латиница + цифры + пунктуация). Дообучена из
parseqdocTR, экспортирована в ONNX. Используется вместе с детектором docTR/OnnxTR для распознавания страниц. Точность на синтетической валидации — 96.2% слов целиком. Есть более быстрая и лёгкая версия: doctr-crnn-vgg16-bn-russian-onnx.
| Architecture | PARSeq (ViT encoder + permuted autoregressive decoder), docTR implementation |
| Parameters | ~24M (92 MB fp32 ONNX) |
| Input | input: float32 [N, 3, 32, 128] |
| Output | logits: float32 [N, 33, 175] — 32 characters max + EOS |
| Vocabulary | 174 characters (russian_latin, see config.json) |
| ONNX opset | 17 |
Files
model.onnx— the model (dynamic batch size).config.json— everything needed for pre/post-processing:vocab,mean,std,input_shape,eos_index, ...
Usage
Full page OCR with OnnxTR
import json
from huggingface_hub import hf_hub_download
from onnxtr.io import DocumentFile
from onnxtr.models import ocr_predictor, parseq
repo = "graycat660/doctr-parseq-russian-onnx"
cfg = json.load(open(hf_hub_download(repo, "config.json"), encoding="utf-8"))
reco = parseq(hf_hub_download(repo, "model.onnx"), vocab=cfg["vocab"])
predictor = ocr_predictor(det_arch="db_resnet50", reco_arch=reco)
doc = DocumentFile.from_images("page.png") # or DocumentFile.from_pdf("doc.pdf")
print(predictor(doc).render())
Single words with onnxruntime only
import json
import numpy as np
import onnxruntime as ort
from huggingface_hub import hf_hub_download
from PIL import Image
repo = "graycat660/doctr-parseq-russian-onnx"
cfg = json.load(open(hf_hub_download(repo, "config.json"), encoding="utf-8"))
sess = ort.InferenceSession(hf_hub_download(repo, "model.onnx"))
def preprocess(img):
# keep aspect ratio, pad bottom/right with zeros, normalize
_, h, w = cfg["input_shape"]
img = img.convert("RGB")
s = min(h / img.height, w / img.width)
nh, nw = max(1, round(img.height * s)), max(1, round(img.width * s))
x = np.zeros((h, w, 3), np.float32)
x[:nh, :nw] = np.asarray(img.resize((nw, nh), Image.BILINEAR), np.float32) / 255
return ((x - cfg["mean"]) / cfg["std"]).transpose(2, 0, 1).astype(np.float32)
def recognize(images):
logits = sess.run(None, {cfg["input_name"]: np.stack([preprocess(i) for i in images])})[0]
words = []
for seq in logits.argmax(-1): # greedy decoding up to the first EOS
eos = seq == cfg["eos_index"]
words.append("".join(cfg["vocab"][i] for i in seq[: eos.argmax() if eos.any() else len(seq)]))
return words
print(recognize([Image.open("word.png")]))
Vocabulary
Russian alphabet (incl. ё, ъ), Latin a–z A–Z, digits, ASCII punctuation, £€¥¢฿₽ and №«»—–…°§.
Spaces are not part of the vocabulary: the model reads single words (as cropped by a text detector).
Letters with stress marks (а́) are not supported.
Training
- Initialization: docTR
parseqpretrained weights (French vocab); embedding and classification head re-initialized for the new vocab. - Data: 1.5M synthetic word images (+20k validation) rendered with ~130 Cyrillic-capable fonts (serif, sans, mono, condensed, italic/bold, decorative).
Text mix: Russian words (58%, sampled from a 50k frequency list with flattened Zipf distribution), English words (10%),
numbers/dates/times/phones/amounts/percentages (12%), random character strings (7%), abbreviations such as
ООО,ИНН,т.д.(5%), hyphenated compounds (4%), initials and list markers (4%); random case and surrounding punctuation («»,(),,.:;!?). Rendering augmentations: background/foreground colors incl. inverted, letter spacing, underline, rotation ±3°, blur, noise, low resolution (24–48 px height), JPEG quality 60–95; plus docTR training augmentations (perspective, shadow, photometric distortion, blur, noise). - Recipe: docTR
references/recognition/train.py, 12 epochs, batch 128, AdamW lr 3e-4 with cosine schedule, bf16 AMP, single RTX 4060 Laptop GPU (~13 h).
Evaluation
On the held-out synthetic validation set (20,000 words, unseen text samples, same font pool):
| Model | Exact match | Case-insensitive match |
|---|---|---|
| this model (PARSeq) | 96.18% | 97.78% |
| CRNN VGG16-BN variant | 95.51% | 97.27% |
| stock docTR model (French vocab) | does not read Cyrillic | — |
ONNX and PyTorch predictions are identical (verified on 512 samples).
CPU latency (onnxruntime 1.30, Intel i9-13900H): ~33 ms per word at batch 1, ~19 ms per word at batch 32.
Note: validation data is synthetic and drawn from the same generator as training; accuracy on real scans/photos will be lower.
Limitations
- Trained on synthetic printed text only: expect lower accuracy on noisy scans, photos, and no support for handwriting.
- Ambiguous glyphs without context: italic
иvs Latinu, case of letters with identical shapes (с/С,о/О,в/Вat small sizes), Cyrillic vs Latin look-alikes in short uppercase words (ВСvsBC). - Footnote superscripts (
[13]) are sometimes misread. - Max 32 characters per word.
Fine-tuning on a few thousand real labeled word crops from your domain is the most effective way to improve accuracy.
License and attribution
Released under Apache-2.0, the same license as docTR and its pretrained weights.
- Base model and training code: mindee/doctr (Apache-2.0).
- Word lists used to generate training text: hermitdave/FrequencyWords (content CC-BY-SA-4.0, derived from OpenSubtitles). Neither the word lists nor the rendered images are redistributed here.
- Fonts were used only to render training images; no font files are included.
- Downloads last month
- -