docTR PARSeq — Russian + Latin text recognition (ONNX)

Word-level text recognition model for printed Russian documents (Cyrillic + Latin + digits + punctuation), fine-tuned from the docTR parseq checkpoint and exported to ONNX.

It is a drop-in recognition model for docTR / OnnxTR OCR pipelines: pair it with any docTR text detector (e.g. db_resnet50) to read full pages.

Кратко по-русски. Модель распознавания слов для печатных документов на русском языке (кириллица + латиница + цифры + пунктуация). Дообучена из parseq docTR, экспортирована в ONNX. Используется вместе с детектором docTR/OnnxTR для распознавания страниц. Точность на синтетической валидации — 96.2% слов целиком. Есть более быстрая и лёгкая версия: doctr-crnn-vgg16-bn-russian-onnx.

Architecture PARSeq (ViT encoder + permuted autoregressive decoder), docTR implementation
Parameters ~24M (92 MB fp32 ONNX)
Input input: float32 [N, 3, 32, 128]
Output logits: float32 [N, 33, 175] — 32 characters max + EOS
Vocabulary 174 characters (russian_latin, see config.json)
ONNX opset 17

Files

  • model.onnx — the model (dynamic batch size).
  • config.json — everything needed for pre/post-processing: vocab, mean, std, input_shape, eos_index, ...

Usage

Full page OCR with OnnxTR

import json
from huggingface_hub import hf_hub_download
from onnxtr.io import DocumentFile
from onnxtr.models import ocr_predictor, parseq

repo = "graycat660/doctr-parseq-russian-onnx"
cfg = json.load(open(hf_hub_download(repo, "config.json"), encoding="utf-8"))
reco = parseq(hf_hub_download(repo, "model.onnx"), vocab=cfg["vocab"])

predictor = ocr_predictor(det_arch="db_resnet50", reco_arch=reco)
doc = DocumentFile.from_images("page.png")  # or DocumentFile.from_pdf("doc.pdf")
print(predictor(doc).render())

Single words with onnxruntime only

import json
import numpy as np
import onnxruntime as ort
from huggingface_hub import hf_hub_download
from PIL import Image

repo = "graycat660/doctr-parseq-russian-onnx"
cfg = json.load(open(hf_hub_download(repo, "config.json"), encoding="utf-8"))
sess = ort.InferenceSession(hf_hub_download(repo, "model.onnx"))

def preprocess(img):
    # keep aspect ratio, pad bottom/right with zeros, normalize
    _, h, w = cfg["input_shape"]
    img = img.convert("RGB")
    s = min(h / img.height, w / img.width)
    nh, nw = max(1, round(img.height * s)), max(1, round(img.width * s))
    x = np.zeros((h, w, 3), np.float32)
    x[:nh, :nw] = np.asarray(img.resize((nw, nh), Image.BILINEAR), np.float32) / 255
    return ((x - cfg["mean"]) / cfg["std"]).transpose(2, 0, 1).astype(np.float32)

def recognize(images):
    logits = sess.run(None, {cfg["input_name"]: np.stack([preprocess(i) for i in images])})[0]
    words = []
    for seq in logits.argmax(-1):  # greedy decoding up to the first EOS
        eos = seq == cfg["eos_index"]
        words.append("".join(cfg["vocab"][i] for i in seq[: eos.argmax() if eos.any() else len(seq)]))
    return words

print(recognize([Image.open("word.png")]))

Vocabulary

Russian alphabet (incl. ё, ъ), Latin a–z A–Z, digits, ASCII punctuation, £€¥¢฿₽ and №«»—–…°§. Spaces are not part of the vocabulary: the model reads single words (as cropped by a text detector). Letters with stress marks (а́) are not supported.

Training

  • Initialization: docTR parseq pretrained weights (French vocab); embedding and classification head re-initialized for the new vocab.
  • Data: 1.5M synthetic word images (+20k validation) rendered with ~130 Cyrillic-capable fonts (serif, sans, mono, condensed, italic/bold, decorative). Text mix: Russian words (58%, sampled from a 50k frequency list with flattened Zipf distribution), English words (10%), numbers/dates/times/phones/amounts/percentages (12%), random character strings (7%), abbreviations such as ООО, ИНН, т.д. (5%), hyphenated compounds (4%), initials and list markers (4%); random case and surrounding punctuation («», (), ,.:;!?). Rendering augmentations: background/foreground colors incl. inverted, letter spacing, underline, rotation ±3°, blur, noise, low resolution (24–48 px height), JPEG quality 60–95; plus docTR training augmentations (perspective, shadow, photometric distortion, blur, noise).
  • Recipe: docTR references/recognition/train.py, 12 epochs, batch 128, AdamW lr 3e-4 with cosine schedule, bf16 AMP, single RTX 4060 Laptop GPU (~13 h).

Evaluation

On the held-out synthetic validation set (20,000 words, unseen text samples, same font pool):

Model Exact match Case-insensitive match
this model (PARSeq) 96.18% 97.78%
CRNN VGG16-BN variant 95.51% 97.27%
stock docTR model (French vocab) does not read Cyrillic —

ONNX and PyTorch predictions are identical (verified on 512 samples).

CPU latency (onnxruntime 1.30, Intel i9-13900H): ~33 ms per word at batch 1, ~19 ms per word at batch 32.

Note: validation data is synthetic and drawn from the same generator as training; accuracy on real scans/photos will be lower.

Limitations

  • Trained on synthetic printed text only: expect lower accuracy on noisy scans, photos, and no support for handwriting.
  • Ambiguous glyphs without context: italic и vs Latin u, case of letters with identical shapes (с/С, о/О, в/В at small sizes), Cyrillic vs Latin look-alikes in short uppercase words (ВС vs BC).
  • Footnote superscripts ([13]) are sometimes misread.
  • Max 32 characters per word.

Fine-tuning on a few thousand real labeled word crops from your domain is the most effective way to improve accuracy.

License and attribution

Released under Apache-2.0, the same license as docTR and its pretrained weights.

  • Base model and training code: mindee/doctr (Apache-2.0).
  • Word lists used to generate training text: hermitdave/FrequencyWords (content CC-BY-SA-4.0, derived from OpenSubtitles). Neither the word lists nor the rendered images are redistributed here.
  • Fonts were used only to render training images; no font files are included.
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support