ScienceSoft multilingual OCR recogniser (en · ru · de · fr · es · it · ar)
One text-line recogniser for printed Latin, Cyrillic and Arabic text in seven languages, including identity-document text — passport / ID-card MRZ lines and field labels, and the identity documents of the Gulf states (UAE, Saudi Arabia, Qatar, Kuwait, Bahrain, Oman). It is the recognition half of the OCR step in a data-loss-prevention agent that reads text in images leaving a device, so personal data in a screenshot can be found and redacted before it reaches an AI assistant.
- PP-OCR mobile architecture (PPLCNetV3 backbone + SVTR neck + CTC head), 1.97 M parameters, re-implemented and trained in PyTorch, exported to ONNX.
- Drop-in for PP-OCR-style pipelines: same input/output contract as a PaddleOCR recognition model exported to ONNX, so it runs in RapidOCR or PaddleOCR's ONNX mode with your own detector (see Usage).
rec-multi.onnx: int8 (static QDQ, convolutions only, percentile-calibrated), 3.2 MB, CPU.rec-multi.fp32.onnx(7.9 MB) is the float twin — use it for further conversion (e.g. X2Paddle) or fine-tuning.- 659-entry dictionary: Latin incl. French/German/Spanish/Italian letters, full Cyrillic, Arabic letters (plus پ چ ژ گ ک ی ڤ for names and loanwords), Western and Arabic-Indic digits, punctuation (incl. ، ؛ ؟ ٪), currency and typographic symbols.
It is a recogniser only: it reads one cropped text line. Pair it with a text detector (e.g. PP-OCR DB).
Not affiliated with or endorsed by PaddlePaddle/Baidu; "PP-OCR" names the architecture it follows.
Input / output
input x |
float32 [N, 3, 48, W], BGR, (pixel/255 − 0.5) / 0.5; resize the line crop to height 48 keeping aspect, right-pad with 0.0 (after normalisation). Trained with W padded to 960; dynamic W works. |
output probs |
float32 [N, W/8, 661], softmax over [blank, dict..., " "] |
| decode | greedy CTC: argmax per step, collapse repeats, drop blank (index 0); index i → dict[i-1], index 660 → space |
Arabic comes out in DRAWN order — reorder it
A CTC recogniser reads left to right, so an Arabic line is emitted in the order
its characters are drawn: الاسم: محمد comes out دمحم :مسالا. Put it back in
written order after decoding. PaddleOCR does this only when the dictionary file
name contains arabic, which dict-multi.txt does not — and its pred_reverse
leaves the separator in the wrong place (الاسم :محمد). This does it for one
line, as a renderer following the Unicode bidi algorithm (UAX #9) draws it:
import re
L = "A-Za-zÀ-ɏЀ-ӿ" # Latin + Cyrillic letters
D = "0-9٠-٩۰-۹" # Western + Arabic-Indic digits
A = "ؠ-يٮ-ۓە" # Arabic letters
_ar, _strong, _letter = re.compile(f"[{A}]"), re.compile(f"[{A}{L}]"), re.compile(f"[{L}]")
_ltr_run = re.compile(f"[{L}{D}](?:[^{A}]*[{L}{D}])?")
_ar_run = re.compile(f"[{A}](?:[^{L}{D}]*[{A}])?")
_number = re.compile(f"[{D}]+(?:[.,:/ ٫٬][{D}]+)*")
_mirror = str.maketrans("()[]{}<>«»", ")(][}{><»«")
def _ltr_back(run):
if _letter.search(run): # Latin text keeps its own order
return run[::-1].translate(_mirror)
return _number.sub(lambda m: m.group()[::-1], run) # after Arabic, each number is its own run
def to_written_order(drawn: str) -> str:
if not _ar.search(drawn):
return drawn # Latin / Cyrillic: unchanged
if _ar.match(_strong.findall(drawn)[-1]): # rightmost letter Arabic: a right-to-left line
return _ltr_run.sub(lambda m: _ltr_back(m.group()), drawn[::-1].translate(_mirror))
return _ar_run.sub(lambda m: m.group()[::-1], drawn) # a left-to-right line with Arabic words
to_written_order("1-1234567-1990-784 :ةيوهلا مقر") # 'رقم الهوية: 784-1990-1234567-1'
It guesses the line's direction from the drawing: a left-to-right line that
ends in Arabic (Name: محمد in an English form) comes back as if it were
right-to-left. The ai-firewall runtime (rust/ocr-core/src/bidi.rs) has the same
limit; both put 12 of 14 mixed test lines back exactly, the other two being that
case.
Usage
RapidOCR (PaddleOCR's ONNX port) — verified
# pip install rapidocr_onnxruntime==1.4.4 opencv-python-headless
from huggingface_hub import hf_hub_download
from rapidocr_onnxruntime import RapidOCR
repo = "ScienceSoft/scnsoft-ocr-rec-multilingual"
engine = RapidOCR(
rec_model_path=hf_hub_download(repo, "rec-multi.onnx"),
rec_keys_path=hf_hub_download(repo, "dict-multi.txt"),
)
result, _ = engine("screenshot.png")
for box, text, score in result or []:
print(f"{score:.2f} {text}") # Arabic: see "drawn order" above
onnxruntime directly (one cropped line)
See examples/onnxruntime_line.py — preprocessing
and CTC decoding in ~40 lines of numpy.
PaddleOCR 2.x ONNX mode
use_onnx=True, rec_model_dir=rec-multi.onnx, rec_char_dict_path=dict-multi.txt,
use_space_char=True, rec_image_shape="3,48,320". Same contract as RapidOCR;
not separately verified. A native Paddle (.pdmodel) version is not provided.
Versions
- r4 (this revision, 2026-10-02): adds Arabic — Modern Standard Arabic and seven dialects, Gulf identity documents, right-to-left interfaces. Fine-tuned from r3 for 30 k steps.
- r3 (2026-09-30, not published separately): detector crops at DB unclip ratio 1.5, browser-rendered screenshot text, OCR-B MRZ.
- r2 (git tag
r2): the first release, trained for a tighter 2 px crop.
Training
- Data: 12.6 M line images of printed text in the seven languages — document layouts, prose, interface text rendered by a real browser at 1×–2×, identity-document text (MRZ in ICAO 9303 TD1/TD2/TD3 with valid check digits, ID-card fields), Gulf identity documents (numbers that pass the issuers' checks; Hijri and Gregorian dates in Western and Arabic-Indic digits) and Arabic interfaces laid out right to left, across ~170 font faces. All personal-data values are synthetic. The training data is not published.
- Crops match a real detector: drawn from offsets measured between ground-truth boxes and a PP-OCRv3 DB detector's boxes (85 % at unclip ratio 1.5, 15 % at a tight 2 px expansion).
- Batch 256, AdamW + cosine, bf16, every batch at width 960, one NVIDIA DGX Spark (GB10). BatchNorm statistics re-estimated before every evaluation. The checkpoint was chosen on 34 separate validation sets, in fp32.
Evaluation
End to end (PP-OCRv3 DB detector, unclip ratio 1.5 → this recogniser → written order), character error rate, on held-out synthetic sets. Baselines are PaddleOCR recognisers (ONNX) through the identical pipeline — PP-OCRv5 mobile for Latin/Cyrillic, PP-OCRv3 for Arabic (the Arabic one available as ONNX).
Arabic
| eval set | PP-OCRv3 arabic | r4 |
|---|---|---|
| prose: MSA / Najdi / Egyptian | 0.178 / 0.177 / 0.178 | 0.068 / 0.066 / 0.074 |
| prose: Levantine / Iraqi / Maghrebi | 0.173 / 0.208 / 0.176 | 0.075 / 0.076 / 0.065 |
| Arabic UI pages, unseen fonts | 0.206 | 0.097 |
| Gulf ID documents (6 states) | 0.082–0.106 | 0.018–0.024 (91–93 % of lines exact) |
| MRZ lines on Gulf documents, exact | 0.00 | 0.52–0.71 |
Latin / Cyrillic
| eval set | best PP-OCRv5 per language | r3 | r4 |
|---|---|---|---|
| rendered en / de / es / fr / it | 0.024 / 0.030 / 0.022 / 0.023 / 0.040 | 0.019 / 0.022 / 0.015 / 0.024 / 0.034 | 0.016 / 0.018 / 0.011 / 0.017 / 0.029 |
| rendered ru | 0.058 | 0.028 | 0.027 |
| document-layout lines en / ru | 0.185 / 0.382 | 0.209 / 0.265 | 0.190 / 0.262 |
| ID-card mock-ups en / de / es / fr / it / ru | 0.050 / 0.046 / 0.062 / 0.058 / 0.043 / 0.093 | 0.046 / 0.027 / 0.040 / 0.037 / 0.024 / 0.049 | 0.057 / 0.029 / 0.042 / 0.036 / 0.024 / 0.071 |
| web pages, unseen fonts | 0.118 | 0.069 | 0.061 |
r4 beats the best per-language PP-OCRv5 model on 13 of 15 Latin/Cyrillic sets and the PP-OCRv3 Arabic model on all 13 Arabic sets (mean 0.144 → 0.050). The int8 file is within 0.002 CER of its fp32 twin on these sets.
Read these numbers for what they are: every set is synthetic and comes from the same kinds of sources the model was trained on (held-out samples, not held-out sources), while the PaddleOCR models saw none of them — that favours this model. The checkpoint was chosen on separate validation pages. The model has not been benchmarked on real scanned or photographed documents or real screenshots.
Limitations
- Evaluated on synthetic data only (see above).
- Arabic output is drawn order — reorder it (see above).
- Latin/Cyrillic ID cards read worse than r3 (en 0.057 vs 0.046, ru 0.071 vs
0.049): small ALL-CAPS lines (under ~20 px) are often read partly or wholly
in lower case (
ВОДИТЕЛЬСКОЕ→водительСкое). Mixed-case text is not affected. - MRZ: 48–71 % of lines read exactly (monospace faces); 22 % in OCR-B. Do not rely on it for identity verification.
- Arabic: more tokens fall under a 0.9 confidence threshold than in Latin text, so a consumer that redacts low-confidence tokens over-masks Arabic more.
- Trained for crops at DB unclip ratio ~1.5 (PaddleOCR's default).
- Printed text only; no handwriting. Horizontal lines only (trained ≤ 2° tilt).
- A line longer than ~110 characters exceeds the 120 CTC steps at width 960.
- No CJK, Greek, Hebrew, Persian/Urdu-specific letters beyond those listed, or other scripts.
Intended use
Reading printed text in screenshots and documents, in particular to find personal data for redaction. Not intended for decisions about individuals (e.g. identity verification).
Licence
Proprietary to ScienceSoft. The weights, the dictionary and the example code are published so that the results on this page can be evaluated and reproduced; no licence to use them in production, to modify them or to redistribute them is granted. For licensing, contact ScienceSoft (https://www.scnsoft.com).
The network follows the PP-OCR mobile text-recognition architecture described by the PaddleOCR project (Apache-2.0, https://github.com/PaddlePaddle/PaddleOCR). It was re-implemented and trained independently; no PaddleOCR weights were used.
Citation
@misc{sciencesoft2026ocrmultilingual,
title = {ScienceSoft multilingual OCR recogniser (en, ru, de, fr, es, it, ar)},
author = {{ScienceSoft}},
year = {2026},
url = {https://huggingface.co/ScienceSoft/scnsoft-ocr-rec-multilingual}
}