MonOCR โ line-level OCR for Mon (mnw)
A CRNN that reads one cropped line of Mon text and returns a string. This repository holds the ONNX and Core ML exports, their charset and their sidecars.
It runs in production on the web at ocr.mondevhub.com. The Android and iOS apps bundle the same artifacts and build from source; neither is in an app store yet.
Held-out CER 0.0100 on 150 unseen lines in a typeface the model never trained on. Read what that number does not cover before quoting it.
Use
Pin a revision. The SDKs pin d3d9d5e; every later revision so far has changed
only this card, not the artifacts.
from huggingface_hub import hf_hub_download
model = hf_hub_download("janakhpon/monocr", "onnx/monocr.onnx", revision="d3d9d5e")
charset = hf_hub_download("janakhpon/monocr", "onnx/charset.txt", revision="d3d9d5e")
MonDevHub/monocr-onnx wraps the model with line segmentation and decoding for Python, JavaScript, Go and Rust. Its bindings do not yet agree on page-level output; its README has the figures.
Files
| Path | What |
|---|---|
onnx/monocr.onnx |
ONNX FP32, opset 17, 46,247,040 bytes |
onnx/monocr.json |
sidecar: charset, geometry, class count, normalization |
onnx/charset.txt |
the 276 characters for indices 1..276, on one line; the first is U+0020 |
coreml/monocr.mlpackage |
Core ML, FP32 |
coreml/monocr.mlpackage.json |
its sidecar |
charset.txt, monocr.json |
copies at the root, for consumers that expect them there |
Only deployment artifacts are published. There is no PyTorch checkpoint: nothing in the SDKs or apps reads one.
Input contract
- Grayscale, 1 channel, float32, shape
[batch, 1, 160, 1024]. The batch axis is dynamic; height and width are static. - Aspect-preserving resize to height 160, then pad the width to 1024.
- Normalize
pixel / 127.5 - 1.0. - Pad with white, which is
+1.0after normalization, not0.0. - Greedy CTC decode over 277 classes; index 0 is the blank.
Feeding raw uint8 in [0, 255] produces confident garbage with no error.
Load the charset from the same revision as the weights. A charset from another
revision decodes every index to the wrong character and raises nothing. Strip
only \n and \r: the first class is a space, and a bare .strip() removes it
and shifts every index by one.
Architecture
MobileNetV3-Large + squeeze-excitation neck โ band pooling โ 2รBiLSTM(512) โ bottleneck self-attention (256-dim, 4 heads) โ Linear(1024 โ 277) โ CTC. 11,553,437 parameters, all trainable.
Moving from v2
v2 is still served at revision a51be11 and will stay there; anything pinned to it
is unaffected.
v3.5 is not a drop-in replacement:
v2 (a51be11) |
v3.5 (d3d9d5e) |
|
|---|---|---|
| Input height | 128 | 160 |
| Input width | dynamic | static 1024 |
| Batch axis | fixed at 1 | dynamic |
| Output classes | 316 | 277 |
| Charset | 315 characters | 276 characters |
| Parameters | 6,575,868 | 11,553,437 |
The width is the change that breaks integrations. nn.MultiheadAttention fixes
the sequence length when the graph is traced, so the graph accepts a width of 1024
and nothing else.
Evaluation
Held-out test
Measured 2026-08-16 on the test split, held out from training and scored once.
| CER | CI95 | sequence accuracy | |
|---|---|---|---|
| Overall | 0.0100 | [0.0056, 0.0147] | 85.33% |
| Mon (n=99) | 0.0113 | [0.0050, 0.0186] | 85.86% |
| Burmese (n=32) | 0.0027 | [0.0000, 0.0064] | 93.75% |
| English (n=14) | 0.0103 | [0.0009, 0.0226] | 71.43% |
| Mon + English (n=5) | 0.0144 | [0.0000, 0.0292] | 60.00% |
n = 150 lines, greedy decoding. Macro-averaged CER 0.0047. Expected calibration error 0.0315. Per typeface: Pyidaungsu-Regular 0.0076 (n=56), Pyidaungsu-Numbers 0.0108 (n=51), Pyidaungsu-Bold 0.0124 (n=43).
Three baselines ran alongside it:
| Baseline | CER | What it rules out |
|---|---|---|
| empty prediction | 1.0000 | A broken scorer. Anything other than exactly 1.0 and the suite refuses to report |
| most-common grapheme | 1.0832 | That the charset's prior alone explains the result |
| 1-nearest-neighbour pixel retrieval over 3,000 training images | 1.3675 | Near-duplicates. Retrieving the closest training image scores worse than predicting nothing, so test images are not near-duplicates of training ones. It does not rule out memorisation in general |
No train/serve skew measured. The PyTorch checkpoint and the published ONNX
graph both scored 0.0100 (cer_delta 0.0) and produced exactly the same string on
150 of 150 lines, so the number describes the artifact in this repository.
Latency p50 134.8 ms/line on CPU, batched. It is a per-batch timing divided by batch size, so it is a mean, not a tail.
What the held-out number does not cover
- n = 150. The interval is [0.0056, 0.0147]; treat the width as real.
- One typeface. All 150 lines are Pyidaungsu, held out from training. A
second held-out design,
yunghkio, is not represented: this split was generated before that design was set aside. - Unseen text, not an unseen renderer. The same generator and augmentation pipeline produced training and test images, so a renderer defect is learned, validated and tested against identically. This is the largest caveat on the card, and only real photographed lines close it.
- Disjointness is argued, not directly verified. The training-side labels were produced on a machine whose state was not fully retained. The partition function is byte-unchanged since generation and none of the 150 test labels appears in the local training or validation sets, but that is an argument from the stability of a hash, not a direct comparison.
Against v2, on identical images
Both generations ship an ONNX export, so the same rendered lines go through both graphs with the same preprocessing and greedy decode. Text is restricted to the 273 characters both charsets can emit, so neither model is charged for a character it has no class for. 600 lines per arm, rendered 2026-08-15.
| Rendered in | n | v2 CER | v3.5 CER | error reduced |
|---|---|---|---|---|
| the 32 trained designs | 600 | 0.1470 | 0.0396 | 73% |
namkhon, held out from v3.5 |
600 | 0.0521 | 0.0188 | 64% |
pyidaungsu, yunghkio, held out from v3.5 |
600 | 0.0342 | 0.0051 | 85% |
v2 predates font-disjoint splits, so the bottom two rows are held out from v3.5 only, which tilts the comparison against v3.5. This is a preview, not an evaluation: one rendering pipeline reading its own output, synthetic for both models. It says v3.5 reads rendered Mon better than v2 did, nothing more.
Wide lines
Which handling wins depends on width. Measured 2026-08-22 over 201 rendered lines: squeezing a whole line into the 1024px canvas wins at 2 tiles, the two are level at 3, and cutting it into canvas-width tiles at whitespace columns wins from 4 up. On a book page at 150 dpi every line fitted one tile. Re-measure before swapping models.
Export checks
Both exports are gated against the PyTorch model on a seeded uniform [-1, 1]
input, and the gate fails the build:
- ONNX: logits within
rtol=atol=1e-3. - Core ML: logits within
1e-3and an identical CTC-decoded string.
The Core ML gate runs on CPU. The Neural Engine computes in fp16 and its compiler may reassociate, so what runs on a device is not what was verified.
Limitations
- Real photographs. Every training sample is synthetic, rendered by one generator. A camera photo of physical text is out of domain and fails confidently: on a video frame of a whiteboard the model returned fluent Mon at confidence 0.83 for text that appears nowhere on the page. Confidence is not a usable filter for this.
- Page layout. The input is a cropped line. There is no detector and no reading-order model.
- Handwriting, except Myanmar digits.
- Beam search. Only greedy decoding has been evaluated.
- Latency. No per-line tail has been measured, and no on-device number exists on any platform.
- Typeface coverage. 35 distinct designs across 79 usable font files bound every generalisation claim above.
Training data
Rendered from a Mon text corpus. The corpus licensing is not fully settled: one source has no established licence, and CC BY-SA attribution for another cannot be reconstructed because per-article URLs were stripped at import. The MIT licence on this repository covers the model weights and the code that produced them; it does not resolve the provenance of the text they were rendered from.
Citation
@misc{monocr,
title = {MonOCR: line-level OCR for the Mon language},
author = {Zin Min},
year = {2026},
url = {https://huggingface.co/janakhpon/monocr}
}
- Downloads last month
- 10