MonOCR โ€” line-level OCR for Mon (mnw)

A CRNN that reads one cropped line of Mon text and returns a string. This repository holds the ONNX and Core ML exports, their charset and their sidecars.

It runs in production on the web at ocr.mondevhub.com. The Android and iOS apps bundle the same artifacts and build from source; neither is in an app store yet.

Held-out CER 0.0100 on 150 unseen lines in a typeface the model never trained on. Read what that number does not cover before quoting it.

Use

Pin a revision. The SDKs pin d3d9d5e; every later revision so far has changed only this card, not the artifacts.

from huggingface_hub import hf_hub_download

model = hf_hub_download("janakhpon/monocr", "onnx/monocr.onnx", revision="d3d9d5e")
charset = hf_hub_download("janakhpon/monocr", "onnx/charset.txt", revision="d3d9d5e")

MonDevHub/monocr-onnx wraps the model with line segmentation and decoding for Python, JavaScript, Go and Rust. Its bindings do not yet agree on page-level output; its README has the figures.

Files

Path What
onnx/monocr.onnx ONNX FP32, opset 17, 46,247,040 bytes
onnx/monocr.json sidecar: charset, geometry, class count, normalization
onnx/charset.txt the 276 characters for indices 1..276, on one line; the first is U+0020
coreml/monocr.mlpackage Core ML, FP32
coreml/monocr.mlpackage.json its sidecar
charset.txt, monocr.json copies at the root, for consumers that expect them there

Only deployment artifacts are published. There is no PyTorch checkpoint: nothing in the SDKs or apps reads one.

Input contract

  • Grayscale, 1 channel, float32, shape [batch, 1, 160, 1024]. The batch axis is dynamic; height and width are static.
  • Aspect-preserving resize to height 160, then pad the width to 1024.
  • Normalize pixel / 127.5 - 1.0.
  • Pad with white, which is +1.0 after normalization, not 0.0.
  • Greedy CTC decode over 277 classes; index 0 is the blank.

Feeding raw uint8 in [0, 255] produces confident garbage with no error.

Load the charset from the same revision as the weights. A charset from another revision decodes every index to the wrong character and raises nothing. Strip only \n and \r: the first class is a space, and a bare .strip() removes it and shifts every index by one.

Architecture

MobileNetV3-Large + squeeze-excitation neck โ†’ band pooling โ†’ 2ร—BiLSTM(512) โ†’ bottleneck self-attention (256-dim, 4 heads) โ†’ Linear(1024 โ†’ 277) โ†’ CTC. 11,553,437 parameters, all trainable.

Moving from v2

v2 is still served at revision a51be11 and will stay there; anything pinned to it is unaffected. v3.5 is not a drop-in replacement:

v2 (a51be11) v3.5 (d3d9d5e)
Input height 128 160
Input width dynamic static 1024
Batch axis fixed at 1 dynamic
Output classes 316 277
Charset 315 characters 276 characters
Parameters 6,575,868 11,553,437

The width is the change that breaks integrations. nn.MultiheadAttention fixes the sequence length when the graph is traced, so the graph accepts a width of 1024 and nothing else.

Evaluation

Held-out test

Measured 2026-08-16 on the test split, held out from training and scored once.

CER CI95 sequence accuracy
Overall 0.0100 [0.0056, 0.0147] 85.33%
Mon (n=99) 0.0113 [0.0050, 0.0186] 85.86%
Burmese (n=32) 0.0027 [0.0000, 0.0064] 93.75%
English (n=14) 0.0103 [0.0009, 0.0226] 71.43%
Mon + English (n=5) 0.0144 [0.0000, 0.0292] 60.00%

n = 150 lines, greedy decoding. Macro-averaged CER 0.0047. Expected calibration error 0.0315. Per typeface: Pyidaungsu-Regular 0.0076 (n=56), Pyidaungsu-Numbers 0.0108 (n=51), Pyidaungsu-Bold 0.0124 (n=43).

Three baselines ran alongside it:

Baseline CER What it rules out
empty prediction 1.0000 A broken scorer. Anything other than exactly 1.0 and the suite refuses to report
most-common grapheme 1.0832 That the charset's prior alone explains the result
1-nearest-neighbour pixel retrieval over 3,000 training images 1.3675 Near-duplicates. Retrieving the closest training image scores worse than predicting nothing, so test images are not near-duplicates of training ones. It does not rule out memorisation in general

No train/serve skew measured. The PyTorch checkpoint and the published ONNX graph both scored 0.0100 (cer_delta 0.0) and produced exactly the same string on 150 of 150 lines, so the number describes the artifact in this repository.

Latency p50 134.8 ms/line on CPU, batched. It is a per-batch timing divided by batch size, so it is a mean, not a tail.

What the held-out number does not cover

  1. n = 150. The interval is [0.0056, 0.0147]; treat the width as real.
  2. One typeface. All 150 lines are Pyidaungsu, held out from training. A second held-out design, yunghkio, is not represented: this split was generated before that design was set aside.
  3. Unseen text, not an unseen renderer. The same generator and augmentation pipeline produced training and test images, so a renderer defect is learned, validated and tested against identically. This is the largest caveat on the card, and only real photographed lines close it.
  4. Disjointness is argued, not directly verified. The training-side labels were produced on a machine whose state was not fully retained. The partition function is byte-unchanged since generation and none of the 150 test labels appears in the local training or validation sets, but that is an argument from the stability of a hash, not a direct comparison.

Against v2, on identical images

Both generations ship an ONNX export, so the same rendered lines go through both graphs with the same preprocessing and greedy decode. Text is restricted to the 273 characters both charsets can emit, so neither model is charged for a character it has no class for. 600 lines per arm, rendered 2026-08-15.

Rendered in n v2 CER v3.5 CER error reduced
the 32 trained designs 600 0.1470 0.0396 73%
namkhon, held out from v3.5 600 0.0521 0.0188 64%
pyidaungsu, yunghkio, held out from v3.5 600 0.0342 0.0051 85%

v2 predates font-disjoint splits, so the bottom two rows are held out from v3.5 only, which tilts the comparison against v3.5. This is a preview, not an evaluation: one rendering pipeline reading its own output, synthetic for both models. It says v3.5 reads rendered Mon better than v2 did, nothing more.

Wide lines

Which handling wins depends on width. Measured 2026-08-22 over 201 rendered lines: squeezing a whole line into the 1024px canvas wins at 2 tiles, the two are level at 3, and cutting it into canvas-width tiles at whitespace columns wins from 4 up. On a book page at 150 dpi every line fitted one tile. Re-measure before swapping models.

Export checks

Both exports are gated against the PyTorch model on a seeded uniform [-1, 1] input, and the gate fails the build:

  • ONNX: logits within rtol=atol=1e-3.
  • Core ML: logits within 1e-3 and an identical CTC-decoded string.

The Core ML gate runs on CPU. The Neural Engine computes in fp16 and its compiler may reassociate, so what runs on a device is not what was verified.

Limitations

  • Real photographs. Every training sample is synthetic, rendered by one generator. A camera photo of physical text is out of domain and fails confidently: on a video frame of a whiteboard the model returned fluent Mon at confidence 0.83 for text that appears nowhere on the page. Confidence is not a usable filter for this.
  • Page layout. The input is a cropped line. There is no detector and no reading-order model.
  • Handwriting, except Myanmar digits.
  • Beam search. Only greedy decoding has been evaluated.
  • Latency. No per-line tail has been measured, and no on-device number exists on any platform.
  • Typeface coverage. 35 distinct designs across 79 usable font files bound every generalisation claim above.

Training data

Rendered from a Mon text corpus. The corpus licensing is not fully settled: one source has no established licence, and CC BY-SA attribution for another cannot be reconstructed because per-article URLs were stripped at import. The MIT licence on this repository covers the model weights and the code that produced them; it does not resolve the provenance of the text they were rendered from.

Citation

@misc{monocr,
  title  = {MonOCR: line-level OCR for the Mon language},
  author = {Zin Min},
  year   = {2026},
  url    = {https://huggingface.co/janakhpon/monocr}
}
Downloads last month
10
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support