Instructions to use Darayut/khmer-text-recognition with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Darayut/khmer-text-recognition with Transformers:
# Use a pipeline as a high-level helper # Warning: Pipeline type "image-to-text" is no longer supported in transformers v5. # You must load the model directly (see below) or downgrade to v4.x with: # 'pip install "transformers<5.0.0' from transformers import pipeline pipe = pipeline("image-to-text", model="Darayut/khmer-text-recognition", trust_remote_code=True)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("Darayut/khmer-text-recognition", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
GitHub | PyPI | Desktop app (Docker) | Datasets | Inference Space
A Squeeze-and-Excitation Transformer Network for Khmer Optical Character Recognition
Character Error Rate (CER %) on KHOB, Legal Documents and Printed Words. Lower is better.
Introduction
Netra-OCR is a 17M-parameter text-line recognizer for Khmer and English, trained from scratch on 1.3M synthetic text-line images. A Squeeze-and-Excitation network and a Transformer encoder read the image; a Transformer decoder writes the text. Khmer is decoded as grapheme clusters rather than single code points, which makes Khmer output sequences about 43% shorter than character-level decoding.
The model reads one text line at a time. To read whole pages (scans, phone photos, PDFs), use the netra-ocr package, which adds text-line detection, page layout analysis and export to Word, HTML, Markdown and JSON, or run the ready-made Netra OCR app with Docker (below).
| Parameters | 17M |
| Input | one text-line image (any width; resized to 48 px high) |
| Languages | Khmer, English |
| Output vocabulary | 2,101 tokens: Khmer grapheme clusters, Latin characters, digits, punctuation |
| Decoders | ar (autoregressive, default) and blockwise (identical output, ~1.55× faster) |
| Training data | 1.3M synthetic text lines |
| Package | pip install netra-ocr (version 1.1.0) |
Files in this repository
| File | What it is |
|---|---|
khmerocr_cluster_ar.pth |
The current model. Grapheme-cluster vocabulary, autoregressive decoder. The default for netra-ocr. |
khmerocr_cluster_blockwise.pth |
The same model with a trained blockwise-parallel decoding head. Produces exactly the same text as greedy ar, about 1.55× faster. |
model.safetensors, config.json, *_khmerocr.py, inference.py, vocab.json, tokenizer files |
An earlier character-level model (124-token vocabulary), loadable with transformers + trust_remote_code=True. Kept for existing users; new projects should use the cluster model above. |
khmerocr_epoch*.pth |
Earlier training checkpoints of the character-level model. |
The cluster model's vocabulary (char2idx_cluster.json) ships inside the netra-ocr package, which downloads the checkpoint from this repository automatically.
Usage
Recognize text lines (Python)
pip install netra-ocr
from netra_ocr.recognition import recognize, recognize_batch
# One cropped text line: a file path or a PIL.Image
text = recognize("line_crop.png")
# Many crops at once; the model stays loaded between calls
texts = recognize_batch(line_crops, batch_size=8)
# Blockwise decoder: same text, usually faster
text = recognize("line_crop.png", decoder="blockwise")
# Beam search: slower, slightly more accurate (autoregressive decoder only)
text = recognize("line_crop.png", beam_width=3)
The checkpoint is downloaded from this repository on first use and cached.
Read whole documents (Python)
pip install "netra-ocr[yolo,layout,pdf,docx]"
from netra_ocr.ocr_engine import KhmerOCRPipeline
from netra_ocr.exporters import save_document
pipeline = KhmerOCRPipeline(detector="yolo", layout=True)
doc = pipeline.process_document("letter.pdf") # images, multi-page TIFF or PDF
save_document(doc, "letter.docx") # .docx .html .md .json .txt
This detects the text lines, analyses the page layout (headings, paragraphs, tables, figures, headers and footers, reading order) with DocLayout-YOLO, and recognizes each line with this model.
Netra OCR app (no Python needed)
A web app and REST API with every model included. It runs on your own computer, on the CPU, with no internet connection needed after the download:
docker run -p 8000:8000 -v netra-data:/data ghcr.io/netra-ai-lab/netra-ocr
# then open http://localhost:8000
Upload a scan, a phone photo or a PDF, correct the recognized document block by block next to the page image, and export it to Word. Works on Windows, Linux and Mac (Intel and Apple Silicon). See the GitHub README for details.
Datasets
The model was trained entirely on synthetic data and evaluated on real-world and synthetic data.
Training data (synthetic, 1.3M text lines)
| Dataset | Lines | Source | Augmentations |
|---|---|---|---|
| khmer-document-synthetic-low-res | 100,000 | Pillow + Khmer corpus, 11 fonts | Erosion, noise, thinning/thickening, perspective distortion |
| khmer-scene-text-synthetic-contrast | 102,500 | SynthTIGER + Stanford Background | Rotation, blur, noise, realistic backgrounds |
| KhmerSynthetic1M | 1,000,000 | SoyVitou (external) | Pre-applied by the source authors |
| khmer-hanuman-100k | 100,000 | seanghay (external) | Pre-applied by the source authors |
Evaluation data
| Dataset | Type | Size | Description |
|---|---|---|---|
| KHOB | Real | 325 | Standard benchmark: clean backgrounds, compression artifacts. |
| Legal Documents | Real | 227 | Smartphone photos of official documents (birth certificates, diplomas, ID cards): varied degradation, lighting and distortion. Not public. |
| Printed Words | Synthetic | 1,000 | Short, isolated words in 10 fonts, to test short sequences. |
Methodology & Architecture
Preprocessing: chunking and merging
To handle variable-length text lines without aggressive resizing, each image is resized to a height of 48 px (keeping its aspect ratio) and split into overlapping 48×100 px chunks with a 16 px overlap.
Model architecture
The input image is resized and split into 48×100 px chunks with 16 px overlaps. Each chunk passes through the Squeeze-and-Excitation network in parallel, producing 512 feature maps of 2×32 px, which become patch embeddings with positional embeddings. The Transformer encoder turns them into vision tokens; the tokens of all chunks are merged, smoothed by a bidirectional LSTM, and passed to the Transformer decoder, which outputs the text one character cluster at a time.
Squeeze-and-Excitation network. Five VGG-style convolution blocks (64 → 512 channels) with height-only pooling in the later stages, so the horizontal axis, where character order lives, is preserved. Blocks 3, 4 and 5 add a 1D Squeeze-and-Excitation module that averages over height only, giving each horizontal column its own channel weights, so background and noise can be suppressed independently at each position along the line.
Patch module. Collapses the feature map's height and projects each column to a 384-dimensional embedding, plus a learnable positional embedding within the chunk.
Transformer encoder. Self-attention among the patches of a chunk resolves local ambiguities, such as visually similar sub-consonant stacks.
Merging module. Concatenates the vision tokens of all chunks of a line into one sequence and adds a second, global positional embedding across the whole line.
BiLSTM context smoother. A bidirectional LSTM runs over the merged sequence so information flows across chunk boundaries in both directions, smoothing the seams where a character is split between two chunks.
Transformer decoder. Generates the output sequence autoregressively, with cross-attention over the smoothed encoder output.
Tokenization
Latin text is decoded character by character. Khmer text is decoded as Khmer Character Clusters (a consonant with its subscripts and vowels, e.g. ក្រ), with any cluster outside the vocabulary falling back to single characters. On a 35,000-line Khmer corpus this shortens sequences by 42.65% with lossless round-tripping, which also means fewer decoding steps per line.
Blockwise-parallel decoding
khmerocr_cluster_blockwise.pth adds a small proposal head (Stern et al. 2018) on top of the frozen autoregressive model. Each step it proposes several tokens ahead and keeps only those that the base model itself would have produced, so the output is provably identical to greedy ar decoding, in fewer steps.
Training
Adam (learning rate 1e-4) with a staged cyclic learning-rate schedule, cross-entropy loss, and 50,000 randomly sampled, augmented lines per epoch. The cluster-vocabulary model was warm-started from the earlier character-level model (vision layers and shared token embeddings transferred), then trained on the full 1.3M-line set.
Evaluation
| Model | KHOB | Legal Documents | Printed Words |
|---|---|---|---|
| Tesseract | 5.1 | 41.8 | 8.0 |
| Surya | 14.5 | 51.2 | 28.4 |
| Qwen2.5-VL (3B) † | 89.2 | 84.6 | 98.6 |
| DeepSeek-OCR (3B) † | 33.2 | 44.2 | 69.8 |
| Netra-OCR (ours, 17M) | 1.0 | 5.3 | 3.0 |
Character Error Rate (%), lower is better. † Qwen2.5-VL and DeepSeek-OCR were fine-tuned with unsloth on the same training set as Netra-OCR before evaluation.
Decoders
The table above is the benchmark evaluation. This one compares the two decoders of the released checkpoints with each other, under identical settings: the 325 KHOB lines, greedy decoding, an NVIDIA RTX 3060.
| Decoder | CER | WER | Exact match | ms / line |
|---|---|---|---|---|
ar (greedy) |
1.17% | 17.11% | 84.92% | 17.6 |
blockwise |
1.17% | 17.11% | 84.92% | 11.4 |
The blockwise decoder accepts on average 2.75 tokens per step (out of 4) and is 1.55× faster end to end, with identical output. Beam search (beam_width=3) lowers KHOB CER to about 0.9% at a higher cost per line.
Examples
Limitations
- The model recognizes single text lines. Whole pages need a line detector first; the
netra-ocrpackage includes one. - It was trained on printed text; handwriting was not part of the training data.
- Only Khmer and English are supported.
License
The code and model weights are released under the MIT license. Note that part of the training data, KhmerSynthetic1M, is licensed by its authors for research and academic use only; check its terms before using this model commercially.
References
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, et al. ICLR 2021. arXiv:2010.11929
TrOCR: Transformer-based Optical Character Recognition with Pre-trained Models Minghao Li, Tengchao Lv, Lei Cui, Yijuan Lu, Dinei Florencio, Cha Zhang, Zhoujun Li, Furu Wei. AAAI 2023. arXiv:2109.10282
Toward a Low-Resource Non-Latin-Complete Baseline: An Exploration of Khmer Optical Character Recognition R. Buoy, M. Iwamura, S. Srun and K. Kise. IEEE Access, vol. 11, pp. 128044–128060, 2023. DOI: 10.1109/ACCESS.2023.3332361
Squeeze-and-Excitation Networks Jie Hu, Li Shen, and Gang Sun. CVPR 2018. arXiv:1709.01507
Bidirectional Recurrent Neural Networks Mike Schuster and Kuldip K. Paliwal. IEEE Transactions on Signal Processing, 1997. DOI: 10.1109/78.650093
Blockwise Parallel Decoding for Deep Autoregressive Models Mitchell Stern, Noam Shazeer, Jakob Uszkoreit. NeurIPS 2018. arXiv:1811.03115
DocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception Zhiyuan Zhao, Hengrui Kang, Bin Wang, Conghui He.
Balraj98. (2018). Stanford background dataset [Data set]. Kaggle. https://www.kaggle.com/datasets/balraj98/stanford-background-dataset
EKYC Solutions. (2022). Khmer OCR benchmark dataset (KHOB) [Data set]. GitHub. https://github.com/EKYCSolutions/khmer-ocr-benchmark-dataset
Em, H., Valy, D., Gosselin, B., & Kong, P. (2024). Khmer text recognition dataset [Data set]. Kaggle. https://www.kaggle.com/datasets/emhengly/khmer-text-recognition-dataset
- Downloads last month
- 401



