Netra Lab

GitHub | PyPI | Desktop app (Docker) | Datasets | Inference Space

A Squeeze-and-Excitation Transformer Network for Khmer Optical Character Recognition

Character Error Rate (CER %) on KHOB, Legal Documents and Printed Words. Lower is better.

Introduction

Netra-OCR is a 17M-parameter text-line recognizer for Khmer and English, trained from scratch on 1.3M synthetic text-line images. A Squeeze-and-Excitation network and a Transformer encoder read the image; a Transformer decoder writes the text. Khmer is decoded as grapheme clusters rather than single code points, which makes Khmer output sequences about 43% shorter than character-level decoding.

The model reads one text line at a time. To read whole pages (scans, phone photos, PDFs), use the netra-ocr package, which adds text-line detection, page layout analysis and export to Word, HTML, Markdown and JSON, or run the ready-made Netra OCR app with Docker (below).

Parameters 17M
Input one text-line image (any width; resized to 48 px high)
Languages Khmer, English
Output vocabulary 2,101 tokens: Khmer grapheme clusters, Latin characters, digits, punctuation
Decoders ar (autoregressive, default) and blockwise (identical output, ~1.55× faster)
Training data 1.3M synthetic text lines
Package pip install netra-ocr (version 1.1.0)

Files in this repository

File What it is
khmerocr_cluster_ar.pth The current model. Grapheme-cluster vocabulary, autoregressive decoder. The default for netra-ocr.
khmerocr_cluster_blockwise.pth The same model with a trained blockwise-parallel decoding head. Produces exactly the same text as greedy ar, about 1.55× faster.
model.safetensors, config.json, *_khmerocr.py, inference.py, vocab.json, tokenizer files An earlier character-level model (124-token vocabulary), loadable with transformers + trust_remote_code=True. Kept for existing users; new projects should use the cluster model above.
khmerocr_epoch*.pth Earlier training checkpoints of the character-level model.

The cluster model's vocabulary (char2idx_cluster.json) ships inside the netra-ocr package, which downloads the checkpoint from this repository automatically.


Usage

Recognize text lines (Python)

pip install netra-ocr
from netra_ocr.recognition import recognize, recognize_batch

# One cropped text line: a file path or a PIL.Image
text = recognize("line_crop.png")

# Many crops at once; the model stays loaded between calls
texts = recognize_batch(line_crops, batch_size=8)

# Blockwise decoder: same text, usually faster
text = recognize("line_crop.png", decoder="blockwise")

# Beam search: slower, slightly more accurate (autoregressive decoder only)
text = recognize("line_crop.png", beam_width=3)

The checkpoint is downloaded from this repository on first use and cached.

Read whole documents (Python)

pip install "netra-ocr[yolo,layout,pdf,docx]"
from netra_ocr.ocr_engine import KhmerOCRPipeline
from netra_ocr.exporters import save_document

pipeline = KhmerOCRPipeline(detector="yolo", layout=True)
doc = pipeline.process_document("letter.pdf")     # images, multi-page TIFF or PDF
save_document(doc, "letter.docx")                 # .docx .html .md .json .txt

This detects the text lines, analyses the page layout (headings, paragraphs, tables, figures, headers and footers, reading order) with DocLayout-YOLO, and recognizes each line with this model.

Netra OCR app (no Python needed)

A web app and REST API with every model included. It runs on your own computer, on the CPU, with no internet connection needed after the download:

docker run -p 8000:8000 -v netra-data:/data ghcr.io/netra-ai-lab/netra-ocr
# then open http://localhost:8000

Upload a scan, a phone photo or a PDF, correct the recognized document block by block next to the page image, and export it to Word. Works on Windows, Linux and Mac (Intel and Apple Silicon). See the GitHub README for details.


Datasets

The model was trained entirely on synthetic data and evaluated on real-world and synthetic data.

Training data (synthetic, 1.3M text lines)

Dataset Lines Source Augmentations
khmer-document-synthetic-low-res 100,000 Pillow + Khmer corpus, 11 fonts Erosion, noise, thinning/thickening, perspective distortion
khmer-scene-text-synthetic-contrast 102,500 SynthTIGER + Stanford Background Rotation, blur, noise, realistic backgrounds
KhmerSynthetic1M 1,000,000 SoyVitou (external) Pre-applied by the source authors
khmer-hanuman-100k 100,000 seanghay (external) Pre-applied by the source authors

Evaluation data

Dataset Type Size Description
KHOB Real 325 Standard benchmark: clean backgrounds, compression artifacts.
Legal Documents Real 227 Smartphone photos of official documents (birth certificates, diplomas, ID cards): varied degradation, lighting and distortion. Not public.
Printed Words Synthetic 1,000 Short, isolated words in 10 fonts, to test short sequences.

Dataset Overview


Methodology & Architecture

Preprocessing: chunking and merging

To handle variable-length text lines without aggressive resizing, each image is resized to a height of 48 px (keeping its aspect ratio) and split into overlapping 48×100 px chunks with a 16 px overlap.

Model architecture

Model Architecture

The input image is resized and split into 48×100 px chunks with 16 px overlaps. Each chunk passes through the Squeeze-and-Excitation network in parallel, producing 512 feature maps of 2×32 px, which become patch embeddings with positional embeddings. The Transformer encoder turns them into vision tokens; the tokens of all chunks are merged, smoothed by a bidirectional LSTM, and passed to the Transformer decoder, which outputs the text one character cluster at a time.

  1. Squeeze-and-Excitation network. Five VGG-style convolution blocks (64 → 512 channels) with height-only pooling in the later stages, so the horizontal axis, where character order lives, is preserved. Blocks 3, 4 and 5 add a 1D Squeeze-and-Excitation module that averages over height only, giving each horizontal column its own channel weights, so background and noise can be suppressed independently at each position along the line.

    SE Module

  2. Patch module. Collapses the feature map's height and projects each column to a 384-dimensional embedding, plus a learnable positional embedding within the chunk.

  3. Transformer encoder. Self-attention among the patches of a chunk resolves local ambiguities, such as visually similar sub-consonant stacks.

  4. Merging module. Concatenates the vision tokens of all chunks of a line into one sequence and adds a second, global positional embedding across the whole line.

  5. BiLSTM context smoother. A bidirectional LSTM runs over the merged sequence so information flows across chunk boundaries in both directions, smoothing the seams where a character is split between two chunks.

    Context Smoothing Module

  6. Transformer decoder. Generates the output sequence autoregressively, with cross-attention over the smoothed encoder output.

Tokenization

Latin text is decoded character by character. Khmer text is decoded as Khmer Character Clusters (a consonant with its subscripts and vowels, e.g. ក្រ), with any cluster outside the vocabulary falling back to single characters. On a 35,000-line Khmer corpus this shortens sequences by 42.65% with lossless round-tripping, which also means fewer decoding steps per line.

Blockwise-parallel decoding

khmerocr_cluster_blockwise.pth adds a small proposal head (Stern et al. 2018) on top of the frozen autoregressive model. Each step it proposes several tokens ahead and keeps only those that the base model itself would have produced, so the output is provably identical to greedy ar decoding, in fewer steps.

Training

Adam (learning rate 1e-4) with a staged cyclic learning-rate schedule, cross-entropy loss, and 50,000 randomly sampled, augmented lines per epoch. The cluster-vocabulary model was warm-started from the earlier character-level model (vision layers and shared token embeddings transferred), then trained on the full 1.3M-line set.


Evaluation

Model KHOB Legal Documents Printed Words
Tesseract 5.1 41.8 8.0
Surya 14.5 51.2 28.4
Qwen2.5-VL (3B) † 89.2 84.6 98.6
DeepSeek-OCR (3B) † 33.2 44.2 69.8
Netra-OCR (ours, 17M) 1.0 5.3 3.0

Character Error Rate (%), lower is better. † Qwen2.5-VL and DeepSeek-OCR were fine-tuned with unsloth on the same training set as Netra-OCR before evaluation.

Decoders

The table above is the benchmark evaluation. This one compares the two decoders of the released checkpoints with each other, under identical settings: the 325 KHOB lines, greedy decoding, an NVIDIA RTX 3060.

Decoder CER WER Exact match ms / line
ar (greedy) 1.17% 17.11% 84.92% 17.6
blockwise 1.17% 17.11% 84.92% 11.4

The blockwise decoder accepts on average 2.75 tokens per step (out of 4) and is 1.55× faster end to end, with identical output. Beam search (beam_width=3) lowers KHOB CER to about 0.9% at a higher cost per line.

Examples


Limitations

  • The model recognizes single text lines. Whole pages need a line detector first; the netra-ocr package includes one.
  • It was trained on printed text; handwriting was not part of the training data.
  • Only Khmer and English are supported.

License

The code and model weights are released under the MIT license. Note that part of the training data, KhmerSynthetic1M, is licensed by its authors for research and academic use only; check its terms before using this model commercially.


References

  1. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, et al. ICLR 2021. arXiv:2010.11929

  2. TrOCR: Transformer-based Optical Character Recognition with Pre-trained Models Minghao Li, Tengchao Lv, Lei Cui, Yijuan Lu, Dinei Florencio, Cha Zhang, Zhoujun Li, Furu Wei. AAAI 2023. arXiv:2109.10282

  3. Toward a Low-Resource Non-Latin-Complete Baseline: An Exploration of Khmer Optical Character Recognition R. Buoy, M. Iwamura, S. Srun and K. Kise. IEEE Access, vol. 11, pp. 128044–128060, 2023. DOI: 10.1109/ACCESS.2023.3332361

  4. Squeeze-and-Excitation Networks Jie Hu, Li Shen, and Gang Sun. CVPR 2018. arXiv:1709.01507

  5. Bidirectional Recurrent Neural Networks Mike Schuster and Kuldip K. Paliwal. IEEE Transactions on Signal Processing, 1997. DOI: 10.1109/78.650093

  6. Blockwise Parallel Decoding for Deep Autoregressive Models Mitchell Stern, Noam Shazeer, Jakob Uszkoreit. NeurIPS 2018. arXiv:1811.03115

  7. DocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception Zhiyuan Zhao, Hengrui Kang, Bin Wang, Conghui He.

    1. arXiv:2410.12628
  8. Balraj98. (2018). Stanford background dataset [Data set]. Kaggle. https://www.kaggle.com/datasets/balraj98/stanford-background-dataset

  9. EKYC Solutions. (2022). Khmer OCR benchmark dataset (KHOB) [Data set]. GitHub. https://github.com/EKYCSolutions/khmer-ocr-benchmark-dataset

  10. Em, H., Valy, D., Gosselin, B., & Kong, P. (2024). Khmer text recognition dataset [Data set]. Kaggle. https://www.kaggle.com/datasets/emhengly/khmer-text-recognition-dataset

Downloads last month
401
Safetensors
Model size
17.6M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train Darayut/khmer-text-recognition

Space using Darayut/khmer-text-recognition 1

Papers for Darayut/khmer-text-recognition