Devika: Malayalam Handwritten Word OCR

Devika is an open-source model for recognizing single handwritten Malayalam words from images. It is a fine-tuned version of Microsoft's trocr-large-handwritten with a new Malayalam character-level decoder vocabulary.

Scope: Devika is trained on word-level images. It reads one cropped word at a time. It is not a line, paragraph or full-page reader. See Limitations and Using Devika on full documents.

Results

Evaluated on the held-out test split (19,635 word images per the dataset description), which was never used for training or checkpoint selection.

Metric Value
Character Error Rate (CER) 1.67%
Character accuracy (1 − CER) 98.33%
Word accuracy (exact match) 89.14%

Word accuracy counts a word as correct only if every character matches, so a single wrong vowel sign or conjunct component fails the whole word. This is why it is much lower than character accuracy.

Checkpoint selection. This release is the checkpoint at step 39,000 (about epoch 29), chosen by lowest validation CER:

Checkpoint Validation CER
step 39,000 (released) 1.67%
step 40,000 2.00%
step 41,000 1.76%

Validation CER, tracked every 1,000 steps, plateaued in the 1.7–2.1% range in the later epochs.

Model details

Architecture VisionEncoderDecoderModel: ViT-Large image encoder (24 layers, hidden size 1024, 384×384 input) + 12-layer TrOCR decoder
Base model microsoft/trocr-large-handwritten
Tokenizer Custom character-level tokenizer (95 symbols + 4 special tokens; vocabulary size 99), built from the training data
Max output length 32 tokens (word level)
Language Malayalam (ml)
Input An image of one handwritten word
Output The Malayalam text of that word

The encoder and decoder transformer layers are initialized from the English TrOCR checkpoint. The decoder's embedding and output layers were resized and re-learned for the Malayalam character vocabulary, since the original English BPE vocabulary cannot represent Malayalam.

Training data

Trained on the Mozhi Malayalam handwritten word-level dataset:

Split Word images
Train 80,146
Validation 9893
Test 9980

The training split comes with a vocabulary of 13,401 unique words. Ground truth is one transcription per cropped word image.

dataset link

Training procedure

  • Method: full fine-tuning (all encoder and decoder weights updated), not LoRA
  • Hardware: 4× NVIDIA A10G (24 GB), AWS EC2 g5.12xlarge, multi-GPU data parallel via 🤗 accelerate
  • Per-GPU batch size: 16 (effective batch size 64)
  • Learning rate: 3e-5, 10% warmup, linear decay
  • Precision: fp16
  • Label handling: an explicit end-of-sequence token is appended to every label so the model learns when to stop generating
  • Checkpointing: evaluated every 1,000 steps on validation CER, and the best checkpoint by CER was kept
  • Decoding used for the reported results: beam search, 5 beams, max length 32

How to use

import torch
from PIL import Image
from transformers import TrOCRProcessor, VisionEncoderDecoderModel

repo_id = "Artificial-Soul/Devika"
device = "cuda" if torch.cuda.is_available() else "cpu"

processor = TrOCRProcessor.from_pretrained(repo_id)
model = VisionEncoderDecoderModel.from_pretrained(repo_id).to(device).eval()

image = Image.open("word.jpg").convert("RGB")   # a crop containing ONE word
pixel_values = processor(images=image, return_tensors="pt").pixel_values.to(device)

with torch.no_grad():
    generated_ids = model.generate(pixel_values, max_length=32, num_beams=5)

print(processor.batch_decode(generated_ids, skip_special_tokens=True)[0])

You may see a warning that encoder.pooler.* weights are newly initialized. This is harmless: TrOCR never uses that layer.

Limitations

  • Single words only. The model was trained exclusively on isolated word crops. Feeding it a sentence, line or paragraph image produces poor output, because the image encoder squeezes the whole image to a fixed 384×384 input and the model has never seen multi-word images.
  • Very wide word crops are compressed horizontally by the fixed-size resize, which can hurt accuracy on fine conjunct details.
  • Evaluation is in-distribution. The reported numbers come from the Bhashini test split. Accuracy on other handwriting styles, low-quality camera photos, unusual pens or backgrounds, or words very different from the training vocabulary (for example rare proper nouns) has not been measured and may be lower.
  • Not for unsupervised use on high-stakes documents. Even at 98.3% character accuracy, errors occur. For legal, official or financial records, keep a human review step and use confidence-based flagging.

Using Devika on full documents

To read a full page, run a text detector first, split the page into lines, split each line into words, and pass each word crop to Devika. Join the words with spaces and the lines with newlines. Text detectors such as PaddleOCR or docTR can provide the line boxes.

Intended use

Malayalam handwritten word recognition for research, education, digitization pipelines (as the recognition stage after word segmentation), and as a base for further fine-tuning on domain-specific handwriting.

Licensing note

The model weights are released under the MIT license, matching the base TrOCR model. Please also respect the terms of the Bhashini dataset used for fine-tuning.

Acknowledgements

  • Microsoft Research for TrOCR and the pretrained trocr-large-handwritten checkpoint. This model builds directly on their work.
  • Bhashini for the Malayalam handwritten word dataset.
  • Compute: training was run on AWS EC2 g5.12xlarge (4× NVIDIA A10G), provided by C-DIT.
  • Hugging Face 🤗 transformers and accelerate for the training and model tooling.

Citation

If you use this model, please also cite the original TrOCR paper:

@inproceedings{li2023trocr,
  title     = {TrOCR: Transformer-based Optical Character Recognition with Pre-trained Models},
  author    = {Li, Minghao and Lv, Tengchao and Chen, Jingye and Cui, Lei and Lu, Yijuan and Florencio, Dinei and Zhang, Cha and Li, Zhoujun and Wei, Furu},
  booktitle = {Proceedings of the AAAI Conference on Artificial Intelligence},
  year      = {2023},
  url       = {https://arxiv.org/abs/2109.10282}
}
Downloads last month
17
Safetensors
Model size
0.5B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Artificial-Soul/Devika

Finetuned
(18)
this model

Paper for Artificial-Soul/Devika

Evaluation results

  • Character Error Rate on Bhashini Malayalam Handwritten Word Dataset (test split)
    self-reported
    0.017
  • Word accuracy (exact match) on Bhashini Malayalam Handwritten Word Dataset (test split)
    self-reported
    0.891