Instructions to use Artificial-Soul/Devika with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Artificial-Soul/Devika with Transformers:
# Use a pipeline as a high-level helper # Warning: Pipeline type "image-to-text" is no longer supported in transformers v5. # You must load the model directly (see below) or downgrade to v4.x with: # 'pip install "transformers<5.0.0' from transformers import pipeline pipe = pipeline("image-to-text", model="Artificial-Soul/Devika")# Load model directly from transformers import AutoTokenizer, AutoModelForMultimodalLM tokenizer = AutoTokenizer.from_pretrained("Artificial-Soul/Devika") model = AutoModelForMultimodalLM.from_pretrained("Artificial-Soul/Devika", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Devika: Malayalam Handwritten Word OCR
Devika is an open-source model for recognizing single handwritten Malayalam words from images. It is a fine-tuned version of Microsoft's trocr-large-handwritten with a new Malayalam character-level decoder vocabulary.
Scope: Devika is trained on word-level images. It reads one cropped word at a time. It is not a line, paragraph or full-page reader. See Limitations and Using Devika on full documents.
Results
Evaluated on the held-out test split (19,635 word images per the dataset description), which was never used for training or checkpoint selection.
| Metric | Value |
|---|---|
| Character Error Rate (CER) | 1.67% |
| Character accuracy (1 − CER) | 98.33% |
| Word accuracy (exact match) | 89.14% |
Word accuracy counts a word as correct only if every character matches, so a single wrong vowel sign or conjunct component fails the whole word. This is why it is much lower than character accuracy.
Checkpoint selection. This release is the checkpoint at step 39,000 (about epoch 29), chosen by lowest validation CER:
| Checkpoint | Validation CER |
|---|---|
| step 39,000 (released) | 1.67% |
| step 40,000 | 2.00% |
| step 41,000 | 1.76% |
Validation CER, tracked every 1,000 steps, plateaued in the 1.7–2.1% range in the later epochs.
Model details
| Architecture | VisionEncoderDecoderModel: ViT-Large image encoder (24 layers, hidden size 1024, 384×384 input) + 12-layer TrOCR decoder |
| Base model | microsoft/trocr-large-handwritten |
| Tokenizer | Custom character-level tokenizer (95 symbols + 4 special tokens; vocabulary size 99), built from the training data |
| Max output length | 32 tokens (word level) |
| Language | Malayalam (ml) |
| Input | An image of one handwritten word |
| Output | The Malayalam text of that word |
The encoder and decoder transformer layers are initialized from the English TrOCR checkpoint. The decoder's embedding and output layers were resized and re-learned for the Malayalam character vocabulary, since the original English BPE vocabulary cannot represent Malayalam.
Training data
Trained on the Mozhi Malayalam handwritten word-level dataset:
| Split | Word images |
|---|---|
| Train | 80,146 |
| Validation | 9893 |
| Test | 9980 |
The training split comes with a vocabulary of 13,401 unique words. Ground truth is one transcription per cropped word image.
Training procedure
- Method: full fine-tuning (all encoder and decoder weights updated), not LoRA
- Hardware: 4× NVIDIA A10G (24 GB), AWS EC2
g5.12xlarge, multi-GPU data parallel via 🤗accelerate - Per-GPU batch size: 16 (effective batch size 64)
- Learning rate: 3e-5, 10% warmup, linear decay
- Precision: fp16
- Label handling: an explicit end-of-sequence token is appended to every label so the model learns when to stop generating
- Checkpointing: evaluated every 1,000 steps on validation CER, and the best checkpoint by CER was kept
- Decoding used for the reported results: beam search, 5 beams, max length 32
How to use
import torch
from PIL import Image
from transformers import TrOCRProcessor, VisionEncoderDecoderModel
repo_id = "Artificial-Soul/Devika"
device = "cuda" if torch.cuda.is_available() else "cpu"
processor = TrOCRProcessor.from_pretrained(repo_id)
model = VisionEncoderDecoderModel.from_pretrained(repo_id).to(device).eval()
image = Image.open("word.jpg").convert("RGB") # a crop containing ONE word
pixel_values = processor(images=image, return_tensors="pt").pixel_values.to(device)
with torch.no_grad():
generated_ids = model.generate(pixel_values, max_length=32, num_beams=5)
print(processor.batch_decode(generated_ids, skip_special_tokens=True)[0])
You may see a warning that encoder.pooler.* weights are newly initialized. This is harmless: TrOCR never uses that layer.
Limitations
- Single words only. The model was trained exclusively on isolated word crops. Feeding it a sentence, line or paragraph image produces poor output, because the image encoder squeezes the whole image to a fixed 384×384 input and the model has never seen multi-word images.
- Very wide word crops are compressed horizontally by the fixed-size resize, which can hurt accuracy on fine conjunct details.
- Evaluation is in-distribution. The reported numbers come from the Bhashini test split. Accuracy on other handwriting styles, low-quality camera photos, unusual pens or backgrounds, or words very different from the training vocabulary (for example rare proper nouns) has not been measured and may be lower.
- Not for unsupervised use on high-stakes documents. Even at 98.3% character accuracy, errors occur. For legal, official or financial records, keep a human review step and use confidence-based flagging.
Using Devika on full documents
To read a full page, run a text detector first, split the page into lines, split each line into words, and pass each word crop to Devika. Join the words with spaces and the lines with newlines. Text detectors such as PaddleOCR or docTR can provide the line boxes.
Intended use
Malayalam handwritten word recognition for research, education, digitization pipelines (as the recognition stage after word segmentation), and as a base for further fine-tuning on domain-specific handwriting.
Licensing note
The model weights are released under the MIT license, matching the base TrOCR model. Please also respect the terms of the Bhashini dataset used for fine-tuning.
Acknowledgements
- Microsoft Research for TrOCR and the pretrained
trocr-large-handwrittencheckpoint. This model builds directly on their work. - Bhashini for the Malayalam handwritten word dataset.
- Compute: training was run on AWS EC2
g5.12xlarge(4× NVIDIA A10G), provided by C-DIT. - Hugging Face 🤗
transformersandacceleratefor the training and model tooling.
Citation
If you use this model, please also cite the original TrOCR paper:
@inproceedings{li2023trocr,
title = {TrOCR: Transformer-based Optical Character Recognition with Pre-trained Models},
author = {Li, Minghao and Lv, Tengchao and Chen, Jingye and Cui, Lei and Lu, Yijuan and Florencio, Dinei and Zhang, Cha and Li, Zhoujun and Wei, Furu},
booktitle = {Proceedings of the AAAI Conference on Artificial Intelligence},
year = {2023},
url = {https://arxiv.org/abs/2109.10282}
}
- Downloads last month
- 17
Model tree for Artificial-Soul/Devika
Base model
microsoft/trocr-large-handwrittenPaper for Artificial-Soul/Devika
Evaluation results
- Character Error Rate on Bhashini Malayalam Handwritten Word Dataset (test split)self-reported0.017
- Word accuracy (exact match) on Bhashini Malayalam Handwritten Word Dataset (test split)self-reported0.891