nko-ocr β Tesseract LSTM models for N'Ko (ίίί) OCR
Also spelled: N'Ko OCR Β· Nko OCR Β· NKO OCR β first open-source OCR for the N'Ko script (ίίί), writing system of Bambara/Manding languages, Mali.
Tesseract LSTM models for N'Ko optical character recognition β the writing system of Manding languages (Bambara, Maninka, Dioula), spoken by ~50 million people in West Africa. First usable open-source N'Ko OCR (the official Tesseract nko.traineddata produced empty output on N'Ko text).
By Fousseyni Diarra β Bamako, Mali. GPLv3.
Models
| File | CER | WER | Notes |
|---|---|---|---|
nko-v2.traineddata (recommended) |
6.26% | 17.85% | v1 + 4,000 targeted images (diacritics, N'Ko digits, 2 fonts) + 2,856-word dictionary wordlist, 32k iterations |
nko-v1.traineddata |
15.84% | 34.84% | 914 pure-N'Ko images, 18k iterations |
Measured with lstmeval on the same 92-line held-out test set (~94% of characters correct with v2).
Quickstart
# Tesseract >= 5.0, N'Ko model files in ./ (or --tessdata-dir)
tesseract page.png output.txt -l nko-v2 --tessdata-dir ./
# single text line
tesseract line.png stdout -l nko-v2 --tessdata-dir ./ --psm 6
Notes:
- Input works best at β₯120 DPI (upscale lower-resolution renders first).
- N'Ko is right-to-left; the models were trained with RTL handling.
- Printed text only β handwriting is not supported. Human review required for critical documents.
Links
- GitHub (full project: corpus, wordlist, training notes): https://github.com/koussedia/nko-ocr
- Live demo: https://huggingface.co/spaces/koussedia/nko-ocr-demo
- Dataset (4,422 image/text pairs): https://huggingface.co/datasets/koussedia/nko-ocr-dataset