nko-ocr β€” Tesseract LSTM models for N'Ko (ί’ίžί) OCR

Also spelled: N'Ko OCR Β· Nko OCR Β· NKO OCR β€” first open-source OCR for the N'Ko script (ί’ίžί), writing system of Bambara/Manding languages, Mali.

Tesseract LSTM models for N'Ko optical character recognition β€” the writing system of Manding languages (Bambara, Maninka, Dioula), spoken by ~50 million people in West Africa. First usable open-source N'Ko OCR (the official Tesseract nko.traineddata produced empty output on N'Ko text).

By Fousseyni Diarra β€” Bamako, Mali. GPLv3.

Models

File CER WER Notes
nko-v2.traineddata (recommended) 6.26% 17.85% v1 + 4,000 targeted images (diacritics, N'Ko digits, 2 fonts) + 2,856-word dictionary wordlist, 32k iterations
nko-v1.traineddata 15.84% 34.84% 914 pure-N'Ko images, 18k iterations

Measured with lstmeval on the same 92-line held-out test set (~94% of characters correct with v2).

Quickstart

# Tesseract >= 5.0, N'Ko model files in ./ (or --tessdata-dir)
tesseract page.png output.txt -l nko-v2 --tessdata-dir ./

# single text line
tesseract line.png stdout -l nko-v2 --tessdata-dir ./ --psm 6

Notes:

  • Input works best at β‰₯120 DPI (upscale lower-resolution renders first).
  • N'Ko is right-to-left; the models were trained with RTL handling.
  • Printed text only β€” handwriting is not supported. Human review required for critical documents.

Links

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support