Instructions to use mahmoudalrefaey/AraSpellX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use mahmoudalrefaey/AraSpellX with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("token-classification", model="mahmoudalrefaey/AraSpellX")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForTokenClassification tokenizer = AutoTokenizer.from_pretrained("mahmoudalrefaey/AraSpellX") model = AutoModelForTokenClassification.from_pretrained("mahmoudalrefaey/AraSpellX", device_map="auto") - Notebooks
- Google Colab
- Kaggle
AraSpellX
Arabic spelling and OCR error correction with a small character-level transformer trained from scratch.
AraSpellX reads Arabic text, typed by people or produced by OCR, and predicts for every character whether to keep it, delete it, replace it or insert something after it. The AraSpellX package turns these predictions into the corrected text and a list of corrections, each with its position in the input, a calibrated confidence and an error category, so that a pipeline can apply only confident edits, show suggestions or flag words.
| Task | Arabic spelling and OCR error correction, as character-level edit tagging |
| Architecture | BERT encoder: 8 layers, hidden size 384, 6 heads, feed-forward 1536, relative positions (relative_key), 512 positions |
| Parameters | 15.1M |
| Input โ output | characters (210-token vocabulary) โ one of 178 edit labels per character |
| Training | from scratch: masked-character pretraining, then correction training |
| Language | Arabic: modern standard and classical text are corrected; dialect and diacritized text are left alone |
| Model class | standard BertForTokenClassification (loads in plain transformers) |
| License | MIT |
| Code | github.com/mahmoudalrefaey/AraSpellX (branch v1) |
Example
Output of this model for a typed sentence:
ุฐูุจุช ุงูู ุงูุฌุงู
ุนู ุตุจุงุญุง ููู ุงุญุถุฑ ุงูู
ุญุงุถุฑู ุงูุงูููุ ุซู
ูุงุจูุช ุตุฏููู ูู ุงูู
ูุชุจู.
ุฐูุจุช ุฅูู ุงูุฌุงู
ุนุฉ ุตุจุงุญุง ููู ุฃุญุถุฑ ุงูู
ุญุงุถุฑุฉ ุงูุฃูููุ ุซู
ูุงุจูุช ุตุฏููู ูู ุงูู
ูุชุจุฉ.
ุงูู -> ุฅูู 98% mixed
ุงูุฌุงู
ุนู -> ุงูุฌุงู
ุนุฉ 100% ta_marbuta
ุงุญุถุฑ -> ุฃุญุถุฑ 94% hamza_alef
ุงูู
ุญุงุถุฑู -> ุงูู
ุญุงุถุฑุฉ 100% ta_marbuta
ุงูุงูููุ -> ุงูุฃูููุ 99% mixed
ุงูู
ูุชุจู. -> ุงูู
ูุชุจุฉ. 100% ta_marbuta
On OCR output it corrects what it is sure about and leaves the rest: in ุชุนุชุจุฑ ุงูู
ุฏููู ู
ู ุฃูู
ุงูู
ุฑุงูุฒ ุงููุฌุงุฑูุฉ ูู ุงูู
ูุทูุฉุ ุญูุซ ููุตุฏูุง ุงูุฒูุงุฑ ู
ู ุญู
ูุน ุฃูุญุงุก ุงูุนุงูู
. it fixes ุงูู
ุฏููู โ ุงูู
ุฏููุฉ and ุญู
ูุน โ ุฌู
ูุน and leaves the garbled ุงููุฌุงุฑูุฉ and ุงูู
ูุทูุฉ as they are.
Intended use
- Search and retrieval-augmented generation: clean Arabic text before indexing it, so that misspelled words match their correct forms.
- OCR post-processing of printed Arabic: fix common OCR confusions without risking the text that was read correctly.
- Text review: suggest or flag likely spelling errors using each correction's confidence and position.
Out of scope: grammar (agreement, case endings), punctuation and style, converting dialect to standard spelling, adding or fixing diacritics, handwriting and layout. The target spelling is standard modern Arabic orthography, as in edited publications and Arabic Wikipedia.
How to use
With the AraSpellX package (recommended)
The package normalizes the input, reads long text in overlapping windows, applies the confidence threshold and the editing rules, maps every correction back to the input, and calibrates the confidences.
git clone -b v1 https://github.com/mahmoudalrefaey/AraSpellX.git
cd AraSpellX
uv sync
from araspellx.correct.corrector import Corrector
corrector = Corrector("mahmoudalrefaey/AraSpellX") # downloads this repository once
result = corrector.correct("ุฐูุจุช ุงูู ุงูุฌุงู
ุนู", source="typed") # source="ocr" for OCR output
print(result.text) # ุฐูุจุช ุฅูู ุงูุฌุงู
ุนุฉ
for c in result.corrections:
print(c.start, c.end, c.original, c.replacement, round(c.confidence, 2), c.category)
A local web page for trying the model (right-to-left, nothing leaves the computer):
python -m araspellx.correct.demo --model mahmoudalrefaey/AraSpellX
With transformers only
The model is a standard BertForTokenClassification and predicts one label per character:
import torch
from transformers import AutoModelForTokenClassification, AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("mahmoudalrefaey/AraSpellX")
model = AutoModelForTokenClassification.from_pretrained("mahmoudalrefaey/AraSpellX").eval()
text = "ุฐูุจุช ุงูู ุงูุฌุงู
ุนู"
inputs = tokenizer(text, return_tensors="pt") # one token per character, plus [CLS] and [SEP]
with torch.no_grad():
probs = model(**inputs).logits.softmax(-1)[0]
for char, p in zip(text, probs[1:-1]):
label = model.config.id2label[p.argmax().item()]
if label != "K":
print(char, label, round(p.max().item(), 2)) # ุง R:ุฅ 0.96 ยท ู R:ู 0.95 ยท ู R:ุฉ 0.99
Labels: K keep, D delete, R:x replace with x, and โฆ+s insert s after the character; the [CLS] position carries insertions before the first character. This raw output is not yet a corrected text: use the package, or reproduce its normalization, windowing, threshold and editing rules.
Outputs of the package
| Field | Meaning |
|---|---|
Result.text |
the corrected text (normalized input with the corrections applied) |
Correction.start, Correction.end |
span of the original word in the input |
Correction.original, Correction.replacement |
the word before and after |
Correction.confidence |
calibrated probability that the correction is right |
Correction.category |
hamza_alef, hamza_seat, ta_marbuta, alef_maqsura, alef_fariqa, dots, spacing, typo or mixed |
How it works
- Normalization: look-alike Persian and Urdu letters are folded into Arabic ones, presentation-form ligatures expanded, letters stored as a base letter plus a combining hamza or madda composed, and tatweel and invisible characters removed. Every character keeps a link to its span in the input.
- Edit labels: one label per character, from the 178 most frequent edits of the training data (99.5% of all edits). Correct text stays correct by default, and every change is an explicit decision with a probability.
- Editing rules: only Arabic letters, diacritics, tatweel and spaces are edited. Digits, Latin text, symbols and existing punctuation are never changed; a letter that OCR made of a punctuation mark may be turned back into it. Spacing next to punctuation follows Arabic typography.
- Threshold and calibration: an edit is applied only above a confidence of 0.9, chosen on development data to keep damage on clean text within 0.05%.
calibration.jsonmaps raw confidences to observed accuracy, with one fit for typed and one for OCR input.
Training data
All training and test data are public. Training text is filtered against every test set (any paragraph sharing an 8-word passage with a test text is dropped), and test pages are held out by a hash of their page id.
| Source | Role | Size | License |
|---|---|---|---|
| Arabic, Egyptian and Moroccan Wikipedia, Arabic Wikisource (2026-10-01 dumps) | pretraining text; clean and noisy correction text | 2.09B characters | CC BY-SA |
| Typed-error noise generated on the fly | dropped hamza, ุฉ/ู, ู/ู, hamza seats, ุธ/ุถ, dropped alef after waw, keyboard typos, merged and split words | โ | โ |
| Wikipedia (70%) and Wikisource (30%) paragraphs rendered in five open fonts, degraded like scans and read by Tesseract | OCR errors | 93,499 pairs | CC BY-SA text |
| Yarmouk Arabic OCR Dataset: real scans, read by Tesseract and by ABBYY | real OCR errors | 104,553 pairs | listed as CC0; the text is Wikipedia's |
| Arabic Wikipedia edit history | real spelling fixes, learned only on the words the editor changed | 253,294 pairs | CC BY-SA |
Training procedure
| Pretraining | Correction | |
|---|---|---|
| Objective | restore 15% hidden characters (whole words and 1โ5 character spans) | one edit label per character |
| Steps | 80,000 ร 16,384 characters (128-character windows, then 512) | 30,000 ร 32 windows of 512 characters (best checkpoint: step 28,000) |
| Mixture | โ | clean 35%, typed noise 25%, OCR 30%, real edits 10% |
| Optimizer | AdamW (ฮฒ 0.9/0.98), learning rate 5e-4, cosine decay | AdamW (ฮฒ 0.9/0.98), learning rate 2e-4, cosine decay |
| Hardware and time | one RTX 3060 Laptop GPU (6 GB), bf16, about 8 hours | same GPU, about 4 hours |
During both stages each attention head's largest score is capped at 50 by scaling its query; this prevents a runaway head that collapsed an earlier pretraining run, and leaves the saved model a standard BERT.
Evaluation
Measured on frozen test sets of real text at the threshold of 0.9:
| Test set | Result |
|---|---|
| T-1: 7,627 Wikipedia paragraphs before and after real spelling fixes | precision 0.962 on the words editors fixed, recall 0.082; elsewhere 3.4 edits per 1,000 words, 76% of them fixes editors made on other pages |
| Benchmark: 27 hand-corrected typed sentences | precision 1.000, recall 0.900, word error rate 38.4% โ 3.8% |
| T-7: clean text, correct words changed | modern 0.058%, classical 0.010%, diacritized 0.020%, dialect 0.007% |
| T-4: real Yarmouk scans, word error rate | 24.0% โ 20.8% (โ13.2%); ABBYY output โ17.2%, Tesseract output โ7.8% |
| T-4: pages no worse than the raw OCR | 99.7% (589 of 591) |
| T-4: harmful edits (a correct word changed, or a wrong word not brought closer) | 5.1% of the model's edits |
| T-6: classical books (OpenITI), error rates | CER 12.6% โ 12.3%, WER 43.3% โ 41.9% |
| Calibration (T-4 and T-6) | expected calibration error 0.048 (0.085 before calibration) |
| Speed, laptop CPU (Intel i5-10500H, 4 threads) | 410 words/s (ONNX int8), 337 (ONNX fp32); with PyTorch about 20 ms per sentence |
T-1 is scored on the words the editors fixed, because a single Wikipedia edit leaves a paragraph's other errors in place; strict precision against such references counts the model's fixes of those errors as damage. References in general contain spelling errors, so some edits counted as harmful are fixes of the reference itself. Full method and per-set results: evaluation.
Limitations
- OCR recall is low. On real scans the model rarely makes a page worse, but it removes only 13% of word errors (the project's target is 20%): most OCR errors are badly garbled words, and it fixes about 1 in 10 of them.
- Errors that produce another valid word need meaning rather than spelling, which a 15M-parameter character model knows little about.
- Quranic text and accepted spelling variants are not protected: verses are corrected like any other text, and either form of a word with two accepted spellings (ู ุณุคูู/ู ุณุฆูู) may be changed.
- Dialect and diacritized text are left alone, not corrected.
- The training text is encyclopedic (Wikipedia) and classical (Wikisource); informal registers such as social media were not evaluated.
Files
| File | Content |
|---|---|
config.json, model.safetensors |
the model (BertForTokenClassification, 178 labels) |
tokenizer.json, tokenizer_config.json, special_tokens_map.json |
the character tokenizer |
labels.json |
the edit labels, in the order of the classifier |
calibration.json |
the threshold and the confidence calibration for typed and OCR input |
License and attribution
The model and the code are released under the MIT License. The training text comes from Wikimedia projects (CC BY-SA) and from the Yarmouk Arabic OCR Dataset (Abu Doush, AlKhateeb and Gharibeh, CSIT 2018, doi:10.1109/CSIT.2018.8486162); test data also include NOD (CC BY 4.0) and the OpenITI OCR gold standard (CC BY-NC-SA 4.0, evaluation only). OCR training pairs were produced with Tesseract (Apache-2.0). The project started from AraSpell (Salhab and Abu-Khzam); the edit-label approach follows GECToR and Alhafni and Habash's Arabic text editing.
Citation
@software{alrefaey2026araspellx,
author = {Al-Refaey, Mahmoud},
title = {{AraSpellX}: Arabic Spelling and {OCR} Error Correction with a Character-Level Transformer},
year = {2026},
url = {https://huggingface.co/mahmoudalrefaey/AraSpellX}
}
- Downloads last month
- 22
Papers for mahmoudalrefaey/AraSpellX
AraSpell: A Deep Learning Approach for Arabic Spelling Correction
Evaluation results
- Precision on the words editors fixed on T-1: real spelling fixes from Arabic Wikipedia's edit historyself-reported0.962
- Recall on the words editors fixed on T-1: real spelling fixes from Arabic Wikipedia's edit historyself-reported0.082
- Word error rate after correction, % (24.0 before) on T-4: Yarmouk real scans (Tesseract and ABBYY output)self-reported20.800