Instructions to use algerian-nlp/NUMALIA with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use algerian-nlp/NUMALIA with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="algerian-nlp/NUMALIA", trust_remote_code=True)# Load model directly from transformers import AutoModelForSequenceClassification model = AutoModelForSequenceClassification.from_pretrained("algerian-nlp/NUMALIA", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
NUMALIA
NUMALIA (نوماليا) is a 5.19M-parameter byte-level dialect identifier for the Maghreb and North Africa, built from scratch and trained end-to-end on the 124,036-row NUMALIA multi-dialect benchmark dataset. The name folds the ancient ancestors of each nation: NUmidia (Algeria) + MAuretania (Morocco & Mauritania) + LIbu (Libya) + Afri (Tunisia).
It distinguishes 8 classes across Arabic script, Arabizi (Latin phoneme-digits 3, 7, 9), and French code-switching — algerian, moroccan, tunisian, libyan, hassaniya, msa, french, other_arabic — at 20.8 MB (float32 safetensors) with zero tokenizer dependencies: input is raw bytes, vocabulary size 258.
On the 12,990-row sealed test split NUMALIA scores 91.79% accuracy and 91.59% macro F1 from scratch, while shipping a mathematically calibrated 99%-coverage conformal prediction set alongside every point estimate.
Results
Evaluated on sealed benchmark splits — one training run, one measurement, seed 42.
Test split (12,990 rows, 8 classes)
| dialect / class | support | precision | recall | F1 |
|---|---|---|---|---|
French (french) |
1,500 | 99.40% | 100.00% | 99.70% |
MSA (msa) |
1,861 | 98.36% | 99.68% | 99.01% |
Tunisian (tunisian) |
2,000 | 96.81% | 92.50% | 94.60% |
Other Arabic (other_arabic) |
2,000 | 94.67% | 86.15% | 90.21% |
Libyan (libyan) |
1,078 | 83.98% | 96.29% | 89.71% |
Hassaniya (hassaniya) |
552 | 81.35% | 94.02% | 87.23% |
Algerian (algerian) |
1,999 | 91.89% | 82.14% | 86.74% |
Moroccan (moroccan) |
2,000 | 81.57% | 89.85% | 85.51% |
| Overall macro avg | 12,990 | 91.00% | 92.58% | 91.59% |
| Overall accuracy | 12,990 | — | — | 91.79% |
Validation split (9,389 rows)
| metric | value |
|---|---|
| Accuracy | 93.64% |
| Macro F1 | 94.08% |
| Conformal empirical coverage (99% target, α=0.01) | 99.13% |
| Conformal mean set size | 2.02 dialects |
Anchor harness. The collective's default reference suites are DziriEval and MADAR. Both evaluate Algerian-only representations; no dialect-identification benchmark spans these 8 classes, so no anchor row applies here — unmeasured with that reason, not by omission.
Single-run constraint. All figures above come from one training run (seed 42) and one evaluation pass. Standard deviations across seeds are unmeasured. A number that cannot be reproduced from the Reproduction section is not a result.
Intended use
Dialect routing and filtering for Maghrebi NLP pipelines: corpus provenance screening, data-collection labelling, and dialect-conditional inference. The conformal set provides a hard per-sample guarantee useful for conservative corpus curation — include a text only when algerian is in the prediction set, and at 99% coverage fewer than 1 in 100 genuine Algerian rows will be rejected.
Not suitable for: generation of any kind (classification only); any language or dialect outside the 8 trained classes; any decision about a person. Arabizi greetings shared across borders (wesh, khoya, cava) produce multi-dialect sets by design — that is the model working correctly, not a failure. It has not been evaluated for bias, toxicity, or factuality.
Usage
transformers and torch, nothing else. The architecture travels with the weights via trust_remote_code=True.
from transformers import AutoModelForSequenceClassification
REPO = "algerian-nlp/NUMALIA"
model = AutoModelForSequenceClassification.from_pretrained(
REPO,
trust_remote_code=True,
).eval()
identify returns point prediction, confidence, and the 99%-coverage conformal set for each row:
texts = [
"واش راك خويا لابس عليك؟", # Algerian Arabic script
"wesh rak khoya cava chwiya?", # Arabizi (Algerian/cross-Maghrebi)
"اش خبارك اخاي لاباس عليك كولشي مزيان؟", # Moroccan Arabic script
"chnawa a7welek ya sahbi?", # Tunisian Arabizi
"شنو جوك اليوم يا غالي؟", # Libyan Arabic script
"يكانك اتدير الفظة ماهي الاوقية يالتك اتكيس البنك الشعبي", # Hassaniya
"عقد مجلس الأمن الدولي جلسة طارئة لمناقشة الأوضاع الراهنة.", # MSA
"Le gouvernement a annoncé de nouvelles mesures économiques.", # French
]
for r in model.identify(texts):
print(f"{r.language} ({r.confidence:.2%}) → {list(r.prediction_set)}")
Representative output (numbers from the staged release):
algerian (93.06%) → ['algerian', 'moroccan']
algerian (77.02%) → ['algerian', 'moroccan', 'tunisian', 'other_arabic']
moroccan (97.20%) → ['moroccan']
tunisian (91.40%) → ['tunisian']
libyan (88.10%) → ['libyan']
hassaniya (83.50%) → ['hassaniya']
msa (99.68%) → ['msa']
french (99.40%) → ['french']
Reading the prediction set. Any class with adjusted probability ≥ 1.96% enters the set. At 99% target coverage this threshold is intentionally low: the model is 93% confident in algerian on the first text; moroccan at 3.3% is included only because discarding it would violate the coverage guarantee on a fraction of Moroccan-adjacent texts. The point prediction language is always the argmax and is the right choice when a single label is needed.
Adjustable coverage. identify accepts an optional target_coverage argument; lower values (e.g., 0.95) tighten the threshold and collapse most clear-cut texts to singletons:
for r in model.identify(texts, target_coverage=0.95):
print(f"{r.language} ({r.confidence:.2%}) → {list(r.prediction_set)}")
Architecture
| Parameters | 5,191,048 (20.8 MB, float32) |
| Input | Raw bytes (UTF-8), vocabulary 258, max 512 bytes |
| Conv stem | 3 parallel depthwise-separable branches, kernel widths 5 / 9 / 13, dim 128 each |
| Transformer | 6 layers, 8 heads, embed dim 256, FFN dim 704 (SwiGLU), RoPE positions |
| Pooling | Attentive pooling over the sequence |
| Classifier | Scaled cosine classifier with LDAM per-class margins |
| Projection head | 128-dim, used during training for contrastive loss only |
| Context | 512 bytes |
| Precision | float32 |
No tokenizer. The byte embedding table (258 × 256) handles any script, Arabizi digit (3, 7, 9), and French character without OOV failures or subword fragmentation. The ConvStem captures character, affix, and clitic n-gram patterns before the Transformer contextualises them globally.
Training data
Trained on the NUMALIA multi-dialect benchmark dataset (146,415 total allocated split rows across 13 sources, seed 42): 124,036 training rows, 9,389 validation rows, and 12,990 sealed test rows. The raw pool contained 337,742 rows extracted across 13 sources; 62,514 were pruned during length/noise filtering (<3 words or >512 bytes), 17,752 exact duplicates were removed, and 111,061 rows were held in surplus.
Training split by source (124,036 rows)
| source | domain / dialect | tier | train rows | share |
|---|---|---|---|---|
evageon/IADD |
North African & Arab world dialects | permissive (CC-BY-4.0) | 29,734 | 23.97% |
wikimedia/wikipedia-ar |
Modern Standard Arabic (MSA) | restricted (CC-BY-SA-4.0) | 16,743 | 13.50% |
wikimedia/wikipedia-fr |
French articles (sample) | restricted (CC-BY-SA-4.0) | 15,000 | 12.09% |
zenodo/TUNIZI-v2 |
Tunisian Arabizi social comments | permissive (CC-BY-4.0) | 14,706 | 11.86% |
ayoubkirouane/Algerian-Darija |
Algerian Darija | permissive (CC-BY-4.0) | 11,061 | 8.92% |
MouadJb/MYC |
Moroccan YouTube comments | unknown (research citation) | 10,501 | 8.47% |
LeMGarouani/MAC |
Moroccan tweets | unknown (research citation) | 6,692 | 5.40% |
Mansour-Essgaer/Libya-Telecom |
Libyan telecom tweets | unknown (research use) | 6,005 | 4.84% |
DarjaCore/algerian-darja-sample |
Algerian Darija Telegram sample | permissive (MIT) | 5,929 | 4.78% |
Mansour-Essgaer/Libyan-Resturant |
Libyan restaurant reviews | unknown (research use) | 2,976 | 2.40% |
Amin-tech99/ai-for-rim |
Hassaniya / Mauritanian text | permissive (MIT) | 2,505 | 2.02% |
tunis-ai/tsac |
Tunisian social media comments | unknown (research use) | 1,378 | 1.11% |
mendeley/HASSANIYA |
Hassaniya Facebook comments | permissive (CC-BY-4.0) | 806 | 0.65% |
| Total | 8 classes across Maghreb & Saharan | mixed | 124,036 | 100.00% |
Training split by class and script
- By class (124,036 rows):
algerian20,001 (16.13%),moroccan20,000 (16.12%),tunisian20,000 (16.12%),other_arabic20,000 (16.12%),msa16,743 (13.50%),french15,000 (12.09%),libyan8,981 (7.24%),hassaniya3,311 (2.67%). - By script: Arabic script (
arab) 84,535 (68.15%), Latin / Arabizi (latn) 36,539 (29.46%), mixed (mixed) 2,962 (2.39%).
Class imbalance in lower-resource dialects (hassaniya at 3,311 and libyan at 8,981) is handled during training via Group-DRO with LDAM margins and compensated at inference through Balanced Softmax prior adjustment (τ = 0.5).
Sealing and decontamination. The test split (12,990 rows) is strictly sealed against training and dev splits (zero hash collisions). Where independent test partitions existed in upstream sources (e.g. tunis-ai/tsac), they were routed directly to the evaluation split.
Licence composition of the training text
The weights are CC-BY-SA 4.0. That grant does not relicense the text they were trained on. The exact measured composition across the training partition:
| tier | rows | share | note |
|---|---|---|---|
| permissive | 64,741 | 52.19% | CC-BY-4.0 (IADD, TUNIZI-v2, Algerian-Darija, Mendeley HASSANIYA) and MIT (DarjaCore, ai-for-rim) |
| restricted | 31,743 | 25.59% | CC-BY-SA-4.0 (Wikipedia Arabic and Wikipedia French) |
| unknown | 27,552 | 22.21% | Research and academic sources without formal open licences (MYC, MAC, Libya-Telecom, Libyan-Resturant, TSAC) |
Across all 146,415 rows in the three splits, the composition is 51.76% permissive (75,777 rows), 25.61% restricted (37,499 rows), and 22.63% unknown (33,139 rows). A permissive-only dataset variant can be constructed mechanically by filtering on the tier field in the split metadata.
Training recipe
| Objective | Cross-entropy + supervised contrastive (SupCon, weight 0.5) + Group-DRO with LDAM margins |
| Optimiser | AdamW, lr 3e-4, weight decay 0.01 |
| Batch | 12 groups × 10 samples = 120 per step (group-structured for contrastive) |
| Schedule | Cosine decay with warm-up (deferred share 0.5) |
| Epochs | 8 (best checkpoint selected on validation macro F1) |
| Contrastive | Temperature 0.05, memory bank 2,048 samples, hard negatives across shared sources |
| Group-DRO | Step size 0.01, tracking worst (dialect, source) cell |
| Augmentation | Phrase-level mixing prob 0.35, case augmentation prob 0.10 |
| Label smoothing | 0.05 |
| Logit adjustment | τ = 0.5 (Balanced Softmax prior shift at inference) |
| Max grad norm | 1.0 |
| Precision | float32 throughout |
| Hardware | 1× Kaggle T4 (16 GB); wall-clock from log: 10,783 s (2.99 h) |
| Seed | 42 |
Best validation checkpoint achieved 93.71% accuracy and 94.10% macro F1 at step 6,966 (worst-group recall 67.86% on Hassaniya). Training completed at step 8,264 with 93.64% accuracy and 94.08% macro F1.
Limitations
Single-run, no seed sweep. All numbers come from one training run and one evaluation pass under seed 42. Standard deviations across seeds are unmeasured.
Cross-Maghrebi Arabizi is genuinely ambiguous. Greetings shared between Algeria, Morocco, and Tunisia (wesh, khoya, chwiya, sahbi) produce multi-dialect prediction sets — a correct reflection of linguistic overlap, not a model defect.
Hassaniya is low-resource. 3,311 training rows with a mean set size of 2.02 and recall 94.02% on the 552-row test partition; this recall will degrade on Hassaniya text that diverges stylistically from the Mendeley scrape. The LDAM + logit-adjustment approach partially compensates but does not replace more data.
Conformal calibration is split-specific. The quantile q̂ = 0.9804 was fitted on the first half of a 9,389-row held-out split. Coverage is guaranteed (≥ 99%) on the second half; distribution shift outside the Maghrebi sources above is not covered by the guarantee.
8-class scope. Levantine, Gulf, Egyptian, and Mesopotamian Arabic collapse into other_arabic. Fine-grained Eastern Arabic identification is out of scope.
No safety evaluation of any kind has been performed.
Files
| file | size | contents |
|---|---|---|
model.safetensors |
20.8 MB | all 5,191,048 parameters (float32) · SHA-256 33ebc03ba642064ce4f3663aa0a46f12dd5dd89eb0e9e065a40286606c5f5e8f |
config.json |
1.4 KB | architecture plus auto_map for trust_remote_code, calibration constants |
modeling_numalia.py |
17 KB | full architecture and identify API in one self-contained file |
calibration.json |
551 B | conformal q̂, target coverage, prior-shift vector, class list |
export_report.json |
980 B | measured sizes, SHA-256 per file, parameter count, reload parity |
modeling_numalia.py is the hub package flattened into one file by the export; the export verifies that the staged directory reloads to identical predictions (max confidence diff = 0.0) before writing the report.
Reproduction
The published weights (model.safetensors, SHA-256 33ebc03ba642064ce4f3663aa0a46f12dd5dd89eb0e9e065a40286606c5f5e8f) load directly with zero missing and zero unexpected keys:
from transformers import AutoModelForSequenceClassification
model = AutoModelForSequenceClassification.from_pretrained(
"algerian-nlp/NUMALIA",
trust_remote_code=True,
).eval()
All training hyperparameters, loss schedules, and architecture parameters are detailed in the Training recipe and Architecture sections above. The evaluation seed is 42 throughout; evaluation passes under the same environment reproduce the reported test and validation metrics to float32 precision.
Citation
If you use NUMALIA in your research or applications, please cite:
@software{ainouche_algerian_nlp_NUMALIA_2026,
title = {NUMALIA: Byte-Level Maghrebi Dialect Identification with Conformal Coverage Guarantees},
author = {Ainouche, Abderahmane and Algerian NLP Collective},
year = {2026},
url = {https://huggingface.co/algerian-nlp/NUMALIA}
}
Licence
CC-BY-SA 4.0 for the weights and card. Read the licence composition of the training text above before redistributing derivatives — a CC-BY-SA grant on the weights makes no claim about the underlying text.
- Downloads last month
- 19
Evaluation results
- Accuracy (test, 8 classes, single run) on LID benchmark (test, 12,990 rows)self-reported0.918
- Macro F1 (test, 8 classes, single run) on LID benchmark (test, 12,990 rows)self-reported0.916
- Accuracy (validation, 8 classes, single run) on LID benchmark (validation, 9,389 rows)self-reported0.936
- Macro F1 (validation, 8 classes, single run) on LID benchmark (validation, 9,389 rows)self-reported0.941