NUMALIA

NUMALIA (نوماليا) is a 5.19M-parameter byte-level dialect identifier for the Maghreb and North Africa, built from scratch and trained end-to-end on the 124,036-row NUMALIA multi-dialect benchmark dataset. The name folds the ancient ancestors of each nation: NUmidia (Algeria) + MAuretania (Morocco & Mauritania) + LIbu (Libya) + Afri (Tunisia).

It distinguishes 8 classes across Arabic script, Arabizi (Latin phoneme-digits 3, 7, 9), and French code-switching — algerian, moroccan, tunisian, libyan, hassaniya, msa, french, other_arabic — at 20.8 MB (float32 safetensors) with zero tokenizer dependencies: input is raw bytes, vocabulary size 258.

On the 12,990-row sealed test split NUMALIA scores 91.79% accuracy and 91.59% macro F1 from scratch, while shipping a mathematically calibrated 99%-coverage conformal prediction set alongside every point estimate.

Results

Evaluated on sealed benchmark splits — one training run, one measurement, seed 42.

Test split (12,990 rows, 8 classes)

dialect / class support precision recall F1
French (french) 1,500 99.40% 100.00% 99.70%
MSA (msa) 1,861 98.36% 99.68% 99.01%
Tunisian (tunisian) 2,000 96.81% 92.50% 94.60%
Other Arabic (other_arabic) 2,000 94.67% 86.15% 90.21%
Libyan (libyan) 1,078 83.98% 96.29% 89.71%
Hassaniya (hassaniya) 552 81.35% 94.02% 87.23%
Algerian (algerian) 1,999 91.89% 82.14% 86.74%
Moroccan (moroccan) 2,000 81.57% 89.85% 85.51%
Overall macro avg 12,990 91.00% 92.58% 91.59%
Overall accuracy 12,990 — — 91.79%

Validation split (9,389 rows)

metric value
Accuracy 93.64%
Macro F1 94.08%
Conformal empirical coverage (99% target, α=0.01) 99.13%
Conformal mean set size 2.02 dialects

Anchor harness. The collective's default reference suites are DziriEval and MADAR. Both evaluate Algerian-only representations; no dialect-identification benchmark spans these 8 classes, so no anchor row applies here — unmeasured with that reason, not by omission.

Single-run constraint. All figures above come from one training run (seed 42) and one evaluation pass. Standard deviations across seeds are unmeasured. A number that cannot be reproduced from the Reproduction section is not a result.

Intended use

Dialect routing and filtering for Maghrebi NLP pipelines: corpus provenance screening, data-collection labelling, and dialect-conditional inference. The conformal set provides a hard per-sample guarantee useful for conservative corpus curation — include a text only when algerian is in the prediction set, and at 99% coverage fewer than 1 in 100 genuine Algerian rows will be rejected.

Not suitable for: generation of any kind (classification only); any language or dialect outside the 8 trained classes; any decision about a person. Arabizi greetings shared across borders (wesh, khoya, cava) produce multi-dialect sets by design — that is the model working correctly, not a failure. It has not been evaluated for bias, toxicity, or factuality.

Usage

transformers and torch, nothing else. The architecture travels with the weights via trust_remote_code=True.

from transformers import AutoModelForSequenceClassification

REPO = "algerian-nlp/NUMALIA"
model = AutoModelForSequenceClassification.from_pretrained(
    REPO,
    trust_remote_code=True,
).eval()

identify returns point prediction, confidence, and the 99%-coverage conformal set for each row:

texts = [
    "واش راك خويا لابس عليك؟",                  # Algerian Arabic script
    "wesh rak khoya cava chwiya?",              # Arabizi (Algerian/cross-Maghrebi)
    "اش خبارك اخاي لاباس عليك كولشي مزيان؟",   # Moroccan Arabic script
    "chnawa a7welek ya sahbi?",                 # Tunisian Arabizi
    "شنو جوك اليوم يا غالي؟",                   # Libyan Arabic script
    "يكانك اتدير الفظة ماهي الاوقية يالتك اتكيس البنك الشعبي",  # Hassaniya
    "عقد مجلس الأمن الدولي جلسة طارئة لمناقشة الأوضاع الراهنة.",  # MSA
    "Le gouvernement a annoncé de nouvelles mesures économiques.",  # French
]

for r in model.identify(texts):
    print(f"{r.language} ({r.confidence:.2%}) → {list(r.prediction_set)}")

Representative output (numbers from the staged release):

algerian (93.06%) → ['algerian', 'moroccan']
algerian (77.02%) → ['algerian', 'moroccan', 'tunisian', 'other_arabic']
moroccan (97.20%) → ['moroccan']
tunisian (91.40%) → ['tunisian']
libyan (88.10%) → ['libyan']
hassaniya (83.50%) → ['hassaniya']
msa (99.68%) → ['msa']
french (99.40%) → ['french']

Reading the prediction set. Any class with adjusted probability ≥ 1.96% enters the set. At 99% target coverage this threshold is intentionally low: the model is 93% confident in algerian on the first text; moroccan at 3.3% is included only because discarding it would violate the coverage guarantee on a fraction of Moroccan-adjacent texts. The point prediction language is always the argmax and is the right choice when a single label is needed.

Adjustable coverage. identify accepts an optional target_coverage argument; lower values (e.g., 0.95) tighten the threshold and collapse most clear-cut texts to singletons:

for r in model.identify(texts, target_coverage=0.95):
    print(f"{r.language} ({r.confidence:.2%}) → {list(r.prediction_set)}")

Architecture

Parameters 5,191,048 (20.8 MB, float32)
Input Raw bytes (UTF-8), vocabulary 258, max 512 bytes
Conv stem 3 parallel depthwise-separable branches, kernel widths 5 / 9 / 13, dim 128 each
Transformer 6 layers, 8 heads, embed dim 256, FFN dim 704 (SwiGLU), RoPE positions
Pooling Attentive pooling over the sequence
Classifier Scaled cosine classifier with LDAM per-class margins
Projection head 128-dim, used during training for contrastive loss only
Context 512 bytes
Precision float32

No tokenizer. The byte embedding table (258 × 256) handles any script, Arabizi digit (3, 7, 9), and French character without OOV failures or subword fragmentation. The ConvStem captures character, affix, and clitic n-gram patterns before the Transformer contextualises them globally.

Training data

Trained on the NUMALIA multi-dialect benchmark dataset (146,415 total allocated split rows across 13 sources, seed 42): 124,036 training rows, 9,389 validation rows, and 12,990 sealed test rows. The raw pool contained 337,742 rows extracted across 13 sources; 62,514 were pruned during length/noise filtering (<3 words or >512 bytes), 17,752 exact duplicates were removed, and 111,061 rows were held in surplus.

Training split by source (124,036 rows)

source domain / dialect tier train rows share
evageon/IADD North African & Arab world dialects permissive (CC-BY-4.0) 29,734 23.97%
wikimedia/wikipedia-ar Modern Standard Arabic (MSA) restricted (CC-BY-SA-4.0) 16,743 13.50%
wikimedia/wikipedia-fr French articles (sample) restricted (CC-BY-SA-4.0) 15,000 12.09%
zenodo/TUNIZI-v2 Tunisian Arabizi social comments permissive (CC-BY-4.0) 14,706 11.86%
ayoubkirouane/Algerian-Darija Algerian Darija permissive (CC-BY-4.0) 11,061 8.92%
MouadJb/MYC Moroccan YouTube comments unknown (research citation) 10,501 8.47%
LeMGarouani/MAC Moroccan tweets unknown (research citation) 6,692 5.40%
Mansour-Essgaer/Libya-Telecom Libyan telecom tweets unknown (research use) 6,005 4.84%
DarjaCore/algerian-darja-sample Algerian Darija Telegram sample permissive (MIT) 5,929 4.78%
Mansour-Essgaer/Libyan-Resturant Libyan restaurant reviews unknown (research use) 2,976 2.40%
Amin-tech99/ai-for-rim Hassaniya / Mauritanian text permissive (MIT) 2,505 2.02%
tunis-ai/tsac Tunisian social media comments unknown (research use) 1,378 1.11%
mendeley/HASSANIYA Hassaniya Facebook comments permissive (CC-BY-4.0) 806 0.65%
Total 8 classes across Maghreb & Saharan mixed 124,036 100.00%

Training split by class and script

  • By class (124,036 rows): algerian 20,001 (16.13%), moroccan 20,000 (16.12%), tunisian 20,000 (16.12%), other_arabic 20,000 (16.12%), msa 16,743 (13.50%), french 15,000 (12.09%), libyan 8,981 (7.24%), hassaniya 3,311 (2.67%).
  • By script: Arabic script (arab) 84,535 (68.15%), Latin / Arabizi (latn) 36,539 (29.46%), mixed (mixed) 2,962 (2.39%).

Class imbalance in lower-resource dialects (hassaniya at 3,311 and libyan at 8,981) is handled during training via Group-DRO with LDAM margins and compensated at inference through Balanced Softmax prior adjustment (τ = 0.5).

Sealing and decontamination. The test split (12,990 rows) is strictly sealed against training and dev splits (zero hash collisions). Where independent test partitions existed in upstream sources (e.g. tunis-ai/tsac), they were routed directly to the evaluation split.

Licence composition of the training text

The weights are CC-BY-SA 4.0. That grant does not relicense the text they were trained on. The exact measured composition across the training partition:

tier rows share note
permissive 64,741 52.19% CC-BY-4.0 (IADD, TUNIZI-v2, Algerian-Darija, Mendeley HASSANIYA) and MIT (DarjaCore, ai-for-rim)
restricted 31,743 25.59% CC-BY-SA-4.0 (Wikipedia Arabic and Wikipedia French)
unknown 27,552 22.21% Research and academic sources without formal open licences (MYC, MAC, Libya-Telecom, Libyan-Resturant, TSAC)

Across all 146,415 rows in the three splits, the composition is 51.76% permissive (75,777 rows), 25.61% restricted (37,499 rows), and 22.63% unknown (33,139 rows). A permissive-only dataset variant can be constructed mechanically by filtering on the tier field in the split metadata.

Training recipe

Objective Cross-entropy + supervised contrastive (SupCon, weight 0.5) + Group-DRO with LDAM margins
Optimiser AdamW, lr 3e-4, weight decay 0.01
Batch 12 groups × 10 samples = 120 per step (group-structured for contrastive)
Schedule Cosine decay with warm-up (deferred share 0.5)
Epochs 8 (best checkpoint selected on validation macro F1)
Contrastive Temperature 0.05, memory bank 2,048 samples, hard negatives across shared sources
Group-DRO Step size 0.01, tracking worst (dialect, source) cell
Augmentation Phrase-level mixing prob 0.35, case augmentation prob 0.10
Label smoothing 0.05
Logit adjustment τ = 0.5 (Balanced Softmax prior shift at inference)
Max grad norm 1.0
Precision float32 throughout
Hardware 1× Kaggle T4 (16 GB); wall-clock from log: 10,783 s (2.99 h)
Seed 42

Best validation checkpoint achieved 93.71% accuracy and 94.10% macro F1 at step 6,966 (worst-group recall 67.86% on Hassaniya). Training completed at step 8,264 with 93.64% accuracy and 94.08% macro F1.

Limitations

Single-run, no seed sweep. All numbers come from one training run and one evaluation pass under seed 42. Standard deviations across seeds are unmeasured.

Cross-Maghrebi Arabizi is genuinely ambiguous. Greetings shared between Algeria, Morocco, and Tunisia (wesh, khoya, chwiya, sahbi) produce multi-dialect prediction sets — a correct reflection of linguistic overlap, not a model defect.

Hassaniya is low-resource. 3,311 training rows with a mean set size of 2.02 and recall 94.02% on the 552-row test partition; this recall will degrade on Hassaniya text that diverges stylistically from the Mendeley scrape. The LDAM + logit-adjustment approach partially compensates but does not replace more data.

Conformal calibration is split-specific. The quantile q̂ = 0.9804 was fitted on the first half of a 9,389-row held-out split. Coverage is guaranteed (≥ 99%) on the second half; distribution shift outside the Maghrebi sources above is not covered by the guarantee.

8-class scope. Levantine, Gulf, Egyptian, and Mesopotamian Arabic collapse into other_arabic. Fine-grained Eastern Arabic identification is out of scope.

No safety evaluation of any kind has been performed.

Files

file size contents
model.safetensors 20.8 MB all 5,191,048 parameters (float32) · SHA-256 33ebc03ba642064ce4f3663aa0a46f12dd5dd89eb0e9e065a40286606c5f5e8f
config.json 1.4 KB architecture plus auto_map for trust_remote_code, calibration constants
modeling_numalia.py 17 KB full architecture and identify API in one self-contained file
calibration.json 551 B conformal q̂, target coverage, prior-shift vector, class list
export_report.json 980 B measured sizes, SHA-256 per file, parameter count, reload parity

modeling_numalia.py is the hub package flattened into one file by the export; the export verifies that the staged directory reloads to identical predictions (max confidence diff = 0.0) before writing the report.

Reproduction

The published weights (model.safetensors, SHA-256 33ebc03ba642064ce4f3663aa0a46f12dd5dd89eb0e9e065a40286606c5f5e8f) load directly with zero missing and zero unexpected keys:

from transformers import AutoModelForSequenceClassification

model = AutoModelForSequenceClassification.from_pretrained(
    "algerian-nlp/NUMALIA",
    trust_remote_code=True,
).eval()

All training hyperparameters, loss schedules, and architecture parameters are detailed in the Training recipe and Architecture sections above. The evaluation seed is 42 throughout; evaluation passes under the same environment reproduce the reported test and validation metrics to float32 precision.

Citation

If you use NUMALIA in your research or applications, please cite:

@software{ainouche_algerian_nlp_NUMALIA_2026,
  title  = {NUMALIA: Byte-Level Maghrebi Dialect Identification with Conformal Coverage Guarantees},
  author = {Ainouche, Abderahmane and Algerian NLP Collective},
  year   = {2026},
  url    = {https://huggingface.co/algerian-nlp/NUMALIA}
}

Licence

CC-BY-SA 4.0 for the weights and card. Read the licence composition of the training text above before redistributing derivatives — a CC-BY-SA grant on the weights makes no claim about the underlying text.

Downloads last month
19
Safetensors
Model size
5.19M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Evaluation results

  • Accuracy (test, 8 classes, single run) on LID benchmark (test, 12,990 rows)
    self-reported
    0.918
  • Macro F1 (test, 8 classes, single run) on LID benchmark (test, 12,990 rows)
    self-reported
    0.916
  • Accuracy (validation, 8 classes, single run) on LID benchmark (validation, 9,389 rows)
    self-reported
    0.936
  • Macro F1 (validation, 8 classes, single run) on LID benchmark (validation, 9,389 rows)
    self-reported
    0.941