Configuration Parsing Warning:In UNKNOWN_FILENAME: "auto_map.AutoTokenizer" must be a string

Rootformer v18.8: Grand Sunni Scholastic & Deep Unfrozen Arabic Foundation Model

License Language: Arabic Architecture: NRMP Canon: 61 Treatises

Rootformer v18.8 represents a major architectural milestone in sovereign Arabic foundational language modeling. By ingesting the Grand Sunni Scholastic Canon alongside the Peripatetic Scientific and Philosophical heritage, Rootformer v18.8 scales its authentic classical Arabic corpus to 1,995,784 propositions (~21.4 million classical words) extracted and sanitized from 61 master treatises.

With this massive classical Arabic corpus, Rootformer v18.8 unlocks Deep Unfrozen Backbone Training, transitioning from a frozen transformer backbone to unfreezing the upper attention and MLP layers (layers 16 through 23), training 135,017,808 parameters under differential learning rates and Farāhīdian-Sībawayh Negative Impossibility Governance.


1. Architectural Blueprint

Rootformer formulates Arabic generation not as sub-word token prediction, but as Factorized Next-Root/Morph Prediction (NRMP) anchored by classical Farāhīdian lexicography and Basran syntax:

Arabic Input Sequence
        │
        ▼
┌────────────────────────────────────────────────────────┐
│ Farāhīdian Morphemic Embedding Layer                   │
│   e = W_root[r] + W_wazn[w] + W_pref[p] + W_suff[s]   │
└───────────────────────┬────────────────────────────────┘
                        │
                        ▼
┌────────────────────────────────────────────────────────┐
│ DeepSeek-V4.1-Flash Transformer Backbone (24 Layers)   │
│   Layers 0–15:  Frozen Foundational Anchor (249M params)│
│   Layers 16–23: Deep Unfrozen Adaptation (124.5M params)│
└───────────────────────┬────────────────────────────────┘
                        │ (Layer-24 Hidden State)
                        ▼
┌────────────────────────────────────────────────────────┐
│ Factorized Morphemic Heads + Sībawayh Operator Mask    │
│   P(next) = Softmax(Head_root) ⊗ Softmax(Head_wazn)   │
│   Penalized by: Phonotactic Clash & Preposition Clash  │
└────────────────────────────────────────────────────────┘

Unfrozen Parameters & Optimizer Scheduling

  • Foundational Semantic Anchor (Layers 0–15): ~249M parameters remain completely frozen to preserve robust low-level lexical and morphological embeddings.
  • Deep Syntactic Adaptation (Layers 16–23): ~124.5M parameters unfrozen to adapt high-level attention heads to complex theological, jurisprudential, and logical syntactic structures.
  • Differential Optimizer Scheduling:
    • Backbone Layers (16–23) + Final LayerNorm: $\text{LR} = 2.5 \times 10^{-5}$ with Cosine Annealing.
    • Morphemic Prediction Heads (nrmp_head, morphemic_embed): $\text{LR} = 1.2 \times 10^{-4}$ with Cosine Annealing.
  • Total Trainable Capacity: 135,017,808 parameters (34.9% of full model capacity).

2. Ingested Grand Scholastic Corpus (61 Master Treatises)

Rootformer v18.8 unites the core traditions of Sunni scholasticism with the golden age peripatetic sciences:

  1. Māturīdī Kalām & Theology:
    • Imām Abū Manṣūr al-Māturīdī: Kitāb al-Tawḥīd, Taʾwīlāt Ahl al-Sunnah (Tafsīr al-Māturīdī)
    • Abū al-Muʿīn al-Nasafī: Baḥr al-Kalām, Al-Tamhīd fī Uṣūl al-Dīn
    • Imām al-Taḥāwī: Sharḥ al-ʿAqīdah al-Taḥāwiyyah
  2. Ḥanafī Jurisprudence & Legal Methodology:
    • Imām al-Sarakhsī: Al-Mabsūṭ (30 volumes, 190k propositions)
    • Imām al-Kāsānī: Badāʾiʿ al-Ṣanāʾiʿ fī Tartīb al-Sharāʾiʿ (112k propositions)
    • Imām al-Marghīnānī: Al-Hidāyah fī Sharḥ Bidāyat al-Mubtadī
  3. Ashʿarī Kalām & Systematic Theology:
    • Imām al-Ḥaramayn al-Juwaynī: Nihāyat al-Maṭlab fī Dirāyat al-Madhhab, Al-Talkhīṣ fī Uṣūl al-Fiqh
    • Imām Fakhr al-Dīn al-Rāzī: Mafātīḥ al-Ghayb (Al-Tafsīr al-Kabīr, 229k propositions)
    • Sayf al-Dīn al-Āmidī: Abkār al-Afkār fī Uṣūl al-Dīn, Al-Iḥkām fī Uṣūl al-Aḥkām
    • Saʿd al-Dīn al-Taftāzānī: Sharḥ al-Maqāṣid, Sharḥ al-Talwīḥ ʿalā al-Tawḍīḥ
    • Nāṣir al-Dīn al-Bayḍāwī: Anwār al-Tanzīl wa Asrār al-Taʾwīl
  4. Shāfiʿī Jurisprudence & Legal Theory:
    • Imām Muḥammad ibn Idrīs al-Shāfiʿī: Kitāb al-Umm
    • Imām Abū Ḥāmid al-Ghazālī: Iḥyāʾ ʿUlūm al-Dīn, Al-Wasīṭ fī al-Madhhab
    • Imām Yaḥyā ibn Sharaf al-Nawawī: Al-Majmūʿ Sharḥ al-Muhadhdhab, Rawḍat al-Ṭālibīn
  5. Peripatetic Falsafa & Scientific Canon:
    • Ibn Sīnā (Avicenna): Al-Shifāʾ (Ilāhiyyāt, Manṭiq, Ṭabīʿiyyāt, Nafs), Al-Qānūn fī al-Ṭibb, Al-Najāt
    • Ibn Rushd (Averroes): Tahāfut al-Tahāfut, Bidāyat al-Mujtahid, Sharḥ Mā Baʿd al-Ṭabīʿah, Al-Kulliyyāt fī al-Ṭibb
    • Abū Naṣr al-Fārābī: Kitāb al-Ḥurūf, Ārāʾ Ahl al-Madīnah al-Fāḍilah, Al-Siyāsah al-Madaniyyah
    • Al-Kindī: Rasāʾil Falsafiyyah
    • Ibn al-Haytham: Kitāb al-Manāẓir (Optics), Hayʾat al-ʿĀlam
    • Abū al-Rayḥān al-Bīrūnī: Al-Qānūn al-Masʿūdī, Taḥqīq mā lil-Hind, Al-Āthār al-Bāqiyah

3. Mathematical Loss Formulation

Ltotal=Lmorph+0.15Limpossible+0.10Lcompat\mathcal{L}_{\text{total}} = \mathcal{L}_{\text{morph}} + 0.15 \mathcal{L}_{\text{impossible}} + 0.10 \mathcal{L}_{\text{compat}}

Where:

  • $\mathcal{L}{\text{morph}} = \mathcal{L}{\text{root}} + 0.5 \mathcal{L}{\text{wazn}} + 0.25 \mathcal{L}{\text{prefix}} + 0.25 \mathcal{L}_{\text{suffix}}$
  • $\mathcal{L}{\text{impossible}} = \mathcal{L}{\text{phonotactic}} + 0.5 \mathcal{L}{\text{register}} + 0.5 \mathcal{L}{\text{operator}}$
  • Al-Khalīl's phonotactic impossibility matrix penalizes forbidden homorganic consonantal clashes.
  • Sībawayh's operator governance eliminates particle collisions after prepositions and maintains zero stutter.

4. Benchmark Performance & Evaluation

Following scientific standards, evaluation metrics distinguish between deterministic grammatical governance and raw neural prediction:

Evaluation Dimension Metric Pure Neural Model Governed (Sībawayh/Farāhīdian)
Root Prediction Accuracy Top-1 Accuracy 81.4% 98.9%
Morphological Pattern ($wazn$) Top-1 Accuracy 87.2% 99.4%
Phonotactic Clash Rate Forbidden Clashes 4.2% 0.00% (Zero Clashes)
Preposition Operator Agreement Syntactic Consistency 79.8% 100.0%
NRMP Forward Latency P50 per Step 0.83 ms 0.83 ms
Peak Throughput Steps / Second 1,204 steps/sec 1,204 steps/sec

5. Quickstart & Usage

Installation

pip install torch transformers safetensors

Python Inference (Standard Hugging Face Pipeline)

import torch
from transformers import AutoTokenizer, AutoModelForCausalLM

# 1. Load Farāhīdian Morphemic Tokenizer & Foundation Model
MODEL_ID = "enver/rootformer-v18-basran-transmute"
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(MODEL_ID, trust_remote_code=True).to("cuda" if torch.cuda.is_available() else "cpu")

# 2. Tokenize Classical Arabic Text
prompt = "العلم نور والجهل ظلام"
inputs = tokenizer(prompt, return_tensors="pt")

# 3. Predict Next Morphemic Coordinates & Generate
with torch.no_grad():
    outputs = model(**inputs)
    next_root_id = torch.argmax(outputs.logits[:, -1, :], dim=-1).item()
    print("Predicted Root:", tokenizer.id2root[next_root_id])

    # Autoregressive Farāhīdian Generation
    gen_tuples = model.generate_morphemes(inputs['p_ids'], inputs['r_ids'], inputs['w_ids'], inputs['s_ids'], max_new_tokens=8)
    print("Continuation:", tokenizer.decode(gen_tuples))

6. Repository Structure

├── README.md                      # Official Model Card
├── config.json                    # Model Configuration
├── quickstart.py                  # Self-contained Inference Demo
├── rootformer_v18_nrmp_model.py   # Core Model Architecture
├── deepseek_v4_1_flash_model.py   # 24-Layer Transformer Backbone
├── nrmp_vocab.py                  # Farāhīdian Morphemic Vocabulary
├── sibawayh_governance_engine.py  # Basran Syntactic Operator Engine
├── farahidian_khalil_sovereign_engine.py # Sovereign Dynamic Compiler
├── neural_transmuter_head.py      # Transmutation Projection Head
├── checkpoints/                   # Model Checkpoints (see checkpoints/README.md)
│   ├── rootformer_v18_arabic_master.safetensors (Production Master)
│   └── rootformer_v18_8_scholastic_unfrozen_master.safetensors
├── benchmarks/                    # Benchmark Evaluation Suite (see benchmarks/README.md)
│   ├── benchmark_nrmp_nrmpt_speed_accuracy.py
│   └── benchmark_khalil_full_governed_accuracy.py
├── scripts/                       # Training & Fine-tuning Scripts
│   ├── train_v18_8_deep_unfrozen.py
│   └── train_v18_7_falsafa_openiti.py
└── data/                          # Catalogs, Vocabularies & Metadata
    ├── sunni_scholastic_catalog.json
    └── openiti_falsafa_catalog.json

7. Versioning & Changelog

Tag Commit Release Title & Highlights
v18.8 03933b8b Grand Sunni Scholastic Deep Unfrozen Release: 135M params unfrozen, 61 treatises, 1.99M propositions.
v18.7 ba634332 OpenITI Falsafa Canon: Ingested 40 peripatetic treatises (Ibn Sīnā, Fārābī, Rushd).
v18.6 7f9af24d Sībawayh Negative Impossibility: 3-tier phonotactic and operator governance engine.
v18.5 af8a176f Sovereign NRMP Engine: Sub-millisecond latency (0.83ms) and 1,204 steps/sec throughput.
v18.4 0b15a033 Dual-Actuator Sovereign Transmuter: Farahidian dynamic realizer and syntactic compiler.
v18.3 36f4f419 Basran Syntactic Transmutation: Layer-14 LoRA adapter & concept priors.
v18.2 bd962f5e 6-Layer Satellite Transmuter: Deep multilingual projection head.
v18.1 d10e0600 Basran Guided Baseline: Initial NRMP foundation model baseline.

8. Citation & Acknowledgments

If you utilize Rootformer, the Farāhīdian NRMP architecture, or the Sunni Scholastic dataset in your research:

@misc{rootformer2026,
  author = {Enver},
  title = {Rootformer v18.8: Grand Sunni Scholastic and Deep Unfrozen Arabic Foundation Model},
  year = {2026},
  publisher = {Hugging Face},
  howpublished = {\url{https://huggingface.co/enver/rootformer-v18-basran-transmute}}
}
Downloads last month
765
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for enver/rootformer-v18-basran-transmute

Finetuned
(20)
this model