Manara-3B: Sovereign SLM for the Kuwaiti Banking Sector
Manara-3B (Ω ΩΨ§Ψ±Ω-3B) is a sovereign, bilingual (Business Arabic / English) Small Language Model (SLM) purpose-built for the Kuwaiti banking and financial sector. Built upon the Jais Arabic foundation model, Manara-3B is designed for on-premise deployment and provides deep, regulation-grounded intelligence across three critical domains.
Architecture Overview
The Manara-3B pipeline is structured as a three-phase system, following a Teacher-Student distillation paradigm.
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β MANARA-3B PIPELINE β
β β
β ββββββββββββββββ ββββββββββββββββββββ ββββββββββββββββββββββββ β
β β STEP 1 β β STEP 2 β β STEP 3 β β
β β Teacher βββββΆβ Student βββββΆβ Self-Reflection β β
β β (Jais API) β β (LoRA/Unsloth) β β (Critic Pass) β β
β β β β β β β β
β β 50,000 pairs β β 3B-param SLM β β 12 Guardrails β β
β β CBK/Sharia/ β β 16GB VRAM β β JSON Correction Log β β
β β Boursa KW β β safetensors out β β Recursive Retraining β β
β ββββββββββββββββ ββββββββββββββββββββ ββββββββββββββββββββββββ β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
Domain Coverage
Manara-3B is trained on three core knowledge domains specific to the Kuwaiti financial ecosystem.
| Domain | Coverage | Weight |
|---|---|---|
| CBK 2026 Regulations | Capital adequacy, AML/CFT, Open Banking, KACH, KDMS, Cybersecurity Framework | 40% |
| Sharia-Compliant Finance | Murabaha, Wakala, Ijara, Mudarabah, Musharakah, Sukuk, AAOIFI standards | 35% |
| Boursa Kuwait Standards | ESG reporting, financial disclosure, corporate governance, IFRS, TCFD | 25% |
Repository Structure
manara-3b/
βββ README.md # This file
βββ model_card.md # Professional HuggingFace model card
βββ requirements.txt # Python dependencies
βββ config.yaml # Central configuration file
β
βββ src/
β βββ distillation/
β β βββ generate_synthetic_data.py # Step 1: Teacher distillation pipeline
β β βββ preprocess_dataset.py # Step 1b: Data validation & Alpaca formatting
β β
β βββ training/
β β βββ fine_tune_lora.py # Step 2: LoRA/Unsloth fine-tuning script
β β βββ inference.py # Inference engine with guardrail integration
β β
β βββ guardrails/
β βββ self_reflective_loop.py # Step 3: 12-guardrail critic engine
β
βββ data/
β βββ raw/ # Raw generated data (gitignored)
β βββ processed/ # Processed Alpaca-format datasets (gitignored)
β
βββ logs/ # Training & guardrail logs (gitignored)
β βββ distillation.log
β βββ training.log
β βββ self_corrections.jsonl # Self-correction events for retraining
β
βββ model_weights/ # Saved model weights (gitignored)
βββ manara-3b-final/ # Merged float16 model (safetensors)
βββ manara-3b-lora/ # LoRA adapters only (lightweight)
Quickstart
1. Installation
git clone https://github.com/your-org/manara-3b.git
cd manara-3b
pip install -r requirements.txt
# Recommended: Install Unsloth for 2x faster training
pip install unsloth
2. Step 1: Generate Synthetic Training Data
export JAIS_API_KEY="your-jais-api-key"
python src/distillation/generate_synthetic_data.py \
--output_dir data/processed \
--num_pairs 50000 \
--model jais-13b-chat
3. Step 1b: Preprocess the Dataset
python src/distillation/preprocess_dataset.py \
--input_file data/processed/manara_train.jsonl \
--output_file data/processed/manara_alpaca.jsonl
4. Step 2: Fine-Tune the Model
python src/training/fine_tune_lora.py \
--config config.yaml \
--train_file data/processed/manara_train_alpaca.jsonl \
--val_file data/processed/manara_val_alpaca.jsonl \
--output_dir model_weights
5. Step 3: Run Inference with Guardrails
python src/training/inference.py \
--model_path model_weights/manara-3b-final \
--prompt "Ω
Ψ§ ΩΩ Ω
ΨͺΨ·ΩΨ¨Ψ§Ψͺ ΩΨ³Ψ¨Ψ© ΩΩΨ§ΩΨ© Ψ±Ψ£Ψ³ Ψ§ΩΩ
Ψ§Ω ΩΩΩ ΨͺΨΉΩΩΩ
Ψ§Ψͺ Ψ¨ΩΩ Ψ§ΩΩΩΩΨͺ Ψ§ΩΩ
Ψ±ΩΨ²ΩΨ" \
--language ar
The 12 Self-Reflective Guardrails
Every response generated by Manara-3B is automatically verified against 12 domain-specific guardrails. Any violation is logged to logs/self_corrections.jsonl for recursive retraining.
| ID | Guardrail Name | Severity |
|---|---|---|
| G-01 | No Riba (Interest) Statements | Critical |
| G-02 | CBK Regulatory Accuracy | Warning |
| G-03 | Sharia Compliance Verification | Critical |
| G-04 | Boursa Kuwait Disclosure Standards | Warning |
| G-05 | No Fabricated Regulatory Citations | Critical |
| G-06 | No Harmful Financial Advice | Critical |
| G-07 | Currency & Jurisdiction Accuracy (KWD) | Warning |
| G-08 | AAOIFI Standard Alignment | Warning |
| G-09 | No Discriminatory Content | Critical |
| G-10 | Data Privacy & Confidentiality (PII) | Critical |
| G-11 | Kuwaiti Dialect & Register Appropriateness | Info |
| G-12 | Factual Consistency & Internal Coherence | Warning |
Hardware Requirements
| Deployment Mode | VRAM | Notes |
|---|---|---|
| Training (LoRA + Unsloth) | 16 GB | Optimized for single GPU on-premise |
| Inference (4-bit quantized) | 8 GB | Suitable for production deployment |
| Inference (float16 merged) | 16 GB | Higher accuracy, more VRAM |
License
This project is licensed under the Apache 2.0 License. See LICENSE for details.
Citation
If you use Manara-3B in your research or applications, please cite:
@misc{manara3b2026,
title = {Manara-3B: A Sovereign Small Language Model for the Kuwaiti Banking Sector},
author = {Manus AI},
year = {2026},
howpublished = {\url{https://github.com/your-org/manara-3b}},
note = {Fine-tuned on Jais, optimized for CBK regulations, Sharia finance, and Boursa Kuwait standards.}
}
Acknowledgements
This project builds upon the foundational work of the Jais team at Inception (G42) and MBZUAI, whose open-source Arabic LLM made this work possible. The fine-tuning methodology draws from the Unsloth and TRL libraries.
- Downloads last month
- -