Instructions to use IJyad/jeb with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use IJyad/jeb with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="IJyad/jeb")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("IJyad/jeb", device_map="auto") - Notebooks
- Google Colab
- Kaggle
jeb · جِب
Arabic-only typed decision model. Give it a state (Arabic text) and typed questions; it returns typed answers with calibrated probabilities in a single forward pass (~9 ms). It never generates text, so there is nothing to parse and nothing to hallucinate.
جِب is Saudi dialect for "bring it — fast." That is the whole design goal.
The jeb family. Same architecture throughout; pick the checkpoint that matches your workload:
| Checkpoint | Backbone Encoder | Params | Context | Best at |
|---|---|---|---|---|
IJyad/jeb (this repo) |
MARBERTv2 | 178M | 512 | Arabic intent routing, topic, sentiment, NLI |
IJyad/jeb-typed-decisions |
MARBERTv2 | 178M | 512 | invoice · security · support · agent-trace workflows (0.8752) |
IJyad/jeb-onnx |
MARBERTv2 | 178M | 512 | ONNX Runtime — CPU, no PyTorch |
Why Arabic-only
General-purpose decision models encode Arabic badly. Measured on identical Arabic text:
| Encoder | MSA | Gulf | Egyptian | tokens vs jeb |
|---|---|---|---|---|
| ModernBERT-large (Laya EN) | 55 | 59 | 54 | 3.2x worse |
| mmBERT-base (Laya multilingual) | 31 | 32 | 31 | 1.8x worse |
| MARBERTv2 (jeb) | 17 | 22 | 17 | — |
ModernBERT has no Arabic vocabulary at all — it shreds each word into raw UTF-8 bytes. A specialist spends its whole capacity on one language; that is where the win comes from.
Benchmarks — identical Arabic tests, 400 cases per task
Byte-identical questions, fixed seed, same prompts for every model.
| Task | jeb | Jev 1.13.0 (live API) | laya-multilingual | laya (English) |
|---|---|---|---|---|
| MASSIVE intent (20 options) | 0.8650 | 0.8075 | 0.3475 | 0.1250 |
| XNLI-ar | 0.7375 | 0.7450 | 0.6450 | 0.4200 |
| SANAD topic (7 labels) | 0.9450 | 0.9325 | 0.8475 | 0.2175 |
| Sentiment | 0.9475 | 0.7475 | 0.7000 | 0.5700 |
| Overall | 0.8738 | 0.8081 | 0.6350 | 0.3331 |
| ECE (lower better) | 0.0324 | — | 0.0810 | — |
| p50 latency | 9.3 ms | 988 ms | 32.8 ms | — |
| Cost | $0 self-hosted | $0.042 / Mtok | $0 | $0 |
jeb beats Jev 1.13.0 — a closed commercial model — by +6.6 points overall on Arabic, at 106x lower latency, self-hosted and free. Jev was measured against its live API across two independent 1,600-case passes (agreeing to ±0.0006), not quoted from published figures.
Half the benchmark (MASSIVE intent, sentiment test split) is held out — never in training.
Quickstart
pip install torch transformers safetensors
huggingface-cli download IJyad/jeb --local-dir jeb
import model # model.py ships in the repo
m = model.load('jeb') # tokenizer + weights come from the repo itself
print(m.predict("أعلنت الشركة عن أرباح قياسية في الربع الثالث من هذا العام", {
"topic": {"type": "choice",
"instructions": "ما هو موضوع هذا المقال؟",
"criteria": {"Tech": "تقنية", "Finance": "اقتصاد ومال", "Politics": "سياسة",
"Religion": "دين", "Medical": "طب وصحة",
"Culture": "ثقافة", "Sports": "رياضة"}}}))
# {'topic': {'answer': 'Finance', 'confidence': 0.978,
# 'probabilities': {'Finance': 0.982, 'Tech': 0.015, ...}}}
print(m.predict("صحيني الساعة سبعة الصبح", {
"intent": {"type": "choice",
"instructions": "ما هو قصد المستخدم من هذه العبارة؟",
"criteria": {"مجموعة التنبيه": "مجموعة التنبيه",
"الاستعلام عن الطقس": "الاستعلام عن الطقس",
"تشغيل الموسيقى": "تشغيل الموسيقى"}}}))
# {'intent': {'answer': 'مجموعة التنبيه', 'confidence': 0.997, ...}}
print(m.predict("والله يا اخوي الخدمة زفت، صار لي اسبوع اتصل وما احد يرد", {
"sentiment": {"type": "choice",
"instructions": "ما هو شعور كاتب النص؟",
"criteria": {"pos": "إيجابي", "neg": "سلبي"}}}))
# {'sentiment': {'answer': 'neg', 'confidence': 0.71, ...}}
The answer space is defined at request time — write different criteria and the model scores
them, no retraining. Every option is scored at its own [MASK] marker and softmaxed within the
question.
Scope: jeb is trained on intent routing, topic, sentiment and NLI. Questions far outside those families (e.g. bespoke CRM fields like churn risk or SLA urgency) will return low-confidence answers — the confidence score is doing its job. Fine-tune on your own labels for those.
Question types
| type | question | answer |
|---|---|---|
choice |
which of these options? | the option, a probability per option, confidence |
score |
rate against ordered levels | expected level, a probability per level, confidence |
noul |
is this proposition true? | P(true) |
Architecture
- Encoder: MARBERTv2 (163M, 12 layers, 100k vocab, trained on ~1B Arabic tweets)
- Head: 2 transformer layers + option-marker scorer, trained from scratch. 178M total.
- Option markers: every option is scored at its own
[MASK]token, then softmaxed over that question's options. The answer space is defined at request time — new schemas need no retraining. - Confidence:
(n * peak - 1) / (n - 1)— normalized max-probability, 0 at chance, 1 at certainty.
Training
Supervised fine-tuning of the decision head and encoder against typed questions built from public Arabic datasets. The reward is plain cross-entropy over the option distribution; the option-marker layout means a question's answer space is data, not architecture.
| source | examples | contributes |
|---|---|---|
| XNLI-ar + human-written Arabic NLI | 260k | inference, entailment |
| SANAD (7-way news topic) | 40k | topic classification |
| Arabic Sentiment Twitter Corpus | 30k | sentiment |
| MASSIVE-ar (60 intents, Arabic labels) | 11.5k | intent routing |
| ASTD + AJGT | 11.5k | dialect sentiment |
Option counts during training span 2, 3, 4, 6, 8, 10, 12, 16 and 20, so the model generalises to label spaces wider than any single dataset provides.
Training runs
Four runs behind the root checkpoint, all measured on the same 1,600-case benchmark. Run 3 ships
as jeb.pt; the others are recorded here because their trade-offs are informative.
| run | overall | MASSIVE | XNLI-ar | topic | sentiment | ECE |
|---|---|---|---|---|---|---|
| 3 (shipped) | 0.8738 | 0.8650 | 0.7375 | 0.9450 | 0.9475 | 0.0324 |
| 4 | 0.8688 | 0.8525 | 0.7600 | 0.9275 | 0.9350 | 0.0283 |
| 5 (NLI-heavy) | 0.8650 | 0.8300 | 0.7700 | 0.9200 | 0.9400 | 0.0316 |
| 6 (speech-DAPT) | 0.8631 | 0.8275 | 0.7575 | 0.9250 | 0.9425 | 0.0228 |
Honest limits
- Arabic XNLI is harder than English XNLI. jeb scores 0.7375 (0.7700 on the run-5 checkpoint); laya scores 0.8825 on English XNLI but only 0.6450 on Arabic XNLI — a 24-point drop on identical architecture. Jev scores 0.7450. Published CAMeLBERT scores 0.557. XNLI-ar is machine-translated and some items are not coherent Arabic. Judge Arabic NLI against Arabic baselines, not English ones.
- Speech-domain pretraining did not help these tasks. Continued MLM on 159k spoken-Saudi podcast segments improved held-out spoken-Arabic perplexity 204.87 -> 62.63 (3.3x) but left downstream accuracy unchanged within ±0.007. The benchmarks are written text; none is speech. It did produce the best calibration of any run.
- More NLI data hits a ceiling. The first 200k human-written Arabic NLI examples moved XNLI +0.0225; the next 360k moved it +0.010 and cost overall accuracy.
- Arabic-only. Do not use it for other languages.
- Fit temperature on your own data before trusting probabilities in production.
- Keep choice questions under ~20 options.
Links
- Playground: https://huggingface.co/spaces/IJyad/jeb
- Full benchmark report: BENCHMARKS.md
- Typed-decisions checkpoint: https://huggingface.co/IJyad/jeb-typed-decisions
- ONNX build: https://huggingface.co/IJyad/jeb-onnx
Apache 2.0 · Arabic-only by design.
- Downloads last month
- 13