jeb · جِب

Arabic-only typed decision model. Give it a state (Arabic text) and typed questions; it returns typed answers with calibrated probabilities in a single forward pass (~9 ms). It never generates text, so there is nothing to parse and nothing to hallucinate.

جِب is Saudi dialect for "bring it — fast." That is the whole design goal.

The jeb family. Same architecture throughout; pick the checkpoint that matches your workload:

Checkpoint Backbone Encoder Params Context Best at
IJyad/jeb (this repo) MARBERTv2 178M 512 Arabic intent routing, topic, sentiment, NLI
IJyad/jeb-typed-decisions MARBERTv2 178M 512 invoice · security · support · agent-trace workflows (0.8752)
IJyad/jeb-onnx MARBERTv2 178M 512 ONNX Runtime — CPU, no PyTorch

Why Arabic-only

General-purpose decision models encode Arabic badly. Measured on identical Arabic text:

Encoder MSA Gulf Egyptian tokens vs jeb
ModernBERT-large (Laya EN) 55 59 54 3.2x worse
mmBERT-base (Laya multilingual) 31 32 31 1.8x worse
MARBERTv2 (jeb) 17 22 17 —

ModernBERT has no Arabic vocabulary at all — it shreds each word into raw UTF-8 bytes. A specialist spends its whole capacity on one language; that is where the win comes from.

Benchmarks — identical Arabic tests, 400 cases per task

Byte-identical questions, fixed seed, same prompts for every model.

Task jeb Jev 1.13.0 (live API) laya-multilingual laya (English)
MASSIVE intent (20 options) 0.8650 0.8075 0.3475 0.1250
XNLI-ar 0.7375 0.7450 0.6450 0.4200
SANAD topic (7 labels) 0.9450 0.9325 0.8475 0.2175
Sentiment 0.9475 0.7475 0.7000 0.5700
Overall 0.8738 0.8081 0.6350 0.3331
ECE (lower better) 0.0324 — 0.0810 —
p50 latency 9.3 ms 988 ms 32.8 ms —
Cost $0 self-hosted $0.042 / Mtok $0 $0

jeb beats Jev 1.13.0 — a closed commercial model — by +6.6 points overall on Arabic, at 106x lower latency, self-hosted and free. Jev was measured against its live API across two independent 1,600-case passes (agreeing to ±0.0006), not quoted from published figures.

Half the benchmark (MASSIVE intent, sentiment test split) is held out — never in training.

Quickstart

pip install torch transformers safetensors
huggingface-cli download IJyad/jeb --local-dir jeb
import model                      # model.py ships in the repo
m = model.load('jeb')             # tokenizer + weights come from the repo itself

print(m.predict("أعلنت الشركة عن أرباح قياسية في الربع الثالث من هذا العام", {
    "topic": {"type": "choice",
              "instructions": "ما هو موضوع هذا المقال؟",
              "criteria": {"Tech": "تقنية", "Finance": "اقتصاد ومال", "Politics": "سياسة",
                           "Religion": "دين", "Medical": "طب وصحة",
                           "Culture": "ثقافة", "Sports": "رياضة"}}}))
# {'topic': {'answer': 'Finance', 'confidence': 0.978,
#            'probabilities': {'Finance': 0.982, 'Tech': 0.015, ...}}}

print(m.predict("صحيني الساعة سبعة الصبح", {
    "intent": {"type": "choice",
               "instructions": "ما هو قصد المستخدم من هذه العبارة؟",
               "criteria": {"مجموعة التنبيه": "مجموعة التنبيه",
                            "الاستعلام عن الطقس": "الاستعلام عن الطقس",
                            "تشغيل الموسيقى": "تشغيل الموسيقى"}}}))
# {'intent': {'answer': 'مجموعة التنبيه', 'confidence': 0.997, ...}}

print(m.predict("والله يا اخوي الخدمة زفت، صار لي اسبوع اتصل وما احد يرد", {
    "sentiment": {"type": "choice",
                  "instructions": "ما هو شعور كاتب النص؟",
                  "criteria": {"pos": "إيجابي", "neg": "سلبي"}}}))
# {'sentiment': {'answer': 'neg', 'confidence': 0.71, ...}}

The answer space is defined at request time — write different criteria and the model scores them, no retraining. Every option is scored at its own [MASK] marker and softmaxed within the question.

Scope: jeb is trained on intent routing, topic, sentiment and NLI. Questions far outside those families (e.g. bespoke CRM fields like churn risk or SLA urgency) will return low-confidence answers — the confidence score is doing its job. Fine-tune on your own labels for those.

Question types

type question answer
choice which of these options? the option, a probability per option, confidence
score rate against ordered levels expected level, a probability per level, confidence
noul is this proposition true? P(true)

Architecture

  • Encoder: MARBERTv2 (163M, 12 layers, 100k vocab, trained on ~1B Arabic tweets)
  • Head: 2 transformer layers + option-marker scorer, trained from scratch. 178M total.
  • Option markers: every option is scored at its own [MASK] token, then softmaxed over that question's options. The answer space is defined at request time — new schemas need no retraining.
  • Confidence: (n * peak - 1) / (n - 1) — normalized max-probability, 0 at chance, 1 at certainty.

Training

Supervised fine-tuning of the decision head and encoder against typed questions built from public Arabic datasets. The reward is plain cross-entropy over the option distribution; the option-marker layout means a question's answer space is data, not architecture.

source examples contributes
XNLI-ar + human-written Arabic NLI 260k inference, entailment
SANAD (7-way news topic) 40k topic classification
Arabic Sentiment Twitter Corpus 30k sentiment
MASSIVE-ar (60 intents, Arabic labels) 11.5k intent routing
ASTD + AJGT 11.5k dialect sentiment

Option counts during training span 2, 3, 4, 6, 8, 10, 12, 16 and 20, so the model generalises to label spaces wider than any single dataset provides.

Training runs

Four runs behind the root checkpoint, all measured on the same 1,600-case benchmark. Run 3 ships as jeb.pt; the others are recorded here because their trade-offs are informative.

run overall MASSIVE XNLI-ar topic sentiment ECE
3 (shipped) 0.8738 0.8650 0.7375 0.9450 0.9475 0.0324
4 0.8688 0.8525 0.7600 0.9275 0.9350 0.0283
5 (NLI-heavy) 0.8650 0.8300 0.7700 0.9200 0.9400 0.0316
6 (speech-DAPT) 0.8631 0.8275 0.7575 0.9250 0.9425 0.0228

Honest limits

  • Arabic XNLI is harder than English XNLI. jeb scores 0.7375 (0.7700 on the run-5 checkpoint); laya scores 0.8825 on English XNLI but only 0.6450 on Arabic XNLI — a 24-point drop on identical architecture. Jev scores 0.7450. Published CAMeLBERT scores 0.557. XNLI-ar is machine-translated and some items are not coherent Arabic. Judge Arabic NLI against Arabic baselines, not English ones.
  • Speech-domain pretraining did not help these tasks. Continued MLM on 159k spoken-Saudi podcast segments improved held-out spoken-Arabic perplexity 204.87 -> 62.63 (3.3x) but left downstream accuracy unchanged within ±0.007. The benchmarks are written text; none is speech. It did produce the best calibration of any run.
  • More NLI data hits a ceiling. The first 200k human-written Arabic NLI examples moved XNLI +0.0225; the next 360k moved it +0.010 and cost overall accuracy.
  • Arabic-only. Do not use it for other languages.
  • Fit temperature on your own data before trusting probabilities in production.
  • Keep choice questions under ~20 options.

Links

Apache 2.0 · Arabic-only by design.

Downloads last month
13
Safetensors
Model size
0.2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for IJyad/jeb

Finetunes
1 model
Quantizations
1 model

Spaces using IJyad/jeb 2