Configuration Parsing Warning:In config.json: "num_experts" must be a number
Bayan meaning judge (Gemma 4 E2B)
Scores whether an Arabic rewrite (a simplification) keeps the meaning of the original sentence. Built for the Bayan project as a fast, local replacement for an API LLM judge:
- reward / penalty while training the simplifier (score each generated rewrite),
- benchmarking simplification systems (meaning-preservation metric),
- data filtering (drop synthetic pairs that change the meaning).
It outputs four scores in 0-1 from a single forward pass:
| head | meaning |
|---|---|
same |
the meaning score: the rewrite keeps every fact, claim and qualifier, adds nothing, reverses nothing |
added |
the rewrite states something the original does not |
missing |
a fact, detail or qualifier of the original is missing |
contradict |
the rewrite contradicts the original (negation, number, roles, statement → question) |
Use same as the score. The other three explain why a pair fails.
This repository: bf16 (reference)
Full-precision weights (8.7 GB, of which ~4.7 GB are Gemma 4's per-layer embedding tables). Quantized variants:
-w8a8 (int8, fastest on Ampere+) and
-w4a16 (int4 weights, smallest).
Usage
# pip install transformers huggingface_hub safetensors (+ vllm for fast batch scoring)
from huggingface_hub import hf_hub_download
import importlib.util, sys
p = hf_hub_download("Congi-libya/bayan-meaning-judge-e2b", "judge.py")
spec = importlib.util.spec_from_file_location("judge", p); judge = importlib.util.module_from_spec(spec); spec.loader.exec_module(judge)
j = judge.MeaningJudge("Congi-libya/bayan-meaning-judge-e2b", backend="vllm") # or backend="transformers"
src = "ذهب الولد إلى المدرسة صباحا لأنه تأخر أمس."
j.score([(src, "ذهب الولد إلى المدرسة صباحا."), # drops the reason
(src, "ذهب الولد صباحا إلى المدرسة، فقد تأخر يوم أمس."), # faithful paraphrase
(src, "لم يذهب الولد إلى المدرسة صباحا لأنه تأخر أمس.")]) # negated
# same 0.00 added 0.00 missing 1.00 contradict 0.00
# same 0.87 added 0.03 missing 0.07 contradict 0.03
# same 0.00 added 0.00 missing 0.00 contradict 1.00 (bf16, rounded)
How it works: the prompt (chat template, thinking off) asks the "same meaning" question; the model is a plain text-only
Gemma4ForCausalLM; the score is sigmoid(W · h + b) where h is the final hidden state of the last prompt token
and W, b are in judge_head/meaning_head.safetensors (4 × 1536). Any runtime that returns last-token hidden states
works (vLLM: runner="pooling", pooling_type="LAST", use_activation=False).
Batching caveat (transformers 5.17). Padded batches (left or right padding with an attention mask) return corrupted hidden states for some rows with Gemma 4.
judge.pytherefore batches only prompts of identical token length (no padding). vLLM is not affected.
Results
Frozen bench of 2,779 Arabic pairs (none of its sources are in the training data): DeepSeek-written adversarial rewrites (add / delete / negate / question), scripted defects, the team's human meaning labels (393 pairs: 52 changed / 341 same), and 1,366 real pipeline pairs labelled same / minor / changed. "Catch" = share of meaning-changing pairs rejected at the threshold that wrongly rejects 5% of faithful pairs.
| scorer | planted errors: catch | human labels: AUC / catch | real subtle errors: AUC / catch | pairs/s |
|---|---|---|---|---|
| Gemma 4 31B (teacher, zero-shot scoring mode) | 97.6% | .979 / 90.4% | .917 / 68.5% | 5.5 |
| Jev API (TypeSafe) | 93.1% | .969 / 84.6% | .870 / 64.4% | ~2 (API) |
| Qwen3.8 27B reasoning judge (previous pipeline) | 92.0% | .914 / 83.9% | .766 / 55.8% | — |
| Gemma 4 E2B zero-shot | 61.7% | .860 / 53.8% | .782 / 39.7% | 69 |
| mDeBERTa-xnli fine-tuned (both directions) | 76.0% | .936 / 80.8% | .832 / 56.2% | 232 |
| this judge, bf16 | 95.5% | .965 / 86.5% | .877 / 61.6% | 165 |
| this judge, W8A8 | 94.6% | .962 / 84.6% | .870 / 61.6% | 187 |
| this judge, W4A16 | 94.2% | .965 / 86.5% | .874 / 60.3% | 158 |
Speeds: one RTX A6000, vLLM (Gemma 4 31B and E2B zero-shot: 4 questions per pair). The judge reaches about 94% of Gemma 4 31B on average at ~30× its speed.
Training (three stages, all from google/gemma-4-E2B-it, LoRA r=16 on the language model)
Training pool: 125k (source, rewrite) pairs from the Bayan synthetic pipelines (DeepSeek- and Qwen-judged candidates), 7.5k SAMER professional rewrites (faithful), and ~10k scripted hard negatives.
- Epoch 1: Jev labels. 60k pairs, soft BCE on the (Yes − No) logit margin of the "same meaning" prompt, labels = Jev
same_meaning(planted rows 0/1). Bench: human AUC .954, planted catch 83%. Jev hedges near 0.5 on genuine Arabic paraphrases (it passes only 54% of SAMER human rewrites), so the student learned "reworded = suspicious". - Epoch 2: Gemma 4 31B + Jev. Gemma 4 31B labelled the same 41k judged pairs in scoring mode (P(yes) in one pass); labels = 0.7 Gemma + 0.3 Jev. Planted catch 83% → 91%, question flips caught 73% → 93%. Failure analysis: misses small deletions inside fluent rewrites, question flips (no such negatives), and falsely rejects faithful sentence-split rewrites.
- Epoch 3: four heads + targeted data. Added 12k certain-label rows (2–5-word phrase drops, adjective drops, statement → question flips, sentence-split positives mined from Gemma-confirmed rewrites); Gemma 4 31B answered the added / missing / contradict questions for 20k judged pairs (labels 0.85 Gemma + 0.15 Jev); a 4-output linear head on the last hidden state (initialised from the LM's Yes − No direction). 52k rows, 1 epoch. Missed changes 120 → 86, missed question flips 19 → 0, missed small deletions 46 → 28, false alarms 99 → 76.
Known issue for a future retrain: training batches used left padding, which (per the caveat above) corrupts a small share of rows; a retrain without padding may score slightly higher.
Training data and acknowledgements
Trained on synthetic simplification pairs from the Bayan pipelines (built on the BAREC corpus, CC BY-SA 4.0), scripted hard negatives, and rewrites from The SAMER Arabic Text Simplification Corpus (New York University Abu Dhabi), used with the corpus authors' permission; the weights are shared for non-commercial research only (see Licence). Teacher labels: Gemma 4 31B (Apache-2.0) and TypeSafe Jev. The training data itself is not redistributed.
If you use this model, please also cite SAMER:
@inproceedings{alhafni-etal-2024-samer,
title = {The {SAMER} {A}rabic Text Simplification Corpus},
author = {Alhafni, Bashar and Hazim, Reem and Liberato, Juan Pi{\~n}eros and Al Khalil, Muhamed and Habash, Nizar},
booktitle = {Proceedings of LREC-COLING 2024},
year = {2024}
}
Licence
CC BY-NC 4.0 (non-commercial). The base model, Gemma 4 E2B, is Apache-2.0, but this judge was trained partly on The SAMER Arabic Text Simplification Corpus, which its authors permit us to share for non-commercial research. Commercial use is not allowed. If you need other terms, contact the SAMER authors (New York University Abu Dhabi).
Limitations
- Catches ~62% of subtle real meaning changes at 5% false rejects (Gemma 4 31B: 68.5%). Remaining misses: negations inside long sentences and subtle reference / generalisation errors, which the teachers also miss.
- Scores are not calibrated probabilities (training labels are soft teacher blends). Pick thresholds on your own data, or use the score continuously as a reward.
- Modern Standard Arabic, sentence-level; long multi-paragraph inputs are untested.
- The human labels come from one annotator (393 pairs); the "real subtle" slice is labelled by Claude.
- Teachers: Gemma 4 31B generated part of the training candidates (no measurable self-preference found: it passes 87% of both its own and other generators' pairs) and Jev (English-first). Their biases carry over.
- Downloads last month
- 16