QARAR
An Arabic-first typed-decision model. Given one state (an Arabic message, record or document) and a typed question, QARAR returns a calibrated probability distribution over the options supplied with the question — it never generates text.
Three question types are supported:
| type | meaning | options |
|---|---|---|
choice |
pick one option | your named options |
score |
place the state on an ordered scale | your ordered levels |
noul |
yes/no question about the state | false, true |
The model also returns a confidence and an abstention flag, so a caller can trade coverage for precision.
Architecture
QARAR is a single self-contained model: a multilingual encoder and a lightweight
typed-decision head, released together as one architecture (model.safetensors).
- Encoder:
silma-ai/silma-embedding-matryoshka-v0.1(768-d, 12 layers). - Decision head: projects the state, question and candidate definitions; four decision blocks attend from the question to the state; candidate/level definitions are encoded by the same encoder and scored against the question.
- The state is encoded once; questions are independent (adding or reordering questions cannot change an answer).
- Calibration: per-
(type × option-count)temperature. - Abstention: per-type threshold.
Usage
from qarar import load # or: from nomeda import load
m = load("nomeda-lab/qarar")
out = m.predict(
state="أنا اتخصم مني مرتين على فاتورة مارس، عايز أرجّع الفلوس النهاردة.",
questions={
"intent": {
"type": "choice",
"instructions": "ماذا يريد العميل؟",
"criteria": {
"refund": "يريد استرداد المال",
"cancel": "يريد إلغاء الطلب",
"info": "يطلب معلومات",
},
},
"needs_human": {
"type": "noul",
"instructions": "هل يحتاج الأمر تدخل موظف؟",
"criteria": {"true": "يحتاج تدخلًا بشريًا", "false": "يمكن التعامل آليًا"},
},
},
)
print(out["answers"]["intent"]["choice"], out["answers"]["intent"]["answer_confidence"])
Evaluation
Measured on the public Arabic Decision Benchmark
(nomeda-lab/arabic-decision-benchmark, 5,635 questions / 22 tasks), against two
external typed-decision systems.
| slice | n | QARAR | JEV-1.13 | Laya |
|---|---|---|---|---|
| Native-Arabic tasks | 3,474 | 0.724 / 0.671 | 0.696 / 0.591 | 0.454 / 0.386 |
| Workflow tasks | 2,161 | 0.626 / 0.449 | 0.875 / 0.747 | 0.426 / 0.230 |
| Whole benchmark | 5,635 | 0.687 / 0.560 | 0.765 / 0.669 | 0.443 / 0.308 |
(accuracy / macro-F1)
| dimension | leader | values |
|---|---|---|
| Calibration (ECE, lower better) | QARAR | 0.035 · 0.072 · 0.337 |
| Selective precision @ 20% coverage | QARAR | 0.955 · 0.898 · 0.570 |
| Latency (CPU-matched) | QARAR | 13.9 · 10.0 q/s (QARAR · Laya) |
QARAR is the best-calibrated system and gives the highest selective precision at low coverage (monotone risk–coverage); it leads on native Arabic and is self-hosted. JEV leads on the workflow tasks and on whole-benchmark accuracy.
Per-task accuracy
| task | n | QARAR | JEV-1.13 | Laya |
|---|---|---|---|---|
| massive_ar / scenario | 484 | 0.870 | 0.758 | 0.498 |
| egy hate speech / hate | 326 | 0.862 | 0.828 | 0.537 |
| arbanking77 / intent | 337 | 0.816 | 0.807 | 0.457 |
| massive_ar / intent | 410 | 0.780 | 0.829 | 0.571 |
| egy fake reviews / sentiment | 279 | 0.746 | 0.674 | 0.355 |
| egy fake reviews / rating | 270 | 0.685 | 0.722 | 0.156 |
| arsarcasm / sentiment | 169 | 0.686 | 0.639 | 0.331 |
| egy fake reviews / authenticity | 470 | 0.670 | 0.491 | 0.487 |
| arsarcasm / sarcasm | 250 | 0.628 | 0.672 | 0.556 |
| shein reviews / rating | 76 | 0.539 | 0.263 | 0.145 |
| stance | 403 | 0.489 | 0.643 | 0.489 |
| openjev / mailroom | 233 | 0.923 | 0.991 | 0.236 |
| openjev / sponsor segment | 225 | 0.280 | 0.964 | 0.267 |
| openjev / phone extraction | 57 | 0.825 | 0.965 | 0.789 |
| openjev / email selection | 165 | 0.897 | 0.945 | 0.552 |
| openjev / silent failure | 240 | 0.504 | 0.921 | 0.479 |
| openjev / entity alignment | 227 | 0.687 | 0.872 | 0.643 |
| openjev / ir decision | 233 | 0.597 | 0.850 | 0.189 |
| openjev / amount extraction | 164 | 0.817 | 0.848 | 0.598 |
| openjev / citation control | 216 | 0.329 | 0.792 | 0.375 |
| openjev / browser drone | 181 | 0.580 | 0.768 | 0.475 |
| openjev / context retention | 220 | 0.700 | 0.750 | 0.450 |
Limitations
- Closed-label. QARAR selects among the options you supply; it degrades on tasks and labels outside its training coverage. It is not a zero-shot label-invention system.
- Ordinal
scoredecisions are the weakest type. - It does not generate text and does not execute tools.
License
Apache-2.0. Encoder: silma-ai/silma-embedding-matryoshka-v0.1 (Apache-2.0).
See LICENSES.yaml.
- Downloads last month
- 9
Model tree for nomeda-lab/qarar
Base model
aubmindlab/bert-base-arabertv02