QARAR

An Arabic-first typed-decision model. Given one state (an Arabic message, record or document) and a typed question, QARAR returns a calibrated probability distribution over the options supplied with the question — it never generates text.

Three question types are supported:

type meaning options
choice pick one option your named options
score place the state on an ordered scale your ordered levels
noul yes/no question about the state false, true

The model also returns a confidence and an abstention flag, so a caller can trade coverage for precision.

Architecture

QARAR is a single self-contained model: a multilingual encoder and a lightweight typed-decision head, released together as one architecture (model.safetensors).

  • Encoder: silma-ai/silma-embedding-matryoshka-v0.1 (768-d, 12 layers).
  • Decision head: projects the state, question and candidate definitions; four decision blocks attend from the question to the state; candidate/level definitions are encoded by the same encoder and scored against the question.
  • The state is encoded once; questions are independent (adding or reordering questions cannot change an answer).
  • Calibration: per-(type × option-count) temperature.
  • Abstention: per-type threshold.

Usage

from qarar import load            # or: from nomeda import load
m = load("nomeda-lab/qarar")

out = m.predict(
    state="أنا اتخصم مني مرتين على فاتورة مارس، عايز أرجّع الفلوس النهاردة.",
    questions={
        "intent": {
            "type": "choice",
            "instructions": "ماذا يريد العميل؟",
            "criteria": {
                "refund": "يريد استرداد المال",
                "cancel": "يريد إلغاء الطلب",
                "info":   "يطلب معلومات",
            },
        },
        "needs_human": {
            "type": "noul",
            "instructions": "هل يحتاج الأمر تدخل موظف؟",
            "criteria": {"true": "يحتاج تدخلًا بشريًا", "false": "يمكن التعامل آليًا"},
        },
    },
)
print(out["answers"]["intent"]["choice"], out["answers"]["intent"]["answer_confidence"])

Evaluation

Measured on the public Arabic Decision Benchmark (nomeda-lab/arabic-decision-benchmark, 5,635 questions / 22 tasks), against two external typed-decision systems.

slice n QARAR JEV-1.13 Laya
Native-Arabic tasks 3,474 0.724 / 0.671 0.696 / 0.591 0.454 / 0.386
Workflow tasks 2,161 0.626 / 0.449 0.875 / 0.747 0.426 / 0.230
Whole benchmark 5,635 0.687 / 0.560 0.765 / 0.669 0.443 / 0.308

(accuracy / macro-F1)

dimension leader values
Calibration (ECE, lower better) QARAR 0.035 · 0.072 · 0.337
Selective precision @ 20% coverage QARAR 0.955 · 0.898 · 0.570
Latency (CPU-matched) QARAR 13.9 · 10.0 q/s (QARAR · Laya)

QARAR is the best-calibrated system and gives the highest selective precision at low coverage (monotone risk–coverage); it leads on native Arabic and is self-hosted. JEV leads on the workflow tasks and on whole-benchmark accuracy.

Per-task accuracy

task n QARAR JEV-1.13 Laya
massive_ar / scenario 484 0.870 0.758 0.498
egy hate speech / hate 326 0.862 0.828 0.537
arbanking77 / intent 337 0.816 0.807 0.457
massive_ar / intent 410 0.780 0.829 0.571
egy fake reviews / sentiment 279 0.746 0.674 0.355
egy fake reviews / rating 270 0.685 0.722 0.156
arsarcasm / sentiment 169 0.686 0.639 0.331
egy fake reviews / authenticity 470 0.670 0.491 0.487
arsarcasm / sarcasm 250 0.628 0.672 0.556
shein reviews / rating 76 0.539 0.263 0.145
stance 403 0.489 0.643 0.489
openjev / mailroom 233 0.923 0.991 0.236
openjev / sponsor segment 225 0.280 0.964 0.267
openjev / phone extraction 57 0.825 0.965 0.789
openjev / email selection 165 0.897 0.945 0.552
openjev / silent failure 240 0.504 0.921 0.479
openjev / entity alignment 227 0.687 0.872 0.643
openjev / ir decision 233 0.597 0.850 0.189
openjev / amount extraction 164 0.817 0.848 0.598
openjev / citation control 216 0.329 0.792 0.375
openjev / browser drone 181 0.580 0.768 0.475
openjev / context retention 220 0.700 0.750 0.450

Limitations

  • Closed-label. QARAR selects among the options you supply; it degrades on tasks and labels outside its training coverage. It is not a zero-shot label-invention system.
  • Ordinal score decisions are the weakest type.
  • It does not generate text and does not execute tools.

License

Apache-2.0. Encoder: silma-ai/silma-embedding-matryoshka-v0.1 (Apache-2.0). See LICENSES.yaml.

Downloads last month
9
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for nomeda-lab/qarar