--- license: apache-2.0 base_model: Qwen/Qwen3-0.6B-Base language: - tr - en tags: - text-classification - decision-making - triage - moderation - turkish pipeline_tag: text-classification --- ![ARBITER banner](./arbiter-banner.svg) [![held-out](https://img.shields.io/badge/held--out_macro-0.844-blue)](./README.md) [![params](https://img.shields.io/badge/params-0.6B-violet)](./README.md) [![langs](https://img.shields.io/badge/langs-TR_%2B_EN-yellow)](./README.md) [![license](https://img.shields.io/badge/license-Apache--2.0-green)](./README.md) # ARBITER-general — 0.6B single-pass decision layer **State in, typed decisions out. One forward pass, zero generated tokens.** ARBITER is a fine-tune of `Qwen/Qwen3-0.6B-Base` that answers structured questions — `choice` (up to 52 options), `score` (ordered levels), `noul` (yes/no) — in a single forward pass, in Turkish or English. It never generates text: option-letter logits (`A`–`Z`, `a`–`z`) are read at a fixed `Answer: (` slot and soft-maxed into calibrated probabilities with a confidence score. ```python from arbiter import Arbiter arb = Arbiter("Utiric/arbiter-general") # a local directory path also works out = arb.decide( state="My package never arrived, I want a refund.", questions=[ {"id": "intent", "type": "choice", "instructions": "What does the customer want?", "options": {"Refund": "wants money back", "Whereabouts": "asks where the package is", "Cancel": "wants to cancel the subscription", "Greeting": "just saying hello"}}, {"id": "urgent", "type": "noul", "instructions": "Does this need immediate action?", "criteria": {"true": "angry, threatening, or time-critical", "false": "routine request"}}, {"id": "tone", "type": "score", "instructions": "Customer anger?", "levels": ["calm", "annoyed", "angry", "furious"]}, ], ) # {"intent": {"choice": "Refund", "probabilities": {...}, "confidence": 0.61}, # "urgent": {"noul": 0.12, "confidence": 0.76, ...}, # "tone": {"level": "2", "score": 2.1, ...}} ``` Turkish works the same way: ```python out = arb.decide( state="Kargo iki günde geldi, ürün fotoğraftaki gibi. Teşekkürler!", questions=[{"id": "sentiment", "type": "choice", "instructions": "What is the overall sentiment polarity expressed in this text?", "options": {"Positive": "satisfaction, praise, thanks", "Neutral": "factual objective information", "Negative": "complaint, anger, disappointment"}}], ) # {"sentiment": {"choice": "Positive", "probabilities": {...}}} ``` > **No pip package (yet).** Copy `arbiter.py` from this repo (~200 lines, only > `torch` + `transformers`) next to your code. No training stack, no server, no API > key needed. The file mirrors the exact training prompt format — do not rephrase > the prompt template inside it. ## Requirements - Python 3.10+, `torch`, `transformers`. That is the whole dependency list. - Weights are 1.2 GB (bf16). Comfortable in ~2 GB of VRAM; CPU works too (float32, slower). - No server, no API key, no Inference Providers deployment: you run `arbiter.py` wherever your code already runs. Measured latency (T4 GPU, September 2026, single questions, temperature 1.6): | Situation | Latency | |---|---| | first call (warmup) | ~1 s | | steady state, per question | 60–90 ms | Latency grows with prompt length (prefill cost) and is flat in the number of questions per pass: asking 6 questions at once costs barely more than asking 1. --- ## How it works ![How it works](./arbiter-how.svg) 1. State and question are rendered into one prompt wrapped in `` tags with `Question (choice|yes/no|score):` labels and `(A) option — description` lines, ending in `Answer: (`. 2. The model runs **once**. The hidden state at the slot position is projected through the LM-head rows of the 52 letter tokens only. 3. Softmax over the valid options gives the distribution. Temperature is a **single global value (1.6)**, fit post-hoc on held-out data and shipped in `s1_config.json` — not per-task, not per-slice. `confidence` = top-1 minus top-2 probability gap. 4. `noul` returns P(yes); `score` returns the argmax level plus the expectation over levels. Latency scales with prompt length only (prefill cost) — there is no per-token generation loop. ## Getting good answers (read this before prompting) ARBITER is prompt-format sensitive: it answers from the exact prompt shape it trained on. Three habits decide whether you get 0.88 or 0.37: 1. **Frame the state like training data.** A short domain prefix plus the raw text works far better than a bare sentence, e.g. `[SUPPORT_TIER: return_desk] | Ticket #43998 Customer purchased a wireless headphones 15 days ago. Reason: 'item defective on arrival'...` We measured this directly on these weights: 8 hand-written free-form probes scored 3/8, while 96 verbatim held-out rows in training format scored 85/96. 2. **Describe every option.** `{"Refund": "wants money back"}` beats `["Refund", "Whereabouts", ...]`. The description is the signal; the name is just a key. 3. **Gate on confidence.** `confidence` is the top-1 minus top-2 probability gap: ```python a = out["intent"] if a["confidence"] < 0.2: escalate(state) # human review or a stronger judge else: route(a["choice"]) ``` ## When to use it (and when not) Good fit: support-ticket triage and routing, Discord/forum moderation pre-filtering, dialogue turn-taking, code-review triage, sentiment buckets — any pipeline step shaped "text in, one of N labels out" in Turkish or English. Not a fit: arithmetic or reasoning (`2 + 2` with options `[3, 4, 5, 22]` picks `22` — it matches strings, it does not compute), free-text generation, standalone medical or legal decisions, more than 52 options in one question, languages beyond TR/EN. ## API reference `Arbiter(path, device=None, temperature=None, max_state_tokens=2048)` | Argument | Meaning | |---|---| | `path` | HF repo id or local directory with weights | | `device` | `"cuda"` / `"cpu"` (auto-detected by default) | | `temperature` | softmax temperature; default read from `s1_config.json` (1.6) | | `max_state_tokens` | state truncation budget (70% head / 30% tail, marked) | `decide(state, questions)` — `state` is a string (or JSON-serializable object). Each question needs `id`, `type` (`choice` \| `score` \| `noul`), and: | Type | Fields | Returns | |---|---|---| | `choice` | `instructions`, `options` (list or `{name: description}`) | `choice`, `probabilities`, `confidence` | | `score` | `instructions`, `levels` (ordered list) | `level`, `score` (expectation), `confidence`, `probabilities` | | `noul` | `instructions`, optional `criteria: {true, false}` | `noul` = P(yes), `confidence`, `probabilities` | Budgets: **2048** state tokens · **52** options per pass · Turkish + English tested. --- ## Evaluation — held-out (8 tasks × 500, full eval, temperature 1.6) Held-out rows never appeared in training (exact-dedup isolated). Macro **0.8438** vs chance 0.2958. | Slice | acc | ECE | Note | |---|---|---|---| | code_pr_review | 1.000 | 0.053 | | | customer_action | 0.796 | 0.073 | weakest link, under active work | | dialogue_turn_taking | 0.992 | 0.014 | | | irony_sarcasm | 0.620 | 0.117 | noisy slice — do not use standalone | | legal_contract_risk | 0.914 | 0.025 | routing signal, not legal advice | | medical_triage | 0.778 | 0.112 | routing signal, never autonomous diagnosis | | moderation_discord | 0.704 | 0.084 | tier-1 pre-filter; escalate on low confidence | | multilingual_sentiment | 0.946 | 0.028 | | | **macro** | **0.8438** | 0.063 | Brier 0.220 | ## Evaluation — public zero-shot (0-shot, slot-logit scoring, no generation) | Benchmark (split, n) | ARBITER | Random / Majority | |---|---|---| | Belebele tur_Latn (900) / eng_Latn (900) | 0.284 / 0.340 | 0.25 / 0.28 | | MMLU 6-subset avg (1,425) | 0.373 | 0.25 / 0.32 | | Banking77 sample-500 | 0.035 | 0.013 | | OffensEval-TR (3,528) | 0.797 | 0.50 / 0.80 | | TweetEval offensive+hate (3,830) | 0.632 | 0.50 / 0.61 | | XCOPA-TR (500) | 0.486 | 0.50 | | XNLI-TR (1k sample) | 0.407 | 0.33 / 0.35 | Read honestly: strong on trained domains, modest zero-shot — reported with random and majority baselines everywhere, as it should be. XCOPA is chance-level; Banking77's 77-way zero-shot is 3× random but low in absolute terms. Treat these as a floor for fine-tuning, not a ceiling. ## Evaluation — permutation consistency (position-bias check) Same question, three different option orders (50 held-out rows × 8 slices, seeded shuffles); a content-driven model must pick the same option every time. | Slice | consistent | |---|---| | code_pr_review | 1.000 | | customer_action | 0.980 | | dialogue_turn_taking | 0.940 | | irony_sarcasm | 0.880 | | legal_contract_risk | 0.980 | | medical_triage | 0.960 | | moderation_discord | 0.900 | | multilingual_sentiment | 0.960 | | **overall (400 rows)** | **0.950** | Position bias is dead: option-shuffle training (uniform Check 4 + per-epoch shuffle) holds, weakest on the already-noisy irony slice. Nothing above is "not yet measured" anymore. ## Training journey (held-out macro) | Stage | macro | What changed | |---|---|---| | v1 SFT | 0.497 | baseline; overfit on 1.5 epochs | | v2 + teacher relabel | 0.768 | 26k teacher-written descriptions | | v3 + moderation data | 0.787 | moderation slice relabeled | | v4 cherry-pick | 0.811 | kept gains, restored casualties | | DPO-1 | 0.822 | 16k preference pairs | | v5–v7 + new domains | 0.818 | legal/medical/multilingual siblings, irony-train contradiction found and dropped | | v8 SFT | 0.840 | customer siblings (+19 pts on customer); moderation dipped 0.750 → 0.626 | | **DPO-4 (this release)** | **0.844** | preference pairs recovered moderation to 0.704, kept the rest | Recipe: full fine-tune of the 0.6B backbone, CE + 0.3·Brier loss, 1 epoch, held-out-Brier early stopping, then Decision-DPO (β=1.0 + 0.1 CE anchor, LR 1e-5). ## Data card (summary) - ~130k training rows across 26 tasks; ~26k teacher-(re)labeled rows. - The teacher model (Ministral-14B) writes **descriptions only**; gold labels are deterministic (rules, dataset labels, star ratings). Teacher-written golds proved unreliable on irony and were reverted to dataset labels there. - Gates on every iteration: exact-dedup train↔held-out, test-set contamination scan (**0 overlaps** across 11.4k public-benchmark rows), option-name leakage scan on teacher-written descriptions. ## Limitations (please read before deploying) - A classifier head — not a reasoner, generator, or judge. No free text, ever. - No arithmetic or multi-step reasoning. `2 + 2` with options `[3, 4, 5, 22]` answers `22`: it matches surface strings, it does not compute. - Medical and legal outputs are triage signals for pipelines **with humans in the loop**. - Moderation output is a tier-1 signal; escalate to a stronger judge and a human on low confidence. - Turkish + English only. Customer-action and irony are the weakest slices. - Not hardened against adversarial prompt injection. ## FAQ **Why is temperature 1.6, above 1.0?** Decision-DPO sharpens the distributions (expected side effect). 1.6 is a single global value fit post-hoc on held-out data to re-calibrate them — not per-task, not per-slice. If you fine-tune or DPO further, refit it. **What does `confidence` mean?** Top-1 probability minus top-2. A 0.90 prediction with a 0.08 runner-up (gap 0.82) is a sure thing; 0.45 vs 0.44 (gap 0.01) is a coin flip. Threshold it. **More than 52 options?** Chunk them: score 40-option windows with a `none of the above` catcher, then re-score the finalists in one pass. `arbiter.py` covers the single-pass case (≤52); port the chunking from `s1/engine.py` for the rest. **Several questions at once?** Yes — pass a list. Every question gets its own answer slot in one forward pass (the quickstart answers intent + urgency + tone together). **Can I fine-tune it?** It is a stock Qwen3-0.6B-Base fine-tune (full, not LoRA), so any HF trainer works. Keep option shuffling on and early-stop on held-out Brier, not train loss — 1.5 epochs already overfit this recipe once (v1: 0.94 in-task, 0.50 held-out). **Why no Inference Providers / widget?** The hosted inference API generates text; this model never generates — it is read at a logit slot. A widget would return nonsense, so the model is deliberately left undeployed there. Run `arbiter.py` locally instead. ## Files | File | What | |---|---| | `model.safetensors` (+ `config.json`, `tokenizer.*`, `s1_config.json`) | weights, Qwen3-0.6B-Base format | | `arbiter.py` | standalone inference client — copy it, nothing to install | | `arbiter-banner.svg`, `arbiter-how.svg` | release artwork | | `README.md` | this card | ## License & citation Apache-2.0 (base-model license carried over). ```bibtex @misc{arbiter2026, title = {ARBITER-general: a 0.6B single-pass decision layer for TR+EN}, author = {Utiric}, year = {2026}, note = {Held-out macro 0.844 across 8 tasks; Qwen3-0.6B-Base fine-tune + Decision-DPO} } ```