bev-decider-0.4B
bev-decider is a 0.4B-parameter System One decision model. It reads a state (text or JSON) and typed questions about it, and returns calibrated probabilities in a single forward pass, with no text generation. The questions use the same format as TypeSafe's Jev:
choice: pick one of several named options.noul: yes/no, returned as P(yes).score: an ordinal level, returned as the expected value over the levels.
Highlights
- Tiny. Roughly 0.4B parameters: 20 layers of Qwen3-0.6B plus a small decision head. It runs on a laptop CPU or Apple Silicon. Every question is answered in one prefill pass, with no generation.
- Choice-order invariant by construction. Every option starts at the same position id and is read in parallel, so shuffling the options cannot change the answer. It has no bias toward the first or last option, unlike prompted LLMs. Inside the backbone, the attention mask lets each option's tokens see only the question and their own earlier tokens, never another option. Each option's embedding is then passed through a small, newly trained self-attention head, which is where the options are compared with each other. Because that head has no position information either, the whole model is order invariant, not just the backbone. This is exact in fp32: over 960 random shuffles of 3–12 options, no probability moved by more than 1e-5. On GPU or Apple Silicon the default bf16 inference adds rounding noise (about 0.001 typical), which can only flip near-ties.
- Close to Jev on typed decisions. On 5,000 held-out questions from
avbiswas/bev-decision, it scores 74.7% against Jev 1.13's 78.0%. It is ahead of Jev on ordinalscorequestions (63.7% against 59.3%) and close on yes/no (84.6% against 85.8%). It is 18–24 points ahead of Kev 0.8B and Laya. - Strong on routing, policy and rule reasoning: 89.7% on support and intent routing and 81.9% on policy and rule reasoning.
- The best open model under 0.5B on public benchmarks. It beats Laya on JevBench (65.8% against 58.4%) and sysone-bench (70.8% against 68.6%). On entailment (mnli) it beats both Kev 0.8B and Laya (62.5% against 48.3% and 55.8%).
- Calibrated. Expected calibration error is 0.065 on sysone-bench, so its probabilities can be thresholded directly.
Architecture
This is a single self-contained model. model.safetensors holds the whole network in one file, and nothing else is downloaded.
| Backbone | The first 20 of 28 layers of Qwen/Qwen3-0.6B, with the fine-tuned LoRA (r=8 on q/k/v/o_proj of layers 8–19) merged into the weights |
| Decision head | 2-layer attention head, 512-dim, with a task-type embedding |
| Parameters | ≈ 0.4B (0.31B in the transformer layers, 0.16B token embeddings, 7.4M head) |
| Weights | model.safetensors (0.97 GB, bf16 backbone + fp32 head). Also here: backbone_config.json (Qwen3 config, 20 layers), config.json, and Qwen's tokenizer files |
| Context | Trained with states up to 1,024 tokens; the defaults and benchmarks use a 2,048-token limit. Options are limited to 64 tokens each |
The head is a custom module, so the model is loaded with the bev-decider package.
Usage
Install the bev-decider package. It downloads this model (about 1 GB) on first use.
pip install bev-decider # library
pip install "bev-decider[serve]" # + local /v1/systemone server
from bev_decider import load
decider = load() # avbiswas/bev-decider-0.4B
state = {
"message": (
"URGENT: you charged my card twice this month. "
"Refund the duplicate within 24 hours or I'm disputing it with my bank."
)
}
questions = {
"intent": {
"type": "choice",
"instructions": "What does the customer want?",
"criteria": {
"refund": "money returned or a duplicate charge reversed",
"technical_help": "a bug, outage or integration problem",
"cancellation": "wants to cancel or downgrade",
},
},
"urgent": {
"type": "noul",
"instructions": "Does the message communicate time pressure or a deadline?",
},
"anger": {
"type": "score",
"instructions": "How angry is the customer?",
"criteria": ["calm", "mildly annoyed", "frustrated", "furious"],
},
}
answers = decider.decide(state, questions)
Output:
{
"intent": {
"type": "choice",
"choice": "refund",
"probabilities": {"refund": 1.0, "technical_help": 0.0, "cancellation": 0.0}
},
"urgent": {"type": "noul", "noul": 0.99},
"anger": {
"type": "score",
"score": 1.90,
"probabilities": {"0": 0.09, "1": 0.10, "2": 0.65, "3": 0.17}
}
}
Run it as a Jev-compatible API:
bev-decider serve --port 8008
curl -s localhost:8008/v1/systemone -H 'content-type: application/json' -d '{
"state": "Order 1182 arrived with a cracked screen.",
"questions": {
"damaged": {"type": "noul", "instructions": "Was the item damaged on arrival?"}
}
}'
Code, CLI and tests: https://github.com/avbiswas/bev-decider
Evaluation
Held-out typed decisions (5,000 questions)
These are 5,000 questions from the avbiswas/bev-decision test split (revision 687e28f), 2,617 states in all. None of these states appears in the training data. Each model received the same states and questions in the /v1/systemone format. The split has the same distribution as our training data, so read these numbers alongside the external benchmarks below.
| Model | Size | All | Choice | Yes/no | Score |
|---|---|---|---|---|---|
| Jev 1.13.0 (closed API) | ? | 78.0 | 77.8 | 85.8 | 59.3 |
| bev-decider-0.4B | 0.4B | 74.7 | 69.8 | 84.6 | 63.7 |
| Kev 0.8B | 0.8B | 56.6 | 54.7 | 65.1 | 40.3 |
| Laya (English) | 0.4B | 50.4 | 44.3 | 66.8 | 26.5 |
| Domain | Jev 1.13 | bev-decider-0.4B | Kev 0.8B | Laya |
|---|---|---|---|---|
| Support and intent routing | 91.6 | 89.7 | 80.0 | 65.2 |
| Policy and rule reasoning | 87.2 | 81.9 | 49.8 | 38.8 |
| Multi-hop reading and evidence | 78.2 | 70.6 | 45.8 | 36.9 |
| Answer and solution verification | 75.1 | 54.1 | 32.6 | 38.1 |
| Temporal and unit reasoning | 66.7 | 54.0 | 54.0 | 40.2 |
sysone-bench v2 (evaluation split, 1,240 decisions)
sysone-bench evaluation split. Jev and Laya are the published numbers; we ran Kev and bev-decider locally with the same grading rules.
| Suite | Jev 1.13 | Kev 0.8B | bev-decider-0.4B | Laya 0.3.11 |
|---|---|---|---|---|
| All | 90.7 | 76.6 | 70.8 | 68.6 |
| agnews | 98.8 | 90.0 | 82.5 | 85.0 |
| banking77 (12 intents) | 94.8 | 89.6 | 77.1 | 81.3 |
| emotion | 84.4 | 68.8 | 63.5 | 65.6 |
| guardrails | 100 | 90.6 | 85.4 | 76.0 |
| mnli | 86.7 | 48.3 | 62.5 | 55.8 |
| moderation | 94.4 | 88.2 | 73.6 | 75.7 |
| multilingual intent | 100 | 82.5 | 69.2 | 45.0 |
| sst5 (score) | 65.0 | 41.7 | 30.8 | 33.3 |
| triage | 93.2 | 87.0 | 87.0 | 87.5 |
Expected calibration error (top-label, 10 bins): 0.065.
JevBench (public, 231 tasks)
The other rows are JevBench's published results.
| Model | Size | All | Easy (48) | Standard (72) | Hard (111) |
|---|---|---|---|---|---|
| Jev 1.13.0 | ? | 86.6 | 100 | 98.6 | 73.0 |
| system-one-open (Gemma 4 E2B LoRA) | E2B | 73.2 | 100 | 93.1 | 48.6 |
| Kev 0.6B (research preview) | 0.6B | 66.7 | 100 | 80.6 | 43.2 |
| bev-decider-0.4B | 0.4B | 65.8 | 97.9 | 73.6 | 46.8 |
| jeff (GLiFormer) | 0.4B | 62.8 | 100 | 75.0 | 38.7 |
| Laya (ModernBERT-large) | 0.4B | 58.4 | 95.8 | 69.4 | 35.1 |
| open-jev-deberta-v3-large | 0.4B | 52.4 | 100 | 43.1 | 37.8 |
| Kev 0.5B | 0.5B | 49.4 | 95.8 | 48.6 | 29.7 |
bev-decider scores higher on the hard tier than every open model its size and larger, except system-one-open. The results use the noul criteria as options and a 2,048-token state limit.
Known weaknesses
- Date and arithmetic comparisons (for example "valid through March 14" against an order on March 15, or price plus tax against a cap) are its least reliable area. Compute such values before asking.
- Stated exceptions to a general rule are sometimes overlooked.
- 5-level sentiment (sst5) is weaker than binary or categorical questions.
- Language and length: it was trained mostly on English, with states up to 1,024 tokens.
Training
The adapter and head were trained with cross-entropy on about 940K typed questions. Sources include rule-conditioned decisions, routing and intent, procedural reasoning, extraction, NLI-style reading, counterfactual pairs, and a range of public classification sets. The 12-layer adapter was initialised from an 8-layer run (the extra layers started with LoRA B = 0) and trained for 12,000 steps at batch size 64. No benchmark test items were used in training; the only exact state overlap found was 1 of 1,030 sysone-bench evaluation texts.
License
The model is CC-BY-NC-4.0 (non-commercial), because the training data includes sources with non-commercial terms. It contains weights derived from Qwen3-0.6B, which is Apache-2.0; Qwen's license is included as LICENSE-Qwen.
- Downloads last month
- 148