bev-decider-0.4B

bev-decider is a 0.4B-parameter System One decision model. It reads a state (text or JSON) and typed questions about it, and returns calibrated probabilities in a single forward pass, with no text generation. The questions use the same format as TypeSafe's Jev:

  • choice: pick one of several named options.
  • noul: yes/no, returned as P(yes).
  • score: an ordinal level, returned as the expected value over the levels.

Highlights

  • Tiny. Roughly 0.4B parameters: 20 layers of Qwen3-0.6B plus a small decision head. It runs on a laptop CPU or Apple Silicon. Every question is answered in one prefill pass, with no generation.
  • Choice-order invariant by construction. Every option starts at the same position id and is read in parallel, so shuffling the options cannot change the answer. It has no bias toward the first or last option, unlike prompted LLMs. Inside the backbone, the attention mask lets each option's tokens see only the question and their own earlier tokens, never another option. Each option's embedding is then passed through a small, newly trained self-attention head, which is where the options are compared with each other. Because that head has no position information either, the whole model is order invariant, not just the backbone. This is exact in fp32: over 960 random shuffles of 3–12 options, no probability moved by more than 1e-5. On GPU or Apple Silicon the default bf16 inference adds rounding noise (about 0.001 typical), which can only flip near-ties.
  • Close to Jev on typed decisions. On 5,000 held-out questions from avbiswas/bev-decision, it scores 74.7% against Jev 1.13's 78.0%. It is ahead of Jev on ordinal score questions (63.7% against 59.3%) and close on yes/no (84.6% against 85.8%). It is 18–24 points ahead of Kev 0.8B and Laya.
  • Strong on routing, policy and rule reasoning: 89.7% on support and intent routing and 81.9% on policy and rule reasoning.
  • The best open model under 0.5B on public benchmarks. It beats Laya on JevBench (65.8% against 58.4%) and sysone-bench (70.8% against 68.6%). On entailment (mnli) it beats both Kev 0.8B and Laya (62.5% against 48.3% and 55.8%).
  • Calibrated. Expected calibration error is 0.065 on sysone-bench, so its probabilities can be thresholded directly.

Architecture

This is a single self-contained model. model.safetensors holds the whole network in one file, and nothing else is downloaded.

Backbone The first 20 of 28 layers of Qwen/Qwen3-0.6B, with the fine-tuned LoRA (r=8 on q/k/v/o_proj of layers 8–19) merged into the weights
Decision head 2-layer attention head, 512-dim, with a task-type embedding
Parameters ≈ 0.4B (0.31B in the transformer layers, 0.16B token embeddings, 7.4M head)
Weights model.safetensors (0.97 GB, bf16 backbone + fp32 head). Also here: backbone_config.json (Qwen3 config, 20 layers), config.json, and Qwen's tokenizer files
Context Trained with states up to 1,024 tokens; the defaults and benchmarks use a 2,048-token limit. Options are limited to 64 tokens each

The head is a custom module, so the model is loaded with the bev-decider package.

Usage

Install the bev-decider package. It downloads this model (about 1 GB) on first use.

pip install bev-decider            # library
pip install "bev-decider[serve]"   # + local /v1/systemone server
from bev_decider import load

decider = load()  # avbiswas/bev-decider-0.4B

state = {
    "message": (
        "URGENT: you charged my card twice this month. "
        "Refund the duplicate within 24 hours or I'm disputing it with my bank."
    )
}

questions = {
    "intent": {
        "type": "choice",
        "instructions": "What does the customer want?",
        "criteria": {
            "refund": "money returned or a duplicate charge reversed",
            "technical_help": "a bug, outage or integration problem",
            "cancellation": "wants to cancel or downgrade",
        },
    },
    "urgent": {
        "type": "noul",
        "instructions": "Does the message communicate time pressure or a deadline?",
    },
    "anger": {
        "type": "score",
        "instructions": "How angry is the customer?",
        "criteria": ["calm", "mildly annoyed", "frustrated", "furious"],
    },
}

answers = decider.decide(state, questions)

Output:

{
  "intent": {
    "type": "choice",
    "choice": "refund",
    "probabilities": {"refund": 1.0, "technical_help": 0.0, "cancellation": 0.0}
  },
  "urgent": {"type": "noul", "noul": 0.99},
  "anger": {
    "type": "score",
    "score": 1.90,
    "probabilities": {"0": 0.09, "1": 0.10, "2": 0.65, "3": 0.17}
  }
}

Run it as a Jev-compatible API:

bev-decider serve --port 8008
curl -s localhost:8008/v1/systemone -H 'content-type: application/json' -d '{
  "state": "Order 1182 arrived with a cracked screen.",
  "questions": {
    "damaged": {"type": "noul", "instructions": "Was the item damaged on arrival?"}
  }
}'

Code, CLI and tests: https://github.com/avbiswas/bev-decider

Evaluation

Held-out typed decisions (5,000 questions)

These are 5,000 questions from the avbiswas/bev-decision test split (revision 687e28f), 2,617 states in all. None of these states appears in the training data. Each model received the same states and questions in the /v1/systemone format. The split has the same distribution as our training data, so read these numbers alongside the external benchmarks below.

Model Size All Choice Yes/no Score
Jev 1.13.0 (closed API) ? 78.0 77.8 85.8 59.3
bev-decider-0.4B 0.4B 74.7 69.8 84.6 63.7
Kev 0.8B 0.8B 56.6 54.7 65.1 40.3
Laya (English) 0.4B 50.4 44.3 66.8 26.5
Domain Jev 1.13 bev-decider-0.4B Kev 0.8B Laya
Support and intent routing 91.6 89.7 80.0 65.2
Policy and rule reasoning 87.2 81.9 49.8 38.8
Multi-hop reading and evidence 78.2 70.6 45.8 36.9
Answer and solution verification 75.1 54.1 32.6 38.1
Temporal and unit reasoning 66.7 54.0 54.0 40.2

sysone-bench v2 (evaluation split, 1,240 decisions)

sysone-bench evaluation split. Jev and Laya are the published numbers; we ran Kev and bev-decider locally with the same grading rules.

Suite Jev 1.13 Kev 0.8B bev-decider-0.4B Laya 0.3.11
All 90.7 76.6 70.8 68.6
agnews 98.8 90.0 82.5 85.0
banking77 (12 intents) 94.8 89.6 77.1 81.3
emotion 84.4 68.8 63.5 65.6
guardrails 100 90.6 85.4 76.0
mnli 86.7 48.3 62.5 55.8
moderation 94.4 88.2 73.6 75.7
multilingual intent 100 82.5 69.2 45.0
sst5 (score) 65.0 41.7 30.8 33.3
triage 93.2 87.0 87.0 87.5

Expected calibration error (top-label, 10 bins): 0.065.

JevBench (public, 231 tasks)

The other rows are JevBench's published results.

Model Size All Easy (48) Standard (72) Hard (111)
Jev 1.13.0 ? 86.6 100 98.6 73.0
system-one-open (Gemma 4 E2B LoRA) E2B 73.2 100 93.1 48.6
Kev 0.6B (research preview) 0.6B 66.7 100 80.6 43.2
bev-decider-0.4B 0.4B 65.8 97.9 73.6 46.8
jeff (GLiFormer) 0.4B 62.8 100 75.0 38.7
Laya (ModernBERT-large) 0.4B 58.4 95.8 69.4 35.1
open-jev-deberta-v3-large 0.4B 52.4 100 43.1 37.8
Kev 0.5B 0.5B 49.4 95.8 48.6 29.7

bev-decider scores higher on the hard tier than every open model its size and larger, except system-one-open. The results use the noul criteria as options and a 2,048-token state limit.

Known weaknesses

  • Date and arithmetic comparisons (for example "valid through March 14" against an order on March 15, or price plus tax against a cap) are its least reliable area. Compute such values before asking.
  • Stated exceptions to a general rule are sometimes overlooked.
  • 5-level sentiment (sst5) is weaker than binary or categorical questions.
  • Language and length: it was trained mostly on English, with states up to 1,024 tokens.

Training

The adapter and head were trained with cross-entropy on about 940K typed questions. Sources include rule-conditioned decisions, routing and intent, procedural reasoning, extraction, NLI-style reading, counterfactual pairs, and a range of public classification sets. The 12-layer adapter was initialised from an 8-layer run (the extra layers started with LoRA B = 0) and trained for 12,000 steps at batch size 64. No benchmark test items were used in training; the only exact state overlap found was 1 of 1,030 sysone-bench evaluation texts.

License

The model is CC-BY-NC-4.0 (non-commercial), because the training data includes sources with non-commercial terms. It contains weights derived from Qwen3-0.6B, which is Apache-2.0; Qwen's license is included as LICENSE-Qwen.

Downloads last month
148
Safetensors
Model size
0.5B params
Tensor type
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for avbiswas/bev-decider-0.4B

Finetuned
Qwen/Qwen3-0.6B
Finetuned
(1336)
this model

Dataset used to train avbiswas/bev-decider-0.4B