MiniCPM5-2B-Jev

License GitHub Base Model Contract

MiniCPM5-2B-Jev is a high-speed, calibrated System 1 Decision Model built on OpenBMB's MiniCPM5-2B (base weights on Hugging Face). Code and evaluation suites are open-sourced at github.com/yuting-ai/minicpm5-2b-jev. Given a shared context document (state) and one or more typed decision specifications (choice, noul, score), it outputs a calibrated probability distribution over the candidate options for every question in a single forward pass — with zero autoregressive token generation.

It natively implements the /v1/systemone structured decision contract (compatible with Jev / OpenJev / Decider / Kev clients) and achieves state-of-the-art accuracy among all ≤ 2B parameter open-weight System 1 / Jev models, while outperforming several 4B–9B models on both JevBench and decision-v7.


Key Highlights

  • #1 Among ≤ 2B Models on JevBench: Achieves 78.79% Overall (97.92% Easy, 94.44% Standard, 60.36% Hard) on the 231-item public JevBench benchmark, surpassing decider-2b v11 (76.2%), Kev-4B (75.8%), open-alternative-jev 4B (74.0%), system-one-open Gemma-4-E2B (73.2%), system-one Qwen3-8B (71.9%), and Bespoke Nimble 9B (67.5%).
  • 86.41% Accuracy & 0.0307 ECE on decision-v7: Outperforms Kev-4B (85.9% Acc, 0.056 ECE) and hosted Jev (84.5% Acc on decision-v7 dev), and approaches Kev-9B (87.2%) and Kev-27B (87.0%) at 1/4th to 1/13th the parameter count.
  • 87.18% on jabr v1 & 79.75% on jabr v2: Strong zero-shot generalization across 49 diverse classification, policy routing, Boolean verification (noul), and ordinal rubric scoring (score) tasks.
  • Exact Multi-Question KV-Cache Branching: Encodes the shared state prefix once (use_cache=True) and evaluates M questions in parallel branches via DynamicCache.crop(), guaranteeing exact question isolation (zero cross-question leakage) and low latency (~260 ms/state on Apple Silicon MPS; <35 ms on CUDA).
  • Calibrated Native LetterReadoutHead: Projects the terminal Answer: ( hidden state onto 255 single-token option letters (A..Z, AA..) initialized from lm_head.weight, combined with class-balanced marginal prior debiasing (noul_bias, choice_k_bias, score_k_bias) and per-type temperatures (choice: 1.1314, noul: 1.2338, score: 0.4213).

Horizontal Benchmark Comparison

1. JevBench Public Benchmark Suite (195 Groups / 231 Decisions)

Evaluated on the official fstandhartinger/jevbench suite (easy: 48, original/standard: 72, hard: 111; total 231 decisions). Peer model scores are from the official JevBench / Decider / Kev / Laya published leaderboards on the exact same frozen items:

Model Base Backbone Params Easy (48) Standard (72) Hard (111) Overall (231) License / Weights
Jev 1.13.0 (Hosted API) Proprietary (TypeSafe) Closed 100.0% 98.6% 73.0% 86.6% Proprietary
decider-35b-a3b v1 Qwen3.5-35B-A3B (MoE) 35B (3B act.) 100.0% 97.2% 67.6% 83.5% Apache-2.0
decider-4b v2.1 Qwen3.5-4B 4.2B 100.0% 98.6% 64.9% 82.7% Apache-2.0
OpenJev (Hosted) DiffusionGemma 26B-A4B 26B (4B act.) 100.0% 97.2% 64.0% 81.8% Proprietary
SemIf Qwen3.5-4B 4.2B 100.0% 98.6% 61.3% 81.0% Open Weights
⭐ MiniCPM5-2B-Jev (Ours) openbmb/MiniCPM5-2B 2.0B 97.9% 94.4% 60.4% 78.8% Apache-2.0
decider-2b v11 Qwen3.5-2B 1.9B 100.0% 88.9% 57.7% 76.2% Apache-2.0
Kev-4B (r10) Qwen3.5-4B-Base 4.2B — — 54.1% 75.8% Apache-2.0
open-alternative-jev Qwen3.5-4B 4.2B 100.0% 83.3% 56.8% 74.0% Open Weights
system-one-open Gemma 4 E2B ~2B 100.0% 93.1% 48.6% 73.2% Open Weights
system-one Qwen3-8B 8.2B 100.0% 88.9% 48.6% 71.9% Open Weights
decider-2b v10 Qwen3.5-2B 1.9B 100.0% 88.9% 45.9% 70.6% Apache-2.0
Bespoke Nimble 9B Bespoke-Nimble-9B 9.0B 100.0% 93.1% 36.9% 67.5% Hosted
Kev-0.8B (r15) Qwen3.5-0.8B-Base 0.8B — — 36.0% 63.6% Apache-2.0
open-jev-deberta-v3-large DeBERTa-v3-Large 0.4B 100.0% 43.1% 37.8% 52.4% Apache-2.0

Key Takeaway: Within the ≤ 2B parameter class, MiniCPM5-2B-Jev achieves 78.79% overall accuracy and 60.36% on the challenging Hard tier (adversarial traps, prompt injections, multi-step arithmetic, negation, and complex policy edge cases), outperforming decider-2b v11 (+2.6 pp overall, +2.7 pp Hard), Kev-4B (+3.0 pp overall, +6.3 pp Hard), and Qwen3-8B system-one (+6.9 pp overall, +11.8 pp Hard).


2. decision-v7 Benchmark (10-Source Typed Decision Evaluation)

Comparison on the standard 10-source decision-v7 benchmark (ag_news, amazon, banking77, boolq, dbpedia_14, imdb, mnli, sst5, trec, yelp) across accuracy, Brier score, and Expected Calibration Error (ECE):

Model Base Backbone Params Accuracy (↑) Brier Score (↓) ECE (↓)
Kev-9B (as served, T=2.30) Qwen3.5-9B-Base 9.0B 87.2% — 0.042
Kev-27B (as served, T=1.38) Qwen3.8-27B 27.0B 87.0% — —
⭐ MiniCPM5-2B-Jev (Ours) openbmb/MiniCPM5-2B 2.0B 86.4% (86.41%) 0.214 (0.2143) 0.031 (0.0307)
Kev-4B (as served, T=1.89) Qwen3.5-4B-Base 4.2B 85.9% 0.256 0.056
Jev (TypeSafe Hosted) Proprietary Closed 84.5% — —
Kev-0.8B (as served, T=1.81) Qwen3.5-0.8B-Base 0.8B 77.1% 0.369 0.088
MiniCPM5-2B (Untuned Base, Step 0) openbmb/MiniCPM5-2B 2.0B 62.5% 0.693 0.272

Per-Source Breakdown on decision-v7 (MiniCPM5-2B-Jev):

  • boolq: 100.00% | dbpedia_14: 97.44% | ag_news: 94.87% | imdb: 94.87% | trec: 94.87% | mnli: 92.31% | banking77: 87.18% (77-way intent!) | sst5: 76.92% | yelp: 64.10% (5-way fine-grained star rating) | amazon: 61.54% (5-way star rating).
  • By Question Type: noul (Boolean verification): 94.97% | choice (multi-class up to 77 options): 87.10% | score (ordinal 0–10 / 5-level): 67.11%.

3. Comprehensive Multi-Suite Results

We evaluate MiniCPM5-2B-Jev across 7 zero-browser decision benchmarks totaling 2,042 states and 2,637 questions:

Benchmark Suite States / Questions Accuracy Choice Acc Noul Acc Score Acc Brier (↓) ECE (↓)
decision-v7 (10-Source Stratified) 300 / 390 86.41% 87.10% 94.97% 67.11% 0.2143 0.0307
jabr v1 (8 Tasks / 78 Cases) 78 / 78 87.18% 96.00% 76.92% 88.89% 0.1771 0.0277
jabr v2 (49 Tasks / 869 Cases) 869 / 869 79.75% 85.16% 79.49% 69.95% 0.2919 0.0663
JevBench (Overall Public Suite) 231 / 231 78.79% 78.42% 81.08% 72.22% 0.3157 0.0699
↳ JevBench Easy Tier 48 / 48 97.92% — — — — —
↳ JevBench Standard/Original Tier 72 / 72 94.44% — — — — —
↳ JevBench Hard Tier 111 / 111 60.36% — — — — —
hard-v1 (Held-Out Template 4) 350 / 543 60.41% 58.35% 70.43% 40.00% 0.5419 0.0603
DecisionBench Medium 80 / 293 60.07% 70.93% 80.52% 34.48% 0.6596 0.2647
DecisionBench Hard 80 / 293 44.71% 53.75% 70.79% 23.33% 0.8768 0.3711

Model Architecture

1. Native lm_head Letter Readout (LetterReadoutHead)

Instead of training a random projection head from scratch, MiniCPM5-2B-Jev extracts the exact rows of MiniCPM5-2B's pretrained lm_head.weight corresponding to 255 single-token option letters (A..Z for indices 0..25, plus two-letter tokens AA.. for indices 26..254, shape [255, 2304]).

Each question branch is formatted as:

Question: [{type}] {instructions}
(A) {option_0}
(B) {option_1}
...
Answer: (

At the final token ( of Answer: (, the model's hidden state h (dimension 2304) is projected onto the K active option letters: zk=s⋅(RMSNorm(h)⋅Wletter[k]⊤)+bletter[k]+bprior(t,k)Ttz_k = \frac{s \cdot (\text{RMSNorm}(h) \cdot W_{\text{letter}}[k]^\top) + b_{\text{letter}}[k] + b_{\text{prior}}(t, k)}{T_t} where:

  • W_letter (shape [255, 2304]) is initialized from lm_head.weight[letter_token_ids] and fine-tuned jointly with LoRA.
  • b_prior(t, k) is the class-balanced marginal prior debiasing offset (noul_bias, choice_k_bias, score_k_bias), which centers marginal log-odds on the validation set to eliminate yes/no or position bias without altering the training loss.
  • T_t is the per-type temperature (choice: 1.1314, noul: 1.2338, score: 0.4213), fitted via NLL + Brier minimization on held-out calibration records.

2. Shared-Prefix KV-Cache Multi-Question Execution

When a request contains multiple questions (q_1, ..., q_M) over a shared state document:

  1. For states with ≥ 256 tokens and M > 1, the shared state prefix S is encoded once (use_cache=True).
  2. Each question branch br_m is evaluated by reusing the cached state key-value tensors and cropping back via pkv.crop(-len(br_m)) after each branch.
  3. For shorter states, each S + br_m row is evaluated with exact causal masking. In both paths, question isolation is mathematically exact: the probability distribution for question q_m never depends on which other questions are asked in the same request.

3. Deterministic date_facts Preprocessor

Small language models often struggle with mental calendar subtraction across months. During to_internal_record(..., add_date_facts=True), if a state contains 2 to 7 absolute dates (e.g., June 26, 2026 and July 4, 2026), a deterministic helper appends explicit day-count differences (date_facts: July 4, 2026 is 8 days after June 26, 2026.). This boosts hard-v1 date/deadline accuracy to 77.00% with zero external overhead.


Training Recipe (jev_s1_clean_v2)

The model was trained on jev_s1_clean_v2, a pure System 1 dataset of 18,604 training records (34,134 questions) and 770 development records (1,134 questions) across 5 balanced pools (100% decoupled from any browser or DOM tasks):

  1. Pool 1 — decision-v7 Enriched 12-Source Pool (6,504 records): Balanced public classification/NLI/sentiment datasets (ag_news, amazon, banking77, boolq, dbpedia_14, imdb, mnli, sst5, trec, yelp, plus programmatic policies), rendered 50% with rich teacher semantic criteria descriptions and 50% with terse labels. All choice options are deterministically shuffled so gold option indices are strictly 1/K uniform.
  2. Pool 2 — hard-v1 Compositional Reasoning Families (3,500 records): 7 programmatic reasoning families (negation, counterfactual, arithmetic, temporal, entity_disambig, multi_hop_chain, distractor_resistance) using Templates 0–3 for training while holding out Template 4 exclusively for evaluation.
  3. Pool 3 — night2 Calibration & Robustness Suite (1,400 records): Date-bearing policy cases with relational day counts, unknowable/evidence-free states trained with uniform 1/K soft targets to prevent hallucinated overconfidence, and statement-form noul assertions.
  4. Pool 4 — Filtered Multi-Branch Teacher Suite (4,200 records): High-confidence (teacher_p >= 0.85) multi-question records covering situation routing, generic-vs-catchall other/none_of_the_above abstention, monotone score rubrics, and dual-question safety checks.
  5. Pool 5 — Programmatic Rule Twins & Level-Balanced Ordinal Rubrics (3,000 records): Counterfactual state_twin and rule_twin pairs (where flipping one condition flips the answer) plus strictly 1:1 level-balanced 3/4/5-level score rubrics and multi-condition noul verification checklists.

Optimization Hyperparameters:

  • LoRA Configuration: r=16, lora_alpha=32, lora_dropout=0.0, all 42 layers, target modules ["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj"] (~25.3M trainable parameters).
  • Loss Function: Hybrid Proper Scoring Rule combining KL divergence, Brier proper scoring rule, and an ordinal distance penalty on score questions: $$\mathcal{L} = \text{KL}(p_{\text{target}} ,|, p_\theta) + 0.25 \cdot |p_\theta - p_{\text{target}}|_2^2 + 0.15 \cdot \mathbb{E}\left[\frac{|i - y|}{K - 1}\right]$$
  • Schedule: Cosine LR schedule with peak lr=2.8e-5 for LoRA and 1.4e-4 for LetterReadoutHead, batch size 12, weight decay 0.01, gradient clipping 1.0.

Quickstart

1. Installation

pip install -r requirements.txt

2. Direct Python Inference

import torch
from model import MiniCPMSystemOne, to_internal_record

device = "cuda" if torch.cuda.is_available() else "mps" if torch.backends.mps.is_available() else "cpu"
model, tok = MiniCPMSystemOne.load_checkpoint(".", device=device, dtype=torch.bfloat16)

record = {
    "state": "Order #8841 placed on June 2, 2026. Customer requested a return on June 11, 2026. Item is unopened in original packaging. Standard return window is 14 days.",
    "questions": {
        "within_window": {
            "type": "noul",
            "instructions": "Is the return request within the 14-day return window?",
        },
        "routing": {
            "type": "choice",
            "instructions": "Select the appropriate resolution action.",
            "criteria": {
                "approve_full_refund": "Item is unopened and within 14 days",
                "approve_store_credit": "Item is unopened but between 15 and 30 days",
                "reject_return": "Item is opened or past 30 days",
            },
        },
        "urgency": {
            "type": "score",
            "instructions": "Rate the escalation risk from 0 to 3.",
            "criteria": [
                "0: Routine automated return",
                "1: Minor policy clarification needed",
                "2: High-value dispute",
                "3: Immediate legal/chargeback threat",
            ],
        },
    },
}

internal = to_internal_record(record, add_date_facts=True)
results = model.predict_record(internal, use_pride=False)

for q, (_, probs) in zip(internal["questions"], results):
    dist = {k: round(float(p), 4) for k, p in zip(q["keys"], probs.tolist())}
    best_key = max(dist, key=dist.get)
    print(f"{q['name']} ({q['type']}) -> argmax={best_key}: {dist}")

3. Serving the /v1/systemone HTTP API

Start the FastAPI server (defaults to port 8013 and loads the checkpoint from the current directory):

python serve.py --host 127.0.0.1 --port 8013

Query with curl (or any /v1/systemone compatible SDK):

curl -s http://127.0.0.1:8013/v1/systemone \
  -H "Content-Type: application/json" \
  -d '{
    "state": "Customer writes: I was charged twice for my monthly Pro subscription this morning (#INV-9921 and #INV-9922).",
    "questions": {
      "department": {
        "type": "choice",
        "instructions": "Route this ticket to the right team.",
        "criteria": {
          "billing": "Duplicate charges, invoices, refunds, payment methods",
          "technical": "Bugs, crashes, API errors, latency",
          "sales": "Enterprise quotes, seat expansion"
        }
      },
      "is_duplicate_charge": {
        "type": "noul",
        "instructions": "Does the customer report being billed more than once?"
      }
    }
  }' | python3 -m json.tool

Response:

{
  "model": "MiniCPM5-2B-Jev",
  "answers": {
    "department": {
      "billing": 0.9961,
      "technical": 0.0021,
      "sales": 0.0018
    },
    "is_duplicate_charge": {
      "false": 0.0034,
      "true": 0.9966
    }
  },
  "usage": {
    "latency_ms": 42.8
  }
}

Repository File Inventory

File Size SHA-256 (Prefix) Description
adapter_model.safetensors 100.5 MB ea364cdb96037814... LoRA adapter weights (all 42 layers, r=16)
adapter_config.json 1.2 KB 2c9f15a8... PEFT LoRA configuration (openbmb/MiniCPM5-2B)
head.pt 2.1 MB 0adbb21a3798b193... LetterReadoutHead weights, per-type temperatures, and marginal prior debiasing buffers
checkpoint_meta.json 1.2 KB — Calibration temperatures and validation summary
benchmark_results.json 11.1 KB — Full evaluation metrics across all 7 benchmark suites (decision-v7, JevBench, jabr v1/v2, hard-v1, DecisionBench)
training_history.json 22.4 KB — Complete training log from Step 0 (untuned base) through training completion
model.py 36.0 KB — Standalone MiniCPMSystemOne + LetterReadoutHead + to_internal_record implementation
serve.py 4.7 KB — Standalone FastAPI /v1/systemone server
requirements.txt 153 B — Python dependencies

License

Released under the Apache-2.0 License. Base model openbmb/MiniCPM5-2B is subject to its upstream license terms.

Downloads last month
42
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ytbai/MiniCPM5-2B-Jev

Adapter
(24)
this model

Datasets used to train ytbai/MiniCPM5-2B-Jev

Evaluation results

  • Overall Micro Accuracy on JevBench (Easy + Standard + Hard)
    self-reported
    0.788
  • Easy Tier Accuracy (48 items) on JevBench (Easy + Standard + Hard)
    self-reported
    0.979
  • Standard/Original Tier Accuracy (72 items) on JevBench (Easy + Standard + Hard)
    self-reported
    0.944
  • Hard Tier Accuracy (111 items) on JevBench (Easy + Standard + Hard)
    self-reported
    0.604
  • brier_score on JevBench (Easy + Standard + Hard)
    self-reported
    0.316
  • expected_calibration_error on JevBench (Easy + Standard + Hard)
    self-reported
    0.070
  • accuracy on decision-v7 (10-Source Stratified, 300 states / 390 questions)
    self-reported
    0.864
  • brier_score on decision-v7 (10-Source Stratified, 300 states / 390 questions)
    self-reported
    0.214