Instructions to use ytbai/MiniCPM5-2B-Jev with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use ytbai/MiniCPM5-2B-Jev with PEFT:
from peft import PeftModel from transformers import AutoModel base_model = AutoModel.from_pretrained("openbmb/MiniCPM5-2B") model = PeftModel.from_pretrained(base_model, "ytbai/MiniCPM5-2B-Jev") - Notebooks
- Google Colab
- Kaggle
MiniCPM5-2B-Jev
MiniCPM5-2B-Jev is a high-speed, calibrated System 1 Decision Model built on OpenBMB's MiniCPM5-2B (base weights on Hugging Face). Code and evaluation suites are open-sourced at github.com/yuting-ai/minicpm5-2b-jev. Given a shared context document (state) and one or more typed decision specifications (choice, noul, score), it outputs a calibrated probability distribution over the candidate options for every question in a single forward pass — with zero autoregressive token generation.
It natively implements the /v1/systemone structured decision contract (compatible with Jev / OpenJev / Decider / Kev clients) and achieves state-of-the-art accuracy among all ≤ 2B parameter open-weight System 1 / Jev models, while outperforming several 4B–9B models on both JevBench and decision-v7.
Key Highlights
- #1 Among ≤ 2B Models on JevBench: Achieves 78.79% Overall (97.92% Easy, 94.44% Standard, 60.36% Hard) on the 231-item public
JevBenchbenchmark, surpassingdecider-2b v11(76.2%),Kev-4B(75.8%),open-alternative-jev 4B(74.0%),system-one-open Gemma-4-E2B(73.2%),system-one Qwen3-8B(71.9%), andBespoke Nimble 9B(67.5%). - 86.41% Accuracy & 0.0307 ECE on
decision-v7: OutperformsKev-4B(85.9% Acc, 0.056 ECE) and hostedJev(84.5% Acc ondecision-v7dev), and approachesKev-9B(87.2%) andKev-27B(87.0%) at 1/4th to 1/13th the parameter count. - 87.18% on
jabr v1& 79.75% onjabr v2: Strong zero-shot generalization across 49 diverse classification, policy routing, Boolean verification (noul), and ordinal rubric scoring (score) tasks. - Exact Multi-Question KV-Cache Branching: Encodes the shared
stateprefix once (use_cache=True) and evaluates M questions in parallel branches viaDynamicCache.crop(), guaranteeing exact question isolation (zero cross-question leakage) and low latency (~260 ms/state on Apple Silicon MPS; <35 ms on CUDA). - Calibrated Native
LetterReadoutHead: Projects the terminalAnswer: (hidden state onto 255 single-token option letters (A..Z, AA..) initialized fromlm_head.weight, combined with class-balanced marginal prior debiasing (noul_bias,choice_k_bias,score_k_bias) and per-type temperatures (choice: 1.1314,noul: 1.2338,score: 0.4213).
Horizontal Benchmark Comparison
1. JevBench Public Benchmark Suite (195 Groups / 231 Decisions)
Evaluated on the official fstandhartinger/jevbench suite (easy: 48, original/standard: 72, hard: 111; total 231 decisions). Peer model scores are from the official JevBench / Decider / Kev / Laya published leaderboards on the exact same frozen items:
| Model | Base Backbone | Params | Easy (48) | Standard (72) | Hard (111) | Overall (231) | License / Weights |
|---|---|---|---|---|---|---|---|
| Jev 1.13.0 (Hosted API) | Proprietary (TypeSafe) | Closed | 100.0% | 98.6% | 73.0% | 86.6% | Proprietary |
| decider-35b-a3b v1 | Qwen3.5-35B-A3B (MoE) | 35B (3B act.) | 100.0% | 97.2% | 67.6% | 83.5% | Apache-2.0 |
| decider-4b v2.1 | Qwen3.5-4B | 4.2B | 100.0% | 98.6% | 64.9% | 82.7% | Apache-2.0 |
| OpenJev (Hosted) | DiffusionGemma 26B-A4B | 26B (4B act.) | 100.0% | 97.2% | 64.0% | 81.8% | Proprietary |
| SemIf | Qwen3.5-4B | 4.2B | 100.0% | 98.6% | 61.3% | 81.0% | Open Weights |
| ⭐ MiniCPM5-2B-Jev (Ours) | openbmb/MiniCPM5-2B | 2.0B | 97.9% | 94.4% | 60.4% | 78.8% | Apache-2.0 |
| decider-2b v11 | Qwen3.5-2B | 1.9B | 100.0% | 88.9% | 57.7% | 76.2% | Apache-2.0 |
| Kev-4B (r10) | Qwen3.5-4B-Base | 4.2B | — | — | 54.1% | 75.8% | Apache-2.0 |
| open-alternative-jev | Qwen3.5-4B | 4.2B | 100.0% | 83.3% | 56.8% | 74.0% | Open Weights |
| system-one-open | Gemma 4 E2B | ~2B | 100.0% | 93.1% | 48.6% | 73.2% | Open Weights |
| system-one | Qwen3-8B | 8.2B | 100.0% | 88.9% | 48.6% | 71.9% | Open Weights |
| decider-2b v10 | Qwen3.5-2B | 1.9B | 100.0% | 88.9% | 45.9% | 70.6% | Apache-2.0 |
| Bespoke Nimble 9B | Bespoke-Nimble-9B | 9.0B | 100.0% | 93.1% | 36.9% | 67.5% | Hosted |
| Kev-0.8B (r15) | Qwen3.5-0.8B-Base | 0.8B | — | — | 36.0% | 63.6% | Apache-2.0 |
| open-jev-deberta-v3-large | DeBERTa-v3-Large | 0.4B | 100.0% | 43.1% | 37.8% | 52.4% | Apache-2.0 |
Key Takeaway: Within the ≤ 2B parameter class,
MiniCPM5-2B-Jevachieves 78.79% overall accuracy and 60.36% on the challengingHardtier (adversarial traps, prompt injections, multi-step arithmetic, negation, and complex policy edge cases), outperformingdecider-2b v11(+2.6 pp overall, +2.7 pp Hard),Kev-4B(+3.0 pp overall, +6.3 pp Hard), andQwen3-8B system-one(+6.9 pp overall, +11.8 pp Hard).
2. decision-v7 Benchmark (10-Source Typed Decision Evaluation)
Comparison on the standard 10-source decision-v7 benchmark (ag_news, amazon, banking77, boolq, dbpedia_14, imdb, mnli, sst5, trec, yelp) across accuracy, Brier score, and Expected Calibration Error (ECE):
| Model | Base Backbone | Params | Accuracy (↑) | Brier Score (↓) | ECE (↓) |
|---|---|---|---|---|---|
| Kev-9B (as served, T=2.30) | Qwen3.5-9B-Base | 9.0B | 87.2% | — | 0.042 |
| Kev-27B (as served, T=1.38) | Qwen3.8-27B | 27.0B | 87.0% | — | — |
| ⭐ MiniCPM5-2B-Jev (Ours) | openbmb/MiniCPM5-2B | 2.0B | 86.4% (86.41%) | 0.214 (0.2143) | 0.031 (0.0307) |
| Kev-4B (as served, T=1.89) | Qwen3.5-4B-Base | 4.2B | 85.9% | 0.256 | 0.056 |
| Jev (TypeSafe Hosted) | Proprietary | Closed | 84.5% | — | — |
| Kev-0.8B (as served, T=1.81) | Qwen3.5-0.8B-Base | 0.8B | 77.1% | 0.369 | 0.088 |
| MiniCPM5-2B (Untuned Base, Step 0) | openbmb/MiniCPM5-2B | 2.0B | 62.5% | 0.693 | 0.272 |
Per-Source Breakdown on decision-v7 (MiniCPM5-2B-Jev):
boolq: 100.00% |dbpedia_14: 97.44% |ag_news: 94.87% |imdb: 94.87% |trec: 94.87% |mnli: 92.31% |banking77: 87.18% (77-way intent!) |sst5: 76.92% |yelp: 64.10% (5-way fine-grained star rating) |amazon: 61.54% (5-way star rating).- By Question Type:
noul(Boolean verification): 94.97% |choice(multi-class up to 77 options): 87.10% |score(ordinal 0–10 / 5-level): 67.11%.
3. Comprehensive Multi-Suite Results
We evaluate MiniCPM5-2B-Jev across 7 zero-browser decision benchmarks totaling 2,042 states and 2,637 questions:
| Benchmark Suite | States / Questions | Accuracy | Choice Acc | Noul Acc | Score Acc | Brier (↓) | ECE (↓) |
|---|---|---|---|---|---|---|---|
decision-v7 (10-Source Stratified) |
300 / 390 | 86.41% | 87.10% | 94.97% | 67.11% | 0.2143 | 0.0307 |
jabr v1 (8 Tasks / 78 Cases) |
78 / 78 | 87.18% | 96.00% | 76.92% | 88.89% | 0.1771 | 0.0277 |
jabr v2 (49 Tasks / 869 Cases) |
869 / 869 | 79.75% | 85.16% | 79.49% | 69.95% | 0.2919 | 0.0663 |
JevBench (Overall Public Suite) |
231 / 231 | 78.79% | 78.42% | 81.08% | 72.22% | 0.3157 | 0.0699 |
| ↳ JevBench Easy Tier | 48 / 48 | 97.92% | — | — | — | — | — |
| ↳ JevBench Standard/Original Tier | 72 / 72 | 94.44% | — | — | — | — | — |
| ↳ JevBench Hard Tier | 111 / 111 | 60.36% | — | — | — | — | — |
hard-v1 (Held-Out Template 4) |
350 / 543 | 60.41% | 58.35% | 70.43% | 40.00% | 0.5419 | 0.0603 |
DecisionBench Medium |
80 / 293 | 60.07% | 70.93% | 80.52% | 34.48% | 0.6596 | 0.2647 |
DecisionBench Hard |
80 / 293 | 44.71% | 53.75% | 70.79% | 23.33% | 0.8768 | 0.3711 |
Model Architecture
1. Native lm_head Letter Readout (LetterReadoutHead)
Instead of training a random projection head from scratch, MiniCPM5-2B-Jev extracts the exact rows of MiniCPM5-2B's pretrained lm_head.weight corresponding to 255 single-token option letters (A..Z for indices 0..25, plus two-letter tokens AA.. for indices 26..254, shape [255, 2304]).
Each question branch is formatted as:
Question: [{type}] {instructions}
(A) {option_0}
(B) {option_1}
...
Answer: (
At the final token ( of Answer: (, the model's hidden state h (dimension 2304) is projected onto the K active option letters:
where:
W_letter(shape[255, 2304]) is initialized fromlm_head.weight[letter_token_ids]and fine-tuned jointly with LoRA.b_prior(t, k)is the class-balanced marginal prior debiasing offset (noul_bias,choice_k_bias,score_k_bias), which centers marginal log-odds on the validation set to eliminate yes/no or position bias without altering the training loss.T_tis the per-type temperature (choice: 1.1314,noul: 1.2338,score: 0.4213), fitted via NLL + Brier minimization on held-out calibration records.
2. Shared-Prefix KV-Cache Multi-Question Execution
When a request contains multiple questions (q_1, ..., q_M) over a shared state document:
- For states with ≥ 256 tokens and
M > 1, the sharedstateprefixSis encoded once (use_cache=True). - Each question branch
br_mis evaluated by reusing the cached state key-value tensors and cropping back viapkv.crop(-len(br_m))after each branch. - For shorter states, each
S + br_mrow is evaluated with exact causal masking. In both paths, question isolation is mathematically exact: the probability distribution for questionq_mnever depends on which other questions are asked in the same request.
3. Deterministic date_facts Preprocessor
Small language models often struggle with mental calendar subtraction across months. During to_internal_record(..., add_date_facts=True), if a state contains 2 to 7 absolute dates (e.g., June 26, 2026 and July 4, 2026), a deterministic helper appends explicit day-count differences (date_facts: July 4, 2026 is 8 days after June 26, 2026.). This boosts hard-v1 date/deadline accuracy to 77.00% with zero external overhead.
Training Recipe (jev_s1_clean_v2)
The model was trained on jev_s1_clean_v2, a pure System 1 dataset of 18,604 training records (34,134 questions) and 770 development records (1,134 questions) across 5 balanced pools (100% decoupled from any browser or DOM tasks):
- Pool 1 —
decision-v7Enriched 12-Source Pool (6,504 records): Balanced public classification/NLI/sentiment datasets (ag_news,amazon,banking77,boolq,dbpedia_14,imdb,mnli,sst5,trec,yelp, plus programmatic policies), rendered 50% with rich teacher semantic criteria descriptions and 50% with terse labels. Allchoiceoptions are deterministically shuffled so gold option indices are strictly 1/K uniform. - Pool 2 —
hard-v1Compositional Reasoning Families (3,500 records): 7 programmatic reasoning families (negation,counterfactual,arithmetic,temporal,entity_disambig,multi_hop_chain,distractor_resistance) using Templates 0–3 for training while holding out Template 4 exclusively for evaluation. - Pool 3 —
night2Calibration & Robustness Suite (1,400 records): Date-bearing policy cases with relational day counts, unknowable/evidence-free states trained with uniform 1/K soft targets to prevent hallucinated overconfidence, and statement-formnoulassertions. - Pool 4 — Filtered Multi-Branch Teacher Suite (4,200 records): High-confidence (
teacher_p >= 0.85) multi-question records covering situation routing, generic-vs-catchallother/none_of_the_aboveabstention, monotone score rubrics, and dual-question safety checks. - Pool 5 — Programmatic Rule Twins & Level-Balanced Ordinal Rubrics (3,000 records): Counterfactual
state_twinandrule_twinpairs (where flipping one condition flips the answer) plus strictly 1:1 level-balanced 3/4/5-levelscorerubrics and multi-conditionnoulverification checklists.
Optimization Hyperparameters:
- LoRA Configuration:
r=16,lora_alpha=32,lora_dropout=0.0, all 42 layers, target modules["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj"](~25.3M trainable parameters). - Loss Function: Hybrid Proper Scoring Rule combining KL divergence, Brier proper scoring rule, and an ordinal distance penalty on
scorequestions: $$\mathcal{L} = \text{KL}(p_{\text{target}} ,|, p_\theta) + 0.25 \cdot |p_\theta - p_{\text{target}}|_2^2 + 0.15 \cdot \mathbb{E}\left[\frac{|i - y|}{K - 1}\right]$$ - Schedule: Cosine LR schedule with peak
lr=2.8e-5for LoRA and1.4e-4forLetterReadoutHead, batch size 12, weight decay0.01, gradient clipping1.0.
Quickstart
1. Installation
pip install -r requirements.txt
2. Direct Python Inference
import torch
from model import MiniCPMSystemOne, to_internal_record
device = "cuda" if torch.cuda.is_available() else "mps" if torch.backends.mps.is_available() else "cpu"
model, tok = MiniCPMSystemOne.load_checkpoint(".", device=device, dtype=torch.bfloat16)
record = {
"state": "Order #8841 placed on June 2, 2026. Customer requested a return on June 11, 2026. Item is unopened in original packaging. Standard return window is 14 days.",
"questions": {
"within_window": {
"type": "noul",
"instructions": "Is the return request within the 14-day return window?",
},
"routing": {
"type": "choice",
"instructions": "Select the appropriate resolution action.",
"criteria": {
"approve_full_refund": "Item is unopened and within 14 days",
"approve_store_credit": "Item is unopened but between 15 and 30 days",
"reject_return": "Item is opened or past 30 days",
},
},
"urgency": {
"type": "score",
"instructions": "Rate the escalation risk from 0 to 3.",
"criteria": [
"0: Routine automated return",
"1: Minor policy clarification needed",
"2: High-value dispute",
"3: Immediate legal/chargeback threat",
],
},
},
}
internal = to_internal_record(record, add_date_facts=True)
results = model.predict_record(internal, use_pride=False)
for q, (_, probs) in zip(internal["questions"], results):
dist = {k: round(float(p), 4) for k, p in zip(q["keys"], probs.tolist())}
best_key = max(dist, key=dist.get)
print(f"{q['name']} ({q['type']}) -> argmax={best_key}: {dist}")
3. Serving the /v1/systemone HTTP API
Start the FastAPI server (defaults to port 8013 and loads the checkpoint from the current directory):
python serve.py --host 127.0.0.1 --port 8013
Query with curl (or any /v1/systemone compatible SDK):
curl -s http://127.0.0.1:8013/v1/systemone \
-H "Content-Type: application/json" \
-d '{
"state": "Customer writes: I was charged twice for my monthly Pro subscription this morning (#INV-9921 and #INV-9922).",
"questions": {
"department": {
"type": "choice",
"instructions": "Route this ticket to the right team.",
"criteria": {
"billing": "Duplicate charges, invoices, refunds, payment methods",
"technical": "Bugs, crashes, API errors, latency",
"sales": "Enterprise quotes, seat expansion"
}
},
"is_duplicate_charge": {
"type": "noul",
"instructions": "Does the customer report being billed more than once?"
}
}
}' | python3 -m json.tool
Response:
{
"model": "MiniCPM5-2B-Jev",
"answers": {
"department": {
"billing": 0.9961,
"technical": 0.0021,
"sales": 0.0018
},
"is_duplicate_charge": {
"false": 0.0034,
"true": 0.9966
}
},
"usage": {
"latency_ms": 42.8
}
}
Repository File Inventory
| File | Size | SHA-256 (Prefix) | Description |
|---|---|---|---|
adapter_model.safetensors |
100.5 MB | ea364cdb96037814... |
LoRA adapter weights (all 42 layers, r=16) |
adapter_config.json |
1.2 KB | 2c9f15a8... |
PEFT LoRA configuration (openbmb/MiniCPM5-2B) |
head.pt |
2.1 MB | 0adbb21a3798b193... |
LetterReadoutHead weights, per-type temperatures, and marginal prior debiasing buffers |
checkpoint_meta.json |
1.2 KB | — | Calibration temperatures and validation summary |
benchmark_results.json |
11.1 KB | — | Full evaluation metrics across all 7 benchmark suites (decision-v7, JevBench, jabr v1/v2, hard-v1, DecisionBench) |
training_history.json |
22.4 KB | — | Complete training log from Step 0 (untuned base) through training completion |
model.py |
36.0 KB | — | Standalone MiniCPMSystemOne + LetterReadoutHead + to_internal_record implementation |
serve.py |
4.7 KB | — | Standalone FastAPI /v1/systemone server |
requirements.txt |
153 B | — | Python dependencies |
License
Released under the Apache-2.0 License. Base model openbmb/MiniCPM5-2B is subject to its upstream license terms.
- Downloads last month
- 42
Model tree for ytbai/MiniCPM5-2B-Jev
Base model
openbmb/MiniCPM5-2BDatasets used to train ytbai/MiniCPM5-2B-Jev
fancyzhx/ag_news
google/boolq
Evaluation results
- Overall Micro Accuracy on JevBench (Easy + Standard + Hard)self-reported0.788
- Easy Tier Accuracy (48 items) on JevBench (Easy + Standard + Hard)self-reported0.979
- Standard/Original Tier Accuracy (72 items) on JevBench (Easy + Standard + Hard)self-reported0.944
- Hard Tier Accuracy (111 items) on JevBench (Easy + Standard + Hard)self-reported0.604
- brier_score on JevBench (Easy + Standard + Hard)self-reported0.316
- expected_calibration_error on JevBench (Easy + Standard + Hard)self-reported0.070
- accuracy on decision-v7 (10-Source Stratified, 300 states / 390 questions)self-reported0.864
- brier_score on decision-v7 (10-Source Stratified, 300 states / 390 questions)self-reported0.214