pjev-2b

A small typed-decision model. Give it evidence, a question and a fixed set of allowed answers; it returns a calibrated probability for every option in one forward pass. It never generates text, so there is nothing to parse and nothing to hallucinate.

The readout is the base model's own output layer. Options are rendered A., B., C., the prompt ends where the answer letter goes, and the logits at that single position are restricted to the letters in play and softmaxed. No decision head is added and no new parameters are introduced — which is why the pretrained model's own sense of its uncertainty survives fine-tuning instead of being relearned from scratch, and why the merged model is stock LlamaForCausalLM that any runtime loads without adaptation.

This repository holds the merged standalone model: base weights with the adapter folded in, tokenizer, and calibration file. No separate adapter, no separate base download.

Base model · JevBench

Accuracy per parameter

Accuracy against parameter count on the JevBench hard tier

JevBench's 231 public items, scored with the benchmark's own modules, one H100, in-process, batch 1. Competitor figures are recomputed from the benchmark's published per-item records on the identical items, except imajev-4b, which publishes aggregates only.

System Params Easy Standard Hard
imajev-4b 4B 100.0 98.6 72.1
djev — 100.0 98.6 67.6
pjev-2b 2.5B 100.0 81.9 61.3
SemIf 4B 100.0 98.6 61.3
jqv 32B 100.0 95.8 61.3
reflex-4B 4B 100.0 94.4 60.4
kev 8B 8B 100.0 93.1 45.0
Laya 0.4B 95.8 69.4 35.1

pjev-2b is not the most accurate model here, and two open systems above it are. What it is: the smallest model that holds the 4B line. At 61.3 on the hard tier it matches a 4B and a 32B, beats every 8B-and-under open model we could pair against, and does it at a quarter to a thirteenth of their parameter count.

Paired on the 111 hard items, McNemar's exact test with a 20,000-sample paired bootstrap:

Opponent Diff 95% CI p Verdict
Laya +26.1 +12.6, +38.7 0.0003 win
the 8B-and-under open field +16.2 to +31.5 — ≤0.0055 win
the 4B class ±0.9 — ≥0.28 tie

These are public-half numbers and are not an official JevBench score. The official board also scores 308 sealed items that only its operator can run, and every ranked system drops substantially from public to sealed accuracy. Treat this as a setup check, not a ranking.

Several questions, one piece of evidence

Latency for one to six questions over a shared state

Real traffic rarely asks one question about a document. Route it, flag its urgency, and check it against a policy, and that is three decisions over the same evidence. The state is encoded once and the per-question suffixes batch: six questions cost 48 ms in total, 8.0 ms each, against 121 ms one at a time. One question is still cheapest on the direct path at 20.5 ms; the crossover is at two.

Quick start

pip install "vllm>=0.30.0"
hf download Kailune-AI/pjev-2b --local-dir pjev-2b
vllm serve ./pjev-2b --port 8000 --served-model-name pjev-2b \
  --dtype bfloat16 --max-model-len 16384 --enable-prefix-caching \
  --logprobs-mode processed_logprobs --max-logprobs 32

It also loads directly with AutoModelForCausalLM.from_pretrained("Kailune-AI/pjev-2b").

A request and its real response:

import requests, json

state = "Ticket: 'Charged twice for one order, need the duplicate refunded.'"
questions = {
    "queue": {"type": "choice", "instructions": "Route this ticket.",
              "criteria": {"billing": None, "shipping": None, "account": None, "other": None}},
    "urgent": {"type": "noul", "instructions": "This needs attention within 24 hours."},
}
r = requests.post("http://127.0.0.1:8000/v1/systemone",
                  json={"state": state, "questions": questions})
print(json.dumps(r.json()["answers"], indent=2))
{
  "queue": {
    "type": "choice",
    "choice": "billing",
    "probabilities": {"billing": 0.981, "shipping": 0.004, "account": 0.011, "other": 0.004},
    "confidence": 0.981
  },
  "urgent": {"type": "noul", "noul": 0.874}
}

Read it as: route to billing, and threshold the urgent probability against your own cost of being wrong. There is no verdict to parse and no confidence the model invented.

What to expect

Accuracy. Strong on easy and standard items, mid-field on hard ones. If your decisions look like routing, classification, policy checks against a short document, or yes/no judgements with clear criteria, this is the right size. If they need multi-hop reasoning over long policies, the hard-tier number above is the honest guide.

Confidence. Calibration is the axis where this model does not lead. Hard-tier expected calibration error is 0.120; several competing systems are better there, and the ones that are fit a temperature on a held-out set. We ship at temperature 1.0 — see below — which is a defensible default and also a measurable cost on the hardest items.

Speed. 20.5 ms per decision on one H100, one question at a time; 8.0 ms per question when several share one state.

Options. Up to 26 options are read directly as letters. Beyond that a tournament round is used, which is an approximation rather than a single joint softmax.

Abstention. There is none. If the evidence supports no option, the model still distributes probability across the options you supplied. Threshold on confidence and route the low end to a person.

How it works

  1. The state, the question and the option list are rendered into a prompt that ends immediately before the answer letter.
  2. One forward pass. At that final position the logits of the tokens A, B, C, … are gathered — each is a single token in this tokenizer, asserted at load — and everything else is discarded.
  3. Softmax over the letters in play gives the distribution. For noul questions the two letters are the yes and no branches; for score questions they are the ordered levels.
  4. Nothing is decoded. A request with several questions about one state runs the shared prefix once and batches the per-question suffixes.

Why temperature 1.0. Our held-out calibration folds turned out easier than genuinely hard items, so every temperature fitted on them came out below 1 and made the model more confident exactly where it should have been less. Rather than ship a temperature that flatters the easy cases, we ship uncalibrated and say so. If you have a few hundred labelled examples of your own traffic, fitting one temperature on them is the first improvement to try, and on our own hard-tier numbers it is worth real points.

The merged weights here were checked against the unmerged base-plus-adapter on 40 JevBench easy and standard items: identical answers on 40 of 40, maximum probability difference 0.0000.

Limitations

  • Mid-field on hard reasoning. Two open systems in the table above beat it on the hard tier. It is a 2B; this is the trade.
  • Not the best calibrated. Hard-tier ECE 0.120, shipped uncalibrated. Fit your own temperature.
  • No abstention output. Nothing detects "the evidence does not answer this".
  • 26 options before a tournament round, which is an approximation.
  • English in practice. Evaluated in English only.
  • Public-benchmark numbers only. We have no measurement on any sealed evaluation, and every ranked system on the public board drops substantially when one is run.
  • Not a safety, medical, legal or hiring certificate. It returns probabilities over options you supply.

Licence and credits

Apache-2.0, matching the openbmb/MiniCPM5-2B base.

Fine-tuned on the public caiovicentino1/eikos-decisions dataset (CC BY 4.0, with an ODC-BY-1.0 subset), used as released. That corpus and the recipe it documents are its author's work; please carry the attribution if you build on this.

Downloads last month
-
Safetensors
Model size
3B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Kailune-AI/pjev-2b

Finetuned
(55)
this model

Dataset used to train Kailune-AI/pjev-2b