Instructions to use Kailune-AI/pjev-2b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Kailune-AI/pjev-2b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Kailune-AI/pjev-2b") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Kailune-AI/pjev-2b") model = AutoModelForCausalLM.from_pretrained("Kailune-AI/pjev-2b", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Kailune-AI/pjev-2b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Kailune-AI/pjev-2b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Kailune-AI/pjev-2b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Kailune-AI/pjev-2b
- SGLang
How to use Kailune-AI/pjev-2b with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Kailune-AI/pjev-2b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Kailune-AI/pjev-2b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Kailune-AI/pjev-2b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Kailune-AI/pjev-2b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Kailune-AI/pjev-2b with Docker Model Runner:
docker model run hf.co/Kailune-AI/pjev-2b
pjev-2b
A small typed-decision model. Give it evidence, a question and a fixed set of allowed answers; it returns a calibrated probability for every option in one forward pass. It never generates text, so there is nothing to parse and nothing to hallucinate.
The readout is the base model's own output layer. Options are rendered A., B., C., the prompt
ends where the answer letter goes, and the logits at that single position are restricted to the
letters in play and softmaxed. No decision head is added and no new parameters are introduced —
which is why the pretrained model's own sense of its uncertainty survives fine-tuning instead of
being relearned from scratch, and why the merged model is stock LlamaForCausalLM that any runtime
loads without adaptation.
This repository holds the merged standalone model: base weights with the adapter folded in, tokenizer, and calibration file. No separate adapter, no separate base download.
Accuracy per parameter
JevBench's 231 public items, scored with the benchmark's own modules, one H100, in-process, batch 1. Competitor figures are recomputed from the benchmark's published per-item records on the identical items, except imajev-4b, which publishes aggregates only.
| System | Params | Easy | Standard | Hard |
|---|---|---|---|---|
| imajev-4b | 4B | 100.0 | 98.6 | 72.1 |
| djev | — | 100.0 | 98.6 | 67.6 |
| pjev-2b | 2.5B | 100.0 | 81.9 | 61.3 |
| SemIf | 4B | 100.0 | 98.6 | 61.3 |
| jqv | 32B | 100.0 | 95.8 | 61.3 |
| reflex-4B | 4B | 100.0 | 94.4 | 60.4 |
| kev 8B | 8B | 100.0 | 93.1 | 45.0 |
| Laya | 0.4B | 95.8 | 69.4 | 35.1 |
pjev-2b is not the most accurate model here, and two open systems above it are. What it is: the smallest model that holds the 4B line. At 61.3 on the hard tier it matches a 4B and a 32B, beats every 8B-and-under open model we could pair against, and does it at a quarter to a thirteenth of their parameter count.
Paired on the 111 hard items, McNemar's exact test with a 20,000-sample paired bootstrap:
| Opponent | Diff | 95% CI | p | Verdict |
|---|---|---|---|---|
| Laya | +26.1 | +12.6, +38.7 | 0.0003 | win |
| the 8B-and-under open field | +16.2 to +31.5 | — | ≤0.0055 | win |
| the 4B class | ±0.9 | — | ≥0.28 | tie |
These are public-half numbers and are not an official JevBench score. The official board also scores 308 sealed items that only its operator can run, and every ranked system drops substantially from public to sealed accuracy. Treat this as a setup check, not a ranking.
Several questions, one piece of evidence
Real traffic rarely asks one question about a document. Route it, flag its urgency, and check it against a policy, and that is three decisions over the same evidence. The state is encoded once and the per-question suffixes batch: six questions cost 48 ms in total, 8.0 ms each, against 121 ms one at a time. One question is still cheapest on the direct path at 20.5 ms; the crossover is at two.
Quick start
pip install "vllm>=0.30.0"
hf download Kailune-AI/pjev-2b --local-dir pjev-2b
vllm serve ./pjev-2b --port 8000 --served-model-name pjev-2b \
--dtype bfloat16 --max-model-len 16384 --enable-prefix-caching \
--logprobs-mode processed_logprobs --max-logprobs 32
It also loads directly with AutoModelForCausalLM.from_pretrained("Kailune-AI/pjev-2b").
A request and its real response:
import requests, json
state = "Ticket: 'Charged twice for one order, need the duplicate refunded.'"
questions = {
"queue": {"type": "choice", "instructions": "Route this ticket.",
"criteria": {"billing": None, "shipping": None, "account": None, "other": None}},
"urgent": {"type": "noul", "instructions": "This needs attention within 24 hours."},
}
r = requests.post("http://127.0.0.1:8000/v1/systemone",
json={"state": state, "questions": questions})
print(json.dumps(r.json()["answers"], indent=2))
{
"queue": {
"type": "choice",
"choice": "billing",
"probabilities": {"billing": 0.981, "shipping": 0.004, "account": 0.011, "other": 0.004},
"confidence": 0.981
},
"urgent": {"type": "noul", "noul": 0.874}
}
Read it as: route to billing, and threshold the urgent probability against your own cost of being
wrong. There is no verdict to parse and no confidence the model invented.
What to expect
Accuracy. Strong on easy and standard items, mid-field on hard ones. If your decisions look like routing, classification, policy checks against a short document, or yes/no judgements with clear criteria, this is the right size. If they need multi-hop reasoning over long policies, the hard-tier number above is the honest guide.
Confidence. Calibration is the axis where this model does not lead. Hard-tier expected calibration error is 0.120; several competing systems are better there, and the ones that are fit a temperature on a held-out set. We ship at temperature 1.0 — see below — which is a defensible default and also a measurable cost on the hardest items.
Speed. 20.5 ms per decision on one H100, one question at a time; 8.0 ms per question when several share one state.
Options. Up to 26 options are read directly as letters. Beyond that a tournament round is used, which is an approximation rather than a single joint softmax.
Abstention. There is none. If the evidence supports no option, the model still distributes probability across the options you supplied. Threshold on confidence and route the low end to a person.
How it works
- The state, the question and the option list are rendered into a prompt that ends immediately before the answer letter.
- One forward pass. At that final position the logits of the tokens
A,B,C, … are gathered — each is a single token in this tokenizer, asserted at load — and everything else is discarded. - Softmax over the letters in play gives the distribution. For
noulquestions the two letters are the yes and no branches; forscorequestions they are the ordered levels. - Nothing is decoded. A request with several questions about one state runs the shared prefix once and batches the per-question suffixes.
Why temperature 1.0. Our held-out calibration folds turned out easier than genuinely hard items, so every temperature fitted on them came out below 1 and made the model more confident exactly where it should have been less. Rather than ship a temperature that flatters the easy cases, we ship uncalibrated and say so. If you have a few hundred labelled examples of your own traffic, fitting one temperature on them is the first improvement to try, and on our own hard-tier numbers it is worth real points.
The merged weights here were checked against the unmerged base-plus-adapter on 40 JevBench easy and standard items: identical answers on 40 of 40, maximum probability difference 0.0000.
Limitations
- Mid-field on hard reasoning. Two open systems in the table above beat it on the hard tier. It is a 2B; this is the trade.
- Not the best calibrated. Hard-tier ECE 0.120, shipped uncalibrated. Fit your own temperature.
- No abstention output. Nothing detects "the evidence does not answer this".
- 26 options before a tournament round, which is an approximation.
- English in practice. Evaluated in English only.
- Public-benchmark numbers only. We have no measurement on any sealed evaluation, and every ranked system on the public board drops substantially when one is run.
- Not a safety, medical, legal or hiring certificate. It returns probabilities over options you supply.
Licence and credits
Apache-2.0, matching the openbmb/MiniCPM5-2B base.
Fine-tuned on the public
caiovicentino1/eikos-decisions
dataset (CC BY 4.0, with an ODC-BY-1.0 subset), used as released. That corpus and the recipe it
documents are its author's work; please carry the attribution if you build on this.
- Downloads last month
- -
Model tree for Kailune-AI/pjev-2b
Base model
openbmb/MiniCPM5-2B
