K2-Type-1B

A 1B decision model in the style of TypeSafe's Jev. You send one state (text or JSON) and any number of typed questions; it returns a probability for every option of every question from one forward pass. It never generates text.

Question type You give You get
noul a statement, optional definitions of false / true P(true)
choice 1-255 named options with optional descriptions the best option, its confidence, all probabilities
score 2-255 ordered levels expected level, probability per level

Questions share the state but cannot see each other (block-causal attention mask), so adding a question never changes another's answer. Built on IFM/K2-Horizon-0.9B.

Results

JevBench public set (231 items, jevbench commit 26eb72d, typesafe adapter against this repo's server, one H200, no network):

Tier Correct
standard (72, original.jsonl) 66
easy (48, easy.jsonl) 47
hard (111, hard.jsonl) 63
Total 176 / 231 = 0.762

Brier 0.328, ECE 0.065; latency p50 27 ms, p95 60 ms per decision; mean 590 input tokens per decision. For reference, public accuracy on the JevBench v1.4 board: Gemma 4 E2B + LoRA (system-one-open) 0.732, decider-2b 0.710, kev 0.6B 0.667, kev 4B 0.662, Qwen3.5-4B entries 0.74-0.82, Jev 1.13.0 0.866. The official JevBench score also uses sealed items; every listed system scores well below its public accuracy there.

Kev suites (jaredpalmer/kev frozen dev/test sets, accuracy):

Model Trained sources, test New sources, test
K2-Type-1B 0.813 0.688
Kev-0.8B (published) 0.834 0.684

Snake: the same weights play Snake from a text board (one choice + four yes/no questions per move): 66 food per game on 12x12 on average (max 101).

Quickstart

pip install -U huggingface_hub
hf download chen11003/K2-Type-1B --local-dir K2-Type-1B && cd K2-Type-1B
pip install -r requirements.txt
python -m jev.serve --run . --port 8000          # from this repo's root; needs one CUDA GPU

The first request after start-up takes a few seconds (CUDA warm-up); later requests take about 20-60 ms.

curl -s localhost:8000/v1/systemone -H 'Content-Type: application/json' -d '{
  "state": {"subject": "Charged twice", "body": "You billed my card twice for March. Refund one or I cancel."},
  "questions": {
    "queue":  {"type": "choice", "instructions": "Which queue handles this?",
               "criteria": {"billing": "Payments and refunds", "technical": "Bugs and login", "general": "Anything else"}},
    "angry":  {"type": "noul", "instructions": "The customer sounds angry."},
    "urgency": {"type": "score", "instructions": "How urgent is it?", "criteria": ["Low", "Normal", "High", "Critical"]}
  }}'

The request and response follow TypeSafe's /v1/systemone wire format, so clients written for Jev or Kev work unchanged. GET /health reports the model name and temperature.

This repository contains the weights and the minimal code needed to serve them (jev/: input encoding, the pointer head, and the HTTP server). Training code and data are not released at this time.

Requires transformers >= 5.17 (remote code); tested on torch 2.8. Do not use the backbone for text generation: its weights were trained for the decision head, and the language-model head is left from the base model. Use the jev/ server.

How it works

  • Input layout: <state> ... | <q> question <opt> option </opt> ... <decide> | <q> ..., using five reserved tokens of the base tokenizer (decision_config.json).
  • A pointer head (pointer_head.safetensors) scores each option's </opt> hidden state against the question's <decide> hidden state; a softmax at temperature 1.478 gives the probabilities.
  • Training, in short: full fine-tune on ~354k decision records (public classification/NLI/QA sets, Kev's decision data, game positions with exact or search labels, and synthetic decision items written and blind-verified by a large model), targets mixed with soft labels from a 7B teacher decision model; then 50 iterations of PPO on Snake; then a temperature refit on held-out calibration data. Training data were checked against every Kev evaluation suite and the 231 JevBench public items: no exact or containment overlap.

Limits

  • One pass, no reasoning: multi-step arithmetic, date differences and very long documents (> 8192 tokens) are weak.
  • Calibration is fitted on Kev's calibration suite; on other distributions probabilities can be over- or under-confident (JevBench public ECE 0.065).
  • Answers are bounded by the options you give; it cannot say "none of these" unless you offer that option.
Downloads last month
-
Safetensors
Model size
1B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for chen11003/K2-Type-1B

Finetuned
(3)
this model