local-system-one-student
"We have Jev at home."
A thought experiment, not a product. TypeSafe's Jev is a "System One" model: state in, typed decisions with calibrated probabilities out, no text generation. Its architecture is not public. This is a 149M-parameter encoder that reproduces that interface and behaviour for one fixed decision schema (a coding agent's next step), distilled from local LLMs on a laptop. It is not a claim of parity with Jev and has no affiliation with TypeSafe.
Code, benchmark, and the full write-up: https://github.com/djhoomin/local-system-one
What it does
Given a JSON agent state (user_message, cwd, recent_tool_calls), one forward pass returns
four typed, calibrated answers:
| head | type | output |
|---|---|---|
tool |
Choice over web_search / run_shell / edit_file / send_email / none |
probabilities + argmax |
urgency |
Score over not urgent / somewhat urgent / urgent / critical |
distribution + expected level |
destructive |
Noul (yes/no) | P(irreversible data/state loss) |
needs_confirmation |
Noul | P(agent should confirm first) |
How it was made
- Student: ModernBERT-base, mean-pooled, four linear heads, 4 epochs, 6 min on an M1 Pro.
- Data: 2,300 synthetic agent states generated by Gemma 3n E4B (317 of them seeded on destructive actions so the rare positive class is learnable);
cwd/recent_tool_callssampled independently of the message so the student cannot route from context alone. - Teachers: Gemma 3 12B for
tool,needs_confirmationanddestructive; Gemma 3n E4B forurgency, each read from first-token logprobs on a multiple-choice prompt. - Targets are the teachers' calibrated distributions (temperature / Platt scaling fitted on a 123-row human-labeled set), so the student is calibrated by construction: a post-hoc temperature fit on its output comes back at T≈0.9.
Results on the 123 human-labeled test rows (never trained or selected on)
| model | ms/state | tool acc | tool ECE (raw) | urgency ±1 | destructive AUROC | needs_confirm AUROC |
|---|---|---|---|---|---|---|
| zero-shot DeBERTa-v3 NLI | 123 | 0.50 | 0.08 | 0.38 | 0.72 | 0.47 |
| Gemma 3n E4B (teacher) | 2062 | 0.72 | 0.28 | 0.99 | 0.95 | 0.71 |
| Gemma 3 12B (teacher) | 5508 | 0.83 | 0.15 | 0.86 | 0.94 | 0.78 |
| this model (v4) | 19 | 0.76 | 0.08 | 0.98 | 0.99 | 0.67 |
Latencies are on an M1 Pro (MPS). The test set is one person's labels on a made-up task; read the numbers as a benchmark of the method, not of general ability.
Usage
The weights are a plain state_dict for the Student class in the GitHub repo; no custom
transformers architecture.
# git clone https://github.com/djhoomin/local-system-one && cd local-system-one && uv sync
from local_systemone import LocalClient
from typesafe_sdk import Choice, Noul
client = LocalClient(backend="student", model="Mannedood/local-system-one-student") # downloads from the Hub
resp = client.system_one(
state={"user_message": "rm -rf the build dir and rerun tests, prod is down",
"cwd": "/work/webapp", "recent_tool_calls": []},
questions={"tool": Choice(instructions="", criteria={...}), # matched by name; instructions are ignored
"destructive": Noul(instructions="")},
)
Or by hand:
import json, torch
from safetensors.torch import load_file
from huggingface_hub import hf_hub_download
from distill.train_student import Student, render # from the GitHub repo
from transformers import AutoTokenizer
cfg = json.load(open(hf_hub_download("Mannedood/local-system-one-student", "config.json")))
model = Student(cfg["encoder"]); model.load_state_dict(load_file(hf_hub_download("Mannedood/local-system-one-student", "model.safetensors"))); model.eval()
tok = AutoTokenizer.from_pretrained(cfg["encoder"])
out = model(**tok(render(state), return_tensors="pt", truncation=True, max_length=cfg["max_len"]))
probs = torch.softmax(out["tool"][0], -1) # over cfg["tool_labels"]
Limitations
- Task-specific: it answers these four questions by name and ignores instructions/criteria. That is the trade for encoder latency; Jev is general.
- Fixed schema, one domain. The input must be the JSON shape it was trained on
(
user_message,cwd,recent_tool_calls); rename or add a key and it is out of distribution without any error. All training states were a coding agent's. A customer-support or ops message in the same JSON shape will still get an answer with plausible-looking probabilities, but nothing in the evaluation says those are calibrated — encoders degrade quietly, not loudly. Inside the box it is a 19 ms calibrated function; outside it, a guess with a confident face. - English only.
needs_confirmationis the weakest head (AUROC 0.67 vs the teacher's 0.78).destructivepositives score 0.28–0.93 on the human test set against a negative median of 0.01 (Brier 0.026). Earlier versions had a compressed scale here; v4 fixed it by adding 317 destructive-seeded training states, since the balanced set had only 1.4% positives.- 123 test rows: ±4 points on accuracy is noise.
License and provenance
Weights and code are MIT. The base model is ModernBERT-base (Apache-2.0). Training data was generated and labeled with Gemma models, so the Gemma Terms of Use apply to that data.
- Downloads last month
- 18
Model tree for Mannedood/local-system-one-student
Base model
answerdotai/ModernBERT-base