Leo 1.7B (leo-1.7b-v3)
Leo is an open-weight decision model. You send a state (text or JSON) and typed questions (choice, score, or noul yes/no). It returns a calibrated probability distribution for every question from a single forward pass, with no text generation. Typical uses: routing, classification, moderation triage, grading on a rubric, policy checks, and the next-action choices of an agent.
The request and response format follows the POST /v1/systemone API that TypeSafe documents for its Jev models, so existing clients work after a base-URL change. Leo is an independent project: it is not affiliated with TypeSafe, contains none of its weights, and was not trained on Jev outputs. Jev was called only to score it next to Leo on identical requests.
- Code, training pipeline, benchmarks: github.com/SuparvaCode/leo
- Versions in this repo:
main= leo-1.7b-v3 (recommended),v4-experimental= leo-1.7b-v4
Model details
| Developer | Suparva Baranwal |
| Version | leo-1.7b-v3 |
| Model type | Decoder-only transformer run prefill-only, LoRA adapter, listwise pointer head |
| Base model | Qwen/Qwen3-1.7B-Base (Apache-2.0), revision ea980cb0a6 |
| Parameters | about 1.7B in the base; 20.3M trained (LoRA r16 on every attention and MLP projection, plus markers and head) |
| Files | adapter/ (LoRA, 70 MB), leo_head.safetensors (head and marker embeddings, 12 MB), leo_config.json (calibration and metadata), leo/ (inference code) |
| Inputs | a state (string, JSON object or array) and many typed questions per request |
| Outputs | per question: choice + probabilities + confidence, or score + probabilities + legend, or noul = P(yes) |
| Context | trained on states up to 2,048 tokens; browser states up to about 12k tokens were used in evaluation |
| Languages | English plus 50+ languages via MASSIVE, SIB-200 and multilingual sentiment data; much weaker on low-resource languages |
| Licence | Apache-2.0 (weights and code); training datasets keep their own terms |
Results at a glance
Measured on one machine, with Jev 1.13 (live API) (jev-1.13.0, September 2026) answering exactly the same requests. Per-task tables and the protocol follow further down.
| benchmark | leo-1.7b-v3 | leo-1.7b-v4 | Jev 1.13 (live API) |
|---|---|---|---|
| Browser tasks passed (jev-ultrafast, 7 tasks x 3 runs) | 18/21 | 12/21 | 18/21 (one per matched session) |
| False DONE rate, held-out browser screens (lower is better) | 0.067 | 0.000 | not measured |
| JevBench, all 231 items | 0.697 | 0.680 | 0.861 |
| JevBench, standard tier | 0.958 | 0.903 | 0.986 |
| JevBench, hard tier | 0.396 | 0.396 | 0.721 |
| Held-out classification, mean accuracy (4 datasets, zero-shot) | 0.628 | 0.627 | 0.689 |
| Multilingual, 2,049 items in 15 languages (blind) | 0.560 | 0.557 | 0.857 |
| Calibration error (ECE), held-out mean (lower is better) | 0.099 | 0.094 | 0.167 |
In short: leo-1.7b-v3 ties Jev on the browser suite (18/21 vs 18/21), is close on JevBench's easy and standard tiers, and beats Jev on tweet topics. It is clearly behind on hard reasoning (long policies, ambiguity, trade-offs, multi-hop), knowledge-heavy exams and low-resource languages. Most of that gap comes from the size of the 1.7B base model.
The v4 experiment: fixing false DONE
Problem in v3. leo-1.7b-v3 declares a browser task DONE once its action history covers every part of the goal, whatever the page shows. On Google Flights it stopped with DONE in all three runs without completing the search, with P(DONE) = 0.986 while only a seat-class list was on screen, and 0.999 on a blank render.
What v4 changed. leo-1.7b-v4 continues v3's weights (1 epoch, half the learning rate) on 30,000 replayed v3 requests plus 9,000 steps from a new simulator family (leo/data/browser_evidence.py) in which DONE is only right with visible evidence: close the open dropdown, wait out blank or loading renders, resubmit when results show an older search, report BLOCKED when a requested value is not offered.
What happened.
- The targeted flaw is fixed in simulation and largely on the real site. The false-DONE rate on held-out screens fell from 0.067 to 0.000. On live Google Flights, v4 no longer claims success: all 3 runs ended with an honest BLOCKED (the seat-class state that fooled v3 went from P(DONE) 0.986 to 0.655).
- It introduced a new failure. On jev-ultrafast's two reading-room tasks (open one article, no search) v4 opens the right article and then goes back to the list, over and over: 0/6 passed, every run hit the 60-action budget. The likely cause: every goal in the new family includes a search, so an item page reached without a search in the history always meant "go back" in training. Browser total: 18/21 for v3, 12/21 for v4.
- JevBench fell from 0.697 to 0.680 (standard tier 0.958 to 0.903), mostly policy and trap items where v4 now answers "yes" or skips "other"/"unknown". Held-out classification and multilingual accuracy are unchanged within noise.
Decision. leo-1.7b-v3 stays the recommended model on main. leo-1.7b-v4 is published on the v4-experimental branch for research and for agent loops that need an honest BLOCKED more than they need open-an-item tasks. The fix for the next version is to mix goals without a search step into the evidence family and to re-check the reading-room tasks before release.
Quick start
pip install torch transformers peft safetensors numpy pydantic huggingface_hub
import sys
from huggingface_hub import snapshot_download
path = snapshot_download("Suparva/leo-1.7b") # revision="v4-experimental" for v4
sys.path.insert(0, path) # the repo ships its own inference code in leo/
from leo.infer import Leo
leo = Leo.load(path, dtype="bf16") # use "fp32" on CPU; the Qwen3 base is fetched on first use
out = leo.system_one(
{
"ticket": {
"subject": "Charged twice",
"body": "I was billed twice this month. Please fix it today."
}
},
{
"route": {
"type": "choice",
"instructions": "Which team should handle `ticket`?",
"criteria": {
"billing": "charges, refunds, invoices",
"technical": "bugs, outages",
"other": None
}
},
"urgent": {
"type": "noul",
"instructions": "Does the customer need an answer today?"
},
"anger": {
"type": "score",
"instructions": "How upset is the customer?",
"criteria": [
"calm",
"annoyed",
"very angry"
]
}
},
)
print(out)
Real output of this request (the exported folder, loaded on its own). The Python API returns 4 decimals; the HTTP server returns TypeSafe's exact shape by default (2 decimals summing to exactly 1).
{
"model": "leo-1.7b-v3",
"answers": {
"route": {
"type": "choice",
"choice": "billing",
"probabilities": {
"billing": 0.9026,
"technical": 0.0374,
"other": 0.06
},
"confidence": 0.8539
},
"urgent": {
"type": "noul",
"noul": 0.8721
},
"anger": {
"type": "score",
"score": 1.1116,
"legend": {
"0": "calm",
"1": "annoyed",
"2": "very angry"
},
"probabilities": {
"0": 0.0432,
"1": 0.802,
"2": 0.1548
},
"confidence": 0.7029
}
},
"usage": {
"input_tokens": 90,
"output_tokens": 136
},
"latency_ms": 847.38
}
leo.predict_many([{"state": ..., "questions": ...}, ...]) batches many requests.
HTTP server
pip install fastapi uvicorn httpx
cd <snapshot path>
python -m leo.serve --model . --port 8000 # binds 127.0.0.1
LEO_API_KEY=change-me python -m leo.serve --model . --host 0.0.0.0 # a key is required off loopback
POST /v1/systemone and GET /v1/models follow TypeSafe's API: bearer auth whenever LEO_API_KEY is set, 422 on invalid requests, body-size and question-count limits. The server refuses to bind a public address without a key. Flags:
--order-views 2averages each choice over two option orders: about half the option-order sensitivity, for 1.4 to 2.8 times the latency.--dtype fp32gives answers that do not depend on which other questions share the request (bf16 moves probabilities by up to about 0.02).--precisereturns 4-decimal probabilities pluslatency_msinstead of the TypeSafe-exact shape.
leo.client.SystemOneClient talks to a Leo server or any other /v1/systemone endpoint.
Question types
| type | criteria |
answer |
|---|---|---|
choice |
object mapping option key to a description (string, object, array or null), 2 to 255 options | choice, probabilities over the keys, confidence = (K·p_max − 1)/(K − 1) |
score |
ordered array of 2 to 10 level descriptions | score = Σ i·pᵢ, probabilities, confidence, legend |
noul |
optional {"true": ..., "false": ...} |
noul = P(yes) |
Instructions can be a string or any JSON value and can refer to parts of the state by path, such as `ticket.body`. Question IDs are never shown to the model. Questions cannot see each other, so adding or removing one does not change the others.
How it works
- Backbone:
Qwen/Qwen3-1.7B-Baserun prefill-only (no decoding), adapted with LoRA r16 / alpha 32. - Packed layout: the state is encoded once; each question follows it in the same sequence under a block-causal mask, with position ids restarting after the state. Questions share the state encoding but cannot attend to each other.
- Reserved markers: state, question, option and decision boundaries are trainable embeddings outside the vocabulary, so text inside the state cannot forge them.
- Listwise pointer readout: a small head scores each option's end marker against the question's decision marker, which comes after the full option list, so options are judged together.
- Proper-scoring-rule training: log loss against the label distribution plus a ranked probability score term for
scorequestions. - Calibration: one temperature per question type and option-count bucket, fitted on dev data (dev top-label ECE 0.014). Confidence uses the formulas TypeSafe publishes.
Evaluation
- Every Jev number comes from live calls on byte-identical requests (same state, instructions, option keys and option order), September 2026. Jev responses were used only for scoring, never for training, prompt tuning or checkpoint selection. Checkpoints were selected on dev loss alone.
- Held-out classification and the multilingual suite are blind: none of those datasets or task families were trained on. JevBench items were read while designing the data, so treat that score as seen.
- The browser suite runs TypeSafe's own agent unchanged and checks the final page independently. Live sites (Wikipedia, Google Flights) can change between sessions; Jev's own Google Flights result differed between the two sessions.
- TYPE_TEXT values in the browser suite come from the same local Qwen3-1.7B helper for both arms.
| benchmark | leo-1.7b-v3 | leo-1.7b-v4 | Jev 1.13 (live API) |
|---|---|---|---|
| Browser tasks passed (jev-ultrafast, 7 tasks x 3 runs) | 18/21 | 12/21 | 18/21 (one per matched session) |
| False DONE rate, held-out browser screens (lower is better) | 0.067 | 0.000 | not measured |
| JevBench, all 231 items | 0.697 | 0.680 | 0.861 |
| JevBench, standard tier | 0.958 | 0.903 | 0.986 |
| JevBench, hard tier | 0.396 | 0.396 | 0.721 |
| Held-out classification, mean accuracy (4 datasets, zero-shot) | 0.628 | 0.627 | 0.689 |
| Multilingual, 2,049 items in 15 languages (blind) | 0.560 | 0.557 | 0.857 |
| Calibration error (ECE), held-out mean (lower is better) | 0.099 | 0.094 | 0.167 |
Browser automation
browser-use/jev-ultrafast is TypeSafe's own open-source browser agent: each step sends one /v1/systemone request with an operation question and one target question per operation. The agent loop, request builder, response validation, DOM snapshot and executor were left unchanged; only the model answering the request differs. A run passes only when the agent stops with DONE and independent checks on the final page hold (a DONE alone is never trusted). 3 runs per task; each Leo version had its own session with Jev, the two arms alternating within each repeat (same Chrome profile, viewport, 60-action budget, 120 s limit). Five tasks use jev-ultrafast's local fixture, two use live websites.
Cells are passed runs, then median time / model decisions.
| task | leo-1.7b-v3 | Jev (same session) | leo-1.7b-v4 | Jev (same session) |
|---|---|---|---|---|
| travel-casa-flora | 3/3 · 5.0 s / 6 | 3/3 · 3.6 s / 6 | 3/3 · 5.3 s / 6 | 3/3 · 3.5 s / 6 |
| research-finite-choices | 3/3 · 5.7 s / 12 | 3/3 · 0.8 s / 2 | 0/3 · 33.0 s / 61 | 3/3 · 0.9 s / 2 |
| travel-glasshouse | 3/3 · 5.0 s / 6 | 3/3 · 3.5 s / 6 | 3/3 · 5.3 s / 6 | 2/3 · 3.5 s / 6 |
| travel-serra-lodge | 3/3 · 4.3 s / 5 | 3/3 · 3.0 s / 5 | 3/3 · 4.4 s / 5 | 3/3 · 3.1 s / 5 |
| research-confidence | 3/3 · 0.7 s / 2 | 3/3 · 0.8 s / 2 | 0/3 · 31.6 s / 61 | 3/3 · 0.9 s / 2 |
| wikipedia-godel (live site) | 3/3 · 11.6 s / 5 | 3/3 · 4.0 s / 4 | 3/3 · 10.1 s / 4 | 3/3 · 4.3 s / 5 |
| google-flights (live site) | 0/3 · 18.7 s / 14 | 0/3 · 1.6 s / 3 | 0/3 · 48.5 s / 36 | 1/3 · 13.5 s / 17 |
| all | 18/21 | 18/21 | 12/21 | 18/21 |
How the failed runs ended:
- leo-1.7b-v3, google-flights: 3 run(s) stopped with DONE
- Jev (session with leo-1.7b-v3), google-flights: 3 run(s) stopped with BLOCKED
- leo-1.7b-v4, research-finite-choices: 3 run(s) hit the 60-action budget
- leo-1.7b-v4, research-confidence: 3 run(s) hit the 60-action budget
- leo-1.7b-v4, google-flights: 3 run(s) stopped with BLOCKED
- Jev (session with leo-1.7b-v4), travel-glasshouse: 1 run(s) stopped with DONE
- Jev (session with leo-1.7b-v4), google-flights: 2 run(s) stopped with BLOCKED
False-DONE probe
scripts/done_probe.py replays 400 simulated browser steps from two website themes that never appear in any training set, on screens where the action history can look finished while the page proves nothing.
| screen | n | leo-1.7b-v3 false DONE | leo-1.7b-v4 false DONE | leo-1.7b-v3 step accuracy | leo-1.7b-v4 step accuracy |
|---|---|---|---|---|---|
| blank page between two renders (right move: WAIT) | 18 | 0.556 | 0.000 | 0.444 | 1.000 |
| opened item (DONE is right if it matches) | 18 | n/a | n/a | 0.667 | 1.000 |
| settled results page (DONE is right if it matches) | 45 | 0.000 | 0.000 | 0.822 | 1.000 |
| results still loading (WAIT) | 18 | 0.111 | 0.000 | 0.889 | 1.000 |
| a dropdown list covers the page | 82 | 0.085 | 0.000 | 0.756 | 1.000 |
| a counter pop-over with its own Done button | 65 | 0.031 | 0.000 | 0.723 | 1.000 |
| results still show the previous search | 19 | 0.053 | 0.000 | 0.474 | 1.000 |
| search form not submitted yet | 135 | 0.015 | 0.000 | 0.556 | 1.000 |
| all screens | 400 | 0.067 | 0.000 | 0.665 | 1.000 |
Real Google Flights states from leo-1.7b-v3's recorded runs, replayed as sent:
| recorded state | leo-1.7b-v3 P(DONE) / answer | leo-1.7b-v4 P(DONE) / answer |
|---|---|---|
| 26 elements: Skip to main content / Accessibility feedback / Explore | 0.378 / CLICK | 0.001 / CLICK |
| 4 elements: Economy / Premium economy / Business / First | 0.986 / DONE | 0.655 / DONE |
| blank page (0 elements) | 0.999 / DONE | 0.092 / WAIT |
None of these pages show flight results, so DONE is wrong on all of them. The held-out themes come from the same simulator family as v4's new training data, so the probe is easier than real sites; the browser suite above is the real test.
JevBench (public items)
231 typed decisions from fstandhartinger/jevbench (commit 1bcc55eb6c). JevBench items were read while designing the training data, so this is not a blind score.
| system | easy (48) | standard (72) | hard (111) | all (231) | ECE |
|---|---|---|---|---|---|
| leo-1.7b-v3 | 1.000 | 0.958 | 0.396 | 0.697 | 0.101 |
| leo-1.7b-v4 | 1.000 | 0.903 | 0.396 | 0.680 | 0.132 |
| Jev 1.13 (live API) | 1.000 | 0.986 | 0.721 | 0.861 | 0.057 |
Accuracy by task family
| tier / family | n | leo-1.7b-v3 | leo-1.7b-v4 | Jev 1.13 (live API) |
|---|---|---|---|---|
| easy/extraction | 12 | 1.00 | 1.00 | 1.00 |
| easy/fact | 12 | 1.00 | 1.00 | 1.00 |
| easy/intent | 12 | 1.00 | 1.00 | 1.00 |
| easy/tool_selection | 12 | 1.00 | 1.00 | 1.00 |
| hard/adversarial | 6 | 0.33 | 0.33 | 1.00 |
| hard/ambiguous | 7 | 0.14 | 0.14 | 0.86 |
| hard/judge_hard | 17 | 0.53 | 0.53 | 0.71 |
| hard/long_policy | 19 | 0.21 | 0.32 | 0.63 |
| hard/multi_hop | 18 | 0.44 | 0.50 | 0.83 |
| hard/probability | 10 | 0.40 | 0.30 | 0.70 |
| hard/routing_hard | 5 | 1.00 | 1.00 | 1.00 |
| hard/temporal_numeric | 15 | 0.20 | 0.20 | 0.27 |
| hard/tradeoff | 6 | 0.17 | 0.17 | 0.83 |
| hard/trap | 8 | 0.88 | 0.62 | 1.00 |
| standard/adequacy | 12 | 0.83 | 0.83 | 1.00 |
| standard/extraction | 12 | 1.00 | 0.92 | 1.00 |
| standard/intent | 12 | 1.00 | 0.92 | 1.00 |
| standard/ordinal | 12 | 1.00 | 1.00 | 1.00 |
| standard/policy | 12 | 1.00 | 0.83 | 0.92 |
| standard/routing | 12 | 0.92 | 0.92 | 1.00 |
Held-out classification (zero-shot)
Four datasets whose sources and task families were never trained on, with the protocol of elcronos/jev-vs-open-decision-models: emotion (6 labels), tweet_topic (6), fin_topic (20 financial-news topics), daily_dialog (7 dialogue emotions, 82% "no emotion"). Cells are accuracy / macro-F1 / ECE.
| system | emotion | tweet_topic | fin_topic | daily_dialog | mean accuracy |
|---|---|---|---|---|---|
| leo-1.7b-v3 | 0.564 / 0.469 / 0.059 | 0.845 / 0.703 / 0.123 | 0.456 / 0.473 / 0.046 | 0.647 / 0.326 / 0.169 | 0.628 |
| leo-1.7b-v4 | 0.557 / 0.473 / 0.070 | 0.826 / 0.684 / 0.141 | 0.479 / 0.496 / 0.031 | 0.644 / 0.324 / 0.136 | 0.627 |
| Jev 1.13 (live API) | 0.587 / 0.503 / 0.280 | 0.790 / 0.693 / 0.064 | 0.669 / 0.627 / 0.168 | 0.710 / 0.385 / 0.156 | 0.689 |
Multilingual (blind, evaluation only)
Belebele reading comprehension, MMMLU (professionally translated MMLU) and INCLUDE (native regional exams): 50 items per language and suite, one choice request each. None were used for training; Belebele passages that share text with the SIB-200 training data were dropped.
| system | Belebele | MMMLU (knowledge) | INCLUDE (knowledge) | all | ECE |
|---|---|---|---|---|---|
| leo-1.7b-v3 | 0.692 | 0.479 | 0.491 | 0.560 | 0.101 |
| leo-1.7b-v4 | 0.680 | 0.471 | 0.503 | 0.557 | 0.133 |
| Jev 1.13 (live API) | 0.917 | 0.861 | 0.775 | 0.857 | 0.013 |
belebele accuracy by language
| language | leo-1.7b-v3 | leo-1.7b-v4 | Jev 1.13 (live API) |
|---|---|---|---|
| Arabic | 0.64 | 0.62 | 0.92 |
| Bengali | 0.64 | 0.60 | 0.92 |
| Chinese | 0.84 | 0.82 | 0.96 |
| English | 0.90 | 0.90 | 0.94 |
| French | 0.82 | 0.78 | 0.94 |
| German | 0.82 | 0.84 | 0.94 |
| Hindi | 0.62 | 0.62 | 0.84 |
| Indonesian | 0.70 | 0.72 | 0.94 |
| Italian | 0.74 | 0.70 | 0.92 |
| Japanese | 0.74 | 0.72 | 0.94 |
| Korean | 0.70 | 0.70 | 0.96 |
| Portuguese | 0.80 | 0.76 | 0.94 |
| Spanish | 0.78 | 0.70 | 0.94 |
| Swahili | 0.44 | 0.42 | 0.92 |
| Yoruba | 0.20 | 0.30 | 0.74 |
mmmlu accuracy by language
| language | leo-1.7b-v3 | leo-1.7b-v4 | Jev 1.13 (live API) |
|---|---|---|---|
| Arabic | 0.48 | 0.44 | 0.82 |
| Bengali | 0.36 | 0.32 | 0.84 |
| Chinese | 0.58 | 0.52 | 0.90 |
| French | 0.50 | 0.56 | 0.90 |
| German | 0.54 | 0.50 | 0.86 |
| Hindi | 0.42 | 0.40 | 0.80 |
| Indonesian | 0.58 | 0.52 | 0.92 |
| Italian | 0.50 | 0.52 | 0.94 |
| Japanese | 0.44 | 0.42 | 0.90 |
| Korean | 0.38 | 0.38 | 0.86 |
| Portuguese | 0.56 | 0.56 | 0.94 |
| Spanish | 0.62 | 0.56 | 0.96 |
| Swahili | 0.36 | 0.34 | 0.76 |
| Yoruba | 0.38 | 0.56 | 0.66 |
include accuracy by language
| language | leo-1.7b-v3 | leo-1.7b-v4 | Jev 1.13 (live API) |
|---|---|---|---|
| Arabic | 0.41 | 0.33 | 0.71 |
| Bengali | 0.42 | 0.44 | 0.66 |
| Chinese | 0.70 | 0.76 | 0.92 |
| French | 0.44 | 0.46 | 0.78 |
| German | 0.48 | 0.52 | 0.58 |
| Hindi | 0.40 | 0.48 | 0.76 |
| Indonesian | 0.60 | 0.62 | 0.94 |
| Italian | 0.54 | 0.50 | 0.90 |
| Japanese | 0.52 | 0.58 | 0.92 |
| Korean | 0.38 | 0.40 | 0.70 |
| Portuguese | 0.46 | 0.46 | 0.72 |
| Spanish | 0.54 | 0.48 | 0.70 |
Robustness and speed (leo-1.7b-v3)
Option-order sensitivity (how often the top answer changes when the options are shuffled), whether questions sharing a request influence each other, and latency. Leo was timed on one NVIDIA GeForce RTX 3050 (8 GB laptop GPU). Jev's model time is the server time TypeSafe reports; its end-to-end time includes the network round trip from this machine. "2 views" is --order-views 2.
| probe | leo-1.7b-v3 bf16 | leo-1.7b-v3 bf16 2 views | Jev (live) |
|---|---|---|---|
| option-order flip rate, emotion (300 rows x 5 shuffles) | 0.100 | 0.055 | 0.018 |
| mean total-variation shift, emotion | 0.063 | 0.032 | 0.036 |
| option-order flip rate, fin_topic (300 rows x 5 shuffles) | 0.217 | 0.116 | 0.069 |
| mean total-variation shift, fin_topic | 0.141 | 0.058 | 0.063 |
| 4 questions together vs one at a time: max prob. difference (100 states) | 0.0236 | 0.0195 | 0.1300 |
| ... top answers that changed | 2 | 3 | 1 |
| model p50 / p95 ms, short state, 1 question(s) | 37 / 46 | 36 / 41 | 56 / 84 |
| end-to-end p50 / p95 ms incl. network, short state, 1 question(s) | - | - | 324 / 365 |
| model p50 / p95 ms, short state, 10 question(s) | 91 / 92 | 142 / 144 | 64 / 83 |
| end-to-end p50 / p95 ms incl. network, short state, 10 question(s) | - | - | 336 / 357 |
| model p50 / p95 ms, short state, 50 question(s) | 392 / 393 | 1082 / 1086 | 78 / 133 |
| end-to-end p50 / p95 ms incl. network, short state, 50 question(s) | - | - | 358 / 439 |
| model p50 / p95 ms, long state, 1 question(s) | 86 / 87 | 91 / 92 | 60 / 105 |
| end-to-end p50 / p95 ms incl. network, long state, 1 question(s) | - | - | 333 / 373 |
| model p50 / p95 ms, long state, 10 question(s) | 142 / 142 | 186 / 187 | 62 / 113 |
| end-to-end p50 / p95 ms incl. network, long state, 10 question(s) | - | - | 336 / 415 |
| model p50 / p95 ms, long state, 50 question(s) | 460 / 461 | 1082 / 1085 | 72 / 110 |
| end-to-end p50 / p95 ms incl. network, long state, 50 question(s) | - | - | 351 / 382 |
Limitations and known flaws
Please read these before relying on Leo.
- False DONE in browser agents. On Google Flights, leo-1.7b-v3 stopped with DONE in 3 of 3 runs without completing the search. It can declare a task done when its action history looks complete even if the page shows no evidence (an open dropdown, a blank render). Always verify outcomes independently, as jev-ultrafast recommends, and never let a DONE trigger anything irreversible. The v4 experiment above targets this.
- Hard reasoning. JevBench hard tier: 0.396 vs Jev 0.721. Weakest on ambiguous items, long policies with exceptions, trade-offs and multi-hop lookups.
- World knowledge. MMMLU 0.479 and INCLUDE 0.491, far below Jev. Do not use it as a knowledge source; put the facts it needs into the state.
- Low-resource languages. Accuracy drops steeply outside the major languages (Yoruba and Swahili are near chance on several suites). Test on your language before deploying.
- Option order. The top answer changes more often than Jev's when options are shuffled, especially with many similar options. Use
--order-views 2when that matters. - Catch-all labels such as "other" or "no emotion" are picked less often than they are right (low macro-F1 on daily_dialog).
- Calibration away from home. Dev ECE is about 0.01; held-out ECE ranges from about 0.03 to 0.17. Re-fit the temperatures on a few hundred labelled examples of your own traffic before thresholding on probabilities.
scorequestions are the weakest type (dev accuracy about 0.65); treat the expected score as a soft signal.- Many questions per request cost more time on Leo than on Jev's servers (see the latency table).
- Long states beyond the 2,048 training tokens work but are less tested.
Intended use
Good fits: ticket and message routing, intent and topic classification, moderation triage, policy checks with the policy in the state, yes/no checks over documents, rubric grading, next-action choices for agents that verify outcomes, research on calibrated decision models, and a local, private backend for /v1/systemone clients.
Out of scope: decisions about people's health, legal status, credit, housing or employment without qualified human review; sole-filter safety moderation; questions that need world knowledge the state does not contain; unattended agents that can spend money or change accounts.
Training
- leo-1.7b-v3: 84,315 requests (130,286 questions), 2.0 epochs, 6,556 optimizer steps, learning rate 0.0002, from
Qwen/Qwen3-1.7B-Base. - leo-1.7b-v4: warm start from leo-1.7b-v3; 39,000 requests (30,000 replayed from v3's training set, 9,000 new evidence-family steps), 1 epoch, learning rate 1e-4; released checkpoint = step 1,200 of 1,911, chosen by calibrated dev loss.
- fp16 mixed precision, data-parallel on 2x NVIDIA T4 (Kaggle). States cut at 2048 tokens, packed rows at 3072 tokens. About 31 T4 GPU-hours.
- Option order shuffled; option keys and descriptions varied (bare, described, snake_case, letters, numbers); none/other options appear both when right and when wrong, so they are not a shortcut.
Training data
Human-labelled public datasets (check each licence before commercial use; some are unknown or custom):
| source | dataset | licence | requests |
|---|---|---|---|
| ag_news | fancyzhx/ag_news |
unknown | 2,500 |
| dbpedia_14 | fancyzhx/dbpedia_14 |
CC BY-SA 3.0 | 2,000 |
| yahoo_answers | community-datasets/yahoo_answers_topics |
unknown | 2,500 |
| banking77 | legacy-datasets/banking77 |
CC BY 4.0 | 2,500 |
| clinc_oos | clinc/clinc_oos |
CC BY 3.0 | 2,500 |
| snli | stanfordnlp/snli |
CC BY-SA 4.0 | 2,000 |
| multi_nli | nyu-mll/multi_nli |
mixed (see dataset card) | 2,500 |
| mrpc | nyu-mll/glue |
other (GLUE / MRPC terms) | 1,000 |
| boolq | google/boolq |
CC BY-SA 3.0 | 2,000 |
| arc_easy | allenai/ai2_arc |
CC BY-SA 4.0 | 1,200 |
| arc_challenge | allenai/ai2_arc |
CC BY-SA 4.0 | 800 |
| commonsense_qa | tau/commonsense_qa |
MIT | 2,000 |
| imdb | stanfordnlp/imdb |
other | 1,500 |
| yelp | Yelp/yelp_review_full |
other (Yelp dataset terms) | 2,500 |
| sms_spam | ucirvine/sms_spam |
unknown on card | 1,200 |
| prompt_injections | deepset/prompt-injections |
Apache-2.0 | 546 |
| jailbreak | jackhhao/jailbreak-classification |
Apache-2.0 | 900 |
| helpsteer2 | nvidia/HelpSteer2 |
CC BY 4.0 | 2,000 |
| massive | AmazonScience/massive |
CC BY 4.0 | 11,146 |
| sib200 | Davlan/sib200 |
CC BY-SA 4.0 | 6,384 |
| ml_sentiment | tyqiangz/multilingual-sentiments |
Apache-2.0 | 4,607 |
| mind2web | osunlp/Mind2Web |
CC BY 4.0 | 7,132 |
Code-generated families (labels computed by code, not by a model):
| family | what it teaches | requests |
|---|---|---|
| syn_browser | browser-agent episodes on simulated sites, in jev-ultrafast's exact request format | 12,000 |
| syn_field_reference | questions that point at a JSON field by path | 1,800 |
| syn_policy | policy compliance checks | 1,800 |
| syn_unknowable | questions the state cannot answer | 500 |
| syn_conversation | multi-turn conversations | 900 |
| syn_multi_hop | multi-hop lookups across records | 1,500 |
| syn_long_state | a relevant fact inside long irrelevant state | 1,000 |
| syn_injection | instructions injected into state | 700 |
| syn_numeric | counting, arithmetic and date comparison | 1,500 |
| syn_policy_exceptions | policies with exceptions | 1,200 |
| syn_browser_evidence (v4 only) | browser screens where DONE needs visible evidence (overlays, pop-overs, blank and loading renders, stale or unsubmitted results, values a site does not offer) | 9,000 |
Never used for training: any output of TypeSafe's models, the four held-out datasets, Belebele, MMMLU, INCLUDE, JevBench, and the sites and goals of the browser benchmark.
Versions
| version | branch | notes |
|---|---|---|
| leo-1.7b-v3 | main |
recommended |
| leo-1.7b-v4 | v4-experimental |
honest BLOCKED instead of false DONE; fails open-an-item tasks; lower JevBench |
| leo-0.6b-v0 / v1 | not released | Qwen3-0.6B prototypes |
Licence
Weights and code: Apache-2.0; the Qwen3 base is Apache-2.0 too. The training mix includes datasets under CC BY-SA, custom and unknown terms (listed above). For a clean licence chain, retrain on the permissive subset with the pipeline on GitHub.
Citation
@misc{leo2026,
title = {Leo: an open-weight decision model with calibrated typed answers},
author = {Baranwal, Suparva},
year = {2026},
url = {https://huggingface.co/Suparva/leo-1.7b},
note = {Code: https://github.com/SuparvaCode/leo}
}
Acknowledgements
Qwen for the base model; browser-use/jev-ultrafast, fstandhartinger/jevbench and elcronos/jev-vs-open-decision-models for the benchmarks; TypeSafe for documenting the /v1/systemone format publicly.
Model tree for Suparva/leo-1.7b
Base model
Qwen/Qwen3-1.7B-Base