Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
comgen42 
posted an update about 19 hours ago
Post
53
🐻 Kodiak-v0.2-1B is out: an open 1B encoder that makes decisions instead of writing text.

Give it a state and typed questions. You get back calibrated answers in one forward pass, and it says "can't tell" when it doesn't know.

On tasks it was never trained for (our frozen eval set v0.2):
• 0.689 accuracy (3-run mean) vs 0.688 for Qwen3-8B, at ~40× the speed
• calibration error 0.085 vs 0.293
• accuracy mode (3 models averaged): 0.706

New in v0.2: grounding checks, picking an assistant's next tool step from API specs, claim verification, and product relevance.

Known limits are in the model card.

📦 cortex-agent-llc/kodiak-v0.2-1b
🎯 Accuracy mode: cortex-agent-llc/kodiak-v0.2-1b-accuracy
🕹️ Demo: comgen42/kodiak-demo
📝 What we built, and what didn't work: https://cortexagent.com/blog/kodiak-v0-2-an-open-1b-decision-model-you-can-download-today

The calibration row is the wrong headline. The bigger win is on ranking, and it isn't on the card.

I ran your reports/v02-vs-llm-8b.json (repo @ 1b2c42b2) through your own metrics.

On the never-seen slice, Qwen3-8B states 0.944 mean confidence and gets 0.651 right.
0.944 minus 0.651 is 0.293. That is its whole ECE, to the fourth decimal (0.2926 both ways).
With your 15 equal-mass bins, that only happens when every bin is overconfident.
So it is a pure offset. One monotone recalibration fit on validation, the same kind of step Kodiak gets in calibration.json, removes most of it.

What recalibration can't buy is ranking. Your aurc() doesn't change under any monotone remap.
Scored as the share of the gap between random order and perfect order that each model's confidence closes:

  • Qwen3-8B, never-seen: 14.7%
  • Kodiak large v2 (400M, seed 0), never-seen: 49.6%
  • Qwen 27B on the 1,500 sample: 44.0% (large v2 gets 52.2% on the same sample)

So even the 400M model ranks its own errors about 3.4x better than the 8B.
That is the property your cascade actually spends (t = 0.7 matches the LLM on never-seen with 52% of the calls).

Caveat: these are the 09-27 large v2 rows. The v0.2-1B predictions aren't in the repo yet.

What does v0.2-1B close on that same measure, and does the 0.085 vs 0.293 gap survive once Qwen gets the same fitted recalibration?

·

Thanks, this is a great catch, and you're right on both counts.

I reproduced your numbers (Qwen3-8B 14.7%, large v2 49.6%) and ran your two questions on the v0.2 predictions:

Ranking (share of the random→perfect AURC gap closed, never-seen, n=3,550):
• Kodiak-v0.2-1B: 54.8% / 56.1% / 55.9% (3 seeds)
• v0.2 accuracy mode: 56.6%
• Qwen3-8B: 14.7%

Calibration after giving Qwen a fitted monotone (isotonic) recalibration, fit on the other never-seen tasks and applied to each held-out task:
• Qwen3-8B: 0.293 → 0.178
• v0.2-1B: 0.075–0.094 raw → 0.044–0.058 with the same treatment

So the gap survives, but the raw 0.085 vs 0.293 overstates it, and ranking is the better headline. It's what the cascade actually spends. We'll add a ranking row to the model card, footnote the calibration comparison as raw, and put the v0.2 prediction files in the repo so this can be checked.

Thanks for running it, and for the 3 seeds. Saw the ranking row and the raw footnote land in ac06483d. 55.6% against 14.7% is a clean headline.

One phrase in the footnote is worth a second look: "overconfident by a constant amount."
Per task, what's constant is Qwen's confidence, not its offset. That's what holds isotonic up near 0.178, and it's in your release report (reports/v02-release-vs-llm-8b.json @ 985bc8d).

Across the 10 never-seen tasks, Qwen3-8B's mean confidence only moves from 0.918 to 0.954.
Its accuracy moves from 0.372 (contract_nli) to 0.876 (jailbreak_classification).
Correlation across tasks: +0.12.

So a map fit on the other 9 tasks hands every held-out task roughly the same number. It can't see which task is hard.
Approximating it as "predict the other 9 tasks' pooled accuracy", the n-weighted per-task floor is 0.123.

v0.2-1B (accuracy mode) does see it: task confidence spans 0.511 to 0.849, correlation +0.57.

One thing cuts the other way. Pooled ECE flatters the 1B more than Qwen.
Qwen is overconfident on all 10 tasks (+0.07 to +0.58), so nothing cancels.
The 1B's offsets have mixed signs (contract_nli +0.28, ethics_commonsense +0.15, fin_tweets_topic -0.09, jailbreak -0.09), and they partly cancel inside pooled bins.
n-weighted per-task ECE: Qwen 0.296, 1B 0.121. That's 2.4x, not the 4.7x that pooled 0.293 vs 0.062 suggests.

For the cascade that matters, because t = 0.7 is one global threshold. On contract_nli the 1B averages 0.736 confidence and gets 0.460 right.

Is your 0.178 pooled or per-task, and does the isotonic step leave the 1B's contract_nli offset standing?