Shane Larson PRO
AI & ML interests
Recent Activity
Organizations
Thanks, this is a great catch, and you're right on both counts.
I reproduced your numbers (Qwen3-8B 14.7%, large v2 49.6%) and ran your two questions on the v0.2 predictions:
Ranking (share of the random→perfect AURC gap closed, never-seen, n=3,550):
• Kodiak-v0.2-1B: 54.8% / 56.1% / 55.9% (3 seeds)
• v0.2 accuracy mode: 56.6%
• Qwen3-8B: 14.7%
Calibration after giving Qwen a fitted monotone (isotonic) recalibration, fit on the other never-seen tasks and applied to each held-out task:
• Qwen3-8B: 0.293 → 0.178
• v0.2-1B: 0.075–0.094 raw → 0.044–0.058 with the same treatment
So the gap survives, but the raw 0.085 vs 0.293 overstates it, and ranking is the better headline. It's what the cascade actually spends. We'll add a ranking row to the model card, footnote the calibration comparison as raw, and put the v0.2 prediction files in the repo so this can be checked.
Give it a state and typed questions. You get back calibrated answers in one forward pass, and it says "can't tell" when it doesn't know.
On tasks it was never trained for (our frozen eval set v0.2):
• 0.689 accuracy (3-run mean) vs 0.688 for Qwen3-8B, at ~40× the speed
• calibration error 0.085 vs 0.293
• accuracy mode (3 models averaged): 0.706
New in v0.2: grounding checks, picking an assistant's next tool step from API specs, claim verification, and product relevance.
Known limits are in the model card.
📦 cortex-agent-llc/kodiak-v0.2-1b
🎯 Accuracy mode: cortex-agent-llc/kodiak-v0.2-1b-accuracy
🕹️ Demo: comgen42/kodiak-demo
📝 What we built, and what didn't work: https://cortexagent.com/blog/kodiak-v0-2-an-open-1b-decision-model-you-can-download-today