Post
21
๐ป Kodiak-v0.2-1B is out: an open 1B encoder that makes decisions instead of writing text.
Give it a state and typed questions. You get back calibrated answers in one forward pass, and it says "can't tell" when it doesn't know.
On tasks it was never trained for (our frozen eval set v0.2):
โข 0.689 accuracy (3-run mean) vs 0.688 for Qwen3-8B, at ~40ร the speed
โข calibration error 0.085 vs 0.293
โข accuracy mode (3 models averaged): 0.706
New in v0.2: grounding checks, picking an assistant's next tool step from API specs, claim verification, and product relevance.
Known limits are in the model card.
๐ฆ cortex-agent-llc/kodiak-v0.2-1b
๐ฏ Accuracy mode: cortex-agent-llc/kodiak-v0.2-1b-accuracy
๐น๏ธ Demo: comgen42/kodiak-demo
๐ What we built, and what didn't work: https://cortexagent.com/blog/kodiak-v0-2-an-open-1b-decision-model-you-can-download-today
Give it a state and typed questions. You get back calibrated answers in one forward pass, and it says "can't tell" when it doesn't know.
On tasks it was never trained for (our frozen eval set v0.2):
โข 0.689 accuracy (3-run mean) vs 0.688 for Qwen3-8B, at ~40ร the speed
โข calibration error 0.085 vs 0.293
โข accuracy mode (3 models averaged): 0.706
New in v0.2: grounding checks, picking an assistant's next tool step from API specs, claim verification, and product relevance.
Known limits are in the model card.
๐ฆ cortex-agent-llc/kodiak-v0.2-1b
๐ฏ Accuracy mode: cortex-agent-llc/kodiak-v0.2-1b-accuracy
๐น๏ธ Demo: comgen42/kodiak-demo
๐ What we built, and what didn't work: https://cortexagent.com/blog/kodiak-v0-2-an-open-1b-decision-model-you-can-download-today