Thanks, this is a great catch, and you're right on both counts.
I reproduced your numbers (Qwen3-8B 14.7%, large v2 49.6%) and ran your two questions on the v0.2 predictions:
Ranking (share of the random→perfect AURC gap closed, never-seen, n=3,550):
• Kodiak-v0.2-1B: 54.8% / 56.1% / 55.9% (3 seeds)
• v0.2 accuracy mode: 56.6%
• Qwen3-8B: 14.7%
Calibration after giving Qwen a fitted monotone (isotonic) recalibration, fit on the other never-seen tasks and applied to each held-out task:
• Qwen3-8B: 0.293 → 0.178
• v0.2-1B: 0.075–0.094 raw → 0.044–0.058 with the same treatment
So the gap survives, but the raw 0.085 vs 0.293 overstates it, and ranking is the better headline. It's what the cascade actually spends. We'll add a ranking row to the model card, footnote the calibration comparison as raw, and put the v0.2 prediction files in the repo so this can be checked.