Thanks for running it, and for the 3 seeds. Saw the ranking row and the raw footnote land in ac06483d. 55.6% against 14.7% is a clean headline.
One phrase in the footnote is worth a second look: "overconfident by a constant amount."
Per task, what's constant is Qwen's confidence, not its offset. That's what holds isotonic up near 0.178, and it's in your release report (reports/v02-release-vs-llm-8b.json @ 985bc8d).
Across the 10 never-seen tasks, Qwen3-8B's mean confidence only moves from 0.918 to 0.954.
Its accuracy moves from 0.372 (contract_nli) to 0.876 (jailbreak_classification).
Correlation across tasks: +0.12.
So a map fit on the other 9 tasks hands every held-out task roughly the same number. It can't see which task is hard.
Approximating it as "predict the other 9 tasks' pooled accuracy", the n-weighted per-task floor is 0.123.
v0.2-1B (accuracy mode) does see it: task confidence spans 0.511 to 0.849, correlation +0.57.
One thing cuts the other way. Pooled ECE flatters the 1B more than Qwen.
Qwen is overconfident on all 10 tasks (+0.07 to +0.58), so nothing cancels.
The 1B's offsets have mixed signs (contract_nli +0.28, ethics_commonsense +0.15, fin_tweets_topic -0.09, jailbreak -0.09), and they partly cancel inside pooled bins.
n-weighted per-task ECE: Qwen 0.296, 1B 0.121. That's 2.4x, not the 4.7x that pooled 0.293 vs 0.062 suggests.
For the cascade that matters, because t = 0.7 is one global threshold. On contract_nli the 1B averages 0.736 confidence and gets 0.460 right.
Is your 0.178 pooled or per-task, and does the isotonic step leave the 1B's contract_nli offset standing?