Your JevBench table already prices the 0.96 gate.
0.8B alone 0.723, 4B alone 0.835, router 0.797 on the 231 public items. A gate that escalated a blind 70% of items would land at 0.801. Per tier it holds too: original 65 vs 65.1 blind, hard 71 vs 72.1 blind.
So on this set the confidence band carries no routing signal. That is what a 0.95 soft label predicts. The cross-entropy optimum puts the top option at 0.95, so a 0.96 threshold sits above the number the model was trained to output. G14 at 0.94 sat below it, and most steps stayed local at p50 0.57 s.
This assumes JevBench escalates near the desktop's 70%. Your routing records would give the real rate.
One check before the distillation round: what does the 4B alone score on bench v25? If it is also 39/39, the G14 to G18b gain is the 4B answering 70% of steps, not the router choosing well. And a 0.8B distilled from the 4B's probabilities takes the 4B's calibration, Brier 0.269, as its target.
Have you swept the threshold on the public items to see where the router actually beats the blind mixture?