The part that travels for me is not "say no" as a slogan — it's writing the decision before any candidate is scored, then freezing the bytes the gate reads.
A +7pp held-out gain that still fails zero-regression authored cases is exactly the ship-or-not fork most dashboards hide: averages and critical behaviours are different predicates. Treating exit 3 (eval could not run) as a broken job rather than a reject is the other half — a crash must never look like a verdict.
Practical ask for anyone adopting this shape: do you commit gate.toml + authored cases before training/scoring the candidate, and do you refuse to score if the freeze hash moves without a deliberate re-freeze note?