The cost of a judging gate is usually quoted as a number. This puts it on a Tetris board.
Three boards get the same piece order, and on every move the same proposal and the same noise β a paired comparison. The gate decides one thing: keep this move, or draw again. Each board gets the same 60 seconds of gate time.
The text-writing gates get through 15β22 moves. The generation-free gate gets through 40β50. The boards that stop simply run out of clock.
It does not win on accuracy: on the same 2,018-question LODO set, JEV scores AUC 0.7350 against ZTC-Judge-27B's 0.7289. The separation is elsewhere. Clock β 2.1 s vs 0.0615 s per call, and on a 200-candidate agent screen one judging call measured 3.206 s generative vs 0.033 s readout, same server. Calibration β a gate is a threshold, and at ECE 0.4985 (vs ZTC 0.0245) a threshold stops carrying information. Mechanism β a text judge can name option 42 when there is no option 42; a scoring readout cannot. Not a lower error rate. No path.
The curve in the ZTC panel is real online fitting, scored prequentially β predict first, learn after β with base weights untouched. Not recursive self-improvement.
Limits, also stated on the page: Laya's AUC and latency are not our measurements and are set equal to JEV's, so calibration is the only measured axis it differs on. The page is a simulation driven by measured constants.
OpenRouter Leaderboard β every model, every provider, one comparable table. Price, precision, uptime, measured latency and language quality on the same axes.
Building it turned up three things.
We graded 330 models on Korean and two axes collapsed.
Honorifics β only 8.5% earn an A Knowledge of Korean institutions β 9.4% Every other axis sits above 31% Fluency hides it. A model can write clean, natural Korean and still attach an honorific to a coffee cup. Fluent and wrong at the same time is worse than obviously broken, because nobody catches it in review.
A 2023 model beats the 2026 flagships. gpt-3.5-turbo-16k scores a perfect 3.00. Korean cannot be inferred from release date, parameter count or English benchmarks β it has to be measured, per model.
Quality, value and speed are three different models. Across five axes, the same model almost never takes two columns.
425 models, latency measured on 329 on a paid API, Korean graded on 330. Three languages, three currencies, daily refresh, open API, no key.