Β·
AI & ML interests
None yet
Recent Activity
reacted to SeaWolf-AI's post with π€ 1 day ago The cost of a judging gate is usually quoted as a number. This puts it on a Tetris board.
Three boards get the same piece order, and on every move the same proposal and the same noise β a paired comparison. The gate decides one thing: keep this move, or draw again. Each board gets the same 60 seconds of gate time.
The text-writing gates get through 15β22 moves. The generation-free gate gets through 40β50. The boards that stop simply run out of clock.
It does not win on accuracy: on the same 2,018-question LODO set, JEV scores AUC 0.7350 against ZTC-Judge-27B's 0.7289. The separation is elsewhere. Clock β 2.1 s vs 0.0615 s per call, and on a 200-candidate agent screen one judging call measured 3.206 s generative vs 0.033 s readout, same server. Calibration β a gate is a threshold, and at ECE 0.4985 (vs ZTC 0.0245) a threshold stops carrying information. Mechanism β a text judge can name option 42 when there is no option 42; a scoring readout cannot. Not a lower error rate. No path.
The curve in the ZTC panel is real online fitting, scored prequentially β predict first, learn after β with base weights untouched. Not recursive self-improvement.
Limits, also stated on the page: Laya's AUC and latency are not our measurements and are set equal to JEV's, so calibration is the only measured axis it differs on. The page is a simulation driven by measured constants.
KO / EN / ZH.
https://huggingface.co/spaces/FINAL-Bench/Tetris-JEV-LAYA-ZTC
https://huggingface.co/FINAL-Bench/ZTC-Judge-27B reacted to SeaWolf-AI's post with π₯ 1 day ago The cost of a judging gate is usually quoted as a number. This puts it on a Tetris board.
Three boards get the same piece order, and on every move the same proposal and the same noise β a paired comparison. The gate decides one thing: keep this move, or draw again. Each board gets the same 60 seconds of gate time.
The text-writing gates get through 15β22 moves. The generation-free gate gets through 40β50. The boards that stop simply run out of clock.
It does not win on accuracy: on the same 2,018-question LODO set, JEV scores AUC 0.7350 against ZTC-Judge-27B's 0.7289. The separation is elsewhere. Clock β 2.1 s vs 0.0615 s per call, and on a 200-candidate agent screen one judging call measured 3.206 s generative vs 0.033 s readout, same server. Calibration β a gate is a threshold, and at ECE 0.4985 (vs ZTC 0.0245) a threshold stops carrying information. Mechanism β a text judge can name option 42 when there is no option 42; a scoring readout cannot. Not a lower error rate. No path.
The curve in the ZTC panel is real online fitting, scored prequentially β predict first, learn after β with base weights untouched. Not recursive self-improvement.
Limits, also stated on the page: Laya's AUC and latency are not our measurements and are set equal to JEV's, so calibration is the only measured axis it differs on. The page is a simulation driven by measured constants.
KO / EN / ZH.
https://huggingface.co/spaces/FINAL-Bench/Tetris-JEV-LAYA-ZTC
https://huggingface.co/FINAL-Bench/ZTC-Judge-27B View all activity Organizations
view article Inside the JEV Ecosystem: 13 Answer Verifiers on One Test Set
mayafree
β’ β’ 12
published an article about 2 months ago view article Model Genome: Fingerprinting Whether an LLM Was Trained From Scratch or Derived
mayafree
β’ β’ 18
view article Open NPC AI: Design Principles of a Proto-AGI Society