·
AI & ML interests
None yet
Recent Activity
repliedto SeaWolf-AI's post 18 minutes ago The cost of a judging gate is usually quoted as a number. This puts it on a Tetris board.
Three boards get the same piece order, and on every move the same proposal and the same noise — a paired comparison. The gate decides one thing: keep this move, or draw again. Each board gets the same 60 seconds of gate time.
The text-writing gates get through 15–22 moves. The generation-free gate gets through 40–50. The boards that stop simply run out of clock.
It does not win on accuracy: on the same 2,018-question LODO set, JEV scores AUC 0.7350 against ZTC-Judge-27B's 0.7289. The separation is elsewhere. Clock — 2.1 s vs 0.0615 s per call, and on a 200-candidate agent screen one judging call measured 3.206 s generative vs 0.033 s readout, same server. Calibration — a gate is a threshold, and at ECE 0.4985 (vs ZTC 0.0245) a threshold stops carrying information. Mechanism — a text judge can name option 42 when there is no option 42; a scoring readout cannot. Not a lower error rate. No path.
The curve in the ZTC panel is real online fitting, scored prequentially — predict first, learn after — with base weights untouched. Not recursive self-improvement.
Limits, also stated on the page: Laya's AUC and latency are not our measurements and are set equal to JEV's, so calibration is the only measured axis it differs on. The page is a simulation driven by measured constants.
KO / EN / ZH.
https://huggingface.co/spaces/FINAL-Bench/Tetris-JEV-LAYA-ZTC
https://huggingface.co/FINAL-Bench/ZTC-Judge-27B repliedto SeaWolf-AI's post about 1 hour ago The cost of a judging gate is usually quoted as a number. This puts it on a Tetris board.
Three boards get the same piece order, and on every move the same proposal and the same noise — a paired comparison. The gate decides one thing: keep this move, or draw again. Each board gets the same 60 seconds of gate time.
The text-writing gates get through 15–22 moves. The generation-free gate gets through 40–50. The boards that stop simply run out of clock.
It does not win on accuracy: on the same 2,018-question LODO set, JEV scores AUC 0.7350 against ZTC-Judge-27B's 0.7289. The separation is elsewhere. Clock — 2.1 s vs 0.0615 s per call, and on a 200-candidate agent screen one judging call measured 3.206 s generative vs 0.033 s readout, same server. Calibration — a gate is a threshold, and at ECE 0.4985 (vs ZTC 0.0245) a threshold stops carrying information. Mechanism — a text judge can name option 42 when there is no option 42; a scoring readout cannot. Not a lower error rate. No path.
The curve in the ZTC panel is real online fitting, scored prequentially — predict first, learn after — with base weights untouched. Not recursive self-improvement.
Limits, also stated on the page: Laya's AUC and latency are not our measurements and are set equal to JEV's, so calibration is the only measured axis it differs on. The page is a simulation driven by measured constants.
KO / EN / ZH.
https://huggingface.co/spaces/FINAL-Bench/Tetris-JEV-LAYA-ZTC
https://huggingface.co/FINAL-Bench/ZTC-Judge-27B View all activity Organizations
seawolf2357/my-lora-20260110-0842
Text-to-Image
• Updated • 16
seawolf2357/my-lora-20260110-0809
Text-to-Image
• Updated • 10
seawolf2357/Qwen-Image-Edit-Rapid-AIO
Text-to-Image
• Updated Text Generation
• 357B • Updated • 14
Text-to-Speech
• 0.7B • Updated • 140
Text Generation
• 229B • Updated • 15
Image-Text-to-Text
• 1.0B • Updated • 12
Image-Text-to-Text
• 3B • Updated • 12
seawolf2357/Darwin-gpt-ernie-20b
21B • Updated • 12
Text-to-Image
• Updated • 15
• Text-to-Image
• Updated • 18
• • 15
seawolf2357/nsfw-detection
Text-to-Image
• Updated • 19
• • 38
seawolf2357/blingone-lani
Text-to-Image
• Updated • 17
• • 24
Text-to-Image
• Updated • 72
• • 25
Text-to-Image
• Updated • 29
• • 179
seawolf2357/audrey-hepburn
Text-to-Image
• Updated • 31
• • 23
Text-to-Image
• Updated • 56
• • 182
Text-to-Image
• Updated • 19
• Text-to-Image
• Updated • 61
• Text-to-Image
• Updated • 29
• • 26
Text-to-Image
• Updated • 68
• • 4
Text-to-Image
• Updated • 16
• • 25
Text-to-Image
• Updated • 39
• Text-to-Image
• Updated • 41
• seawolf2357/flux-lora-military-artillery-k9
Text-to-Image
• Updated • 106
• • 159
seawolf2357/flux-lora-car-rolls-royce
Text-to-Image
• Updated • 86
• • 188
Text-to-Image
• Updated • 13
• seawolf2357/LLAMA-3-70B-Ko
Text Generation
• 29B • Updated • 14