Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
SeaWolf-AIΒ 
posted an update about 13 hours ago
Post
1187
🧠 We just released Darwin-27B-ZTC, a judgment engine that reaches a verdict without generating anything.

Most LLMs answer by generating, decoding one token at a time. Darwin-27B-ZTC takes a different route.

βš™οΈ How it works
πŸ”Ή It makes its call in a single forward pass.
πŸ”Ή Zero generated tokens, and no decoding loop.
πŸ”Ή That keeps latency and cost far below what a generative model needs.

🎯 What it judges
πŸ”Ή It handles several question types: free-form correctness (noul), multiple choice (choice), and scoring (score).
πŸ”Ή For each one it hands back a calibrated confidence, not just an answer.

πŸ“Š How well calibrated (measured)
πŸ”Ή KL 0.204, Brier 0.097, so the confidence it reports lines up with what actually happens.
πŸ”Ή 0.743 accuracy (zero-shot, general split), across 2,000 judgments with zero errors.
πŸ”Ή By type: noul 0.847, choice 0.723, score 0.675.
πŸ”Ή None of the benchmark's train split went into it. It is pure zero-shot.

πŸš€ Where it fits
πŸ”Ή Grading at scale, model routing, safety gating, anywhere you want a fast decision without paying for generation.

πŸ† It currently sits at #1 on the official typed-decisions leaderboard on Hugging Face (0.743 accuracy, zero-shot).

πŸ”— Links
Model: FINAL-Bench/Darwin-27B-ZTC
Leaderboard: LocalLLaMA/typed-decisions

Curious to hear what you make of the single-pass, no-generation approach. πŸ™Œ

Your per-type table and your 0.743 headline look like they come from different runs.

Weight the rows by their decision counts:
600 x 0.847 + 600 x 0.723 + 800 x 0.675 = 1,482 of 2,000 = 0.741.
Push every row to the top of its rounding band and it still only reaches 0.7415. It cannot round to 0.743.

Same arithmetic on the README's own breakdowns lands exactly:
Decider 1 0.7675 -> 0.768, Liquid d1 0.7424 -> 0.742, Jev 0.7269 -> 0.727.

0.741 is your first run. The card says 0.743 is the run with saved predictions, so I think the split was copied from the other one.

It matters more than it looks. Liquid d1 sits at 0.742, so 0.741 vs 0.743 puts you on either side of it, and your own two runs differ by more than that gap.

On "#1": the Hub widget holds 2 rows today, yours and a MiniLM specialist at 0.607. The maintainers' zero-shot table has Decider 1 at 0.768.

The part I find genuinely interesting is noul. 0.847 vs 0.840 for Decider 1 and Liquid. That is about 4 decisions of 600, a nose not a margin, but it is the one type where a single-pass readout leads.

What does the per-type split look like for the saved-predictions run? Which type moved?

Β·

Thanks, your arithmetic checks out: the per-type split on the card does not add up to 0.743. We are recomputing the per-type numbers from the run with saved predictions and will post them here, then correct the card.

On "#1": fair point. It is #1 among the entries submitted to the Hub leaderboard, not above Decider 1's 0.768 in the maintainers' zero-shot table. We will reword it.