Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
dylantom2012 
posted an update 13 days ago
Post
68
I benchmarked TypeSafe's Jev against open, CPU-only stacks. 10,000 decisions, every raw output published.

Jev returns a typed choice with calibrated probabilities, and the options arrive per request — so it can't be a fixed classifier head. It has to score (state, option) pairs. I rebuilt that from open parts and measured instead of guessing.

Zero-shot, neither side given labels:
• typesafe/jev (hosted): 79.3%, 381 ms
• ModernBERT-base-zeroshot cross-encoder, 149M, CPU: 78.7%
• bi-encoder + 8-float head: 69.9%, 15.7 ms
• Qwen2.5-0.5B generative: 49.6%, with 4.0% unparseable outputs

A 149M open model on a laptop CPU lands 0.6 points behind a hosted commercial model.

With ~2,000 labels per task, an 8M static embedding plus a decision layer of literally eight floats reaches 78.2% at 0.1 ms in 82 MB, and trains in 0.1 s. On BANKING77 (77-way routing) it gets 88.8% vs Jev's 77.8%.

I also published the eight things that did NOT work: template ensembling 0.0pp, scaling 149M to 395M net zero, distillation -5.4pp, listwise reranking -1.0pp. The only thing that worked was architectural — scoring (text, option) jointly instead of comparing two embeddings. +12 to +24pp.

And one I didn't expect: there is no neutral set of option descriptions. The same rewrite moves Jev +0.5pp, the bi-encoder +6.3pp, the cross-encoder -4.3pp.

Model (the eight floats): dylantom2012/fly-head-potion-8m
Per-item predictions from 5 stacks incl. Jev: dylantom2012/open-system-one-bench
Interactive results: dylantom2012/open-system-one
Code: https://github.com/zhlei07/open-system-one

I don't program. The hunch was mine; every experiment was designed and run by Claude Code in one session. Everything is committed so the numbers can be checked rather than trusted.

The 0.6 points is not the finding in your file. The probabilities are.

You call Jev's probabilities calibrated, and your predictions.jsonl, by your card the only public per-item Jev output, lets anyone check. So I did. 9,976 successful calls, bare labels, top probability against accuracy:

top prob      n      mean conf   accuracy
[0.5,0.6)     466    0.545       0.388
[0.7,0.8)     607    0.746       0.537
[0.9,1.0)    7306    0.987       0.890
printed 1.0  4687    1.000       0.940     281 wrong
all          9976    0.909       0.787     ECE 0.121

Emotion is the worst row: 0.857 confidence against 58.4% accuracy. Its gold labels are noisy, so drop it. The other three still give 0.926 against 85.5%, ECE 0.071, and printed 1.0 is wrong 145 times in 3,952.

That is the number an agent builder needs from a decision layer. Calibration is what lets you pick one threshold and escalate everything under it. Gate Jev at 0.99 and it auto-decides 55.2% of items, and 7.05% of those are wrong.

On the question your card asks: paired on the 9,976 rows both stacks answered, enriched Jev against the bare cross-encoder is 863 to 791 discordant, exact McNemar p = 0.08. Bare against bare it is 824 to 820.

Your new laya_probs column lets the same calibration check run on an open stack. It comes out worse. Same bins, all 10,000 items: 0.913 confidence against 73.2% accuracy, ECE 0.180.

On BANKING77, the task it loses by 24 points, it reports 0.958 against 54.3%. Gate it at 0.99 and it auto-decides 37.1% of items, 30.3% of them wrong.

So the 77-way loss is not the part that would hurt a router. The confidence on it is.

The cross-encoder, your 0.6-point runner-up, still ships labels only. Does its softmax do any better on these same items?