File size: 1,866 Bytes
7c85c7e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
# Pivot performance

All figures here refer to checkpoint `14bf8c26bf344ebdf88e22a4b6152dc5f75f3578`, public JevBench v1.4.1 commit `24b9b5c1609a7a9e8fa14f49e5985a836c9dc842`, FP32 and the same frozen 512/128-token input contract.

## Accuracy on public tasks

| Tier | Correct | Tasks | Accuracy | Top-label ECE, 10 bins |
|---|---:|---:|---:|---:|
| Original | 27 | 72 | 37.50% | 0.5247 |
| Easy | 39 | 48 | 81.25% | 0.1139 |
| Hard | 41 | 111 | 36.94% | 0.3792 |
| **Total** | **107** | **231** | **46.32%** | — |

![Public-tier accuracy](../evaluation/2026-09-24/accuracy.png)

The official JevBench composite score is **unavailable** because the sealed and judge tasks and official cost input were not measured.

## Local speed

| Warm local FP32 measure | H200 GPU | Xeon CPU, 4 threads |
|---|---:|---:|
| Single decision p50, 32 measured | 15.7668 ms | 797.5558 ms |
| Single decision p95, 32 measured | 19.8708 ms | 1089.0266 ms |
| Single decision mean | 16.1614 ms | 770.7159 ms |
| Batch size for throughput | 32 | 4 |
| Throughput, median of 3 × 64 decisions | 545.2833 decisions/s | 3.7708 decisions/s |

![Warm local latency and throughput](../evaluation/2026-09-24/latency_throughput.png)

The CPU p50 single-decision time is **50.6×** the H200 p50 for these two machines. CPU and GPU batch throughput used different batch sizes and should not be read as a same-batch comparison. The timings cover tokenizer + inference + scoring with 5 warmup singles and 2 warmup bulk passes. Reproduce them on your own hardware using the [CPU script](../cpu-speed/README.md) or [full public runner](../benchmarks/README.md).

Source: [full structured summary](../evaluation/2026-09-24/performance.json), [original public benchmark result](../evaluation/2026-09-24/jevbench_public.json), and [original CPU measurement](../evaluation/2026-09-24/cpu_speed.json).