Independent measurements of the six external board rows (GPT-X2-125M, JugnuLM-110M/53M, Haidass-143M, Supra-50M, Hydrion-v1)
Following up on #1 β the six external board rows, measured the same way as the five GoLLeM rows there: the board's own protocol (glint_parity_eval.py / exact port of Glint-1.3/benchmark.py), RTX 5090, BLiMP 67,000 pairs, ARC-Easy 2,376 questions, WikiText-2 test in 256-token windows. Each loaded from its public checkpoint via from_pretrained. All six repos are unchanged since before this thread (GPT-X2-125M 2026-08-07, JugnuLM-110M-R2plus 2026-09-19, Haidass-143M-v1 2026-08-12, JugnuLM-53M 2026-09-05, Supra-50M-Base 2026-05-27, Hydrion-v1-Base 2026-07-24), so current main = the revision listed in #1.
| model (HF) | params (measured) | vocab | BLiMP (n=67,000) | ARC-Easy (n=2,376) | WikiText-2 token-PPL |
|---|---|---|---|---|---|
| AxiomicLabs/GPT-X2-125M | 125.1M | 32,768 | 79.08 | 55.51 | 30.6748 |
| altslate/JugnuLM-110M-R2plus | 109.7M | 49,152 | 76.43 | 53.91 | 31.5084 |
| DALabCommunity/Haidass-143M-v1 | 143.1M | 64,000 | 75.23 | 55.77 | 25.8358 |
| altslate/JugnuLM-53M | 53.5M | 49,152 | 73.65 | 50.88 | 50.2773 |
| SupraLabs/Supra-50M-Base | 51.8M | 32,000 | 75.51 | 50.63 | 46.5707 |
| OPENGCM/Hydrion-v1-Base | 114.1M | 50,280 | 75.37 | 45.62 | 59.4294 |
vs the "ours" column in #1: BLiMP and ARC-Easy agree closely on all six β BLiMP within ~1.2 pts (mine 1.13β1.30 lower on five, exact on JugnuLM-53M), ARC within ~0.5 pts (mine lower on five, exact on JugnuLM-53M). So this is a clean independent confirmation of the "ours" numbers, same as the five GoLLeM rows in #1.
Two honest caveats, so the comparison is read correctly:
- The WikiText-2 column is token-PPL (the protocol's raw output, each model's own tokenizer), not the board's byte-PPL. The six use different tokenizers (32kβ64k vocab), so their bytes/token differ and I have not converted to byte-PPL here β that column is therefore not directly comparable to the board's self-reported byte-PPL (1.86β2.04). To convert I'd compute each model's bytes/token on the WikiText-2 test and take token-PPL^(1/bytes-token); happy to add that column if you want the byte-PPL comparison directly.
- These are my measurements of the public checkpoints with your harness β they confirm the checkpoints reproduce your "ours" numbers under the board's protocol, not an independent check of the protocol itself.
On the Supra-50M-Base 2.0400: I raised it with SupraLabs on their model page (SupraLabs/Supra-50M-Base#2) as suggested in #1 β their card reports a token-level "Final loss 3.259" and an lm-eval table but no WikiText-2 byte-perplexity, so the board's 2.04 for that row is unsupported by its own card. Awaiting their reply; I'll report back here when they respond.
Reporting back as I said I would: SupraLabs responded on their model page (Supra-50M-Base#2), and it changes my earlier "unsupported by its own card" line.
Their benchmarks.md does carry the number β byte_perplexity 2.0374 (bits_per_byte 1.0267, word_perplexity 44.9548) on wikitext-2-raw-v1 test, via lm-eval's wikitext task. I'd only read the card, which doesn't repeat it, so my "no WikiText-2 byte-perplexity in the card" was literally true but pointed at the wrong file. The board's 2.0400 β their 2.0374, so the row is backed by their own measurement.
The remaining gap (their 2.0374 vs my 2.4262) is protocol, not model or text β I confirmed it by re-running Supra-50M-Base under lm-eval's exact loglikelihood_rolling and getting 2.0374 / 1.0267, matching their row exactly:
- their 2.0374 = lm-eval sliding-window (every token scored with full context)
- my 2.4262 = the board's own non-overlapping 256-token-chunk protocol (first token of each chunk scored with ~0 context, which inflates PPL)
So the Supra-50M-Base row in the external column should read as verified, not unsupported β I'll correct that in my #2 table if you want me to. The one thing still worth fixing on the board: the "WikiText-2 byte-PPL" column doesn't say which protocol, so 2.0374 and 2.4262 get read as a contradiction when they aren't. Happy to open a PR labeling the column (or adding a second column for the board-protocol number) if that's useful.