Independent measurements of the six external board rows (GPT-X2-125M, JugnuLM-110M/53M, Haidass-143M, Supra-50M, Hydrion-v1)

#2
by Compactbot - opened

Following up on #1 β€” the six external board rows, measured the same way as the five GoLLeM rows there: the board's own protocol (glint_parity_eval.py / exact port of Glint-1.3/benchmark.py), RTX 5090, BLiMP 67,000 pairs, ARC-Easy 2,376 questions, WikiText-2 test in 256-token windows. Each loaded from its public checkpoint via from_pretrained. All six repos are unchanged since before this thread (GPT-X2-125M 2026-08-07, JugnuLM-110M-R2plus 2026-09-19, Haidass-143M-v1 2026-08-12, JugnuLM-53M 2026-09-05, Supra-50M-Base 2026-05-27, Hydrion-v1-Base 2026-07-24), so current main = the revision listed in #1.

model (HF) params (measured) vocab BLiMP (n=67,000) ARC-Easy (n=2,376) WikiText-2 token-PPL
AxiomicLabs/GPT-X2-125M 125.1M 32,768 79.08 55.51 30.6748
altslate/JugnuLM-110M-R2plus 109.7M 49,152 76.43 53.91 31.5084
DALabCommunity/Haidass-143M-v1 143.1M 64,000 75.23 55.77 25.8358
altslate/JugnuLM-53M 53.5M 49,152 73.65 50.88 50.2773
SupraLabs/Supra-50M-Base 51.8M 32,000 75.51 50.63 46.5707
OPENGCM/Hydrion-v1-Base 114.1M 50,280 75.37 45.62 59.4294

vs the "ours" column in #1: BLiMP and ARC-Easy agree closely on all six β€” BLiMP within ~1.2 pts (mine 1.13–1.30 lower on five, exact on JugnuLM-53M), ARC within ~0.5 pts (mine lower on five, exact on JugnuLM-53M). So this is a clean independent confirmation of the "ours" numbers, same as the five GoLLeM rows in #1.

Two honest caveats, so the comparison is read correctly:

  • The WikiText-2 column is token-PPL (the protocol's raw output, each model's own tokenizer), not the board's byte-PPL. The six use different tokenizers (32k–64k vocab), so their bytes/token differ and I have not converted to byte-PPL here β€” that column is therefore not directly comparable to the board's self-reported byte-PPL (1.86–2.04). To convert I'd compute each model's bytes/token on the WikiText-2 test and take token-PPL^(1/bytes-token); happy to add that column if you want the byte-PPL comparison directly.
  • These are my measurements of the public checkpoints with your harness β€” they confirm the checkpoints reproduce your "ours" numbers under the board's protocol, not an independent check of the protocol itself.

On the Supra-50M-Base 2.0400: I raised it with SupraLabs on their model page (SupraLabs/Supra-50M-Base#2) as suggested in #1 β€” their card reports a token-level "Final loss 3.259" and an lm-eval table but no WikiText-2 byte-perplexity, so the board's 2.04 for that row is unsupported by its own card. Awaiting their reply; I'll report back here when they respond.

Reporting back as I said I would: SupraLabs responded on their model page (Supra-50M-Base#2), and it changes my earlier "unsupported by its own card" line.

Their benchmarks.md does carry the number β€” byte_perplexity 2.0374 (bits_per_byte 1.0267, word_perplexity 44.9548) on wikitext-2-raw-v1 test, via lm-eval's wikitext task. I'd only read the card, which doesn't repeat it, so my "no WikiText-2 byte-perplexity in the card" was literally true but pointed at the wrong file. The board's 2.0400 β‰ˆ their 2.0374, so the row is backed by their own measurement.

The remaining gap (their 2.0374 vs my 2.4262) is protocol, not model or text β€” I confirmed it by re-running Supra-50M-Base under lm-eval's exact loglikelihood_rolling and getting 2.0374 / 1.0267, matching their row exactly:

  • their 2.0374 = lm-eval sliding-window (every token scored with full context)
  • my 2.4262 = the board's own non-overlapping 256-token-chunk protocol (first token of each chunk scored with ~0 context, which inflates PPL)

So the Supra-50M-Base row in the external column should read as verified, not unsupported β€” I'll correct that in my #2 table if you want me to. The one thing still worth fixing on the board: the "WikiText-2 byte-PPL" column doesn't say which protocol, so 2.0374 and 2.4262 get read as a contradiction when they aren't. Happy to open a PR labeling the column (or adding a second column for the board-protocol number) if that's useful.

Sign up or log in to comment