Correct held-out perplexity to full-split test 1.4369 / 4.21 (was 60-batch log numbers)

#7
by Compactbot - opened
Files changed (1) hide show
  1. README.md +5 -3
README.md CHANGED
@@ -59,12 +59,11 @@ nanoGPT-style GPT-2, all bias-free except LayerNorm:
59
 
60
  ## Quality — what it is and is not
61
 
62
- Held-out perplexities (from training log):
63
 
64
  | split | loss | perplexity |
65
  |---|---|---|
66
- | val | 1.9046 | **6.72** |
67
- | test | 1.9473 | **7.01** |
68
 
69
  It captures TinyStories' surface style (short declarative sentences, simple
70
  vocabulary, character names) but it is a **1.2M-parameter model on ~1M
@@ -72,6 +71,9 @@ characters** — it does not grasp meaning, it repeats and drifts, and it will
72
  produce the kind of plausible-looking-but-nonsense text in `sample.txt`.
73
  Treat it as a working toy / reference architecture, not a useful language model.
74
 
 
 
 
75
  ## Files
76
 
77
  - `model.safetensors` — 4,869,136 B (53 tensors, F32)
 
59
 
60
  ## Quality — what it is and is not
61
 
62
+ Held-out perplexities (measured on the full held-out test split, 2026-09-20):
63
 
64
  | split | loss | perplexity |
65
  |---|---|---|
66
+ | test (49,674 tokens) | 1.4369 | **4.21** |
 
67
 
68
  It captures TinyStories' surface style (short declarative sentences, simple
69
  vocabulary, character names) but it is a **1.2M-parameter model on ~1M
 
71
  produce the kind of plausible-looking-but-nonsense text in `sample.txt`.
72
  Treat it as a working toy / reference architecture, not a useful language model.
73
 
74
+ > The original training log reported val 1.9046 / test 1.9473 from a
75
+ > 60-batch random evaluation; the full-split number above is the honest one.
76
+
77
  ## Files
78
 
79
  - `model.safetensors` — 4,869,136 B (53 tensors, F32)