Fix Results section to report the measured val loss/ppl from the shipped checkpoint's own eval (eval_fresh.json: 3.8719 / 48.03). The previous card cited 3.8775 / 48.30 / "step 20000", but the training run diverged to NaN at step 14300 and the log died at step 16000 — step 20000 was never reached. Also correct the eval filename reference (eval_shipped.json -> eval_fresh.json, the file actually in the repo).

#8
Files changed (1) hide show
  1. README.md +6 -3
README.md CHANGED
@@ -56,8 +56,11 @@ Standard LLaMA block layout: `RMSNorm -> Attention(q/k/v/o) -> residual`,
56
 
57
  ## Results (measured, not asserted)
58
 
59
- - **Validation loss:** 3.8775 (final checkpoint, step 20000)
60
- - **Validation perplexity:** 48.30 (over held-out fineweb-edu text)
 
 
 
61
  - **Degeneracy check:** 0 / 15 samples flagged by the repeated-3-gram loop
62
  detector (a single 3-gram covering >60% of the 40-word tail). Note this
63
  detector only catches exact token-loops; it does **not** catch the more
@@ -65,7 +68,7 @@ Standard LLaMA block layout: `RMSNorm -> Attention(q/k/v/o) -> residual`,
65
  a sentence), which the samples show clearly.
66
 
67
  Representative samples (temperature 0.8, top-k 40, **verbatim from the shipped
68
- `model.safetensors`**, from `eval_shipped.json`):
69
 
70
  > "The cat sat on the center of the church in the center of the church. The
71
  > catalog is the same as the Bishop of the church, which includes the church."
 
56
 
57
  ## Results (measured, not asserted)
58
 
59
+ - **Validation loss:** 3.8719 (measured on the shipped checkpoint, held-out fineweb-edu)
60
+ - **Validation perplexity:** 48.03 (over held-out fineweb-edu text)
61
+ - **Training note:** the run was budgeted for 20000 steps but diverged to NaN
62
+ loss at step 14300 and the log died at step 16000; the shipped
63
+ `model.safetensors` is the checkpoint that was evaluated (numbers above).
64
  - **Degeneracy check:** 0 / 15 samples flagged by the repeated-3-gram loop
65
  detector (a single 3-gram covering >60% of the 40-word tail). Note this
66
  detector only catches exact token-loops; it does **not** catch the more
 
68
  a sentence), which the samples show clearly.
69
 
70
  Representative samples (temperature 0.8, top-k 40, **verbatim from the shipped
71
+ `model.safetensors`**, from `eval_fresh.json`):
72
 
73
  > "The cat sat on the center of the church in the center of the church. The
74
  > catalog is the same as the Bishop of the church, which includes the church."