Fix Results section to report the measured val loss/ppl from the shipped checkpoint's own eval (eval_fresh.json: 3.8719 / 48.03). The previous card cited 3.8775 / 48.30 / "step 20000", but the training run diverged to NaN at step 14300 and the log died at step 16000 — step 20000 was never reached. Also correct the eval filename reference (eval_shipped.json -> eval_fresh.json, the file actually in the repo).
#8
by Compactbot - opened
README.md
CHANGED
|
@@ -56,8 +56,11 @@ Standard LLaMA block layout: `RMSNorm -> Attention(q/k/v/o) -> residual`,
|
|
| 56 |
|
| 57 |
## Results (measured, not asserted)
|
| 58 |
|
| 59 |
-
- **Validation loss:** 3.
|
| 60 |
-
- **Validation perplexity:** 48.
|
|
|
|
|
|
|
|
|
|
| 61 |
- **Degeneracy check:** 0 / 15 samples flagged by the repeated-3-gram loop
|
| 62 |
detector (a single 3-gram covering >60% of the 40-word tail). Note this
|
| 63 |
detector only catches exact token-loops; it does **not** catch the more
|
|
@@ -65,7 +68,7 @@ Standard LLaMA block layout: `RMSNorm -> Attention(q/k/v/o) -> residual`,
|
|
| 65 |
a sentence), which the samples show clearly.
|
| 66 |
|
| 67 |
Representative samples (temperature 0.8, top-k 40, **verbatim from the shipped
|
| 68 |
-
`model.safetensors`**, from `
|
| 69 |
|
| 70 |
> "The cat sat on the center of the church in the center of the church. The
|
| 71 |
> catalog is the same as the Bishop of the church, which includes the church."
|
|
|
|
| 56 |
|
| 57 |
## Results (measured, not asserted)
|
| 58 |
|
| 59 |
+
- **Validation loss:** 3.8719 (measured on the shipped checkpoint, held-out fineweb-edu)
|
| 60 |
+
- **Validation perplexity:** 48.03 (over held-out fineweb-edu text)
|
| 61 |
+
- **Training note:** the run was budgeted for 20000 steps but diverged to NaN
|
| 62 |
+
loss at step 14300 and the log died at step 16000; the shipped
|
| 63 |
+
`model.safetensors` is the checkpoint that was evaluated (numbers above).
|
| 64 |
- **Degeneracy check:** 0 / 15 samples flagged by the repeated-3-gram loop
|
| 65 |
detector (a single 3-gram covering >60% of the 40-word tail). Note this
|
| 66 |
detector only catches exact token-loops; it does **not** catch the more
|
|
|
|
| 68 |
a sentence), which the samples show clearly.
|
| 69 |
|
| 70 |
Representative samples (temperature 0.8, top-k 40, **verbatim from the shipped
|
| 71 |
+
`model.safetensors`**, from `eval_fresh.json`):
|
| 72 |
|
| 73 |
> "The cat sat on the center of the church in the center of the church. The
|
| 74 |
> catalog is the same as the Bishop of the church, which includes the church."
|