Update card to v8 (current weights): 79,375 steps / 2.60B tokens (hits target), val 3.63. Honest eval: 40-sample sweep shows mean loop frac 0.219 (up from v7's 0.167) and 14/40 degenerate — more tokens did not buy coherence. Replaces the stale v7-only card that described the weights as 'well short of 2.6B'.

#15
by Compactbot - opened
Files changed (1) hide show
  1. README.md +31 -28
README.md CHANGED
@@ -44,47 +44,51 @@ Standard LLaMA block: RMSNorm → MHA (RoPE) → residual → RMSNorm → SwiGLU
44
 
45
  ## Training
46
 
47
- | | v1 (first ckpt) | v4 | **v7 (current)** |
48
- |---|---|---|---|
49
- | Steps | 4,000 | 16,000 | **30,000** |
50
- | Batch size | 64 | 64 | 64 |
51
- | LR | 3e-4 → 3e-5 (cosine) | 1e-4 → 1e-5 (cosine, fresh optimizer) | 1e-4 → 1e-5 (cosine, fresh optimizer) |
52
- | Data | FineWeb-Edu + DCLM | + FineWeb-Edu + DCLM | + FineWeb-Edu + DCLM (continued) |
53
- | Hardware | RTX 5090 (32 GB) | RTX 5090 (32 GB) | RTX 5090 (32 GB) |
54
- | Final val loss | 4.6020 (ppl 99.68) | 3.8943 (ppl 49.12) | **3.6892 (ppl 40.01)** |
 
 
55
 
56
- v7 is a continuation of v4 (resumed from step 12,000 to step 30,000). The data file was cleaned up after training, so the exact cumulative token count is **not re-derived here**; the step count and val loss are the measured values.
57
 
58
- ### ⚠️ Honest caveat: undertrained and still incoherent
59
 
60
- DedeProGames requested **2.6B tokens**. This checkpoint is well short of that. At this token budget the model has learned the **surface shape** of English — real words, parseable sentences, occasional token loops — but the prose is still **semantically incoherent** (word salad): grammatically plausible sentences that don't mean what they say. The val loss (3.689) is well below the 7.38 unigram floor, so it genuinely uses context; the improvement 4.60 → 3.89 → 3.69 is real.
61
 
62
- This is a **continued checkpoint**, not the final deliverable. More data is the fix; continued training toward the 2.6B budget is planned.
63
 
64
  ## Eval (40 samples: 8 prompts × 5 seeds, temp 0.8, top-k 40, 160 new tokens)
65
 
66
- | Metric | v1 | v4 | **v7** |
67
- |---|---|---|---|
68
- | val loss | 4.6020 | 3.8943 | **3.6892** |
69
- | perplexity | 99.68 | 49.12 | **40.01** |
70
- | Below unigram floor (7.38)? | Yes | Yes | Yes |
71
- | mean 4-gram loop fraction | n/a | ~0.11 | **0.167** |
72
- | Degenerate (mean loop > 0.30)? | No | No | **No** |
73
- | Semantically coherent? | No | No | No (improved, still word salad) |
74
 
75
- ### Sample outputs (v7, real generation from the weights)
76
 
77
- > "Once upon a time a human brain will experience a normal amount of memory. The body needs to be removed from the brain due to the fact that the brain is not working properly, and therefore the brain needs to get a lot of memory back to the very first and then there is a chance to think in the future."
78
 
79
- > "Once upon a time, you can easily make a decision about not being involved. - If you are not involved in a job, go to the office, and let your decision be made, then ask you to be a job as soon as possible."
80
 
81
- These are real outputs from the v7 weights (not hand-picked for coherence). The text is grammatically structured — real words, parseable sentences, occasional token loops, no broken tokens — but semantically incoherent. That is the honest state of a 10M model at this token budget.
 
 
82
 
83
  ## Usage
84
 
85
  The model uses a custom `LDT` architecture (standard LLaMA block, no special tricks). The safetensors file contains 47 tensors with tied embeddings (lm_head is not stored separately; it shares `tok.weight`).
86
 
87
- To load with a custom model class, you need a small LLaMA-style implementation matching the config above. The training script (`train_ldt10m_fixed.py`) contains the full architecture definition.
88
 
89
  ```python
90
  from model import LDT
@@ -100,7 +104,6 @@ print(tok.decode(out[0].tolist(), skip_special_tokens=True))
100
 
101
  ## What this is NOT
102
 
103
- - Not a 2.6B-token model (that's the target; this is a continued checkpoint well short of it)
104
- - Not a coherent-text model (it produces grammatically-structured word salad at this token budget)
105
  - Not a general-purpose assistant (it's a raw LM, no instruction tuning)
106
- - Not a replacement for anything larger — it's a research checkpoint in a from-scratch training run
 
44
 
45
  ## Training
46
 
47
+ | | v1 (first ckpt) | v4 | v7 | **v8 (current)** |
48
+ |---|---|---|---|---|
49
+ | Steps | 4,000 | 16,000 | 30,000 | **79,375** |
50
+ | Batch size | 64 | 64 | 64 | 64 |
51
+ | Seq length | 512 | 512 | 512 | 512 |
52
+ | Cumulative tokens | ~131M | ~308M | ~983M | **2.60B** |
53
+ | LR | 3e-4 → 3e-5 (cosine) | 1e-4 → 1e-5 (cosine, fresh optimizer) | 1e-4 → 1e-5 (cosine, fresh optimizer) | 1e-4 → 1e-5 (cosine, fresh optimizer) |
54
+ | Data | FineWeb-Edu + DCLM | + FineWeb-Edu + DCLM | + FineWeb-Edu + DCLM (continued) | + FineWeb-Edu + DCLM (continued) |
55
+ | Hardware | RTX 5090 (32 GB) | RTX 5090 (32 GB) | RTX 5090 (32 GB) | RTX 5090 (32 GB) |
56
+ | Final val loss | 4.6020 (ppl 99.68) | 3.8943 (ppl 49.12) | 3.6892 (ppl 40.01) | **3.63 (ppl 37.8)** |
57
 
58
+ v8 is the final checkpoint of the from-scratch run: 79,375 steps × 64 × 512 = **2,600,960,000 tokens (2.60B)**, which **hits DedeProGames' 2.6B-token target**. The val loss is read from the final checkpoint's recorded `val_loss` (3.63); the earlier columns' token counts are step-derived (steps × 64 × 512).
59
 
60
+ ### ⚠️ Honest caveat: token target hit, but still incoherent and loop-prone
61
 
62
+ v8 reaches the requested 2.6B-token budget. The val loss (3.63) is well below the 7.38 unigram floor, so the model genuinely uses context; the improvement 4.60 → 3.89 → 3.69 → 3.63 is real but **marginal in the last stage** (3.69 → 3.63).
63
 
64
+ However, the 40-sample generation sweep below shows that **more tokens did not buy coherence — it bought more token loops.** The mean 4-gram loop fraction went **up** from v7 (0.167) to v8 (0.219), and the degenerate-sample rate (loop > 0.30) went from 0/40 to **14/40**. At this scale the model has learned the surface shape of English — real words, parseable sentence frames — but the prose is still **semantically incoherent** (word salad) and increasingly prone to hard repetition loops. This is the honest state of a 10M model at 2.6B tokens: the token budget is spent, and the quality ceiling of a 10M-param LM is what it is.
65
 
66
  ## Eval (40 samples: 8 prompts × 5 seeds, temp 0.8, top-k 40, 160 new tokens)
67
 
68
+ | Metric | v1 | v4 | v7 | **v8** |
69
+ |---|---|---|---|---|
70
+ | val loss | 4.6020 | 3.8943 | 3.6892 | **3.63** |
71
+ | perplexity | 99.68 | 49.12 | 40.01 | **37.8** |
72
+ | Below unigram floor (7.38)? | Yes | Yes | Yes | Yes |
73
+ | mean 4-gram loop fraction | n/a | ~0.11 | 0.167 | **0.219** |
74
+ | Degenerate samples (loop > 0.30) | n/a | 0/40 | 0/40 | **14/40** |
75
+ | Semantically coherent? | No | No | No | No (word salad, more loops) |
76
 
77
+ ### Sample outputs (v8, real generation from the published weights — not hand-picked)
78
 
79
+ > "Once upon a time, the time of the Lord, the Lord, the Lord, the Lord, the Lord, the Lord, the Lord, the Lord, the Lord, the Lord, the Lord, the Lord, the Lord, the Lord, the Lord, the Lord, the Lord, the Lord, the Lord" (greedy)
80
 
81
+ > "She walked down the street. When she said she said she said she said she said she said she said she said she said she said she wa…" (temp 0.8, seed 0, loop frac 0.595)
82
 
83
+ > "A long time ago in a galaxy far far away from an galaxy that is almost no point from the distant distant galaxy called a…" (temp 0.8, seed 2, loop frac 0.587)
84
+
85
+ These are real outputs from the v8 weights. The text is grammatically structured — real words, parseable sentence frames — but semantically incoherent and loop-prone. That is the honest state of a 10M model at this scale.
86
 
87
  ## Usage
88
 
89
  The model uses a custom `LDT` architecture (standard LLaMA block, no special tricks). The safetensors file contains 47 tensors with tied embeddings (lm_head is not stored separately; it shares `tok.weight`).
90
 
91
+ To load with a custom model class, you need a small LLaMA-style implementation matching the config above. The training script (`train_ldt10m_v8.py`) contains the full architecture definition.
92
 
93
  ```python
94
  from model import LDT
 
104
 
105
  ## What this is NOT
106
 
107
+ - Not a coherent-text model (it produces grammatically-structured word salad with frequent token loops at this scale)
 
108
  - Not a general-purpose assistant (it's a raw LM, no instruction tuning)
109
+ - Not a replacement for anything larger — it is the final from-scratch checkpoint of a 10M-param run that hit its 2.6B-token budget