File size: 26,939 Bytes
f345921
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
# 01 β€” Plan: architecture, hyperparameters, throughput, targets

Phase 1 deliverable. Every choice below cites a source actually inspected on 2026-09-19, as Gate 1
requires. Items marked **β†’ test** are decisions whose *numbers* Phase 3 must validate on the real
pipeline before the run freezes; the design itself is settled.

---

## 1. The envelope the design has to fit

Measured in Phase 0 (`docs/00-platform-notes.md`), not assumed:

| Constraint | Value | Consequence |
|---|---|---|
| Accelerator | 2x Tesla T4, cc 7.5, **14.56 GiB usable each** | fp16 only; no bf16 tensor cores; no flash-attention (sm80+ only) |
| GPU quota | 108,000 s/week, billed at **~1x container wall-clock** (3 samples) | β‰ˆ30 wall-clock h/week, both cards |
| Test budget | ≀6 GPU-h lifetime | 0.073 h spent; **5.93 h remains for all of Phase 3** |
| Working disk in a job | **19.5 GB** `/kaggle/working` | shards consumed a few at a time, never whole (Β§3.13) |
| RAM / CPU | 30 GiB cgroup, 4 vCPU, **no swap** | input pipeline memory-bounded by design |
| Image | torch 2.10.0+cu128, **transformers 5.0.0**, datasets 5.0.0, accelerate 1.13.0, triton 3.6.0; **no** trl/deepspeed/flash-attn/xformers/bnb | prefer the preinstalled stack; each install is risk + round trip |
| Hub from a job | reads anonymous & fine, 35–89 MB/s; writes **401** until given a token | credential path: private Kaggle dataset mount (Β§8) |

**Parameter convention.** Β§2's 90–110M is *including embeddings* β€” the Pythia convention (Pythia was
explicitly renamed to include embedding + unembedding, [pythia README](https://raw.githubusercontent.com/pythia/main/README.md)).
This matters more at 100M than anywhere else: the tied embedding is 21–37 % of the model across the
candidates below. `code/config/param_count.py` computes it two ways and both were checked to agree
exactly against transformers 5.0.0 in `dodosoomro/ounce100m-p1-param-count`.

## 2. Architecture

### 2.1 The shape: deep-thin, tied, GQA, SwiGLU, pre-norm

**Choice: `hidden 576, layers 22, 9 Q heads / 3 KV heads (GQA 3:1), SwiGLU intermediate 1536, RMSNorm
pre-norm eps 1e-5, attention_bias false, RoPE theta 10,000, tied embeddings, vocab 49,152, seq 2048`.**

Evidence, and it is convergent rather than a single citation:

- **Deep-thin beats wide at this size.** Meta's MobileLLM ablated exactly this at a fixed 135M: 30
  layers Γ— d512 scored **44.8 avg vs 43.9 for 12 Γ— d768**, with embedding-sharing removing 13 % of
  parameters at ~equal accuracy ([arXiv 2402.14905](https://arxiv.org/abs/2402.14905), HIGH). Its
  shipped 125M model is **30 layers, d576, 9Q/3KV GQA, SwiGLU, tied**, 124.6M total.
- **The same shape is what HF's SmolLM2-135M uses**: 30 layers, d576, 9Q/**3KV**, `intermediate_size`
  1536 (2.67Γ—d), silu, `attention_bias=false`, **`tie_word_embeddings=true`**, RMSNorm 1e-5,
  [model config](https://huggingface.co/HuggingFaceTB/SmolLM2-135M/blob/main/config.json) +
  [arXiv 2502.02737 Β§6](https://arxiv.org/abs/2502.02737). Two independent labs landing on
  d576 + GQA 3:1 + SwiGLU at ~130M is the strongest available signal.
- Our budget is 110M, not 135M, so the same d576/GQA3:1/SwiGLU-1536 stack is cut to **22 layers**
  rather than 30 β€” verified count below β€” and RoPE ΞΈ drops to 10,000 to match the shorter context
  (SmolLM **v1** used ΞΈ=10,000 at ctx 2048; SmolLM2's 100,000 accompanies ctx 8192, which we cannot
  afford and do not need for these benchmarks).
- **GQA at 3:1 is a parameter choice as much as an attention choice**, since k/v are `d_head Γ— n_kv`
  wide; MobileLLM reports GQA-plus-enlarged-d as +0.4 avg.

**Not chosen: Mamba/SSM or hybrid.** The official kernels *do* build for `sm_75`
([state-spaces/mamba setup.py](https://github.com/state-spaces/mamba)), which corrects the common
assumption that they are Ampere-only β€” but they are absent from the image (source install only), and
the research found **no published 50–200M SSM/hybrid with released hyperparameters** to imitate. That is
research risk, not a default, and Β§4 forbids the kind of custom code where autonomous projects die.

### 2.2 Verified parameter counts

Closed form and `sum(p.numel())` agreed exactly on all eight shapes, in the target environment.
Tying confirmed genuinely tied, not merely declared.

| candidate | H | L | heads | kv | FFN | vocab | **total params** | emb share | in 90–110M |
|---|---|---|---|---|---|---|---|---|---|
| Phase-0 probe shape | 768 | 12 | 12 | 3 | 2048 | 50257 | 112,934,400 | 34.2 % | βœ— |
| A | 768 | 11 | 12 | 3 | 2048 | 50257 | 106,739,712 | 36.2 % | βœ“ |
| B | 768 | 12 | 12 | 3 | 1792 | 50257 | 105,856,512 | 36.5 % | βœ“ |
| C | 768 | 12 | 12 | 2 | 1824 | 50257 | 105,561,600 | 36.6 % | βœ“ |
| D | 640 | 16 | 10 | 4 | 2048 | 50257 | 113,450,240 | 28.4 % | βœ— |
| E | 768 | 12 | 12 | 3 | 2048 | 32768 | 99,502,848 | 25.3 % | βœ“ |
| F | 896 | 9 | 14 | 2 | 2432 | 50257 | 120,397,312 | 37.4 % | βœ— |
| G | 1024 | 8 | 8 | 2 | 2816 | 50257 | 141,658,112 | 36.3 % | βœ— |

**The binding insight: the tokenizer decides how much model you are allowed to have.** At 50,257 vocab
and d768 the tied embedding is 38.6M β€” a third of the entire budget β€” which is why every 50k candidate
clusters at 105–113M and must trim depth or FFN. MobileLLM's own ablation reached the same conclusion
(embedding sharing: βˆ’13 % parameters at ~equal accuracy). Choosing a smaller `hidden` is therefore not
a weakening, it is how the embedding tax gets paid for.

**β†’ test:** the frozen shape's exact count was re-derived from the constructed model β€” see Β§2.3.

### 2.3 The frozen shape, counted to the unit

`dodosoomro/ounce100m-p1-param-count` v2 built each candidate as a real `LlamaForCausalLM` under
transformers 5.0.0. Closed form and `sum(p.numel())` agree exactly on every row.

| candidate | H | L | Q/KV | FFN | vocab | **total params** | emb share | in 90–110M |
|---|---|---|---|---|---|---|---|---|
| H | 576 | 20 | 9/3 | 1536 | 49,152 | 99,114,048 | 28.6 % | βœ“ |
| **I β€” CHOSEN** | **576** | **22** | **9/3** | **1536** | **49,152** | **106,194,240** | **26.7 %** | βœ“ |
| J | 576 | 24 | 9/3 | 1536 | 49,152 | 113,274,432 | 25.0 % | βœ— over |
| K | 576 | 26 | 9/3 | 1536 | 49,152 | 120,354,624 | 23.5 % | βœ— over |
| L | 640 | 20 | 10/4 | 1728 | 49,152 | 120,776,320 | 26.0 % | βœ— over |

**Frozen: 106,194,240 parameters (embeddings included, tied), counted as
`vocabΓ—h` once + `L Γ— (attn + SwiGLU)` + norms + final norm, and confirmed by constructing the model.**
The counting method is stated here because Β§2 requires the number *and* how it was counted.

Two consequences worth recording:

- **Depth is capped at 22 by the budget, not by taste.** Each layer at d576 costs 3,538,944 parameters,
  so 24 layers is 113.3M β€” over. That is a hard stop, and it means the Phase-3 "deep-thin vs wide"
  throughput test (T1) must compare **H (20 layers, 99.1M) against I (22 layers, 106.2M)**, both
  in-budget, rather than the 12-layer d768 probe shape which cannot fit the SmolLM2 vocab without
  trimming elsewhere.
- **Widening is not a free alternative to deepening.** Candidate L at d640 Γ— 20 layers is *over* budget
  (120.8M) despite having fewer layers than I, because attention and MLP cost scale with `hΒ²`/`hΒ·ff`.
  So the deep-thin choice is forced from two directions: it is better per parameter (MobileLLM's
  ablation) and it is the only way to spend the 78M non-embedding budget on 22 layers.


## 3. Tokenizer

**Choice: reuse the SmolLM2 BPE tokenizer** β€” vocab 49,152, byte-level, **Apache-2.0**, from
[`HuggingFaceTB/SmolLM2-135M`](https://huggingface.co/HuggingFaceTB/SmolLM2-135M)'s `tokenizer.json`;
trained on SmolCorpus per [arXiv 2502.02737](https://arxiv.org/abs/2502.02737). Not trained by us.

Β§4 says maximum reuse, minimal custom code, and training a tokenizer from scratch is precisely the
kind of optional machinery that adds failure surface without adding capability. Candidates checked today:
SmolLM2 **49,152 / Apache-2.0**; `gpt2` 50,257 / MIT; `pythia-160m` 50,304 / Apache-2.0;
`t5-v1_1-small` 32,128 Unigram (not byte-level, carries `extra_ids` sentinels); `Qwen3-0.6B` 151,936
(disqualifying β€” its embedding alone would be ~78M at d512); `TinyLlama` 32,000 but Llama-2-derived and
**redistribution status unverified** β†’ excluded.

SmolLM2 wins on three counts: licence is clean and permissive, it is byte-level and English-curated at
exactly this model scale, and sharing it with a family of released 135M models gives the benchmark
numbers somebody else's token statistics to be compared against. Its 49,152 Γ— 576 tied embedding is
28.3M = **21 %** of budget, versus 34–37 % for the 50k/d768 shapes.

## 4. Optimizer, batch, schedule, precision

| | choice | justification |
|---|---|---|
| Optimizer | AdamW, Ξ²=(0.9, 0.95), wd 0.1, clip 1.0 | Pythia-160M and SmolLM2 both use Ξ²β‚‚=0.95 at small scale ([2304.01373](https://arxiv.org/abs/2304.01373), [2502.02737](https://arxiv.org/abs/2502.02737)) |
| Peak LR | **6e-4** | bracketed from three directions: Cerebras-111M 6.0e-4, Pythia-160M 6e-4, and the 150M low-tokens/param run of Sardana et al. at 4.6e-4 ([2304.03208](https://arxiv.org/abs/2304.03208), [2401.00448](https://arxiv.org/abs/2401.00448)). The 2e-3/3e-3 figures that appear in small-model recipes belong to 1–2 T-token runs, which this is not. |
| Warmup | 2 % of steps, floor 100 steps | Pythia used 1 %; SmolLM2 used 2,000 steps at 2M tok/batch (~0.2 %). Warmer than 1 % because **fp16 without bf16 needs the ramp** β€” Phase 0 showed an unramped/unscaled fp16 run going straight to NaN (E-006). |
| Schedule | **trapezoid / WSD: constant, then linear decay to 0 over the final 20 %** β€” *not* cosine-to-10 % | SmolLM1 used a trapezoid with 20 % cooldown; SmolLM2 used WSD ([blog/smollm](https://huggingface.co/blog/smollm), [2502.02737 App. A](https://arxiv.org/abs/2502.02737)). *Straight to Zero* finds linear-decay-to-zero beats cosine-to-10 % when peak LR is optimal, across sizes/batch/data ([2502.15938](https://arxiv.org/abs/2502.15938)); COLT 2026 theory says decay *shape* barely matters but **overly slow terminal decay causes schedule-induced capacity saturation** ([2602.06797](https://arxiv.org/abs/2602.06797)). Its caveat applies: their evidence is at/above compute-optimal tokens/param, we are at half. |
| Global batch | **β‰ˆ262 k tokens/step** (~3,815 steps for 1 B tokens) | Critical batch size fits `B* = 621.341Β·N^0.087`, i.e. **nearly independent of model size** ([2410.21676](https://arxiv.org/abs/2410.21676)) β€” no reason to chase 2–4M-token batches. Sardana et al. deliberately used *smaller* batches for smaller models "so that low-token-count training runs see enough steps". At 10 tokens/param we are such a run. |
| Micro-batch | **2 Γ— 2048 per GPU Γ— 2 GPUs Γ— 32 accumulation** | Set by memory, not taste: Phase 0 measured bs4Γ—1024 at 11.5–13.8 GB of 14.56 GB and **bs8 OOMed outright**. β†’ test |
| Precision | **fp16 autocast + fp32 master weights + `GradScaler` + grad clip** | Only mixed-precision option on Turing; transformers v5 documents this path by name ("fall back to fp16 on older hardware like V100 or T4") and warns **load the model in fp32 or autocast is a no-op** ([v5 mixed_precision docs](https://huggingface.co/docs/transformers/perf_training)). Verified finite in Phase 0 probe C; bf16 *runs* but is emulated β†’ must be asserted off, not defaulted. |
| Seed / RNG | single fixed seed, recorded | Β§3.1 requires "same RNG semantics" across resumes |

**Undertrained by design, and honestly so.** 1 B tokens on ~100M β‰ˆ **10 tokens/param**, about half
Chinchilla-optimal (Cerebras-111M's compute-optimal point was 2.2 B tokens = 20 t/p). That is mildly
undertrained, not pathological β€” the genuine anomaly is 2025-26 practice, which overtrains tiny models
by three orders of magnitude (SmolLM2-135M at 2 T tokens β‰ˆ 14,800 t/p). Sardana et al. trained a
150M/d768/12L model across 3β†’10,000 t/p and quality kept improving; Gadre et al. show scaling laws
extrapolate across over-training ([2403.08540](https://arxiv.org/abs/2403.08540)). We are on the
left-hand side of that curve because the quota puts us there, and the report should say so.

## 5. Training stack

**Choice: `torchrun --nproc_per_node=2` + Hugging Face `Trainer` / `TrainingArguments`, `fp16=True`,
on a pre-tokenised **map-style** `Dataset`.** Minimal glue, no framework authored here (Β§4).

- Everything needed is already in the image: transformers 5.0.0, accelerate 1.13.0, datasets 5.0.0.
- v5 `Trainer` restores **RNG (`_save_rng_state`/`_load_rng_load`), optimizer, LR scheduler and the fp16
  GradScaler** β€” the exact set Β§3.1 demands β€” via `Accelerator.save_state`. Verified against
  v5.17.0 source; the image's 5.0.0 is ~8 patch releases behind but the same major. **β†’ test at Gate 3.**
- **The resume trap, stated plainly:** v5 recovers data position with `skip_first_batches`, which
  **replays** discarded batches β€” O(steps) wasted work on an `IterableDataset`, up to 3,815 steps of
  pure I/O after every interruption. Mitigation is architectural, not a flag: a **map-style, concatenated
  token store** makes skipping an index advance rather than a decode. A `dataset_shard` layout with an
  explicit recorded cursor is what Gate 3 must prove, and Β§Phase 2's "exact positional resumption"
  requirement is the reason it is designed that way up front.
- **v5 migration traps to check in preflight** ([MIGRATION_GUIDE_V5.md](https://github.com/huggingface/transformers/blob/main/MIGRATION_GUIDE_V5.md), HF *Transformers v5* blog): PyTorch-only backend;
  slow tokenizers gone (irrelevant, we use `tokenizers`); several `TrainingArguments` removed **without
  a deprecation cycle**; attention moved to `AttentionInterface`; `remove_unused_columns` **still
  defaults True** (silently drops dataset columns a loss function needs); `lr_scheduler_type` default is
  still `"linear"`, *not* cosine β€” so an unstated schedule is a linear decay to zero over the whole run;
  new `train_sampling_strategy` replaces the `group_by_length` bool;
  `restore_callback_states_from_checkpoint` is required to restore scheduler state properly.
- **Disqualified:** **torchtitan** β€” `training_dtype` is typed `Literal["bfloat16","float32"]`, i.e.
  **no fp16 path at all**, plus Hopper-oriented CI β†’ unusable on T4 ([torchtitan config](https://github.com/pytorch/torchtitan)).
  **unsloth** β€” fine-tuning/GRPO only, no from-scratch pretraining path. **nanoGPT** β€” dormant since
  2025-11, no HF export, no distributed input pipeline. **FSDP** β€” ~100M params in fp32 master +
  grads + AdamW is ~1.8 GB; DDP on 2Γ—16 GB does not need sharding, and sharding would add failure
  surface for nothing. **DeepSpeed / flash-attn / xformers / trl** β€” absent from the image; installing
  them is risk without a capability we lack.

## 6. Throughput, and the token target it implies

Measured on a main-run-shaped model, `seq 1024`, bs4/card, DDP over both T4s: **11,062 tok/s
aggregate** (5,531/rank; single-GPU 6,478 SDPA / 7,239 eager; **85 % DDP efficiency**).

Sanity-checking that against compute: `6ND β‰ˆ 6 Γ— 1.13e8 Γ— 11,062 β‰ˆ 7.5e12 FLOP/s`, which is **~6 % of a
2Γ—T4 fp16 tensor peak** (65 TFLOPS/card β€” an *unverified* datasheet prior; the spec page could not be
retrieved). Independent corroboration: a documented 163M / ctx-1024 run reached ~19.9k tok/s on a single
RTX 3090 ([gilesthomas.com](https://gilesthomas.com/2025/12/llm-from-scratch-28-training-a-base-model-from-scratch)),
and Pythia-160M's 1,030 A100-hours for 300 B tokens implies ~81k tok/s per A100. So ~11k tok/s on two
Turing cards is the right order of magnitude, and **the low implied MFU means this is a floor with headroom,
not a ceiling** β€” which is the opposite of the risk I most wanted to rule out.

**The tension to resolve by measurement, not argument:** deep-thin is better per *parameter*
(MobileLLM's ablation) but worse per *second* on 2018 hardware, because 22 small layers issue more
sequential kernel launches than 12 fat ones, and we are wall-clock-bound rather than
parameter-bound. Phase 3 therefore measures both shapes.

| token target | at 11.1k tok/s (measured) | at 9k tok/s (planning floor) |
|---|---|---|
| 1.10 B | 27.6 h | 34.0 h |
| **1.00 B** | **25.1 h** | **30.9 h** |
| 0.90 B | 22.6 h | 27.8 h |

**Decision: target 1.0 B tokens, with a pre-registered fallback to 0.9 B if Gate 3's end-to-end rate on
the frozen config and real input pipeline is below ~10.5k tok/s.** Both are inside Β§2's band, and the
fallback is chosen *now* rather than later so it cannot become a post-hoc excuse. Even the optimistic
column leaves under 5 h of slack in a 30 h week, which the interruption budget in Β§3.1 will eat; so the
plan assumes **a two-week run crossing the 2026-09-26 quota reset**, with a checkpoint at every 10 % and
a rolling `latest`, exactly as Β§2 requires.

## 6.1 Amendment after Gate 3's measurements β€” the fallback trigger fired, and here is what was done about it

**Read Β§6 above first: it is kept verbatim, including its numbers, which are now known to be wrong.** The
11,062 tok/s figure came from p0c's *raw* training loop at seq 1024 on a 113M model β€” no `Trainer`, no
real input pipeline, no fp16-master/GradScaler bookkeeping, no seq-2048 activations. Gate 3 measured the
frozen 106,194,240-parameter shape inside the actual stack on 2Γ—T4 (ledger rows 10-14, `memory/QUOTA.md`):

| config | tok/s | peak GB | 1.0 B tokens |
|---|---|---|---|
| sdpa @2048 micro 2 (the Β§6 plan-of-record) | 4,071 | 14.22 | **68.2 h** |
| eager + grad-ckpt @2048 micro 2 | 4,679 | 10.04 | 59.4 h |
| eager @1024 micro 4 | 7,732 | 12.25 | 35.9 h |
| eager + grad-ckpt @1024 micro 8 | 7,828 | 7.12 | 35.5 h |
| eager @2048 micro 2 | OOM | β€” | β€” |

Three consequences, in the order that matters:

1. **Sequence length changed to 1024** β€” that is **D-011**, decided on these numbers and documented in
   `memory/DECISIONS.md`. Micro-batch 4 Γ— accum 32 Γ— 2 cards keeps tokens/step at **262,144**, so the
   batch size in tokens, the ~3,815 steps, the LR and the trapezoid are all unchanged. Β§2.3's parameter
   table is unaffected β€” it never depended on the measurement.
2. **The pre-registered fallback condition in Β§6 fired.** "Below ~10.5k tok/s" is satisfied: every
   measured seq-1024 cell is ~7.7-7.8k. Mechanically that means 0.9 B.
3. **The premise of that trigger is gone.** 10.5k was chosen so 1.0 B would fit *one* 30 h week with
   enough slack for interruption overhead. At the measured rate neither 1.0 B (35.9 h) nor 0.9 B (32.3 h)
   fits week 1, so the trigger no longer discriminates between the two β€” it only buys 3.6 GPU-hours for a
   10 % cut in training tokens. Against the two-week schedule in `docs/04-run-log.md` Β§2 (60 h of quota,
   35.9 h of training, **~22 h of slack**) the 1.0 B target is affordable.

**Therefore, re-registered before launch and before any checkpoint or score exists (D-012): the token
target stays 1.0 B, and the fallback trigger is restated against measured throughput rather than against
p0c's number.** New mechanical rule, same purpose as the old one: **if preflight P3's end-to-end rate at
the frozen geometry (22L/576, seq 1024, eager, grad-ckpt on, micro 4, accum 32) is below 5,150 tok/s β€”
i.e. 1 B tokens needs more than ~54 h of the ~60 h available across the two weeks after allowing 6 h for
interruption overhead and Phase 6's GPU evaluation β€” the target drops to 0.9 B at launch, not later.**
5,150 tok/s is 33 % below the planning rate, so this is a genuine tripwire, not a formality. What did
*not* change: D-005's benchmark bands, the harness pin (D-010), the mix, the model shape, the optimiser,
the schedule, or the one-run rule. If the 0.9 B fallback fires, it fires as a documented arithmetic
consequence, and this section is where a reader can check it.

## 7. Benchmark targets β€” pre-registered before the main run (Β§4)

Recorded now, before any training, so the goalposts cannot move. Full anchor table, protocol and noise
floor in Β§7.1–7.3. Sources: archived Open LLM Leaderboard run metadata and Pythia Appendix G
([arXiv 2304.01373](https://arxiv.org/pdf/2304.01373)), read 2026-09-19.

**The dominant fact: every published β‰₯100M base anchor was trained on ~300 B tokens.** Our budget is
~1 B β€” a **300Γ— gap**. Those anchors are therefore ceilings, not expectations. No harness score for any
50–350M base trained at ~1–3 B tokens could be found published anywhere, so these bands are
extrapolations from the low end of the size curve and are labelled as such.

| Task | metric | band | chance |
|---|---|---|---|
| ARC-Easy | `acc` | **26–34** | 25 |
| ARC-Challenge | `acc` / `acc_norm` | **17–22** / **20–25** | 25 |
| HellaSwag | `acc_norm` (report `acc` too) | **26–32** (acc 25–30) | 25 |
| PIQA | `acc` | **52–60** | 50 |
| WinoGrande | `acc` | **49–53** | 50 |
| MMLU | `acc` | **24–27** | 25 |
| TruthfulQA | `mc2` / `mc1` | **36–46** / **21–26** | β‰ˆ38 / β‰ˆ22 |
| GSM8K | `exact_match,strict-match` | **0.0–1.5** | β‰ˆ0 |

- **MMLU, WinoGrande, GSM8K and ARC-Challenge are expected to be indistinguishable from chance.** MMLU
  measures flat 24.9–27.3 from 70M to 1.4B β€” OPT-1.3B scores *lower* than OPT-125M β€” so it has almost
  no discriminative power at this scale. GSM8K anchors: 0.23 (OPT-125M), 0.68 (GPT-2, Pythia-410M),
  1.52 (Pythia-1.4B), 1.4 even for SmolLM2-135M at 11.2 T tokens. Multi-step arithmetic does not emerge
  from a 100M base at 1 B tokens; predicting otherwise would be the inflation Β§3.12 forbids.
- **The real signals are ARC-Easy, HellaSwag and PIQA** β€” the three where anchors clear chance by a
  visible margin. Failing to beat chance there is a finding about the mix or the run, not the benchmarks.
- **Noise floor:** SE β‰ˆ 1.3 pp ARC-C (n=1172), 1.4 pp WinoGrande (1267), 1.2 pp PIQA (1838), 1.5 pp
  TruthfulQA (817), 0.45 pp HellaSwag (10042). **Differences under ~2 pp on ARC/WinoGrande/TruthfulQA
  are not results**, and no shot count or prompt format gets chosen per-task after seeing scores (Β§3.3).

### 7.1 Protocol, pinned for Phase 6

Harness **EleutherAI `lm-evaluation-harness` v0.4.13** (2026-08-31) β€” still what published anchors
report against; Lighteval v0.13.0 is the active HF alternative but switching would break comparability.
Record the git SHA per run; cadence is real (v0.4.10 2026-01 β†’ v0.4.13 2026-08).

| Task | dataset / config | split | shots |
|---|---|---|---|
| `arc_easy`, `arc_challenge` | `allenai/ai2_arc` | **test** (2376 / 1172) | **25** |
| `hellaswag` | `Rowan/hellaswag` | **validation** (10042; no test split) | **10** |
| `piqa` | `baber/piqa` default | **validation** (1838) | **10** |
| `winogrande` | `allenai/winogrande`, **`winogrande_xl`** | **validation** (1267) | **5** |
| `mmlu` | group of 57 `mmlu_<subject>` | test | **5** |
| `truthfulqa_mc1`, `_mc2` | `sylinrl/TruthfulQA` | validation | **0** (pinned) |
| `gsm8k` | `openai/gsm8k` `main` | test | **5** |

### 7.2 Traps that would silently falsify the numbers

1. **Shots are NOT pinned in the ARC/HellaSwag/PIQA/WinoGrande/MMLU YAMLs** β€” an omitted
   `--num_fewshot` silently yields 0-shot. Pass shots explicitly for every task.
2. **No chat templates.** `--apply_chat_template` / `--fewshot_as_multiturn` are for fine-tuned models
   (Β§3.9 makes this a base model). `gsm8k_cot_llama` requires them β†’ not used.
3. Everything but GSM8K is `output_type: multiple_choice` = loglikelihood ranking; the model never
   generates, so only GSM8K is sensitive to generation settings.
4. **MMLU aggregation:** the `mmlu` group uses `weight_by_size: True`, while archived leaderboard numbers
   were an *unweighted* subject mean. Compute and state which was used.
5. **TruthfulQA mc2 changed definition** 2024-03-11 (PR #2768); pre-April-2024 mc2 is not comparable.
   mc1/mc2 differ by ~20 pp and both report under the metric name `acc`.
6. **Harness PIQA is mean per-item accuracy, not AI2's `p_win`** β€” the harness never computes `p_win`.
7. **Winogrande**: `winogrande_xl`, single fold, validation; the official leaderboard averages 5 folds,
   and `winogrande_debiased` is a different number.
8. **GSM8K**: `temperature 0`, `do_sample false`, `max_gen_toks` inherits the global 256; tiny bases loop
   repetitively and never emit `#### `, so strict-match β‰ˆ0 while flexible-extract looks inflated. Report
   both filters or neither; do **not** add `repetition_penalty`, which silently deviates from everyone.
9. **Splits are mixed** β€” PIQA/HellaSwag/WinoGrande/TruthfulQA run on *validation*. Never call them test.
10. **Parameter-count conventions differ between suites**; Pythia's includes embeddings. Quote ours with
    the convention attached.

## 8. Still open before Gate 1 closes

- ~~Deep-thin-vs-wide throughput result~~ β†’ scheduled as Phase 3 preflight test T1.
- Session wall-clock cap, from the still-running `p0e-session-cap`: it sets how much progress one session
  can make and therefore how often the checkpoint cycle interrupts training. **Known so far: a CPU session
  was still alive at 51 min** with no cap hit, so short sessions are not the failure mode; the ceiling is
  somewhere above that.
- **Closed this session:** the credential path (Β§1 last row) β€” jobs retrieve the HF token from the
  account's own private Kaggle dataset via `code/ounce100m_credentials.py`; a Hub write from inside a job
  was verified by anonymous readback (D-006). And the checkpoint arithmetic below.

### 8.1 Checkpoint size and cadence β€” now measured, and it is cheap

A `latest`-quality exact-resume checkpoint for a ~100M model is **~1.6 GB** (fp32 weights 400 MB + two
fp32 Adam moments 800 MB + master/grad copies + scheduler/scaler/RNG state, which are all negligible
next to the optimizer). Pushed from a Kaggle session at **42.7 MB/s β†’ 37.4 s**, and pulled back in
**13.8 s**, with byte-exact anonymous readback at 200 MB / 800 MB / 1.6 GB
(`dodosoomro/ounce100m-p1-push-bench`, D-007).

Consequences that change design choices rather than just confirming them:

- 10 checkpoints + `latest` β‰ˆ **7 min total, ~0.12 GPU-h of a 30 h week**. Checkpoint cadence is not a
  budget problem, so there is no reason to economise on it β€” and rolling `latest` *more* often than the
  required 10 % is nearly free, which is worth doing since every interruption costs at most one roll.
- A cold resume with an empty disk costs seconds of transfer, not minutes. The expensive part of a resume
  is the `skip_first_batches` replay discussed in Β§5, which is why the input format in
  `docs/02-mix-plan.md` Β§5 step 6 is designed to make it an index advance.
- Peak disk: 1.6 GB written + 1.6 GB staged elsewhere on the 19.5 GB volume is comfortable, and the
  Β§3.13 prune-after-verify step keeps at most one copy resident.