File size: 26,939 Bytes
f345921 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 | # 01 β Plan: architecture, hyperparameters, throughput, targets
Phase 1 deliverable. Every choice below cites a source actually inspected on 2026-09-19, as Gate 1
requires. Items marked **β test** are decisions whose *numbers* Phase 3 must validate on the real
pipeline before the run freezes; the design itself is settled.
---
## 1. The envelope the design has to fit
Measured in Phase 0 (`docs/00-platform-notes.md`), not assumed:
| Constraint | Value | Consequence |
|---|---|---|
| Accelerator | 2x Tesla T4, cc 7.5, **14.56 GiB usable each** | fp16 only; no bf16 tensor cores; no flash-attention (sm80+ only) |
| GPU quota | 108,000 s/week, billed at **~1x container wall-clock** (3 samples) | β30 wall-clock h/week, both cards |
| Test budget | β€6 GPU-h lifetime | 0.073 h spent; **5.93 h remains for all of Phase 3** |
| Working disk in a job | **19.5 GB** `/kaggle/working` | shards consumed a few at a time, never whole (Β§3.13) |
| RAM / CPU | 30 GiB cgroup, 4 vCPU, **no swap** | input pipeline memory-bounded by design |
| Image | torch 2.10.0+cu128, **transformers 5.0.0**, datasets 5.0.0, accelerate 1.13.0, triton 3.6.0; **no** trl/deepspeed/flash-attn/xformers/bnb | prefer the preinstalled stack; each install is risk + round trip |
| Hub from a job | reads anonymous & fine, 35β89 MB/s; writes **401** until given a token | credential path: private Kaggle dataset mount (Β§8) |
**Parameter convention.** Β§2's 90β110M is *including embeddings* β the Pythia convention (Pythia was
explicitly renamed to include embedding + unembedding, [pythia README](https://raw.githubusercontent.com/pythia/main/README.md)).
This matters more at 100M than anywhere else: the tied embedding is 21β37 % of the model across the
candidates below. `code/config/param_count.py` computes it two ways and both were checked to agree
exactly against transformers 5.0.0 in `dodosoomro/ounce100m-p1-param-count`.
## 2. Architecture
### 2.1 The shape: deep-thin, tied, GQA, SwiGLU, pre-norm
**Choice: `hidden 576, layers 22, 9 Q heads / 3 KV heads (GQA 3:1), SwiGLU intermediate 1536, RMSNorm
pre-norm eps 1e-5, attention_bias false, RoPE theta 10,000, tied embeddings, vocab 49,152, seq 2048`.**
Evidence, and it is convergent rather than a single citation:
- **Deep-thin beats wide at this size.** Meta's MobileLLM ablated exactly this at a fixed 135M: 30
layers Γ d512 scored **44.8 avg vs 43.9 for 12 Γ d768**, with embedding-sharing removing 13 % of
parameters at ~equal accuracy ([arXiv 2402.14905](https://arxiv.org/abs/2402.14905), HIGH). Its
shipped 125M model is **30 layers, d576, 9Q/3KV GQA, SwiGLU, tied**, 124.6M total.
- **The same shape is what HF's SmolLM2-135M uses**: 30 layers, d576, 9Q/**3KV**, `intermediate_size`
1536 (2.67Γd), silu, `attention_bias=false`, **`tie_word_embeddings=true`**, RMSNorm 1e-5,
[model config](https://huggingface.co/HuggingFaceTB/SmolLM2-135M/blob/main/config.json) +
[arXiv 2502.02737 Β§6](https://arxiv.org/abs/2502.02737). Two independent labs landing on
d576 + GQA 3:1 + SwiGLU at ~130M is the strongest available signal.
- Our budget is 110M, not 135M, so the same d576/GQA3:1/SwiGLU-1536 stack is cut to **22 layers**
rather than 30 β verified count below β and RoPE ΞΈ drops to 10,000 to match the shorter context
(SmolLM **v1** used ΞΈ=10,000 at ctx 2048; SmolLM2's 100,000 accompanies ctx 8192, which we cannot
afford and do not need for these benchmarks).
- **GQA at 3:1 is a parameter choice as much as an attention choice**, since k/v are `d_head Γ n_kv`
wide; MobileLLM reports GQA-plus-enlarged-d as +0.4 avg.
**Not chosen: Mamba/SSM or hybrid.** The official kernels *do* build for `sm_75`
([state-spaces/mamba setup.py](https://github.com/state-spaces/mamba)), which corrects the common
assumption that they are Ampere-only β but they are absent from the image (source install only), and
the research found **no published 50β200M SSM/hybrid with released hyperparameters** to imitate. That is
research risk, not a default, and Β§4 forbids the kind of custom code where autonomous projects die.
### 2.2 Verified parameter counts
Closed form and `sum(p.numel())` agreed exactly on all eight shapes, in the target environment.
Tying confirmed genuinely tied, not merely declared.
| candidate | H | L | heads | kv | FFN | vocab | **total params** | emb share | in 90β110M |
|---|---|---|---|---|---|---|---|---|---|
| Phase-0 probe shape | 768 | 12 | 12 | 3 | 2048 | 50257 | 112,934,400 | 34.2 % | β |
| A | 768 | 11 | 12 | 3 | 2048 | 50257 | 106,739,712 | 36.2 % | β |
| B | 768 | 12 | 12 | 3 | 1792 | 50257 | 105,856,512 | 36.5 % | β |
| C | 768 | 12 | 12 | 2 | 1824 | 50257 | 105,561,600 | 36.6 % | β |
| D | 640 | 16 | 10 | 4 | 2048 | 50257 | 113,450,240 | 28.4 % | β |
| E | 768 | 12 | 12 | 3 | 2048 | 32768 | 99,502,848 | 25.3 % | β |
| F | 896 | 9 | 14 | 2 | 2432 | 50257 | 120,397,312 | 37.4 % | β |
| G | 1024 | 8 | 8 | 2 | 2816 | 50257 | 141,658,112 | 36.3 % | β |
**The binding insight: the tokenizer decides how much model you are allowed to have.** At 50,257 vocab
and d768 the tied embedding is 38.6M β a third of the entire budget β which is why every 50k candidate
clusters at 105β113M and must trim depth or FFN. MobileLLM's own ablation reached the same conclusion
(embedding sharing: β13 % parameters at ~equal accuracy). Choosing a smaller `hidden` is therefore not
a weakening, it is how the embedding tax gets paid for.
**β test:** the frozen shape's exact count was re-derived from the constructed model β see Β§2.3.
### 2.3 The frozen shape, counted to the unit
`dodosoomro/ounce100m-p1-param-count` v2 built each candidate as a real `LlamaForCausalLM` under
transformers 5.0.0. Closed form and `sum(p.numel())` agree exactly on every row.
| candidate | H | L | Q/KV | FFN | vocab | **total params** | emb share | in 90β110M |
|---|---|---|---|---|---|---|---|---|
| H | 576 | 20 | 9/3 | 1536 | 49,152 | 99,114,048 | 28.6 % | β |
| **I β CHOSEN** | **576** | **22** | **9/3** | **1536** | **49,152** | **106,194,240** | **26.7 %** | β |
| J | 576 | 24 | 9/3 | 1536 | 49,152 | 113,274,432 | 25.0 % | β over |
| K | 576 | 26 | 9/3 | 1536 | 49,152 | 120,354,624 | 23.5 % | β over |
| L | 640 | 20 | 10/4 | 1728 | 49,152 | 120,776,320 | 26.0 % | β over |
**Frozen: 106,194,240 parameters (embeddings included, tied), counted as
`vocabΓh` once + `L Γ (attn + SwiGLU)` + norms + final norm, and confirmed by constructing the model.**
The counting method is stated here because Β§2 requires the number *and* how it was counted.
Two consequences worth recording:
- **Depth is capped at 22 by the budget, not by taste.** Each layer at d576 costs 3,538,944 parameters,
so 24 layers is 113.3M β over. That is a hard stop, and it means the Phase-3 "deep-thin vs wide"
throughput test (T1) must compare **H (20 layers, 99.1M) against I (22 layers, 106.2M)**, both
in-budget, rather than the 12-layer d768 probe shape which cannot fit the SmolLM2 vocab without
trimming elsewhere.
- **Widening is not a free alternative to deepening.** Candidate L at d640 Γ 20 layers is *over* budget
(120.8M) despite having fewer layers than I, because attention and MLP cost scale with `hΒ²`/`hΒ·ff`.
So the deep-thin choice is forced from two directions: it is better per parameter (MobileLLM's
ablation) and it is the only way to spend the 78M non-embedding budget on 22 layers.
## 3. Tokenizer
**Choice: reuse the SmolLM2 BPE tokenizer** β vocab 49,152, byte-level, **Apache-2.0**, from
[`HuggingFaceTB/SmolLM2-135M`](https://huggingface.co/HuggingFaceTB/SmolLM2-135M)'s `tokenizer.json`;
trained on SmolCorpus per [arXiv 2502.02737](https://arxiv.org/abs/2502.02737). Not trained by us.
Β§4 says maximum reuse, minimal custom code, and training a tokenizer from scratch is precisely the
kind of optional machinery that adds failure surface without adding capability. Candidates checked today:
SmolLM2 **49,152 / Apache-2.0**; `gpt2` 50,257 / MIT; `pythia-160m` 50,304 / Apache-2.0;
`t5-v1_1-small` 32,128 Unigram (not byte-level, carries `extra_ids` sentinels); `Qwen3-0.6B` 151,936
(disqualifying β its embedding alone would be ~78M at d512); `TinyLlama` 32,000 but Llama-2-derived and
**redistribution status unverified** β excluded.
SmolLM2 wins on three counts: licence is clean and permissive, it is byte-level and English-curated at
exactly this model scale, and sharing it with a family of released 135M models gives the benchmark
numbers somebody else's token statistics to be compared against. Its 49,152 Γ 576 tied embedding is
28.3M = **21 %** of budget, versus 34β37 % for the 50k/d768 shapes.
## 4. Optimizer, batch, schedule, precision
| | choice | justification |
|---|---|---|
| Optimizer | AdamW, Ξ²=(0.9, 0.95), wd 0.1, clip 1.0 | Pythia-160M and SmolLM2 both use Ξ²β=0.95 at small scale ([2304.01373](https://arxiv.org/abs/2304.01373), [2502.02737](https://arxiv.org/abs/2502.02737)) |
| Peak LR | **6e-4** | bracketed from three directions: Cerebras-111M 6.0e-4, Pythia-160M 6e-4, and the 150M low-tokens/param run of Sardana et al. at 4.6e-4 ([2304.03208](https://arxiv.org/abs/2304.03208), [2401.00448](https://arxiv.org/abs/2401.00448)). The 2e-3/3e-3 figures that appear in small-model recipes belong to 1β2 T-token runs, which this is not. |
| Warmup | 2 % of steps, floor 100 steps | Pythia used 1 %; SmolLM2 used 2,000 steps at 2M tok/batch (~0.2 %). Warmer than 1 % because **fp16 without bf16 needs the ramp** β Phase 0 showed an unramped/unscaled fp16 run going straight to NaN (E-006). |
| Schedule | **trapezoid / WSD: constant, then linear decay to 0 over the final 20 %** β *not* cosine-to-10 % | SmolLM1 used a trapezoid with 20 % cooldown; SmolLM2 used WSD ([blog/smollm](https://huggingface.co/blog/smollm), [2502.02737 App. A](https://arxiv.org/abs/2502.02737)). *Straight to Zero* finds linear-decay-to-zero beats cosine-to-10 % when peak LR is optimal, across sizes/batch/data ([2502.15938](https://arxiv.org/abs/2502.15938)); COLT 2026 theory says decay *shape* barely matters but **overly slow terminal decay causes schedule-induced capacity saturation** ([2602.06797](https://arxiv.org/abs/2602.06797)). Its caveat applies: their evidence is at/above compute-optimal tokens/param, we are at half. |
| Global batch | **β262 k tokens/step** (~3,815 steps for 1 B tokens) | Critical batch size fits `B* = 621.341Β·N^0.087`, i.e. **nearly independent of model size** ([2410.21676](https://arxiv.org/abs/2410.21676)) β no reason to chase 2β4M-token batches. Sardana et al. deliberately used *smaller* batches for smaller models "so that low-token-count training runs see enough steps". At 10 tokens/param we are such a run. |
| Micro-batch | **2 Γ 2048 per GPU Γ 2 GPUs Γ 32 accumulation** | Set by memory, not taste: Phase 0 measured bs4Γ1024 at 11.5β13.8 GB of 14.56 GB and **bs8 OOMed outright**. β test |
| Precision | **fp16 autocast + fp32 master weights + `GradScaler` + grad clip** | Only mixed-precision option on Turing; transformers v5 documents this path by name ("fall back to fp16 on older hardware like V100 or T4") and warns **load the model in fp32 or autocast is a no-op** ([v5 mixed_precision docs](https://huggingface.co/docs/transformers/perf_training)). Verified finite in Phase 0 probe C; bf16 *runs* but is emulated β must be asserted off, not defaulted. |
| Seed / RNG | single fixed seed, recorded | Β§3.1 requires "same RNG semantics" across resumes |
**Undertrained by design, and honestly so.** 1 B tokens on ~100M β **10 tokens/param**, about half
Chinchilla-optimal (Cerebras-111M's compute-optimal point was 2.2 B tokens = 20 t/p). That is mildly
undertrained, not pathological β the genuine anomaly is 2025-26 practice, which overtrains tiny models
by three orders of magnitude (SmolLM2-135M at 2 T tokens β 14,800 t/p). Sardana et al. trained a
150M/d768/12L model across 3β10,000 t/p and quality kept improving; Gadre et al. show scaling laws
extrapolate across over-training ([2403.08540](https://arxiv.org/abs/2403.08540)). We are on the
left-hand side of that curve because the quota puts us there, and the report should say so.
## 5. Training stack
**Choice: `torchrun --nproc_per_node=2` + Hugging Face `Trainer` / `TrainingArguments`, `fp16=True`,
on a pre-tokenised **map-style** `Dataset`.** Minimal glue, no framework authored here (Β§4).
- Everything needed is already in the image: transformers 5.0.0, accelerate 1.13.0, datasets 5.0.0.
- v5 `Trainer` restores **RNG (`_save_rng_state`/`_load_rng_load`), optimizer, LR scheduler and the fp16
GradScaler** β the exact set Β§3.1 demands β via `Accelerator.save_state`. Verified against
v5.17.0 source; the image's 5.0.0 is ~8 patch releases behind but the same major. **β test at Gate 3.**
- **The resume trap, stated plainly:** v5 recovers data position with `skip_first_batches`, which
**replays** discarded batches β O(steps) wasted work on an `IterableDataset`, up to 3,815 steps of
pure I/O after every interruption. Mitigation is architectural, not a flag: a **map-style, concatenated
token store** makes skipping an index advance rather than a decode. A `dataset_shard` layout with an
explicit recorded cursor is what Gate 3 must prove, and Β§Phase 2's "exact positional resumption"
requirement is the reason it is designed that way up front.
- **v5 migration traps to check in preflight** ([MIGRATION_GUIDE_V5.md](https://github.com/huggingface/transformers/blob/main/MIGRATION_GUIDE_V5.md), HF *Transformers v5* blog): PyTorch-only backend;
slow tokenizers gone (irrelevant, we use `tokenizers`); several `TrainingArguments` removed **without
a deprecation cycle**; attention moved to `AttentionInterface`; `remove_unused_columns` **still
defaults True** (silently drops dataset columns a loss function needs); `lr_scheduler_type` default is
still `"linear"`, *not* cosine β so an unstated schedule is a linear decay to zero over the whole run;
new `train_sampling_strategy` replaces the `group_by_length` bool;
`restore_callback_states_from_checkpoint` is required to restore scheduler state properly.
- **Disqualified:** **torchtitan** β `training_dtype` is typed `Literal["bfloat16","float32"]`, i.e.
**no fp16 path at all**, plus Hopper-oriented CI β unusable on T4 ([torchtitan config](https://github.com/pytorch/torchtitan)).
**unsloth** β fine-tuning/GRPO only, no from-scratch pretraining path. **nanoGPT** β dormant since
2025-11, no HF export, no distributed input pipeline. **FSDP** β ~100M params in fp32 master +
grads + AdamW is ~1.8 GB; DDP on 2Γ16 GB does not need sharding, and sharding would add failure
surface for nothing. **DeepSpeed / flash-attn / xformers / trl** β absent from the image; installing
them is risk without a capability we lack.
## 6. Throughput, and the token target it implies
Measured on a main-run-shaped model, `seq 1024`, bs4/card, DDP over both T4s: **11,062 tok/s
aggregate** (5,531/rank; single-GPU 6,478 SDPA / 7,239 eager; **85 % DDP efficiency**).
Sanity-checking that against compute: `6ND β 6 Γ 1.13e8 Γ 11,062 β 7.5e12 FLOP/s`, which is **~6 % of a
2ΓT4 fp16 tensor peak** (65 TFLOPS/card β an *unverified* datasheet prior; the spec page could not be
retrieved). Independent corroboration: a documented 163M / ctx-1024 run reached ~19.9k tok/s on a single
RTX 3090 ([gilesthomas.com](https://gilesthomas.com/2025/12/llm-from-scratch-28-training-a-base-model-from-scratch)),
and Pythia-160M's 1,030 A100-hours for 300 B tokens implies ~81k tok/s per A100. So ~11k tok/s on two
Turing cards is the right order of magnitude, and **the low implied MFU means this is a floor with headroom,
not a ceiling** β which is the opposite of the risk I most wanted to rule out.
**The tension to resolve by measurement, not argument:** deep-thin is better per *parameter*
(MobileLLM's ablation) but worse per *second* on 2018 hardware, because 22 small layers issue more
sequential kernel launches than 12 fat ones, and we are wall-clock-bound rather than
parameter-bound. Phase 3 therefore measures both shapes.
| token target | at 11.1k tok/s (measured) | at 9k tok/s (planning floor) |
|---|---|---|
| 1.10 B | 27.6 h | 34.0 h |
| **1.00 B** | **25.1 h** | **30.9 h** |
| 0.90 B | 22.6 h | 27.8 h |
**Decision: target 1.0 B tokens, with a pre-registered fallback to 0.9 B if Gate 3's end-to-end rate on
the frozen config and real input pipeline is below ~10.5k tok/s.** Both are inside Β§2's band, and the
fallback is chosen *now* rather than later so it cannot become a post-hoc excuse. Even the optimistic
column leaves under 5 h of slack in a 30 h week, which the interruption budget in Β§3.1 will eat; so the
plan assumes **a two-week run crossing the 2026-09-26 quota reset**, with a checkpoint at every 10 % and
a rolling `latest`, exactly as Β§2 requires.
## 6.1 Amendment after Gate 3's measurements β the fallback trigger fired, and here is what was done about it
**Read Β§6 above first: it is kept verbatim, including its numbers, which are now known to be wrong.** The
11,062 tok/s figure came from p0c's *raw* training loop at seq 1024 on a 113M model β no `Trainer`, no
real input pipeline, no fp16-master/GradScaler bookkeeping, no seq-2048 activations. Gate 3 measured the
frozen 106,194,240-parameter shape inside the actual stack on 2ΓT4 (ledger rows 10-14, `memory/QUOTA.md`):
| config | tok/s | peak GB | 1.0 B tokens |
|---|---|---|---|
| sdpa @2048 micro 2 (the Β§6 plan-of-record) | 4,071 | 14.22 | **68.2 h** |
| eager + grad-ckpt @2048 micro 2 | 4,679 | 10.04 | 59.4 h |
| eager @1024 micro 4 | 7,732 | 12.25 | 35.9 h |
| eager + grad-ckpt @1024 micro 8 | 7,828 | 7.12 | 35.5 h |
| eager @2048 micro 2 | OOM | β | β |
Three consequences, in the order that matters:
1. **Sequence length changed to 1024** β that is **D-011**, decided on these numbers and documented in
`memory/DECISIONS.md`. Micro-batch 4 Γ accum 32 Γ 2 cards keeps tokens/step at **262,144**, so the
batch size in tokens, the ~3,815 steps, the LR and the trapezoid are all unchanged. Β§2.3's parameter
table is unaffected β it never depended on the measurement.
2. **The pre-registered fallback condition in Β§6 fired.** "Below ~10.5k tok/s" is satisfied: every
measured seq-1024 cell is ~7.7-7.8k. Mechanically that means 0.9 B.
3. **The premise of that trigger is gone.** 10.5k was chosen so 1.0 B would fit *one* 30 h week with
enough slack for interruption overhead. At the measured rate neither 1.0 B (35.9 h) nor 0.9 B (32.3 h)
fits week 1, so the trigger no longer discriminates between the two β it only buys 3.6 GPU-hours for a
10 % cut in training tokens. Against the two-week schedule in `docs/04-run-log.md` Β§2 (60 h of quota,
35.9 h of training, **~22 h of slack**) the 1.0 B target is affordable.
**Therefore, re-registered before launch and before any checkpoint or score exists (D-012): the token
target stays 1.0 B, and the fallback trigger is restated against measured throughput rather than against
p0c's number.** New mechanical rule, same purpose as the old one: **if preflight P3's end-to-end rate at
the frozen geometry (22L/576, seq 1024, eager, grad-ckpt on, micro 4, accum 32) is below 5,150 tok/s β
i.e. 1 B tokens needs more than ~54 h of the ~60 h available across the two weeks after allowing 6 h for
interruption overhead and Phase 6's GPU evaluation β the target drops to 0.9 B at launch, not later.**
5,150 tok/s is 33 % below the planning rate, so this is a genuine tripwire, not a formality. What did
*not* change: D-005's benchmark bands, the harness pin (D-010), the mix, the model shape, the optimiser,
the schedule, or the one-run rule. If the 0.9 B fallback fires, it fires as a documented arithmetic
consequence, and this section is where a reader can check it.
## 7. Benchmark targets β pre-registered before the main run (Β§4)
Recorded now, before any training, so the goalposts cannot move. Full anchor table, protocol and noise
floor in Β§7.1β7.3. Sources: archived Open LLM Leaderboard run metadata and Pythia Appendix G
([arXiv 2304.01373](https://arxiv.org/pdf/2304.01373)), read 2026-09-19.
**The dominant fact: every published β₯100M base anchor was trained on ~300 B tokens.** Our budget is
~1 B β a **300Γ gap**. Those anchors are therefore ceilings, not expectations. No harness score for any
50β350M base trained at ~1β3 B tokens could be found published anywhere, so these bands are
extrapolations from the low end of the size curve and are labelled as such.
| Task | metric | band | chance |
|---|---|---|---|
| ARC-Easy | `acc` | **26β34** | 25 |
| ARC-Challenge | `acc` / `acc_norm` | **17β22** / **20β25** | 25 |
| HellaSwag | `acc_norm` (report `acc` too) | **26β32** (acc 25β30) | 25 |
| PIQA | `acc` | **52β60** | 50 |
| WinoGrande | `acc` | **49β53** | 50 |
| MMLU | `acc` | **24β27** | 25 |
| TruthfulQA | `mc2` / `mc1` | **36β46** / **21β26** | β38 / β22 |
| GSM8K | `exact_match,strict-match` | **0.0β1.5** | β0 |
- **MMLU, WinoGrande, GSM8K and ARC-Challenge are expected to be indistinguishable from chance.** MMLU
measures flat 24.9β27.3 from 70M to 1.4B β OPT-1.3B scores *lower* than OPT-125M β so it has almost
no discriminative power at this scale. GSM8K anchors: 0.23 (OPT-125M), 0.68 (GPT-2, Pythia-410M),
1.52 (Pythia-1.4B), 1.4 even for SmolLM2-135M at 11.2 T tokens. Multi-step arithmetic does not emerge
from a 100M base at 1 B tokens; predicting otherwise would be the inflation Β§3.12 forbids.
- **The real signals are ARC-Easy, HellaSwag and PIQA** β the three where anchors clear chance by a
visible margin. Failing to beat chance there is a finding about the mix or the run, not the benchmarks.
- **Noise floor:** SE β 1.3 pp ARC-C (n=1172), 1.4 pp WinoGrande (1267), 1.2 pp PIQA (1838), 1.5 pp
TruthfulQA (817), 0.45 pp HellaSwag (10042). **Differences under ~2 pp on ARC/WinoGrande/TruthfulQA
are not results**, and no shot count or prompt format gets chosen per-task after seeing scores (Β§3.3).
### 7.1 Protocol, pinned for Phase 6
Harness **EleutherAI `lm-evaluation-harness` v0.4.13** (2026-08-31) β still what published anchors
report against; Lighteval v0.13.0 is the active HF alternative but switching would break comparability.
Record the git SHA per run; cadence is real (v0.4.10 2026-01 β v0.4.13 2026-08).
| Task | dataset / config | split | shots |
|---|---|---|---|
| `arc_easy`, `arc_challenge` | `allenai/ai2_arc` | **test** (2376 / 1172) | **25** |
| `hellaswag` | `Rowan/hellaswag` | **validation** (10042; no test split) | **10** |
| `piqa` | `baber/piqa` default | **validation** (1838) | **10** |
| `winogrande` | `allenai/winogrande`, **`winogrande_xl`** | **validation** (1267) | **5** |
| `mmlu` | group of 57 `mmlu_<subject>` | test | **5** |
| `truthfulqa_mc1`, `_mc2` | `sylinrl/TruthfulQA` | validation | **0** (pinned) |
| `gsm8k` | `openai/gsm8k` `main` | test | **5** |
### 7.2 Traps that would silently falsify the numbers
1. **Shots are NOT pinned in the ARC/HellaSwag/PIQA/WinoGrande/MMLU YAMLs** β an omitted
`--num_fewshot` silently yields 0-shot. Pass shots explicitly for every task.
2. **No chat templates.** `--apply_chat_template` / `--fewshot_as_multiturn` are for fine-tuned models
(Β§3.9 makes this a base model). `gsm8k_cot_llama` requires them β not used.
3. Everything but GSM8K is `output_type: multiple_choice` = loglikelihood ranking; the model never
generates, so only GSM8K is sensitive to generation settings.
4. **MMLU aggregation:** the `mmlu` group uses `weight_by_size: True`, while archived leaderboard numbers
were an *unweighted* subject mean. Compute and state which was used.
5. **TruthfulQA mc2 changed definition** 2024-03-11 (PR #2768); pre-April-2024 mc2 is not comparable.
mc1/mc2 differ by ~20 pp and both report under the metric name `acc`.
6. **Harness PIQA is mean per-item accuracy, not AI2's `p_win`** β the harness never computes `p_win`.
7. **Winogrande**: `winogrande_xl`, single fold, validation; the official leaderboard averages 5 folds,
and `winogrande_debiased` is a different number.
8. **GSM8K**: `temperature 0`, `do_sample false`, `max_gen_toks` inherits the global 256; tiny bases loop
repetitively and never emit `#### `, so strict-match β0 while flexible-extract looks inflated. Report
both filters or neither; do **not** add `repetition_penalty`, which silently deviates from everyone.
9. **Splits are mixed** β PIQA/HellaSwag/WinoGrande/TruthfulQA run on *validation*. Never call them test.
10. **Parameter-count conventions differ between suites**; Pythia's includes embeddings. Quote ours with
the convention attached.
## 8. Still open before Gate 1 closes
- ~~Deep-thin-vs-wide throughput result~~ β scheduled as Phase 3 preflight test T1.
- Session wall-clock cap, from the still-running `p0e-session-cap`: it sets how much progress one session
can make and therefore how often the checkpoint cycle interrupts training. **Known so far: a CPU session
was still alive at 51 min** with no cap hit, so short sessions are not the failure mode; the ceiling is
somewhere above that.
- **Closed this session:** the credential path (Β§1 last row) β jobs retrieve the HF token from the
account's own private Kaggle dataset via `code/ounce100m_credentials.py`; a Hub write from inside a job
was verified by anonymous readback (D-006). And the checkpoint arithmetic below.
### 8.1 Checkpoint size and cadence β now measured, and it is cheap
A `latest`-quality exact-resume checkpoint for a ~100M model is **~1.6 GB** (fp32 weights 400 MB + two
fp32 Adam moments 800 MB + master/grad copies + scheduler/scaler/RNG state, which are all negligible
next to the optimizer). Pushed from a Kaggle session at **42.7 MB/s β 37.4 s**, and pulled back in
**13.8 s**, with byte-exact anonymous readback at 200 MB / 800 MB / 1.6 GB
(`dodosoomro/ounce100m-p1-push-bench`, D-007).
Consequences that change design choices rather than just confirming them:
- 10 checkpoints + `latest` β **7 min total, ~0.12 GPU-h of a 30 h week**. Checkpoint cadence is not a
budget problem, so there is no reason to economise on it β and rolling `latest` *more* often than the
required 10 % is nearly free, which is worth doing since every interruption costs at most one roll.
- A cold resume with an empty disk costs seconds of transfer, not minutes. The expensive part of a resume
is the `skip_first_batches` replay discussed in Β§5, which is why the input format in
`docs/02-mix-plan.md` Β§5 step 6 is designed to make it an index advance.
- Peak disk: 1.6 GB written + 1.6 GB staged elsewhere on the 19.5 GB volume is comfortable, and the
Β§3.13 prune-after-verify step keeps at most one copy resident.
|