ounce100m-code / docs /01-plan.md
Cion-lab's picture
publish docs/: plan, mix rationale, preflight report, run log, frozen eval protocol, final report
f345921 verified
|
Raw History Blame Contribute Delete
26.9 kB
# 01 β€” Plan: architecture, hyperparameters, throughput, targets
Phase 1 deliverable. Every choice below cites a source actually inspected on 2026-09-19, as Gate 1
requires. Items marked **β†’ test** are decisions whose *numbers* Phase 3 must validate on the real
pipeline before the run freezes; the design itself is settled.
---
## 1. The envelope the design has to fit
Measured in Phase 0 (`docs/00-platform-notes.md`), not assumed:
| Constraint | Value | Consequence |
|---|---|---|
| Accelerator | 2x Tesla T4, cc 7.5, **14.56 GiB usable each** | fp16 only; no bf16 tensor cores; no flash-attention (sm80+ only) |
| GPU quota | 108,000 s/week, billed at **~1x container wall-clock** (3 samples) | β‰ˆ30 wall-clock h/week, both cards |
| Test budget | ≀6 GPU-h lifetime | 0.073 h spent; **5.93 h remains for all of Phase 3** |
| Working disk in a job | **19.5 GB** `/kaggle/working` | shards consumed a few at a time, never whole (Β§3.13) |
| RAM / CPU | 30 GiB cgroup, 4 vCPU, **no swap** | input pipeline memory-bounded by design |
| Image | torch 2.10.0+cu128, **transformers 5.0.0**, datasets 5.0.0, accelerate 1.13.0, triton 3.6.0; **no** trl/deepspeed/flash-attn/xformers/bnb | prefer the preinstalled stack; each install is risk + round trip |
| Hub from a job | reads anonymous & fine, 35–89 MB/s; writes **401** until given a token | credential path: private Kaggle dataset mount (Β§8) |
**Parameter convention.** Β§2's 90–110M is *including embeddings* β€” the Pythia convention (Pythia was
explicitly renamed to include embedding + unembedding, [pythia README](https://raw.githubusercontent.com/pythia/main/README.md)).
This matters more at 100M than anywhere else: the tied embedding is 21–37 % of the model across the
candidates below. `code/config/param_count.py` computes it two ways and both were checked to agree
exactly against transformers 5.0.0 in `dodosoomro/ounce100m-p1-param-count`.
## 2. Architecture
### 2.1 The shape: deep-thin, tied, GQA, SwiGLU, pre-norm
**Choice: `hidden 576, layers 22, 9 Q heads / 3 KV heads (GQA 3:1), SwiGLU intermediate 1536, RMSNorm
pre-norm eps 1e-5, attention_bias false, RoPE theta 10,000, tied embeddings, vocab 49,152, seq 2048`.**
Evidence, and it is convergent rather than a single citation:
- **Deep-thin beats wide at this size.** Meta's MobileLLM ablated exactly this at a fixed 135M: 30
layers Γ— d512 scored **44.8 avg vs 43.9 for 12 Γ— d768**, with embedding-sharing removing 13 % of
parameters at ~equal accuracy ([arXiv 2402.14905](https://arxiv.org/abs/2402.14905), HIGH). Its
shipped 125M model is **30 layers, d576, 9Q/3KV GQA, SwiGLU, tied**, 124.6M total.
- **The same shape is what HF's SmolLM2-135M uses**: 30 layers, d576, 9Q/**3KV**, `intermediate_size`
1536 (2.67Γ—d), silu, `attention_bias=false`, **`tie_word_embeddings=true`**, RMSNorm 1e-5,
[model config](https://huggingface.co/HuggingFaceTB/SmolLM2-135M/blob/main/config.json) +
[arXiv 2502.02737 Β§6](https://arxiv.org/abs/2502.02737). Two independent labs landing on
d576 + GQA 3:1 + SwiGLU at ~130M is the strongest available signal.
- Our budget is 110M, not 135M, so the same d576/GQA3:1/SwiGLU-1536 stack is cut to **22 layers**
rather than 30 β€” verified count below β€” and RoPE ΞΈ drops to 10,000 to match the shorter context
(SmolLM **v1** used ΞΈ=10,000 at ctx 2048; SmolLM2's 100,000 accompanies ctx 8192, which we cannot
afford and do not need for these benchmarks).
- **GQA at 3:1 is a parameter choice as much as an attention choice**, since k/v are `d_head Γ— n_kv`
wide; MobileLLM reports GQA-plus-enlarged-d as +0.4 avg.
**Not chosen: Mamba/SSM or hybrid.** The official kernels *do* build for `sm_75`
([state-spaces/mamba setup.py](https://github.com/state-spaces/mamba)), which corrects the common
assumption that they are Ampere-only β€” but they are absent from the image (source install only), and
the research found **no published 50–200M SSM/hybrid with released hyperparameters** to imitate. That is
research risk, not a default, and Β§4 forbids the kind of custom code where autonomous projects die.
### 2.2 Verified parameter counts
Closed form and `sum(p.numel())` agreed exactly on all eight shapes, in the target environment.
Tying confirmed genuinely tied, not merely declared.
| candidate | H | L | heads | kv | FFN | vocab | **total params** | emb share | in 90–110M |
|---|---|---|---|---|---|---|---|---|---|
| Phase-0 probe shape | 768 | 12 | 12 | 3 | 2048 | 50257 | 112,934,400 | 34.2 % | βœ— |
| A | 768 | 11 | 12 | 3 | 2048 | 50257 | 106,739,712 | 36.2 % | βœ“ |
| B | 768 | 12 | 12 | 3 | 1792 | 50257 | 105,856,512 | 36.5 % | βœ“ |
| C | 768 | 12 | 12 | 2 | 1824 | 50257 | 105,561,600 | 36.6 % | βœ“ |
| D | 640 | 16 | 10 | 4 | 2048 | 50257 | 113,450,240 | 28.4 % | βœ— |
| E | 768 | 12 | 12 | 3 | 2048 | 32768 | 99,502,848 | 25.3 % | βœ“ |
| F | 896 | 9 | 14 | 2 | 2432 | 50257 | 120,397,312 | 37.4 % | βœ— |
| G | 1024 | 8 | 8 | 2 | 2816 | 50257 | 141,658,112 | 36.3 % | βœ— |
**The binding insight: the tokenizer decides how much model you are allowed to have.** At 50,257 vocab
and d768 the tied embedding is 38.6M β€” a third of the entire budget β€” which is why every 50k candidate
clusters at 105–113M and must trim depth or FFN. MobileLLM's own ablation reached the same conclusion
(embedding sharing: βˆ’13 % parameters at ~equal accuracy). Choosing a smaller `hidden` is therefore not
a weakening, it is how the embedding tax gets paid for.
**β†’ test:** the frozen shape's exact count was re-derived from the constructed model β€” see Β§2.3.
### 2.3 The frozen shape, counted to the unit
`dodosoomro/ounce100m-p1-param-count` v2 built each candidate as a real `LlamaForCausalLM` under
transformers 5.0.0. Closed form and `sum(p.numel())` agree exactly on every row.
| candidate | H | L | Q/KV | FFN | vocab | **total params** | emb share | in 90–110M |
|---|---|---|---|---|---|---|---|---|
| H | 576 | 20 | 9/3 | 1536 | 49,152 | 99,114,048 | 28.6 % | βœ“ |
| **I β€” CHOSEN** | **576** | **22** | **9/3** | **1536** | **49,152** | **106,194,240** | **26.7 %** | βœ“ |
| J | 576 | 24 | 9/3 | 1536 | 49,152 | 113,274,432 | 25.0 % | βœ— over |
| K | 576 | 26 | 9/3 | 1536 | 49,152 | 120,354,624 | 23.5 % | βœ— over |
| L | 640 | 20 | 10/4 | 1728 | 49,152 | 120,776,320 | 26.0 % | βœ— over |
**Frozen: 106,194,240 parameters (embeddings included, tied), counted as
`vocabΓ—h` once + `L Γ— (attn + SwiGLU)` + norms + final norm, and confirmed by constructing the model.**
The counting method is stated here because Β§2 requires the number *and* how it was counted.
Two consequences worth recording:
- **Depth is capped at 22 by the budget, not by taste.** Each layer at d576 costs 3,538,944 parameters,
so 24 layers is 113.3M β€” over. That is a hard stop, and it means the Phase-3 "deep-thin vs wide"
throughput test (T1) must compare **H (20 layers, 99.1M) against I (22 layers, 106.2M)**, both
in-budget, rather than the 12-layer d768 probe shape which cannot fit the SmolLM2 vocab without
trimming elsewhere.
- **Widening is not a free alternative to deepening.** Candidate L at d640 Γ— 20 layers is *over* budget
(120.8M) despite having fewer layers than I, because attention and MLP cost scale with `hΒ²`/`hΒ·ff`.
So the deep-thin choice is forced from two directions: it is better per parameter (MobileLLM's
ablation) and it is the only way to spend the 78M non-embedding budget on 22 layers.
## 3. Tokenizer
**Choice: reuse the SmolLM2 BPE tokenizer** β€” vocab 49,152, byte-level, **Apache-2.0**, from
[`HuggingFaceTB/SmolLM2-135M`](https://huggingface.co/HuggingFaceTB/SmolLM2-135M)'s `tokenizer.json`;
trained on SmolCorpus per [arXiv 2502.02737](https://arxiv.org/abs/2502.02737). Not trained by us.
Β§4 says maximum reuse, minimal custom code, and training a tokenizer from scratch is precisely the
kind of optional machinery that adds failure surface without adding capability. Candidates checked today:
SmolLM2 **49,152 / Apache-2.0**; `gpt2` 50,257 / MIT; `pythia-160m` 50,304 / Apache-2.0;
`t5-v1_1-small` 32,128 Unigram (not byte-level, carries `extra_ids` sentinels); `Qwen3-0.6B` 151,936
(disqualifying β€” its embedding alone would be ~78M at d512); `TinyLlama` 32,000 but Llama-2-derived and
**redistribution status unverified** β†’ excluded.
SmolLM2 wins on three counts: licence is clean and permissive, it is byte-level and English-curated at
exactly this model scale, and sharing it with a family of released 135M models gives the benchmark
numbers somebody else's token statistics to be compared against. Its 49,152 Γ— 576 tied embedding is
28.3M = **21 %** of budget, versus 34–37 % for the 50k/d768 shapes.
## 4. Optimizer, batch, schedule, precision
| | choice | justification |
|---|---|---|
| Optimizer | AdamW, Ξ²=(0.9, 0.95), wd 0.1, clip 1.0 | Pythia-160M and SmolLM2 both use Ξ²β‚‚=0.95 at small scale ([2304.01373](https://arxiv.org/abs/2304.01373), [2502.02737](https://arxiv.org/abs/2502.02737)) |
| Peak LR | **6e-4** | bracketed from three directions: Cerebras-111M 6.0e-4, Pythia-160M 6e-4, and the 150M low-tokens/param run of Sardana et al. at 4.6e-4 ([2304.03208](https://arxiv.org/abs/2304.03208), [2401.00448](https://arxiv.org/abs/2401.00448)). The 2e-3/3e-3 figures that appear in small-model recipes belong to 1–2 T-token runs, which this is not. |
| Warmup | 2 % of steps, floor 100 steps | Pythia used 1 %; SmolLM2 used 2,000 steps at 2M tok/batch (~0.2 %). Warmer than 1 % because **fp16 without bf16 needs the ramp** β€” Phase 0 showed an unramped/unscaled fp16 run going straight to NaN (E-006). |
| Schedule | **trapezoid / WSD: constant, then linear decay to 0 over the final 20 %** β€” *not* cosine-to-10 % | SmolLM1 used a trapezoid with 20 % cooldown; SmolLM2 used WSD ([blog/smollm](https://huggingface.co/blog/smollm), [2502.02737 App. A](https://arxiv.org/abs/2502.02737)). *Straight to Zero* finds linear-decay-to-zero beats cosine-to-10 % when peak LR is optimal, across sizes/batch/data ([2502.15938](https://arxiv.org/abs/2502.15938)); COLT 2026 theory says decay *shape* barely matters but **overly slow terminal decay causes schedule-induced capacity saturation** ([2602.06797](https://arxiv.org/abs/2602.06797)). Its caveat applies: their evidence is at/above compute-optimal tokens/param, we are at half. |
| Global batch | **β‰ˆ262 k tokens/step** (~3,815 steps for 1 B tokens) | Critical batch size fits `B* = 621.341Β·N^0.087`, i.e. **nearly independent of model size** ([2410.21676](https://arxiv.org/abs/2410.21676)) β€” no reason to chase 2–4M-token batches. Sardana et al. deliberately used *smaller* batches for smaller models "so that low-token-count training runs see enough steps". At 10 tokens/param we are such a run. |
| Micro-batch | **2 Γ— 2048 per GPU Γ— 2 GPUs Γ— 32 accumulation** | Set by memory, not taste: Phase 0 measured bs4Γ—1024 at 11.5–13.8 GB of 14.56 GB and **bs8 OOMed outright**. β†’ test |
| Precision | **fp16 autocast + fp32 master weights + `GradScaler` + grad clip** | Only mixed-precision option on Turing; transformers v5 documents this path by name ("fall back to fp16 on older hardware like V100 or T4") and warns **load the model in fp32 or autocast is a no-op** ([v5 mixed_precision docs](https://huggingface.co/docs/transformers/perf_training)). Verified finite in Phase 0 probe C; bf16 *runs* but is emulated β†’ must be asserted off, not defaulted. |
| Seed / RNG | single fixed seed, recorded | Β§3.1 requires "same RNG semantics" across resumes |
**Undertrained by design, and honestly so.** 1 B tokens on ~100M β‰ˆ **10 tokens/param**, about half
Chinchilla-optimal (Cerebras-111M's compute-optimal point was 2.2 B tokens = 20 t/p). That is mildly
undertrained, not pathological β€” the genuine anomaly is 2025-26 practice, which overtrains tiny models
by three orders of magnitude (SmolLM2-135M at 2 T tokens β‰ˆ 14,800 t/p). Sardana et al. trained a
150M/d768/12L model across 3β†’10,000 t/p and quality kept improving; Gadre et al. show scaling laws
extrapolate across over-training ([2403.08540](https://arxiv.org/abs/2403.08540)). We are on the
left-hand side of that curve because the quota puts us there, and the report should say so.
## 5. Training stack
**Choice: `torchrun --nproc_per_node=2` + Hugging Face `Trainer` / `TrainingArguments`, `fp16=True`,
on a pre-tokenised **map-style** `Dataset`.** Minimal glue, no framework authored here (Β§4).
- Everything needed is already in the image: transformers 5.0.0, accelerate 1.13.0, datasets 5.0.0.
- v5 `Trainer` restores **RNG (`_save_rng_state`/`_load_rng_load`), optimizer, LR scheduler and the fp16
GradScaler** β€” the exact set Β§3.1 demands β€” via `Accelerator.save_state`. Verified against
v5.17.0 source; the image's 5.0.0 is ~8 patch releases behind but the same major. **β†’ test at Gate 3.**
- **The resume trap, stated plainly:** v5 recovers data position with `skip_first_batches`, which
**replays** discarded batches β€” O(steps) wasted work on an `IterableDataset`, up to 3,815 steps of
pure I/O after every interruption. Mitigation is architectural, not a flag: a **map-style, concatenated
token store** makes skipping an index advance rather than a decode. A `dataset_shard` layout with an
explicit recorded cursor is what Gate 3 must prove, and Β§Phase 2's "exact positional resumption"
requirement is the reason it is designed that way up front.
- **v5 migration traps to check in preflight** ([MIGRATION_GUIDE_V5.md](https://github.com/huggingface/transformers/blob/main/MIGRATION_GUIDE_V5.md), HF *Transformers v5* blog): PyTorch-only backend;
slow tokenizers gone (irrelevant, we use `tokenizers`); several `TrainingArguments` removed **without
a deprecation cycle**; attention moved to `AttentionInterface`; `remove_unused_columns` **still
defaults True** (silently drops dataset columns a loss function needs); `lr_scheduler_type` default is
still `"linear"`, *not* cosine β€” so an unstated schedule is a linear decay to zero over the whole run;
new `train_sampling_strategy` replaces the `group_by_length` bool;
`restore_callback_states_from_checkpoint` is required to restore scheduler state properly.
- **Disqualified:** **torchtitan** β€” `training_dtype` is typed `Literal["bfloat16","float32"]`, i.e.
**no fp16 path at all**, plus Hopper-oriented CI β†’ unusable on T4 ([torchtitan config](https://github.com/pytorch/torchtitan)).
**unsloth** β€” fine-tuning/GRPO only, no from-scratch pretraining path. **nanoGPT** β€” dormant since
2025-11, no HF export, no distributed input pipeline. **FSDP** β€” ~100M params in fp32 master +
grads + AdamW is ~1.8 GB; DDP on 2Γ—16 GB does not need sharding, and sharding would add failure
surface for nothing. **DeepSpeed / flash-attn / xformers / trl** β€” absent from the image; installing
them is risk without a capability we lack.
## 6. Throughput, and the token target it implies
Measured on a main-run-shaped model, `seq 1024`, bs4/card, DDP over both T4s: **11,062 tok/s
aggregate** (5,531/rank; single-GPU 6,478 SDPA / 7,239 eager; **85 % DDP efficiency**).
Sanity-checking that against compute: `6ND β‰ˆ 6 Γ— 1.13e8 Γ— 11,062 β‰ˆ 7.5e12 FLOP/s`, which is **~6 % of a
2Γ—T4 fp16 tensor peak** (65 TFLOPS/card β€” an *unverified* datasheet prior; the spec page could not be
retrieved). Independent corroboration: a documented 163M / ctx-1024 run reached ~19.9k tok/s on a single
RTX 3090 ([gilesthomas.com](https://gilesthomas.com/2025/12/llm-from-scratch-28-training-a-base-model-from-scratch)),
and Pythia-160M's 1,030 A100-hours for 300 B tokens implies ~81k tok/s per A100. So ~11k tok/s on two
Turing cards is the right order of magnitude, and **the low implied MFU means this is a floor with headroom,
not a ceiling** β€” which is the opposite of the risk I most wanted to rule out.
**The tension to resolve by measurement, not argument:** deep-thin is better per *parameter*
(MobileLLM's ablation) but worse per *second* on 2018 hardware, because 22 small layers issue more
sequential kernel launches than 12 fat ones, and we are wall-clock-bound rather than
parameter-bound. Phase 3 therefore measures both shapes.
| token target | at 11.1k tok/s (measured) | at 9k tok/s (planning floor) |
|---|---|---|
| 1.10 B | 27.6 h | 34.0 h |
| **1.00 B** | **25.1 h** | **30.9 h** |
| 0.90 B | 22.6 h | 27.8 h |
**Decision: target 1.0 B tokens, with a pre-registered fallback to 0.9 B if Gate 3's end-to-end rate on
the frozen config and real input pipeline is below ~10.5k tok/s.** Both are inside Β§2's band, and the
fallback is chosen *now* rather than later so it cannot become a post-hoc excuse. Even the optimistic
column leaves under 5 h of slack in a 30 h week, which the interruption budget in Β§3.1 will eat; so the
plan assumes **a two-week run crossing the 2026-09-26 quota reset**, with a checkpoint at every 10 % and
a rolling `latest`, exactly as Β§2 requires.
## 6.1 Amendment after Gate 3's measurements β€” the fallback trigger fired, and here is what was done about it
**Read Β§6 above first: it is kept verbatim, including its numbers, which are now known to be wrong.** The
11,062 tok/s figure came from p0c's *raw* training loop at seq 1024 on a 113M model β€” no `Trainer`, no
real input pipeline, no fp16-master/GradScaler bookkeeping, no seq-2048 activations. Gate 3 measured the
frozen 106,194,240-parameter shape inside the actual stack on 2Γ—T4 (ledger rows 10-14, `memory/QUOTA.md`):
| config | tok/s | peak GB | 1.0 B tokens |
|---|---|---|---|
| sdpa @2048 micro 2 (the Β§6 plan-of-record) | 4,071 | 14.22 | **68.2 h** |
| eager + grad-ckpt @2048 micro 2 | 4,679 | 10.04 | 59.4 h |
| eager @1024 micro 4 | 7,732 | 12.25 | 35.9 h |
| eager + grad-ckpt @1024 micro 8 | 7,828 | 7.12 | 35.5 h |
| eager @2048 micro 2 | OOM | β€” | β€” |
Three consequences, in the order that matters:
1. **Sequence length changed to 1024** β€” that is **D-011**, decided on these numbers and documented in
`memory/DECISIONS.md`. Micro-batch 4 Γ— accum 32 Γ— 2 cards keeps tokens/step at **262,144**, so the
batch size in tokens, the ~3,815 steps, the LR and the trapezoid are all unchanged. Β§2.3's parameter
table is unaffected β€” it never depended on the measurement.
2. **The pre-registered fallback condition in Β§6 fired.** "Below ~10.5k tok/s" is satisfied: every
measured seq-1024 cell is ~7.7-7.8k. Mechanically that means 0.9 B.
3. **The premise of that trigger is gone.** 10.5k was chosen so 1.0 B would fit *one* 30 h week with
enough slack for interruption overhead. At the measured rate neither 1.0 B (35.9 h) nor 0.9 B (32.3 h)
fits week 1, so the trigger no longer discriminates between the two β€” it only buys 3.6 GPU-hours for a
10 % cut in training tokens. Against the two-week schedule in `docs/04-run-log.md` Β§2 (60 h of quota,
35.9 h of training, **~22 h of slack**) the 1.0 B target is affordable.
**Therefore, re-registered before launch and before any checkpoint or score exists (D-012): the token
target stays 1.0 B, and the fallback trigger is restated against measured throughput rather than against
p0c's number.** New mechanical rule, same purpose as the old one: **if preflight P3's end-to-end rate at
the frozen geometry (22L/576, seq 1024, eager, grad-ckpt on, micro 4, accum 32) is below 5,150 tok/s β€”
i.e. 1 B tokens needs more than ~54 h of the ~60 h available across the two weeks after allowing 6 h for
interruption overhead and Phase 6's GPU evaluation β€” the target drops to 0.9 B at launch, not later.**
5,150 tok/s is 33 % below the planning rate, so this is a genuine tripwire, not a formality. What did
*not* change: D-005's benchmark bands, the harness pin (D-010), the mix, the model shape, the optimiser,
the schedule, or the one-run rule. If the 0.9 B fallback fires, it fires as a documented arithmetic
consequence, and this section is where a reader can check it.
## 7. Benchmark targets β€” pre-registered before the main run (Β§4)
Recorded now, before any training, so the goalposts cannot move. Full anchor table, protocol and noise
floor in Β§7.1–7.3. Sources: archived Open LLM Leaderboard run metadata and Pythia Appendix G
([arXiv 2304.01373](https://arxiv.org/pdf/2304.01373)), read 2026-09-19.
**The dominant fact: every published β‰₯100M base anchor was trained on ~300 B tokens.** Our budget is
~1 B β€” a **300Γ— gap**. Those anchors are therefore ceilings, not expectations. No harness score for any
50–350M base trained at ~1–3 B tokens could be found published anywhere, so these bands are
extrapolations from the low end of the size curve and are labelled as such.
| Task | metric | band | chance |
|---|---|---|---|
| ARC-Easy | `acc` | **26–34** | 25 |
| ARC-Challenge | `acc` / `acc_norm` | **17–22** / **20–25** | 25 |
| HellaSwag | `acc_norm` (report `acc` too) | **26–32** (acc 25–30) | 25 |
| PIQA | `acc` | **52–60** | 50 |
| WinoGrande | `acc` | **49–53** | 50 |
| MMLU | `acc` | **24–27** | 25 |
| TruthfulQA | `mc2` / `mc1` | **36–46** / **21–26** | β‰ˆ38 / β‰ˆ22 |
| GSM8K | `exact_match,strict-match` | **0.0–1.5** | β‰ˆ0 |
- **MMLU, WinoGrande, GSM8K and ARC-Challenge are expected to be indistinguishable from chance.** MMLU
measures flat 24.9–27.3 from 70M to 1.4B β€” OPT-1.3B scores *lower* than OPT-125M β€” so it has almost
no discriminative power at this scale. GSM8K anchors: 0.23 (OPT-125M), 0.68 (GPT-2, Pythia-410M),
1.52 (Pythia-1.4B), 1.4 even for SmolLM2-135M at 11.2 T tokens. Multi-step arithmetic does not emerge
from a 100M base at 1 B tokens; predicting otherwise would be the inflation Β§3.12 forbids.
- **The real signals are ARC-Easy, HellaSwag and PIQA** β€” the three where anchors clear chance by a
visible margin. Failing to beat chance there is a finding about the mix or the run, not the benchmarks.
- **Noise floor:** SE β‰ˆ 1.3 pp ARC-C (n=1172), 1.4 pp WinoGrande (1267), 1.2 pp PIQA (1838), 1.5 pp
TruthfulQA (817), 0.45 pp HellaSwag (10042). **Differences under ~2 pp on ARC/WinoGrande/TruthfulQA
are not results**, and no shot count or prompt format gets chosen per-task after seeing scores (Β§3.3).
### 7.1 Protocol, pinned for Phase 6
Harness **EleutherAI `lm-evaluation-harness` v0.4.13** (2026-08-31) β€” still what published anchors
report against; Lighteval v0.13.0 is the active HF alternative but switching would break comparability.
Record the git SHA per run; cadence is real (v0.4.10 2026-01 β†’ v0.4.13 2026-08).
| Task | dataset / config | split | shots |
|---|---|---|---|
| `arc_easy`, `arc_challenge` | `allenai/ai2_arc` | **test** (2376 / 1172) | **25** |
| `hellaswag` | `Rowan/hellaswag` | **validation** (10042; no test split) | **10** |
| `piqa` | `baber/piqa` default | **validation** (1838) | **10** |
| `winogrande` | `allenai/winogrande`, **`winogrande_xl`** | **validation** (1267) | **5** |
| `mmlu` | group of 57 `mmlu_<subject>` | test | **5** |
| `truthfulqa_mc1`, `_mc2` | `sylinrl/TruthfulQA` | validation | **0** (pinned) |
| `gsm8k` | `openai/gsm8k` `main` | test | **5** |
### 7.2 Traps that would silently falsify the numbers
1. **Shots are NOT pinned in the ARC/HellaSwag/PIQA/WinoGrande/MMLU YAMLs** β€” an omitted
`--num_fewshot` silently yields 0-shot. Pass shots explicitly for every task.
2. **No chat templates.** `--apply_chat_template` / `--fewshot_as_multiturn` are for fine-tuned models
(Β§3.9 makes this a base model). `gsm8k_cot_llama` requires them β†’ not used.
3. Everything but GSM8K is `output_type: multiple_choice` = loglikelihood ranking; the model never
generates, so only GSM8K is sensitive to generation settings.
4. **MMLU aggregation:** the `mmlu` group uses `weight_by_size: True`, while archived leaderboard numbers
were an *unweighted* subject mean. Compute and state which was used.
5. **TruthfulQA mc2 changed definition** 2024-03-11 (PR #2768); pre-April-2024 mc2 is not comparable.
mc1/mc2 differ by ~20 pp and both report under the metric name `acc`.
6. **Harness PIQA is mean per-item accuracy, not AI2's `p_win`** β€” the harness never computes `p_win`.
7. **Winogrande**: `winogrande_xl`, single fold, validation; the official leaderboard averages 5 folds,
and `winogrande_debiased` is a different number.
8. **GSM8K**: `temperature 0`, `do_sample false`, `max_gen_toks` inherits the global 256; tiny bases loop
repetitively and never emit `#### `, so strict-match β‰ˆ0 while flexible-extract looks inflated. Report
both filters or neither; do **not** add `repetition_penalty`, which silently deviates from everyone.
9. **Splits are mixed** β€” PIQA/HellaSwag/WinoGrande/TruthfulQA run on *validation*. Never call them test.
10. **Parameter-count conventions differ between suites**; Pythia's includes embeddings. Quote ours with
the convention attached.
## 8. Still open before Gate 1 closes
- ~~Deep-thin-vs-wide throughput result~~ β†’ scheduled as Phase 3 preflight test T1.
- Session wall-clock cap, from the still-running `p0e-session-cap`: it sets how much progress one session
can make and therefore how often the checkpoint cycle interrupts training. **Known so far: a CPU session
was still alive at 51 min** with no cap hit, so short sessions are not the failure mode; the ceiling is
somewhere above that.
- **Closed this session:** the credential path (Β§1 last row) β€” jobs retrieve the HF token from the
account's own private Kaggle dataset via `code/ounce100m_credentials.py`; a Hub write from inside a job
was verified by anonymous readback (D-006). And the checkpoint arithmetic below.
### 8.1 Checkpoint size and cadence β€” now measured, and it is cheap
A `latest`-quality exact-resume checkpoint for a ~100M model is **~1.6 GB** (fp32 weights 400 MB + two
fp32 Adam moments 800 MB + master/grad copies + scheduler/scaler/RNG state, which are all negligible
next to the optimizer). Pushed from a Kaggle session at **42.7 MB/s β†’ 37.4 s**, and pulled back in
**13.8 s**, with byte-exact anonymous readback at 200 MB / 800 MB / 1.6 GB
(`dodosoomro/ounce100m-p1-push-bench`, D-007).
Consequences that change design choices rather than just confirming them:
- 10 checkpoints + `latest` β‰ˆ **7 min total, ~0.12 GPU-h of a 30 h week**. Checkpoint cadence is not a
budget problem, so there is no reason to economise on it β€” and rolling `latest` *more* often than the
required 10 % is nearly free, which is worth doing since every interruption costs at most one roll.
- A cold resume with an empty disk costs seconds of transfer, not minutes. The expensive part of a resume
is the `skip_first_batches` replay discussed in Β§5, which is why the input format in
`docs/02-mix-plan.md` Β§5 step 6 is designed to make it an index advance.
- Peak disk: 1.6 GB written + 1.6 GB staged elsewhere on the 19.5 GB volume is comfortable, and the
Β§3.13 prune-after-verify step keeps at most one copy resident.