|
Download docs/01-plan.md from Cion-lab/ounce100m-code: direct link, hf CLI and curl.
- Browser
- Download file 26.9 kB
-
https://huggingface.co/Cion-lab/ounce100m-code/resolve/main/docs/01-plan.md
- Command line
-
hf download hf://Cion-lab/ounce100m-code/docs/01-plan.md
-
curl -L -o 01-plan.md https://huggingface.co/Cion-lab/ounce100m-code/resolve/main/docs/01-plan.md
26.9 kB
| # 01 β Plan: architecture, hyperparameters, throughput, targets | |
| Phase 1 deliverable. Every choice below cites a source actually inspected on 2026-09-19, as Gate 1 | |
| requires. Items marked **β test** are decisions whose *numbers* Phase 3 must validate on the real | |
| pipeline before the run freezes; the design itself is settled. | |
| --- | |
| ## 1. The envelope the design has to fit | |
| Measured in Phase 0 (`docs/00-platform-notes.md`), not assumed: | |
| | Constraint | Value | Consequence | | |
| |---|---|---| | |
| | Accelerator | 2x Tesla T4, cc 7.5, **14.56 GiB usable each** | fp16 only; no bf16 tensor cores; no flash-attention (sm80+ only) | | |
| | GPU quota | 108,000 s/week, billed at **~1x container wall-clock** (3 samples) | β30 wall-clock h/week, both cards | | |
| | Test budget | β€6 GPU-h lifetime | 0.073 h spent; **5.93 h remains for all of Phase 3** | | |
| | Working disk in a job | **19.5 GB** `/kaggle/working` | shards consumed a few at a time, never whole (Β§3.13) | | |
| | RAM / CPU | 30 GiB cgroup, 4 vCPU, **no swap** | input pipeline memory-bounded by design | | |
| | Image | torch 2.10.0+cu128, **transformers 5.0.0**, datasets 5.0.0, accelerate 1.13.0, triton 3.6.0; **no** trl/deepspeed/flash-attn/xformers/bnb | prefer the preinstalled stack; each install is risk + round trip | | |
| | Hub from a job | reads anonymous & fine, 35β89 MB/s; writes **401** until given a token | credential path: private Kaggle dataset mount (Β§8) | | |
| **Parameter convention.** Β§2's 90β110M is *including embeddings* β the Pythia convention (Pythia was | |
| explicitly renamed to include embedding + unembedding, [pythia README](https://raw.githubusercontent.com/pythia/main/README.md)). | |
| This matters more at 100M than anywhere else: the tied embedding is 21β37 % of the model across the | |
| candidates below. `code/config/param_count.py` computes it two ways and both were checked to agree | |
| exactly against transformers 5.0.0 in `dodosoomro/ounce100m-p1-param-count`. | |
| ## 2. Architecture | |
| ### 2.1 The shape: deep-thin, tied, GQA, SwiGLU, pre-norm | |
| **Choice: `hidden 576, layers 22, 9 Q heads / 3 KV heads (GQA 3:1), SwiGLU intermediate 1536, RMSNorm | |
| pre-norm eps 1e-5, attention_bias false, RoPE theta 10,000, tied embeddings, vocab 49,152, seq 2048`.** | |
| Evidence, and it is convergent rather than a single citation: | |
| - **Deep-thin beats wide at this size.** Meta's MobileLLM ablated exactly this at a fixed 135M: 30 | |
| layers Γ d512 scored **44.8 avg vs 43.9 for 12 Γ d768**, with embedding-sharing removing 13 % of | |
| parameters at ~equal accuracy ([arXiv 2402.14905](https://arxiv.org/abs/2402.14905), HIGH). Its | |
| shipped 125M model is **30 layers, d576, 9Q/3KV GQA, SwiGLU, tied**, 124.6M total. | |
| - **The same shape is what HF's SmolLM2-135M uses**: 30 layers, d576, 9Q/**3KV**, `intermediate_size` | |
| 1536 (2.67Γd), silu, `attention_bias=false`, **`tie_word_embeddings=true`**, RMSNorm 1e-5, | |
| [model config](https://huggingface.co/HuggingFaceTB/SmolLM2-135M/blob/main/config.json) + | |
| [arXiv 2502.02737 Β§6](https://arxiv.org/abs/2502.02737). Two independent labs landing on | |
| d576 + GQA 3:1 + SwiGLU at ~130M is the strongest available signal. | |
| - Our budget is 110M, not 135M, so the same d576/GQA3:1/SwiGLU-1536 stack is cut to **22 layers** | |
| rather than 30 β verified count below β and RoPE ΞΈ drops to 10,000 to match the shorter context | |
| (SmolLM **v1** used ΞΈ=10,000 at ctx 2048; SmolLM2's 100,000 accompanies ctx 8192, which we cannot | |
| afford and do not need for these benchmarks). | |
| - **GQA at 3:1 is a parameter choice as much as an attention choice**, since k/v are `d_head Γ n_kv` | |
| wide; MobileLLM reports GQA-plus-enlarged-d as +0.4 avg. | |
| **Not chosen: Mamba/SSM or hybrid.** The official kernels *do* build for `sm_75` | |
| ([state-spaces/mamba setup.py](https://github.com/state-spaces/mamba)), which corrects the common | |
| assumption that they are Ampere-only β but they are absent from the image (source install only), and | |
| the research found **no published 50β200M SSM/hybrid with released hyperparameters** to imitate. That is | |
| research risk, not a default, and Β§4 forbids the kind of custom code where autonomous projects die. | |
| ### 2.2 Verified parameter counts | |
| Closed form and `sum(p.numel())` agreed exactly on all eight shapes, in the target environment. | |
| Tying confirmed genuinely tied, not merely declared. | |
| | candidate | H | L | heads | kv | FFN | vocab | **total params** | emb share | in 90β110M | | |
| |---|---|---|---|---|---|---|---|---|---| | |
| | Phase-0 probe shape | 768 | 12 | 12 | 3 | 2048 | 50257 | 112,934,400 | 34.2 % | β | | |
| | A | 768 | 11 | 12 | 3 | 2048 | 50257 | 106,739,712 | 36.2 % | β | | |
| | B | 768 | 12 | 12 | 3 | 1792 | 50257 | 105,856,512 | 36.5 % | β | | |
| | C | 768 | 12 | 12 | 2 | 1824 | 50257 | 105,561,600 | 36.6 % | β | | |
| | D | 640 | 16 | 10 | 4 | 2048 | 50257 | 113,450,240 | 28.4 % | β | | |
| | E | 768 | 12 | 12 | 3 | 2048 | 32768 | 99,502,848 | 25.3 % | β | | |
| | F | 896 | 9 | 14 | 2 | 2432 | 50257 | 120,397,312 | 37.4 % | β | | |
| | G | 1024 | 8 | 8 | 2 | 2816 | 50257 | 141,658,112 | 36.3 % | β | | |
| **The binding insight: the tokenizer decides how much model you are allowed to have.** At 50,257 vocab | |
| and d768 the tied embedding is 38.6M β a third of the entire budget β which is why every 50k candidate | |
| clusters at 105β113M and must trim depth or FFN. MobileLLM's own ablation reached the same conclusion | |
| (embedding sharing: β13 % parameters at ~equal accuracy). Choosing a smaller `hidden` is therefore not | |
| a weakening, it is how the embedding tax gets paid for. | |
| **β test:** the frozen shape's exact count was re-derived from the constructed model β see Β§2.3. | |
| ### 2.3 The frozen shape, counted to the unit | |
| `dodosoomro/ounce100m-p1-param-count` v2 built each candidate as a real `LlamaForCausalLM` under | |
| transformers 5.0.0. Closed form and `sum(p.numel())` agree exactly on every row. | |
| | candidate | H | L | Q/KV | FFN | vocab | **total params** | emb share | in 90β110M | | |
| |---|---|---|---|---|---|---|---|---| | |
| | H | 576 | 20 | 9/3 | 1536 | 49,152 | 99,114,048 | 28.6 % | β | | |
| | **I β CHOSEN** | **576** | **22** | **9/3** | **1536** | **49,152** | **106,194,240** | **26.7 %** | β | | |
| | J | 576 | 24 | 9/3 | 1536 | 49,152 | 113,274,432 | 25.0 % | β over | | |
| | K | 576 | 26 | 9/3 | 1536 | 49,152 | 120,354,624 | 23.5 % | β over | | |
| | L | 640 | 20 | 10/4 | 1728 | 49,152 | 120,776,320 | 26.0 % | β over | | |
| **Frozen: 106,194,240 parameters (embeddings included, tied), counted as | |
| `vocabΓh` once + `L Γ (attn + SwiGLU)` + norms + final norm, and confirmed by constructing the model.** | |
| The counting method is stated here because Β§2 requires the number *and* how it was counted. | |
| Two consequences worth recording: | |
| - **Depth is capped at 22 by the budget, not by taste.** Each layer at d576 costs 3,538,944 parameters, | |
| so 24 layers is 113.3M β over. That is a hard stop, and it means the Phase-3 "deep-thin vs wide" | |
| throughput test (T1) must compare **H (20 layers, 99.1M) against I (22 layers, 106.2M)**, both | |
| in-budget, rather than the 12-layer d768 probe shape which cannot fit the SmolLM2 vocab without | |
| trimming elsewhere. | |
| - **Widening is not a free alternative to deepening.** Candidate L at d640 Γ 20 layers is *over* budget | |
| (120.8M) despite having fewer layers than I, because attention and MLP cost scale with `hΒ²`/`hΒ·ff`. | |
| So the deep-thin choice is forced from two directions: it is better per parameter (MobileLLM's | |
| ablation) and it is the only way to spend the 78M non-embedding budget on 22 layers. | |
| ## 3. Tokenizer | |
| **Choice: reuse the SmolLM2 BPE tokenizer** β vocab 49,152, byte-level, **Apache-2.0**, from | |
| [`HuggingFaceTB/SmolLM2-135M`](https://huggingface.co/HuggingFaceTB/SmolLM2-135M)'s `tokenizer.json`; | |
| trained on SmolCorpus per [arXiv 2502.02737](https://arxiv.org/abs/2502.02737). Not trained by us. | |
| Β§4 says maximum reuse, minimal custom code, and training a tokenizer from scratch is precisely the | |
| kind of optional machinery that adds failure surface without adding capability. Candidates checked today: | |
| SmolLM2 **49,152 / Apache-2.0**; `gpt2` 50,257 / MIT; `pythia-160m` 50,304 / Apache-2.0; | |
| `t5-v1_1-small` 32,128 Unigram (not byte-level, carries `extra_ids` sentinels); `Qwen3-0.6B` 151,936 | |
| (disqualifying β its embedding alone would be ~78M at d512); `TinyLlama` 32,000 but Llama-2-derived and | |
| **redistribution status unverified** β excluded. | |
| SmolLM2 wins on three counts: licence is clean and permissive, it is byte-level and English-curated at | |
| exactly this model scale, and sharing it with a family of released 135M models gives the benchmark | |
| numbers somebody else's token statistics to be compared against. Its 49,152 Γ 576 tied embedding is | |
| 28.3M = **21 %** of budget, versus 34β37 % for the 50k/d768 shapes. | |
| ## 4. Optimizer, batch, schedule, precision | |
| | | choice | justification | | |
| |---|---|---| | |
| | Optimizer | AdamW, Ξ²=(0.9, 0.95), wd 0.1, clip 1.0 | Pythia-160M and SmolLM2 both use Ξ²β=0.95 at small scale ([2304.01373](https://arxiv.org/abs/2304.01373), [2502.02737](https://arxiv.org/abs/2502.02737)) | | |
| | Peak LR | **6e-4** | bracketed from three directions: Cerebras-111M 6.0e-4, Pythia-160M 6e-4, and the 150M low-tokens/param run of Sardana et al. at 4.6e-4 ([2304.03208](https://arxiv.org/abs/2304.03208), [2401.00448](https://arxiv.org/abs/2401.00448)). The 2e-3/3e-3 figures that appear in small-model recipes belong to 1β2 T-token runs, which this is not. | | |
| | Warmup | 2 % of steps, floor 100 steps | Pythia used 1 %; SmolLM2 used 2,000 steps at 2M tok/batch (~0.2 %). Warmer than 1 % because **fp16 without bf16 needs the ramp** β Phase 0 showed an unramped/unscaled fp16 run going straight to NaN (E-006). | | |
| | Schedule | **trapezoid / WSD: constant, then linear decay to 0 over the final 20 %** β *not* cosine-to-10 % | SmolLM1 used a trapezoid with 20 % cooldown; SmolLM2 used WSD ([blog/smollm](https://huggingface.co/blog/smollm), [2502.02737 App. A](https://arxiv.org/abs/2502.02737)). *Straight to Zero* finds linear-decay-to-zero beats cosine-to-10 % when peak LR is optimal, across sizes/batch/data ([2502.15938](https://arxiv.org/abs/2502.15938)); COLT 2026 theory says decay *shape* barely matters but **overly slow terminal decay causes schedule-induced capacity saturation** ([2602.06797](https://arxiv.org/abs/2602.06797)). Its caveat applies: their evidence is at/above compute-optimal tokens/param, we are at half. | | |
| | Global batch | **β262 k tokens/step** (~3,815 steps for 1 B tokens) | Critical batch size fits `B* = 621.341Β·N^0.087`, i.e. **nearly independent of model size** ([2410.21676](https://arxiv.org/abs/2410.21676)) β no reason to chase 2β4M-token batches. Sardana et al. deliberately used *smaller* batches for smaller models "so that low-token-count training runs see enough steps". At 10 tokens/param we are such a run. | | |
| | Micro-batch | **2 Γ 2048 per GPU Γ 2 GPUs Γ 32 accumulation** | Set by memory, not taste: Phase 0 measured bs4Γ1024 at 11.5β13.8 GB of 14.56 GB and **bs8 OOMed outright**. β test | | |
| | Precision | **fp16 autocast + fp32 master weights + `GradScaler` + grad clip** | Only mixed-precision option on Turing; transformers v5 documents this path by name ("fall back to fp16 on older hardware like V100 or T4") and warns **load the model in fp32 or autocast is a no-op** ([v5 mixed_precision docs](https://huggingface.co/docs/transformers/perf_training)). Verified finite in Phase 0 probe C; bf16 *runs* but is emulated β must be asserted off, not defaulted. | | |
| | Seed / RNG | single fixed seed, recorded | Β§3.1 requires "same RNG semantics" across resumes | | |
| **Undertrained by design, and honestly so.** 1 B tokens on ~100M β **10 tokens/param**, about half | |
| Chinchilla-optimal (Cerebras-111M's compute-optimal point was 2.2 B tokens = 20 t/p). That is mildly | |
| undertrained, not pathological β the genuine anomaly is 2025-26 practice, which overtrains tiny models | |
| by three orders of magnitude (SmolLM2-135M at 2 T tokens β 14,800 t/p). Sardana et al. trained a | |
| 150M/d768/12L model across 3β10,000 t/p and quality kept improving; Gadre et al. show scaling laws | |
| extrapolate across over-training ([2403.08540](https://arxiv.org/abs/2403.08540)). We are on the | |
| left-hand side of that curve because the quota puts us there, and the report should say so. | |
| ## 5. Training stack | |
| **Choice: `torchrun --nproc_per_node=2` + Hugging Face `Trainer` / `TrainingArguments`, `fp16=True`, | |
| on a pre-tokenised **map-style** `Dataset`.** Minimal glue, no framework authored here (Β§4). | |
| - Everything needed is already in the image: transformers 5.0.0, accelerate 1.13.0, datasets 5.0.0. | |
| - v5 `Trainer` restores **RNG (`_save_rng_state`/`_load_rng_load`), optimizer, LR scheduler and the fp16 | |
| GradScaler** β the exact set Β§3.1 demands β via `Accelerator.save_state`. Verified against | |
| v5.17.0 source; the image's 5.0.0 is ~8 patch releases behind but the same major. **β test at Gate 3.** | |
| - **The resume trap, stated plainly:** v5 recovers data position with `skip_first_batches`, which | |
| **replays** discarded batches β O(steps) wasted work on an `IterableDataset`, up to 3,815 steps of | |
| pure I/O after every interruption. Mitigation is architectural, not a flag: a **map-style, concatenated | |
| token store** makes skipping an index advance rather than a decode. A `dataset_shard` layout with an | |
| explicit recorded cursor is what Gate 3 must prove, and Β§Phase 2's "exact positional resumption" | |
| requirement is the reason it is designed that way up front. | |
| - **v5 migration traps to check in preflight** ([MIGRATION_GUIDE_V5.md](https://github.com/huggingface/transformers/blob/main/MIGRATION_GUIDE_V5.md), HF *Transformers v5* blog): PyTorch-only backend; | |
| slow tokenizers gone (irrelevant, we use `tokenizers`); several `TrainingArguments` removed **without | |
| a deprecation cycle**; attention moved to `AttentionInterface`; `remove_unused_columns` **still | |
| defaults True** (silently drops dataset columns a loss function needs); `lr_scheduler_type` default is | |
| still `"linear"`, *not* cosine β so an unstated schedule is a linear decay to zero over the whole run; | |
| new `train_sampling_strategy` replaces the `group_by_length` bool; | |
| `restore_callback_states_from_checkpoint` is required to restore scheduler state properly. | |
| - **Disqualified:** **torchtitan** β `training_dtype` is typed `Literal["bfloat16","float32"]`, i.e. | |
| **no fp16 path at all**, plus Hopper-oriented CI β unusable on T4 ([torchtitan config](https://github.com/pytorch/torchtitan)). | |
| **unsloth** β fine-tuning/GRPO only, no from-scratch pretraining path. **nanoGPT** β dormant since | |
| 2025-11, no HF export, no distributed input pipeline. **FSDP** β ~100M params in fp32 master + | |
| grads + AdamW is ~1.8 GB; DDP on 2Γ16 GB does not need sharding, and sharding would add failure | |
| surface for nothing. **DeepSpeed / flash-attn / xformers / trl** β absent from the image; installing | |
| them is risk without a capability we lack. | |
| ## 6. Throughput, and the token target it implies | |
| Measured on a main-run-shaped model, `seq 1024`, bs4/card, DDP over both T4s: **11,062 tok/s | |
| aggregate** (5,531/rank; single-GPU 6,478 SDPA / 7,239 eager; **85 % DDP efficiency**). | |
| Sanity-checking that against compute: `6ND β 6 Γ 1.13e8 Γ 11,062 β 7.5e12 FLOP/s`, which is **~6 % of a | |
| 2ΓT4 fp16 tensor peak** (65 TFLOPS/card β an *unverified* datasheet prior; the spec page could not be | |
| retrieved). Independent corroboration: a documented 163M / ctx-1024 run reached ~19.9k tok/s on a single | |
| RTX 3090 ([gilesthomas.com](https://gilesthomas.com/2025/12/llm-from-scratch-28-training-a-base-model-from-scratch)), | |
| and Pythia-160M's 1,030 A100-hours for 300 B tokens implies ~81k tok/s per A100. So ~11k tok/s on two | |
| Turing cards is the right order of magnitude, and **the low implied MFU means this is a floor with headroom, | |
| not a ceiling** β which is the opposite of the risk I most wanted to rule out. | |
| **The tension to resolve by measurement, not argument:** deep-thin is better per *parameter* | |
| (MobileLLM's ablation) but worse per *second* on 2018 hardware, because 22 small layers issue more | |
| sequential kernel launches than 12 fat ones, and we are wall-clock-bound rather than | |
| parameter-bound. Phase 3 therefore measures both shapes. | |
| | token target | at 11.1k tok/s (measured) | at 9k tok/s (planning floor) | | |
| |---|---|---| | |
| | 1.10 B | 27.6 h | 34.0 h | | |
| | **1.00 B** | **25.1 h** | **30.9 h** | | |
| | 0.90 B | 22.6 h | 27.8 h | | |
| **Decision: target 1.0 B tokens, with a pre-registered fallback to 0.9 B if Gate 3's end-to-end rate on | |
| the frozen config and real input pipeline is below ~10.5k tok/s.** Both are inside Β§2's band, and the | |
| fallback is chosen *now* rather than later so it cannot become a post-hoc excuse. Even the optimistic | |
| column leaves under 5 h of slack in a 30 h week, which the interruption budget in Β§3.1 will eat; so the | |
| plan assumes **a two-week run crossing the 2026-09-26 quota reset**, with a checkpoint at every 10 % and | |
| a rolling `latest`, exactly as Β§2 requires. | |
| ## 6.1 Amendment after Gate 3's measurements β the fallback trigger fired, and here is what was done about it | |
| **Read Β§6 above first: it is kept verbatim, including its numbers, which are now known to be wrong.** The | |
| 11,062 tok/s figure came from p0c's *raw* training loop at seq 1024 on a 113M model β no `Trainer`, no | |
| real input pipeline, no fp16-master/GradScaler bookkeeping, no seq-2048 activations. Gate 3 measured the | |
| frozen 106,194,240-parameter shape inside the actual stack on 2ΓT4 (ledger rows 10-14, `memory/QUOTA.md`): | |
| | config | tok/s | peak GB | 1.0 B tokens | | |
| |---|---|---|---| | |
| | sdpa @2048 micro 2 (the Β§6 plan-of-record) | 4,071 | 14.22 | **68.2 h** | | |
| | eager + grad-ckpt @2048 micro 2 | 4,679 | 10.04 | 59.4 h | | |
| | eager @1024 micro 4 | 7,732 | 12.25 | 35.9 h | | |
| | eager + grad-ckpt @1024 micro 8 | 7,828 | 7.12 | 35.5 h | | |
| | eager @2048 micro 2 | OOM | β | β | | |
| Three consequences, in the order that matters: | |
| 1. **Sequence length changed to 1024** β that is **D-011**, decided on these numbers and documented in | |
| `memory/DECISIONS.md`. Micro-batch 4 Γ accum 32 Γ 2 cards keeps tokens/step at **262,144**, so the | |
| batch size in tokens, the ~3,815 steps, the LR and the trapezoid are all unchanged. Β§2.3's parameter | |
| table is unaffected β it never depended on the measurement. | |
| 2. **The pre-registered fallback condition in Β§6 fired.** "Below ~10.5k tok/s" is satisfied: every | |
| measured seq-1024 cell is ~7.7-7.8k. Mechanically that means 0.9 B. | |
| 3. **The premise of that trigger is gone.** 10.5k was chosen so 1.0 B would fit *one* 30 h week with | |
| enough slack for interruption overhead. At the measured rate neither 1.0 B (35.9 h) nor 0.9 B (32.3 h) | |
| fits week 1, so the trigger no longer discriminates between the two β it only buys 3.6 GPU-hours for a | |
| 10 % cut in training tokens. Against the two-week schedule in `docs/04-run-log.md` Β§2 (60 h of quota, | |
| 35.9 h of training, **~22 h of slack**) the 1.0 B target is affordable. | |
| **Therefore, re-registered before launch and before any checkpoint or score exists (D-012): the token | |
| target stays 1.0 B, and the fallback trigger is restated against measured throughput rather than against | |
| p0c's number.** New mechanical rule, same purpose as the old one: **if preflight P3's end-to-end rate at | |
| the frozen geometry (22L/576, seq 1024, eager, grad-ckpt on, micro 4, accum 32) is below 5,150 tok/s β | |
| i.e. 1 B tokens needs more than ~54 h of the ~60 h available across the two weeks after allowing 6 h for | |
| interruption overhead and Phase 6's GPU evaluation β the target drops to 0.9 B at launch, not later.** | |
| 5,150 tok/s is 33 % below the planning rate, so this is a genuine tripwire, not a formality. What did | |
| *not* change: D-005's benchmark bands, the harness pin (D-010), the mix, the model shape, the optimiser, | |
| the schedule, or the one-run rule. If the 0.9 B fallback fires, it fires as a documented arithmetic | |
| consequence, and this section is where a reader can check it. | |
| ## 7. Benchmark targets β pre-registered before the main run (Β§4) | |
| Recorded now, before any training, so the goalposts cannot move. Full anchor table, protocol and noise | |
| floor in Β§7.1β7.3. Sources: archived Open LLM Leaderboard run metadata and Pythia Appendix G | |
| ([arXiv 2304.01373](https://arxiv.org/pdf/2304.01373)), read 2026-09-19. | |
| **The dominant fact: every published β₯100M base anchor was trained on ~300 B tokens.** Our budget is | |
| ~1 B β a **300Γ gap**. Those anchors are therefore ceilings, not expectations. No harness score for any | |
| 50β350M base trained at ~1β3 B tokens could be found published anywhere, so these bands are | |
| extrapolations from the low end of the size curve and are labelled as such. | |
| | Task | metric | band | chance | | |
| |---|---|---|---| | |
| | ARC-Easy | `acc` | **26β34** | 25 | | |
| | ARC-Challenge | `acc` / `acc_norm` | **17β22** / **20β25** | 25 | | |
| | HellaSwag | `acc_norm` (report `acc` too) | **26β32** (acc 25β30) | 25 | | |
| | PIQA | `acc` | **52β60** | 50 | | |
| | WinoGrande | `acc` | **49β53** | 50 | | |
| | MMLU | `acc` | **24β27** | 25 | | |
| | TruthfulQA | `mc2` / `mc1` | **36β46** / **21β26** | β38 / β22 | | |
| | GSM8K | `exact_match,strict-match` | **0.0β1.5** | β0 | | |
| - **MMLU, WinoGrande, GSM8K and ARC-Challenge are expected to be indistinguishable from chance.** MMLU | |
| measures flat 24.9β27.3 from 70M to 1.4B β OPT-1.3B scores *lower* than OPT-125M β so it has almost | |
| no discriminative power at this scale. GSM8K anchors: 0.23 (OPT-125M), 0.68 (GPT-2, Pythia-410M), | |
| 1.52 (Pythia-1.4B), 1.4 even for SmolLM2-135M at 11.2 T tokens. Multi-step arithmetic does not emerge | |
| from a 100M base at 1 B tokens; predicting otherwise would be the inflation Β§3.12 forbids. | |
| - **The real signals are ARC-Easy, HellaSwag and PIQA** β the three where anchors clear chance by a | |
| visible margin. Failing to beat chance there is a finding about the mix or the run, not the benchmarks. | |
| - **Noise floor:** SE β 1.3 pp ARC-C (n=1172), 1.4 pp WinoGrande (1267), 1.2 pp PIQA (1838), 1.5 pp | |
| TruthfulQA (817), 0.45 pp HellaSwag (10042). **Differences under ~2 pp on ARC/WinoGrande/TruthfulQA | |
| are not results**, and no shot count or prompt format gets chosen per-task after seeing scores (Β§3.3). | |
| ### 7.1 Protocol, pinned for Phase 6 | |
| Harness **EleutherAI `lm-evaluation-harness` v0.4.13** (2026-08-31) β still what published anchors | |
| report against; Lighteval v0.13.0 is the active HF alternative but switching would break comparability. | |
| Record the git SHA per run; cadence is real (v0.4.10 2026-01 β v0.4.13 2026-08). | |
| | Task | dataset / config | split | shots | | |
| |---|---|---|---| | |
| | `arc_easy`, `arc_challenge` | `allenai/ai2_arc` | **test** (2376 / 1172) | **25** | | |
| | `hellaswag` | `Rowan/hellaswag` | **validation** (10042; no test split) | **10** | | |
| | `piqa` | `baber/piqa` default | **validation** (1838) | **10** | | |
| | `winogrande` | `allenai/winogrande`, **`winogrande_xl`** | **validation** (1267) | **5** | | |
| | `mmlu` | group of 57 `mmlu_<subject>` | test | **5** | | |
| | `truthfulqa_mc1`, `_mc2` | `sylinrl/TruthfulQA` | validation | **0** (pinned) | | |
| | `gsm8k` | `openai/gsm8k` `main` | test | **5** | | |
| ### 7.2 Traps that would silently falsify the numbers | |
| 1. **Shots are NOT pinned in the ARC/HellaSwag/PIQA/WinoGrande/MMLU YAMLs** β an omitted | |
| `--num_fewshot` silently yields 0-shot. Pass shots explicitly for every task. | |
| 2. **No chat templates.** `--apply_chat_template` / `--fewshot_as_multiturn` are for fine-tuned models | |
| (Β§3.9 makes this a base model). `gsm8k_cot_llama` requires them β not used. | |
| 3. Everything but GSM8K is `output_type: multiple_choice` = loglikelihood ranking; the model never | |
| generates, so only GSM8K is sensitive to generation settings. | |
| 4. **MMLU aggregation:** the `mmlu` group uses `weight_by_size: True`, while archived leaderboard numbers | |
| were an *unweighted* subject mean. Compute and state which was used. | |
| 5. **TruthfulQA mc2 changed definition** 2024-03-11 (PR #2768); pre-April-2024 mc2 is not comparable. | |
| mc1/mc2 differ by ~20 pp and both report under the metric name `acc`. | |
| 6. **Harness PIQA is mean per-item accuracy, not AI2's `p_win`** β the harness never computes `p_win`. | |
| 7. **Winogrande**: `winogrande_xl`, single fold, validation; the official leaderboard averages 5 folds, | |
| and `winogrande_debiased` is a different number. | |
| 8. **GSM8K**: `temperature 0`, `do_sample false`, `max_gen_toks` inherits the global 256; tiny bases loop | |
| repetitively and never emit `#### `, so strict-match β0 while flexible-extract looks inflated. Report | |
| both filters or neither; do **not** add `repetition_penalty`, which silently deviates from everyone. | |
| 9. **Splits are mixed** β PIQA/HellaSwag/WinoGrande/TruthfulQA run on *validation*. Never call them test. | |
| 10. **Parameter-count conventions differ between suites**; Pythia's includes embeddings. Quote ours with | |
| the convention attached. | |
| ## 8. Still open before Gate 1 closes | |
| - ~~Deep-thin-vs-wide throughput result~~ β scheduled as Phase 3 preflight test T1. | |
| - Session wall-clock cap, from the still-running `p0e-session-cap`: it sets how much progress one session | |
| can make and therefore how often the checkpoint cycle interrupts training. **Known so far: a CPU session | |
| was still alive at 51 min** with no cap hit, so short sessions are not the failure mode; the ceiling is | |
| somewhere above that. | |
| - **Closed this session:** the credential path (Β§1 last row) β jobs retrieve the HF token from the | |
| account's own private Kaggle dataset via `code/ounce100m_credentials.py`; a Hub write from inside a job | |
| was verified by anonymous readback (D-006). And the checkpoint arithmetic below. | |
| ### 8.1 Checkpoint size and cadence β now measured, and it is cheap | |
| A `latest`-quality exact-resume checkpoint for a ~100M model is **~1.6 GB** (fp32 weights 400 MB + two | |
| fp32 Adam moments 800 MB + master/grad copies + scheduler/scaler/RNG state, which are all negligible | |
| next to the optimizer). Pushed from a Kaggle session at **42.7 MB/s β 37.4 s**, and pulled back in | |
| **13.8 s**, with byte-exact anonymous readback at 200 MB / 800 MB / 1.6 GB | |
| (`dodosoomro/ounce100m-p1-push-bench`, D-007). | |
| Consequences that change design choices rather than just confirming them: | |
| - 10 checkpoints + `latest` β **7 min total, ~0.12 GPU-h of a 30 h week**. Checkpoint cadence is not a | |
| budget problem, so there is no reason to economise on it β and rolling `latest` *more* often than the | |
| required 10 % is nearly free, which is worth doing since every interruption costs at most one roll. | |
| - A cold resume with an empty disk costs seconds of transfer, not minutes. The expensive part of a resume | |
| is the `skip_first_batches` replay discussed in Β§5, which is why the input format in | |
| `docs/02-mix-plan.md` Β§5 step 6 is designed to make it an index advance. | |
| - Peak disk: 1.6 GB written + 1.6 GB staged elsewhere on the 19.5 GB volume is comfortable, and the | |
| Β§3.13 prune-after-verify step keeps at most one copy resident. | |