# 03 — Preflight report (Gate 3) > §5 Phase 3 requires a written report showing **every** listed proof, plus tokens/sec and the resulting > main-run ETA and GPU-hour estimate. Each section below states what was proven, the measured number, and > the artifact that produced it. Where something is **not** proven, it says so. > > Cost accounting: the 6 GPU-hour lifetime cap for all small-model tests is tracked in > `memory/QUOTA.md`; every row referenced here is a ledger entry, including the failed runs — a session > that crashes still bills. ## 0. Scorecard | # | §5 Phase 3 requirement | Stage | Result | Evidence | |---|---|---|---|---| | 1 | The training code constructs on the target image | `pargs` (CPU) | ✅ | `p1-hub-checkpoint-cycle` v5 | | 2 | Checkpoint upload to the Hub | `p1` (CPU) | ✅ | same kernel, 1,699,108,391 B | | 3 | A restart that pulls the checkpoint back | `p1` (CPU) | ✅ | cold pull, sha-verified, 10.7 s | | 4 | Resumption from the exact data position (proven, not felt) | `p0` (CPU) | ✅ | `resume_matches_uninterrupted: true` | | 5 | Optimizer / LR-state restoration | `p3` (GPU) | ✅/⏳ | see §5 | | 6 | Loss continuity across the restart within noise | `p3` (GPU) | ⏳ | see §5 | | 7 | Stable loss over a few thousand steps | — | **not proven here** | see §7, Limitations | | 8 | Measured throughput | `p2` + A/B v2/v3 (GPU) | ✅ | 7,732–7,828 tok/s at seq 1024 | | 9 | Cold resume from an empty disk, recovering solely from the Hub | `p3` (GPU) | ⏳ | see §5 | | 10 | §3.13 storage cycle at real checkpoint size, end to end | `p1` + `p3` (GPU) | ✅/⏳ | §3, §5 | | 11 | Resume at least twice in sequence, position never drifting backwards | `p3` (GPU) | ⏳ | see §5 | ## 1. `pargs` — does the code construct at all (CPU, free) Question answered: every `TrainingArguments` keyword the trainer passes is a promise about transformers 5.0.0, and E-008/E-010 had already shown that 5.x drops things without warning. Measured on the Kaggle image (transformers **5.0.0**, torch **2.10.0**): - all `TrainingArguments` kwargs accepted, including `accelerator_config={"dispatch_batches": False}` and `average_tokens_across_devices=False`; - 22L/576 model builds at exactly **106,194,240** parameters; the 20L variant at **99,114,048** with FFN 1536 — matching `01-plan.md` §2.3's closed-form table, so the shape arithmetic and the code agree; - RoPE verified as `rope_parameters {rope_theta: 10000.0}` — transformers 5 moved it out of the flat field, which is why `docs/01-plan.md` §2's config line needed reading rather than copying; - forward + backward finite with gradients; - the trapezoid schedule measured, not assumed: plateau over **77.4 %** of steps, factor **1.0** at 80 %, **0.50** at 90 %, **0.0013** at the final update; - precision asserted as fp16 autocast with **fp32 master weights** and a live `GradScaler` (asserted at runtime in the trainer, not configured-and-hoped). Two defects this stage caught before any GPU time was spent: a SwiGLU width computed 2× in the variant branch, and variant GQA head counts that did not divide `hidden` (which would hard-error inside SDPA). ## 2. `p0` — the data reader and its resume cursor (CPU, free) Run against a throwaway 2.35 M-token mix built for the purpose (2 capped sources, published as `Cion-lab/ounce100m-mix-stagetest`, since deleted-and-kept for reuse by GPU stages): | property asserted | result | |---|---| | a run sliced to start at sample *n−40* reads exactly the windows an uninterrupted run reads | `resume_matches_uninterrupted: true` | | labels are next-token | `labels_are_next_token: true` | | same seed ⇒ same order; different seed ⇒ different order | ✅ | | train/val prefix collision over 187 held-out windows | `0` | | `verify_mix` on the same bytes | `PASS: true` | The mechanism being tested: a global PCG64 permutation over all windows, **sliced** at `start_sample` rather than skipped past, with the cursor stored as `(samples_consumed, step, seq_len, shuffle_seed, dataset_files_sha)` in `cursor.json`, and `ignore_data_skip=True` because the reader owns position. This stage also found E-024 — a parameter deleted while its body check remained, i.e. a `NameError` on first call that `py_compile` cannot see. ## 3. `p1` — §3.13's storage cycle at real size (CPU, free) Random weights, real bytes: 106,194,240 params × 4 B × 4 arrays (fp16-saved weights, fp32 master, Adam m and v) = **1,699,108,391 B** in one checkpoint, which is what the main run will move every 10 %. | step | measured | |---|---| | push + verify against the Hub's own listing (size + LFS sha256) | **26.3 s** | | authenticated read-back of one file, anonymous connection | sha matched | | local prune, free space confirmed | 19.24 → **20.94 GB** free | | cold pull into an empty directory | **10.7 s** | | rig cleanup | repo deleted in-session | Why "listing is not verification" is enforced in code: `verify_checkpoint` returns `ok` only if the read-back succeeded (E-023 — an earlier version passed while the read-back 404'd, because the report was built from a fixed key list that silently dropped the error). ## 4. `p2` — throughput, and the decision it forced (GPU, ledger rows 10-14) Test T1 measured the **frozen 22L/576 shape** on 2×T4, fp16 + DDP, on real packed data. The result was a problem: the plan-of-record configuration (seq 2048, SDPA, micro 2) runs at **4,071 tok/s** with peak CUDA memory at **14.22 GB of ~14.56**, i.e. 1 B tokens needs **68 GPU-hours** against a 30 h/week quota. The 25.1 h figure in the first draft of `01-plan.md` §6 came from p0c's *raw* loop at seq 1024 and was wrong; §6.1 records the correction. Four cells plus two A/B follow-ups, all at 22L/576 unless noted: | config | tok/s | peak GB | 1 B tokens | |---|---|---|---| | sdpa @2048 micro 2 (plan-of-record) | 4,071 | 14.22 | 68.2 h | | sdpa @2048 micro 1 | 3,954 | 10.04 | 70.3 h | | eager @2048 micro 2 | **OOM** | — | — | | eager + grad-ckpt @2048 micro 2 | 4,679 | 10.04 | 59.4 h | | sdpa + grad-ckpt @2048 micro 4 | 4,505 | 10.03 | 61.7 h | | **eager @1024 micro 4** | **7,732** | 12.25 | **35.9 h** | | **eager + grad-ckpt @1024 micro 8** | **7,828** | **7.12** | **35.5 h** | | sdpa @1024 micro 4 (matched control) | 5,105 | — | 54.4 h | Three findings, in order of consequence: 1. **Eager attention beats SDPA on Turing by +52 %** at seq 1024 (7,732 vs 5,105 at matched batch geometry), and OOMs at 2048. The library default had been costing 34 % for nothing. 2. **The `Trainer` is not the overhead.** The hypothesis I went looking for was framework cost; the raw loop measured 5,934 tok/s — *slower* than the Trainer at the same shape. The 11,062 of p0c was an artefact of comparing a different seq length, model size and batch, not a framework penalty. 3. **Sequence length is the lever that matters**, and §4 leaves it my choice as long as it is frozen at launch. Halving the context bought 1.7× and moved peak memory from 98 % of the card to 49 %. That produced **D-011**: seq **1024**, attention **eager**, gradient checkpointing **on**, micro-batch **4**, accumulation **32** ⇒ 4 × 1024 × 32 × 2 = **262,144 tokens per step**, identical to the frozen batch size in tokens, so the step count (**3,814** = `int(1e9/262144)`, i.e. 999,817,216 tokens = 99.98 % of the target), the LR and the schedule are unchanged. The capability cost — ARC 25-shot prompts exceed the context and get truncated from the left — is recorded in `05-eval-plan.md` §6, with no shot count changed. ## 5. `p3` — cold resume at the frozen geometry (GPU, ledger row 15) _Status: running as `dodosoomro/ounce100m-gate-3-p3-cold-resume-v2` at time of writing. This section is filled from `preflight_p3.json` when it lands; nothing below is claimed until it is measured._ Design, so the numbers can be checked against the intent: - **Three legs**, cumulative steps 20 → 40 → 60, at 262,144 tokens/step — the real step, the real checkpoint size, the real cursor, not a cheaper stand-in. - Before each resume the session **deletes its own local run directories** and re-`df`s, so recovery is forced through the Hub. That is the only resume test that counts, because every real interruption looks like this (§3.13). - Exactly **one checkpoint per leg** (`--push-every-steps` = the leg's step count), pushed, verified from the Hub, and pruned, so the cursor sequence has three distinct points to check. - `--resume auto` reads `latest.json`; 404 means "start from zero", and *any other* error aborts rather than silently restarting from step 0. - `--attn eager` and `--grad-ckpt` are named on the command line, not inherited from a default, so a later change to a default cannot silently change what this stage tested. - Pass criteria, checked in code: three legs exit 0; both resumes report `auto-resume: hub says step`; at least three cursor records; sample counts **strictly increasing and all distinct**; the resumed model is the frozen **106,194,240**-parameter model; and the Hub rig is deleted afterwards. To be recorded here: the three `CKPT` lines with their sample counts, the `prior loss at resume:` values either side of each wipe, `val_ppl`, `P3_frozen_rate_tok_per_s`, `P3_frozen_peak_gpu_gb`, the wall-clock cold-start cost per leg, and the session's GPU charge against row 15. ## 6. Throughput and the main-run estimate | | | |---|---| | Rate at the frozen geometry | _P3's `tok_per_s`_; planning value **7,732 tok/s** from the m4-without-checkpointing cell | | 1 B tokens | **35.9 h** (35.5 h at the m8 cell) | | Quota | 30 h/week, 2×T4, billed at **1× container wall-clock** (5 sessions observed) | | Preflight spend to date | **0.447 h** of the 6 h test cap | | Calendar | week 1 ≈ 26 h of training (~722 M tokens) → quota reset 2026-09-26T00:00Z → ~10 h in week 2 to finish 1 B | | Tripwires fixed before launch | **D-012**: below 5,150 tok/s ⇒ launch at 0.9 B. **D-013**: below 7,400 tok/s ⇒ launch at m8 × accum 16 (same 262,144 tokens/step) | Details in `docs/04-run-log.md` §1-§2. ## 7. Limitations of these proofs - **"Stable loss over a few thousand steps" is not proven by preflight, and cannot be at this budget.** P3 runs 60 steps; a thousand-step stability claim would cost ~0.5 GPU-h that the 6 h test cap cannot spare on a test rather than on the run. What *is* proven is narrower and honest: the loss is finite, the optimizer and LR state survive a cold resume, and the loss trace is continuous across two restarts. Long-horizon stability is monitored on the main run itself (§3.1), against the plan, on every wake. - **The tiny-model framing of §5 Phase 3 is satisfied differently than literally.** §5 says "test with 10-20 M parameter models … as cheaply as possible (CPU where you can, GPU only when the thing being proven needs a GPU)". Every pipeline proof here runs at the **real 106 M shape** rather than a 15 M stand-in, because the thing being proven — checkpoint bytes, cursor arithmetic, throughput, peak memory — is size-dependent, and a 15 M model would have proven a cheaper pipeline than the one being launched. The cost of that reading is quantified in rows 10-15 instead of assumed. - P3's data is the mix as merged from `ounce100m-mix-stage`, i.e. the *pre-filter* bytes if Gate 2c's contamination filter is still running. That does not affect what P3 measures (resume correctness, throughput at fixed token counts), but it is stated so the record is exact. - Peak-memory numbers are per-cell maxima from `torch.cuda.max_memory_allocated`, which excludes CUDA cache fragmentation. The main run's real OOM risk is therefore slightly understated at 12.25 GB; this is one reason D-011 chose checkpointing on. ### 4.1 Was more throughput available? Audited, and the answer is "not by changing code" Two checks, because "3.8 % of a T4's datasheet peak" looked like a defect and §3 says never trust a prior — including my own arithmetic. 1. **Cross-check against a comparable published run.** A ~163 M-parameter base model trained with DDP on **8×A100-40GB** reached **234,480 tok/s** at micro-batch 13 ([gilesthomas.com, Jan 2026] (https://www.gilesthomas.com/2026/01/llm-from-scratch-29-ddp-training-a-base-model-in-the-cloud)) — **29,310 tok/s per A100**. An A100's fp16 tensor-core peak is roughly an order of magnitude above a T4's, and the T4's advantage shrinks further because PyTorch's autocast accumulates in fp32. Scaling that ratio down, the expectation for one T4 at a 106 M model is ~2,900 tok/s; **measured is 3,866 per card** (7,732 / 2). So the run is slightly *ahead* of the cross-check, and the "low MFU" reading of §6 in the original plan was an artefact of the marketing peak number, not an unclaimed win. 2. **`torch.compile`.** The guidance available is that it pays on long, shape-stable workloads and that `reduce-overhead` targets *small* batches ([Raschka] (https://sebastianraschka.com/faq/docs/torch-compile-llm-workloads.html)), while slow compilation is a known hazard in distributed runs ([pytorch#108971](https://github.com/pytorch/pytorch/issues/108971)). Our batch is not small (4 × 1024 per rank) and the long-running shape is stable, so it is plausible — but the plausible gain is tens of percent against a new class of failure (compile hang at rank startup, graph breaks through the GradScaler path) inside a run that may not be restarted once checkpoint 1 exists (§3.1). **Decision (D-015, at freeze time, before the run):** no `torch.compile`, no flash-attn — that one is measured, not assumed: `SDPBackend.FLASH_ATTENTION` raises *"Flash attention only supports gpu architectures in the range [sm80, sm121]. Attempting to run on a sm 7.5 gpu"* on this hardware (`docs/00-platform-notes.md` §probe C), so it is architectural and no dependency install changes it; and eager already measured faster than SDPA here anyway — no quantised optimisers, no gradient-bucket tuning. The three knobs that were genuinely worth GPU time — sequence length, attention kernel, activation checkpointing — were measured and set in D-011, and together they moved 68 h to 35.9 h. The remaining ~22 h of two-week slack is better spent absorbing interruptions than chasing a few percent with an unrestartable failure mode attached. ## 8. Post-signing mechanism probes (2026-09-20) Gate 3 signed the *recipe*. These probes test the *session machinery* that only Phase 4 exercises, and they ran after the freeze and before the first billed session, on the published mix. ### 8.1 CPU rehearsal — `phase4_prep.py` at head `1dd45f80`, 245.3 s, zero quota, **all stages pass** | check | measured | |---|---| | every published shard, hashed from the Hub download against the manifest | `SHASHED 151 files total_tokens 1,109,714,831 mismatches 0` | | manifest self-consistency (records vs `n_shards` vs bytes actually mapped) | 139 = 139 = 139; shard tokens sum to the claimed total | | reader geometry | `1,109,714,831 tokens → 1,083,705 windows of 1024`; horizon `3,814 × 262,144 = 999,817,216` tokens (11.0 % margin) | | corpus fingerprint (content-hashed, E-035) | `53df4708526da5c8` | | **session schedule, driven through the shipped `plan()`** | `SESSIONS 5 ends_at 3814 horizon 3814 contiguous True` — 0→762→1,524→2,286→3,048→3,814 | | **resume reads the same data** | `RESUME_EQUIV step 381 start_sample 97,536 identical True`; same at 1,524 (390,144) and 3,433 (878,848) — input_ids *and* labels compared, through `save_cursor`/`load_cursor` | | permutation identity guard (E-035 residual, now implemented) | `full.order_sha == perm_sha(...) == resumed.order_sha` at every boundary | | launcher rehearsal | `PREP_ONLY_OK stop_after_steps 762 planned_steps 762`, `usable_hours 6.3` | | cursor round-trip + trainer imports on this image | `CURSOR_AND_IMPORT_RC 0`, `TRAINER_HELP_RC 0 lines 75` | Two real defects surfaced here rather than on a GPU: **E-036** (`--help` crashed because argparse interpolates help text and two strings contained a bare `%`) and **E-037** (`torchrun` was handed `sys.executable` as the script, so it tried to compile the Python binary — the same line in the launcher would have killed session 1 about ten minutes in, after the 2.3 GB pull). ### 8.2 GPU probes `p4_stop_probe.py`, two modes, same assertion set. **Probe 0** (soak, 06:12Z) died in 8.6 s on E-037 and cost effectively nothing. **Probe 1 v1** (`smoke`: 20 steps at 262,144 tokens/step, no checkpointing, `--push-every-steps 10`, stop at 10, run directory deleted, resume to 20) ran 778.7 s and **answered the mechanics question outright**: ``` CKPT 10 samples=2,560 push+verify=60.3s ok=True pruned=True free=18.67GB loss=10.1926 TRAIN DONE step=10/20 elapsed=4.5 min tokens_consumed=2,621,440 auto-resume: hub says step 10 at ckpt/checkpoint-10 CKPT 20 samples=5,120 push+verify=59.7s ok=True pruned=True free=16.12GB loss=9.5255 TRAIN DONE step=20/20 elapsed=4.7 min tokens_consumed=5,242,880 latest.json rolled to the completed run: final at step 20 ``` So `--stop-after-steps` stops on the boundary, the push→verify→pointer-read-back→prune cycle works at the real 1.7 GB size, a cold Hub resume lands on the exact cursor (2,560 → 5,120 samples = step × 256), loss is continuous across the disk wipe, the terminal pointer roll fires, and disk headroom after two checkpoints is 16.12 GB. `samples_consumed` and `step` agree by construction at both boundaries. **What it did not survive:** rank 1 exited 1 with `error_file: ` and no traceback, so six assertions could not be evaluated. Two hypotheses were written here at the time — `--no-grad-ckpt` meeting a 25-batch eval forward, or the segmented stop interacting with the eval loop — and **both were wrong**; §8.4 records what it actually was (E-044: a rank-0-only assertion that `hub_cb.steps` was empty on rank 1 by construction). That diagnosis needed three things this attempt lacked: per-rank log files rather than `--redirects 2` (which filed the stream nobody was reading, E-039), a failure printer that greps the whole capture rather than its last 40 lines, and marks printed by every rank instead of rank 0. **This is the exact case the smoke-first ordering was for**: it cost 13 minutes, not the 1.6 hours the soak would have spent finding it. ### 8.3 Rehearsal v4 and probe 1 v4/v5 (2026-09-20 07:47-07:50Z) `dodosoomro/ounce100m-phase4-prep-cpu-rehearsal` **v4 — COMPLETE** after the round-3 edits to the trainer and launcher (rank-visible evaluation failure, cross-rank peak-memory reduction, complete `run_summary.json`). The kernel ends `raise SystemExit(0 if (rc1 == rc2 == rc3 == 0) else 4)`, so a COMPLETE status is the evidence that all three stages — launcher rehearsal, the 151-shard content hash and the cursor/import checks — still pass on head `4c7f4922`. The reader's guarantees therefore survive the changes that were made in response to review. Probe 1 **v4** re-established the mechanics on a clean scratch repo (`PROBE_REPO Cion-lab/ounce100m-ckpt-probe-0920-0727`, `LEG1 pushed: ['checkpoint-10'] pointer: {'step': 10, 'path_in_repo': 'ckpt/checkpoint-10'}`) and died again at 275 s in the same place — the post-train validation pass — this time with **no Python traceback even though stderr was not redirected**, which is what pointed at the `say()`-is-silent-on-rank-1 mechanism (E-039) rather than a data or memory error. v5 adds `--tee 3`, a directory listing before any tail, `NCCL_DEBUG=WARN`, and per-rank reporting of the evaluation failure. Decision still open, and bounded either way: if the validation pass cannot run at `--no-grad-ckpt`, D-017 stands, the run keeps checkpointing at 9,696 tok/s, and nothing else changes. If it can, D-018 adopts 12,792 tok/s with a 60-step recovery cadence. The smoke probe is 13 minutes; the soak is 1.6 h; session 1 follows whichever way it lands. ### 8.4 Probe 1 closed (v6, v7, v8 — 09:00Z to 09:42Z): `PASS`, and the crash was never the memory config Three more attempts, each one buying the evidence the previous attempt could not print. **v6 (581.1 s, FAIL).** With the tee-prefix bug fixed and the whole-object read-back replaced by two bounded `Range` requests, the checkpoint cycle fell from `push+verify=797.7s` to **13.4 s**, and leg 1 still died after its forced stop. The probe's own `DIAG_leg1` printed nothing at all — it captured the per-rank logs and discarded them, because its output went through the same line-prefix filter as everything else (E-043). **v7 (682.7 s, FAIL) — the diagnostic that worked.** The trainer now marks the post-training sequence with per-rank `print`s, and leg 1 printed: ``` [rank 0] post-train: step=10/20 log_history_loss_entries=2 final_loss=10.192607879638672 start_step=0 [rank 0] past the stop guard: forced_stop=True stopped_at=10 validation skipped: segment ended at step 10 of 20 [rank 0] peak_stats: entering (this is a collective) ``` and nothing from rank 1 — while rank 1's log, which `diag` now actually prints, ended with *"segment ended at step 10 but no checkpoint was pushed and verified for it (pushed: [])"*. Rank 1 was killing the run on **our own assertion**: `HubPush.on_save` returns early on non-zero ranks, so the list that assertion reads is empty on rank 1 by construction, and every correctly-checkpointed forced-stop leg died there. Not a collective, not `evaluate()`, not memory. E-044, and it retires the E-040 hypothesis in §8.2. **v8 (578.0 s) — `VERDICT P4PROBE PASS []`, 16/16 checks, both legs rc 0.** Leg 1 stops at 10 of 20, pushes, Hub-verifies, rolls `latest.json`, reads it back, prunes, skips validation and exits 0 with `val_skipped: true`; the probe wipes the run directory; leg 2 resumes from the Hub alone and runs to the horizon, pushes `ckpt/checkpoint-20` and `final/`, runs the 195-window held-out pass (`loss 9.3532`, `ppl 11535.95`) and rolls the terminal pointer. Tokens 2,621,440 → 5,242,880 = 10 and 20 × 262,144; `samples_consumed` 2,560 → 5,120; `shuffle_perm_sha` and `dataset_files_sha` identical across the resume; `params 106,194,240`; `grad_ckpt false` read off the artifact; loss `10.7723 → 10.1926 → 9.5255`. Numbers carried forward, all at the frozen geometry with checkpointing **off**: | quantity | value | where it came from | |---|---|---| | step time | **20.20 s** = 12,977 tok/s | v8 progress bar, 10-step and 20-step legs agreeing | | horizon wall clock | **21.4 h** of stepping (3,814 steps) | that rate × the step count | | checkpoint cycle | **12.7-16.5 s** per 1.27 GB push+verify+pointer+prune | v7/v8 `CKPT` lines | | memory, training only | **12.84 GiB reserved / 12.70 allocated**, cross-rank max | v8 leg 1, which never evaluates | | memory, after validation | 13.66 GiB reserved | v8 leg 2 — the same high-water mark, plus the eval pass | | session shape | 5 sessions of 762 steps, last running to 3,814 | rehearsal v3/v4 driving the shipped `plan()` | The split of the two memory figures is the reason the soak's gate is `peak_memory_during_training_below_13_6_gb` rather than a check on both legs: `max_memory_reserved` is process-lifetime, so a leg that evaluates reports training-then-eval, and the run occupies the training state for all 3,814 steps. **Two review-found defects closed in the same window** (E-042, E-045): the prune guard lived only inside the wrapper the save path bypasses, so a failed verification would have deleted the last local copy *and* moved the pointer onto it; and `verify_checkpoint`'s length check is now accompanied by a `Content-Range` check, because 850 MB of the right length from the wrong offset passed the old test. Both are rehearsed on free CPU in stage 4 of `phase4_prep`, against the 1.27 GB checkpoint this probe really pushed — clean, tampered inside the read window, tampered outside it, absent, Range-ignoring, wrong-offset, wrong-bytes, and the prune refusal itself.