# 03 — Preflight The discipline of this phase is one sentence: **spend the cheapest evidence first, and make each rung answer a different question.** Every rung below caught, in the project this guide comes from, a bug that would have destroyed a multi-hour session. ## 1. The ladder | Rung | Budget | Question it answers | |---|---|---| | 0. Static | seconds, local | Does it import? Does the launch command line parse? Do the hashes of the files I'm about to run match the files I reviewed? | | 1. Free-CPU rehearsal | minutes, no accelerator | Does the *whole control flow* work with the real modules — plan, resume decision, cursor math, checkpoint cycle — with training stubbed? | | 2. Cheap GPU smoke (20–30 steps) | **budgeted** ~10 min | Does the machinery run on the accelerator: install, device, one real optimiser step, one push, one verified resume? | | 3. Measurement probe | 1–1.5 h billed | Does the *config* hold for hours: sustained rate, flat memory, cadence cost, cold resume from the Hub on an empty disk? | | 4. The run | everything | — | Two honest corrections about rung 2, because they are the argument *for* the ladder, not against it: - It cost **1.25 h across eight attempts**, not ten minutes. Every attempt but the last failed on a real defect: `torchrun` handed the interpreter instead of the script (E-037), a stale-stop guard refusing correctly on a dirty repo (E-038), an argument-order slip in `delete_folder` (E-039), a harness that could not read its own tee'd output (E-041), and finally a rank-only assertion that killed every *successful* leg (E-044). A smoke rung's *budget* is ten minutes; its *cost* is the number of bugs in the code it is the first to execute. That is still cheaper than finding them in a 4.5-hour session — but plan the budget as "10 min, or 1.5 h if the code is new", not as "10 min". - The CPU rehearsal caught the plumbing defects (E-036, E-037). It did **not** catch the expensive one: the rank-only-`on_save` crash needed GPU legs to reproduce and took seven attempts to diagnose. Don't oversell the free rung — it buys you command lines and control flow, not distributed semantics. **Order by price-per-question, and re-run rungs 0–1 on every edit to run code.** They are free; the only cost of a rehearsal is the nine minutes you didn't spend on something else. ## 2. Make the rehearsal rehearse the real thing - Import the *real* modules — a rehearsal of a reimplementation proves nothing about the original. - Print the command line the launcher would execute and parse it in the rehearsal, so a flag typo fails there. - Point it at real artefacts where it can: hash the actually-published shards, read the actually-verified checkpoint. The Phase 4 rehearsal re-hashed all 151 mix files, which is why the run started knowing its data. - Assert the same invariants the run will assert, so a rehearsal pass means the run's guards are armed, not merely present. ## 3. Model size: prove size-dependent things at the real size "Preflight on tiny models" is the instruction, and taken literally it is wrong. **Memory, throughput, checkpoint byte-size, DDP rank behaviour and optimizer-state resume are all size-dependent** — a 15M model proves the code path and nothing about the run's numbers. This project therefore ran every Gate 3 proof at the **real 106M geometry on 2×T4**, and used CPU/tiny cells only for properties that genuinely don't scale (reader semantics, order determinism, the storage cycle's logic). Choose per proof: if the quantity you're measuring is a function of size, the size has to be right; if it isn't, spend nothing. ## 4. What preflight must demonstrate, with logs The eight proofs. Each is a line in the preflight report, with the log excerpt that shows it. 1. **Upload** a checkpoint to the Hub; **restart** and pull it back. 2. **Exact data position** on resume — recorded index/hash *and* a token-offset check, not "loss looked continuous". 3. **Optimizer and LR state** restored: print the scheduler value on the first resumed step and compare it to what the interrupted run would have produced. 4. **Loss continuity** across the restart within noise, with numbers. 5. **Measured throughput** on the target hardware, and the ETA it implies. 6. **Cold resume**: interrupt, *kill the instance*, start fresh with an empty disk, recover solely from the Hub. This is the only resume test that counts, because every real interruption looks like this. A resume that reads a local leftover file is testing something that will not exist. 7. **The storage cycle**: upload → verify by fetching bytes back → delete locally → confirm free space → keep training, at sizes representative of the real model **plus optimizer state** (`07`: ~12 bytes/param, not the weight file alone). 8. **Two sequential resumes** — position must never drift backwards or re-read consumed data. Anything unproven here is a bug you have not found yet. ## 5. Then measure the config, and let the measurement change it Preflight is the last cheap moment to *choose*. Throughput is not a spec sheet; it's a distribution on your hardware with your batch shape, and small-N numbers lie: - A 20-step average is not a rate. The 20-step figure here was **6 % optimistic** — the difference between fitting the week and not. Take the sustained number over a long leg (two legs of 120 and 60 steps measured 12,312 and 12,221 tok/s; the planning number adopted was 12,300). - Measure memory as a **time series**, not a peak. `max_memory_reserved` is process-lifetime, so a post-training validation pass "raises the peak" without any leak. Gate on the leg that does not evaluate, and compare step 10 against step 120 in the same process. - Report the **cross-rank max**, obtained by an explicit `all_reduce(MAX)`. Rank-0-only peak statistics under- report, and that is the number you will otherwise make an OOM decision on. - Cost the checkpoint cycle in seconds per push at real conditions: whole-object read-back measured 797 s off a cold CDN; bounded ranged verification runs inside a 7–10 s push→point gap on the live run. That single measurement decides how often you may push, hence how much work an interruption costs. - If the measurement shows headroom, ask what your current setting is still paying for. 5.99 GB peak on a 15.4 GB card meant gradient checkpointing was buying memory nobody needed; since it computes *identical gradients*, dropping it was plumbing, not a loss-landscape change. It bought ~30 % throughput and turned an infeasible 29.7-hour run into a 22.6-hour one. **Rule for reopening a frozen choice:** only before results exist, only when the change provably leaves gradients identical, and shrink the blast radius first (finer cadence so a failure costs less). Never after a number is on the board — that is tuning, and it invalidates the report. ## 6. Prove the failure paths, and prefer a discriminating job to an argument Rungs 0–3 are where you *inject* faults, because the alternative is discovering them mid-run: - Serve checkpoint bytes back wrong (flipped byte inside the ranged window, wrong span, a server that ignores `Range`) and confirm verification **fails**. A verifier that has never rejected anything is decoration. - Point prune at an unverified result and confirm it refuses to delete. - Kill a run mid-leg: confirm the stop guard fires for the right reason and the resume lands where it should. - Test the absence case: an empty repo must read as "nothing here", not as a timeout or a 403. - **Test rank-blind guards on every rank.** `on_save` ran on rank 0 only, so a "did you push?" assertion was unsatisfiable on rank 1 and killed every correct leg (E-044). Two rules follow: assert only predicates the rank can actually satisfy, and print diagnostics on *all* ranks — a rank-0-only logger turns a crash into silence, which is how E-044 survived seven attempts. - **When an error fits a plausible story, spend five minutes disproving the story directly** (E-017). A cheap, targeted job settled in one run what a morning of reasoning about token scopes could not. The plausible explanation is the expensive one to trust. ## 7. Sign the report Write the phase's report with each of the eight proofs (§4) and its evidence, plus measured tokens/sec, the main-run ETA, the GPU-hour estimate, and the arithmetic showing it fits the remaining quota. Named files, not conversation — a reader without this project's history must be able to check every claim. If a proof cannot state a number, it is missing, and saying so is cheaper than discovering it later.