Download guide/03-preflight.md from Cion-lab/ounce100m-code: direct link, hf CLI and curl.
- Browser
- Download file 8.7 kB
-
https://huggingface.co/Cion-lab/ounce100m-code/resolve/main/guide/03-preflight.md
- Command line
-
hf download hf://Cion-lab/ounce100m-code/guide/03-preflight.md
-
curl -L -o 03-preflight.md https://huggingface.co/Cion-lab/ounce100m-code/resolve/main/guide/03-preflight.md
03 β Preflight
The discipline of this phase is one sentence: spend the cheapest evidence first, and make each rung answer a different question. Every rung below caught, in the project this guide comes from, a bug that would have destroyed a multi-hour session.
1. The ladder
| Rung | Budget | Question it answers |
|---|---|---|
| 0. Static | seconds, local | Does it import? Does the launch command line parse? Do the hashes of the files I'm about to run match the files I reviewed? |
| 1. Free-CPU rehearsal | minutes, no accelerator | Does the whole control flow work with the real modules β plan, resume decision, cursor math, checkpoint cycle β with training stubbed? |
| 2. Cheap GPU smoke (20β30 steps) | budgeted ~10 min | Does the machinery run on the accelerator: install, device, one real optimiser step, one push, one verified resume? |
| 3. Measurement probe | 1β1.5 h billed | Does the config hold for hours: sustained rate, flat memory, cadence cost, cold resume from the Hub on an empty disk? |
| 4. The run | everything | β |
Two honest corrections about rung 2, because they are the argument for the ladder, not against it:
- It cost 1.25 h across eight attempts, not ten minutes. Every attempt but the last failed on a real defect:
torchrunhanded the interpreter instead of the script (E-037), a stale-stop guard refusing correctly on a dirty repo (E-038), an argument-order slip indelete_folder(E-039), a harness that could not read its own tee'd output (E-041), and finally a rank-only assertion that killed every successful leg (E-044). A smoke rung's budget is ten minutes; its cost is the number of bugs in the code it is the first to execute. That is still cheaper than finding them in a 4.5-hour session β but plan the budget as "10 min, or 1.5 h if the code is new", not as "10 min". - The CPU rehearsal caught the plumbing defects (E-036, E-037). It did not catch the expensive one: the
rank-only-
on_savecrash needed GPU legs to reproduce and took seven attempts to diagnose. Don't oversell the free rung β it buys you command lines and control flow, not distributed semantics.
Order by price-per-question, and re-run rungs 0β1 on every edit to run code. They are free; the only cost of a rehearsal is the nine minutes you didn't spend on something else.
2. Make the rehearsal rehearse the real thing
- Import the real modules β a rehearsal of a reimplementation proves nothing about the original.
- Print the command line the launcher would execute and parse it in the rehearsal, so a flag typo fails there.
- Point it at real artefacts where it can: hash the actually-published shards, read the actually-verified checkpoint. The Phase 4 rehearsal re-hashed all 151 mix files, which is why the run started knowing its data.
- Assert the same invariants the run will assert, so a rehearsal pass means the run's guards are armed, not merely present.
3. Model size: prove size-dependent things at the real size
"Preflight on tiny models" is the instruction, and taken literally it is wrong. Memory, throughput, checkpoint byte-size, DDP rank behaviour and optimizer-state resume are all size-dependent β a 15M model proves the code path and nothing about the run's numbers. This project therefore ran every Gate 3 proof at the real 106M geometry on 2ΓT4, and used CPU/tiny cells only for properties that genuinely don't scale (reader semantics, order determinism, the storage cycle's logic). Choose per proof: if the quantity you're measuring is a function of size, the size has to be right; if it isn't, spend nothing.
4. What preflight must demonstrate, with logs
The eight proofs. Each is a line in the preflight report, with the log excerpt that shows it.
- Upload a checkpoint to the Hub; restart and pull it back.
- Exact data position on resume β recorded index/hash and a token-offset check, not "loss looked continuous".
- Optimizer and LR state restored: print the scheduler value on the first resumed step and compare it to what the interrupted run would have produced.
- Loss continuity across the restart within noise, with numbers.
- Measured throughput on the target hardware, and the ETA it implies.
- Cold resume: interrupt, kill the instance, start fresh with an empty disk, recover solely from the Hub. This is the only resume test that counts, because every real interruption looks like this. A resume that reads a local leftover file is testing something that will not exist.
- The storage cycle: upload β verify by fetching bytes back β delete locally β confirm free space β keep
training, at sizes representative of the real model plus optimizer state (
07: ~12 bytes/param, not the weight file alone). - Two sequential resumes β position must never drift backwards or re-read consumed data.
Anything unproven here is a bug you have not found yet.
5. Then measure the config, and let the measurement change it
Preflight is the last cheap moment to choose. Throughput is not a spec sheet; it's a distribution on your hardware with your batch shape, and small-N numbers lie:
- A 20-step average is not a rate. The 20-step figure here was 6 % optimistic β the difference between fitting the week and not. Take the sustained number over a long leg (two legs of 120 and 60 steps measured 12,312 and 12,221 tok/s; the planning number adopted was 12,300).
- Measure memory as a time series, not a peak.
max_memory_reservedis process-lifetime, so a post-training validation pass "raises the peak" without any leak. Gate on the leg that does not evaluate, and compare step 10 against step 120 in the same process. - Report the cross-rank max, obtained by an explicit
all_reduce(MAX). Rank-0-only peak statistics under- report, and that is the number you will otherwise make an OOM decision on. - Cost the checkpoint cycle in seconds per push at real conditions: whole-object read-back measured 797 s off a cold CDN; bounded ranged verification runs inside a 7β10 s pushβpoint gap on the live run. That single measurement decides how often you may push, hence how much work an interruption costs.
- If the measurement shows headroom, ask what your current setting is still paying for. 5.99 GB peak on a 15.4 GB card meant gradient checkpointing was buying memory nobody needed; since it computes identical gradients, dropping it was plumbing, not a loss-landscape change. It bought ~30 % throughput and turned an infeasible 29.7-hour run into a 22.6-hour one.
Rule for reopening a frozen choice: only before results exist, only when the change provably leaves gradients identical, and shrink the blast radius first (finer cadence so a failure costs less). Never after a number is on the board β that is tuning, and it invalidates the report.
6. Prove the failure paths, and prefer a discriminating job to an argument
Rungs 0β3 are where you inject faults, because the alternative is discovering them mid-run:
- Serve checkpoint bytes back wrong (flipped byte inside the ranged window, wrong span, a server that ignores
Range) and confirm verification fails. A verifier that has never rejected anything is decoration. - Point prune at an unverified result and confirm it refuses to delete.
- Kill a run mid-leg: confirm the stop guard fires for the right reason and the resume lands where it should.
- Test the absence case: an empty repo must read as "nothing here", not as a timeout or a 403.
- Test rank-blind guards on every rank.
on_saveran on rank 0 only, so a "did you push?" assertion was unsatisfiable on rank 1 and killed every correct leg (E-044). Two rules follow: assert only predicates the rank can actually satisfy, and print diagnostics on all ranks β a rank-0-only logger turns a crash into silence, which is how E-044 survived seven attempts. - When an error fits a plausible story, spend five minutes disproving the story directly (E-017). A cheap, targeted job settled in one run what a morning of reasoning about token scopes could not. The plausible explanation is the expensive one to trust.
7. Sign the report
Write the phase's report with each of the eight proofs (Β§4) and its evidence, plus measured tokens/sec, the main-run ETA, the GPU-hour estimate, and the arithmetic showing it fits the remaining quota. Named files, not conversation β a reader without this project's history must be able to check every claim. If a proof cannot state a number, it is missing, and saying so is cheaper than discovering it later.