Download docs/03-preflight-report.md from Cion-lab/ounce100m-code: direct link, hf CLI and curl.
- Browser
- Download file 24.1 kB
-
https://huggingface.co/Cion-lab/ounce100m-code/resolve/main/docs/03-preflight-report.md
- Command line
-
hf download hf://Cion-lab/ounce100m-code/docs/03-preflight-report.md
-
curl -L -o 03-preflight-report.md https://huggingface.co/Cion-lab/ounce100m-code/resolve/main/docs/03-preflight-report.md
03 β Preflight report (Gate 3)
Β§5 Phase 3 requires a written report showing every listed proof, plus tokens/sec and the resulting main-run ETA and GPU-hour estimate. Each section below states what was proven, the measured number, and the artifact that produced it. Where something is not proven, it says so.
Cost accounting: the 6 GPU-hour lifetime cap for all small-model tests is tracked in
memory/QUOTA.md; every row referenced here is a ledger entry, including the failed runs β a session that crashes still bills.
0. Scorecard
| # | Β§5 Phase 3 requirement | Stage | Result | Evidence |
|---|---|---|---|---|
| 1 | The training code constructs on the target image | pargs (CPU) |
β | p1-hub-checkpoint-cycle v5 |
| 2 | Checkpoint upload to the Hub | p1 (CPU) |
β | same kernel, 1,699,108,391 B |
| 3 | A restart that pulls the checkpoint back | p1 (CPU) |
β | cold pull, sha-verified, 10.7 s |
| 4 | Resumption from the exact data position (proven, not felt) | p0 (CPU) |
β | resume_matches_uninterrupted: true |
| 5 | Optimizer / LR-state restoration | p3 (GPU) |
β /β³ | see Β§5 |
| 6 | Loss continuity across the restart within noise | p3 (GPU) |
β³ | see Β§5 |
| 7 | Stable loss over a few thousand steps | β | not proven here | see Β§7, Limitations |
| 8 | Measured throughput | p2 + A/B v2/v3 (GPU) |
β | 7,732β7,828 tok/s at seq 1024 |
| 9 | Cold resume from an empty disk, recovering solely from the Hub | p3 (GPU) |
β³ | see Β§5 |
| 10 | Β§3.13 storage cycle at real checkpoint size, end to end | p1 + p3 (GPU) |
β /β³ | Β§3, Β§5 |
| 11 | Resume at least twice in sequence, position never drifting backwards | p3 (GPU) |
β³ | see Β§5 |
1. pargs β does the code construct at all (CPU, free)
Question answered: every TrainingArguments keyword the trainer passes is a promise about transformers
5.0.0, and E-008/E-010 had already shown that 5.x drops things without warning.
Measured on the Kaggle image (transformers 5.0.0, torch 2.10.0):
- all
TrainingArgumentskwargs accepted, includingaccelerator_config={"dispatch_batches": False}andaverage_tokens_across_devices=False; - 22L/576 model builds at exactly 106,194,240 parameters; the 20L variant at 99,114,048 with
FFN 1536 β matching
01-plan.mdΒ§2.3's closed-form table, so the shape arithmetic and the code agree; - RoPE verified as
rope_parameters {rope_theta: 10000.0}β transformers 5 moved it out of the flat field, which is whydocs/01-plan.mdΒ§2's config line needed reading rather than copying; - forward + backward finite with gradients;
- the trapezoid schedule measured, not assumed: plateau over 77.4 % of steps, factor 1.0 at 80 %, 0.50 at 90 %, 0.0013 at the final update;
- precision asserted as fp16 autocast with fp32 master weights and a live
GradScaler(asserted at runtime in the trainer, not configured-and-hoped).
Two defects this stage caught before any GPU time was spent: a SwiGLU width computed 2Γ in the variant
branch, and variant GQA head counts that did not divide hidden (which would hard-error inside SDPA).
2. p0 β the data reader and its resume cursor (CPU, free)
Run against a throwaway 2.35 M-token mix built for the purpose (2 capped sources, published as
Cion-lab/ounce100m-mix-stagetest, since deleted-and-kept for reuse by GPU stages):
| property asserted | result |
|---|---|
| a run sliced to start at sample nβ40 reads exactly the windows an uninterrupted run reads | resume_matches_uninterrupted: true |
| labels are next-token | labels_are_next_token: true |
| same seed β same order; different seed β different order | β |
| train/val prefix collision over 187 held-out windows | 0 |
verify_mix on the same bytes |
PASS: true |
The mechanism being tested: a global PCG64 permutation over all windows, sliced at
start_sample rather than skipped past, with the cursor stored as
(samples_consumed, step, seq_len, shuffle_seed, dataset_files_sha) in cursor.json, and
ignore_data_skip=True because the reader owns position.
This stage also found E-024 β a parameter deleted while its body check remained, i.e. a NameError on
first call that py_compile cannot see.
3. p1 β Β§3.13's storage cycle at real size (CPU, free)
Random weights, real bytes: 106,194,240 params Γ 4 B Γ 4 arrays (fp16-saved weights, fp32 master, Adam m and v) = 1,699,108,391 B in one checkpoint, which is what the main run will move every 10 %.
| step | measured |
|---|---|
| push + verify against the Hub's own listing (size + LFS sha256) | 26.3 s |
| authenticated read-back of one file, anonymous connection | sha matched |
| local prune, free space confirmed | 19.24 β 20.94 GB free |
| cold pull into an empty directory | 10.7 s |
| rig cleanup | repo deleted in-session |
Why "listing is not verification" is enforced in code: verify_checkpoint returns ok only if the
read-back succeeded (E-023 β an earlier version passed while the read-back 404'd, because the report was
built from a fixed key list that silently dropped the error).
4. p2 β throughput, and the decision it forced (GPU, ledger rows 10-14)
Test T1 measured the frozen 22L/576 shape on 2ΓT4, fp16 + DDP, on real packed data. The result was a
problem: the plan-of-record configuration (seq 2048, SDPA, micro 2) runs at 4,071 tok/s with peak CUDA
memory at 14.22 GB of ~14.56, i.e. 1 B tokens needs 68 GPU-hours against a 30 h/week quota. The
25.1 h figure in the first draft of 01-plan.md Β§6 came from p0c's raw loop at seq 1024 and was wrong;
Β§6.1 records the correction.
Four cells plus two A/B follow-ups, all at 22L/576 unless noted:
| config | tok/s | peak GB | 1 B tokens |
|---|---|---|---|
| sdpa @2048 micro 2 (plan-of-record) | 4,071 | 14.22 | 68.2 h |
| sdpa @2048 micro 1 | 3,954 | 10.04 | 70.3 h |
| eager @2048 micro 2 | OOM | β | β |
| eager + grad-ckpt @2048 micro 2 | 4,679 | 10.04 | 59.4 h |
| sdpa + grad-ckpt @2048 micro 4 | 4,505 | 10.03 | 61.7 h |
| eager @1024 micro 4 | 7,732 | 12.25 | 35.9 h |
| eager + grad-ckpt @1024 micro 8 | 7,828 | 7.12 | 35.5 h |
| sdpa @1024 micro 4 (matched control) | 5,105 | β | 54.4 h |
Three findings, in order of consequence:
- Eager attention beats SDPA on Turing by +52 % at seq 1024 (7,732 vs 5,105 at matched batch geometry), and OOMs at 2048. The library default had been costing 34 % for nothing.
- The
Traineris not the overhead. The hypothesis I went looking for was framework cost; the raw loop measured 5,934 tok/s β slower than the Trainer at the same shape. The 11,062 of p0c was an artefact of comparing a different seq length, model size and batch, not a framework penalty. - Sequence length is the lever that matters, and Β§4 leaves it my choice as long as it is frozen at launch. Halving the context bought 1.7Γ and moved peak memory from 98 % of the card to 49 %.
That produced D-011: seq 1024, attention eager, gradient checkpointing on, micro-batch
4, accumulation 32 β 4 Γ 1024 Γ 32 Γ 2 = 262,144 tokens per step, identical to the frozen
batch size in tokens, so the step count (3,814 = int(1e9/262144), i.e. 999,817,216 tokens = 99.98 %
of the target), the LR and the schedule are unchanged. The capability cost β ARC
25-shot prompts exceed the context and get truncated from the left β is recorded in 05-eval-plan.md Β§6,
with no shot count changed.
5. p3 β cold resume at the frozen geometry (GPU, ledger row 15)
Status: running as dodosoomro/ounce100m-gate-3-p3-cold-resume-v2 at time of writing. This section is
filled from preflight_p3.json when it lands; nothing below is claimed until it is measured.
Design, so the numbers can be checked against the intent:
- Three legs, cumulative steps 20 β 40 β 60, at 262,144 tokens/step β the real step, the real checkpoint size, the real cursor, not a cheaper stand-in.
- Before each resume the session deletes its own local run directories and re-
dfs, so recovery is forced through the Hub. That is the only resume test that counts, because every real interruption looks like this (Β§3.13). - Exactly one checkpoint per leg (
--push-every-steps= the leg's step count), pushed, verified from the Hub, and pruned, so the cursor sequence has three distinct points to check. --resume autoreadslatest.json; 404 means "start from zero", and any other error aborts rather than silently restarting from step 0.--attn eagerand--grad-ckptare named on the command line, not inherited from a default, so a later change to a default cannot silently change what this stage tested.- Pass criteria, checked in code: three legs exit 0; both resumes report
auto-resume: hub says step; at least three cursor records; sample counts strictly increasing and all distinct; the resumed model is the frozen 106,194,240-parameter model; and the Hub rig is deleted afterwards.
To be recorded here: the three CKPT lines with their sample counts, the prior loss at resume: values
either side of each wipe, val_ppl, P3_frozen_rate_tok_per_s, P3_frozen_peak_gpu_gb, the wall-clock
cold-start cost per leg, and the session's GPU charge against row 15.
6. Throughput and the main-run estimate
| Rate at the frozen geometry | P3's tok_per_s; planning value 7,732 tok/s from the m4-without-checkpointing cell |
| 1 B tokens | 35.9 h (35.5 h at the m8 cell) |
| Quota | 30 h/week, 2ΓT4, billed at 1Γ container wall-clock (5 sessions observed) |
| Preflight spend to date | 0.447 h of the 6 h test cap |
| Calendar | week 1 β 26 h of training (~722 M tokens) β quota reset 2026-09-26T00:00Z β ~10 h in week 2 to finish 1 B |
| Tripwires fixed before launch | D-012: below 5,150 tok/s β launch at 0.9 B. D-013: below 7,400 tok/s β launch at m8 Γ accum 16 (same 262,144 tokens/step) |
Details in docs/04-run-log.md Β§1-Β§2.
7. Limitations of these proofs
- "Stable loss over a few thousand steps" is not proven by preflight, and cannot be at this budget. P3 runs 60 steps; a thousand-step stability claim would cost ~0.5 GPU-h that the 6 h test cap cannot spare on a test rather than on the run. What is proven is narrower and honest: the loss is finite, the optimizer and LR state survive a cold resume, and the loss trace is continuous across two restarts. Long-horizon stability is monitored on the main run itself (Β§3.1), against the plan, on every wake.
- The tiny-model framing of Β§5 Phase 3 is satisfied differently than literally. Β§5 says "test with 10-20 M parameter models β¦ as cheaply as possible (CPU where you can, GPU only when the thing being proven needs a GPU)". Every pipeline proof here runs at the real 106 M shape rather than a 15 M stand-in, because the thing being proven β checkpoint bytes, cursor arithmetic, throughput, peak memory β is size-dependent, and a 15 M model would have proven a cheaper pipeline than the one being launched. The cost of that reading is quantified in rows 10-15 instead of assumed.
- P3's data is the mix as merged from
ounce100m-mix-stage, i.e. the pre-filter bytes if Gate 2c's contamination filter is still running. That does not affect what P3 measures (resume correctness, throughput at fixed token counts), but it is stated so the record is exact. - Peak-memory numbers are per-cell maxima from
torch.cuda.max_memory_allocated, which excludes CUDA cache fragmentation. The main run's real OOM risk is therefore slightly understated at 12.25 GB; this is one reason D-011 chose checkpointing on.
4.1 Was more throughput available? Audited, and the answer is "not by changing code"
Two checks, because "3.8 % of a T4's datasheet peak" looked like a defect and Β§3 says never trust a prior β including my own arithmetic.
- Cross-check against a comparable published run. A ~163 M-parameter base model trained with DDP on 8ΓA100-40GB reached 234,480 tok/s at micro-batch 13 ([gilesthomas.com, Jan 2026] (https://www.gilesthomas.com/2026/01/llm-from-scratch-29-ddp-training-a-base-model-in-the-cloud)) β 29,310 tok/s per A100. An A100's fp16 tensor-core peak is roughly an order of magnitude above a T4's, and the T4's advantage shrinks further because PyTorch's autocast accumulates in fp32. Scaling that ratio down, the expectation for one T4 at a 106 M model is ~2,900 tok/s; measured is 3,866 per card (7,732 / 2). So the run is slightly ahead of the cross-check, and the "low MFU" reading of Β§6 in the original plan was an artefact of the marketing peak number, not an unclaimed win.
torch.compile. The guidance available is that it pays on long, shape-stable workloads and thatreduce-overheadtargets small batches ([Raschka] (https://sebastianraschka.com/faq/docs/torch-compile-llm-workloads.html)), while slow compilation is a known hazard in distributed runs (pytorch#108971). Our batch is not small (4 Γ 1024 per rank) and the long-running shape is stable, so it is plausible β but the plausible gain is tens of percent against a new class of failure (compile hang at rank startup, graph breaks through the GradScaler path) inside a run that may not be restarted once checkpoint 1 exists (Β§3.1).
Decision (D-015, at freeze time, before the run): no torch.compile, no flash-attn β that one is measured, not
assumed: SDPBackend.FLASH_ATTENTION raises "Flash attention only supports gpu architectures in the
range [sm80, sm121]. Attempting to run on a sm 7.5 gpu" on this hardware (docs/00-platform-notes.md
Β§probe C), so it is architectural and no dependency install changes it; and eager already measured faster
than SDPA here anyway β no quantised optimisers, no gradient-bucket tuning. The three knobs that were genuinely worth GPU time β
sequence length, attention kernel, activation checkpointing β were measured and set in D-011, and together
they moved 68 h to 35.9 h. The remaining ~22 h of two-week slack is better spent absorbing interruptions
than chasing a few percent with an unrestartable failure mode attached.
8. Post-signing mechanism probes (2026-09-20)
Gate 3 signed the recipe. These probes test the session machinery that only Phase 4 exercises, and they ran after the freeze and before the first billed session, on the published mix.
8.1 CPU rehearsal β phase4_prep.py at head 1dd45f80, 245.3 s, zero quota, all stages pass
| check | measured |
|---|---|
| every published shard, hashed from the Hub download against the manifest | SHASHED 151 files total_tokens 1,109,714,831 mismatches 0 |
manifest self-consistency (records vs n_shards vs bytes actually mapped) |
139 = 139 = 139; shard tokens sum to the claimed total |
| reader geometry | 1,109,714,831 tokens β 1,083,705 windows of 1024; horizon 3,814 Γ 262,144 = 999,817,216 tokens (11.0 % margin) |
| corpus fingerprint (content-hashed, E-035) | 53df4708526da5c8 |
session schedule, driven through the shipped plan() |
SESSIONS 5 ends_at 3814 horizon 3814 contiguous True β 0β762β1,524β2,286β3,048β3,814 |
| resume reads the same data | RESUME_EQUIV step 381 start_sample 97,536 identical True; same at 1,524 (390,144) and 3,433 (878,848) β input_ids and labels compared, through save_cursor/load_cursor |
| permutation identity guard (E-035 residual, now implemented) | full.order_sha == perm_sha(...) == resumed.order_sha at every boundary |
| launcher rehearsal | PREP_ONLY_OK stop_after_steps 762 planned_steps 762, usable_hours 6.3 |
| cursor round-trip + trainer imports on this image | CURSOR_AND_IMPORT_RC 0, TRAINER_HELP_RC 0 lines 75 |
Two real defects surfaced here rather than on a GPU: E-036 (--help crashed because argparse
interpolates help text and two strings contained a bare %) and E-037 (torchrun was handed
sys.executable as the script, so it tried to compile the Python binary β the same line in the launcher
would have killed session 1 about ten minutes in, after the 2.3 GB pull).
8.2 GPU probes
p4_stop_probe.py, two modes, same assertion set. Probe 0 (soak, 06:12Z) died in 8.6 s on E-037 and
cost effectively nothing.
Probe 1 v1 (smoke: 20 steps at 262,144 tokens/step, no checkpointing, --push-every-steps 10, stop at
10, run directory deleted, resume to 20) ran 778.7 s and answered the mechanics question outright:
CKPT 10 samples=2,560 push+verify=60.3s ok=True pruned=True free=18.67GB loss=10.1926
TRAIN DONE step=10/20 elapsed=4.5 min tokens_consumed=2,621,440
auto-resume: hub says step 10 at ckpt/checkpoint-10
CKPT 20 samples=5,120 push+verify=59.7s ok=True pruned=True free=16.12GB loss=9.5255
TRAIN DONE step=20/20 elapsed=4.7 min tokens_consumed=5,242,880
latest.json rolled to the completed run: final at step 20
So --stop-after-steps stops on the boundary, the pushβverifyβpointer-read-backβprune cycle works at the
real 1.7 GB size, a cold Hub resume lands on the exact cursor (2,560 β 5,120 samples = step Γ 256), loss is
continuous across the disk wipe, the terminal pointer roll fires, and disk headroom after two checkpoints is
16.12 GB. samples_consumed and step agree by construction at both boundaries.
What it did not survive: rank 1 exited 1 with error_file: <N/A> and no traceback, so six assertions
could not be evaluated. Two hypotheses were written here at the time β --no-grad-ckpt meeting a 25-batch
eval forward, or the segmented stop interacting with the eval loop β and both were wrong; Β§8.4 records
what it actually was (E-044: a rank-0-only assertion that hub_cb.steps was empty on rank 1 by
construction). That diagnosis needed three things this attempt lacked: per-rank log files rather than
--redirects 2 (which filed the stream nobody was reading, E-039), a failure printer that greps the whole
capture rather than its last 40 lines, and marks printed by every rank instead of rank 0.
This is the exact case the smoke-first ordering was for: it cost 13 minutes, not the 1.6 hours the soak
would have spent finding it.
8.3 Rehearsal v4 and probe 1 v4/v5 (2026-09-20 07:47-07:50Z)
dodosoomro/ounce100m-phase4-prep-cpu-rehearsal v4 β COMPLETE after the round-3 edits to the trainer
and launcher (rank-visible evaluation failure, cross-rank peak-memory reduction, complete
run_summary.json). The kernel ends raise SystemExit(0 if (rc1 == rc2 == rc3 == 0) else 4), so a COMPLETE
status is the evidence that all three stages β launcher rehearsal, the 151-shard content hash and the
cursor/import checks β still pass on head 4c7f4922. The reader's guarantees therefore survive the changes
that were made in response to review.
Probe 1 v4 re-established the mechanics on a clean scratch repo (PROBE_REPO Cion-lab/ounce100m-ckpt-probe-0920-0727, LEG1 pushed: ['checkpoint-10'] pointer: {'step': 10, 'path_in_repo': 'ckpt/checkpoint-10'}) and died again at 275 s in the same place β the post-train validation pass β this
time with no Python traceback even though stderr was not redirected, which is what pointed at the
say()-is-silent-on-rank-1 mechanism (E-039) rather than a data or memory error. v5 adds --tee 3, a
directory listing before any tail, NCCL_DEBUG=WARN, and per-rank reporting of the evaluation failure.
Decision still open, and bounded either way: if the validation pass cannot run at --no-grad-ckpt, D-017
stands, the run keeps checkpointing at 9,696 tok/s, and nothing else changes. If it can, D-018 adopts
12,792 tok/s with a 60-step recovery cadence. The smoke probe is 13 minutes; the soak is 1.6 h; session 1
follows whichever way it lands.
8.4 Probe 1 closed (v6, v7, v8 β 09:00Z to 09:42Z): PASS, and the crash was never the memory config
Three more attempts, each one buying the evidence the previous attempt could not print.
v6 (581.1 s, FAIL). With the tee-prefix bug fixed and the whole-object read-back replaced by two bounded
Range requests, the checkpoint cycle fell from push+verify=797.7s to 13.4 s, and leg 1 still died
after its forced stop. The probe's own DIAG_leg1 printed nothing at all β it captured the per-rank logs and
discarded them, because its output went through the same line-prefix filter as everything else (E-043).
v7 (682.7 s, FAIL) β the diagnostic that worked. The trainer now marks the post-training sequence with
per-rank prints, and leg 1 printed:
[rank 0] post-train: step=10/20 log_history_loss_entries=2 final_loss=10.192607879638672 start_step=0
[rank 0] past the stop guard: forced_stop=True stopped_at=10
validation skipped: segment ended at step 10 of 20
[rank 0] peak_stats: entering (this is a collective)
and nothing from rank 1 β while rank 1's log, which diag now actually prints, ended with "segment ended at
step 10 but no checkpoint was pushed and verified for it (pushed: [])". Rank 1 was killing the run on
our own assertion: HubPush.on_save returns early on non-zero ranks, so the list that assertion reads is
empty on rank 1 by construction, and every correctly-checkpointed forced-stop leg died there. Not a
collective, not evaluate(), not memory. E-044, and it retires the E-040 hypothesis in Β§8.2.
v8 (578.0 s) β VERDICT P4PROBE PASS [], 16/16 checks, both legs rc 0. Leg 1 stops at 10 of 20, pushes,
Hub-verifies, rolls latest.json, reads it back, prunes, skips validation and exits 0 with
val_skipped: true; the probe wipes the run directory; leg 2 resumes from the Hub alone and runs to the
horizon, pushes ckpt/checkpoint-20 and final/, runs the 195-window held-out pass (loss 9.3532,
ppl 11535.95) and rolls the terminal pointer. Tokens 2,621,440 β 5,242,880 = 10 and 20 Γ 262,144;
samples_consumed 2,560 β 5,120; shuffle_perm_sha and dataset_files_sha identical across the resume;
params 106,194,240; grad_ckpt false read off the artifact; loss 10.7723 β 10.1926 β 9.5255.
Numbers carried forward, all at the frozen geometry with checkpointing off:
| quantity | value | where it came from |
|---|---|---|
| step time | 20.20 s = 12,977 tok/s | v8 progress bar, 10-step and 20-step legs agreeing |
| horizon wall clock | 21.4 h of stepping (3,814 steps) | that rate Γ the step count |
| checkpoint cycle | 12.7-16.5 s per 1.27 GB push+verify+pointer+prune | v7/v8 CKPT lines |
| memory, training only | 12.84 GiB reserved / 12.70 allocated, cross-rank max | v8 leg 1, which never evaluates |
| memory, after validation | 13.66 GiB reserved | v8 leg 2 β the same high-water mark, plus the eval pass |
| session shape | 5 sessions of 762 steps, last running to 3,814 | rehearsal v3/v4 driving the shipped plan() |
The split of the two memory figures is the reason the soak's gate is peak_memory_during_training_below_13_6_gb
rather than a check on both legs: max_memory_reserved is process-lifetime, so a leg that evaluates reports
training-then-eval, and the run occupies the training state for all 3,814 steps.
Two review-found defects closed in the same window (E-042, E-045): the prune guard lived only inside the
wrapper the save path bypasses, so a failed verification would have deleted the last local copy and moved
the pointer onto it; and verify_checkpoint's length check is now accompanied by a Content-Range check,
because 850 MB of the right length from the wrong offset passed the old test. Both are rehearsed on free CPU
in stage 4 of phase4_prep, against the 1.27 GB checkpoint this probe really pushed β clean, tampered inside
the read window, tampered outside it, absent, Range-ignoring, wrong-offset, wrong-bytes, and the prune
refusal itself.