|
Download docs/03-preflight-report.md from Cion-lab/ounce100m-code: direct link, hf CLI and curl.
- Browser
- Download file 24.1 kB
-
https://huggingface.co/Cion-lab/ounce100m-code/resolve/main/docs/03-preflight-report.md
- Command line
-
hf download hf://Cion-lab/ounce100m-code/docs/03-preflight-report.md
-
curl -L -o 03-preflight-report.md https://huggingface.co/Cion-lab/ounce100m-code/resolve/main/docs/03-preflight-report.md
24.1 kB
| # 03 — Preflight report (Gate 3) | |
| > §5 Phase 3 requires a written report showing **every** listed proof, plus tokens/sec and the resulting | |
| > main-run ETA and GPU-hour estimate. Each section below states what was proven, the measured number, and | |
| > the artifact that produced it. Where something is **not** proven, it says so. | |
| > | |
| > Cost accounting: the 6 GPU-hour lifetime cap for all small-model tests is tracked in | |
| > `memory/QUOTA.md`; every row referenced here is a ledger entry, including the failed runs — a session | |
| > that crashes still bills. | |
| ## 0. Scorecard | |
| | # | §5 Phase 3 requirement | Stage | Result | Evidence | | |
| |---|---|---|---|---| | |
| | 1 | The training code constructs on the target image | `pargs` (CPU) | ✅ | `p1-hub-checkpoint-cycle` v5 | | |
| | 2 | Checkpoint upload to the Hub | `p1` (CPU) | ✅ | same kernel, 1,699,108,391 B | | |
| | 3 | A restart that pulls the checkpoint back | `p1` (CPU) | ✅ | cold pull, sha-verified, 10.7 s | | |
| | 4 | Resumption from the exact data position (proven, not felt) | `p0` (CPU) | ✅ | `resume_matches_uninterrupted: true` | | |
| | 5 | Optimizer / LR-state restoration | `p3` (GPU) | ✅/⏳ | see §5 | | |
| | 6 | Loss continuity across the restart within noise | `p3` (GPU) | ⏳ | see §5 | | |
| | 7 | Stable loss over a few thousand steps | — | **not proven here** | see §7, Limitations | | |
| | 8 | Measured throughput | `p2` + A/B v2/v3 (GPU) | ✅ | 7,732–7,828 tok/s at seq 1024 | | |
| | 9 | Cold resume from an empty disk, recovering solely from the Hub | `p3` (GPU) | ⏳ | see §5 | | |
| | 10 | §3.13 storage cycle at real checkpoint size, end to end | `p1` + `p3` (GPU) | ✅/⏳ | §3, §5 | | |
| | 11 | Resume at least twice in sequence, position never drifting backwards | `p3` (GPU) | ⏳ | see §5 | | |
| ## 1. `pargs` — does the code construct at all (CPU, free) | |
| Question answered: every `TrainingArguments` keyword the trainer passes is a promise about transformers | |
| 5.0.0, and E-008/E-010 had already shown that 5.x drops things without warning. | |
| Measured on the Kaggle image (transformers **5.0.0**, torch **2.10.0**): | |
| - all `TrainingArguments` kwargs accepted, including `accelerator_config={"dispatch_batches": False}` and | |
| `average_tokens_across_devices=False`; | |
| - 22L/576 model builds at exactly **106,194,240** parameters; the 20L variant at **99,114,048** with | |
| FFN 1536 — matching `01-plan.md` §2.3's closed-form table, so the shape arithmetic and the code agree; | |
| - RoPE verified as `rope_parameters {rope_theta: 10000.0}` — transformers 5 moved it out of the flat | |
| field, which is why `docs/01-plan.md` §2's config line needed reading rather than copying; | |
| - forward + backward finite with gradients; | |
| - the trapezoid schedule measured, not assumed: plateau over **77.4 %** of steps, factor **1.0** at 80 %, | |
| **0.50** at 90 %, **0.0013** at the final update; | |
| - precision asserted as fp16 autocast with **fp32 master weights** and a live `GradScaler` (asserted at | |
| runtime in the trainer, not configured-and-hoped). | |
| Two defects this stage caught before any GPU time was spent: a SwiGLU width computed 2× in the variant | |
| branch, and variant GQA head counts that did not divide `hidden` (which would hard-error inside SDPA). | |
| ## 2. `p0` — the data reader and its resume cursor (CPU, free) | |
| Run against a throwaway 2.35 M-token mix built for the purpose (2 capped sources, published as | |
| `Cion-lab/ounce100m-mix-stagetest`, since deleted-and-kept for reuse by GPU stages): | |
| | property asserted | result | | |
| |---|---| | |
| | a run sliced to start at sample *n−40* reads exactly the windows an uninterrupted run reads | `resume_matches_uninterrupted: true` | | |
| | labels are next-token | `labels_are_next_token: true` | | |
| | same seed ⇒ same order; different seed ⇒ different order | ✅ | | |
| | train/val prefix collision over 187 held-out windows | `0` | | |
| | `verify_mix` on the same bytes | `PASS: true` | | |
| The mechanism being tested: a global PCG64 permutation over all windows, **sliced** at | |
| `start_sample` rather than skipped past, with the cursor stored as | |
| `(samples_consumed, step, seq_len, shuffle_seed, dataset_files_sha)` in `cursor.json`, and | |
| `ignore_data_skip=True` because the reader owns position. | |
| This stage also found E-024 — a parameter deleted while its body check remained, i.e. a `NameError` on | |
| first call that `py_compile` cannot see. | |
| ## 3. `p1` — §3.13's storage cycle at real size (CPU, free) | |
| Random weights, real bytes: 106,194,240 params × 4 B × 4 arrays (fp16-saved weights, fp32 master, Adam | |
| m and v) = **1,699,108,391 B** in one checkpoint, which is what the main run will move every 10 %. | |
| | step | measured | | |
| |---|---| | |
| | push + verify against the Hub's own listing (size + LFS sha256) | **26.3 s** | | |
| | authenticated read-back of one file, anonymous connection | sha matched | | |
| | local prune, free space confirmed | 19.24 → **20.94 GB** free | | |
| | cold pull into an empty directory | **10.7 s** | | |
| | rig cleanup | repo deleted in-session | | |
| Why "listing is not verification" is enforced in code: `verify_checkpoint` returns `ok` only if the | |
| read-back succeeded (E-023 — an earlier version passed while the read-back 404'd, because the report was | |
| built from a fixed key list that silently dropped the error). | |
| ## 4. `p2` — throughput, and the decision it forced (GPU, ledger rows 10-14) | |
| Test T1 measured the **frozen 22L/576 shape** on 2×T4, fp16 + DDP, on real packed data. The result was a | |
| problem: the plan-of-record configuration (seq 2048, SDPA, micro 2) runs at **4,071 tok/s** with peak CUDA | |
| memory at **14.22 GB of ~14.56**, i.e. 1 B tokens needs **68 GPU-hours** against a 30 h/week quota. The | |
| 25.1 h figure in the first draft of `01-plan.md` §6 came from p0c's *raw* loop at seq 1024 and was wrong; | |
| §6.1 records the correction. | |
| Four cells plus two A/B follow-ups, all at 22L/576 unless noted: | |
| | config | tok/s | peak GB | 1 B tokens | | |
| |---|---|---|---| | |
| | sdpa @2048 micro 2 (plan-of-record) | 4,071 | 14.22 | 68.2 h | | |
| | sdpa @2048 micro 1 | 3,954 | 10.04 | 70.3 h | | |
| | eager @2048 micro 2 | **OOM** | — | — | | |
| | eager + grad-ckpt @2048 micro 2 | 4,679 | 10.04 | 59.4 h | | |
| | sdpa + grad-ckpt @2048 micro 4 | 4,505 | 10.03 | 61.7 h | | |
| | **eager @1024 micro 4** | **7,732** | 12.25 | **35.9 h** | | |
| | **eager + grad-ckpt @1024 micro 8** | **7,828** | **7.12** | **35.5 h** | | |
| | sdpa @1024 micro 4 (matched control) | 5,105 | — | 54.4 h | | |
| Three findings, in order of consequence: | |
| 1. **Eager attention beats SDPA on Turing by +52 %** at seq 1024 (7,732 vs 5,105 at matched batch | |
| geometry), and OOMs at 2048. The library default had been costing 34 % for nothing. | |
| 2. **The `Trainer` is not the overhead.** The hypothesis I went looking for was framework cost; the raw | |
| loop measured 5,934 tok/s — *slower* than the Trainer at the same shape. The 11,062 of p0c was an | |
| artefact of comparing a different seq length, model size and batch, not a framework penalty. | |
| 3. **Sequence length is the lever that matters**, and §4 leaves it my choice as long as it is frozen at | |
| launch. Halving the context bought 1.7× and moved peak memory from 98 % of the card to 49 %. | |
| That produced **D-011**: seq **1024**, attention **eager**, gradient checkpointing **on**, micro-batch | |
| **4**, accumulation **32** ⇒ 4 × 1024 × 32 × 2 = **262,144 tokens per step**, identical to the frozen | |
| batch size in tokens, so the step count (**3,814** = `int(1e9/262144)`, i.e. 999,817,216 tokens = 99.98 % | |
| of the target), the LR and the schedule are unchanged. The capability cost — ARC | |
| 25-shot prompts exceed the context and get truncated from the left — is recorded in `05-eval-plan.md` §6, | |
| with no shot count changed. | |
| ## 5. `p3` — cold resume at the frozen geometry (GPU, ledger row 15) | |
| _Status: running as `dodosoomro/ounce100m-gate-3-p3-cold-resume-v2` at time of writing. This section is | |
| filled from `preflight_p3.json` when it lands; nothing below is claimed until it is measured._ | |
| Design, so the numbers can be checked against the intent: | |
| - **Three legs**, cumulative steps 20 → 40 → 60, at 262,144 tokens/step — the real step, the real | |
| checkpoint size, the real cursor, not a cheaper stand-in. | |
| - Before each resume the session **deletes its own local run directories** and re-`df`s, so recovery is | |
| forced through the Hub. That is the only resume test that counts, because every real interruption looks | |
| like this (§3.13). | |
| - Exactly **one checkpoint per leg** (`--push-every-steps` = the leg's step count), pushed, verified from | |
| the Hub, and pruned, so the cursor sequence has three distinct points to check. | |
| - `--resume auto` reads `latest.json`; 404 means "start from zero", and *any other* error aborts rather | |
| than silently restarting from step 0. | |
| - `--attn eager` and `--grad-ckpt` are named on the command line, not inherited from a default, so a later | |
| change to a default cannot silently change what this stage tested. | |
| - Pass criteria, checked in code: three legs exit 0; both resumes report `auto-resume: hub says step`; | |
| at least three cursor records; sample counts **strictly increasing and all distinct**; the resumed model | |
| is the frozen **106,194,240**-parameter model; and the Hub rig is deleted afterwards. | |
| To be recorded here: the three `CKPT` lines with their sample counts, the `prior loss at resume:` values | |
| either side of each wipe, `val_ppl`, `P3_frozen_rate_tok_per_s`, `P3_frozen_peak_gpu_gb`, the wall-clock | |
| cold-start cost per leg, and the session's GPU charge against row 15. | |
| ## 6. Throughput and the main-run estimate | |
| | | | | |
| |---|---| | |
| | Rate at the frozen geometry | _P3's `tok_per_s`_; planning value **7,732 tok/s** from the m4-without-checkpointing cell | | |
| | 1 B tokens | **35.9 h** (35.5 h at the m8 cell) | | |
| | Quota | 30 h/week, 2×T4, billed at **1× container wall-clock** (5 sessions observed) | | |
| | Preflight spend to date | **0.447 h** of the 6 h test cap | | |
| | Calendar | week 1 ≈ 26 h of training (~722 M tokens) → quota reset 2026-09-26T00:00Z → ~10 h in week 2 to finish 1 B | | |
| | Tripwires fixed before launch | **D-012**: below 5,150 tok/s ⇒ launch at 0.9 B. **D-013**: below 7,400 tok/s ⇒ launch at m8 × accum 16 (same 262,144 tokens/step) | | |
| Details in `docs/04-run-log.md` §1-§2. | |
| ## 7. Limitations of these proofs | |
| - **"Stable loss over a few thousand steps" is not proven by preflight, and cannot be at this budget.** | |
| P3 runs 60 steps; a thousand-step stability claim would cost ~0.5 GPU-h that the 6 h test cap cannot | |
| spare on a test rather than on the run. What *is* proven is narrower and honest: the loss is finite, | |
| the optimizer and LR state survive a cold resume, and the loss trace is continuous across two restarts. | |
| Long-horizon stability is monitored on the main run itself (§3.1), against the plan, on every wake. | |
| - **The tiny-model framing of §5 Phase 3 is satisfied differently than literally.** §5 says "test with | |
| 10-20 M parameter models … as cheaply as possible (CPU where you can, GPU only when the thing being | |
| proven needs a GPU)". Every pipeline proof here runs at the **real 106 M shape** rather than a 15 M | |
| stand-in, because the thing being proven — checkpoint bytes, cursor arithmetic, throughput, peak memory | |
| — is size-dependent, and a 15 M model would have proven a cheaper pipeline than the one being launched. | |
| The cost of that reading is quantified in rows 10-15 instead of assumed. | |
| - P3's data is the mix as merged from `ounce100m-mix-stage`, i.e. the *pre-filter* bytes if Gate 2c's | |
| contamination filter is still running. That does not affect what P3 measures (resume correctness, | |
| throughput at fixed token counts), but it is stated so the record is exact. | |
| - Peak-memory numbers are per-cell maxima from `torch.cuda.max_memory_allocated`, which excludes CUDA | |
| cache fragmentation. The main run's real OOM risk is therefore slightly understated at 12.25 GB; this | |
| is one reason D-011 chose checkpointing on. | |
| ### 4.1 Was more throughput available? Audited, and the answer is "not by changing code" | |
| Two checks, because "3.8 % of a T4's datasheet peak" looked like a defect and §3 says never trust a prior — | |
| including my own arithmetic. | |
| 1. **Cross-check against a comparable published run.** A ~163 M-parameter base model trained with DDP on | |
| **8×A100-40GB** reached **234,480 tok/s** at micro-batch 13 ([gilesthomas.com, Jan 2026] | |
| (https://www.gilesthomas.com/2026/01/llm-from-scratch-29-ddp-training-a-base-model-in-the-cloud)) — | |
| **29,310 tok/s per A100**. An A100's fp16 tensor-core peak is roughly an order of magnitude above a | |
| T4's, and the T4's advantage shrinks further because PyTorch's autocast accumulates in fp32. Scaling | |
| that ratio down, the expectation for one T4 at a 106 M model is ~2,900 tok/s; **measured is 3,866 per | |
| card** (7,732 / 2). So the run is slightly *ahead* of the cross-check, and the "low MFU" reading of | |
| §6 in the original plan was an artefact of the marketing peak number, not an unclaimed win. | |
| 2. **`torch.compile`.** The guidance available is that it pays on long, shape-stable workloads and that | |
| `reduce-overhead` targets *small* batches ([Raschka] | |
| (https://sebastianraschka.com/faq/docs/torch-compile-llm-workloads.html)), while slow compilation is a | |
| known hazard in distributed runs | |
| ([pytorch#108971](https://github.com/pytorch/pytorch/issues/108971)). Our batch is not small | |
| (4 × 1024 per rank) and the long-running shape is stable, so it is plausible — but the plausible gain is | |
| tens of percent against a new class of failure (compile hang at rank startup, graph breaks through the | |
| GradScaler path) inside a run that may not be restarted once checkpoint 1 exists (§3.1). | |
| **Decision (D-015, at freeze time, before the run):** no `torch.compile`, no flash-attn — that one is measured, not | |
| assumed: `SDPBackend.FLASH_ATTENTION` raises *"Flash attention only supports gpu architectures in the | |
| range [sm80, sm121]. Attempting to run on a sm 7.5 gpu"* on this hardware (`docs/00-platform-notes.md` | |
| §probe C), so it is architectural and no dependency install changes it; and eager already measured faster | |
| than SDPA here anyway — no quantised optimisers, no gradient-bucket tuning. The three knobs that were genuinely worth GPU time — | |
| sequence length, attention kernel, activation checkpointing — were measured and set in D-011, and together | |
| they moved 68 h to 35.9 h. The remaining ~22 h of two-week slack is better spent absorbing interruptions | |
| than chasing a few percent with an unrestartable failure mode attached. | |
| ## 8. Post-signing mechanism probes (2026-09-20) | |
| Gate 3 signed the *recipe*. These probes test the *session machinery* that only Phase 4 exercises, and they | |
| ran after the freeze and before the first billed session, on the published mix. | |
| ### 8.1 CPU rehearsal — `phase4_prep.py` at head `1dd45f80`, 245.3 s, zero quota, **all stages pass** | |
| | check | measured | | |
| |---|---| | |
| | every published shard, hashed from the Hub download against the manifest | `SHASHED 151 files total_tokens 1,109,714,831 mismatches 0` | | |
| | manifest self-consistency (records vs `n_shards` vs bytes actually mapped) | 139 = 139 = 139; shard tokens sum to the claimed total | | |
| | reader geometry | `1,109,714,831 tokens → 1,083,705 windows of 1024`; horizon `3,814 × 262,144 = 999,817,216` tokens (11.0 % margin) | | |
| | corpus fingerprint (content-hashed, E-035) | `53df4708526da5c8` | | |
| | **session schedule, driven through the shipped `plan()`** | `SESSIONS 5 ends_at 3814 horizon 3814 contiguous True` — 0→762→1,524→2,286→3,048→3,814 | | |
| | **resume reads the same data** | `RESUME_EQUIV step 381 start_sample 97,536 identical True`; same at 1,524 (390,144) and 3,433 (878,848) — input_ids *and* labels compared, through `save_cursor`/`load_cursor` | | |
| | permutation identity guard (E-035 residual, now implemented) | `full.order_sha == perm_sha(...) == resumed.order_sha` at every boundary | | |
| | launcher rehearsal | `PREP_ONLY_OK stop_after_steps 762 planned_steps 762`, `usable_hours 6.3` | | |
| | cursor round-trip + trainer imports on this image | `CURSOR_AND_IMPORT_RC 0`, `TRAINER_HELP_RC 0 lines 75` | | |
| Two real defects surfaced here rather than on a GPU: **E-036** (`--help` crashed because argparse | |
| interpolates help text and two strings contained a bare `%`) and **E-037** (`torchrun` was handed | |
| `sys.executable` as the script, so it tried to compile the Python binary — the same line in the launcher | |
| would have killed session 1 about ten minutes in, after the 2.3 GB pull). | |
| ### 8.2 GPU probes | |
| `p4_stop_probe.py`, two modes, same assertion set. **Probe 0** (soak, 06:12Z) died in 8.6 s on E-037 and | |
| cost effectively nothing. | |
| **Probe 1 v1** (`smoke`: 20 steps at 262,144 tokens/step, no checkpointing, `--push-every-steps 10`, stop at | |
| 10, run directory deleted, resume to 20) ran 778.7 s and **answered the mechanics question outright**: | |
| ``` | |
| CKPT 10 samples=2,560 push+verify=60.3s ok=True pruned=True free=18.67GB loss=10.1926 | |
| TRAIN DONE step=10/20 elapsed=4.5 min tokens_consumed=2,621,440 | |
| auto-resume: hub says step 10 at ckpt/checkpoint-10 | |
| CKPT 20 samples=5,120 push+verify=59.7s ok=True pruned=True free=16.12GB loss=9.5255 | |
| TRAIN DONE step=20/20 elapsed=4.7 min tokens_consumed=5,242,880 | |
| latest.json rolled to the completed run: final at step 20 | |
| ``` | |
| So `--stop-after-steps` stops on the boundary, the push→verify→pointer-read-back→prune cycle works at the | |
| real 1.7 GB size, a cold Hub resume lands on the exact cursor (2,560 → 5,120 samples = step × 256), loss is | |
| continuous across the disk wipe, the terminal pointer roll fires, and disk headroom after two checkpoints is | |
| 16.12 GB. `samples_consumed` and `step` agree by construction at both boundaries. | |
| **What it did not survive:** rank 1 exited 1 with `error_file: <N/A>` and no traceback, so six assertions | |
| could not be evaluated. Two hypotheses were written here at the time — `--no-grad-ckpt` meeting a 25-batch | |
| eval forward, or the segmented stop interacting with the eval loop — and **both were wrong**; §8.4 records | |
| what it actually was (E-044: a rank-0-only assertion that `hub_cb.steps` was empty on rank 1 by | |
| construction). That diagnosis needed three things this attempt lacked: per-rank log files rather than | |
| `--redirects 2` (which filed the stream nobody was reading, E-039), a failure printer that greps the whole | |
| capture rather than its last 40 lines, and marks printed by every rank instead of rank 0. | |
| **This is the exact case the smoke-first ordering was for**: it cost 13 minutes, not the 1.6 hours the soak | |
| would have spent finding it. | |
| ### 8.3 Rehearsal v4 and probe 1 v4/v5 (2026-09-20 07:47-07:50Z) | |
| `dodosoomro/ounce100m-phase4-prep-cpu-rehearsal` **v4 — COMPLETE** after the round-3 edits to the trainer | |
| and launcher (rank-visible evaluation failure, cross-rank peak-memory reduction, complete | |
| `run_summary.json`). The kernel ends `raise SystemExit(0 if (rc1 == rc2 == rc3 == 0) else 4)`, so a COMPLETE | |
| status is the evidence that all three stages — launcher rehearsal, the 151-shard content hash and the | |
| cursor/import checks — still pass on head `4c7f4922`. The reader's guarantees therefore survive the changes | |
| that were made in response to review. | |
| Probe 1 **v4** re-established the mechanics on a clean scratch repo (`PROBE_REPO | |
| Cion-lab/ounce100m-ckpt-probe-0920-0727`, `LEG1 pushed: ['checkpoint-10'] pointer: {'step': 10, 'path_in_repo': | |
| 'ckpt/checkpoint-10'}`) and died again at 275 s in the same place — the post-train validation pass — this | |
| time with **no Python traceback even though stderr was not redirected**, which is what pointed at the | |
| `say()`-is-silent-on-rank-1 mechanism (E-039) rather than a data or memory error. v5 adds `--tee 3`, a | |
| directory listing before any tail, `NCCL_DEBUG=WARN`, and per-rank reporting of the evaluation failure. | |
| Decision still open, and bounded either way: if the validation pass cannot run at `--no-grad-ckpt`, D-017 | |
| stands, the run keeps checkpointing at 9,696 tok/s, and nothing else changes. If it can, D-018 adopts | |
| 12,792 tok/s with a 60-step recovery cadence. The smoke probe is 13 minutes; the soak is 1.6 h; session 1 | |
| follows whichever way it lands. | |
| ### 8.4 Probe 1 closed (v6, v7, v8 — 09:00Z to 09:42Z): `PASS`, and the crash was never the memory config | |
| Three more attempts, each one buying the evidence the previous attempt could not print. | |
| **v6 (581.1 s, FAIL).** With the tee-prefix bug fixed and the whole-object read-back replaced by two bounded | |
| `Range` requests, the checkpoint cycle fell from `push+verify=797.7s` to **13.4 s**, and leg 1 still died | |
| after its forced stop. The probe's own `DIAG_leg1` printed nothing at all — it captured the per-rank logs and | |
| discarded them, because its output went through the same line-prefix filter as everything else (E-043). | |
| **v7 (682.7 s, FAIL) — the diagnostic that worked.** The trainer now marks the post-training sequence with | |
| per-rank `print`s, and leg 1 printed: | |
| ``` | |
| [rank 0] post-train: step=10/20 log_history_loss_entries=2 final_loss=10.192607879638672 start_step=0 | |
| [rank 0] past the stop guard: forced_stop=True stopped_at=10 | |
| validation skipped: segment ended at step 10 of 20 | |
| [rank 0] peak_stats: entering (this is a collective) | |
| ``` | |
| and nothing from rank 1 — while rank 1's log, which `diag` now actually prints, ended with *"segment ended at | |
| step 10 but no checkpoint was pushed and verified for it (pushed: [])"*. Rank 1 was killing the run on | |
| **our own assertion**: `HubPush.on_save` returns early on non-zero ranks, so the list that assertion reads is | |
| empty on rank 1 by construction, and every correctly-checkpointed forced-stop leg died there. Not a | |
| collective, not `evaluate()`, not memory. E-044, and it retires the E-040 hypothesis in §8.2. | |
| **v8 (578.0 s) — `VERDICT P4PROBE PASS []`, 16/16 checks, both legs rc 0.** Leg 1 stops at 10 of 20, pushes, | |
| Hub-verifies, rolls `latest.json`, reads it back, prunes, skips validation and exits 0 with | |
| `val_skipped: true`; the probe wipes the run directory; leg 2 resumes from the Hub alone and runs to the | |
| horizon, pushes `ckpt/checkpoint-20` and `final/`, runs the 195-window held-out pass (`loss 9.3532`, | |
| `ppl 11535.95`) and rolls the terminal pointer. Tokens 2,621,440 → 5,242,880 = 10 and 20 × 262,144; | |
| `samples_consumed` 2,560 → 5,120; `shuffle_perm_sha` and `dataset_files_sha` identical across the resume; | |
| `params 106,194,240`; `grad_ckpt false` read off the artifact; loss `10.7723 → 10.1926 → 9.5255`. | |
| Numbers carried forward, all at the frozen geometry with checkpointing **off**: | |
| | quantity | value | where it came from | | |
| |---|---|---| | |
| | step time | **20.20 s** = 12,977 tok/s | v8 progress bar, 10-step and 20-step legs agreeing | | |
| | horizon wall clock | **21.4 h** of stepping (3,814 steps) | that rate × the step count | | |
| | checkpoint cycle | **12.7-16.5 s** per 1.27 GB push+verify+pointer+prune | v7/v8 `CKPT` lines | | |
| | memory, training only | **12.84 GiB reserved / 12.70 allocated**, cross-rank max | v8 leg 1, which never evaluates | | |
| | memory, after validation | 13.66 GiB reserved | v8 leg 2 — the same high-water mark, plus the eval pass | | |
| | session shape | 5 sessions of 762 steps, last running to 3,814 | rehearsal v3/v4 driving the shipped `plan()` | | |
| The split of the two memory figures is the reason the soak's gate is `peak_memory_during_training_below_13_6_gb` | |
| rather than a check on both legs: `max_memory_reserved` is process-lifetime, so a leg that evaluates reports | |
| training-then-eval, and the run occupies the training state for all 3,814 steps. | |
| **Two review-found defects closed in the same window** (E-042, E-045): the prune guard lived only inside the | |
| wrapper the save path bypasses, so a failed verification would have deleted the last local copy *and* moved | |
| the pointer onto it; and `verify_checkpoint`'s length check is now accompanied by a `Content-Range` check, | |
| because 850 MB of the right length from the wrong offset passed the old test. Both are rehearsed on free CPU | |
| in stage 4 of `phase4_prep`, against the 1.27 GB checkpoint this probe really pushed — clean, tampered inside | |
| the read window, tampered outside it, absent, Range-ignoring, wrong-offset, wrong-bytes, and the prune | |
| refusal itself. | |