|
Download guide/06-checklists.md from Cion-lab/ounce100m-code: direct link, hf CLI and curl.
- Browser
- Download file 8.89 kB
-
https://huggingface.co/Cion-lab/ounce100m-code/resolve/main/guide/06-checklists.md
- Command line
-
hf download hf://Cion-lab/ounce100m-code/guide/06-checklists.md
-
curl -L -o 06-checklists.md https://huggingface.co/Cion-lab/ounce100m-code/resolve/main/guide/06-checklists.md
8.89 kB
| # 06 β Checklists | |
| The working page. Tick nothing from memory. Each gate names the error/decision records it came from, so a claim | |
| here is checkable against the ledger rather than taken on trust. | |
| ## Gate 0 β Orientation Β· `E-001..E-019` | |
| - [ ] Credentials verified by a call returning identity, not by a file existing | |
| - [ ] **A secret store's privacy proved by two different calls** (authenticated listing *and* anonymous read) β | |
| 403/404 alone cannot distinguish "private" from "never created" (`E-011`) | |
| - [ ] Remote quota read live; hours remaining + reset **instant** written in the ledger (look up the weekday) | |
| - [ ] One throwaway CPU job and one throwaway GPU job completed end to end; mechanics recorded | |
| - [ ] Any async job from a previous session found and accounted for (running / finished / dead) | |
| - [ ] Memory files match reality; every discrepancy logged, not silently corrected | |
| ## Gate 1 β Plan Β· `E-020..E-026` | |
| - [ ] Architecture, hyperparameters, token total, step count written and dated | |
| - [ ] Throughput estimate with its source, and the GPU-hour arithmetic it implies | |
| - [ ] **Budget has a line for every remaining phase, including publish and benchmarks** (`D-019`) | |
| - [ ] Per-benchmark target bands pre-registered, each with a citation | |
| - [ ] Eval split + metric named per task **before** any eval code exists | |
| - [ ] What you may change mid-run vs what is frozen, stated as one sentence | |
| - [ ] Every choice traceable to a source you actually looked at | |
| ## Gate 2 β Data Β· `E-027..E-035` | |
| - [ ] Assembled above target; validation split reserved before sharding | |
| - [ ] Deduplicated; mechanical overlap audit at `overlap_total 0` | |
| - [ ] **Audit touched `validation`/`dev` only; no `test` split opened** | |
| - [ ] Audit denominators (rows per task) asserted nonzero and printed beside the counts | |
| - [ ] No benchmark item text in any published audit artefact, including exception strings | |
| - [ ] Shard layout proven small enough that a job reads a few shards, not the mix | |
| - [ ] Packing boundary choice *measured*, not assumed (seam rate) | |
| - [ ] Published public; token count re-verified **from the Hub copy** against the manifest | |
| - [ ] Manifest self-consistent (sums, counts, no empties); per-file + set fingerprints recorded | |
| - [ ] Cursor design supports exact positional resume; fingerprint keys chosen and persisted | |
| ## Gate 3 β Preflight Β· `E-036..E-044` Β· `D-012..D-018` | |
| - [ ] **Rehearsal re-run after every edit to run code** β it's free | |
| - [ ] Rehearsal imports the real modules and parses the real launch command line | |
| - [ ] Each proof run **at the size the property depends on** (`03` Β§3), not at a token tiny model | |
| - [ ] Fault injection: wrong bytes, ignored `Range`, unverified prune, absent repo β all *rejected* | |
| - [ ] Guards tested **on every rank**; diagnostics printed on every rank, not just rank 0 | |
| - [ ] Upload β restart β pull-back, on the Hub | |
| - [ ] Exact data position proven by index/hash **and** token offset | |
| - [ ] Optimizer/LR state restored, values printed; loss continuous across restart | |
| - [ ] **Cold resume** on a fresh instance, empty disk, recovering solely from the Hub | |
| - [ ] Storage cycle end to end at representative size (**weights + optimizer state**), free space logged | |
| - [ ] Two sequential resumes, no backward drift, no re-read data | |
| - [ ] Sustained tokens/sec over a long leg (not 20 steps); memory as a time series, cross-rank max | |
| - [ ] Checkpoint-cycle cost measured in seconds per push | |
| - [ ] Report: the eight proofs (`03` Β§4) each with evidence + ETA + GPU-hour arithmetic that fits the quota | |
| ## Gate 4 β Run Β· `E-042`, `D-018` | |
| - [ ] Config frozen; hash recorded and asserted at every session start | |
| - [ ] Cadence divides session length; session tiling ends exactly on the horizon | |
| - [ ] Per checkpoint: write β push β **verify bytes** β point β read pointer back β prune | |
| - [ ] Prune gated by the verification result *inside* the prune function | |
| - [ ] Free space logged beside loss, every cycle | |
| - [ ] Ledger row booked **before** each GPU job; hours remaining re-read **at each launch** | |
| - [ ] `tokens == steps Γ tokens_per_step` and cursor sample math checked at each checkpoint | |
| - [ ] Loss anomalies documented with grad-norm and LR context, config unchanged | |
| - [ ] **Results written before trailing cross-checks**; verification cannot destroy a completed report (`E-012`) | |
| - [ ] Multi-week plan written if the estimate exceeds remaining hours | |
| ## Gate 5 β Publish Β· `E-045`, `E-046` | |
| - [ ] Card fields generated from artefacts; required-field guard demonstrably fails on an empty run | |
| - [ ] Parameter/vocab/geometry assertions recomputed from the pushed files | |
| - [ ] Clean-room `from_pretrained` + generate: env set **before** import, `token=False`, **and every token var | |
| popped from the environment** | |
| - [ ] Anonymous re-hash of every published file | |
| - [ ] 404 vs timeout vs 403 distinguished in the "nothing there" path | |
| ## Gate 6 β Benchmarks Β· `E-045..E-049` | |
| - [ ] Harness version pinned; task names validated against the harness's own registry, by executing it | |
| - [ ] Smoke pass with `--limit` green: metric keys present, splits correct, installs complete **with deps** | |
| - [ ] Per task: effective shot count + provenance, metric name verbatim, split, `n_input` | |
| - [ ] Failed / missing-metric / no-denominator cells carry a `status` and fail the "all ok" check | |
| - [ ] Incremental saving; completed tasks skipped on re-run only if config matches | |
| - [ ] Run against the published repo, not local weights | |
| - [ ] Task order pre-declared; a partial table never published as a result | |
| - [ ] Table + exact reproduction config pushed into the model repo | |
| ## Gate 7 β Report | |
| - [ ] Measured vs. targeted per task | |
| - [ ] What went wrong, with steps and dates | |
| - [ ] Contamination statement: method, splits used (never test), denominators, exclusions, residual risk | |
| - [ ] Limitations, incl. any mid-run finding that touched the loss landscape and was *not* fixed | |
| - [ ] Which week each number came from, if the schedule spanned a reset | |
| - [ ] Status file set to COMPLETE; re-verify and stop | |
| ## The five recurring failure patterns | |
| Nearly every ledger entry is one of these (`E-001..E-049`). Check for them before checking for anything exotic. | |
| 1. **A claim written from memory instead of from the artefact.** A doc says a function does X; it doesn't. A log | |
| line says a gate passed; the gate read a default. *Fix: assert against the file, the bytes, or the executed | |
| call β never against the note.* | |
| 2. **A guard that cannot fail.** Empty-list denominators, `dict.get(k, 0)`, `.get("ok", True)`, a required-field | |
| check over defaults, a verifier that has never rejected anything. *Fix: every guard gets one negative test | |
| that makes it fire.* | |
| 3. **Verification by listing instead of by bytes.** Commit returned, file in the tree, HTTP 200 on a pointer β | |
| none of them authorise deleting the local copy. | |
| 4. **A check that is blind to its own context.** Guards and logs written rank-0-only in a 2-rank job (`E-039`, | |
| `E-044`); clean rooms that inherit an installed token (`E-045`); "which environment runs this step" never | |
| decided, so a job depends on packages the other image doesn't have (`E-007`, `E-010`, `E-028`, `E-047`). | |
| *Fix: state the executor β rank, container, interpreter β for every assertion and every install.* | |
| 5. **Assuming the platform matches its docs.** A CLI subcommand that isn't; a field that moved; a split that | |
| isn't there; billing on one clock and reporting on another. *Fix: execute before you rely.* | |
| Meta-pattern underneath all five: **the check itself needs a test.** Twice here a fix for a verified bug was | |
| itself wrong (E-046: four of thirteen fixes), and it passed because the check was written by the same reasoning | |
| that wrote the bug. Independent review of the *test*, not just the code, is the cheap way out. | |
| ## Runbook: a session died at 04:00Z and nobody is awake | |
| 1. Read the remote pointer and the commit list. `latest -> N` tells you exactly what survives. Do not read the | |
| job status as the authority. | |
| 2. Compute what is lost: at most one cadence interval (`04` Β§3). Nothing before `N` needs re-running. | |
| 3. Check remaining quota against remaining steps from `N`. Both numbers come from live reads. | |
| 4. If quota covers it: relaunch the same kernel, unchanged. It resumes from the pointer β this path is | |
| idempotent by construction, which is why it was tested. | |
| 5. If quota doesn't: park at the boundary, write the resume instructions in the status file with the step | |
| number and the reset instant, and stop. **Do not change configuration to fit** β that's Β§8's prohibition, | |
| and the pause is a documented outcome. | |
| 6. Only after the run is safe: read the dead session's log for the cause, and write it up in the error ledger | |
| with the symptom that will let the next session recognise it in ten seconds instead of seven attempts. | |