|
Download guide/06-checklists.md from Cion-lab/ounce100m-code: direct link, hf CLI and curl.
- Browser
- Download file 8.89 kB
-
https://huggingface.co/Cion-lab/ounce100m-code/resolve/main/guide/06-checklists.md
- Command line
-
hf download hf://Cion-lab/ounce100m-code/guide/06-checklists.md
-
curl -L -o 06-checklists.md https://huggingface.co/Cion-lab/ounce100m-code/resolve/main/guide/06-checklists.md
8.89 kB
06 — Checklists
The working page. Tick nothing from memory. Each gate names the error/decision records it came from, so a claim here is checkable against the ledger rather than taken on trust.
Gate 0 — Orientation · E-001..E-019
- Credentials verified by a call returning identity, not by a file existing
- A secret store's privacy proved by two different calls (authenticated listing and anonymous read) —
403/404 alone cannot distinguish "private" from "never created" (
E-011) - Remote quota read live; hours remaining + reset instant written in the ledger (look up the weekday)
- One throwaway CPU job and one throwaway GPU job completed end to end; mechanics recorded
- Any async job from a previous session found and accounted for (running / finished / dead)
- Memory files match reality; every discrepancy logged, not silently corrected
Gate 1 — Plan · E-020..E-026
- Architecture, hyperparameters, token total, step count written and dated
- Throughput estimate with its source, and the GPU-hour arithmetic it implies
- Budget has a line for every remaining phase, including publish and benchmarks (
D-019) - Per-benchmark target bands pre-registered, each with a citation
- Eval split + metric named per task before any eval code exists
- What you may change mid-run vs what is frozen, stated as one sentence
- Every choice traceable to a source you actually looked at
Gate 2 — Data · E-027..E-035
- Assembled above target; validation split reserved before sharding
- Deduplicated; mechanical overlap audit at
overlap_total 0 - Audit touched
validation/devonly; notestsplit opened - Audit denominators (rows per task) asserted nonzero and printed beside the counts
- No benchmark item text in any published audit artefact, including exception strings
- Shard layout proven small enough that a job reads a few shards, not the mix
- Packing boundary choice measured, not assumed (seam rate)
- Published public; token count re-verified from the Hub copy against the manifest
- Manifest self-consistent (sums, counts, no empties); per-file + set fingerprints recorded
- Cursor design supports exact positional resume; fingerprint keys chosen and persisted
Gate 3 — Preflight · E-036..E-044 · D-012..D-018
- Rehearsal re-run after every edit to run code — it's free
- Rehearsal imports the real modules and parses the real launch command line
- Each proof run at the size the property depends on (
03§3), not at a token tiny model - Fault injection: wrong bytes, ignored
Range, unverified prune, absent repo — all rejected - Guards tested on every rank; diagnostics printed on every rank, not just rank 0
- Upload → restart → pull-back, on the Hub
- Exact data position proven by index/hash and token offset
- Optimizer/LR state restored, values printed; loss continuous across restart
- Cold resume on a fresh instance, empty disk, recovering solely from the Hub
- Storage cycle end to end at representative size (weights + optimizer state), free space logged
- Two sequential resumes, no backward drift, no re-read data
- Sustained tokens/sec over a long leg (not 20 steps); memory as a time series, cross-rank max
- Checkpoint-cycle cost measured in seconds per push
- Report: the eight proofs (
03§4) each with evidence + ETA + GPU-hour arithmetic that fits the quota
Gate 4 — Run · E-042, D-018
- Config frozen; hash recorded and asserted at every session start
- Cadence divides session length; session tiling ends exactly on the horizon
- Per checkpoint: write → push → verify bytes → point → read pointer back → prune
- Prune gated by the verification result inside the prune function
- Free space logged beside loss, every cycle
- Ledger row booked before each GPU job; hours remaining re-read at each launch
-
tokens == steps × tokens_per_stepand cursor sample math checked at each checkpoint - Loss anomalies documented with grad-norm and LR context, config unchanged
- Results written before trailing cross-checks; verification cannot destroy a completed report (
E-012) - Multi-week plan written if the estimate exceeds remaining hours
Gate 5 — Publish · E-045, E-046
- Card fields generated from artefacts; required-field guard demonstrably fails on an empty run
- Parameter/vocab/geometry assertions recomputed from the pushed files
- Clean-room
from_pretrained+ generate: env set before import,token=False, and every token var popped from the environment - Anonymous re-hash of every published file
- 404 vs timeout vs 403 distinguished in the "nothing there" path
Gate 6 — Benchmarks · E-045..E-049
- Harness version pinned; task names validated against the harness's own registry, by executing it
- Smoke pass with
--limitgreen: metric keys present, splits correct, installs complete with deps - Per task: effective shot count + provenance, metric name verbatim, split,
n_input - Failed / missing-metric / no-denominator cells carry a
statusand fail the "all ok" check - Incremental saving; completed tasks skipped on re-run only if config matches
- Run against the published repo, not local weights
- Task order pre-declared; a partial table never published as a result
- Table + exact reproduction config pushed into the model repo
Gate 7 — Report
- Measured vs. targeted per task
- What went wrong, with steps and dates
- Contamination statement: method, splits used (never test), denominators, exclusions, residual risk
- Limitations, incl. any mid-run finding that touched the loss landscape and was not fixed
- Which week each number came from, if the schedule spanned a reset
- Status file set to COMPLETE; re-verify and stop
The five recurring failure patterns
Nearly every ledger entry is one of these (E-001..E-049). Check for them before checking for anything exotic.
- A claim written from memory instead of from the artefact. A doc says a function does X; it doesn't. A log line says a gate passed; the gate read a default. Fix: assert against the file, the bytes, or the executed call — never against the note.
- A guard that cannot fail. Empty-list denominators,
dict.get(k, 0),.get("ok", True), a required-field check over defaults, a verifier that has never rejected anything. Fix: every guard gets one negative test that makes it fire. - Verification by listing instead of by bytes. Commit returned, file in the tree, HTTP 200 on a pointer — none of them authorise deleting the local copy.
- A check that is blind to its own context. Guards and logs written rank-0-only in a 2-rank job (
E-039,E-044); clean rooms that inherit an installed token (E-045); "which environment runs this step" never decided, so a job depends on packages the other image doesn't have (E-007,E-010,E-028,E-047). Fix: state the executor — rank, container, interpreter — for every assertion and every install. - Assuming the platform matches its docs. A CLI subcommand that isn't; a field that moved; a split that isn't there; billing on one clock and reporting on another. Fix: execute before you rely.
Meta-pattern underneath all five: the check itself needs a test. Twice here a fix for a verified bug was itself wrong (E-046: four of thirteen fixes), and it passed because the check was written by the same reasoning that wrote the bug. Independent review of the test, not just the code, is the cheap way out.
Runbook: a session died at 04:00Z and nobody is awake
- Read the remote pointer and the commit list.
latest -> Ntells you exactly what survives. Do not read the job status as the authority. - Compute what is lost: at most one cadence interval (
04§3). Nothing beforeNneeds re-running. - Check remaining quota against remaining steps from
N. Both numbers come from live reads. - If quota covers it: relaunch the same kernel, unchanged. It resumes from the pointer — this path is idempotent by construction, which is why it was tested.
- If quota doesn't: park at the boundary, write the resume instructions in the status file with the step number and the reset instant, and stop. Do not change configuration to fit — that's §8's prohibition, and the pause is a documented outcome.
- Only after the run is safe: read the dead session's log for the cause, and write it up in the error ledger with the symptom that will let the next session recognise it in ten seconds instead of seven attempts.