ounce100m-code / ops /SESSION_LOOP.md
Cion-lab's picture
mirror ops/ so the kept code is self-contained
b2ba2a9 verified
|
Raw History Blame Contribute Delete
7.47 kB
# SESSION_LOOP β€” unattended Phase 4 session turnover
You are a fresh agent session, probably at an odd hour, with no conversation history. Your only job is to
advance the ounce100m main training run by **one** step of this loop. Do nothing else. Read
`memory/STATE.md` first; it overrides this file if they disagree.
**Standing instruction from the user (2026-09-20T14:02Z):** run the sleep β†’ wake β†’ check β†’ launch loop across
all remaining sessions and do not start any other work unless it is an emergency as defined in Β§5.
## 1. Check (always do this, ~30 s, no GPU)
```
curl -s "https://huggingface.co/api/datasets/Cion-lab/ounce100m-ckpt/commits/main"
curl -sL "https://huggingface.co/datasets/Cion-lab/ounce100m-ckpt/resolve/main/latest.json"
curl -sL "https://huggingface.co/datasets/Cion-lab/ounce100m-ckpt/resolve/main/ckpt/checkpoint-<step>/cursor.json"
curl -sL "https://huggingface.co/datasets/Cion-lab/ounce100m-ckpt/resolve/main/ckpt/checkpoint-<step>/trainer_state.json"
```
Pass criteria, all of them:
| Check | Expected |
|---|---|
| head commit | `latest -> N`, and `ckpt/checkpoint-N` exists 7–10 s earlier |
| `latest.json.step` | equals `N`, and `path_in_repo` is `ckpt/checkpoint-N` |
| `cursor.json.tokens_consumed` | exactly `N Γ— 262,144` |
| `cursor.json.samples_consumed` | exactly `N Γ— 256` |
| `dataset_files_sha` / `shuffle_perm_sha` | `53df4708526da5c8` / `c6f617b96f502d96` β€” **unchanged** |
| `trainer_state.json` | `global_step == N`, `max_steps == 3814`, `save_steps == 127`, loss lower than the previous checkpoint |
Also read the live quota (`get_accelerator_quota` β†’ `gpu_quota.time_used`, `total_time_allowed`) and the session
status for `dodosoomro/ounce100m-p4-sessionK`.
## 2. Decide which of the three states you are in
- **A session is still running** and the pointer is advancing on the 127 grid β†’ nothing to do. Report the step
and stop. Do not launch anything.
- **A session ended and `latest -> 762 Γ— k`** for k = 1..4, k < 5 β†’ launch session k+1 (Β§3).
- **`latest -> 3814`** β†’ the run is finished (Β§4).
Session grid: S1 0β†’762, S2 β†’1,524, S3 β†’2,286, S4 β†’3,048, S5 β†’3,814. A session that stopped short of its
boundary (platform kill) is **not** an emergency β€” the next session resumes from wherever `latest` points,
because `phase4_session.py` plans from the pointer. Launch the next session with the *remaining* step count it
would need; if the planner's own arithmetic disagrees, that is an emergency (Β§5).
## 3. Launch the next session
**The one editable line.** `code/kernels/phase4_bootstrap.py` is the whole kernel body; it pins
`REV = "c237b478076f68da43dabd3403daa5639286899e"` and `LAUNCHER_SHA = "5fb14bc3…"` and fetches
`kernels/phase4_session.py` at that REV, sha256-asserting it. **Only line 13 changes per session:**
`os.environ["QUOTA_LEFT_HOURS"] = "<live hours from the Β§1 read>"`. Nothing else β€” it resumes from
`latest.json` and needs no other input.
1. **Book the ledger row first** (`memory/QUOTA.md`, next row number, before the job starts). Β§8 forbids an
unbooked GPU job. State the live hours remaining and the CPU-alternative check (Β§3.7: there is none β€” the
quantity is 762 steps of real T4 training).
2. **Confirm session K has actually exited** before launching K+1 β€” two concurrent GPU sessions would
double-bill and the second hits the stale-stop guard. Use `get_notebook_session_status`, not the Hub
pointer, as the liveness answer (`time_reserved` in the quota read is always `0s` and means nothing).
3. Submit with `mcp__kaggle__save_notebook`, `kernelExecutionType: "SaveAndRunAll"` (this creates *and* runs β€”
there is no separate run call; `create_notebook_session` takes only a slug and a machine shape, so it is not
how a script is uploaded). Send the **full verified envelope** every time; a missing `language` fails with a
message-less error (E-020):
```
slug = "dodosoomro/ounce100m-p4-sessionK"
text = <contents of code/kernels/phase4_bootstrap.py, line 13 updated>
language = "python"
kernelType = "script"
machineShape = "GPU" # + enableGpu: true β‡’ 2Γ—T4. Do not invent a T4-specific shape string.
kernelExecutionType = "SaveAndRunAll"
enableGpu = true
isPrivate = true
enableInternet = true
sessionTimeoutSeconds = 25200
```
`newTitle` is **required for a kernel that does not exist yet** β€” session 2 was refused with
"New kernels must have a title specified" until it was sent. For a new kernel set it **equal to the slug
suffix** so the returned `ref` is the slug you asked for; for an **existing** kernel never send a title that
differs from its current suffix, because `newTitle` rewrites the slug (E-020) and breaks every reference.
Always poll the returned `ref`, never the requested slug.
4. **Poll the returned slug, never the requested one**, and record slug + version + kernel id + submission
time in `ASSETS.md`. Expect ~60-90 s of queue and boot before anything happens; don't poll sooner than 60 s.
Note that `get_notebook_info` may be denied by API scope β€” **liveness comes from the quota delta**: read
`get_accelerator_quota` twice a few minutes apart and see whether `time_used` is advancing (it is, at 1Γ—
wall clock, from the moment the container starts).
5. Confirm within ~3 minutes that it did **not** exit fast. An exit inside ~2 min is a `REFUSED_TO_START`
verdict (free space / pointer read / manifest arithmetic) β€” read the log, and treat it as Β§5 only if the
reason is not self-evidently transient.
6. Do **not** babysit it. Write "next check at <time>" into `STATE.md` and stop, or arm one long-lived waiter
(`watch_ckpt.sh <boundary_step> <seconds>`, or a read-only watcher agent for a whole session leg).
## 4. When `latest -> 3814`
1. Verify all six final checkpoints and `final/` exist and are byte-verified (compare oids against the local
manifest the run wrote).
2. Run **Phase 5 on CPU** (`code/build/publish_model.py`) β€” 0 GPU hours. Gate 5's clean-room check must pop
`HF_TOKEN`, not just pass `token=False` (E-045/E-046).
3. Then Phase 6 per `docs/05-eval-plan.md`, tasks in constitution order, incrementally saved, smoke `--limit 5`
first, installed **with** deps. Spend whatever GPU hours remain this week; the rest waits for the reset
(D-019). **Never publish a partial eight-task table as a result.**
4. Fill the four PENDING blocks in `docs/final-report.md`, set `STATE.md` to COMPLETE, and stop. Do not start
new work after that.
## 5. Emergencies β€” the only reasons to do more than the above
- The pointer **moved backwards**, or a cursor hash changed, or `tokens != step Γ— 262,144`. Stop everything;
this is data corruption. Do not launch another session. Write the evidence into `ERRORS.md` and `STATE.md`
as a blocker and surface it to the user.
- Two consecutive launches that exit inside ~2 min for the same reason.
- Quota cannot fund a full remaining session. Then **park**: do not change configuration to fit (Β§8). Write the
resume step, the reset instant (2026-09-26T00:00Z, a Saturday) and the exact next command into `STATE.md`.
- Loss discontinuity beyond noise at a checkpoint boundary, or grad-norm moving with it.
Anything else is not an emergency. Refactoring, "improving throughput", re-auditing settled decisions, and
re-running completed probes are all explicitly out of scope for these turns.