# SESSION_LOOP — unattended Phase 4 session turnover You are a fresh agent session, probably at an odd hour, with no conversation history. Your only job is to advance the ounce100m main training run by **one** step of this loop. Do nothing else. Read `memory/STATE.md` first; it overrides this file if they disagree. **Standing instruction from the user (2026-09-20T14:02Z):** run the sleep → wake → check → launch loop across all remaining sessions and do not start any other work unless it is an emergency as defined in §5. ## 1. Check (always do this, ~30 s, no GPU) ``` curl -s "https://huggingface.co/api/datasets/Cion-lab/ounce100m-ckpt/commits/main" curl -sL "https://huggingface.co/datasets/Cion-lab/ounce100m-ckpt/resolve/main/latest.json" curl -sL "https://huggingface.co/datasets/Cion-lab/ounce100m-ckpt/resolve/main/ckpt/checkpoint-/cursor.json" curl -sL "https://huggingface.co/datasets/Cion-lab/ounce100m-ckpt/resolve/main/ckpt/checkpoint-/trainer_state.json" ``` Pass criteria, all of them: | Check | Expected | |---|---| | head commit | `latest -> N`, and `ckpt/checkpoint-N` exists 7–10 s earlier | | `latest.json.step` | equals `N`, and `path_in_repo` is `ckpt/checkpoint-N` | | `cursor.json.tokens_consumed` | exactly `N × 262,144` | | `cursor.json.samples_consumed` | exactly `N × 256` | | `dataset_files_sha` / `shuffle_perm_sha` | `53df4708526da5c8` / `c6f617b96f502d96` — **unchanged** | | `trainer_state.json` | `global_step == N`, `max_steps == 3814`, `save_steps == 127`, loss lower than the previous checkpoint | Also read the live quota (`get_accelerator_quota` → `gpu_quota.time_used`, `total_time_allowed`) and the session status for `dodosoomro/ounce100m-p4-sessionK`. ## 2. Decide which of the three states you are in - **A session is still running** and the pointer is advancing on the 127 grid → nothing to do. Report the step and stop. Do not launch anything. - **A session ended and `latest -> 762 × k`** for k = 1..4, k < 5 → launch session k+1 (§3). - **`latest -> 3814`** → the run is finished (§4). Session grid: S1 0→762, S2 →1,524, S3 →2,286, S4 →3,048, S5 →3,814. A session that stopped short of its boundary (platform kill) is **not** an emergency — the next session resumes from wherever `latest` points, because `phase4_session.py` plans from the pointer. Launch the next session with the *remaining* step count it would need; if the planner's own arithmetic disagrees, that is an emergency (§5). ## 3. Launch the next session **The one editable line.** `code/kernels/phase4_bootstrap.py` is the whole kernel body; it pins `REV = "c237b478076f68da43dabd3403daa5639286899e"` and `LAUNCHER_SHA = "5fb14bc3…"` and fetches `kernels/phase4_session.py` at that REV, sha256-asserting it. **Only line 13 changes per session:** `os.environ["QUOTA_LEFT_HOURS"] = ""`. Nothing else — it resumes from `latest.json` and needs no other input. 1. **Book the ledger row first** (`memory/QUOTA.md`, next row number, before the job starts). §8 forbids an unbooked GPU job. State the live hours remaining and the CPU-alternative check (§3.7: there is none — the quantity is 762 steps of real T4 training). 2. **Confirm session K has actually exited** before launching K+1 — two concurrent GPU sessions would double-bill and the second hits the stale-stop guard. Use `get_notebook_session_status`, not the Hub pointer, as the liveness answer (`time_reserved` in the quota read is always `0s` and means nothing). 3. Submit with `mcp__kaggle__save_notebook`, `kernelExecutionType: "SaveAndRunAll"` (this creates *and* runs — there is no separate run call; `create_notebook_session` takes only a slug and a machine shape, so it is not how a script is uploaded). Send the **full verified envelope** every time; a missing `language` fails with a message-less error (E-020): ``` slug = "dodosoomro/ounce100m-p4-sessionK" text = language = "python" kernelType = "script" machineShape = "GPU" # + enableGpu: true ⇒ 2×T4. Do not invent a T4-specific shape string. kernelExecutionType = "SaveAndRunAll" enableGpu = true isPrivate = true enableInternet = true sessionTimeoutSeconds = 25200 ``` `newTitle` is **required for a kernel that does not exist yet** — session 2 was refused with "New kernels must have a title specified" until it was sent. For a new kernel set it **equal to the slug suffix** so the returned `ref` is the slug you asked for; for an **existing** kernel never send a title that differs from its current suffix, because `newTitle` rewrites the slug (E-020) and breaks every reference. Always poll the returned `ref`, never the requested slug. 4. **Poll the returned slug, never the requested one**, and record slug + version + kernel id + submission time in `ASSETS.md`. Expect ~60-90 s of queue and boot before anything happens; don't poll sooner than 60 s. Note that `get_notebook_info` may be denied by API scope — **liveness comes from the quota delta**: read `get_accelerator_quota` twice a few minutes apart and see whether `time_used` is advancing (it is, at 1× wall clock, from the moment the container starts). 5. Confirm within ~3 minutes that it did **not** exit fast. An exit inside ~2 min is a `REFUSED_TO_START` verdict (free space / pointer read / manifest arithmetic) — read the log, and treat it as §5 only if the reason is not self-evidently transient. 6. Do **not** babysit it. Write "next check at