ounce100m-code / ops /SESSION_LOOP.md
Cion-lab's picture
mirror ops/ so the kept code is self-contained
b2ba2a9 verified
|
Raw History Blame Contribute Delete
7.47 kB

SESSION_LOOP β€” unattended Phase 4 session turnover

You are a fresh agent session, probably at an odd hour, with no conversation history. Your only job is to advance the ounce100m main training run by one step of this loop. Do nothing else. Read memory/STATE.md first; it overrides this file if they disagree.

Standing instruction from the user (2026-09-20T14:02Z): run the sleep β†’ wake β†’ check β†’ launch loop across all remaining sessions and do not start any other work unless it is an emergency as defined in Β§5.

1. Check (always do this, ~30 s, no GPU)

curl -s "https://huggingface.co/api/datasets/Cion-lab/ounce100m-ckpt/commits/main"
curl -sL "https://huggingface.co/datasets/Cion-lab/ounce100m-ckpt/resolve/main/latest.json"
curl -sL "https://huggingface.co/datasets/Cion-lab/ounce100m-ckpt/resolve/main/ckpt/checkpoint-<step>/cursor.json"
curl -sL "https://huggingface.co/datasets/Cion-lab/ounce100m-ckpt/resolve/main/ckpt/checkpoint-<step>/trainer_state.json"

Pass criteria, all of them:

Check Expected
head commit latest -> N, and ckpt/checkpoint-N exists 7–10 s earlier
latest.json.step equals N, and path_in_repo is ckpt/checkpoint-N
cursor.json.tokens_consumed exactly N Γ— 262,144
cursor.json.samples_consumed exactly N Γ— 256
dataset_files_sha / shuffle_perm_sha 53df4708526da5c8 / c6f617b96f502d96 β€” unchanged
trainer_state.json global_step == N, max_steps == 3814, save_steps == 127, loss lower than the previous checkpoint

Also read the live quota (get_accelerator_quota β†’ gpu_quota.time_used, total_time_allowed) and the session status for dodosoomro/ounce100m-p4-sessionK.

2. Decide which of the three states you are in

  • A session is still running and the pointer is advancing on the 127 grid β†’ nothing to do. Report the step and stop. Do not launch anything.
  • A session ended and latest -> 762 Γ— k for k = 1..4, k < 5 β†’ launch session k+1 (Β§3).
  • latest -> 3814 β†’ the run is finished (Β§4).

Session grid: S1 0β†’762, S2 β†’1,524, S3 β†’2,286, S4 β†’3,048, S5 β†’3,814. A session that stopped short of its boundary (platform kill) is not an emergency β€” the next session resumes from wherever latest points, because phase4_session.py plans from the pointer. Launch the next session with the remaining step count it would need; if the planner's own arithmetic disagrees, that is an emergency (Β§5).

3. Launch the next session

The one editable line. code/kernels/phase4_bootstrap.py is the whole kernel body; it pins REV = "c237b478076f68da43dabd3403daa5639286899e" and LAUNCHER_SHA = "5fb14bc3…" and fetches kernels/phase4_session.py at that REV, sha256-asserting it. Only line 13 changes per session: os.environ["QUOTA_LEFT_HOURS"] = "<live hours from the Β§1 read>". Nothing else β€” it resumes from latest.json and needs no other input.

  1. Book the ledger row first (memory/QUOTA.md, next row number, before the job starts). Β§8 forbids an unbooked GPU job. State the live hours remaining and the CPU-alternative check (Β§3.7: there is none β€” the quantity is 762 steps of real T4 training).

  2. Confirm session K has actually exited before launching K+1 β€” two concurrent GPU sessions would double-bill and the second hits the stale-stop guard. Use get_notebook_session_status, not the Hub pointer, as the liveness answer (time_reserved in the quota read is always 0s and means nothing).

  3. Submit with mcp__kaggle__save_notebook, kernelExecutionType: "SaveAndRunAll" (this creates and runs β€” there is no separate run call; create_notebook_session takes only a slug and a machine shape, so it is not how a script is uploaded). Send the full verified envelope every time; a missing language fails with a message-less error (E-020):

    slug      = "dodosoomro/ounce100m-p4-sessionK"
    text      = <contents of code/kernels/phase4_bootstrap.py, line 13 updated>
    language  = "python"
    kernelType = "script"
    machineShape = "GPU"          # + enableGpu: true β‡’ 2Γ—T4. Do not invent a T4-specific shape string.
    kernelExecutionType = "SaveAndRunAll"
    enableGpu = true
    isPrivate = true
    enableInternet = true
    sessionTimeoutSeconds = 25200
    

    newTitle is required for a kernel that does not exist yet β€” session 2 was refused with "New kernels must have a title specified" until it was sent. For a new kernel set it equal to the slug suffix so the returned ref is the slug you asked for; for an existing kernel never send a title that differs from its current suffix, because newTitle rewrites the slug (E-020) and breaks every reference. Always poll the returned ref, never the requested slug.

  4. Poll the returned slug, never the requested one, and record slug + version + kernel id + submission time in ASSETS.md. Expect ~60-90 s of queue and boot before anything happens; don't poll sooner than 60 s. Note that get_notebook_info may be denied by API scope β€” liveness comes from the quota delta: read get_accelerator_quota twice a few minutes apart and see whether time_used is advancing (it is, at 1Γ— wall clock, from the moment the container starts).

  5. Confirm within ~3 minutes that it did not exit fast. An exit inside ~2 min is a REFUSED_TO_START verdict (free space / pointer read / manifest arithmetic) β€” read the log, and treat it as Β§5 only if the reason is not self-evidently transient.

  6. Do not babysit it. Write "next check at

4. When latest -> 3814

  1. Verify all six final checkpoints and final/ exist and are byte-verified (compare oids against the local manifest the run wrote).
  2. Run Phase 5 on CPU (code/build/publish_model.py) β€” 0 GPU hours. Gate 5's clean-room check must pop HF_TOKEN, not just pass token=False (E-045/E-046).
  3. Then Phase 6 per docs/05-eval-plan.md, tasks in constitution order, incrementally saved, smoke --limit 5 first, installed with deps. Spend whatever GPU hours remain this week; the rest waits for the reset (D-019). Never publish a partial eight-task table as a result.
  4. Fill the four PENDING blocks in docs/final-report.md, set STATE.md to COMPLETE, and stop. Do not start new work after that.

5. Emergencies β€” the only reasons to do more than the above

  • The pointer moved backwards, or a cursor hash changed, or tokens != step Γ— 262,144. Stop everything; this is data corruption. Do not launch another session. Write the evidence into ERRORS.md and STATE.md as a blocker and surface it to the user.
  • Two consecutive launches that exit inside ~2 min for the same reason.
  • Quota cannot fund a full remaining session. Then park: do not change configuration to fit (Β§8). Write the resume step, the reset instant (2026-09-26T00:00Z, a Saturday) and the exact next command into STATE.md.
  • Loss discontinuity beyond noise at a checkpoint boundary, or grad-norm moving with it.

Anything else is not an emergency. Refactoring, "improving throughput", re-auditing settled decisions, and re-running completed probes are all explicitly out of scope for these turns.