stackcraft-clef-flash-lora / code /docs /tutorials /04-gpu-feasibility.md
nima1's picture
Publish verified checkpoint and losslessly compressed study evidence
4be6a52 verified
|
Raw History Blame Contribute Delete
16.4 kB
# Tutorial 04 — Prove training and reload before running an experiment
This milestone asks a narrow engineering question: can the real pinned Clef-flash
model take gradient steps on this GPU, save all learned parameters, and reproduce
its probabilities in a fresh process? It does not ask whether the model plays
better. The latter requires complete games on the held-out sequences in Milestone 5.
**The bounded real-weight training and fresh-process reload gate passed.** Both
head-only and LoRA-plus-head artifacts reproduced their reference probabilities
exactly. The first LoRA reload exposed a saved-configuration compatibility issue;
the corrected loader and successful retry are documented below. This establishes
local training feasibility, not improved game performance.
## Start from the native decision model
The base is `Cloudflare/clef-flash`, revision
`17f0b0ad64efb65d273590632833508766b2aae6`. The local loader verifies the pinned
`joint_schema_model.py` source hash and loads an explicitly resolved snapshot.
It requires `trust_pinned_code=True`: pinning and inspecting Python source controls
which code executes; it does not turn that code into a sandbox.
An observation becomes one native choice question containing every legal
placement. The answer is a distribution over action IDs, not a generated command
or natural-language explanation. `encode_observation` checks the complete token
budget before native encoding, because silently truncating a board would change
the task. Training uses only the visible board, current piece, one preview and
legal placements. Dataset seed and episode metadata remain outside the input.
The native encoder sorts option IDs lexicographically. `decision_loss` locates the
teacher's action in `encoded.questions[0].option_ids`; it must not reuse the
engine's rotation/column index as a target index. Both orders contain the same
legal moves, but their order can differ.
Do not train through `systemone()` or `ClefPlayer.choose()`: those are inference
paths. The probe calls `native.collate_records` and then `model(batch)` with
normal autograd enabled. `choose()` is appropriate for the before/after reference
probabilities, where gradients are unnecessary.
## Keep the backbone small in memory and the head stable
`src/stackcraft/training.py` implements two bounded configurations:
| Configuration | Updated parameters | Purpose |
| --- | --- | --- |
| Head only | Native joint decision head | Isolate head training and checkpoint handling before a backbone backward pass. |
| Rank-4 LoRA plus head | Small added text-layer matrices and the native head | Check whether the task can adapt both text representations and decision scoring. |
The backbone remains BF16. The native decision head is converted to FP32, and
`FP32DecisionHead` disables autocast inside the head. Its floating inputs are
converted to FP32 while token IDs and masks retain their original types. This
keeps the trainable scoring operations in full precision without converting the
whole multi-billion-parameter backbone.
The head reads selected rows of the output vocabulary embedding. Converting the
entire vocabulary matrix to FP32 first creates a large temporary allocation
(roughly 3.79 GiB for this release). `GatheredFloat32Embedding` instead implements
`weight[indices].float()`: select the needed rows first, then cast. This preserves
the actual native head computation and autograd connections. A tiny actual-native-
head test compares outputs and gradients with the full-cast reference exactly.
That test checks the optimization, not full-model training feasibility.
LoRA targets are full module paths under the text transformer layers. They include
both ordinary attention projections and Qwen3.5's linear-attention projections,
as well as the MLP projections. Matching only suffixes such as `q_proj` could
accidentally select the vision encoder; matching only ordinary attention would
miss the hybrid backbone. The vision tower, vocabulary output layer and unrelated
modules are excluded. Original backbone parameters are frozen. The probe uses
rank 4, alpha 8, no LoRA dropout and no trained backbone biases.
Non-reentrant gradient checkpointing recomputes intermediate activations during
backward instead of retaining all of them. This trades time for memory while
preserving gradients through frozen parts of the backbone into LoRA matrices.
A tiny actual Qwen3.5 hybrid-backbone test checks nonzero LoRA gradients in ordinary
attention, linear attention and MLP layers. These are stronger plumbing checks
than testing a stand-in linear model alone, but still do not establish the real
9B GPU memory requirement.
We start with BF16 rather than immediately adding 4-bit quantized training. The
native model has a custom head and loader; another quantization layer would add
an unverified compatibility dependency. If BF16 fails, record the failure and
make a bounded next decision rather than silently changing the experiment.
## Prepare before borrowing the GPU
Run these commands from the repository root or the public bundle's `code/`
directory. Install the ML tools and cache the pinned release before testing:
```bash
uv sync --locked --extra ml --group dev
uv run --locked --extra ml hf download Cloudflare/clef-flash \
--revision 17f0b0ad64efb65d273590632833508766b2aae6
```
Downloading only populates the cache; it does not load GPU weights. It can run
while other GPU services remain available. The actual native-head tests need
this cached source. Explicitly enable the native tokenizer/encoder check too:
```bash
STACKCRAFT_TEST_NATIVE_ENCODING=1 uv run --locked --extra ml \
pytest tests/test_training.py tests/test_clef.py
```
These checks use CPU fixtures and the real tokenizer/source; they do not load
the full backbone or establish GPU feasibility. Inspect the test summary: a
skipped test is not a passed native-model test.
The training probe expects the audited `data/study-v1` bundle established by
Tutorial03's exact-study reconstruction or a verified dataset download. If you
used a different directory, pass that same path with `--dataset` below.
It validates the train/validation bundle, then uses only the first four training
positions. It does not sample held-out test seeds. Record the current source
commit and any uncommitted probe changes in `record.md` before running.
This workstation normally runs a local GLM Docker service. The user explicitly
authorized temporarily stopping that service and restoring it after these
Stackcraft experiments. `scripts/gpu_session.py` is specific to that approved
service; it is not a generic instruction to stop another machine's workloads.
It first requires the configured container to be running, then stops it, runs a
bounded child job, and attempts restoration during cleanup. The training script
itself never stops services. It refuses to load unless CUDA is available and at
least **25 GiB is free**. Free-memory admission is a precondition, not a guarantee
that backward will fit.
## Run the bounded real-weight probe
The initial budget is two configurations and at most 30 minutes of child-job
time. The commands below allocate 20 minutes to the training probe and 10 minutes
to the separate reload check. Choose fresh output and session-record paths on a
retry; neither existing evidence nor checkpoints should be overwritten.
```bash
uv run --locked --extra ml python scripts/gpu_session.py \
--timeout 1200 --record runs/gpu-training-probe-v1.json -- \
uv run --locked --extra ml python scripts/probe_clef_training.py \
--dataset data/study-v1 --output runs/clef-training-probe-v1
```
The probe first records unchanged native probabilities. It measures the initial
probability drift introduced by moving the head to FP32, before training; a
precision change must not be mistaken for a learned improvement. It then runs
three head-only steps. After saving that checkpoint, it releases the model and
loads a fresh unchanged base for five LoRA-plus-head steps. The LoRA experiment
therefore does not inherit the head-only warm-up.
Both configurations use batch size one, AdamW with learning rate `1e-5`, norm
clipping at 1.0, and this loss:
smoothed_cross_entropy(label_smoothing=0.05)
+ 0.1 * sum_over_options((probability - one_hot_teacher_label)**2)
The cross-entropy term learns the expert action. Smoothing avoids assigning the
entire training target to one action. The Brier term also penalizes the predicted
probability distribution's distance from the teacher target. These fixed choices
are a probe configuration, not evidence of calibrated confidence or optimal
hyperparameters.
For every step, the probe checks finite loss and gradients and records nonzero
head/LoRA gradient sums, tokens, elapsed time, and CUDA peak allocated/reserved
bytes. Before/after parameter hashes check that intended parameters changed and
frozen parameters did not. Hashing the entire frozen backbone costs time but
checks something that `requires_grad=False` alone cannot prove: what bytes actually
changed during this run.
Inspect `runs/clef-training-probe-v1/report.json` and the session record. A timeout,
CUDA error, absent gradient or failed hash check is a failed gate. A report left at
`running` after forced termination is incomplete evidence, even if some checkpoint
files exist. Do not start long training merely because loss fell on four examples.
## Save the head as well as the adapter
A LoRA checkpoint has this structure:
```text
checkpoint/
joint_head.safetensors
adapter/
adapter_config.json
adapter_model.safetensors
training_config.json
reference.json
```
The head is an independently trained part of Clef; saving only the PEFT adapter
would omit learned parameters. `joint_head.safetensors` stores FP32 native head
keys without wrapper-specific prefixes. The adapter stores only the added LoRA
weights, not another copy of the frozen base. Metadata identifies the base
revision, native source hash, input encoding, head format, LoRA settings and
teacher rows. The loader checks this contract, head shapes and dtypes, and actual
saved adapter configuration before restoring weights.
The head-only probe writes `head-checkpoint/` without an adapter. It is diagnostic
output, not a selected release model.
## Require a separate fresh-process reload
Only after the first probe succeeds, run:
```bash
uv run --locked --extra ml python scripts/gpu_session.py \
--timeout 600 --record runs/gpu-training-reload-v1.json -- \
uv run --locked --extra ml python scripts/probe_clef_training.py \
--dataset data/study-v1 --output runs/clef-training-reload-v1 \
--reload runs/clef-training-probe-v1/checkpoint
```
This starts another Python process, loads the unchanged pinned base, restores both
adapter and head, and evaluates the same four training observations. Require the
same option IDs and a **maximum absolute probability difference of at most
`1e-4`** across all recorded options. The tolerance was declared before the run.
Compare unrounded probabilities; rounded display values can conceal a mismatch.
A same-process test could accidentally reuse trained parameters that were never
saved. The fresh process tests the actual inference artifact. A passing reload
establishes serialization fidelity on these references, not general game skill
or correctness on every possible input.
## Restore service and record the limits
Inspect both `restored_running` and `restored_healthy` after success and failure.
The wrapper checks `/health` inside the restored container and allows up to 120
seconds for readiness, retrying transient timeouts and malformed startup responses.
It does not query the unrelated service on host port 8080. A running process alone
would not establish readiness. Nested cleanup ensures a child-termination error
still reaches restoration; repeated ordinary interrupts are ignored during that
cleanup. Preserve the wrapper record, probe reports and console logs, including
errors. Mocked tests exercise these recovery paths without touching Docker.
Process cleanup cannot guarantee recovery after `SIGKILL`, host power loss or a
Docker daemon failure. If the wrapper was forcibly terminated, inspect the exact
approved container and any remaining job before acting. Restore the existing
service only after the experiment has released GPU memory; do not start another
probe simply because the last terminal stopped printing output.
The earlier M2 real-weight development pilot is useful context: on this RTX 5090,
it recorded 55 unchanged-Clef decisions at a mean of about **0.144 seconds per
decision**, with **19,686,539,776 bytes** peak CUDA allocation (about 19.7 GB,
18.3 GiB). The benchmark script took 12.60 seconds; the stop/job/restore wrapper
interval was about 14.20 seconds and recorded successful restoration. These are
inference measurements on two short development games, not training memory or
held-out quality results. Sources are `runs/clef-development-v1.json` and
`runs/gpu-baseline-session.json`.
## Measured training probe and the first reload failure
The real-weight probe passed on 2026-10-05 using the RTX 5090 and Torch
`2.14.1+cu130`. Evidence is `runs/probe-v1/report.json`; the actual run directory
differs from the fresh reproduction paths shown above. It used training positions
`seed-10000-turn-0` through `seed-10000-turn-3`, with 1,302–2,252 encoded tokens.
| Measurement | Head-only probe | LoRA-plus-head probe |
| --- | ---: | ---: |
| Optimization steps | 3 | 5 |
| Observed step times | 0.154–0.328 s | 0.670–1.322 s |
| Peak allocated CUDA bytes | 21,323,760,640 | 23,180,389,888 |
| Peak reserved CUDA bytes | 21,541,945,344 | 23,595,057,152 |
| Intended parameter changes | Verified | Verified |
| Frozen parameter hashes | Unchanged | Unchanged |
All reported losses and gradients were finite, with nonzero gradients in the
intended groups. Total probe time was **51.81 seconds**, including model reload,
parameter hashing and checkpoint work. This is not an estimate of a full epoch:
only four short training positions were exercised. Peak reserved memory is the
PyTorch allocator's reservation, not the same quantity as all memory reported by
`nvidia-smi`.
The FP32 head changed initial probabilities by at most **0.00074412** before any
training. That is a precision effect relative to the original BF16 head, not a
failed serialization comparison. The `1e-4` reload threshold compares the saved
trained FP32-head model with its own fresh reload.
The first fresh reload, `runs/probe-reload-v1/report.json`, failed after 4.15
seconds because the saved adapter configuration differed from our metadata. PEFT
automatically shortened **248 full target paths to 12 module suffixes**. The strict
loader correctly refused a mismatch; the failure does not by itself mean the
weight tensors are corrupt. Accepting suffixes without checking what they match
could attach an adapter to unintended modules. The correction must expand saved
target selectors against the actual pinned backbone and require exactly the
intended full target set, rejecting any extra destinations. This is why fresh
reload is a milestone gate rather than a final packaging chore.
The corrected loader resolves saved selectors against the pinned backbone and
requires the exact intended destinations. The retry preserved the original
checkpoint and failed report; it changed validation of equivalent target selectors,
not the trained weights. A regression test covers target shortening and rejects
selectors that also match unintended modules.
Both independent fresh-process checks then passed:
| Artifact | Report | Maximum probability difference | Process time |
| --- | --- | ---: | ---: |
| Head only | `runs/probe-head-reload-v1/report.json` | 0.0 | 5.44 s |
| LoRA plus head | `runs/probe-reload-v2/report.json` | 0.0 | 5.65 s |
The LoRA retry used source commit
`f7be7ceee9c38566dfebad0ebf5f1c04958faee3`. The dataset manifest SHA-256 was
`aeddcfc1ea5390f12122d2e786c1e2030dcc97d3790b6d7523432548198b1f94` throughout.
The approved GLM service was restored and passed its health check after these
probes. **M4 is complete.** Milestone 5 can now train a selected checkpoint and
measure complete games; the four-position probe does not demonstrate that the
trained policy wins or generalizes.