Download code/docs/tutorials/04-gpu-feasibility.md from nima1/stackcraft-clef-flash-lora: direct link, hf CLI and curl.
- Browser
- Download file 16.4 kB
-
https://huggingface.co/nima1/stackcraft-clef-flash-lora/resolve/main/code/docs/tutorials/04-gpu-feasibility.md
- Command line
-
hf download hf://nima1/stackcraft-clef-flash-lora/code/docs/tutorials/04-gpu-feasibility.md
-
curl -L -o 04-gpu-feasibility.md https://huggingface.co/nima1/stackcraft-clef-flash-lora/resolve/main/code/docs/tutorials/04-gpu-feasibility.md
Tutorial 04 — Prove training and reload before running an experiment
This milestone asks a narrow engineering question: can the real pinned Clef-flash model take gradient steps on this GPU, save all learned parameters, and reproduce its probabilities in a fresh process? It does not ask whether the model plays better. The latter requires complete games on the held-out sequences in Milestone 5.
The bounded real-weight training and fresh-process reload gate passed. Both head-only and LoRA-plus-head artifacts reproduced their reference probabilities exactly. The first LoRA reload exposed a saved-configuration compatibility issue; the corrected loader and successful retry are documented below. This establishes local training feasibility, not improved game performance.
Start from the native decision model
The base is Cloudflare/clef-flash, revision
17f0b0ad64efb65d273590632833508766b2aae6. The local loader verifies the pinned
joint_schema_model.py source hash and loads an explicitly resolved snapshot.
It requires trust_pinned_code=True: pinning and inspecting Python source controls
which code executes; it does not turn that code into a sandbox.
An observation becomes one native choice question containing every legal
placement. The answer is a distribution over action IDs, not a generated command
or natural-language explanation. encode_observation checks the complete token
budget before native encoding, because silently truncating a board would change
the task. Training uses only the visible board, current piece, one preview and
legal placements. Dataset seed and episode metadata remain outside the input.
The native encoder sorts option IDs lexicographically. decision_loss locates the
teacher's action in encoded.questions[0].option_ids; it must not reuse the
engine's rotation/column index as a target index. Both orders contain the same
legal moves, but their order can differ.
Do not train through systemone() or ClefPlayer.choose(): those are inference
paths. The probe calls native.collate_records and then model(batch) with
normal autograd enabled. choose() is appropriate for the before/after reference
probabilities, where gradients are unnecessary.
Keep the backbone small in memory and the head stable
src/stackcraft/training.py implements two bounded configurations:
| Configuration | Updated parameters | Purpose |
|---|---|---|
| Head only | Native joint decision head | Isolate head training and checkpoint handling before a backbone backward pass. |
| Rank-4 LoRA plus head | Small added text-layer matrices and the native head | Check whether the task can adapt both text representations and decision scoring. |
The backbone remains BF16. The native decision head is converted to FP32, and
FP32DecisionHead disables autocast inside the head. Its floating inputs are
converted to FP32 while token IDs and masks retain their original types. This
keeps the trainable scoring operations in full precision without converting the
whole multi-billion-parameter backbone.
The head reads selected rows of the output vocabulary embedding. Converting the
entire vocabulary matrix to FP32 first creates a large temporary allocation
(roughly 3.79 GiB for this release). GatheredFloat32Embedding instead implements
weight[indices].float(): select the needed rows first, then cast. This preserves
the actual native head computation and autograd connections. A tiny actual-native-
head test compares outputs and gradients with the full-cast reference exactly.
That test checks the optimization, not full-model training feasibility.
LoRA targets are full module paths under the text transformer layers. They include
both ordinary attention projections and Qwen3.5's linear-attention projections,
as well as the MLP projections. Matching only suffixes such as q_proj could
accidentally select the vision encoder; matching only ordinary attention would
miss the hybrid backbone. The vision tower, vocabulary output layer and unrelated
modules are excluded. Original backbone parameters are frozen. The probe uses
rank 4, alpha 8, no LoRA dropout and no trained backbone biases.
Non-reentrant gradient checkpointing recomputes intermediate activations during backward instead of retaining all of them. This trades time for memory while preserving gradients through frozen parts of the backbone into LoRA matrices. A tiny actual Qwen3.5 hybrid-backbone test checks nonzero LoRA gradients in ordinary attention, linear attention and MLP layers. These are stronger plumbing checks than testing a stand-in linear model alone, but still do not establish the real 9B GPU memory requirement.
We start with BF16 rather than immediately adding 4-bit quantized training. The native model has a custom head and loader; another quantization layer would add an unverified compatibility dependency. If BF16 fails, record the failure and make a bounded next decision rather than silently changing the experiment.
Prepare before borrowing the GPU
Run these commands from the repository root or the public bundle's code/
directory. Install the ML tools and cache the pinned release before testing:
uv sync --locked --extra ml --group dev
uv run --locked --extra ml hf download Cloudflare/clef-flash \
--revision 17f0b0ad64efb65d273590632833508766b2aae6
Downloading only populates the cache; it does not load GPU weights. It can run while other GPU services remain available. The actual native-head tests need this cached source. Explicitly enable the native tokenizer/encoder check too:
STACKCRAFT_TEST_NATIVE_ENCODING=1 uv run --locked --extra ml \
pytest tests/test_training.py tests/test_clef.py
These checks use CPU fixtures and the real tokenizer/source; they do not load the full backbone or establish GPU feasibility. Inspect the test summary: a skipped test is not a passed native-model test.
The training probe expects the audited data/study-v1 bundle established by
Tutorial03's exact-study reconstruction or a verified dataset download. If you
used a different directory, pass that same path with --dataset below.
It validates the train/validation bundle, then uses only the first four training
positions. It does not sample held-out test seeds. Record the current source
commit and any uncommitted probe changes in record.md before running.
This workstation normally runs a local GLM Docker service. The user explicitly
authorized temporarily stopping that service and restoring it after these
Stackcraft experiments. scripts/gpu_session.py is specific to that approved
service; it is not a generic instruction to stop another machine's workloads.
It first requires the configured container to be running, then stops it, runs a
bounded child job, and attempts restoration during cleanup. The training script
itself never stops services. It refuses to load unless CUDA is available and at
least 25 GiB is free. Free-memory admission is a precondition, not a guarantee
that backward will fit.
Run the bounded real-weight probe
The initial budget is two configurations and at most 30 minutes of child-job time. The commands below allocate 20 minutes to the training probe and 10 minutes to the separate reload check. Choose fresh output and session-record paths on a retry; neither existing evidence nor checkpoints should be overwritten.
uv run --locked --extra ml python scripts/gpu_session.py \
--timeout 1200 --record runs/gpu-training-probe-v1.json -- \
uv run --locked --extra ml python scripts/probe_clef_training.py \
--dataset data/study-v1 --output runs/clef-training-probe-v1
The probe first records unchanged native probabilities. It measures the initial probability drift introduced by moving the head to FP32, before training; a precision change must not be mistaken for a learned improvement. It then runs three head-only steps. After saving that checkpoint, it releases the model and loads a fresh unchanged base for five LoRA-plus-head steps. The LoRA experiment therefore does not inherit the head-only warm-up.
Both configurations use batch size one, AdamW with learning rate 1e-5, norm
clipping at 1.0, and this loss:
smoothed_cross_entropy(label_smoothing=0.05)
+ 0.1 * sum_over_options((probability - one_hot_teacher_label)**2)
The cross-entropy term learns the expert action. Smoothing avoids assigning the entire training target to one action. The Brier term also penalizes the predicted probability distribution's distance from the teacher target. These fixed choices are a probe configuration, not evidence of calibrated confidence or optimal hyperparameters.
For every step, the probe checks finite loss and gradients and records nonzero
head/LoRA gradient sums, tokens, elapsed time, and CUDA peak allocated/reserved
bytes. Before/after parameter hashes check that intended parameters changed and
frozen parameters did not. Hashing the entire frozen backbone costs time but
checks something that requires_grad=False alone cannot prove: what bytes actually
changed during this run.
Inspect runs/clef-training-probe-v1/report.json and the session record. A timeout,
CUDA error, absent gradient or failed hash check is a failed gate. A report left at
running after forced termination is incomplete evidence, even if some checkpoint
files exist. Do not start long training merely because loss fell on four examples.
Save the head as well as the adapter
A LoRA checkpoint has this structure:
checkpoint/
joint_head.safetensors
adapter/
adapter_config.json
adapter_model.safetensors
training_config.json
reference.json
The head is an independently trained part of Clef; saving only the PEFT adapter
would omit learned parameters. joint_head.safetensors stores FP32 native head
keys without wrapper-specific prefixes. The adapter stores only the added LoRA
weights, not another copy of the frozen base. Metadata identifies the base
revision, native source hash, input encoding, head format, LoRA settings and
teacher rows. The loader checks this contract, head shapes and dtypes, and actual
saved adapter configuration before restoring weights.
The head-only probe writes head-checkpoint/ without an adapter. It is diagnostic
output, not a selected release model.
Require a separate fresh-process reload
Only after the first probe succeeds, run:
uv run --locked --extra ml python scripts/gpu_session.py \
--timeout 600 --record runs/gpu-training-reload-v1.json -- \
uv run --locked --extra ml python scripts/probe_clef_training.py \
--dataset data/study-v1 --output runs/clef-training-reload-v1 \
--reload runs/clef-training-probe-v1/checkpoint
This starts another Python process, loads the unchanged pinned base, restores both
adapter and head, and evaluates the same four training observations. Require the
same option IDs and a maximum absolute probability difference of at most
1e-4 across all recorded options. The tolerance was declared before the run.
Compare unrounded probabilities; rounded display values can conceal a mismatch.
A same-process test could accidentally reuse trained parameters that were never saved. The fresh process tests the actual inference artifact. A passing reload establishes serialization fidelity on these references, not general game skill or correctness on every possible input.
Restore service and record the limits
Inspect both restored_running and restored_healthy after success and failure.
The wrapper checks /health inside the restored container and allows up to 120
seconds for readiness, retrying transient timeouts and malformed startup responses.
It does not query the unrelated service on host port 8080. A running process alone
would not establish readiness. Nested cleanup ensures a child-termination error
still reaches restoration; repeated ordinary interrupts are ignored during that
cleanup. Preserve the wrapper record, probe reports and console logs, including
errors. Mocked tests exercise these recovery paths without touching Docker.
Process cleanup cannot guarantee recovery after SIGKILL, host power loss or a
Docker daemon failure. If the wrapper was forcibly terminated, inspect the exact
approved container and any remaining job before acting. Restore the existing
service only after the experiment has released GPU memory; do not start another
probe simply because the last terminal stopped printing output.
The earlier M2 real-weight development pilot is useful context: on this RTX 5090,
it recorded 55 unchanged-Clef decisions at a mean of about 0.144 seconds per
decision, with 19,686,539,776 bytes peak CUDA allocation (about 19.7 GB,
18.3 GiB). The benchmark script took 12.60 seconds; the stop/job/restore wrapper
interval was about 14.20 seconds and recorded successful restoration. These are
inference measurements on two short development games, not training memory or
held-out quality results. Sources are runs/clef-development-v1.json and
runs/gpu-baseline-session.json.
Measured training probe and the first reload failure
The real-weight probe passed on 2026-10-05 using the RTX 5090 and Torch
2.14.1+cu130. Evidence is runs/probe-v1/report.json; the actual run directory
differs from the fresh reproduction paths shown above. It used training positions
seed-10000-turn-0 through seed-10000-turn-3, with 1,302–2,252 encoded tokens.
| Measurement | Head-only probe | LoRA-plus-head probe |
|---|---|---|
| Optimization steps | 3 | 5 |
| Observed step times | 0.154–0.328 s | 0.670–1.322 s |
| Peak allocated CUDA bytes | 21,323,760,640 | 23,180,389,888 |
| Peak reserved CUDA bytes | 21,541,945,344 | 23,595,057,152 |
| Intended parameter changes | Verified | Verified |
| Frozen parameter hashes | Unchanged | Unchanged |
All reported losses and gradients were finite, with nonzero gradients in the
intended groups. Total probe time was 51.81 seconds, including model reload,
parameter hashing and checkpoint work. This is not an estimate of a full epoch:
only four short training positions were exercised. Peak reserved memory is the
PyTorch allocator's reservation, not the same quantity as all memory reported by
nvidia-smi.
The FP32 head changed initial probabilities by at most 0.00074412 before any
training. That is a precision effect relative to the original BF16 head, not a
failed serialization comparison. The 1e-4 reload threshold compares the saved
trained FP32-head model with its own fresh reload.
The first fresh reload, runs/probe-reload-v1/report.json, failed after 4.15
seconds because the saved adapter configuration differed from our metadata. PEFT
automatically shortened 248 full target paths to 12 module suffixes. The strict
loader correctly refused a mismatch; the failure does not by itself mean the
weight tensors are corrupt. Accepting suffixes without checking what they match
could attach an adapter to unintended modules. The correction must expand saved
target selectors against the actual pinned backbone and require exactly the
intended full target set, rejecting any extra destinations. This is why fresh
reload is a milestone gate rather than a final packaging chore.
The corrected loader resolves saved selectors against the pinned backbone and requires the exact intended destinations. The retry preserved the original checkpoint and failed report; it changed validation of equivalent target selectors, not the trained weights. A regression test covers target shortening and rejects selectors that also match unintended modules.
Both independent fresh-process checks then passed:
| Artifact | Report | Maximum probability difference | Process time |
|---|---|---|---|
| Head only | runs/probe-head-reload-v1/report.json |
0.0 | 5.44 s |
| LoRA plus head | runs/probe-reload-v2/report.json |
0.0 | 5.65 s |
The LoRA retry used source commit
f7be7ceee9c38566dfebad0ebf5f1c04958faee3. The dataset manifest SHA-256 was
aeddcfc1ea5390f12122d2e786c1e2030dcc97d3790b6d7523432548198b1f94 throughout.
The approved GLM service was restored and passed its health check after these
probes. M4 is complete. Milestone 5 can now train a selected checkpoint and
measure complete games; the four-position probe does not demonstrate that the
trained policy wins or generalizes.