Download code/README.md from nima1/stackcraft-clef-flash-lora: direct link, hf CLI and curl.
- Browser
- Download file 7.44 kB
-
https://huggingface.co/nima1/stackcraft-clef-flash-lora/resolve/main/code/README.md
- Command line
-
hf download hf://nima1/stackcraft-clef-flash-lora/code/README.md
-
curl -L -o README.md https://huggingface.co/nima1/stackcraft-clef-flash-lora/resolve/main/code/README.md
Stackcraft
A playable, deterministic falling-block game for studying decision models. The goal is a fair comparison of a heuristic player, unchanged Clef-flash, and a fine-tuned model on identical piece sequences.
Status: training, validation selection, fresh-process reload and the 1,000-game held-out evaluation are complete. The final report and recorded demo passed review. See the release tutorial for milestones and acceptance criteria.
Releases
- Run the playable demo locally
- Model adapter and decision head
- Synthetic dataset
- Source and release verification
The demo replays actual model decisions alongside live human play. It needs no GPU hosting. GitHub, the model and dataset are public. Hosted deployment was deferred by the owner; the game and Docker demo remain runnable locally.
What the experiment found
On 200 paired sequences per player, with a 200-piece cap:
| Player | Mean lines | Mean pieces placed | Mean decision time |
|---|---|---|---|
| Unchanged Clef-flash | 0.070 | 26.020 | 149.11 ms |
| Unchanged FP32-head control | 0.105 | 26.000 | 150.19 ms |
| Fine-tuned Clef-flash | 16.810 | 83.565 | 206.81 ms |
| Random legal placement | 0.140 | 26.645 | 0.0015 ms |
| Fixed heuristic | 76.530 | 200.000 | 0.1985 ms |
Fine-tuning improved mean lines over native Clef by 16.74, with a paired 95% bootstrap interval of [15.74, 17.77]. The same improvement remains against the unchanged FP32-head control. The heuristic cleared substantially more lines than the trained model and was about 1,042 times faster per decision on this machine. All players had zero errors or invalid decisions. The heuristic reached the cap in every game; its uncapped survival was not measured.
This is a successful adaptation experiment and a clear example of why a cheap baseline matters. It does not establish that a large model is the best tool for this game. See the full study report for medians, uncertainty, validation results, limitations and reproducible evidence.
Run locally
Install uv, then from this directory:
uv sync --locked
uv run --locked stackcraft serve --port 8087
Open http://127.0.0.1:8087. The game needs no GPU or model credential. Sessions are local, in memory, and expire when the server restarts; download a replay to retain a game. The local service retains 128 sessions and limits each to 2,000 placements. Use a single server worker. The game does not require an account.
Development
uv run --locked ruff check .
uv run --locked ruff format --check .
uv run --locked ty check src/stackcraft --exclude src/stackcraft/training.py
uv run --locked pytest
These commands check the game without installing ML packages. Full-repository type checking also covers training scripts and requires the ML extra:
uv run --locked --extra ml ty check
Installing that extra does not run GPU inference. When keeping an ML environment,
use --extra ml on uv commands to prevent synchronization from removing those
optional packages.
The committed lockfile fixes dependency versions; uv sync --locked refuses to silently update it. Python 3.13 is selected by .python-version. No pip activation or global dependency installation is needed.
Read tutorial 00 for the setup rationale, the rules for the exact game, and the API contract for integration. This is a simplified turn-based placement game, with no gravity timer, hold or wall kicks. It is not affiliated with Tetris.
Offline baseline tournament
uv run --locked stackcraft tournament --seed-start 1000 --episodes 20 --max-pieces 100 --output runs/development.json
This runs random and fixed heuristic players on identical seeds, records every decision and replay, and refuses to overwrite an existing result. See tutorial02. Development scores are not held-out model results.
Expert-labelled data
uv run --locked stackcraft generate-data --output data/study-v1
This generates training/validation episodes from fixed disjoint seed pools, labels moves with bounded search, removes duplicate observations and writes a hashed manifest. Final test seeds remain reserved. See tutorial03.
Optional Clef tooling uses uv sync --locked --extra ml. Its native encoder, real GPU inference, finite head/LoRA gradients and exact fresh-process checkpoint reload have been verified. Full-study quality is a separate evaluation. See native adapter documentation.
The model used here
The base is Cloudflare/clef-flash, built on Qwen3.5-9B, pinned at
17f0b0ad64efb65d273590632833508766b2aae6. It produces probabilities over every
supplied legal move instead of generating text. Training updates rank-4 text-layer
LoRA matrices and the native joint decision head; the original backbone is frozen.
The checkpoint must save and load both the adapter and head.
On the RTX 5090, the short native inference pilot allocated 19.7 GB and averaged 144 ms per decision. The training probe allocated 23.2 GB; both head-only and LoRA-plus-head checkpoints reproduced their reference probabilities exactly in fresh processes. These are engineering measurements, not evidence of better game play. See the measured feasibility report.
Follow the milestones
- Foundation and modern Python tools
- Deterministic game and replay
- Cheap and unchanged-model baselines
- Search-teacher data and leakage checks
- Real GPU training and checkpoint fidelity
- Validation selection and paired game evaluation
- Release, public reproduction and fresh-download checks
Each tutorial explains the commands, evidence, design choices and alternatives. The evaluation protocol freezes 200 test sequences and a 200-piece cap before results. It compares the selected model with the native base, an unchanged FP32-head ablation, random moves and a fast heuristic.
The training runner intentionally accepts only the exact frozen study dataset. A newly generated dataset can have different provenance hashes even if its labels match. For exact reproduction use the released dataset or Tutorial03's source-hash checked reconstruction. A new dataset or hyperparameter study should be recorded as a new experiment, not quietly substituted into this study.
Demo hosting
The Docker deployment guide builds a CPU-only image with no model credentials or Torch. Human gameplay is live; model comparisons replay recorded decisions and probabilities. Playback speed is separate from measured inference latency. Local packaging has been verified. Release verification is recorded in the GitHub repository.
Apache-2.0. See LICENSE and NOTICE. Stackcraft is independent of Cloudflare, Qwen and Tetris.