|
Download code/docs/evaluation-protocol.md from nima1/stackcraft-clef-flash-lora: direct link, hf CLI and curl.
- Browser
- Download file 10.1 kB
-
https://huggingface.co/nima1/stackcraft-clef-flash-lora/resolve/main/code/docs/evaluation-protocol.md
- Command line
-
hf download hf://nima1/stackcraft-clef-flash-lora/code/docs/evaluation-protocol.md
-
curl -L -o evaluation-protocol.md https://huggingface.co/nima1/stackcraft-clef-flash-lora/resolve/main/code/docs/evaluation-protocol.md
10.1 kB
| # Stackcraft evaluation protocol — completed study, frozen procedure | |
| The final test used all **200 episode seeds, 30000–30199**, at the frozen 200-piece | |
| cap, with all five players and no errors. Training, validation selection, reload | |
| verification, final evaluation and the independent CPU audit are complete. The | |
| [final study](../reports/final-study.md) reports the measured outcomes and | |
| [compact JSON](../reports/final-summary.json) preserves exact statistics and hashes. | |
| Release verification is tracked separately in the source repository. | |
| The procedure below records decisions frozen before test results were opened; | |
| completion does not authorize tuning against these now-disclosed sequences. | |
| `evaluation.py` consumes existing episode artifacts and position predictions; | |
| it does not invoke models, select checkpoints, or start games. | |
| ## Decisions that must be frozen before opening the test results | |
| The final episode cap is now **frozen at 200 pieces**, using all 200 reserved seeds. | |
| This was chosen before inspecting any test results. The unchanged-model development | |
| pilot measured approximately 144 ms per decision. If all three neural players reach | |
| every cap, 120,000 decisions would take roughly 4.8 hours at that rate, excluding | |
| loading and reporting. Earlier base development games died much sooner; runtime | |
| must still be measured and reported rather than promised. Each episode is saved | |
| atomically so interrupted evaluation can resume without discarding difficult games. | |
| Freeze the selected trained checkpoint and original pinned Clef revision, input | |
| encoding, precision, hardware/runtime configuration, heuristic revision, all | |
| seeds, cap, inference failure behavior, bootstrap parameters and compute budget. | |
| The same legal placements and one-piece preview must reach every player. Primary | |
| base is the pinned native model with its original BF16 head. Training uses an FP32 | |
| head; therefore an explicitly named `base-fp32` ablation keeps the original weights | |
| but uses the same FP32 wrapper as training. Report trained-versus-native-base as | |
| the primary comparison and trained-versus-base-fp32 to separate training from head | |
| precision effects. Backbone BF16, input encoding, context limit, device and source | |
| revision must agree. A changed base quantization would be a separate comparison. | |
| Checkpoint selection uses validation only. Teacher-action agreement and target | |
| cross-entropy on validation positions are diagnostic selection metrics, separate | |
| from complete-game outcomes. A model that imitates the expert more closely may | |
| still play worse. Preregister any training configuration search and tie rule; | |
| `paired_report` never makes those choices from test results. | |
| ## Paired episode outcomes | |
| Primary outcome is lines cleared. Secondary outcomes are game score and pieces | |
| placed. Retain every planned seed for each player, including failed episodes. | |
| Duplicate player/seed pairs, mismatched seed sets, caps or rules, and changing | |
| player identity/runtime within a comparison are rejected by the report function. | |
| Its schema checks are not a replacement for verifying replay integrity when | |
| episodes are collected. | |
| An inference exception or illegal decision ends that player's episode. For the | |
| primary analysis, assign **zero lines, score and pieces to any errored episode**. | |
| Keep its original partial achievements in `observed_before_error`, retain its seed, | |
| and report its error. This explicit reliability penalty prevents a model from | |
| benefiting statistically from dropping difficult failures. The same policy applies | |
| to base, trained and heuristic. Never describe a failed episode as a successful | |
| zero-line game; status and error counts remain separate. | |
| Report mean and median metrics, cap-hit rate, failed episodes, error rate and | |
| invalid-decision count. Survival is censored at the cap: a cap hit means at least | |
| that many pieces, not a measured death at that turn. This protocol cannot establish | |
| uncapped longevity or full competitive Tetris performance. | |
| For each seed, subtract base outcome from trained outcome. Resample these paired | |
| episode differences with replacement, **10,000 bootstrap replicates, RNG seed | |
| 2026**, and report their mean plus the percentile 95% interval. Resampling turns | |
| would falsely treat correlated moves within a game as independent observations. | |
| Pieces and score get descriptive intervals; lines are the single primary endpoint. | |
| The intervals are conditional on the tested checkpoint and episode distribution; | |
| one trained seed does not characterize training-seed variability. | |
| `positive_mean_lines_signal` is true only if there are at least two pairs, the | |
| primary interval's lower bound exceeds zero, and both base and trained episode | |
| error rates are at or below the declared threshold (**default 0%**). This flag is | |
| evidence for that measured comparison, not an automatic claim of general | |
| superiority, approval to publish, or a way to select a checkpoint on the test set. | |
| The heuristic remains a separate required comparator. Run the same paired analysis | |
| against it and disclose whether the model actually beats that cheap baseline. | |
| ## Timing and input lengths | |
| Summarize decision latency mean, median and nearest-rank p95 from all attempts, | |
| including failed attempts. These are timings recorded around `Player.choose`; | |
| they include tokenization and policy overhead, exclude game rendering/replay, and | |
| should not be advertised as pure GPU kernel time. CUDA probability transfer must | |
| complete before `choose` returns. Record cold loading separately and identify any | |
| warmup excluded before measured episodes. Never mix cached-replay speed with live | |
| model latency. Input-token summaries use only events containing `input_tokens` | |
| and report their count, so missing token metadata remains visible. | |
| ## Position predictions | |
| `summarize_positions` reports agreement with the fixed teacher label, target NLL | |
| and multiclass Brier score (sum over choices, averaged over positions). The fixed | |
| label uses engine-order tie-breaking; agreement therefore penalizes choosing a | |
| different equally valued action. Consider a separately named optimal-value tie | |
| agreement metric later if preregistered; do not silently change the definition. | |
| Missing or invalid predictions stay in the agreement denominator as failures. | |
| Aggregate NLL/Brier are available only with a complete legal probability | |
| distribution for every position. Partial diagnostic means have explicit subset | |
| denominators and cannot be presented as full-set results. A zero probability on | |
| the teacher label means infinite NLL; JSON stores a null mean plus an explicit | |
| `nll_is_infinite` flag and zero-target count, rather than silently clipping it. | |
| ```python | |
| from stackcraft.evaluation import paired_report, summarize_positions | |
| # Load complete, previously collected artifacts; this runs no model or game. | |
| report = paired_report(episodes, trained_id="trained", base_id="base") | |
| position_report = summarize_positions(validation_rows, predictions_by_record_id) | |
| ``` | |
| Toy tests verify known constant paired improvements, mismatched-pair rejection, | |
| failure retention, metadata consistency and proper probability-score arithmetic. | |
| Passing them verifies report mechanics, not any model's game performance. | |
| ## Execute and resume a frozen evaluation | |
| `scripts/evaluate_clef.py positions --dataset data/study-v1 --checkpoint PATH | |
| --output NEWDIR` evaluates the entire validation split and writes unrounded | |
| predictions plus NLL/agreement/Brier evidence. Its default players are native base, | |
| base-fp32 and trained. Model loading is sequential to fit one GPU; this is real | |
| inference and must follow memory admission. `--players trained` evaluates another | |
| candidate without rerunning the unchanged baselines. | |
| After validation, run `scripts/select_checkpoint.py` with both explicit epoch | |
| checkpoint/report pairs and a new output directory, as shown in Tutorial 05. | |
| It recomputes metrics from the raw predictions, verifies both checkpoint trees, | |
| selects by the preregistered rule, and copies self-contained validation evidence. | |
| Its `selection.json` contains schema_version 1, | |
| selection_split `validation`, selected_key, decision_rule, checkpoint_sha256 mapping, | |
| validation_metrics, validation_evidence `{path,sha256}`, test_seeds, their canonical | |
| JSON SHA256 as test_seeds_sha256, max_pieces 200, encoding_version, max_length, | |
| players, head_dtypes, bootstrap_samples 10000, bootstrap_seed 2026, max_error_rate | |
| 0.0. Validation evidence must be the script's position report for the same checkpoint; | |
| its path is relative to selection.json. Checkpoint hashes include every checkpoint | |
| file. Set head_dtypes to `{base: bfloat16, base-fp32: float32, trained: float32}` | |
| when all three neural players are included. | |
| Both epoch candidates and their eligibility reasons are mandatory in the selection | |
| artifact. Before opening test seeds, the evaluator checks all 215 positions, | |
| complete finite probability metrics, zero inference errors, checkpoint/dataset | |
| binding, and the lowest-NLL winner with an earlier-epoch tie rule. All five players, | |
| including the FP32-head ablation, are required by this study. | |
| ```sh | |
| uv run --locked --extra ml python scripts/evaluate_clef.py tournament \ | |
| --checkpoint PATH --seeds 30000:30200 --max-pieces 200 \ | |
| --final-test --selection-file selection.json --output runs/final-evaluation | |
| ``` | |
| The default includes all five players. The final-test switch is an explicit | |
| acknowledgment that these held-out seeds will now be consumed. Without it, the | |
| script rejects any overlap with the reserved pool before loading a player. | |
| Every completed episode is stored as `PLAYER/seed-N.json`; `progress.json` records | |
| progress, and the final `report.json` includes all episodes and paired comparisons. | |
| Repeat exactly the same command with `--resume` after an interruption. The script | |
| checks request, source hashes, selected checkpoint, frozen protocol and saved replay | |
| outcomes, skips completed players without loading them, and resumes missing seeds. | |
| It never retries or deletes a recorded error episode. Loading is timed separately; | |
| there are no silently excluded warmup decisions. | |