Instructions to use BennyLin01/pact-rl-tetris with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use BennyLin01/pact-rl-tetris with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
PACT-RL: a 9B language model that plays Tetris
This repository contains two LoRA adapters for Qwen/Qwen3.5-9B, the data they were trained on, and all the code needed to use or retrain them.
- PACT (
results/model/) is a supervised adapter for single-token typed decisions: the model reads a context and a schema question and answers by scoring one letter code per allowed choice. It is trained with Pairwise-Anchored Calibrated Tuning, described inpaper/main.pdf. - PACT-RL (
pact-rl-tetris/) starts from PACT and is fine-tuned with REINFORCE to choose Tetris moves. Each move is asked as an ordinary typed question, so the game is played through exactly the same single-token interface; no new heads or architecture changes.
The full game (browser and Windows desktop version) is included, and the model can play it live on a single 16 GB GPU.
| Path | Contents |
|---|---|
pact-rl-tetris/ |
PACT-RL adapter (step 1,550) and the complete training logs |
results/model/ |
PACT adapter, temperature calibrator, prompt contract |
data/ |
train (2,676 examples) and frozen holdout (324 examples) |
tetris/, desktop/ |
game server for AI play, game UI, desktop app |
pact/, nimble/, configs/ |
training, evaluation and inference code |
results/, paper/ |
every run's metrics, tables, figures, and the paper |
media/ |
gameplay recording and the scripts that produced it |
How the model plays Tetris
A language model cannot be expected to simulate falling blocks from a raw board, so the game engine does that part:
- For the current piece, the engine lists every legal placement.
- For each one it also tries every placement of the (already known) next piece and keeps the best combined result. This two-ply lookahead is scored with the Dellacherie/Lee heuristic plus a column-height variance term.
- The six best placements, and any placement that clears a row, are written out in plain English: rows cleared, new holes, stack height, bumpiness, and what the board looks like after the next piece.
- The model answers the question "which placement?" with a single forward pass. The highest-scoring answer code is played.
Every candidate is a legal move, so the model can only choose a weak move, never an illegal one.
Reinforcement learning
The supervised PACT adapter was trained on business documents, not games. When its moves were sampled at step 1 of RL training, it chose the lookahead's top-ranked candidate only 19% of the time.
PACT-RL treats each move as a one-step contextual bandit. The lookahead already scores every candidate, so the reward of the sampled move is its lookahead score minus the mean score of all candidates at that state. This baseline is exact and free, so plain REINFORCE works without a value network:
loss = -log p(chosen move) * (score(chosen) - mean score) - 0.01 * entropy
Training details (pact/train_tetris_rl.py, logs in pact-rl-tetris/logs/):
| Starting point | PACT adapter (results/model/) |
| Hardware | 2 x 40 GB GPUs, BF16, device_map=auto |
| Batch | 32 on-policy moves per gradient step |
| Optimizer | AdamW, lr 1e-5, 20 warm-up steps, grad-norm clip 1.0 |
| Steps | 1,574 (about 33 s per step, 14.5 h); released checkpoint: step 1,550 |
Results
| Metric | Value |
|---|---|
| Top-ranked move chosen, step 1 | 18.8% |
| Top-ranked move chosen, step 16 | 87.5% |
| Top-ranked move chosen, mean of the last 200 steps | 94.9% |
| Greedy eval, 5 games, 400-piece cap (31 evaluations) | every game reached the cap; mean score 16,940 to 17,940 |
| Greedy eval, 1 game, 1,500-piece cap (final checkpoint) | reached the cap: 598 lines, score 65,700 |
The capped evaluation stayed flat during training because each game ends at the piece cap rather than by topping out, so it does not measure the change in move quality. The agreement rate with the lookahead is the clearer signal.
The recording in media/ shows six minutes of uncut play (sped up
2x) on one RTX 5060 Ti 16 GB with the 4-bit config: more than 130 lines and
level 13 without topping out.
Quick start
Install PyTorch for your CUDA version first, then:
pip install -r requirements.txt
pip install bitsandbytes # for the 4-bit config used below
python download_all.py # base model into ./.cache
Play against the model
python tetris/server.py
# open http://127.0.0.1:8848/index_en.html (or index_zh.html) and press "Let PACT play"
The server loads pact-rl-tetris/ in 4-bit by default, which fits on a
16 GB GPU. The first move takes 10 to 20 s to load the model; after that each
move is one forward pass of about 0.3 to 0.5 s. To use the supervised adapter
or a different config:
PACT_ADAPTER_DIR=results/model PACT_CONFIG=configs/server_2x40gb_rl.json python tetris/server.py
GET http://127.0.0.1:8848/health shows which adapter and config are loaded.
Manual play needs no server at all: open desktop/app/index_en.html, or run
the Windows build described in desktop/README.md.
Use the adapter in Python
import sys
sys.path.insert(0, ".") # repository root
sys.path.insert(0, "tetris")
import torch
from peft import PeftModel
from pact.config import PactConfig
from pact.model import load_tokenizer, _load_base, LogitReader
from nimble.scoring.parallel_schema import prepare_prompts
import server # tetris/server.py: build_context, build_schema
cfg = PactConfig.load("configs/server_9b_24gb_qlora.json")
tokenizer = load_tokenizer(cfg)
base, _ = _load_base(cfg, device="cuda", dtype=torch.bfloat16, log=print)
model = PeftModel.from_pretrained(base, "pact-rl-tetris").eval()
# payload has the same shape as the JSON the game posts to /ai_move
context = server.build_context(payload)
schema = server.build_schema(payload["candidates"])
prompt = prepare_prompts(tokenizer, context, schema, cfg.data.max_length)
ids = torch.tensor([list(prompt.full_ids[0])], device="cuda")
with torch.no_grad(), torch.autocast("cuda", dtype=torch.bfloat16):
logits = LogitReader()(model, ids, torch.ones_like(ids))[0]
scores = [float(logits[code]) for code in prompt.candidate_ids[0]]
move = max(range(len(scores)), key=scores.__getitem__)
pact/train_tetris_rl.py (build_candidates, to_payload) builds the same
payload from a headless game in pact/tetris_env.py.
Retrain PACT-RL
python train_rl.py # 2 x 40 GB GPUs, BF16
python train_rl.py --config configs/server_9b_24gb_qlora.json # 1 x 24 GB GPU, 4-bit
train_rl.py first times two steps on your hardware and picks the batch size
from that, then runs 2,000 steps. Progress goes to
runs/tetris-rl/train_log.jsonl (every step) and eval_log.jsonl (greedy
games every 50 steps); checkpoints go to runs/tetris-rl/step-<N>/ and
runs/tetris-rl/latest/. Ctrl+C saves a checkpoint before exiting. To
continue:
python train_rl.py --skip-probe --batch-size 32 --resume-from runs/tetris-rl/latest --max-steps 4000
To measure how long a checkpoint survives:
python pact/train_tetris_rl.py --eval-only --adapter-dir pact-rl-tetris --eval-games 3 --eval-max-pieces 20000
The training prompt is produced by the same functions as the serving prompt
(tetris/server.py), and pact/tetris_env.py is a Python port of the
JavaScript engine, so training and play see the same states and text.
PACT: the supervised adapter
PACT changes only the training objective, the batching and the probability head of the Nimble recipe; the prompt, the answer codes and the adapter format stay the same. The training data comes in contrastive pairs: two nearly identical contexts that differ in one fact, which changes the answer. PACT adds four terms to cross-entropy:
| Term | What it does |
|---|---|
L_CF pair margin |
Both members of a pair share a batch; a difference-in-differences margin rewards responding to the edited fact and cancels anything the two contexts share. |
L_PC permutation consistency |
Each example is also encoded with its answer letters shuffled; the two distributions must agree after mapping back. Targets letter-position bias. |
L_NR necessity |
With the decisive evidence sentence removed, the two answers it separated must score equally. |
L_EMD ordinal |
Earth mover's distance on rubric (score) fields, so being one level off costs less than three. |
L = L_CE + λ_cf·L_CF + λ_pc·L_PC + λ_nr·L_NR + λ_emd·L_EMD
It also fits a contextual temperature T = softplus(a + b·log C + c·log(L/1000))
on an inner split and uses rsLoRA with LoRA+ learning rates.
Results on the 324-example holdout
Single-pass decoding, mean ± s.d. over seeds 17, 18 and 19:
| Run | Accuracy | Pair acc. | ECE → with T(x) | Score MAE | Answer flips under relabelling (K=4) |
|---|---|---|---|---|---|
| Base model (no adapter) | 66.4 | 35.8 | 24.6 → – | 0.571 | 23.5 |
| Nimble recipe (reproduced) | 85.2 ± 2.1 | 71.0 ± 3.9 | 9.5 → 5.4 | 0.311 | 13.8 |
| CE only (same optimizer, no new terms) | 82.2 ± 6.0 | 70.2 ± 6.2 | 14.3 → 10.3 | 0.322 | 12.1 |
| PACT | 84.6 ± 2.8 | 73.3 ± 6.2 | 10.8 → 7.9 | 0.232 | 9.8 |
- Accuracy is not significantly different from the reproduced Nimble recipe (paired McNemar p ≥ 0.50 at every seed).
- Against the cross-entropy-only control with the same optimizer and schedule, PACT is significantly more accurate at 2 of 3 seeds (p = 0.008, 0.035), halves the spread across seeds and lowers NLL by 26%.
- PACT has the lowest letter-position bias and the lowest ordinal error on score fields of all runs. The Nimble recipe remains the best calibrated.
The released adapter is seed 18, chosen by inner-validation NLL (never by
holdout score). Per-run metrics are in results/, the full analysis in the
paper.
from pact.inference import PactScorer
scorer = PactScorer("results/model", ensemble=1)
result = scorer.score("The payment service is down for all customers.", schema)
Reproduce the sweep
CUDA_VISIBLE_DEVICES=0,1 nohup python train_all.py > train.log 2>&1 &
This runs preflight, view building, training, calibration, evaluation,
reporting and packaging for all 14 run × seed jobs
(configs/server_2x48gb.json, one job per GPU) and writes everything to
results/. Useful variants:
python train_all.py --dry-run # print the plan
python train_all.py --only-run pact_full ce_only # a subset of runs
python train_all.py --stages evaluate report # rescore existing adapters
python train_all.py --stages preflight # data checks, no GPU
python -m unittest discover -s tests -t . # 44 CPU-only tests
Configs: server_2x48gb.json (default, 2 × 48 GB), server_9b_h100.json,
server_9b_24gb_qlora.json (one 24 GB GPU, 4-bit), server_2x40gb_rl.json
(RL), smoke_cpu.json (plumbing check with a 135M model).
The figures and tables of the paper are rebuilt from results/ and
pact-rl-tetris/logs/ by python paper/build_assets.py.
Data
data/train.jsonl (2,676 examples, 1,338 contrastive pairs, 10 domains) and
data/eval.jsonl (324 examples, held-out source families) come from the
Nimble project by Bespoke Labs.
They are synthetic and model-labelled; no example has been reviewed by a
person. Each row has the context and schema (input), the gold answer
(reference), its pair family, and an evidence_certificate that records
which sentences decide the answer. PACT uses the certificates to rebuild the
evidence-removed contexts for L_NR. data/manifest.json has the counts and
hashes that preflight checks. The holdout is used only for final scoring.
The nimble/ package contains unmodified copies of the upstream modules that
define the prompt format, so prompts cannot drift from what the adapters were
trained with.
Limitations
- The holdout is small (324 examples from six source families) and synthetic.
L_CFandL_NRdepend on the same model-written certificates as the labels, so certificate errors carry over. - PACT-RL learns to agree with a fixed heuristic. It does not discover strategy beyond what the lookahead ranks highly, and the game engine, not the model, does all the spatial reasoning.
- The RL adapter is specialized for the Tetris prompt; use the PACT adapter for general typed decisions.
- Permutation ensembling multiplies inference cost by K; all headline numbers use K = 1.
Citation
@misc{lin2026pact,
author = {Yida Lin},
title = {{PACT}: Pairwise-Anchored Calibrated Tuning for Single-Token Typed Decisions},
year = {2026},
url = {https://github.com/BennyLinntu/PACT-Pairwise-Anchored-Calibrated-Tuning-for-Single-Token-Typed-Decisions}
}
- Downloads last month
- -
