PACT-RL: a 9B language model that plays Tetris

PACT-RL playing Tetris

This repository contains two LoRA adapters for Qwen/Qwen3.5-9B, the data they were trained on, and all the code needed to use or retrain them.

  • PACT (results/model/) is a supervised adapter for single-token typed decisions: the model reads a context and a schema question and answers by scoring one letter code per allowed choice. It is trained with Pairwise-Anchored Calibrated Tuning, described in paper/main.pdf.
  • PACT-RL (pact-rl-tetris/) starts from PACT and is fine-tuned with REINFORCE to choose Tetris moves. Each move is asked as an ordinary typed question, so the game is played through exactly the same single-token interface; no new heads or architecture changes.

The full game (browser and Windows desktop version) is included, and the model can play it live on a single 16 GB GPU.

Path Contents
pact-rl-tetris/ PACT-RL adapter (step 1,550) and the complete training logs
results/model/ PACT adapter, temperature calibrator, prompt contract
data/ train (2,676 examples) and frozen holdout (324 examples)
tetris/, desktop/ game server for AI play, game UI, desktop app
pact/, nimble/, configs/ training, evaluation and inference code
results/, paper/ every run's metrics, tables, figures, and the paper
media/ gameplay recording and the scripts that produced it

How the model plays Tetris

A language model cannot be expected to simulate falling blocks from a raw board, so the game engine does that part:

  1. For the current piece, the engine lists every legal placement.
  2. For each one it also tries every placement of the (already known) next piece and keeps the best combined result. This two-ply lookahead is scored with the Dellacherie/Lee heuristic plus a column-height variance term.
  3. The six best placements, and any placement that clears a row, are written out in plain English: rows cleared, new holes, stack height, bumpiness, and what the board looks like after the next piece.
  4. The model answers the question "which placement?" with a single forward pass. The highest-scoring answer code is played.

Every candidate is a legal move, so the model can only choose a weak move, never an illegal one.

Reinforcement learning

The supervised PACT adapter was trained on business documents, not games. When its moves were sampled at step 1 of RL training, it chose the lookahead's top-ranked candidate only 19% of the time.

PACT-RL treats each move as a one-step contextual bandit. The lookahead already scores every candidate, so the reward of the sampled move is its lookahead score minus the mean score of all candidates at that state. This baseline is exact and free, so plain REINFORCE works without a value network:

loss = -log p(chosen move) * (score(chosen) - mean score) - 0.01 * entropy

Training details (pact/train_tetris_rl.py, logs in pact-rl-tetris/logs/):

Starting point PACT adapter (results/model/)
Hardware 2 x 40 GB GPUs, BF16, device_map=auto
Batch 32 on-policy moves per gradient step
Optimizer AdamW, lr 1e-5, 20 warm-up steps, grad-norm clip 1.0
Steps 1,574 (about 33 s per step, 14.5 h); released checkpoint: step 1,550

Results

Metric Value
Top-ranked move chosen, step 1 18.8%
Top-ranked move chosen, step 16 87.5%
Top-ranked move chosen, mean of the last 200 steps 94.9%
Greedy eval, 5 games, 400-piece cap (31 evaluations) every game reached the cap; mean score 16,940 to 17,940
Greedy eval, 1 game, 1,500-piece cap (final checkpoint) reached the cap: 598 lines, score 65,700

The capped evaluation stayed flat during training because each game ends at the piece cap rather than by topping out, so it does not measure the change in move quality. The agreement rate with the lookahead is the clearer signal.

The recording in media/ shows six minutes of uncut play (sped up 2x) on one RTX 5060 Ti 16 GB with the 4-bit config: more than 130 lines and level 13 without topping out.

Quick start

Install PyTorch for your CUDA version first, then:

pip install -r requirements.txt
pip install bitsandbytes                 # for the 4-bit config used below
python download_all.py                   # base model into ./.cache

Play against the model

python tetris/server.py
# open http://127.0.0.1:8848/index_en.html (or index_zh.html) and press "Let PACT play"

The server loads pact-rl-tetris/ in 4-bit by default, which fits on a 16 GB GPU. The first move takes 10 to 20 s to load the model; after that each move is one forward pass of about 0.3 to 0.5 s. To use the supervised adapter or a different config:

PACT_ADAPTER_DIR=results/model PACT_CONFIG=configs/server_2x40gb_rl.json python tetris/server.py

GET http://127.0.0.1:8848/health shows which adapter and config are loaded. Manual play needs no server at all: open desktop/app/index_en.html, or run the Windows build described in desktop/README.md.

Use the adapter in Python

import sys
sys.path.insert(0, ".")          # repository root
sys.path.insert(0, "tetris")

import torch
from peft import PeftModel
from pact.config import PactConfig
from pact.model import load_tokenizer, _load_base, LogitReader
from nimble.scoring.parallel_schema import prepare_prompts
import server  # tetris/server.py: build_context, build_schema

cfg = PactConfig.load("configs/server_9b_24gb_qlora.json")
tokenizer = load_tokenizer(cfg)
base, _ = _load_base(cfg, device="cuda", dtype=torch.bfloat16, log=print)
model = PeftModel.from_pretrained(base, "pact-rl-tetris").eval()

# payload has the same shape as the JSON the game posts to /ai_move
context = server.build_context(payload)
schema = server.build_schema(payload["candidates"])
prompt = prepare_prompts(tokenizer, context, schema, cfg.data.max_length)
ids = torch.tensor([list(prompt.full_ids[0])], device="cuda")
with torch.no_grad(), torch.autocast("cuda", dtype=torch.bfloat16):
    logits = LogitReader()(model, ids, torch.ones_like(ids))[0]
scores = [float(logits[code]) for code in prompt.candidate_ids[0]]
move = max(range(len(scores)), key=scores.__getitem__)

pact/train_tetris_rl.py (build_candidates, to_payload) builds the same payload from a headless game in pact/tetris_env.py.

Retrain PACT-RL

python train_rl.py                                    # 2 x 40 GB GPUs, BF16
python train_rl.py --config configs/server_9b_24gb_qlora.json   # 1 x 24 GB GPU, 4-bit

train_rl.py first times two steps on your hardware and picks the batch size from that, then runs 2,000 steps. Progress goes to runs/tetris-rl/train_log.jsonl (every step) and eval_log.jsonl (greedy games every 50 steps); checkpoints go to runs/tetris-rl/step-<N>/ and runs/tetris-rl/latest/. Ctrl+C saves a checkpoint before exiting. To continue:

python train_rl.py --skip-probe --batch-size 32 --resume-from runs/tetris-rl/latest --max-steps 4000

To measure how long a checkpoint survives:

python pact/train_tetris_rl.py --eval-only --adapter-dir pact-rl-tetris --eval-games 3 --eval-max-pieces 20000

The training prompt is produced by the same functions as the serving prompt (tetris/server.py), and pact/tetris_env.py is a Python port of the JavaScript engine, so training and play see the same states and text.

PACT: the supervised adapter

PACT changes only the training objective, the batching and the probability head of the Nimble recipe; the prompt, the answer codes and the adapter format stay the same. The training data comes in contrastive pairs: two nearly identical contexts that differ in one fact, which changes the answer. PACT adds four terms to cross-entropy:

Term What it does
L_CF pair margin Both members of a pair share a batch; a difference-in-differences margin rewards responding to the edited fact and cancels anything the two contexts share.
L_PC permutation consistency Each example is also encoded with its answer letters shuffled; the two distributions must agree after mapping back. Targets letter-position bias.
L_NR necessity With the decisive evidence sentence removed, the two answers it separated must score equally.
L_EMD ordinal Earth mover's distance on rubric (score) fields, so being one level off costs less than three.
L = L_CE + λ_cf·L_CF + λ_pc·L_PC + λ_nr·L_NR + λ_emd·L_EMD

It also fits a contextual temperature T = softplus(a + b·log C + c·log(L/1000)) on an inner split and uses rsLoRA with LoRA+ learning rates.

Results on the 324-example holdout

Single-pass decoding, mean ± s.d. over seeds 17, 18 and 19:

Run Accuracy Pair acc. ECE → with T(x) Score MAE Answer flips under relabelling (K=4)
Base model (no adapter) 66.4 35.8 24.6 → – 0.571 23.5
Nimble recipe (reproduced) 85.2 ± 2.1 71.0 ± 3.9 9.5 → 5.4 0.311 13.8
CE only (same optimizer, no new terms) 82.2 ± 6.0 70.2 ± 6.2 14.3 → 10.3 0.322 12.1
PACT 84.6 ± 2.8 73.3 ± 6.2 10.8 → 7.9 0.232 9.8
  • Accuracy is not significantly different from the reproduced Nimble recipe (paired McNemar p ≥ 0.50 at every seed).
  • Against the cross-entropy-only control with the same optimizer and schedule, PACT is significantly more accurate at 2 of 3 seeds (p = 0.008, 0.035), halves the spread across seeds and lowers NLL by 26%.
  • PACT has the lowest letter-position bias and the lowest ordinal error on score fields of all runs. The Nimble recipe remains the best calibrated.

The released adapter is seed 18, chosen by inner-validation NLL (never by holdout score). Per-run metrics are in results/, the full analysis in the paper.

from pact.inference import PactScorer

scorer = PactScorer("results/model", ensemble=1)
result = scorer.score("The payment service is down for all customers.", schema)

Reproduce the sweep

CUDA_VISIBLE_DEVICES=0,1 nohup python train_all.py > train.log 2>&1 &

This runs preflight, view building, training, calibration, evaluation, reporting and packaging for all 14 run × seed jobs (configs/server_2x48gb.json, one job per GPU) and writes everything to results/. Useful variants:

python train_all.py --dry-run                      # print the plan
python train_all.py --only-run pact_full ce_only   # a subset of runs
python train_all.py --stages evaluate report       # rescore existing adapters
python train_all.py --stages preflight             # data checks, no GPU
python -m unittest discover -s tests -t .          # 44 CPU-only tests

Configs: server_2x48gb.json (default, 2 × 48 GB), server_9b_h100.json, server_9b_24gb_qlora.json (one 24 GB GPU, 4-bit), server_2x40gb_rl.json (RL), smoke_cpu.json (plumbing check with a 135M model).

The figures and tables of the paper are rebuilt from results/ and pact-rl-tetris/logs/ by python paper/build_assets.py.

Data

data/train.jsonl (2,676 examples, 1,338 contrastive pairs, 10 domains) and data/eval.jsonl (324 examples, held-out source families) come from the Nimble project by Bespoke Labs. They are synthetic and model-labelled; no example has been reviewed by a person. Each row has the context and schema (input), the gold answer (reference), its pair family, and an evidence_certificate that records which sentences decide the answer. PACT uses the certificates to rebuild the evidence-removed contexts for L_NR. data/manifest.json has the counts and hashes that preflight checks. The holdout is used only for final scoring.

The nimble/ package contains unmodified copies of the upstream modules that define the prompt format, so prompts cannot drift from what the adapters were trained with.

Limitations

  • The holdout is small (324 examples from six source families) and synthetic. L_CF and L_NR depend on the same model-written certificates as the labels, so certificate errors carry over.
  • PACT-RL learns to agree with a fixed heuristic. It does not discover strategy beyond what the lookahead ranks highly, and the game engine, not the model, does all the spatial reasoning.
  • The RL adapter is specialized for the Tetris prompt; use the PACT adapter for general typed decisions.
  • Permutation ensembling multiplies inference cost by K; all headline numbers use K = 1.

Citation

@misc{lin2026pact,
  author = {Yida Lin},
  title  = {{PACT}: Pairwise-Anchored Calibrated Tuning for Single-Token Typed Decisions},
  year   = {2026},
  url    = {https://github.com/BennyLinntu/PACT-Pairwise-Anchored-Calibrated-Tuning-for-Single-Token-Typed-Decisions}
}
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for BennyLin01/pact-rl-tetris

Finetuned
Qwen/Qwen3.5-9B
Adapter
(763)
this model