File size: 3,620 Bytes
2111e43 ee84263 2111e43 ee84263 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 | ---
license: apache-2.0
tags:
- magic-the-gathering
- draft
- game-ai
- zero-shot
- model-evaluation
- reproducibility
---
# DraftFM: zero-shot drafting of unseen Magic: The Gathering sets
DraftFM is a pick model for *Magic: The Gathering* drafting that represents
every card as a frozen vector of public information (structured Scryfall
features plus a sentence embedding of the rules text), with no card- or
set-specific parameters. It can therefore score a set the moment its card
list goes public, before any human has drafted it.
- **Paper:** [DraftFM: Zero-Shot Drafting of Unseen *Magic: The Gathering*
Sets from Public Card Features](https://github.com/brianward92/mtga/blob/main/paper/draftfm.pdf)
(Brian Ward, 2026; arXiv submission in progress)
- **Code:** https://github.com/brianward92/mtga
- **Per-pick prediction archive:** [brianward92/draftfm-frozen-eval](https://huggingface.co/datasets/brianward92/draftfm-frozen-eval)
## Headline results
Trained on 149.4M picks from 28 [17Lands](https://www.17lands.com) sets:
| Evaluation | Result |
|---|---|
| Three held-out dev sets, top-1 agreement with high-win-rate players | 54.3% |
| Same, as a fraction of a model trained directly on each set (pre-registered normalized score) | 78.6% |
| **MSH frozen evaluation** (licensed-IP set, untouched by training or tuning; single pre-registered pass over its first public snapshot) | **57.0% top-1**, 87.7% top-3, log-loss 1.132, ECE 0.005 |
The deployed recipe, F-full, has 1.6M parameters.
## What is in this repository
| Path | Contents |
|---|---|
| `runs/<run_id>/best.pt` | The 14 pinned PyTorch checkpoints evaluated in the paper: F-dev, F-full, scaling rungs s1–s16, and ablations (no-text, no-context, proportional, top-filter, no-UB) |
| `onnx/` | ONNX exports of the deployed model (fdev-20260704, f-full-20260705) |
| `run_manifest.json` | Role → run_id → checkpoint sha256. The authoritative pins: `make_paper_tables.py` refuses mismatched or missing runs |
| `frozen_battery.json` | The pre-registered evaluation battery (protocol v1.1), frozen before the MSH snapshot download |
| `ledger.jsonl` | Append-only experiment ledger |
| `paper-data/runs/` | Run-level JSONs (configs, per-epoch metrics, eval summaries) consumed by `scripts/make_paper_tables.py` |
Every checkpoint's sha256 is recorded in `run_manifest.json`; verify after
download. The frozen protocol, including the pre-registration chronology and
the post-day-one MSH ceiling, is documented in
[`docs/eval_protocol.md`](https://github.com/brianward92/mtga/blob/main/docs/eval_protocol.md).
## Reproducing the paper tables
```bash
git clone https://github.com/brianward92/mtga
cd mtga
python -m venv .venv && .venv/bin/pip install -r requirements-foundation.txt
# place this repo's paper-data/runs/ at paper/data/runs/
.venv/bin/python scripts/make_paper_tables.py
```
To re-score picks or run the battery, set `MTGA_DATA_ROOT` to a directory
with this repository's `runs/` under `foundation/` and see
`scripts/run_frozen_eval.py`.
## Data and licensing
The code is Apache-2.0 (see the GitHub repository's LICENSE and NOTICE).
Training and evaluation data derive from 17Lands public datasets
(CC BY 4.0; "Data from 17Lands.com") and Scryfall bulk data. No raw
third-party data is redistributed here: the checkpoints, exports, and
manifests are self-generated artifacts. Unofficial Fan Content per the
Wizards of the Coast Fan Content Policy; not approved or endorsed by
Wizards. *Magic: The Gathering* is a trademark of Wizards of the Coast LLC.
## Contact
Brian Ward — brian.ward.92@gmail.com
|