File size: 3,620 Bytes
2111e43
 
ee84263
 
 
 
 
 
 
2111e43
ee84263
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
---
license: apache-2.0
tags:
  - magic-the-gathering
  - draft
  - game-ai
  - zero-shot
  - model-evaluation
  - reproducibility
---

# DraftFM: zero-shot drafting of unseen Magic: The Gathering sets

DraftFM is a pick model for *Magic: The Gathering* drafting that represents
every card as a frozen vector of public information (structured Scryfall
features plus a sentence embedding of the rules text), with no card- or
set-specific parameters. It can therefore score a set the moment its card
list goes public, before any human has drafted it.

- **Paper:** [DraftFM: Zero-Shot Drafting of Unseen *Magic: The Gathering*
  Sets from Public Card Features](https://github.com/brianward92/mtga/blob/main/paper/draftfm.pdf)
  (Brian Ward, 2026; arXiv submission in progress)
- **Code:** https://github.com/brianward92/mtga
- **Per-pick prediction archive:** [brianward92/draftfm-frozen-eval](https://huggingface.co/datasets/brianward92/draftfm-frozen-eval)

## Headline results

Trained on 149.4M picks from 28 [17Lands](https://www.17lands.com) sets:

| Evaluation | Result |
|---|---|
| Three held-out dev sets, top-1 agreement with high-win-rate players | 54.3% |
| Same, as a fraction of a model trained directly on each set (pre-registered normalized score) | 78.6% |
| **MSH frozen evaluation** (licensed-IP set, untouched by training or tuning; single pre-registered pass over its first public snapshot) | **57.0% top-1**, 87.7% top-3, log-loss 1.132, ECE 0.005 |

The deployed recipe, F-full, has 1.6M parameters.

## What is in this repository

| Path | Contents |
|---|---|
| `runs/<run_id>/best.pt` | The 14 pinned PyTorch checkpoints evaluated in the paper: F-dev, F-full, scaling rungs s1–s16, and ablations (no-text, no-context, proportional, top-filter, no-UB) |
| `onnx/` | ONNX exports of the deployed model (fdev-20260704, f-full-20260705) |
| `run_manifest.json` | Role → run_id → checkpoint sha256. The authoritative pins: `make_paper_tables.py` refuses mismatched or missing runs |
| `frozen_battery.json` | The pre-registered evaluation battery (protocol v1.1), frozen before the MSH snapshot download |
| `ledger.jsonl` | Append-only experiment ledger |
| `paper-data/runs/` | Run-level JSONs (configs, per-epoch metrics, eval summaries) consumed by `scripts/make_paper_tables.py` |

Every checkpoint's sha256 is recorded in `run_manifest.json`; verify after
download. The frozen protocol, including the pre-registration chronology and
the post-day-one MSH ceiling, is documented in
[`docs/eval_protocol.md`](https://github.com/brianward92/mtga/blob/main/docs/eval_protocol.md).

## Reproducing the paper tables

```bash
git clone https://github.com/brianward92/mtga
cd mtga
python -m venv .venv && .venv/bin/pip install -r requirements-foundation.txt
# place this repo's paper-data/runs/ at paper/data/runs/
.venv/bin/python scripts/make_paper_tables.py
```

To re-score picks or run the battery, set `MTGA_DATA_ROOT` to a directory
with this repository's `runs/` under `foundation/` and see
`scripts/run_frozen_eval.py`.

## Data and licensing

The code is Apache-2.0 (see the GitHub repository's LICENSE and NOTICE).
Training and evaluation data derive from 17Lands public datasets
(CC BY 4.0; "Data from 17Lands.com") and Scryfall bulk data. No raw
third-party data is redistributed here: the checkpoints, exports, and
manifests are self-generated artifacts. Unofficial Fan Content per the
Wizards of the Coast Fan Content Policy; not approved or endorsed by
Wizards. *Magic: The Gathering* is a trademark of Wizards of the Coast LLC.

## Contact

Brian Ward — brian.ward.92@gmail.com