Model collection · Benchmark code and results · Tournament software
Results · Insights · Training · Architecture · Use the model
A compact, 26-character Melee policy refined through Fox-only reinforcement learning.
Task: Melee control from structured game state. Faynt predicts the next GameCube controller command from game observations and recent controller history. One set of weights controls all 26 characters.
Offline use only. Do not use or adapt Faynt for Slippi Online. The Slippi Online rules prohibit macros and bots. Use local matches and offline research environments.
98.4%Supported mirrors240/244 games | 68/68Zero-delay Slippi-AIFox mirrors · two conditioning settings | 26CharactersOne shared checkpoint |
Results
Timing matters: Faynt has zero added policy delay; these opponents retain 21- or 24-frame action queues. The report does not isolate the effect of that difference. In supported mirrors, both agents use the same character from the opponent’s deployed roster. Extended-roster games keep the opponent on a supported fighter. Forced mirrors put both agents on a character outside the opponent’s deployed roster. Each Arena checkpoint plays 1,312 games across the three conditions.
The 10M checkpoint follows 1,950 RL steps across two runs; the 75M follows 980 steps with different training settings.
| Expanded-suite condition | 10M Arena | 75M Arena |
|---|---|---|
| Supported mirrors | 240/244 · 98.4% | 149/244 · 61.1% |
| Extended roster | 427/534 · 80.0% | 198/534 · 37.1% |
| Forced mirrors | 531/534 · 99.4% | 503/534 · 94.2% |
A separate zero-delay comparison
Both sides use zero added policy delay in these panels. The privately supplied Slippi-AI checkpoint is evaluated in Fox mirrors under “Master Player” and “Cody” conditioning.
Against the private zero-delay Slippi-AI checkpoint:
| Faynt stage | 10M | 75M |
|---|---|---|
| Expert | 23/68 | 25/68 |
| Arena | 68/68 | 58/68 |
Against the seven zero-delay Phillip specialists:
| Arena evaluation | 10M | 75M |
|---|---|---|
| Final Destination | 112/112 | 108/112 |
| Six stages | 126/126 | 126/126 |
The private Slippi-AI rows form a matched Expert-to-Arena comparison: 23 to 68 wins at 10M, and 25 to 58 at 75M. Phillip chooses an action every 2, 3, or 4 frames, depending on the specialist; Faynt chooses each frame. Several opponents in the wider benchmark also informed RL run or checkpoint selection.
Source: Faynt: Scaling and Optimizing Policies for Competitive Melee, Sections 5 and 6 and the benchmark appendices. The evaluation manifest records the specific source tables, denominators, and qualifications.
Insights
01 / Strong on opponents’ supported characters
The selected policy wins 240/244 supported mirrors and has a winning record against all 14 opponent releases. Both agents use the same fighter from the opponent’s deployed roster.
02 / 68 wins in 68 zero-delay Fox mirrors
On the matched private zero-delay Slippi-AI Fox-mirror panel, Expert wins 23/68 games and Arena wins 68/68 across two conditioning settings, including 61 four-stock wins. This smaller panel complements the delayed-release suite.
03 / A compact model with two RL runs
The 75M policy has about 7.4 times as many parameters as the 10M. The selected 10M lineage spans 1,950 RL steps across two runs, with different training settings from the 75M.
Training
Arena continues 10M Expert with PPO on Fox mirror matches across six stages. Its selected lineage is Expert 195,248 → first RL run 1,318 → second RL run 632, totaling 1,950 RL steps along that lineage. Step 632 counts the second run from its own start.
The second run uses self-play, an eight-second reward half-life, and forward/reverse penalties toward the supervised policy with weight 0.003 each. Rewards combine stock and damage advantage with movement and positioning terms. The checkpoint still supports all 26 characters; Fox supplies its RL experience.
| Stage | Training signal | Selected step |
|---|---|---|
| Base | Human replay pretraining | 122,064 |
| Expert | Curriculum + 75M distillation | 195,248 |
| Arena, run 1 | PPO on Fox mirrors | 1,318 |
| Arena, run 2 | Self-play with reference KL | 632 |
Data, objectives, and checkpoint selection
The fixed pretraining snapshot contains 839,942 human ranked replays, 17.84B valid targets, and 17.48B training targets from Melee Ranked Replays. A target is one player-perspective transition from frame t to t + 1. Both perspectives of a game stay in the same approximately 98/1/1 train/validation/test split. Deduplication covers both identical files and identical parsed training content.
Pretraining minimizes controller negative log-likelihood using 256-frame windows. Muon updates the backbone hidden matrices; auxiliary AdamW updates the remaining parameter groups. Corpus size and processed training targets describe different quantities because sampling can revisit frames.
Expert emphasizes winning demonstrations through a rank/outcome curriculum, then a mixture of approximately 90% Master-winner and 10% Diamond-winner replay visits. Checkpoint selection uses W = 0.9 × Master-winner NLL + 0.1 × Diamond-winner NLL on fixed held-out slices. The 10M Expert adds two distillation rounds with a frozen 75M Expert teacher, weight 0.5 and temperature 1.
Arena uses PPO with stock, damage, movement, and positioning reward terms. Its RL experience is restricted to Fox mirrors on six stages. The other 25 characters share the updated weights. The selected 10M and 75M policies have distinct RL schedules and exposure.
Architecture
The controller distribution factors into a 728-way joint-control choice, followed by an 85-way main-stick choice conditioned on that choice. The joint category includes buttons, shoulder pressure, and C-stick position. The codec converts both categories into a complete controller command.
Dimensions and architectural choices
| Component | Configuration |
|---|---|
| Trainable parameters | 10,163,629 |
| Model width / blocks | 384 / 5 |
| Query heads / key-value heads | 6 / 2 |
| Head dimension / feed-forward width | 64 / 768 |
| Native ring-cache capacity | 256 frames |
| Actor trajectory context | 128 frames for history/rollout bookkeeping |
| Stored weights / default compute / cache | FP32 / FP32 / FP32 |
| Prediction offset | One frame |
Learned character, action-state, and character/action embeddings represent the state categories. A shared item MLP and masked sum combine item features. Causal grouped-query attention processes the frame history with RoPE positions, query/key RMS normalization, and elementwise gated attention outputs. Full Attention Residuals learn how to mix earlier depth representations at each frame. RMSNorm and SwiGLU complete the backbone.
Use
python -m pip install "torch>=2.5" "transformers>=4.57,<5" safetensors "PyYAML>=6"
import torch
from transformers import AutoModel
model = AutoModel.from_pretrained(
"frisson-labs/Faynt-10M-Arena",
trust_remote_code=True,
).eval()
batch = model.example_inputs(batch_size=1, sequence_length=1)
with torch.inference_mode():
output = model.sample(**batch, temperature=1.0)
print(output.controller.as_packed_tensor().shape) # torch.Size([1, 1, 13])
example_inputs creates synthetic structured states for a loading check. Live play requires a Slippi/Dolphin adapter that parses the game, assigns player perspective, synchronizes frames, and executes commands. trust_remote_code=True loads the custom model source included here. Authenticate with hf auth login when repository access requires it.
Streaming inference, game resets, and the 256-frame ring cache
The saved configuration uses a 256-frame continuous ring cache, temperature 1, FP32 compute/cache, and zero added policy delay. The separate 128-frame actor trajectory context describes history/rollout bookkeeping. Load the saved configuration directly:
import torch
from transformers import AutoModel
repo_id = "frisson-labs/Faynt-10M-Arena"
model = AutoModel.from_pretrained(
repo_id, trust_remote_code=True,
).eval()
cache = model.init_cache(batch_size=1)
frame = model.example_inputs(batch_size=1, sequence_length=None)
current_controller = frame["controller_t"]
with torch.inference_mode():
for index in range(3):
# Three synthetic frames; index 2 starts another game.
new_game = index in (0, 2)
if new_game:
current_controller = frame["controller_t"]
output = model.step(
game_state_t=frame["game_state_t"],
controller_t=current_controller,
cache=cache,
reset_mask=torch.tensor([new_game], device=model.device),
temperature=1.0,
)
current_controller = output.controller
print(cache.valid_length.item()) # 1, then 2, then 1
Supply a fresh observed state on each live iteration and feed back the controller actually executed. Reset the appropriate batch slots when a new game begins. The cache updates in place and stays within its capacity. Reproducing the reported match scores also requires the complete opponent, character, stage, port, seed, delay, and execution settings.
Input and output reference
| Field | Shape and meaning |
|---|---|
game_state_t.p0 / .p1 |
Controlled player / opponent: character, action, position, damage, shield, jumps, facing, controller and companion state. |
| Stage / platforms / items | Stage category, Randall and Fountain of Dreams platform coordinates, and 15 ordered item slots. |
controller_t |
The current executed controller as a structured record or codec labels. |
| Sequence tensors | Scalar fields [batch, time]; item fields [batch, time, 15]. |
step tensors |
Scalar fields [batch]; item fields [batch, 15]. |
output.logits |
buttons: [..., 728]; main_stick: [..., 85]. |
output.controller |
Decoded logical controller. as_packed_tensor() returns [..., 13]. |
| Packed order | Main x/y, C-stick x/y, shoulder, A, B, X, Y, Z, L, R, D_UP. Stick coordinates use [0, 1]. |
model(...) performs deterministic or teacher-forced prediction. model.sample(...) samples commands. model.step(...) processes one frame using a rolling cache. model.policy exposes the native encoder, backbone, controller head, and loss. The complete field definitions are in tensor_batch.py and controller_codec.py.
Scope and reproducibility
Use this checkpoint for structured-state game-agent research, replay prediction, training-stage comparisons, and further adaptation. Results depend on the evaluated opponent and character distribution, the timing contract, and the selected training history. Post-training changes the demonstration distribution; Arena receives Fox-only RL experience. In the expanded suite, each additional-character matchup has two games, so character-level results have limited samples.
The report’s optimized T4 decision loop averages 5.2 ms at 10M and 8.7 ms at 75M, excluding emulator execution and communication. Those measurements use recorded-state inputs, random weights, and a separate optimized runtime with CUDA graphs. The Transformers examples above provide the portable inference interface and have their own runtime performance.
Release files and verification
| File | Purpose |
|---|---|
| model.safetensors | Policy tensors in their original FP32 values. |
| config.json | Complete model configuration and AutoModel mapping. |
| modeling_faynt.py, configuration_faynt.py | Transformers interface. |
| faynt_native.py, controller_codec.py, tensor_batch.py | Native policy, codec, and structured tensor contract. |
| checkpoint.pt | Original checkpoint, retained byte for byte. |
| provenance.json, runtime_provenance.json | Training lineage and native source hashes. |
| card-evaluation.json | Reported card metrics and their evaluation context. |
The release was checked for exact tensor equality, strict native state loading, local AutoModel loading, and native forward-output parity. Sampling, cached inference, and independent game resets were also exercised. The uploaded safetensors SHA-256 matches the verified local file.
model.safetensors SHA-256
d3a841e6f18eb86ed86d1e94ad0ffc8decebf8dfbb26cc5e670c0fd68a5409b2
Validated with PyTorch 2.13.0, Transformers 4.57.6, and safetensors 0.8.0. The requirements record the supported dependency range.
Acknowledgments
We thank the Slippi-AI developers for the open-source tools, policy representation and learning methods that this work builds on, and for privately supplying the zero-delay checkpoint used in our evaluations. We thank Project Slippi for its replay infrastructure, the Slippi/ranked community for the original anonymized replay collections, and Erick Martinez for preparing and hosting the Melee Ranked Replays redistribution used in this work.
Explore the family
Base learns from human replays. Expert concentrates on high-ranked winning play, with teacher distillation for 10M. Arena continues with gameplay rewards.
Frisson Labs · Faynt
Model code and weights: MIT license · Third-party notices. Melee and the emulator are obtained separately under their respective terms.
- Downloads last month
- 5
Model tree for frisson-labs/Faynt-10M-Arena
Base model
frisson-labs/Faynt-10M-Base

