Faynt-10M-Expert: Frisson Labs Melee policy, 10,163,629 parameters, 26 characters.

Model collection · Benchmark code and results · Tournament software

Results · Insights · Training · Architecture · Use the model

A compact Melee policy refined through curriculum learning and distillation from 75M Expert.

Task: Melee control from structured game state. Faynt predicts the next GameCube controller command from game observations and recent controller history. One set of weights controls all 26 characters.

Offline use only. Do not use or adapt Faynt for Slippi Online. The Slippi Online rules prohibit macros and bots. Use local matches and offline research environments.

69.7%

Initial-suite wins
106/152 games

0.76974

Test controller NLL
Nats per target

195,248

Selected step
Supervised optimizer counter

Results

The four Faynt checkpoints share opponents, character assignments, policy sampling seeds, and ports. Games use Final Destination, four stocks, and an eight-minute timer. The schedule mixes native-roster and roster-extension matchups. All initial-panel outcomes are verified. Faynt uses zero added policy delay; the reference agents retain their own timing settings.

Reported evaluation results for Faynt-10M-Expert, with explicit game counts and protocol conditions.

The chart compares the initial suite. Arena is evaluated in a separate expanded protocol.

Checkpoint Wins / games Win rate Test controller NLL ↓
10M Base 49/152 32.2% 0.79766
75M Base 69/152 45.4% 0.76543
10M Expert 106/152 69.7% 0.76974
75M Expert 123/152 80.9% 0.74206

Source: Faynt: Scaling and Optimizing Policies for Competitive Melee, Sections 5 and 6 and the benchmark appendices. The evaluation manifest records the specific source tables, denominators, and qualifications.

Insights

01 / 49 to 106 wins after supervised post-training

The initial-suite win rate rises from 32.2% at Base to 69.7% here, a gain of 37.5 percentage points. Both checkpoints use the same 152-game schedule.

02 / Winner-weighted loss falls during distillation

Two 900M-target distillation rounds reduce winner-weighted validation NLL from 0.76443 to 0.75746. Equal objective weights combine human actions with the 75M teacher’s action distributions.

03 / The starting policy and reference for Arena

This final supervised 10M checkpoint initializes Arena and remains its frozen reference during reinforcement learning.

Training

Expert continues Base through a curriculum that progressively favors winning, higher-ranked demonstrations. A second phase samples approximately 90% Master-winner and 10% Diamond-winner replay visits. Two subsequent distillation rounds transfer action distributions from 75M Expert while retaining recorded-action supervision, with teacher weight 0.5 and temperature 1.

The selected checkpoint is step 195,248. It is the supervised endpoint before RL and a compact starting policy for reinforcement-learning experiments.

Stage Training signal Selected step
Base Human replay pretraining 122,064
Expert Curriculum + 75M distillation 195,248
Data, objectives, and checkpoint selection

The fixed pretraining snapshot contains 839,942 human ranked replays, 17.84B valid targets, and 17.48B training targets from Melee Ranked Replays. A target is one player-perspective transition from frame t to t + 1. Both perspectives of a game stay in the same approximately 98/1/1 train/validation/test split. Deduplication covers both identical files and identical parsed training content.

Pretraining minimizes controller negative log-likelihood using 256-frame windows. Muon updates the backbone hidden matrices; auxiliary AdamW updates the remaining parameter groups. Corpus size and processed training targets describe different quantities because sampling can revisit frames.

Expert emphasizes winning demonstrations through a rank/outcome curriculum, then a mixture of approximately 90% Master-winner and 10% Diamond-winner replay visits. Checkpoint selection uses W = 0.9 × Master-winner NLL + 0.1 × Diamond-winner NLL on fixed held-out slices. The 10M Expert adds two distillation rounds with a frozen 75M Expert teacher, weight 0.5 and temperature 1.

Arena uses PPO with stock, damage, movement, and positioning reward terms. Its RL experience is restricted to Fox mirrors on six stages. The other 25 characters share the updated weights. The selected 10M and 75M policies have distinct RL schedules and exposure.

Architecture

10M architecture: 2,091 frame features, 384-wide projection, 5 Transformer blocks, 728 joint-controller categories, then 85 main-stick categories.

The controller distribution factors into a 728-way joint-control choice, followed by an 85-way main-stick choice conditioned on that choice. The joint category includes buttons, shoulder pressure, and C-stick position. The codec converts both categories into a complete controller command.

Dimensions and architectural choices
Component Configuration
Trainable parameters 10,163,629
Model width / blocks 384 / 5
Query heads / key-value heads 6 / 2
Head dimension / feed-forward width 64 / 768
Native ring-cache capacity 256 frames
Actor trajectory context 128 frames for history/rollout bookkeeping
Stored weights / default compute / cache FP32 / FP32 / FP32
Prediction offset One frame

Learned character, action-state, and character/action embeddings represent the state categories. A shared item MLP and masked sum combine item features. Causal grouped-query attention processes the frame history with RoPE positions, query/key RMS normalization, and elementwise gated attention outputs. Full Attention Residuals learn how to mix earlier depth representations at each frame. RMSNorm and SwiGLU complete the backbone.

Use

python -m pip install "torch>=2.5" "transformers>=4.57,<5" safetensors "PyYAML>=6"
import torch
from transformers import AutoModel

model = AutoModel.from_pretrained(
    "frisson-labs/Faynt-10M-Expert",
    trust_remote_code=True,
).eval()

batch = model.example_inputs(batch_size=1, sequence_length=1)
with torch.inference_mode():
    output = model.sample(**batch, temperature=1.0)

print(output.controller.as_packed_tensor().shape)  # torch.Size([1, 1, 13])

example_inputs creates synthetic structured states for a loading check. Live play requires a Slippi/Dolphin adapter that parses the game, assigns player perspective, synchronizes frames, and executes commands. trust_remote_code=True loads the custom model source included here. Authenticate with hf auth login when repository access requires it.

Streaming inference, game resets, and the 256-frame ring cache

The saved configuration uses a 256-frame continuous ring cache, temperature 1, FP32 compute/cache, and zero added policy delay. The separate 128-frame actor trajectory context describes history/rollout bookkeeping. Load the saved configuration directly:

import torch
from transformers import AutoModel

repo_id = "frisson-labs/Faynt-10M-Expert"
model = AutoModel.from_pretrained(
    repo_id, trust_remote_code=True,
).eval()

cache = model.init_cache(batch_size=1)
frame = model.example_inputs(batch_size=1, sequence_length=None)
current_controller = frame["controller_t"]

with torch.inference_mode():
    for index in range(3):
        # Three synthetic frames; index 2 starts another game.
        new_game = index in (0, 2)
        if new_game:
            current_controller = frame["controller_t"]
        output = model.step(
            game_state_t=frame["game_state_t"],
            controller_t=current_controller,
            cache=cache,
            reset_mask=torch.tensor([new_game], device=model.device),
            temperature=1.0,
        )
        current_controller = output.controller
        print(cache.valid_length.item())  # 1, then 2, then 1

Supply a fresh observed state on each live iteration and feed back the controller actually executed. Reset the appropriate batch slots when a new game begins. The cache updates in place and stays within its capacity. Reproducing the reported match scores also requires the complete opponent, character, stage, port, seed, delay, and execution settings.

Input and output reference
Field Shape and meaning
game_state_t.p0 / .p1 Controlled player / opponent: character, action, position, damage, shield, jumps, facing, controller and companion state.
Stage / platforms / items Stage category, Randall and Fountain of Dreams platform coordinates, and 15 ordered item slots.
controller_t The current executed controller as a structured record or codec labels.
Sequence tensors Scalar fields [batch, time]; item fields [batch, time, 15].
step tensors Scalar fields [batch]; item fields [batch, 15].
output.logits buttons: [..., 728]; main_stick: [..., 85].
output.controller Decoded logical controller. as_packed_tensor() returns [..., 13].
Packed order Main x/y, C-stick x/y, shoulder, A, B, X, Y, Z, L, R, D_UP. Stick coordinates use [0, 1].

model(...) performs deterministic or teacher-forced prediction. model.sample(...) samples commands. model.step(...) processes one frame using a rolling cache. model.policy exposes the native encoder, backbone, controller head, and loss. The complete field definitions are in tensor_batch.py and controller_codec.py.

Scope and reproducibility

Use this checkpoint for structured-state game-agent research, replay prediction, training-stage comparisons, and further adaptation. Results depend on the evaluated opponent and character distribution, the timing contract, and the selected training history. Post-training changes the demonstration distribution; Arena receives Fox-only RL experience. In the expanded suite, each additional-character matchup has two games, so character-level results have limited samples.

The report’s optimized T4 decision loop averages 5.2 ms at 10M and 8.7 ms at 75M, excluding emulator execution and communication. Those measurements use recorded-state inputs, random weights, and a separate optimized runtime with CUDA graphs. The Transformers examples above provide the portable inference interface and have their own runtime performance.

Release files and verification
File Purpose
model.safetensors Policy tensors in their original FP32 values.
config.json Complete model configuration and AutoModel mapping.
modeling_faynt.py, configuration_faynt.py Transformers interface.
faynt_native.py, controller_codec.py, tensor_batch.py Native policy, codec, and structured tensor contract.
checkpoint.pt Original checkpoint, retained byte for byte.
provenance.json, runtime_provenance.json Training lineage and native source hashes.
card-evaluation.json Reported card metrics and their evaluation context.

The release was checked for exact tensor equality, strict native state loading, local AutoModel loading, and native forward-output parity. Sampling, cached inference, and independent game resets were also exercised. The uploaded safetensors SHA-256 matches the verified local file.

model.safetensors SHA-256
f56d7055cdcc9e2a4187f2ed9475d89fc51ec2c61ccbb7ae98f68ed9361d27b0

Validated with PyTorch 2.13.0, Transformers 4.57.6, and safetensors 0.8.0. The requirements record the supported dependency range.

Acknowledgments

We thank the Slippi-AI developers for the open-source tools, policy representation and learning methods that this work builds on, and for privately supplying the zero-delay checkpoint used in our evaluations. We thank Project Slippi for its replay infrastructure, the Slippi/ranked community for the original anonymized replay collections, and Erick Martinez for preparing and hosting the Melee Ranked Replays redistribution used in this work.

Explore the family

Base learns from human replays. Expert concentrates on high-ranked winning play, with teacher distillation for 10M. Arena continues with gameplay rewards.

Size Base Expert Arena
10M Base Expert · this model Arena
75M Base Expert Arena

Frisson Labs · Faynt

Model code and weights: MIT license · Third-party notices. Melee and the emulator are obtained separately under their respective terms.

Downloads last month
3
Safetensors
Model size
10.2M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for frisson-labs/Faynt-10M-Expert

Finetuned
(1)
this model
Finetunes
1 model

Dataset used to train frisson-labs/Faynt-10M-Expert

Collection including frisson-labs/Faynt-10M-Expert