AI & ML interests
None defined yet.
Recent Activity
FlowCraft
Action-conditioned video world models for Minecraft. FlowCraft models are causal diffusion transformers trained with flow matching in the latent space of a fine-tuned WAN 2.2 VAE. They generate autoregressively, one latent at a time, conditioned on the latents and actions so far.
This organisation holds the data, the VAEs and the model checkpoints from How to Train Your Video World Model for Long-Horizon Rollout (arXiv link coming soon). The paper compares pre-training recipes at equal compute, then post-trains three of them (Diffusion Forcing, clean-prefix teacher forcing and chunked teacher forcing): flow map finetuning followed by Self Forcing against the model's own pre-trained teacher. Every stage is evaluated on 1000-frame rollouts, far beyond the attention window.
How to Train Your Video World Model for Long-Horizon Rollout
Aidan Scannell*, Samuel Garcin*, Paul Chang, Peter Bell and Amos Storkey
Toyota Research Institute · University of Edinburgh · Verda. *Equal contribution.
Code: github.com/aidanscannell/flowcraft
What is here
| Resolution | Latents | VAE | Models |
|---|---|---|---|
| 128×224 | minecraft-latents-128-224-v2 (142 GB) |
wan-vae-minecraft-128-224 |
all 16 checkpoints |
| 352×640 | minecraft-latents-352-640-v2 (1.1 TB) |
wan-vae-minecraft-352-640 |
none yet |
Collections: Models · VAEs · Datasets
The latents are the gameplay recordings from OpenAI's Video Pre-Training (VPT) contractor data, encoded once with the VAE of the same resolution, with the per-frame actions alongside. You never need the raw videos to train or evaluate.
Checkpoint names
{pre,post}-train-<recipe>[-quarter-budget]
- pre-train: trained from scratch with the given recipe.
- post-train: the final post-trained model. Stage 1 finetunes the
pre-trained model into a few-step flow map; stage 2 is Self Forcing with
distribution matching distillation, with the pre-trained model of the same
recipe as the teacher (
post-train-dfcomes frompre-train-df, and so on). Only stage 2 is released. - Budget: full is ~680 EFLOP (350k steps for single-stream recipes, 185k for
the two-stream ones:
pf,chunk,jump).-quarter-budgetis 166 EFLOP (85k / 45k steps).
| Recipe | Paper name | What it trains on |
|---|---|---|
df |
Diffusion Forcing | An independent flow time for every latent, with a loss on each. Uniform flow-time prior. |
pf |
Clean-prefix teacher forcing | Every latent against its own clean prefix, all in one pass over a doubled clean/noisy sequence. |
chunk |
Chunked teacher forcing (B=4) | As pf, but targets come in blocks of 4 latents, each seeing the clean prefix and the earlier members of its block at their own flow times. |
jump |
Jump prediction (hmax=3) | As pf, but up to 3 latents between a target and its clean prefix are deleted instead of noised. |
ctx |
Context-length sampling | One context length per clip: the latents before it clean, a loss on every latent after it. |
ctx-notail |
ctx − tail | As ctx, with a loss only on the first latent after the cut. |
All recipes except df use a logit-normal flow-time prior. In pre-training,
every recipe except df also uses history corruption: on a fraction p=0.1 of clips the clean
history is presented at flow times τ ~ U[τmin, 1], with τmin=0.7.
A name without a suffix uses these defaults, so pre-train-pf and
pre-train-pf-quarter-budget differ only in budget. The suffixed variants
sweep the history corruption (quarter budget only):
| Suffix | History corruption |
|---|---|
pf-clean |
none (p=0) |
pf |
p=0.1, τmin=0.7 (default) |
pf-t05 |
p=0.1, τmin=0.5 |
pf-t00 |
p=0.1, τmin=0 |
Every checkpoint repo carries its full training config (config.json), the
step, and the sha256 of the original training checkpoint (meta.json).
Quickstart
git clone https://github.com/aidanscannell/flowcraft && cd flowcraft
uv sync --extra gpu --extra cu130
# Evaluate a released model (downloads the weights, VAE and test split)
uv run fc-eval ckpt_path=flow-craft/pre-train-pf
from omegaconf import OmegaConf
from flowcraft.config import ExperimentConfig
from flowcraft.model import build_world_model
from flowcraft.train.checkpoint import CheckpointManager, load_checkpoint_config
ref = "flow-craft/pre-train-pf@iclr2027-submission"
cfg = OmegaConf.merge(OmegaConf.structured(ExperimentConfig), load_checkpoint_config(ref))
model = build_world_model(cfg)
CheckpointManager(cfg.train.ckpt, "/tmp/unused").load(ref, model)
Only need a few tasks of the data?
from huggingface_hub import snapshot_download
snapshot_download(
"flow-craft/minecraft-latents-128-224-v2", repo_type="dataset",
allow_patterns=["build-house/**", "manifest.json", "splits/**"],
)
Reproducing the paper
Every checkpoint has an iclr2027-submission tag: the exact version the
paper evaluated. Pin it (@iclr2027-submission) for reproduction; main may
move.
License
- Code: MIT.
- Checkpoints and VAEs: Apache-2.0. The VAEs are fine-tuned from WAN 2.2 (Apache-2.0), and every checkpoint contains one.
- Datasets: our encoding and packaging are MIT. The gameplay footage comes from OpenAI's VPT contractor data and is subject to OpenAI's terms. Minecraft is a trademark of Mojang Synergies AB; this project is not affiliated with Mojang or Microsoft.
Citation
If you find FlowCraft, its checkpoints or its data useful, please consider citing us:
@article{scannell2026flowcraft,
title = {How to Train Your Video World Model for Long-Horizon Rollout},
author = {Scannell, Aidan and Garcin, Samuel and Chang, Paul and Bell, Peter and Storkey, Amos},
journal = {arXiv preprint arXiv:<ARXIV-ID>},
year = {2026}
}