FlowCraft

Team
community
Activity Feed

AI & ML interests

None defined yet.

Recent Activity

aidanscannell  updated a Space 4 days ago
flow-craft/README
aidanscannell  published a Space 4 days ago
flow-craft/README
View all activity

Organization Card

FlowCraft

Action-conditioned video world models for Minecraft. FlowCraft models are causal diffusion transformers trained with flow matching in the latent space of a fine-tuned WAN 2.2 VAE. They generate autoregressively, one latent at a time, conditioned on the latents and actions so far.

This organisation holds the data, the VAEs and the model checkpoints from How to Train Your Video World Model for Long-Horizon Rollout (arXiv link coming soon). The paper compares pre-training recipes at equal compute, then post-trains three of them (Diffusion Forcing, clean-prefix teacher forcing and chunked teacher forcing): flow map finetuning followed by Self Forcing against the model's own pre-trained teacher. Every stage is evaluated on 1000-frame rollouts, far beyond the attention window.

How to Train Your Video World Model for Long-Horizon Rollout
Aidan Scannell*, Samuel Garcin*, Paul Chang, Peter Bell and Amos Storkey
Toyota Research Institute · University of Edinburgh · Verda. *Equal contribution.

Code: github.com/aidanscannell/flowcraft

What is here

Resolution Latents VAE Models
128×224 minecraft-latents-128-224-v2 (142 GB) wan-vae-minecraft-128-224 all 16 checkpoints
352×640 minecraft-latents-352-640-v2 (1.1 TB) wan-vae-minecraft-352-640 none yet

Collections: Models · VAEs · Datasets

The latents are the gameplay recordings from OpenAI's Video Pre-Training (VPT) contractor data, encoded once with the VAE of the same resolution, with the per-frame actions alongside. You never need the raw videos to train or evaluate.

Checkpoint names

{pre,post}-train-<recipe>[-quarter-budget]

  • pre-train: trained from scratch with the given recipe.
  • post-train: the final post-trained model. Stage 1 finetunes the pre-trained model into a few-step flow map; stage 2 is Self Forcing with distribution matching distillation, with the pre-trained model of the same recipe as the teacher (post-train-df comes from pre-train-df, and so on). Only stage 2 is released.
  • Budget: full is ~680 EFLOP (350k steps for single-stream recipes, 185k for the two-stream ones: pf, chunk, jump). -quarter-budget is 166 EFLOP (85k / 45k steps).
Recipe Paper name What it trains on
df Diffusion Forcing An independent flow time for every latent, with a loss on each. Uniform flow-time prior.
pf Clean-prefix teacher forcing Every latent against its own clean prefix, all in one pass over a doubled clean/noisy sequence.
chunk Chunked teacher forcing (B=4) As pf, but targets come in blocks of 4 latents, each seeing the clean prefix and the earlier members of its block at their own flow times.
jump Jump prediction (hmax=3) As pf, but up to 3 latents between a target and its clean prefix are deleted instead of noised.
ctx Context-length sampling One context length per clip: the latents before it clean, a loss on every latent after it.
ctx-notail ctx − tail As ctx, with a loss only on the first latent after the cut.

All recipes except df use a logit-normal flow-time prior. In pre-training, every recipe except df also uses history corruption: on a fraction p=0.1 of clips the clean history is presented at flow times τ ~ U[τmin, 1], with τmin=0.7. A name without a suffix uses these defaults, so pre-train-pf and pre-train-pf-quarter-budget differ only in budget. The suffixed variants sweep the history corruption (quarter budget only):

Suffix History corruption
pf-clean none (p=0)
pf p=0.1, τmin=0.7 (default)
pf-t05 p=0.1, τmin=0.5
pf-t00 p=0.1, τmin=0

Every checkpoint repo carries its full training config (config.json), the step, and the sha256 of the original training checkpoint (meta.json).

Quickstart

git clone https://github.com/aidanscannell/flowcraft && cd flowcraft
uv sync --extra gpu --extra cu130

# Evaluate a released model (downloads the weights, VAE and test split)
uv run fc-eval ckpt_path=flow-craft/pre-train-pf
from omegaconf import OmegaConf
from flowcraft.config import ExperimentConfig
from flowcraft.model import build_world_model
from flowcraft.train.checkpoint import CheckpointManager, load_checkpoint_config

ref = "flow-craft/pre-train-pf@iclr2027-submission"
cfg = OmegaConf.merge(OmegaConf.structured(ExperimentConfig), load_checkpoint_config(ref))
model = build_world_model(cfg)
CheckpointManager(cfg.train.ckpt, "/tmp/unused").load(ref, model)

Only need a few tasks of the data?

from huggingface_hub import snapshot_download

snapshot_download(
    "flow-craft/minecraft-latents-128-224-v2", repo_type="dataset",
    allow_patterns=["build-house/**", "manifest.json", "splits/**"],
)

Reproducing the paper

Every checkpoint has an iclr2027-submission tag: the exact version the paper evaluated. Pin it (@iclr2027-submission) for reproduction; main may move.

License

  • Code: MIT.
  • Checkpoints and VAEs: Apache-2.0. The VAEs are fine-tuned from WAN 2.2 (Apache-2.0), and every checkpoint contains one.
  • Datasets: our encoding and packaging are MIT. The gameplay footage comes from OpenAI's VPT contractor data and is subject to OpenAI's terms. Minecraft is a trademark of Mojang Synergies AB; this project is not affiliated with Mojang or Microsoft.

Citation

If you find FlowCraft, its checkpoints or its data useful, please consider citing us:

@article{scannell2026flowcraft,
  title   = {How to Train Your Video World Model for Long-Horizon Rollout},
  author  = {Scannell, Aidan and Garcin, Samuel and Chang, Paul and Bell, Peter and Storkey, Amos},
  journal = {arXiv preprint arXiv:<ARXIV-ID>},
  year    = {2026}
}

datasets 0

None public yet