Mask2Real-WM checkpoints
Checkpoints of the paper Mask2Real-WM: Controllable Dexterous World Models via Segmentation Masks as a Sim-to-Real Bridge by Riccardo O. Feingold, Davide Liconti, Chenyu Yang and Robert K. Katzschmann.
Mask2Real-WM is an action-conditioned video world model for a dexterous robot hand (an ORCA hand on a Franka arm, seen from a side camera and a wrist camera). The model splits video prediction into two stages: WM1 predicts a future segmentation-mask video from the past frames and the actions, and WM2 renders the RGB video conditioned on those masks and the actions. The monolithic baselines predict RGB directly.
- Code, configs and evaluation scripts: https://github.com/srl-ethz/Mask2Real-WM
- Paper: https://arxiv.org/abs/2607.04546
- Project page: https://srl-ethz.github.io/Mask2Real-WM/
Checkpoints
| Folder | Model | Training | Training step | Action statistics |
|---|---|---|---|---|
wm1_cascade_r |
WM1 of Cascade-R | real data only | 15,000 | real |
wm1_cascade_s |
WM1 of Cascade-S | simulation only | 55,000 | simulation |
wm1_cascade_sr |
WM1 of Cascade-SR | wm1_cascade_s, then LoRA (rank 16) on real data |
45,000 (LoRA stage) | simulation |
wm2 |
WM2, shared by all cascades | real data; ControlNet branch and UNet LoRA (rank 16) | 70,000 | simulation |
mono_r |
Mono-R | real data only | 22,000 | real |
mono_s_base |
stage 1 of Mono-SR | simulation only | 55,000 | simulation (RGB set) |
mono_sr_57500 |
Mono-SR | mono_s_base, then LoRA (rank 16) on real data |
57,500 | real |
mono_sr_58000 |
Mono-SR | mono_s_base, then LoRA (rank 16) on real data |
58,000 | real |
Two Mono-SR checkpoints are provided because the evaluations used two: step 57,500 for the
controllability evaluation and step 58,000 for the video-fidelity tables and the sine-sweep
study. mono_r is the checkpoint of the controllability evaluation and the sine-sweep study.
Every folder contains
model.safetensors: the full model (UNet, VAE, image encoder, text encoder, action encoder and, forwm2, the ControlNet branch) as 32-bit floats in one file; 9.3 GB each, 12.0 GB forwm2. The LoRA-stage checkpoints include their base weights and the adapters.model.safetensors.sha256: checksum of that file.stat.json: the action statistics the model was trained with. Pass this file to the model at inference; several models use the simulation statistics, not those of the evaluation dataset.resolved_config.yaml: the fully resolved training configuration.
All models work on two camera views at 135x240 and 5 fps with 23-dimensional actions, take 5 history frames and predict 5 future frames per step.
Usage
git clone --single-branch --branch main https://github.com/srl-ethz/Mask2Real-WM.git
cd Mask2Real-WM
hf download riccardofeingold/Mask2Real-WM --local-dir checkpoints
The repository's README describes installation, inference, training and the evaluations. The models are built on the Stable Video Diffusion and CLIP weights, which the code downloads from the Hugging Face Hub; Stable Video Diffusion requires accepting its license there.
License
The weights are derivative works of Stable Video Diffusion Image-to-Video by Stability AI and are distributed under the Stability AI Community License; see also the NOTICE file, which lists the changes. The files also contain the text encoder of CLIP ViT-B/32 by OpenAI (MIT License). The code is released separately under the MIT License.
Powered by Stability AI.
Acknowledgements
The code builds on Ctrl-World (Guo, Shi, Chen, Finn, arXiv:2510.10125).
Citation
@article{feingold2026mask2realwm,
title = {Mask2Real-WM: Controllable Dexterous World Models via Segmentation Masks as a Sim-to-Real Bridge},
author = {Feingold, Riccardo O. and Liconti, Davide and Yang, Chenyu and Katzschmann, Robert K.},
journal = {arXiv preprint arXiv:2607.04546},
year = {2026}
}
Model tree for riccardofeingold/Mask2Real-WM
Base model
stabilityai/stable-video-diffusion-img2vid