Robotics
Safetensors
world-model
video-prediction

Mask2Real-WM checkpoints

Checkpoints of the paper Mask2Real-WM: Controllable Dexterous World Models via Segmentation Masks as a Sim-to-Real Bridge by Riccardo O. Feingold, Davide Liconti, Chenyu Yang and Robert K. Katzschmann.

Mask2Real-WM is an action-conditioned video world model for a dexterous robot hand (an ORCA hand on a Franka arm, seen from a side camera and a wrist camera). The model splits video prediction into two stages: WM1 predicts a future segmentation-mask video from the past frames and the actions, and WM2 renders the RGB video conditioned on those masks and the actions. The monolithic baselines predict RGB directly.

Checkpoints

Folder Model Training Training step Action statistics
wm1_cascade_r WM1 of Cascade-R real data only 15,000 real
wm1_cascade_s WM1 of Cascade-S simulation only 55,000 simulation
wm1_cascade_sr WM1 of Cascade-SR wm1_cascade_s, then LoRA (rank 16) on real data 45,000 (LoRA stage) simulation
wm2 WM2, shared by all cascades real data; ControlNet branch and UNet LoRA (rank 16) 70,000 simulation
mono_r Mono-R real data only 22,000 real
mono_s_base stage 1 of Mono-SR simulation only 55,000 simulation (RGB set)
mono_sr_57500 Mono-SR mono_s_base, then LoRA (rank 16) on real data 57,500 real
mono_sr_58000 Mono-SR mono_s_base, then LoRA (rank 16) on real data 58,000 real

Two Mono-SR checkpoints are provided because the evaluations used two: step 57,500 for the controllability evaluation and step 58,000 for the video-fidelity tables and the sine-sweep study. mono_r is the checkpoint of the controllability evaluation and the sine-sweep study.

Every folder contains

  • model.safetensors: the full model (UNet, VAE, image encoder, text encoder, action encoder and, for wm2, the ControlNet branch) as 32-bit floats in one file; 9.3 GB each, 12.0 GB for wm2. The LoRA-stage checkpoints include their base weights and the adapters.
  • model.safetensors.sha256: checksum of that file.
  • stat.json: the action statistics the model was trained with. Pass this file to the model at inference; several models use the simulation statistics, not those of the evaluation dataset.
  • resolved_config.yaml: the fully resolved training configuration.

All models work on two camera views at 135x240 and 5 fps with 23-dimensional actions, take 5 history frames and predict 5 future frames per step.

Usage

git clone --single-branch --branch main https://github.com/srl-ethz/Mask2Real-WM.git
cd Mask2Real-WM
hf download riccardofeingold/Mask2Real-WM --local-dir checkpoints

The repository's README describes installation, inference, training and the evaluations. The models are built on the Stable Video Diffusion and CLIP weights, which the code downloads from the Hugging Face Hub; Stable Video Diffusion requires accepting its license there.

License

The weights are derivative works of Stable Video Diffusion Image-to-Video by Stability AI and are distributed under the Stability AI Community License; see also the NOTICE file, which lists the changes. The files also contain the text encoder of CLIP ViT-B/32 by OpenAI (MIT License). The code is released separately under the MIT License.

Powered by Stability AI.

Acknowledgements

The code builds on Ctrl-World (Guo, Shi, Chen, Finn, arXiv:2510.10125).

Citation

@article{feingold2026mask2realwm,
  title   = {Mask2Real-WM: Controllable Dexterous World Models via Segmentation Masks as a Sim-to-Real Bridge},
  author  = {Feingold, Riccardo O. and Liconti, Davide and Yang, Chenyu and Katzschmann, Robert K.},
  journal = {arXiv preprint arXiv:2607.04546},
  year    = {2026}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Model tree for riccardofeingold/Mask2Real-WM

Finetuned
(6)
this model

Papers for riccardofeingold/Mask2Real-WM