--- tags: - robotics - world-model - action-conditioned-video-generation - robot-manipulation - coachworld pipeline_tag: robotics --- # CoachWorld — Pretrained Robot World Model [**Paper**](https://arxiv.org/abs/2609.39685) · [**Code**](https://github.com/RoboCoach-AI/CoachWorld) · [**Project Page & Videos**](https://robocoach-ai.github.io/) **CoachWorld predicts future robot observations conditioned on visual history and robot actions.** It is the world-model component of **RoboCoach: World Models as Active Coaches for Compositional Robot Skills**. This repository hosts the **CoachWorld pretrained checkpoint**, trained on heterogeneous single-arm and dual-arm robot data. ## Overview CoachWorld is an action-conditioned video world model built on an adapted Wan2.2 TI2V-5B backbone. It uses visual history and end-effector action conditioning to generate future observations, with camera and action conventions that support learning across robot datasets. The model supports autoregressive rollout: generated observations become part of the context for subsequent predictions. In the full RoboCoach framework, skill policies interact with the world model, while a separate progress judge identifies subtask failures. These diagnoses guide targeted demonstration collection and policy updates. ## Model at a Glance | Item | Description | | --- | --- | | Release | CoachWorld pretrain | | Model type | Action-conditioned video world model | | Backbone | Adapted Wan2.2 | | Conditioning | Visual history and robot end-effector actions, using the prescribed camera/action representation | | Output | Predicted future RGB observations | | Rollout mode | Autoregressive video prediction | | Training scope | Heterogeneous single-arm and dual-arm robot data | | Implementation | Custom `coachworld` Python package | ## Getting Started Download the weights from the [Files and versions](https://huggingface.co/JEdward/CoachWorld/tree/main) tab and use the [CoachWorld code repository](https://github.com/RoboCoach-AI/CoachWorld) for environment setup and model integration. The implementation includes: - Model conditioning, inference, and training components. - Modified Wan2.2 model and VAE components. - Video-latent data interfaces and action/camera conventions. - Camera and robot geometry helpers. - A distributed training entry point and example configuration. **Input preparation matters.** Robot actions must follow the implementation’s coordinate-frame, normalization, and temporal-sampling conventions. Raw joint commands or pixel-space trajectories are not interchangeable with the model’s expected conditioning. Use the CoachWorld implementation to load and run this checkpoint; this release does not claim compatibility with a generic Transformers or Diffusers loading pipeline. ## Citation If you use CoachWorld in your research, please cite: ```bibtex @article{liu2026robocoach, title={RoboCoach: World Models as Active Coaches for Compositional Robot Skills}, author={Liu, Jiajun and Chen, Yifan and Liu, Yichao and Zhang, Jiayi and Chen, Ruoqu and Xie, Shaoxuan and Yao, Guocai and Xu, Mengdi and Cui, Sen and Zhang, Changshui}, journal={arXiv preprint arXiv:2609.39685}, year={2026}, url={https://arxiv.org/abs/2609.39685} } ``` ## Acknowledgments CoachWorld builds on Wan2.2 and other open-source components. See the code repository’s [third-party notices](https://github.com/RoboCoach-AI/CoachWorld/blob/main/THIRD_PARTY.md) for component attribution and applicable terms. ## Contact For implementation questions and reproducibility issues, please open an issue in the [CoachWorld repository](https://github.com/RoboCoach-AI/CoachWorld/issues).