MoSE3: Learning World-Space SE(3) at Every Pixel

Paper | Project page | Code

MoSE3 is a feed-forward model that predicts world-space SE(3) motion for every pixel of a monocular RGB video, generalizing across rigid, articulated and deformable objects.

Usage

from mose3.models.mose3 import MoSE3

model = MoSE3.from_pretrained("JoannaCCCCCC/MoSE3", strict=True).cuda().eval()

The code repository has the full inference and visualization scripts.

License

The weights are released under CC BY-NC 4.0 (non-commercial). They include the weights of π³, which are released under the same license.

Citation

@misc{cheng2026mose3learningworldspacese3,
      title={MoSE3: Learning World-Space SE(3) at Every Pixel},
      author={Jiahuan Cheng and Zhiyi Li and Tian Xia and Ruojin Cai and Yilun Du and Qianqian Wang},
      year={2026},
      eprint={2610.03716},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2610.03716},
}
Downloads last month
31
Safetensors
Model size
2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for JoannaCCCCCC/MoSE3