MoSE3: Learning World-Space SE(3) at Every Pixel
Paper • 2610.03716 • Published • 1
Paper | Project page | Code
MoSE3 is a feed-forward model that predicts world-space SE(3) motion for every pixel of a monocular RGB video, generalizing across rigid, articulated and deformable objects.
from mose3.models.mose3 import MoSE3
model = MoSE3.from_pretrained("JoannaCCCCCC/MoSE3", strict=True).cuda().eval()
The code repository has the full inference and visualization scripts.
The weights are released under CC BY-NC 4.0 (non-commercial). They include the weights of π³, which are released under the same license.
@misc{cheng2026mose3learningworldspacese3,
title={MoSE3: Learning World-Space SE(3) at Every Pixel},
author={Jiahuan Cheng and Zhiyi Li and Tian Xia and Ruojin Cai and Yilun Du and Qianqian Wang},
year={2026},
eprint={2610.03716},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2610.03716},
}