OTRetarget: Joint Robot and Object Motion Retargeting via Optimal Transport
Abstract
Transferring human motion to humanoid robots requires adapting the demonstrated motion to the robot morphology while preserving interactions with the environment. This is particularly challenging for loco-manipulation tasks, where contacts with the ground and manipulated objects must remain consistent despite differences in body proportions. Yet, skeletal motion alone does not fully describe these interactions, and fixing object trajectories limits the adaptation to a new embodiment. In this paper, we introduce OTR ETARGET, a unified approach to jointly retarget robot and multi-object motion from human demonstrations. Our approach represents surface interactions through signed distances, closest surface points, and relative directions, and uses entropic optimal transport to transfer these quantities across human, robot, and object geometries. We incorporate the resulting interaction targets into a constrained inverse kinematics formulation that balances contact preservation with motion style and jointly optimizes robot and object poses at each frame. This formulation accommodates robot-object and object-object interactions without rescaling the scene or the demonstration. We validate the proposed approach on OMOMO, where it achieves a robot- object interaction Jaccard score of 87% and a depth error of 8.7 mm, compared with 28% and 29.3 mm for OmniRetarget. Finally, we demonstrate transfer to a physical G1 humanoid using whole-body policies trained with reinforcement learning on the retargeted references, across motions including two-handed box pick-and-place onto a table.
Community
OTRetarget retargets human demonstrations to a humanoid robot together with the objects it manipulates. Contacts are described by signed distances, closest surface points and relative directions, transferred across human, robot and object geometries with entropic optimal transport, then enforced in a constrained IK that optimizes robot and object poses jointly at each frame, without rescaling the scene or the demonstration. On OMOMO, it reaches a robot-object interaction Jaccard of 87% and a depth error of 8.7 mm, against 28% and 29.3 mm for OmniRetarget. Whole-body RL policies trained on the retargeted references transfer to a real Unitree G1, including two-handed box pick-and-place onto a table.
Get this paper in your agent:
hf papers read 2609.36602 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper