HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone
Abstract
Learning deployable manipulation policies is bottlenecked by the scarcity of data that is both high-fidelity and scalable. Real-robot teleoperation is accurate but costly to scale; robot-free UMI capture scales readily, and current practice uses the resulting data mainly for pre-training, adding a small real-robot "anchor" at post-training. We ask whether raising the fidelity of robot-free UMI data, rather than shrinking the real-robot fraction, can remove that anchor. We present HiFi-UMI, a portable UMI data-production system co-designed for trajectory accuracy, inter-gripper relative pose, synchronization, and field of view: head-mounted offline stereo-inertial SLAM, native rather than reconstructed relative pose, a shared microsecond GPIO trigger, and two wide-angle cameras per hand covering ~200 degrees. It reaches 3 mm workspace-local end-effector accuracy without external tracking infrastructure. Using this corpus, we demonstrate zero-robot post-training: a policy post-trained solely on HiFi-UMI demonstrations deploys directly on a real robot and matches in-domain teleoperation across three backbones spanning the vision-language-action and world-action-model families, with success-rate differences of -2.5, +3.1, and -0.6 percentage points on StarVLA-QwenPI, OpenPI-pi_0.5, and LingBot-VA; the strongest policy reaches 85% on a precision insertion task, even though the teleoperation baseline is collected in the evaluation scene and no HiFi-UMI trajectory is. Pre-training on 4,000 hours from the same corpus lowers action error on ten unseen tasks by 41% and, on StarVLA-QwenPI, raises real-robot success by a further 18.1 percentage points. We open-source HiFi-UMI-2K, 2,000 hours of microsecond-synchronized, ultra-wide-FoV demonstrations, each automatically reconstructed and validated through simulation replay, as a large-scale, high-fidelity resource for the robot-learning community.
Community
Hi everyone! We’re excited to share HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone.
The central question is simple: can we eliminate target-task robot teleoperation from post-training, rather than merely reduce it?
HiFi-UMI is a portable, robot-free data-production system co-designed for action fidelity. It achieves 3 mm workspace-local end-effector accuracy, <40 μs cross-sensor synchronization, and ultra-wide six-view sensing, together with automated trajectory reconstruction, simulation replay, and quality validation.
Our main findings:
Across three VLA and WAM backbones—StarVLA-QwenPI, OpenPI-π0.5, and LingBot-VA—post-training using only HiFi-UMI demonstrations matches in-domain robot teleoperation, with success-rate differences of −2.5, +3.1, and −0.6 percentage points.
The strongest policy reaches 85% success on precision insertion, despite no HiFi-UMI demonstration being collected in the evaluation scene.
Pre-training on 4,000 hours reduces action error on ten unseen tasks by 41% and improves real-robot success by 18.1 percentage points.
We release HiFi-UMI-2K: 2,000 hours and 482K+ replayable demonstrations across 110+ scenes under CC BY 4.0.
Our key takeaway: robot-free data can support deployment—not only pre-training—when it is sufficiently high-fidelity and action-aligned.
📄 Paper:https://arxiv.org/abs/2607.25895
🌐 主页:https://cloud.simpleai.tech/simple-world-lab/hifi-umi/
🤗 数据集:https://huggingface.co/datasets/simple-world-lab/HiFi-UMI-2K
🔥 Daily Paper:https://huggingface.co/papers/2607.25895
Get this paper in your agent:
hf papers read 2607.25895 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 1
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper