HERO — humanoid end-effector controller (Unitree G1 + Dex3)

Exported policy of HERO (Learning Humanoid End-Effector Control for Visual Whole-Body Open-Vocabulary Object Grasping, CoRL 2026): a whole-body controller that tracks world-fixed palm targets with the Unitree G1 humanoid and the Dex3 hand, trained in Isaac Sim with dual-actor PPO and residual end-effector feedback, exported to ONNX for deployment, browser inference and the MuJoCo benchmark.

file what it is
model.onnx the policy graph: two inputs (actor_obs_lower_body, actor_obs_upper_body, both fed the same 1,000-dim frame-major observation of 5 history frames), one output action (29 joint targets)
model_hero.json the metadata sidecar (hero_export_v1): observation layout, scales, history layout, action contract and joint tables, PD gains, odometry/anchor terms — everything the readers need to build the observation

This is the model bundled with the HERO release as checkpoints/example (the browser demo and the benchmark baseline use it). It uses the delta-anchor root feedback, so its observation layout (1,000 inputs) differs from a policy trained with the default without_delta_anchor configuration (675 inputs); the sidecar is authoritative, never assume a layout.

Usage

git clone https://github.com/RunpeiDong/HERO && cd HERO
pip install -e ".[bench]"
hf download RunpeiDong/HERO model.onnx model_hero.json --local-dir checkpoints/example
# score it on the benchmark (archive URL and sha256: see https://huggingface.co/datasets/RunpeiDong/hero_bench)
python scripts/hero_bench.py fetch --url https://huggingface.co/datasets/RunpeiDong/hero_bench/resolve/main/hero_bench_v1_corpus_33542601.tar.gz --sha256 33542601ed456ac606f8177ab75892d2e616585ec562f01976cb6af238c9c5d0 --out data/hero_bench_v1
python scripts/hero_bench.py run --corpus data/hero_bench_v1 --onnx-dir checkpoints/example --sidecar checkpoints/example/model_hero.json --out results/example --tiers core
python scripts/hero_bench.py report --corpus data/hero_bench_v1 --results results/example --card
# browser demo
python -m sim2sim.interactive_client.export_policy_assets --hero checkpoints/example/model.onnx --parity-fixture
python -m sim2sim.interactive_client.export_scene --hero checkpoints/example/model.onnx

Deployment contract (from the sidecar): 50 Hz policy, actions are residuals around the reference arm joints (hero_residual_upper_v1), joint targets = default pose + 0.25 · clip(action) with the arm reference added, PD gains and effort limits as listed in model_hero.json. See docs/benchmark.md and checkpoints/README.md in the code release.

Training

Trained with the released code (scripts/train.py, Isaac Sim 5.1 / Isaac Lab 2.3, 4,096 environments) on synthetic IK reaching references and retargeted human motion; exported with scripts/export.py. The paper's data recipe is described in docs/data.md of the code release.

License

MIT (same as the code release).

Citation

@inproceedings{dong2026hero,
  title     = {{HERO}: Learning Humanoid End-Effector Control for Visual Whole-Body Open-Vocabulary Object Grasping},
  author    = {Dong, Runpei and Li, Ziyan and Gupta, Arjun and He, Xialin and Gupta, Saurabh},
  booktitle = {10th Annual Conference on Robot Learning},
  year      = {2026},
  url       = {https://openreview.net/forum?id=gbchkYm28k}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Collection including RunpeiDong/HERO