Download README.md from Claire0730/executable-goals: direct link, hf CLI and curl.
- Browser
- Download file 5.02 kB
-
https://huggingface.co/Claire0730/executable-goals/resolve/main/README.md
- Command line
-
hf download hf://Claire0730/executable-goals/README.md
-
curl -L -o README.md https://huggingface.co/Claire0730/executable-goals/resolve/main/README.md
license: apache-2.0
library_name: pytorch
pipeline_tag: robotics
base_model: JayLee131/TraceGen
base_model_relation: finetune
tags:
- robotics
- manipulation
- maniskill
- trajectory-prediction
- imitation-learning
Predicted Futures Are Not Enough: Learning Executable Goals for Robot Manipulation
Code · Project page · Paper (arXiv, soon)
Weights and frozen goal banks behind the paper. A 3D trace world model predicts one future per episode; the Entity-Level Goal Readout turns that prediction into a single executable goal in SE(3), and a shared Pose-Native Executor runs it closed loop at 20 Hz across five ManiSkill3 tasks.
What is here
| Folder | Contents | Size |
|---|---|---|
planner/ |
Three 3D Trace Planners fine-tuned from TraceGen: mix4_realcam_n2400 (Rigid Readout source), mix5_t2k_n3000 (entity branch), mix5_t2k_gmap (entity + map branch) |
0.30–0.31 GB each |
student/ |
The three Pose-Native Executors, 804,002 parameters each: mt5_rciid_gmpc_s0, mt5_rcfz_gmpc_s0, mt5_rcfz_t2k_s0 |
3.2 MB each |
teacher/ |
Five privileged PPO teachers, one per task | ~4 MB each |
banks/ |
49 frozen goal banks — one SE(3) goal per scene, for every reported row, task and evaluation seed | 58 MB |
Digests: SHA256SUMS (weights) and banks/SHA256SUMS.
Use
Reproducing the main table needs the executors and the banks only — no planner, no gated licence, one GPU:
git clone https://github.com/Claire0730/executable-goals && cd executable-goals
hf download Claire0730/executable-goals --local-dir checkpoints --include "student/*" --include "teacher/*" --include "SHA256SUMS"
hf download Claire0730/executable-goals --local-dir . --include "banks/*"
bash scripts/00_link_checkpoints.sh
ROW=final bash scripts/verify_main_table.sh 999
Everything else — the full install, the inference and training chains, which record backs which number — is in the
code repository's README and docs/REPRODUCTION.md.
Results
Closed-loop success, %, PickCube / LiftPegUpright / PegInsertionSide / StackCube / PushCube. The readout rows share one set of executor weights; only the goal pipeline differs. "Recorded" pools evaluation seeds 999 / 997 / 998 (768 episodes per task).
| Row | Executor | Printed | Recorded |
|---|---|---|---|
| Rigid Readout K = 1 | mt5_rcfz_gmpc_s0 |
29.80 / 70.18 / 20.05 / 52.47 / 99.74 (mean 54.45) | 29.82 / 70.18 / 20.05 / 52.47 / 99.74 |
| Rigid Readout K = 4 | mt5_rcfz_gmpc_s0, PickCube cell mt5_rcfz_t2k_s0 |
46.48 / 76.95 / 21.88 / 54.82 / 99.22 (mean 59.87) | 33.07 / 76.95 / 21.88 / 54.82 / 99.22 with one executor; the printed PickCube cell is the second, at seed 999 |
| Entity-Level Goal Readout | mt5_rciid_gmpc_s0 |
81.50 / 98.35 / 32.84 / 86.54 / 99.20 (mean 79.69) | 81.51 / 98.31 / 32.81 / 86.46 / 99.22 |
| Oracle Goal | mt5_rcfz_gmpc_s0 |
94.66 / 97.27 / 36.72 / 87.11 / 99.74 | identical |
Terminal goal position error (mm, seed 999, PickCube / PegInsertionSide / StackCube): Rigid K = 1 41.1 / 42.6 / 28.9, Rigid K = 4 37.2 / 40.8 / 25.4, Entity-Level Goal Readout 31.5 / 29.4 / 12.8.
Before you rely on these
- The planner files hold trained parameters only. Every frozen-encoder tensor was bitwise identical to the
published Hub weights and was removed, along with the optimizer state. The encoders are re-created from the Hub
when the planner is built, so DINOv3 is required and gated: accept its licence and run
hf auth loginonce.verify_main_table.shbuilds no planner and needs none of this. - Training paths inside the planner
configare placeholders.dataset_dirs,cache_dirandcheckpoint_dirread[path-to-the-repository-root-here]/.... They are metadata — inference and evaluation never read them — but set them to real directories before resuming training. - PegInsertionSide is not bitwise reproducible. Repeated evaluation of the same checkpoint moves by a couple of episodes in 256, which is why the verification script compares within ±0.04.
- Simulation only, one training seed. The paper's real-robot results are not established for these checkpoints by this release, and the three evaluation seeds vary scenes, not training.
- Nothing here is a third-party weight. CoTracker3 and SAM 2 are not redistributed; the code repository's README lists every external model, its licence and which stage needs it.
License and citation
Apache-2.0. The released weights contain only parameters trained by the authors; the upstream TraceGen Generalist checkpoint they are fine-tuned from is subject to its own terms.
@article{chuang2026executablegoals,
title = {Predicted Futures Are Not Enough: Learning Executable Goals for Robot Manipulation},
author = {Chuang, Tzu-Yu and Chang, Ching-Hsiang and Lee, Yi-Hsiu and Sun, Min and Yang, YuanFu},
year = {2026}
}