executable-goals / README.md
Claire0730's picture
Condense the model card and link code, project page and paper
e4fa2cb verified
|
Raw History Blame Contribute Delete
5.02 kB
metadata
license: apache-2.0
library_name: pytorch
pipeline_tag: robotics
base_model: JayLee131/TraceGen
base_model_relation: finetune
tags:
  - robotics
  - manipulation
  - maniskill
  - trajectory-prediction
  - imitation-learning

Predicted Futures Are Not Enough: Learning Executable Goals for Robot Manipulation

Code · Project page · Paper (arXiv, soon)

Weights and frozen goal banks behind the paper. A 3D trace world model predicts one future per episode; the Entity-Level Goal Readout turns that prediction into a single executable goal in SE(3), and a shared Pose-Native Executor runs it closed loop at 20 Hz across five ManiSkill3 tasks.

What is here

Folder Contents Size
planner/ Three 3D Trace Planners fine-tuned from TraceGen: mix4_realcam_n2400 (Rigid Readout source), mix5_t2k_n3000 (entity branch), mix5_t2k_gmap (entity + map branch) 0.30–0.31 GB each
student/ The three Pose-Native Executors, 804,002 parameters each: mt5_rciid_gmpc_s0, mt5_rcfz_gmpc_s0, mt5_rcfz_t2k_s0 3.2 MB each
teacher/ Five privileged PPO teachers, one per task ~4 MB each
banks/ 49 frozen goal banks — one SE(3) goal per scene, for every reported row, task and evaluation seed 58 MB

Digests: SHA256SUMS (weights) and banks/SHA256SUMS.

Use

Reproducing the main table needs the executors and the banks only — no planner, no gated licence, one GPU:

git clone https://github.com/Claire0730/executable-goals && cd executable-goals
hf download Claire0730/executable-goals --local-dir checkpoints --include "student/*" --include "teacher/*" --include "SHA256SUMS"
hf download Claire0730/executable-goals --local-dir . --include "banks/*"
bash scripts/00_link_checkpoints.sh
ROW=final bash scripts/verify_main_table.sh 999

Everything else — the full install, the inference and training chains, which record backs which number — is in the code repository's README and docs/REPRODUCTION.md.

Results

Closed-loop success, %, PickCube / LiftPegUpright / PegInsertionSide / StackCube / PushCube. The readout rows share one set of executor weights; only the goal pipeline differs. "Recorded" pools evaluation seeds 999 / 997 / 998 (768 episodes per task).

Row Executor Printed Recorded
Rigid Readout K = 1 mt5_rcfz_gmpc_s0 29.80 / 70.18 / 20.05 / 52.47 / 99.74 (mean 54.45) 29.82 / 70.18 / 20.05 / 52.47 / 99.74
Rigid Readout K = 4 mt5_rcfz_gmpc_s0, PickCube cell mt5_rcfz_t2k_s0 46.48 / 76.95 / 21.88 / 54.82 / 99.22 (mean 59.87) 33.07 / 76.95 / 21.88 / 54.82 / 99.22 with one executor; the printed PickCube cell is the second, at seed 999
Entity-Level Goal Readout mt5_rciid_gmpc_s0 81.50 / 98.35 / 32.84 / 86.54 / 99.20 (mean 79.69) 81.51 / 98.31 / 32.81 / 86.46 / 99.22
Oracle Goal mt5_rcfz_gmpc_s0 94.66 / 97.27 / 36.72 / 87.11 / 99.74 identical

Terminal goal position error (mm, seed 999, PickCube / PegInsertionSide / StackCube): Rigid K = 1 41.1 / 42.6 / 28.9, Rigid K = 4 37.2 / 40.8 / 25.4, Entity-Level Goal Readout 31.5 / 29.4 / 12.8.

Before you rely on these

  • The planner files hold trained parameters only. Every frozen-encoder tensor was bitwise identical to the published Hub weights and was removed, along with the optimizer state. The encoders are re-created from the Hub when the planner is built, so DINOv3 is required and gated: accept its licence and run hf auth login once. verify_main_table.sh builds no planner and needs none of this.
  • Training paths inside the planner config are placeholders. dataset_dirs, cache_dir and checkpoint_dir read [path-to-the-repository-root-here]/.... They are metadata — inference and evaluation never read them — but set them to real directories before resuming training.
  • PegInsertionSide is not bitwise reproducible. Repeated evaluation of the same checkpoint moves by a couple of episodes in 256, which is why the verification script compares within ±0.04.
  • Simulation only, one training seed. The paper's real-robot results are not established for these checkpoints by this release, and the three evaluation seeds vary scenes, not training.
  • Nothing here is a third-party weight. CoTracker3 and SAM 2 are not redistributed; the code repository's README lists every external model, its licence and which stage needs it.

License and citation

Apache-2.0. The released weights contain only parameters trained by the authors; the upstream TraceGen Generalist checkpoint they are fine-tuned from is subject to its own terms.

@article{chuang2026executablegoals,
  title  = {Predicted Futures Are Not Enough: Learning Executable Goals for Robot Manipulation},
  author = {Chuang, Tzu-Yu and Chang, Ching-Hsiang and Lee, Yi-Hsiu and Sun, Min and Yang, YuanFu},
  year   = {2026}
}