executable-goals / README.md
Claire0730's picture
Add Yi-Ting Chen to the citation
cbcb80e verified
|
Raw History Blame Contribute Delete
5.04 kB
---
license: apache-2.0
library_name: pytorch
pipeline_tag: robotics
base_model: JayLee131/TraceGen
base_model_relation: finetune
tags:
- robotics
- manipulation
- maniskill
- trajectory-prediction
- imitation-learning
---
# Predicted Futures Are Not Enough: Learning Executable Goals for Robot Manipulation
[Code](https://github.com/Claire0730/executable-goals) · [Project page](https://claire0730.github.io/executable-goals/) · Paper (arXiv, soon)
Weights and frozen goal banks behind the paper. A 3D trace world model predicts one future per episode; the
**Entity-Level Goal Readout** turns that prediction into a single executable goal in SE(3), and a shared
**Pose-Native Executor** runs it closed loop at 20 Hz across five ManiSkill3 tasks.
## What is here
| Folder | Contents | Size |
|---|---|---|
| `planner/` | Three 3D Trace Planners fine-tuned from TraceGen: `mix4_realcam_n2400` (Rigid Readout source), `mix5_t2k_n3000` (entity branch), `mix5_t2k_gmap` (entity + map branch) | 0.30–0.31 GB each |
| `student/` | The three Pose-Native Executors, 804,002 parameters each: `mt5_rciid_gmpc_s0`, `mt5_rcfz_gmpc_s0`, `mt5_rcfz_t2k_s0` | 3.2 MB each |
| `teacher/` | Five privileged PPO teachers, one per task | ~4 MB each |
| `banks/` | 49 frozen goal banks — one SE(3) goal per scene, for every reported row, task and evaluation seed | 58 MB |
Digests: `SHA256SUMS` (weights) and `banks/SHA256SUMS`.
## Use
Reproducing the main table needs the executors and the banks only — no planner, no gated licence, one GPU:
```bash
git clone https://github.com/Claire0730/executable-goals && cd executable-goals
hf download Claire0730/executable-goals --local-dir checkpoints --include "student/*" --include "teacher/*" --include "SHA256SUMS"
hf download Claire0730/executable-goals --local-dir . --include "banks/*"
bash scripts/00_link_checkpoints.sh
ROW=final bash scripts/verify_main_table.sh 999
```
Everything else — the full install, the inference and training chains, which record backs which number — is in the
code repository's README and `docs/REPRODUCTION.md`.
## Results
Closed-loop success, %, PickCube / LiftPegUpright / PegInsertionSide / StackCube / PushCube. The readout rows share
one set of executor weights; only the goal pipeline differs. "Recorded" pools evaluation seeds 999 / 997 / 998
(768 episodes per task).
| Row | Executor | Printed | Recorded |
|---|---|---|---|
| Rigid Readout K = 1 | `mt5_rcfz_gmpc_s0` | 29.80 / 70.18 / 20.05 / 52.47 / 99.74 (mean 54.45) | 29.82 / 70.18 / 20.05 / 52.47 / 99.74 |
| Rigid Readout K = 4 | `mt5_rcfz_gmpc_s0`, PickCube cell `mt5_rcfz_t2k_s0` | 46.48 / 76.95 / 21.88 / 54.82 / 99.22 (mean 59.87) | 33.07 / 76.95 / 21.88 / 54.82 / 99.22 with one executor; the printed PickCube cell is the second, at seed 999 |
| **Entity-Level Goal Readout** | `mt5_rciid_gmpc_s0` | **81.50 / 98.35 / 32.84 / 86.54 / 99.20 (mean 79.69)** | 81.51 / 98.31 / 32.81 / 86.46 / 99.22 |
| Oracle Goal | `mt5_rcfz_gmpc_s0` | 94.66 / 97.27 / 36.72 / 87.11 / 99.74 | identical |
Terminal goal position error (mm, seed 999, PickCube / PegInsertionSide / StackCube): Rigid K = 1 41.1 / 42.6 / 28.9,
Rigid K = 4 37.2 / 40.8 / 25.4, Entity-Level Goal Readout 31.5 / 29.4 / 12.8.
## Before you rely on these
- **The planner files hold trained parameters only.** Every frozen-encoder tensor was bitwise identical to the
published Hub weights and was removed, along with the optimizer state. The encoders are re-created from the Hub
when the planner is built, so **DINOv3 is required and gated**: accept its licence and run `hf auth login` once.
`verify_main_table.sh` builds no planner and needs none of this.
- **Training paths inside the planner `config` are placeholders.** `dataset_dirs`, `cache_dir` and `checkpoint_dir`
read `[path-to-the-repository-root-here]/...`. They are metadata — inference and evaluation never read them — but
set them to real directories before resuming training.
- **PegInsertionSide is not bitwise reproducible.** Repeated evaluation of the same checkpoint moves by a couple of
episodes in 256, which is why the verification script compares within ±0.04.
- **Simulation only, one training seed.** The paper's real-robot results are not established for these checkpoints
by this release, and the three evaluation seeds vary scenes, not training.
- Nothing here is a third-party weight. CoTracker3 and SAM 2 are not redistributed; the code repository's README
lists every external model, its licence and which stage needs it.
## License and citation
Apache-2.0. The released weights contain only parameters trained by the authors; the upstream TraceGen Generalist
checkpoint they are fine-tuned from is subject to its own terms.
```bibtex
@article{chuang2026executablegoals,
title = {Predicted Futures Are Not Enough: Learning Executable Goals for Robot Manipulation},
author = {Chuang, Tzu-Yu and Chang, Ching-Hsiang and Lee, Yi-Hsiu and Chen, Yi-Ting and Sun, Min and Yang, YuanFu},
year = {2026}
}
```