File size: 5,037 Bytes
10f6b6a
 
 
e4fa2cb
 
 
 
 
 
 
 
 
10f6b6a
 
e4fa2cb
10f6b6a
e4fa2cb
10f6b6a
e4fa2cb
 
 
10f6b6a
e4fa2cb
10f6b6a
e4fa2cb
 
 
 
 
 
10f6b6a
e4fa2cb
10f6b6a
e4fa2cb
10f6b6a
e4fa2cb
10f6b6a
e4fa2cb
 
 
 
 
 
 
10f6b6a
e4fa2cb
 
10f6b6a
e4fa2cb
10f6b6a
e4fa2cb
 
 
10f6b6a
e4fa2cb
 
 
 
 
 
10f6b6a
e4fa2cb
 
10f6b6a
e4fa2cb
10f6b6a
e4fa2cb
 
 
 
 
 
 
 
 
 
 
 
 
10f6b6a
e4fa2cb
10f6b6a
e4fa2cb
 
10f6b6a
 
 
 
cbcb80e
10f6b6a
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
---
license: apache-2.0
library_name: pytorch
pipeline_tag: robotics
base_model: JayLee131/TraceGen
base_model_relation: finetune
tags:
  - robotics
  - manipulation
  - maniskill
  - trajectory-prediction
  - imitation-learning
---

# Predicted Futures Are Not Enough: Learning Executable Goals for Robot Manipulation

[Code](https://github.com/Claire0730/executable-goals) · [Project page](https://claire0730.github.io/executable-goals/) · Paper (arXiv, soon)

Weights and frozen goal banks behind the paper. A 3D trace world model predicts one future per episode; the
**Entity-Level Goal Readout** turns that prediction into a single executable goal in SE(3), and a shared
**Pose-Native Executor** runs it closed loop at 20 Hz across five ManiSkill3 tasks.

## What is here

| Folder | Contents | Size |
|---|---|---|
| `planner/` | Three 3D Trace Planners fine-tuned from TraceGen: `mix4_realcam_n2400` (Rigid Readout source), `mix5_t2k_n3000` (entity branch), `mix5_t2k_gmap` (entity + map branch) | 0.30–0.31 GB each |
| `student/` | The three Pose-Native Executors, 804,002 parameters each: `mt5_rciid_gmpc_s0`, `mt5_rcfz_gmpc_s0`, `mt5_rcfz_t2k_s0` | 3.2 MB each |
| `teacher/` | Five privileged PPO teachers, one per task | ~4 MB each |
| `banks/` | 49 frozen goal banks — one SE(3) goal per scene, for every reported row, task and evaluation seed | 58 MB |

Digests: `SHA256SUMS` (weights) and `banks/SHA256SUMS`.

## Use

Reproducing the main table needs the executors and the banks only — no planner, no gated licence, one GPU:

```bash
git clone https://github.com/Claire0730/executable-goals && cd executable-goals
hf download Claire0730/executable-goals --local-dir checkpoints --include "student/*" --include "teacher/*" --include "SHA256SUMS"
hf download Claire0730/executable-goals --local-dir . --include "banks/*"
bash scripts/00_link_checkpoints.sh
ROW=final bash scripts/verify_main_table.sh 999
```

Everything else — the full install, the inference and training chains, which record backs which number — is in the
code repository's README and `docs/REPRODUCTION.md`.

## Results

Closed-loop success, %, PickCube / LiftPegUpright / PegInsertionSide / StackCube / PushCube. The readout rows share
one set of executor weights; only the goal pipeline differs. "Recorded" pools evaluation seeds 999 / 997 / 998
(768 episodes per task).

| Row | Executor | Printed | Recorded |
|---|---|---|---|
| Rigid Readout K = 1 | `mt5_rcfz_gmpc_s0` | 29.80 / 70.18 / 20.05 / 52.47 / 99.74 (mean 54.45) | 29.82 / 70.18 / 20.05 / 52.47 / 99.74 |
| Rigid Readout K = 4 | `mt5_rcfz_gmpc_s0`, PickCube cell `mt5_rcfz_t2k_s0` | 46.48 / 76.95 / 21.88 / 54.82 / 99.22 (mean 59.87) | 33.07 / 76.95 / 21.88 / 54.82 / 99.22 with one executor; the printed PickCube cell is the second, at seed 999 |
| **Entity-Level Goal Readout** | `mt5_rciid_gmpc_s0` | **81.50 / 98.35 / 32.84 / 86.54 / 99.20 (mean 79.69)** | 81.51 / 98.31 / 32.81 / 86.46 / 99.22 |
| Oracle Goal | `mt5_rcfz_gmpc_s0` | 94.66 / 97.27 / 36.72 / 87.11 / 99.74 | identical |

Terminal goal position error (mm, seed 999, PickCube / PegInsertionSide / StackCube): Rigid K = 1 41.1 / 42.6 / 28.9,
Rigid K = 4 37.2 / 40.8 / 25.4, Entity-Level Goal Readout 31.5 / 29.4 / 12.8.

## Before you rely on these

- **The planner files hold trained parameters only.** Every frozen-encoder tensor was bitwise identical to the
  published Hub weights and was removed, along with the optimizer state. The encoders are re-created from the Hub
  when the planner is built, so **DINOv3 is required and gated**: accept its licence and run `hf auth login` once.
  `verify_main_table.sh` builds no planner and needs none of this.
- **Training paths inside the planner `config` are placeholders.** `dataset_dirs`, `cache_dir` and `checkpoint_dir`
  read `[path-to-the-repository-root-here]/...`. They are metadata — inference and evaluation never read them — but
  set them to real directories before resuming training.
- **PegInsertionSide is not bitwise reproducible.** Repeated evaluation of the same checkpoint moves by a couple of
  episodes in 256, which is why the verification script compares within ±0.04.
- **Simulation only, one training seed.** The paper's real-robot results are not established for these checkpoints
  by this release, and the three evaluation seeds vary scenes, not training.
- Nothing here is a third-party weight. CoTracker3 and SAM 2 are not redistributed; the code repository's README
  lists every external model, its licence and which stage needs it.

## License and citation

Apache-2.0. The released weights contain only parameters trained by the authors; the upstream TraceGen Generalist
checkpoint they are fine-tuned from is subject to its own terms.

```bibtex
@article{chuang2026executablegoals,
  title  = {Predicted Futures Are Not Enough: Learning Executable Goals for Robot Manipulation},
  author = {Chuang, Tzu-Yu and Chang, Ching-Hsiang and Lee, Yi-Hsiu and Chen, Yi-Ting and Sun, Min and Yang, YuanFu},
  year   = {2026}
}
```