Condense the model card and link code, project page and paper
Browse files
README.md
CHANGED
|
@@ -1,137 +1,87 @@
|
|
| 1 |
---
|
| 2 |
license: apache-2.0
|
| 3 |
-
tags: [robotics, manipulation, maniskill, trajectory-prediction, imitation-learning]
|
| 4 |
library_name: pytorch
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 5 |
---
|
| 6 |
|
| 7 |
-
#
|
| 8 |
|
| 9 |
-
|
| 10 |
|
| 11 |
-
|
|
|
|
|
|
|
| 12 |
|
| 13 |
-
|
| 14 |
|
| 15 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 16 |
|
| 17 |
-
|
| 18 |
-
- `planner/mix5_t2k_n3000.pth` — the entity branch of the Entity-Level Goal Readout (`MSGEN_T2K=1`). It supplies the full terminal pose for PegInsertionSide and the terminal position for LiftPegUpright and PushCube, and serves as the Rigid Readout source for PushCube.
|
| 19 |
-
- `planner/mix5_t2k_gmap.pth` — entity branch plus map branch / Spatial Goal Map (`MSGEN_T2K=1 MSGEN_T2K_GMAP=1`). It supplies the terminal position of PickCube (map peak plus ray, bank suffix `gmappeakNC`; with the SAM 2 marker localiser prompted at the map peak, `sam2mk6`) and StackCube (weighted depth-centroid readout, `gmapdcc`).
|
| 20 |
|
| 21 |
-
|
| 22 |
|
| 23 |
-
|
| 24 |
|
| 25 |
-
|
| 26 |
-
|
| 27 |
-
|
| 28 |
-
|
| 29 |
-
|
| 30 |
-
|
| 31 |
-
|
| 32 |
-
|
| 33 |
-
## Files
|
| 34 |
-
|
| 35 |
-
| Path | Role | Training data and recipe | Warm start | SHA-256 (first 16 hex) |
|
| 36 |
-
|---|---|---|---|---|
|
| 37 |
-
| `planner/mix4_realcam_n2400.pth` (302,007,919 B) | Rigid Readout source and rotation source: K = 4 trace mean + RANSAC Kabsch for PickCube, LiftPegUpright, PegInsertionSide, StackCube; K = 1 / K = 4 goal banks; psi banks; warm start of the two readout planners | `data/ds/realcam_n2400`: four tasks (PickCube, StackCube, PegInsertionSide, LiftPegUpright), 200 official demos x 3 camera views each, wall backdrop, dense 3D traces with 16 object-aware queries; `msgen.run_train`, 15 epochs, batch 8, `lr_decoder` 1.5e-4; no `MSGEN_*` variable recorded | TraceGen Generalist `tracegen_model.pth` (SHA-256 `f595ad249cbd59ce`) | `81f25468df5de89f` |
|
| 38 |
-
| `planner/mix5_t2k_n3000.pth` (314,213,705 B) | Entity branch: PegInsertionSide full pose; LiftPegUpright and PushCube position; PushCube Rigid Readout source | `data/ds/realcam_t2k_n3000`: five tasks (the four above plus PushCube), same replays, plus entity-level labels (`msgen.labels_t2k`); 8 epochs, batch 8, `lr_decoder` 1.5e-4, `MSGEN_T2K=1 MSGEN_T2K_W=0.3`; code git `ba7de4d2` | `mix4_realcam_n2400` final | `6b5b6544a0b7b502` |
|
| 39 |
-
| `planner/mix5_t2k_gmap.pth` (314,217,163 B) | Entity branch + map branch (Spatial Goal Map): PickCube (`gmappeakNC`, `sam2mk6`) and StackCube (`gmapdcc`) goal positions | `data/ds/realcam_t2k_n3000` with goal-map labels **v1** (6 cm hard ball); 8 epochs, batch 8, `lr_decoder` 1.5e-4, `MSGEN_T2K=1 MSGEN_T2K_GMAP=1 MSGEN_T2K_W=0.3 MSGEN_KEEP_ALL_CKPT=1`; code git `ba7de4d2` | `mix4_realcam_n2400` final (not `mix5_t2k_n3000`) | `46643abdd9744d3c` |
|
| 40 |
-
| `student/mt5_rciid_gmpc_s0/{student.pt, run.json}` (student 3,239,060 B) | Pose-Native Executor of the Entity-Level Goal Readout row of Table II; five tasks, 804,002 parameters (`--ckpt final`) | `msppo.multi_distill` (DAgger) from the five teachers below; 40,000 iterations, 320 envs (64 per task), lr 3e-4, seed 0, no scene channel, psi token, goal = K = 4 mean, goal perturbation redrawn i.i.d. at every control step (`MSPPO_IID_INJECT=1`, uniform scale) from the relbanks StackCube `gmapdcc`, PickCube `gmappeak`, others `t2kdcc`; trained 2026-09-13 | not recorded (no warm-start field in `run.json`) | `5ed8f4361eb6a676` |
|
| 41 |
-
| `student/mt5_rcfz_gmpc_s0/{student.pt, run.json}` (student 3,239,060 B) | Pose-Native Executor of the Rigid Readout rows, the Oracle Goal references and Fig. 7 | same as above with one goal-perturbation draw per episode, frozen (`MSPPO_FROZEN_INJECT=1`, uniform scale); trained 2026-09-07 | not recorded | `50467d34909e0d82` |
|
| 42 |
-
| `student/mt5_rcfz_t2k_s0/{student.pt, run.json}` (student 3,239,060 B) | Earlier Pose-Native Executor; the PickCube cell of the Rigid Readout K = 4 row (seed 999) | same recipe with episode-fixed injection and `t2kdcc` relbanks for all five tasks (its PickCube and StackCube relbanks are not shipped); trained 2026-09-01 | not recorded | `99aaa1f6eb50ac93` |
|
| 43 |
-
| `teacher/pc_v9_nz_s0/{agent.pt, run.json, patches.json}` (agent 4,052,398 B) | PickCube-v1 privileged PPO teacher, 1,011,473 parameters | `msppo.kp_teacher`; 12,000,000 steps, 1024 envs, lr 5e-5, `target_kl` 0.01, noise-v2 injector (`noise_v2` 0.4, `obj_v2` 1.0); patch `msppo/patch_frame.py` with `MSPPO_FRAME_WPUSH=0.25` | `runs_rl/pc_v9_frame4_s0_padsig` (not released) | `702259e46a7ce08b` |
|
| 44 |
-
| `teacher/lp_v9_nz_s0/{agent.pt, run.json, patches.json}` (agent 4,011,438 B) | LiftPegUpright-v1 privileged PPO teacher, 1,001,233 parameters | 12,000,000 steps, 256 envs, lr 5e-5, `target_kl` 0.01, `noise_v2` 1.0, `obj_v2` 1.0; patch `msppo/patch_frame.py` | `runs_rl/lp_v9_frame_s0_padsig` (not released) | `b3449a618fdf5e21` |
|
| 45 |
-
| `teacher/pi_v9_frame4_s0/{agent.pt, run.json, patches.json}` (agent 4,056,494 B) | PegInsertionSide-v1 privileged PPO teacher (psi-conditioned), 1,012,497 parameters | 8,000,000 steps, 1024 envs, lr 2e-5, `target_kl` 0.005, `max_episode_steps` 100, `noise_v2` 1.0, `obj_v2` 1.0; patch `msppo/patch_frame.py` with `MSPPO_FRAME_WH=0 MSPPO_FRAME_WPUSH=0 MSPPO_FRAME_K1_BASE=-0.28,-0.56,-0.78` | `runs_rl/pi_v5_hi_s0` (not released) | `ef0559ebcfdc9c42` |
|
| 46 |
-
| `teacher/sc_v9_nz03b_s0/{agent.pt, run.json, patches.json}` (agent 3,995,054 B) | StackCube-v1 privileged PPO teacher, `panda_wristcam` robot, 997,137 parameters | 12,000,000 steps, 1024 envs, lr 1e-4, `target_kl` 0.02, `noise_v2` 0.3, `obj_v2` 0.15, bump penalty `w_bump` 0.3, `w_kp` 0.125; patch `msppo/patch_stack_frame.py` | `runs_rl/sc_v9_frame_s0_padsig` (not released) | `fd8ef9911d8e153d` |
|
| 47 |
-
| `teacher/push_v9_nz_s0/{agent.pt, run.json}` (agent 4,023,726 B) | PushCube-v1 privileged PPO teacher, 1,004,305 parameters (this run carries no `patches.json`) | 12,000,000 steps, 1024 envs, lr 5e-5, `target_kl` 0.01, `noise_v2` 1.0, `obj_v2` 1.0 | `runs_rl/push_kp_s0_psipad_padsig` (not released) | `34c0175e65cbe84b` |
|
| 48 |
-
| `SHA256SUMS` | SHA-256 of every weight file above (`sha256sum -c SHA256SUMS`) | — | — | — |
|
| 49 |
-
| `README.md` | This model card | — | — | — |
|
| 50 |
-
|
| 51 |
-
Every teacher run and every distillation were launched with seed 0; the planner `run.json` files record no seed, and no executor `run.json` records a git revision. The teacher `patches.json` files record the code patch, the environment variables and the code git revision (`ba7de4d2af700d4d4f02a8d2c2f8f8d350751108`, a commit of the private research repository, not resolvable from this release) that were active when the teacher was trained; the same revision is recorded in the provenance block of the two readout planners. `mix4_realcam_n2400/run.json` carries no provenance or git field.
|
| 52 |
-
|
| 53 |
-
## Intended use and how to load
|
| 54 |
-
|
| 55 |
-
These weights reproduce the closed-loop tables of the paper inside ManiSkill3 under the production visual protocol (front camera `eye=(0.574,-0.051,0.378)`, `target=(-0.4751,0.0562,0.0200)`, `fov=0.754`, wall backdrop). They are research artefacts for that setting; no other use has been evaluated.
|
| 56 |
-
|
| 57 |
-
1. Clone the code repository `Claire0730/executable-goals` and set up the two conda environments described in its `requirements/README.md` (`trace_gen`, Python 3.10, for the planner; `maniskill`, Python 3.11, for the simulator, teachers, distillation and evaluation; both frozen at `torch 2.11.0+cu128`). Clone the upstream TraceGen repository at the pinned commit and apply `third_party/tracegen_local.patch` as described in `third_party/README.md`.
|
| 58 |
-
2. Download this repository and expose the weights under the layout the code expects:
|
| 59 |
-
|
| 60 |
-
```bash
|
| 61 |
-
hf download Claire0730/executable-goals --local-dir checkpoints
|
| 62 |
-
bash scripts/00_link_checkpoints.sh
|
| 63 |
-
```
|
| 64 |
-
|
| 65 |
-
`00_link_checkpoints.sh` symlinks `checkpoints/student/*` (the three executors) and `checkpoints/teacher/*` to `runs_rl/<tag>`, checks that the three planner files named by `CK_MIX4`, `CK_HEAD` and `CK_GMAP` in `scripts/config.sh` exist, and runs `sha256sum -c SHA256SUMS`.
|
| 66 |
-
3. **Planners** are loaded through `msgen.predict` by path (`--ckpt checkpoints/planner/<tag>.pth`). The readout that is active is selected by environment variables, not by the file: `mix5_t2k_n3000.pth` must be run with `MSGEN_T2K=1`, `mix5_t2k_gmap.pth` with `MSGEN_T2K=1 MSGEN_T2K_GMAP=1`, and `mix4_realcam_n2400.pth` without either flag. The sampler differs per prediction family, as in production: the four-task `mix4` K = 4 Rigid Readout source uses the native TraceGen sampler (no `MSGEN_STEPS`), whereas the entity-branch, map-branch and PushCube K = 4 predictions use `MSGEN_STEPS=20 MSGEN_DT=fix`. The wall and camera variables are not passed to the planner process; they affect rendering only. The K = 4 Rigid Readout uses the flow-sampler seeds 1234-1237; the readouts use seed 1234 (for PushCube the seed-1234 K-sample file is also the entity-branch file). Because the encoders are re-created from the Hub, the planner process needs the authenticated Hugging Face session described above. `scripts/20_predict.sh` encodes this procedure.
|
| 67 |
-
4. **Executors and teachers** are not self-describing: the network is rebuilt from the `run.json` next to each weight file (executor: `d_model`, `layers`, `heads`, observation layout `student_recon`, `num_kp` 64, `qdim` 9, `tcp_obs`, `pose_obs`, `psi`/`psi_token`, `no_scene`; teacher: `token_dim` 128, `hidden` 1024,1024,512,512, `pn_hidden` 128,128, `obs_dim`, `qdim` 9, `num_kp` 64). `msppo.multi_eval --run runs_rl/<executor tag> --ckpt final` loads an executor; `msppo.multi_distill --teachers <tag,...>` loads `runs_rl/<tag>/{agent.pt, run.json}` for the teachers. The teacher `patches.json` names the code patch (`msppo/patch_frame.py` or `msppo/patch_stack_frame.py`) and the environment variables that must be active before the environment is built.
|
| 68 |
-
5. To reproduce a row of Table II from the frozen goal banks shipped in `Claire0730/executable-goals/banks/` without running a planner, use `ROW=final|k1|k4|k4pick|oracle bash scripts/verify_main_table.sh [seed]` (it selects the executor of that row); to regenerate the banks from the planners, run `scripts/10_render_banks.sh`, `20_predict.sh`, `30_goals.sh`, optionally `31_goals_sam2marker.sh` (SAM 2), and `40_eval_row4.sh`.
|
| 69 |
-
|
| 70 |
-
## Evaluation results
|
| 71 |
-
|
| 72 |
-
Protocol: `msppo.multi_eval`, N = 256 episodes per task and evaluation seed, training seed 0, `--ckpt final`, metric `success_once`, production camera and wall, psi banks derived from `mix4_realcam_n2400` predictions, the simulator's per-step object correspondence for the executor's pose feedback. The paper pools the evaluation seeds 999, 997 and 998 (768 episodes per task). Values are success rates in percent in the task order PickCube / LiftPegUpright / PegInsertionSide / StackCube / PushCube; "printed" is the value in the paper, "recorded" the mean over the three seeds of the `per_task` field of the records in `Claire0730/executable-goals/paper_results/table2/`. Where printed and recorded values differ by at most 0.08 percentage points, the difference is rounding of the printed values.
|
| 73 |
-
|
| 74 |
-
| Table II row | Executor | Goal banks | Printed | Recorded (pooled) | Records |
|
| 75 |
-
|---|---|---|---|---|---|
|
| 76 |
-
| Rigid Readout K = 1 | `mt5_rcfz_gmpc_s0` | `k1ransac` (PushCube from `mix5_t2k_n3000`) | 29.80 / 70.18 / 20.05 / 52.47 / 99.74 (mean 54.45) | 29.82 / 70.18 / 20.05 / 52.47 / 99.74 (mean 54.45) | `rigid_k1_gmpc_<seed>.json` |
|
| 77 |
-
| Rigid Readout K = 4 | `mt5_rcfz_gmpc_s0`; PickCube cell `mt5_rcfz_t2k_s0` (seed 999) | `kmean_ransac`; PickCube cell `pickcube_goals_mix5_t2k_n3000_999_k4ransac.npz` | 46.48 / 76.95 / 21.88 / 54.82 / 99.22 (mean 59.87) | 33.07 / 76.95 / 21.88 / 54.82 / 99.22 with `mt5_rcfz_gmpc_s0`; the printed PickCube cell 46.48 (119 of 256) is the single-seed record of `mt5_rcfz_t2k_s0` | `rigid_k4_gmpc_<seed>.json`, `rigid_k4_pickcube_t2k_s0_999.json` |
|
| 78 |
-
| Entity-Level Goal Readout | `mt5_rciid_gmpc_s0` | PickCube `sam2mk6`, StackCube `gmapdcc`, others `t2kpos` | 81.50 / 98.35 / 32.84 / 86.54 / 99.20 (mean 79.69) | 81.51 / 98.31 / 32.81 / 86.46 / 99.22 (mean 79.66) | `final_pickcube_sam2_rciid_<seed>.json`, `final_rciid_<seed>.json` |
|
| 79 |
-
| Oracle Goal (Fig. 5a) | `mt5_rcfz_gmpc_s0` | simulator goal | 94.66 / 97.27 / 36.72 / 87.11 / 99.74 | 94.66 / 97.27 / 36.72 / 87.11 / 99.74 | `oracle_gmpc_<seed>.json` |
|
| 80 |
-
|
| 81 |
-
Per seed (999 / 997 / 998): Rigid K = 1 30.86 / 68.36 / 17.58 / 52.34 / 100.00, 32.42 / 71.88 / 18.75 / 51.56 / 99.22, 26.17 / 70.31 / 23.83 / 53.52 / 100.00; Rigid K = 4 (`mt5_rcfz_gmpc_s0`) 30.47 / 78.52 / 18.36 / 53.91 / 100.00, 39.06 / 76.95 / 23.83 / 53.91 / 99.22, 29.69 / 75.39 / 23.44 / 56.64 / 98.44; Entity-Level Goal Readout 84.38 / 98.05 / 36.72 / 88.28 / 99.61, 80.86 / 98.44 / 29.30 / 86.33 / 99.22, 79.30 / 98.44 / 32.42 / 84.77 / 98.83; Oracle Goal 93.75 / 96.88 / 37.11 / 88.28 / 100.00, 94.53 / 98.44 / 32.81 / 86.72 / 99.61, 95.70 / 96.48 / 40.23 / 86.33 / 99.61. The Fig. 7 curves (Oracle Goal translated by 0 to 120 mm, `mt5_rcfz_gmpc_s0`, seeds 999 and 997) are in `paper_results/fig7_perturbation/`. Not in the paper: the final routing with PickCube `gmappeakNC` evaluated with `mt5_rcfz_gmpc_s0` gives 70.31 / 96.35 / 30.34 / 82.94 / 100.00 pooled (`paper_results/supplementary/final_routing_gmpc_<seed>.json`).
|
| 82 |
-
|
| 83 |
-
Table I (terminal goal position error, mm, mean +- sd, seed 999, 256 scenes; PickCube / PegInsertionSide / StackCube) and Fig. 6 are goal-bank statistics of the planners, recorded in `Claire0730/executable-goals/evidence/04_goal_summary.json` (`mean_all_finite`, `sd_all_finite`) and `evidence/goal_table.json` (median / 90th percentile): Rigid Readout K = 1 41.1 +- 21.7 / 42.6 +- 31.6 / 28.9 +- 38.7 (`rigid_K1_ransac|<task>|999`); Rigid Readout K = 4 37.2 +- 19.3 / 40.8 +- 33.0 / 25.4 +- 38.8 (`rigid_K4_ransac|<task>|999`); Entity-Level Goal Readout 31.5 +- 20.1 / 29.4 +- 16.4 / 12.8 +- 34.0 (`final_pick_gmappeakNC|pickcube|999`, `t2k_head_full|peginsert|999`, `final_stack_gmapdcc|stack|999`). The PickCube error of Table I is measured on the `gmappeakNC` bank, whereas the PickCube success of Table II is measured on the `sam2mk6` bank.
|
| 84 |
-
|
| 85 |
-
Planner cost, measured on the planner models built from these checkpoints (batch 1, 20 ODE steps, `MSGEN_STEPS=20 MSGEN_DT=fix`, RTX 5090). The `mix4` row is a like-for-like measurement at 20 steps; the production `mix4` predictions run the native 100-step sampler, whose cost is not recorded:
|
| 86 |
-
|
| 87 |
-
| Checkpoint | Parameters (total / trainable) | s per scene | TFLOPs per scene | Peak VRAM |
|
| 88 |
-
|---|---|---|---|---|
|
| 89 |
-
| `mix4_realcam_n2400.pth` | 674,546,334 / 75,486,238 | 0.253 | 4.51 | 2.62 GiB |
|
| 90 |
-
| `mix5_t2k_gmap.pth` | 677,596,675 / 78,536,579 | 0.277 | 5.06 | 2.63 GiB |
|
| 91 |
-
| `mix5_t2k_n3000.pth` | 677,595,906 / 78,535,810 | 0.278 | 5.06 | 2.63 GiB |
|
| 92 |
-
|
| 93 |
-
A timing of the full K = 4 planning call (four forward passes plus the readout pass and the Rigid Readout) was not recorded; the 1.27 s and 5.60 GB of the paper's Table III do not correspond to these measurements and are not backed by a record in the release.
|
| 94 |
-
|
| 95 |
-
## Training summary
|
| 96 |
-
|
| 97 |
-
**Planner data.** The 200 official ManiSkill demonstrations of each task are replayed under the production camera in three views (nominal, and two jittered views with eye offsets +-(0.02, 0.02, 0.01) m and target offsets +-(0.04, 0.03, 0) m, about +-2-3 degrees of re-aim) with the wall backdrop, and labelled with dense 3D traces on 16 object-aware queries (`msgen.labels --n-obj 16 --stride 2 --min-future 8 --time-mode arclen`). `realcam_n2400` indexes the four tasks without PushCube (4 x 200 x 3 = 2,400 clips); `realcam_t2k_n3000` indexes all five tasks (3,000 clips) and additionally carries the entity-level labels written by `msgen.labels_t2k` (keyframes, per-segment twists, contact descriptor, structure-relative terminal pose, entity membership of the 400 queries, goal-map target).
|
| 98 |
|
| 99 |
-
|
|
|
|
| 100 |
|
| 101 |
-
|
| 102 |
|
| 103 |
-
|
| 104 |
-
|
| 105 |
-
|
| 106 |
-
| `lp_v9_nz_s0` | LiftPegUpright-v1 | panda | 12,000,000 | 256 | 5e-5 | 0.01 | 50 | `lp_v9_frame_s0_padsig` |
|
| 107 |
-
| `pi_v9_frame4_s0` | PegInsertionSide-v1 | panda | 8,000,000 | 1024 | 2e-5 | 0.005 | 100 | `pi_v5_hi_s0` |
|
| 108 |
-
| `sc_v9_nz03b_s0` | StackCube-v1 | panda_wristcam | 12,000,000 | 1024 | 1e-4 | 0.02 | 50 | `sc_v9_frame_s0_padsig` |
|
| 109 |
-
| `push_v9_nz_s0` | PushCube-v1 | panda | 12,000,000 | 1024 | 5e-5 | 0.01 | 50 | `push_kp_s0_psipad_padsig` |
|
| 110 |
|
| 111 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 112 |
|
| 113 |
-
|
|
|
|
| 114 |
|
| 115 |
-
##
|
| 116 |
|
| 117 |
-
- **
|
| 118 |
-
|
| 119 |
-
|
| 120 |
-
|
| 121 |
-
- **
|
| 122 |
-
|
| 123 |
-
|
| 124 |
-
- **
|
| 125 |
-
|
| 126 |
-
- **
|
| 127 |
-
|
| 128 |
-
-
|
| 129 |
-
|
| 130 |
-
- **Records.** `mix4_realcam_n2400/run.json` has no provenance block (no argv or git field); no executor `run.json` has a git field; `teacher/push_v9_nz_s0/` has no `patches.json`.
|
| 131 |
|
| 132 |
-
##
|
| 133 |
|
| 134 |
-
|
|
|
|
| 135 |
|
| 136 |
```bibtex
|
| 137 |
@article{chuang2026executablegoals,
|
|
@@ -140,5 +90,3 @@ If you use these weights, please cite the paper:
|
|
| 140 |
year = {2026}
|
| 141 |
}
|
| 142 |
```
|
| 143 |
-
|
| 144 |
-
Please also cite the upstream TraceGen work whose Generalist checkpoint is the warm start of every planner released here (see the upstream repository for its reference).
|
|
|
|
| 1 |
---
|
| 2 |
license: apache-2.0
|
|
|
|
| 3 |
library_name: pytorch
|
| 4 |
+
pipeline_tag: robotics
|
| 5 |
+
base_model: JayLee131/TraceGen
|
| 6 |
+
base_model_relation: finetune
|
| 7 |
+
tags:
|
| 8 |
+
- robotics
|
| 9 |
+
- manipulation
|
| 10 |
+
- maniskill
|
| 11 |
+
- trajectory-prediction
|
| 12 |
+
- imitation-learning
|
| 13 |
---
|
| 14 |
|
| 15 |
+
# Predicted Futures Are Not Enough: Learning Executable Goals for Robot Manipulation
|
| 16 |
|
| 17 |
+
[Code](https://github.com/Claire0730/executable-goals) · [Project page](https://claire0730.github.io/executable-goals/) · Paper (arXiv, soon)
|
| 18 |
|
| 19 |
+
Weights and frozen goal banks behind the paper. A 3D trace world model predicts one future per episode; the
|
| 20 |
+
**Entity-Level Goal Readout** turns that prediction into a single executable goal in SE(3), and a shared
|
| 21 |
+
**Pose-Native Executor** runs it closed loop at 20 Hz across five ManiSkill3 tasks.
|
| 22 |
|
| 23 |
+
## What is here
|
| 24 |
|
| 25 |
+
| Folder | Contents | Size |
|
| 26 |
+
|---|---|---|
|
| 27 |
+
| `planner/` | Three 3D Trace Planners fine-tuned from TraceGen: `mix4_realcam_n2400` (Rigid Readout source), `mix5_t2k_n3000` (entity branch), `mix5_t2k_gmap` (entity + map branch) | 0.30–0.31 GB each |
|
| 28 |
+
| `student/` | The three Pose-Native Executors, 804,002 parameters each: `mt5_rciid_gmpc_s0`, `mt5_rcfz_gmpc_s0`, `mt5_rcfz_t2k_s0` | 3.2 MB each |
|
| 29 |
+
| `teacher/` | Five privileged PPO teachers, one per task | ~4 MB each |
|
| 30 |
+
| `banks/` | 49 frozen goal banks — one SE(3) goal per scene, for every reported row, task and evaluation seed | 58 MB |
|
| 31 |
|
| 32 |
+
Digests: `SHA256SUMS` (weights) and `banks/SHA256SUMS`.
|
|
|
|
|
|
|
| 33 |
|
| 34 |
+
## Use
|
| 35 |
|
| 36 |
+
Reproducing the main table needs the executors and the banks only — no planner, no gated licence, one GPU:
|
| 37 |
|
| 38 |
+
```bash
|
| 39 |
+
git clone https://github.com/Claire0730/executable-goals && cd executable-goals
|
| 40 |
+
hf download Claire0730/executable-goals --local-dir checkpoints --include "student/*" --include "teacher/*" --include "SHA256SUMS"
|
| 41 |
+
hf download Claire0730/executable-goals --local-dir . --include "banks/*"
|
| 42 |
+
bash scripts/00_link_checkpoints.sh
|
| 43 |
+
ROW=final bash scripts/verify_main_table.sh 999
|
| 44 |
+
```
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 45 |
|
| 46 |
+
Everything else — the full install, the inference and training chains, which record backs which number — is in the
|
| 47 |
+
code repository's README and `docs/REPRODUCTION.md`.
|
| 48 |
|
| 49 |
+
## Results
|
| 50 |
|
| 51 |
+
Closed-loop success, %, PickCube / LiftPegUpright / PegInsertionSide / StackCube / PushCube. The readout rows share
|
| 52 |
+
one set of executor weights; only the goal pipeline differs. "Recorded" pools evaluation seeds 999 / 997 / 998
|
| 53 |
+
(768 episodes per task).
|
|
|
|
|
|
|
|
|
|
|
|
|
| 54 |
|
| 55 |
+
| Row | Executor | Printed | Recorded |
|
| 56 |
+
|---|---|---|---|
|
| 57 |
+
| Rigid Readout K = 1 | `mt5_rcfz_gmpc_s0` | 29.80 / 70.18 / 20.05 / 52.47 / 99.74 (mean 54.45) | 29.82 / 70.18 / 20.05 / 52.47 / 99.74 |
|
| 58 |
+
| Rigid Readout K = 4 | `mt5_rcfz_gmpc_s0`, PickCube cell `mt5_rcfz_t2k_s0` | 46.48 / 76.95 / 21.88 / 54.82 / 99.22 (mean 59.87) | 33.07 / 76.95 / 21.88 / 54.82 / 99.22 with one executor; the printed PickCube cell is the second, at seed 999 |
|
| 59 |
+
| **Entity-Level Goal Readout** | `mt5_rciid_gmpc_s0` | **81.50 / 98.35 / 32.84 / 86.54 / 99.20 (mean 79.69)** | 81.51 / 98.31 / 32.81 / 86.46 / 99.22 |
|
| 60 |
+
| Oracle Goal | `mt5_rcfz_gmpc_s0` | 94.66 / 97.27 / 36.72 / 87.11 / 99.74 | identical |
|
| 61 |
|
| 62 |
+
Terminal goal position error (mm, seed 999, PickCube / PegInsertionSide / StackCube): Rigid K = 1 41.1 / 42.6 / 28.9,
|
| 63 |
+
Rigid K = 4 37.2 / 40.8 / 25.4, Entity-Level Goal Readout 31.5 / 29.4 / 12.8.
|
| 64 |
|
| 65 |
+
## Before you rely on these
|
| 66 |
|
| 67 |
+
- **The planner files hold trained parameters only.** Every frozen-encoder tensor was bitwise identical to the
|
| 68 |
+
published Hub weights and was removed, along with the optimizer state. The encoders are re-created from the Hub
|
| 69 |
+
when the planner is built, so **DINOv3 is required and gated**: accept its licence and run `hf auth login` once.
|
| 70 |
+
`verify_main_table.sh` builds no planner and needs none of this.
|
| 71 |
+
- **Training paths inside the planner `config` are placeholders.** `dataset_dirs`, `cache_dir` and `checkpoint_dir`
|
| 72 |
+
read `[path-to-the-repository-root-here]/...`. They are metadata — inference and evaluation never read them — but
|
| 73 |
+
set them to real directories before resuming training.
|
| 74 |
+
- **PegInsertionSide is not bitwise reproducible.** Repeated evaluation of the same checkpoint moves by a couple of
|
| 75 |
+
episodes in 256, which is why the verification script compares within ±0.04.
|
| 76 |
+
- **Simulation only, one training seed.** The paper's real-robot results are not established for these checkpoints
|
| 77 |
+
by this release, and the three evaluation seeds vary scenes, not training.
|
| 78 |
+
- Nothing here is a third-party weight. CoTracker3 and SAM 2 are not redistributed; the code repository's README
|
| 79 |
+
lists every external model, its licence and which stage needs it.
|
|
|
|
| 80 |
|
| 81 |
+
## License and citation
|
| 82 |
|
| 83 |
+
Apache-2.0. The released weights contain only parameters trained by the authors; the upstream TraceGen Generalist
|
| 84 |
+
checkpoint they are fine-tuned from is subject to its own terms.
|
| 85 |
|
| 86 |
```bibtex
|
| 87 |
@article{chuang2026executablegoals,
|
|
|
|
| 90 |
year = {2026}
|
| 91 |
}
|
| 92 |
```
|
|
|
|
|
|