File size: 9,876 Bytes
43e55d7 96efff3 43e55d7 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 | ---
license: mit
tags:
- robotics
- imitation-learning
- robocasa
- arxiv:2608.22591
---
# WorldToken Checkpoints
[Paper](https://arxiv.org/abs/2608.22591) | [Code](https://github.com/me271828/WorldToken) | [Experiment records](https://huggingface.co/datasets/mepi31415/WorldToken_Experiment_Records) | [Checkpoints](https://huggingface.co/mepi31415/WorldToken_Checkpoints) | [Hugging Face paper page](https://huggingface.co/papers/2608.22591)
Model card and index for the **26 selected RoboCasa checkpoints** accompanying
**WorldToken: Time-First Sequence Modeling for Robotic Imitation Learning**.
WorldToken maps RGB observations, proprioception and task language into tokens,
models their history with a causal Transformer, and predicts action chunks with
a diffusion head. The BC-Transformer entry uses the native robomimic baseline.
| Paper section | Assets | Checkpoints |
|---|---|---:|
| [§4 — Multitask control and scaling](04_robocasa_scaling/README.md) | N1–N3 scaling grid, seed 0, D50/D100/D300/D1000/D2900; BC-Transformer at D300, seed 123 | 15 + 1 |
| [§5 — Token interface](05_token_interface/README.md) | N2, K=4 and K=50, seed 0, all five data settings | 10 |
Each run has its own directory and `run_info.json` with its parameters,
original training name and final step. The section indexes list the expected
weight filenames; BC-Transformer uses `models/model_epoch_1000.pth`.
Each directory also includes a `config.json`, copied from the matching
training record with HDF5 data paths normalized to
`${DATA_ROOT}/robocasa/mg_im/v0.1/...`. The 25 WorldToken configs also carry
the matching public recipe's `eval_window_spec` for offline holdout RMSE;
the BC-Transformer uses its separate native evaluation pipeline.
## Release availability
All 26 selected checkpoints are available: 15 scaling-grid policies, the BC-Transformer baseline, and 10 token-interface policies.
## Training and evaluation records
These policies were trained by imitation learning on the paper's 23-task
RoboCasa MG demonstration subsets and evaluated on 23 tasks × 50 episodes,
with three full-history repeats per checkpoint.
Full details are in the companion [WorldToken Experiment Records](https://huggingface.co/datasets/mepi31415/WorldToken_Experiment_Records) package
(`WorldToken_Experiment_Records`), under the same `<section>/<run>/` path:
- `config.json` and split files: architecture, training settings and data selection.
- Training logs: optimization and holdout metrics.
- `rollouts/`: episode outcomes, summary statistics and videos.
Use the matching run's results when assessing a checkpoint; paper averages over
multiple training seeds are not the score of a single released weight file.
## Usage
Download all currently available weights and their configurations with the
Hugging Face CLI:
```bash
hf download mepi31415/WorldToken_Checkpoints --local-dir ./WorldToken_Checkpoints
```
For only the N2/D300 example below, append
`--include "04_robocasa_scaling/scaling_n2_d300_seed0/*"`.
Set `CHECKPOINT_ROOT` to the absolute path of the downloaded directory.
Follow the code repository's [environment setup](https://github.com/me271828/WorldToken/blob/main/environments/README.md)
and [data preparation](https://github.com/me271828/WorldToken/blob/main/environments/DATA.md).
Evaluation requires RoboCasa assets and demonstration HDF5 metadata as well as
the weight file. From the WorldToken code repository, after setting
`DATA_ROOT` and `CHECKPOINT_ROOT` to your local downloads:
```bash
RUN=scaling_n2_d300_seed0
SECTION=04_robocasa_scaling
EVAL_RUN="$PWD/runs/checkpoint_eval/$RUN"
mkdir -p "$EVAL_RUN"
cp "$CHECKPOINT_ROOT/$SECTION/$RUN/config.json" "$EVAL_RUN/config.json"
python -m experiments.common.rollout \
--config "experiments/$SECTION/configs/$RUN.json" \
--run-dir "$EVAL_RUN" --data-root "$DATA_ROOT" \
--checkpoint "$CHECKPOINT_ROOT/$SECTION/$RUN/checkpoint_step_00030000.pt"
```
This runs one full-history evaluation and saves new outputs under `EVAL_RUN`.
The section READMEs in the code repository describe other variants and the baseline.
### Offline holdout RMSE (WorldToken)
Use the code repository's Linux/CUDA model environment and prepare the frozen
CLIP text encoder as described in its environment instructions. Download the
matching experiment-records package as well as the checkpoint. From the code
repository root, set these paths to your local copies:
```bash
export CODE_ROOT="$PWD"
export DATA_ROOT=/absolute/path/to/data
export CHECKPOINT_ROOT=/absolute/path/to/WorldToken_Checkpoints
export RECORDS_ROOT=/absolute/path/to/WorldToken_Experiment_Records
export RUN=scaling_n2_d300_seed0
export SECTION=04_robocasa_scaling
export EVAL_RUN="$CODE_ROOT/runs/checkpoint_eval/$RUN"
# Optional: an existing CLIP cache from the matching training environment.
# export LANG_EMB_CACHE=/absolute/path/to/robocasa_clip.npz
```
Prepare an evaluation copy with the snippet below. Historical split files may
use `robocasa/v0.1/...`; it maps each HDF5 path to the published config by its
portable `single_stage/...` identity, preserving the demonstration list.
It copies an optional existing language cache, or builds one from the holdout
language strings using the pinned robomimic CLIP wrapper. `ROBOMIMIC_SRC` is
optional when that source is already installed in the environment. The
original records and source cache are not modified.
```bash
python - <<'PY'
import json
import os
import shutil
from pathlib import Path
from worldtoken.data import build_lang_embeddings, portable_episode_key
from worldtoken.eval_holdout_rmse import load_persisted_holdout_refs
from worldtoken.train_utils import expand_path_placeholders
run = Path(os.environ["EVAL_RUN"]).resolve()
relative = Path(os.environ["SECTION"]) / os.environ["RUN"]
records = Path(os.environ["RECORDS_ROOT"]) / relative
config_path = Path(os.environ["CHECKPOINT_ROOT"]) / relative / "config.json"
config = json.loads(config_path.read_text(encoding="utf-8"))
run.mkdir(parents=True, exist_ok=True)
paths = config.get("hdf5_paths") or config["dataset"]
by_file = {portable_episode_key(p): str(expand_path_placeholders(p).resolve()) for p in paths}
if len(by_file) != len(paths):
raise ValueError("Duplicate portable HDF5 identities in config")
split = json.loads((records / "holdout_demos.json").read_text(encoding="utf-8"))
for demo in split["demos"]:
demo["hdf5_path"] = by_file[portable_episode_key(demo["hdf5_path"])]
(run / "holdout_demos.json").write_text(json.dumps(split, indent=2) + "\n", encoding="utf-8")
shutil.copy2(records / "metrics.jsonl.gz", run / "metrics.jsonl.gz")
cache = run / "holdout_clip.npz"
if os.environ.get("LANG_EMB_CACHE"):
source = expand_path_placeholders(os.environ["LANG_EMB_CACHE"]).resolve()
if source != cache:
shutil.copy2(source, cache)
src = os.environ.get("ROBOMIMIC_SRC") # Use the local installation, not an archived source path.
src = expand_path_placeholders(src).resolve() if src else None
config["lang_emb_cache"] = str(cache)
config["robomimic_src"] = str(src) if src else None
spec = expand_path_placeholders(config["eval_window_spec"]).resolve()
if not spec.is_file():
raise FileNotFoundError(spec)
config["eval_window_spec"] = str(spec)
(run / "config.json").write_text(json.dumps(config, indent=2) + "\n", encoding="utf-8")
build_lang_embeddings(
load_persisted_holdout_refs(run), device_arg="cpu", cache_path=cache,
robomimic_src=src, mode=config["lang_emb_mode"], write_cache=True,
)
print(f"Prepared {run}; fixed windows: {spec}")
PY
python -m worldtoken.eval_holdout_rmse \
--run-dir "$EVAL_RUN" \
--checkpoint "$CHECKPOINT_ROOT/$SECTION/$RUN/checkpoint_step_00030000.pt" \
--require-fixed-windows
```
For another WorldToken run, change `SECTION`, `RUN` and the weight filename
together. Always use the window spec from its matching recipe; the two C=10
files have different historical start frames and cannot be substituted based
on history length alone. If evaluating an already prepared historical run
whose config lacks this field, obtain the path from that run's public recipe
and pass `--eval-window-spec PATH` with `--require-fixed-windows`. The latter
prevents silent window reselection.
The result under `EVAL_RUN/holdout_grouped_rmse_v2/` includes the actual window
path, evaluation protocol, and `legacy_v1_parity` against the
final-step row in the copied `metrics.jsonl.gz`. An existing `metrics.jsonl`
takes precedence, so use a dedicated evaluation directory for each run.
Add `--require-legacy-parity` to fail when no reference is available or the
comparison exceeds `--legacy-parity-atol` (default `5e-6`); the result is saved
for inspection before a parity failure. Use `--force` to recompute an existing
result after changing inputs.
Keep `eval_batch_size`, `eval_seed`, `holdout_rmse_samplers`, denoising settings
and `precision` unchanged for a paper comparison. Diffusion samples depend on
batch boundaries, so reducing batch size can change the answer despite fixed
windows. Use a complete matching language cache when available; rebuilding it
or changing hardware/software can introduce numerical differences, which
should be reported with the parity result. The fixed-window check validates
sample selection, not numerical equivalence of the complete evaluator.
The `--debug-max-demos` and `--debug-crops-per-demo` subset options cannot be
combined with fixed windows, which require the full saved split and crop count.
## Scope and license
These are simulation research policies; real-robot performance and transfer to
other tasks or observation/action interfaces have not been established.
The weights and documentation in this package are licensed under the
[MIT License](LICENSE). Third-party materials retain their own licenses.
The separate experiment-records package specifies its own license.
|