mepi31415's picture
Mark all 26 planned checkpoints as available
96efff3 verified
|
Raw History Blame Contribute Delete
9.88 kB
---
license: mit
tags:
- robotics
- imitation-learning
- robocasa
- arxiv:2608.22591
---
# WorldToken Checkpoints
[Paper](https://arxiv.org/abs/2608.22591) | [Code](https://github.com/me271828/WorldToken) | [Experiment records](https://huggingface.co/datasets/mepi31415/WorldToken_Experiment_Records) | [Checkpoints](https://huggingface.co/mepi31415/WorldToken_Checkpoints) | [Hugging Face paper page](https://huggingface.co/papers/2608.22591)
Model card and index for the **26 selected RoboCasa checkpoints** accompanying
**WorldToken: Time-First Sequence Modeling for Robotic Imitation Learning**.
WorldToken maps RGB observations, proprioception and task language into tokens,
models their history with a causal Transformer, and predicts action chunks with
a diffusion head. The BC-Transformer entry uses the native robomimic baseline.
| Paper section | Assets | Checkpoints |
|---|---|---:|
| [§4 — Multitask control and scaling](04_robocasa_scaling/README.md) | N1–N3 scaling grid, seed 0, D50/D100/D300/D1000/D2900; BC-Transformer at D300, seed 123 | 15 + 1 |
| [§5 — Token interface](05_token_interface/README.md) | N2, K=4 and K=50, seed 0, all five data settings | 10 |
Each run has its own directory and `run_info.json` with its parameters,
original training name and final step. The section indexes list the expected
weight filenames; BC-Transformer uses `models/model_epoch_1000.pth`.
Each directory also includes a `config.json`, copied from the matching
training record with HDF5 data paths normalized to
`${DATA_ROOT}/robocasa/mg_im/v0.1/...`. The 25 WorldToken configs also carry
the matching public recipe's `eval_window_spec` for offline holdout RMSE;
the BC-Transformer uses its separate native evaluation pipeline.
## Release availability
All 26 selected checkpoints are available: 15 scaling-grid policies, the BC-Transformer baseline, and 10 token-interface policies.
## Training and evaluation records
These policies were trained by imitation learning on the paper's 23-task
RoboCasa MG demonstration subsets and evaluated on 23 tasks × 50 episodes,
with three full-history repeats per checkpoint.
Full details are in the companion [WorldToken Experiment Records](https://huggingface.co/datasets/mepi31415/WorldToken_Experiment_Records) package
(`WorldToken_Experiment_Records`), under the same `<section>/<run>/` path:
- `config.json` and split files: architecture, training settings and data selection.
- Training logs: optimization and holdout metrics.
- `rollouts/`: episode outcomes, summary statistics and videos.
Use the matching run's results when assessing a checkpoint; paper averages over
multiple training seeds are not the score of a single released weight file.
## Usage
Download all currently available weights and their configurations with the
Hugging Face CLI:
```bash
hf download mepi31415/WorldToken_Checkpoints --local-dir ./WorldToken_Checkpoints
```
For only the N2/D300 example below, append
`--include "04_robocasa_scaling/scaling_n2_d300_seed0/*"`.
Set `CHECKPOINT_ROOT` to the absolute path of the downloaded directory.
Follow the code repository's [environment setup](https://github.com/me271828/WorldToken/blob/main/environments/README.md)
and [data preparation](https://github.com/me271828/WorldToken/blob/main/environments/DATA.md).
Evaluation requires RoboCasa assets and demonstration HDF5 metadata as well as
the weight file. From the WorldToken code repository, after setting
`DATA_ROOT` and `CHECKPOINT_ROOT` to your local downloads:
```bash
RUN=scaling_n2_d300_seed0
SECTION=04_robocasa_scaling
EVAL_RUN="$PWD/runs/checkpoint_eval/$RUN"
mkdir -p "$EVAL_RUN"
cp "$CHECKPOINT_ROOT/$SECTION/$RUN/config.json" "$EVAL_RUN/config.json"
python -m experiments.common.rollout \
--config "experiments/$SECTION/configs/$RUN.json" \
--run-dir "$EVAL_RUN" --data-root "$DATA_ROOT" \
--checkpoint "$CHECKPOINT_ROOT/$SECTION/$RUN/checkpoint_step_00030000.pt"
```
This runs one full-history evaluation and saves new outputs under `EVAL_RUN`.
The section READMEs in the code repository describe other variants and the baseline.
### Offline holdout RMSE (WorldToken)
Use the code repository's Linux/CUDA model environment and prepare the frozen
CLIP text encoder as described in its environment instructions. Download the
matching experiment-records package as well as the checkpoint. From the code
repository root, set these paths to your local copies:
```bash
export CODE_ROOT="$PWD"
export DATA_ROOT=/absolute/path/to/data
export CHECKPOINT_ROOT=/absolute/path/to/WorldToken_Checkpoints
export RECORDS_ROOT=/absolute/path/to/WorldToken_Experiment_Records
export RUN=scaling_n2_d300_seed0
export SECTION=04_robocasa_scaling
export EVAL_RUN="$CODE_ROOT/runs/checkpoint_eval/$RUN"
# Optional: an existing CLIP cache from the matching training environment.
# export LANG_EMB_CACHE=/absolute/path/to/robocasa_clip.npz
```
Prepare an evaluation copy with the snippet below. Historical split files may
use `robocasa/v0.1/...`; it maps each HDF5 path to the published config by its
portable `single_stage/...` identity, preserving the demonstration list.
It copies an optional existing language cache, or builds one from the holdout
language strings using the pinned robomimic CLIP wrapper. `ROBOMIMIC_SRC` is
optional when that source is already installed in the environment. The
original records and source cache are not modified.
```bash
python - <<'PY'
import json
import os
import shutil
from pathlib import Path
from worldtoken.data import build_lang_embeddings, portable_episode_key
from worldtoken.eval_holdout_rmse import load_persisted_holdout_refs
from worldtoken.train_utils import expand_path_placeholders
run = Path(os.environ["EVAL_RUN"]).resolve()
relative = Path(os.environ["SECTION"]) / os.environ["RUN"]
records = Path(os.environ["RECORDS_ROOT"]) / relative
config_path = Path(os.environ["CHECKPOINT_ROOT"]) / relative / "config.json"
config = json.loads(config_path.read_text(encoding="utf-8"))
run.mkdir(parents=True, exist_ok=True)
paths = config.get("hdf5_paths") or config["dataset"]
by_file = {portable_episode_key(p): str(expand_path_placeholders(p).resolve()) for p in paths}
if len(by_file) != len(paths):
raise ValueError("Duplicate portable HDF5 identities in config")
split = json.loads((records / "holdout_demos.json").read_text(encoding="utf-8"))
for demo in split["demos"]:
demo["hdf5_path"] = by_file[portable_episode_key(demo["hdf5_path"])]
(run / "holdout_demos.json").write_text(json.dumps(split, indent=2) + "\n", encoding="utf-8")
shutil.copy2(records / "metrics.jsonl.gz", run / "metrics.jsonl.gz")
cache = run / "holdout_clip.npz"
if os.environ.get("LANG_EMB_CACHE"):
source = expand_path_placeholders(os.environ["LANG_EMB_CACHE"]).resolve()
if source != cache:
shutil.copy2(source, cache)
src = os.environ.get("ROBOMIMIC_SRC") # Use the local installation, not an archived source path.
src = expand_path_placeholders(src).resolve() if src else None
config["lang_emb_cache"] = str(cache)
config["robomimic_src"] = str(src) if src else None
spec = expand_path_placeholders(config["eval_window_spec"]).resolve()
if not spec.is_file():
raise FileNotFoundError(spec)
config["eval_window_spec"] = str(spec)
(run / "config.json").write_text(json.dumps(config, indent=2) + "\n", encoding="utf-8")
build_lang_embeddings(
load_persisted_holdout_refs(run), device_arg="cpu", cache_path=cache,
robomimic_src=src, mode=config["lang_emb_mode"], write_cache=True,
)
print(f"Prepared {run}; fixed windows: {spec}")
PY
python -m worldtoken.eval_holdout_rmse \
--run-dir "$EVAL_RUN" \
--checkpoint "$CHECKPOINT_ROOT/$SECTION/$RUN/checkpoint_step_00030000.pt" \
--require-fixed-windows
```
For another WorldToken run, change `SECTION`, `RUN` and the weight filename
together. Always use the window spec from its matching recipe; the two C=10
files have different historical start frames and cannot be substituted based
on history length alone. If evaluating an already prepared historical run
whose config lacks this field, obtain the path from that run's public recipe
and pass `--eval-window-spec PATH` with `--require-fixed-windows`. The latter
prevents silent window reselection.
The result under `EVAL_RUN/holdout_grouped_rmse_v2/` includes the actual window
path, evaluation protocol, and `legacy_v1_parity` against the
final-step row in the copied `metrics.jsonl.gz`. An existing `metrics.jsonl`
takes precedence, so use a dedicated evaluation directory for each run.
Add `--require-legacy-parity` to fail when no reference is available or the
comparison exceeds `--legacy-parity-atol` (default `5e-6`); the result is saved
for inspection before a parity failure. Use `--force` to recompute an existing
result after changing inputs.
Keep `eval_batch_size`, `eval_seed`, `holdout_rmse_samplers`, denoising settings
and `precision` unchanged for a paper comparison. Diffusion samples depend on
batch boundaries, so reducing batch size can change the answer despite fixed
windows. Use a complete matching language cache when available; rebuilding it
or changing hardware/software can introduce numerical differences, which
should be reported with the parity result. The fixed-window check validates
sample selection, not numerical equivalence of the complete evaluator.
The `--debug-max-demos` and `--debug-crops-per-demo` subset options cannot be
combined with fixed windows, which require the full saved split and crop count.
## Scope and license
These are simulation research policies; real-robot performance and transfer to
other tasks or observation/action interfaces have not been established.
The weights and documentation in this package are licensed under the
[MIT License](LICENSE). Third-party materials retain their own licenses.
The separate experiment-records package specifies its own license.