--- license: mit tags: - robotics - imitation-learning - robocasa - arxiv:2608.22591 --- # WorldToken Checkpoints [Paper](https://arxiv.org/abs/2608.22591) | [Code](https://github.com/me271828/WorldToken) | [Experiment records](https://huggingface.co/datasets/mepi31415/WorldToken_Experiment_Records) | [Checkpoints](https://huggingface.co/mepi31415/WorldToken_Checkpoints) | [Hugging Face paper page](https://huggingface.co/papers/2608.22591) Model card and index for the **26 selected RoboCasa checkpoints** accompanying **WorldToken: Time-First Sequence Modeling for Robotic Imitation Learning**. WorldToken maps RGB observations, proprioception and task language into tokens, models their history with a causal Transformer, and predicts action chunks with a diffusion head. The BC-Transformer entry uses the native robomimic baseline. | Paper section | Assets | Checkpoints | |---|---|---:| | [§4 — Multitask control and scaling](04_robocasa_scaling/README.md) | N1–N3 scaling grid, seed 0, D50/D100/D300/D1000/D2900; BC-Transformer at D300, seed 123 | 15 + 1 | | [§5 — Token interface](05_token_interface/README.md) | N2, K=4 and K=50, seed 0, all five data settings | 10 | Each run has its own directory and `run_info.json` with its parameters, original training name and final step. The section indexes list the expected weight filenames; BC-Transformer uses `models/model_epoch_1000.pth`. Each directory also includes a `config.json`, copied from the matching training record with HDF5 data paths normalized to `${DATA_ROOT}/robocasa/mg_im/v0.1/...`. The 25 WorldToken configs also carry the matching public recipe's `eval_window_spec` for offline holdout RMSE; the BC-Transformer uses its separate native evaluation pipeline. ## Release availability All 26 selected checkpoints are available: 15 scaling-grid policies, the BC-Transformer baseline, and 10 token-interface policies. ## Training and evaluation records These policies were trained by imitation learning on the paper's 23-task RoboCasa MG demonstration subsets and evaluated on 23 tasks × 50 episodes, with three full-history repeats per checkpoint. Full details are in the companion [WorldToken Experiment Records](https://huggingface.co/datasets/mepi31415/WorldToken_Experiment_Records) package (`WorldToken_Experiment_Records`), under the same `
//` path: - `config.json` and split files: architecture, training settings and data selection. - Training logs: optimization and holdout metrics. - `rollouts/`: episode outcomes, summary statistics and videos. Use the matching run's results when assessing a checkpoint; paper averages over multiple training seeds are not the score of a single released weight file. ## Usage Download all currently available weights and their configurations with the Hugging Face CLI: ```bash hf download mepi31415/WorldToken_Checkpoints --local-dir ./WorldToken_Checkpoints ``` For only the N2/D300 example below, append `--include "04_robocasa_scaling/scaling_n2_d300_seed0/*"`. Set `CHECKPOINT_ROOT` to the absolute path of the downloaded directory. Follow the code repository's [environment setup](https://github.com/me271828/WorldToken/blob/main/environments/README.md) and [data preparation](https://github.com/me271828/WorldToken/blob/main/environments/DATA.md). Evaluation requires RoboCasa assets and demonstration HDF5 metadata as well as the weight file. From the WorldToken code repository, after setting `DATA_ROOT` and `CHECKPOINT_ROOT` to your local downloads: ```bash RUN=scaling_n2_d300_seed0 SECTION=04_robocasa_scaling EVAL_RUN="$PWD/runs/checkpoint_eval/$RUN" mkdir -p "$EVAL_RUN" cp "$CHECKPOINT_ROOT/$SECTION/$RUN/config.json" "$EVAL_RUN/config.json" python -m experiments.common.rollout \ --config "experiments/$SECTION/configs/$RUN.json" \ --run-dir "$EVAL_RUN" --data-root "$DATA_ROOT" \ --checkpoint "$CHECKPOINT_ROOT/$SECTION/$RUN/checkpoint_step_00030000.pt" ``` This runs one full-history evaluation and saves new outputs under `EVAL_RUN`. The section READMEs in the code repository describe other variants and the baseline. ### Offline holdout RMSE (WorldToken) Use the code repository's Linux/CUDA model environment and prepare the frozen CLIP text encoder as described in its environment instructions. Download the matching experiment-records package as well as the checkpoint. From the code repository root, set these paths to your local copies: ```bash export CODE_ROOT="$PWD" export DATA_ROOT=/absolute/path/to/data export CHECKPOINT_ROOT=/absolute/path/to/WorldToken_Checkpoints export RECORDS_ROOT=/absolute/path/to/WorldToken_Experiment_Records export RUN=scaling_n2_d300_seed0 export SECTION=04_robocasa_scaling export EVAL_RUN="$CODE_ROOT/runs/checkpoint_eval/$RUN" # Optional: an existing CLIP cache from the matching training environment. # export LANG_EMB_CACHE=/absolute/path/to/robocasa_clip.npz ``` Prepare an evaluation copy with the snippet below. Historical split files may use `robocasa/v0.1/...`; it maps each HDF5 path to the published config by its portable `single_stage/...` identity, preserving the demonstration list. It copies an optional existing language cache, or builds one from the holdout language strings using the pinned robomimic CLIP wrapper. `ROBOMIMIC_SRC` is optional when that source is already installed in the environment. The original records and source cache are not modified. ```bash python - <<'PY' import json import os import shutil from pathlib import Path from worldtoken.data import build_lang_embeddings, portable_episode_key from worldtoken.eval_holdout_rmse import load_persisted_holdout_refs from worldtoken.train_utils import expand_path_placeholders run = Path(os.environ["EVAL_RUN"]).resolve() relative = Path(os.environ["SECTION"]) / os.environ["RUN"] records = Path(os.environ["RECORDS_ROOT"]) / relative config_path = Path(os.environ["CHECKPOINT_ROOT"]) / relative / "config.json" config = json.loads(config_path.read_text(encoding="utf-8")) run.mkdir(parents=True, exist_ok=True) paths = config.get("hdf5_paths") or config["dataset"] by_file = {portable_episode_key(p): str(expand_path_placeholders(p).resolve()) for p in paths} if len(by_file) != len(paths): raise ValueError("Duplicate portable HDF5 identities in config") split = json.loads((records / "holdout_demos.json").read_text(encoding="utf-8")) for demo in split["demos"]: demo["hdf5_path"] = by_file[portable_episode_key(demo["hdf5_path"])] (run / "holdout_demos.json").write_text(json.dumps(split, indent=2) + "\n", encoding="utf-8") shutil.copy2(records / "metrics.jsonl.gz", run / "metrics.jsonl.gz") cache = run / "holdout_clip.npz" if os.environ.get("LANG_EMB_CACHE"): source = expand_path_placeholders(os.environ["LANG_EMB_CACHE"]).resolve() if source != cache: shutil.copy2(source, cache) src = os.environ.get("ROBOMIMIC_SRC") # Use the local installation, not an archived source path. src = expand_path_placeholders(src).resolve() if src else None config["lang_emb_cache"] = str(cache) config["robomimic_src"] = str(src) if src else None spec = expand_path_placeholders(config["eval_window_spec"]).resolve() if not spec.is_file(): raise FileNotFoundError(spec) config["eval_window_spec"] = str(spec) (run / "config.json").write_text(json.dumps(config, indent=2) + "\n", encoding="utf-8") build_lang_embeddings( load_persisted_holdout_refs(run), device_arg="cpu", cache_path=cache, robomimic_src=src, mode=config["lang_emb_mode"], write_cache=True, ) print(f"Prepared {run}; fixed windows: {spec}") PY python -m worldtoken.eval_holdout_rmse \ --run-dir "$EVAL_RUN" \ --checkpoint "$CHECKPOINT_ROOT/$SECTION/$RUN/checkpoint_step_00030000.pt" \ --require-fixed-windows ``` For another WorldToken run, change `SECTION`, `RUN` and the weight filename together. Always use the window spec from its matching recipe; the two C=10 files have different historical start frames and cannot be substituted based on history length alone. If evaluating an already prepared historical run whose config lacks this field, obtain the path from that run's public recipe and pass `--eval-window-spec PATH` with `--require-fixed-windows`. The latter prevents silent window reselection. The result under `EVAL_RUN/holdout_grouped_rmse_v2/` includes the actual window path, evaluation protocol, and `legacy_v1_parity` against the final-step row in the copied `metrics.jsonl.gz`. An existing `metrics.jsonl` takes precedence, so use a dedicated evaluation directory for each run. Add `--require-legacy-parity` to fail when no reference is available or the comparison exceeds `--legacy-parity-atol` (default `5e-6`); the result is saved for inspection before a parity failure. Use `--force` to recompute an existing result after changing inputs. Keep `eval_batch_size`, `eval_seed`, `holdout_rmse_samplers`, denoising settings and `precision` unchanged for a paper comparison. Diffusion samples depend on batch boundaries, so reducing batch size can change the answer despite fixed windows. Use a complete matching language cache when available; rebuilding it or changing hardware/software can introduce numerical differences, which should be reported with the parity result. The fixed-window check validates sample selection, not numerical equivalence of the complete evaluator. The `--debug-max-demos` and `--debug-crops-per-demo` subset options cannot be combined with fixed windows, which require the full saved split and crop count. ## Scope and license These are simulation research policies; real-robot performance and transfer to other tasks or observation/action interfaces have not been established. The weights and documentation in this package are licensed under the [MIT License](LICENSE). Third-party materials retain their own licenses. The separate experiment-records package specifies its own license.