viscore / README.md
BooBooWu's picture
viscore release
9c67758
|
Raw
History Blame Contribute Delete
8.48 kB
metadata
license: mit
library_name: pytorch
tags:
  - world-model
  - jepa
  - planning
  - model-predictive-control
  - representation-evaluation
pipeline_tag: robotics

VIScore: Diagnosing Planning-Relevant Quality in Latent World Models

arXiv Project Page GitHub Data

Key results:

  • πŸ” Covers all four stages that decide planning success: encoder, predictor, planner, task tolerance
  • πŸ”₯ High correlation with success rate: pooled ρ +0.90, against +0.82 for raw empowerment
  • πŸ’ͺ Best calibration across planners and datasets: below the constant-predictor reference on every pool
  • πŸš€ 20Γ— faster than the real evaluation: ~7 s per checkpoint, no environment and no planner rollouts
  • 🧩 Model-agnostic: three methods (encode, action_embed, predict_next) are the whole interface
  • ♻️ CPU-reproducible: every table in the paper recomputes in seconds

Available Checkpoints

Each file is a full-object torch.save pickle of jepa.JEPA: ViT-Tiny encoder + action-conditioned AdaLN transformer predictor, ~15M parameters, 72 MB. Filenames say which model the weights are β€” vis-wm_epoch_<N>.ckpt, lewm_epoch_<N>.ckpt (the SIGReg baseline) or dino-cls_epoch_<N>.ckpt.

Path Contents Files Size
vis-wm/<task>/seed<S>/ VIS-WM at the epoch the paper reports, + config.yaml 39 0.88 GiB
baselines-lewm/maze2d/seed729/ SIGReg (LeWM) baseline on MAZE 3 0.07 GiB
pools/<pool>/<env>/<run>/ every checkpoint behind a reported number, plus the calibration pool 547 36.7 GiB

VIS-WM: PushT / Reacher / Two-Room / Cube Γ— seeds 403, 46, 729, plus MAZE at seed 729. Reported epoch per arm β€” maximum success on the single-seed d=25 evaluation, ties to the later epoch:

Task seed 403 seed 46 seed 729
PushT 9 10 10
Reacher 5 10 4
Two-Room 5 10 9
Cube 6 8 4
MAZE β€” β€” 8 (SIGReg baseline: 10)

Pools. The split is by training run, so no checkpoint in a test pool comes from a run used to fit anything. Every directory says what its checkpoints are for:

Directory Ckpt / runs Which reported number it produces Evaluation seeds
pools/development/ 137 / 14 block 1 of tab:viscore; the constants are fitted here all available
pools/heldout/ 103 / 33 block 2 -- the headline test, runs disjoint from development 5501, 60601, 90210
pools/heldout-method/ 3 / 3 block 3, our DINO-CLS arms (the other 27 are other groups', not re-hosted) its own
pools/heldout-dataset/ 20 / 2 block 4 -- the unseen MAZE task its own
pools/cube-reference/ 50 / 5 the parenthesised Cube column, and nothing else 5501, 60601, 90210
pools/calibration-fit/ 232 / 47 no reported number: these fit the frozen map blocks 3--4 apply all available

So: 340 checkpoints produce every number in the paper (313 hosted here + 27 fetched with reproduce/download_external.py), 572 if you also want to refit the frozen calibration map from weights, and 0 if you only want to recompute the tables from the shipped metric values.

Per-task counts in the two ranked pools (PushT / Reacher / Two-Room): development 57 / 50 / 30, held-out 27 / 44 / 32.

heldout absorbed what earlier versions reported as a separate terminal pool: 14 of that pool's 17 runs were also in the held-out fold, and restricted to converged epochs its labels no longer resolve individual checkpoints.

cube-reference is separate from heldout because Cube entered no fold. Its label spread does not exceed its own binomial standard error, so nothing there is rankable; the paper prints its correlation in parentheses and excludes it from every pooled and calibration number. These checkpoints are published because that exclusion argument is measured on them.

calibration-fit is the part of the 472-cell calibration pool that is not already in development or heldout (both of which are subsets of it). Its checkpoints enter no correlation and no reported error -- they only determine the metric-to-success map that the two transfer blocks apply without refitting. They are published so that a competing metric can be calibrated on the same 472 cells.

Three SIGReg Two-Room runs landed in neither fold and are not released, which is why Two-Room carries no SIGReg checkpoint in heldout.

The rest of the held-out-method pool is other groups' released checkpoints (Qantara, RC-aux, INTACT), which are not re-hosted here; fetch them with reproduce/download_external.py.

pools/pool_manifest.csv gives each checkpoint's pool membership, success labels and seven metric values; pools/pool_assignment.csv gives run β†’ fold.

Training

LeWM architecture with the VISReg regularizer, Ξ» = 4.5. AdamW, linear warmup + cosine decay, 10 epochs, batch 128, lr 1e-4. Two-Room uses history_size=1, the other tasks 3.

Usage

Score a checkpoint

from viscore import score_checkpoint

f = score_checkpoint("vis-wm/pusht/seed403/vis-wm_epoch_9.ckpt",
                     "probes/probe_pusht.npz", task="pusht")
print(f)          # VIS + veracity / influence / sobriety + their inputs

The pickle names jepa and module as top-level modules; viscore binds the shipped copies automatically, so no LeWorldModel checkout is needed.

Download with huggingface_hub

from huggingface_hub import hf_hub_download, snapshot_download

# one checkpoint
path = hf_hub_download(repo_id="BooBooWu/viscore",
                       filename="vis-wm/pusht/seed403/vis-wm_epoch_9.ckpt")

# a whole tree
snapshot_download(repo_id="BooBooWu/viscore", allow_patterns=["vis-wm/*"])

Build a probe

A probe is one frozen slice of training data (rng(0), 300 episodes, frameskip 5) shared by every checkpoint compared.

viscore probe --tasks pusht --data-home $STABLEWM_HOME --out-dir probes
viscore score --task pusht --probe probes/probe_pusht.npz \
                --run-dir vis-wm/pusht/seed403 --epochs 9 --csv out.csv

Evaluation

Both of the paper's main tables recompute from files shipped in the GitHub repo β€” no GPU, dataset or checkpoint download:

git clone https://github.com/HaiyuWu/viscore && cd viscore && pip install -e .

python reproduce/tables.py                     # metric vs success, three pools
python reproduce/planning_tables.py --strict    # planning tables, gated against published values

Notes: peak epoch = maximum evaluation success rate, ties to the later epoch. Cube is reported in parentheses and excluded from pooled and calibration columns. All published labels and sobriety probes use CEM. A single 50-episode evaluation has a success sd of β‰ˆ6 points.

Citation

@article{wu2026viscore,
  title         = {VIScore: Diagnosing Planning-Relevant Quality in Latent World Models},
  author        = {Wu, Haiyu and Balestriero, Randall and Levine, Morgan},
  journal       = {arXiv preprint arXiv:2608.11174},
  year          = {2026},
  eprint        = {2608.11174},
  archivePrefix = {arXiv}
}

Architecture and training recipe from LeWorldModel; regularizer from VISReg.

License

This project (code and checkpoints) is released under the MIT License.