Robotics
world-model
jepa
point-cloud
planning

Point-cloud world models: Two-Room

World-model checkpoints for the Two-Room environment from the paper Does Latent Planning Survive Point Clouds? Action-Conditioned JEPA World Models for Geometric Observations and Goals (arXiv:2608.29434). Trained on fafraob/point-cloud-tworoom. Code, evaluation scripts and full details: github.com/fafraob/point-lewm.

folder model observation
point-lewm/ Point-LeWM (LeWM objective, PointViT encoder) point cloud
point-delta-jepa/ Point-Delta-JEPA (Delta-JEPA objective, PointViT encoder) point cloud
image-delta-jepa/ Delta-JEPA with a ViT-tiny image encoder trained from scratch RGB
utonia-wm/ JEPA predictor on frozen Utonia voxel features point cloud
dino-wm/ DINO-WM baseline, frozen DINOv2 ViT-S/14, no proprioception RGB

Every folder except dino-wm/ holds weights.pt (state dict) and config.json (Hydra instantiation spec) and loads with the stable-worldmodel folder loader from the code repository:

from huggingface_hub import snapshot_download
import stable_worldmodel as swm

local = snapshot_download("fafraob/point-cloud-tworoom", allow_patterns=["point-lewm/*"])  # or any other folder
model = swm.wm.utils.load_pretrained(f"{local}/point-lewm/")

utonia-wm/ needs the public Utonia backbone (utonia.pth from Pointcept/Utonia, not redistributed here) placed at checkpoints/utonia/utonia.pth in the code repository. dino-wm/ is in the layout the official dino_wm code loads (hydra.yaml, checkpoints/model_latest.pth, swm_h5_norm_stats.json with the action normalisation); see the code repository for the evaluation wrapper.

target2latent goal heads

point-lewm/target2latent/ and point-delta-jepa/target2latent/ hold the typed-3-D-goal heads that map an object pose (cube position, T pose + ball position, agent position, or wrist + fingertip position) to the goal latent of the encoder they sit under, so planning needs no recorded goal observation. Four heads each: mlp, mlp_z (also conditioned on the current latent), shortcut (flow-matching shortcut model, samples in 1 to 64 steps), shortcut_z. Each folder holds one model.pt (state dict + architecture config); load with target2latent.models.load_goal_model from the code repository.

Citation

For the most recent citation see the code repository.

@misc{oberweger2026doeslatentplanningsurvive,
  title         = {Does Latent Planning Survive Point Clouds? Action-Conditioned JEPA World Models for Geometric Observations and Goals},
  author        = {Fabio F. Oberweger and Michael Schwingshackl and Markus Murschitz},
  year          = {2026},
  eprint        = {2608.29434},
  archivePrefix = {arXiv},
  primaryClass  = {cs.LG},
  url           = {https://arxiv.org/abs/2608.29434},
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Dataset used to train fafraob/point-cloud-tworoom

Collection including fafraob/point-cloud-tworoom

Papers for fafraob/point-cloud-tworoom