Point-cloud world models: Push-T
World-model checkpoints for the Push-T environment from the paper
Does Latent Planning Survive Point Clouds? Action-Conditioned JEPA World Models for Geometric Observations and Goals
(arXiv:2608.29434).
Trained on fafraob/point-cloud-pusht. Code, evaluation scripts and full details:
github.com/fafraob/point-lewm.
| folder | model | observation |
|---|---|---|
point-lewm/ |
Point-LeWM (LeWM objective, PointViT encoder) | point cloud |
point-delta-jepa/ |
Point-Delta-JEPA (Delta-JEPA objective, PointViT encoder) | point cloud |
image-delta-jepa/ |
Delta-JEPA with a ViT-tiny image encoder trained from scratch | RGB |
utonia-wm/ |
JEPA predictor on frozen Utonia voxel features | point cloud |
dino-wm/ |
DINO-WM baseline, frozen DINOv2 ViT-S/14, no proprioception | RGB |
Every folder except dino-wm/ holds weights.pt (state dict) and config.json (Hydra instantiation spec) and loads
with the stable-worldmodel folder loader from the code repository:
from huggingface_hub import snapshot_download
import stable_worldmodel as swm
local = snapshot_download("fafraob/point-cloud-pusht", allow_patterns=["point-lewm/*"]) # or any other folder
model = swm.wm.utils.load_pretrained(f"{local}/point-lewm/")
utonia-wm/ needs the public Utonia backbone (utonia.pth from Pointcept/Utonia, not redistributed here) placed at
checkpoints/utonia/utonia.pth in the code repository. dino-wm/ is in the layout the official
dino_wm code loads (hydra.yaml, checkpoints/model_latest.pth,
swm_h5_norm_stats.json with the action normalisation); see the code repository for the evaluation wrapper.
target2latent goal heads
point-lewm/target2latent/ and point-delta-jepa/target2latent/ hold the typed-3-D-goal heads that map an object
pose (cube position, T pose + ball position, agent position, or wrist + fingertip position) to the goal latent of the
encoder they sit under, so planning needs no recorded goal observation. Four heads each: mlp, mlp_z (also
conditioned on the current latent), shortcut (flow-matching shortcut model, samples in 1 to 64 steps), shortcut_z.
Each folder holds one model.pt (state dict + architecture config); load with target2latent.models.load_goal_model
from the code repository.
Citation
For the most recent citation see the code repository.
@misc{oberweger2026doeslatentplanningsurvive,
title = {Does Latent Planning Survive Point Clouds? Action-Conditioned JEPA World Models for Geometric Observations and Goals},
author = {Fabio F. Oberweger and Michael Schwingshackl and Markus Murschitz},
year = {2026},
eprint = {2608.29434},
archivePrefix = {arXiv},
primaryClass = {cs.LG},
url = {https://arxiv.org/abs/2608.29434},
}