Outcome-Grounded World Modeling for Autonomous Driving
Project page · Code (GitHub) · Film · Paper (arXiv)
Jieyuan Pei, Meiyi Lu, Sining Ang, Yubo Zhao, Zhangyi Hu, Mingwei Xu, Haokai Ding, Wei Li, Zihan You, Jianwei Zheng, Li Yu, Yifeng Pan, Ji Tao, Rongjunchen Zhang, Yan Wang
The 95-second overview film, with sound. Also on the project page.
About
Checkpoints of World4Scorer (paper, code).
World4Scorer makes the candidate scorer of a generate-and-select planner a trajectory-conditioned latent world model. One shared predictor gives every candidate trajectory a predicted state, and score heads read its outcomes (collision, drivable area, progress, time to collision, comfort) from that state. Simulator outcomes supervise every candidate; the one future the driving log actually recorded anchors the shared predictor during training. No future is needed at test time.
Results
| Benchmark | Metric | World4Scorer | Best other method in the paper |
|---|---|---|---|
| NAVSIM-v2 navtest | EPDMS | 93.0 (91.4 without inertial re-ranking) | Drive-JEPA 90.8 |
| NAVSIM-v1 navtest | PDMS | 94.0 | DrivoR 93.7 · Drive-JEPA 93.7 |
| Bench2Drive-220, closed loop | Driving Score | 73.27 | ReCogDrive 71.36 |
| OGBench-Cube, LeWM planner | Success rate | 73.64% | LeWM 68.91% (same world model and budget) |
The Bench2Drive system is adapted (official 1000-clip training subset, a route point and two deployment rules). Cube numbers are our runs over 11 seeds × 50 episodes.
Demos
![]() NAVSIM, Las Vegas. 64 candidates every 0.5 s, coloured by predicted score. |
![]() OGBench-Cube. Left: LeWM alone fails. Right: with the outcome score it succeeds. |
![]() Bench2Drive. A taxi cuts into the lane (closed loop, 1.5×). |
![]() Bench2Drive. A pedestrian steps into the road at night (closed loop, 1.5×). |
Checkpoints
| File | Model | Result | SHA-256 |
|---|---|---|---|
weights/w4s_navsim_ep23.ckpt |
NAVSIM planner, seed 3, epoch 23, 35.5M parameters | 91.4 EPDMS · 93.0 with re-ranking · 94.0 PDMS | 22604f83dda8a2bd2dd05356fca8457f0c7dc78439f47db396e5c85d74971a44 |
weights/w4s_b2d_route_tp_ep14.ckpt |
Bench2Drive planner, route point, epoch 14 | 73.27 Driving Score | ed3f61a9ac850627c07bbf1d37cc0a9e123d5b41a9c1cecefe820644e14aff5b |
weights/w4s_b2d_route_blind_ep14.ckpt |
Bench2Drive route-blind model, epoch 14 | initialization of the route-point model | 640dd42934b57d9305c336db7234747533977ad916c18723a9bb97b385407d31 |
ogbench_cube/heads/n10_r1/head.pt |
OGBench-Cube outcome head n10_r1 |
73.64% success | fce54aa53576587f002a8a515db393bc346b1daee2716630ecf14352fb08bf0e |
weights/realized_future_emb_navtrain_f7_camf0.pt |
NAVSIM future targets (navtrain) | used only in training | 2d64716d8f07a0dbc94e3f240108167acf9c9ef5093635eacb857725a13f0d82 |
From the root of the code repository, this puts every file where the scripts expect it:
hf download pei2333/World4Scorer --include "weights/*" --include "ogbench_cube/*" --local-dir .
The SHA-256 values identify the exact files behind the reported numbers; check a download with python tools/verify_release.py --artifact KIND=PATH in the code repository.
Citation
@article{pei2026world4scorer,
title = {World4Scorer: Outcome-Grounded World Modeling for Autonomous Driving},
author = {Pei, Jieyuan and Lu, Meiyi and Ang, Sining and Zhao, Yubo and Hu, Zhangyi and Xu, Mingwei and Ding, Haokai and Li, Wei and You, Zihan and Zheng, Jianwei and Yu, Li and Pan, Yifeng and Tao, Ji and Zhang, Rongjunchen and Wang, Yan},
journal = {arXiv preprint arXiv:2609.36438},
year = {2026}
}



