eVTA0
Collection
Dataset and model collection for eVTA0 (https://arxiv.org/abs/2609.33653) • 2 items • Updated
This repository hosts the released checkpoints of the paper "Demonstration-Free Success-Probability Reward Learning for Generalist Robot Policies".
reward_model/): the eVTA₀ success-probability reward
model trained on LIBERO and MetaWorld rollouts (checkpoint_adapter.pth).
The checkpoints contain only the LoRA adapter, the reward-token embedding,
and the prediction head — the backbone is not included; download
Qwen/Qwen3-VL-4B-Instruct
separately and load the checkpoint on top of it. These checkpoints were
serialized under transformers 5.8.0 / peft 0.19.1; loading them requires
transformers >= 5.8.0 and peft >= 0.19.1.pi05_policy/): pi0.5 policies trained with GRPO driven by
the eVTA₀ reward, one per LIBERO suite
(evta0-grpo-6400ep-{suite}-pi05.pt).Code: https://github.com/duowuyms/eVTA0
Paper: https://arxiv.org/abs/2609.33653
Author: Duo Wu
If you find these models useful, please cite our paper:
@article{wu2026evta0,
title={Demonstration-Free Success-Probability Reward Learning for Generalist Robot Policies},
author={Wu, Duo and Wang, Haifeng and Lu, Rongwei and Wang, Jinghe and Xiong, Tianyi and Wang, Zhimin and Yu, Chao and Ma, Shuai and Wang, Zhi},
journal={arXiv preprint arXiv:2609.33653},
year={2026}
}