--- license: mit --- # eVTA₀ Models This repository hosts the released checkpoints of the paper ["Demonstration-Free Success-Probability Reward Learning for Generalist Robot Policies"](https://arxiv.org/abs/2609.33653). - **Reward models** (`reward_model/`): the eVTA₀ success-probability reward model trained on LIBERO and MetaWorld rollouts (`checkpoint_adapter.pth`). The checkpoints contain only the LoRA adapter, the reward-token embedding, and the prediction head — the backbone is not included; download [Qwen/Qwen3-VL-4B-Instruct](https://huggingface.co/Qwen/Qwen3-VL-4B-Instruct) separately and load the checkpoint on top of it. These checkpoints were serialized under transformers 5.8.0 / peft 0.19.1; loading them requires `transformers >= 5.8.0` and `peft >= 0.19.1`. - **RL policies** (`pi05_policy/`): pi0.5 policies trained with GRPO driven by the eVTA₀ reward, one per LIBERO suite (`evta0-grpo-6400ep-{suite}-pi05.pt`). **Code:** https://github.com/duowuyms/eVTA0 **Paper:** https://arxiv.org/abs/2609.33653 **Author:** [Duo Wu](https://duowuyms.github.io/) ## Citation If you find these models useful, please cite our paper: ```bibtex @article{wu2026evta0, title={Demonstration-Free Success-Probability Reward Learning for Generalist Robot Policies}, author={Wu, Duo and Wang, Haifeng and Lu, Rongwei and Wang, Jinghe and Xiong, Tianyi and Wang, Zhimin and Yu, Chao and Ma, Shuai and Wang, Zhi}, journal={arXiv preprint arXiv:2609.33653}, year={2026} } ```