eVTA₀ Models

This repository hosts the released checkpoints of the paper "Demonstration-Free Success-Probability Reward Learning for Generalist Robot Policies".

  • Reward models (reward_model/): the eVTA₀ success-probability reward model trained on LIBERO and MetaWorld rollouts (checkpoint_adapter.pth). The checkpoints contain only the LoRA adapter, the reward-token embedding, and the prediction head — the backbone is not included; download Qwen/Qwen3-VL-4B-Instruct separately and load the checkpoint on top of it. These checkpoints were serialized under transformers 5.8.0 / peft 0.19.1; loading them requires transformers >= 5.8.0 and peft >= 0.19.1.
  • RL policies (pi05_policy/): pi0.5 policies trained with GRPO driven by the eVTA₀ reward, one per LIBERO suite (evta0-grpo-6400ep-{suite}-pi05.pt).

Code: https://github.com/duowuyms/eVTA0

Paper: https://arxiv.org/abs/2609.33653

Author: Duo Wu

Citation

If you find these models useful, please cite our paper:

@article{wu2026evta0,
  title={Demonstration-Free Success-Probability Reward Learning for Generalist Robot Policies},
  author={Wu, Duo and Wang, Haifeng and Lu, Rongwei and Wang, Jinghe and Xiong, Tianyi and Wang, Zhimin and Yu, Chao and Ma, Shuai and Wang, Zhi},
  journal={arXiv preprint arXiv:2609.33653},
  year={2026}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including notmuch2/eVTA0_models

Paper for notmuch2/eVTA0_models