ERVLA LIBERO-Plus

Revisiting Embodied Chain-of-Thought for Generalizable Robot Manipulation · Project page

ERVLA is a vision-language-action model for robot manipulation. It learns from embodied chain-of-thought that links task understanding to end-effector movements and image-space trajectories. With reasoning dropout, the model can generate actions directly at inference without decoding a chain of thought.

This model is post-trained on LIBERO and achieves 86.9% overall success rate on LIBERO-Plus through zero-shot transfer.

Downloads last month
22
Safetensors
Model size
5B params
Tensor type
BF16
·
Video Preview
loading

Paper for ERVLA/liberoplus