GALA: Geometry-Aware Latent Action Modeling for Vision-Language-Action Model Pretraining across Embodiments
Paper • 2609.21948 • Published
GALA inference checkpoint checkpoint_robocasa_gr1 for RoboCasa GR1 tabletop evaluation.
See the GitHub repository for environment setup and evaluation instructions.
@article{liu2026gala,
title={GALA: Geometry-Aware Latent Action Modeling for Vision-Language-Action Model Pretraining across Embodiments},
author={Liu, Yichen and Yuan, Puzhen and Zhu, Xiang and Guo, Yanjiang and Chen, Jianyu},
journal={arXiv preprint arXiv:2609.21948},
year={2026}
}
Base model
Qwen/Qwen2.5-VL-3B-Instruct