GeneralVLA-2 / README.md
wwhyyyyyy's picture
Add selected runtime weights and GeoFuse-MV3D examples
3a1fe8b verified
|
Raw History Blame Contribute Delete
3.08 kB
---
license: apache-2.0
language:
- en
- zh
library_name: pytorch
tags:
- robotics
- vision-language-action
- robot-planning
- 3d-reconstruction
- long-term-memory
- training-free
---
# GeneralVLA-2: Geometry-Aware Reconstruction and Governed Memory for Robot Planning
GeneralVLA-2 is a training-free framework that strengthens robot planning with two complementary components:
1. **GeoFuse-MV3D** produces conservative, mask-verified and appearance-preserving object geometry from calibrated multi-view observations.
2. **Governed KnowledgeBank** provides quality-, confidence-, lifecycle-, conflict- and geometry-aware experience retrieval.
GeneralVLA-2 does not introduce a newly trained monolithic checkpoint. This repository bundles selected runtime weights and representative GeoFuse-MV3D 3D Gaussian reconstruction outputs so that the commonly used vision and grasping components can be obtained from one repository.
## Repository contents
- `weights/sam/sam_vit_h_4b8939.pth`: SAM ViT-H segmentation checkpoint.
- `weights/clip-vit-large-patch14/`: CLIP ViT-L/14 vision-language encoder and tokenizer files.
- `weights/graspnet/checkpoint-rs.tar`: GraspNet checkpoint used by the robot stack.
- `geofuse_mv3d/examples/alarm_refined.ply`: refined 3D Gaussian reconstruction for the alarm scene.
- `geofuse_mv3d/examples/shoe_refined.ply`: refined 3D Gaussian reconstruction for the shoe scene.
- `generalvla2.yaml`: compact deployment-oriented configuration.
- `asset_manifest.json`: machine-readable description of the bundled assets.
The PLY files are method outputs rather than neural-network checkpoints. They contain position, appearance, opacity, scale and rotation attributes for the reconstructed 3D Gaussians.
## Download
```bash
hf download AIGeeksGroup/GeneralVLA-2 \
--local-dir vendor/GeneralVLA-2
```
## Code and project page
- Code: https://github.com/AIGeeksGroup/GeneralVLA-2
- Project page: https://aigeeksgroup.github.io/GeneralVLA-2
- Paper: https://arxiv.org/abs/2606.17480
## Reported results
- On GSO-30, GeoFuse-MV3D reduces CD and LPIPS by 2.20% and 2.02%, while increasing PSNR and SSIM by 2.36% and 1.03% over the MV-SAM3D baseline under the same input-view protocol.
- Governed KnowledgeBank improves over ReasoningBank by 4.53% on Terminal-Bench 2.0 success rate and 3.73% on SWE-Bench Verified resolve rate, while reducing average steps by 4.95% and 5.65%.
- GeneralVLA-2 produces successful trajectories on all 14 evaluated RLBench tasks without parameter training and is best among the compared methods on 10 tasks.
## License notice
The Apache-2.0 metadata applies to the GeneralVLA-2 configuration and documentation. Bundled third-party checkpoints retain their respective upstream licenses and terms.
## Citation
```bibtex
@article{wang2026generalvla2,
title={GeneralVLA-2: Geometry-Aware Reconstruction and Governed Memory for Robot Planning},
author={Wang, Haoyu and Ma, Guoqing and Zhang, Zeyu and Guo, Yandong and Shi, Boxin and Tang, Hao},
journal={arXiv preprint arXiv:2606.17480},
year={2026}
}
```