RoboColosseum/BimanualYAM-datasets
Viewer • Updated • 335k • 74
How to use RoboColosseum/BimanualYAM-models with LeRobot:
These are vision-language-action (VLA) models fine-tuned on BimanualYAM manipulation tasks for the RoboColosseum benchmark. Each VLA is fine-tuned from its public base checkpoint with its repo-default recipe: default LRs and frozen parts, with the LR scaled linearly with the effective batch size, for 5 epochs. Each folder holds the final-epoch inference weights plus the small config files needed to load the model.
<model>/<task>/
| Model | Base checkpoint | Dustpan | Microwave | Drawer | Cups |
|---|---|---|---|---|---|
gr00t/ GR00T N1.7 |
nvidia/GR00T-N1.7-3B |
✅ | ⏳ | ⏳ | ⏳ |
pi05/ π0.5 |
gs://openpi-assets/checkpoints/pi05_base |
✅ | ⏳ | ⏳ | ⏳ |
g05/ G0.5 |
OpenGalaxea/G05 (g05-base) |
✅ | ⏳ | ⏳ | ⏳ |
molmoact2/ MolmoAct2 |
allenai/MolmoAct2 |
✅ | ⏳ | ⏳ | ⏳ |
lingbot-vla-v2/ LingBot-VLA v2 6B |
robbyant/lingbot-vla-v2-6b |
✅ | ⏳ | ⏳ | ⏳ |
Each <model>/<task>/README.md gives the recipe, the files, and how to load the model.
RoboColosseum/BimanualYAM-datasets, LeRobot v3.0, 30 fps)
| Task | Instruction | Episodes | Frames |
|---|---|---|---|
| Dustpan | "Clean the table." | 101 | 55,046 |
| Microwave | "Open the microwave and take out the bowl." | 100 | 108,487 |
| Drawer | "Put the cup into the drawer." | 101 | 82,305 |
| Cups | "Stack the cups." | 100 | 39,184 |
Common setup:
[left arm 6, left gripper 1, right arm 6, right gripper 1].| Model | Trained parts | GBS | Peak LR |
|---|---|---|---|
| GR00T N1.7 | Full: VLM backbone unfrozen + action head | 64 | 1e-4, cosine, 5% warmup |
| π0.5 | Full (openpi), EMA 0.99 | 64 | 2.5e-5, cosine to 2.5e-6, 1k warmup |
| G0.5 | Full (vision LR x0.1), FSDP | 32 | 4e-5, cosine, 200 warmup, wd 0.03 |
| MolmoAct2 | Full | 128 | LLM 2e-5, ViT 1e-5, connector 1e-5, action expert 1e-4 |
| LingBot-VLA v2 | Full | 64 | 1.25e-5 (5e-5 x 64/256), constant, Muon |
Training logs are in the W&B project shaileshxml-nus/RoboColosseum.