Robotics
LeRobot
Safetensors
vla
bimanual
robocolosseum

BimanualYAM-models (RoboColosseum)

These are vision-language-action (VLA) models fine-tuned on BimanualYAM manipulation tasks for the RoboColosseum benchmark. Each VLA is fine-tuned from its public base checkpoint with its repo-default recipe: default LRs and frozen parts, with the LR scaled linearly with the effective batch size, for 5 epochs. Each folder holds the final-epoch inference weights plus the small config files needed to load the model.

Layout: <model>/<task>/

Model Base checkpoint Dustpan Microwave Drawer Cups
gr00t/ GR00T N1.7 nvidia/GR00T-N1.7-3B ✅ ⏳ ⏳ ⏳
pi05/ π0.5 gs://openpi-assets/checkpoints/pi05_base ✅ ⏳ ⏳ ⏳
g05/ G0.5 OpenGalaxea/G05 (g05-base) ✅ ⏳ ⏳ ⏳
molmoact2/ MolmoAct2 allenai/MolmoAct2 ✅ ⏳ ⏳ ⏳
lingbot-vla-v2/ LingBot-VLA v2 6B robbyant/lingbot-vla-v2-6b ✅ ⏳ ⏳ ⏳

Each <model>/<task>/README.md gives the recipe, the files, and how to load the model.

Tasks (RoboColosseum/BimanualYAM-datasets, LeRobot v3.0, 30 fps)

Task Instruction Episodes Frames
Dustpan "Clean the table." 101 55,046
Microwave "Open the microwave and take out the bowl." 100 108,487
Drawer "Put the cup into the drawer." 101 82,305
Cups "Stack the cups." 100 39,184

Common setup:

  • Cameras: top (360x640), left wrist and right wrist (480x640).
  • State and action: 14-D absolute joint angles [left arm 6, left gripper 1, right arm 6, right gripper 1].
  • Training: fp32 master weights with bf16 compute, on 4x H100.

Recipes (same for every task)

Model Trained parts GBS Peak LR
GR00T N1.7 Full: VLM backbone unfrozen + action head 64 1e-4, cosine, 5% warmup
π0.5 Full (openpi), EMA 0.99 64 2.5e-5, cosine to 2.5e-6, 1k warmup
G0.5 Full (vision LR x0.1), FSDP 32 4e-5, cosine, 200 warmup, wd 0.03
MolmoAct2 Full 128 LLM 2e-5, ViT 1e-5, connector 1e-5, action expert 1e-4
LingBot-VLA v2 Full 64 1.25e-5 (5e-5 x 64/256), constant, Muon

Training logs are in the W&B project shaileshxml-nus/RoboColosseum.

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Dataset used to train RoboColosseum/BimanualYAM-models