--- license: apache-2.0 library_name: lerobot pipeline_tag: robotics tags: - LeRobot - act - yamlab - bimanual - yam - isaac-sim - simulation - modality-masking datasets: - kabilanKB/yam_put_pot --- # ACT + M3 for bimanual YAM: PutPotOnCooktop An [ACT](https://arxiv.org/abs/2304.13705) policy for the simulated bimanual YAM robot, trained on [`kabilanKB/yam_put_pot`](https://huggingface.co/datasets/kabilanKB/yam_put_pot) with M3 modality masking, adapted from "Robust Bimanual Vision-Language-Action Models via Embarrassingly Simple Modality Masking" ([arXiv 2608.22419](https://arxiv.org/abs/2608.22419)). The task is PutPotOnCooktop in Isaac Sim 6 / Isaac Lab 3: both arms grasp a pot by its handles, lift it and place it on a cooktop. Code: [kabilankb/Isaac-Lab-3-and-add-SO-101-leader-teleoperation](https://github.com/kabilankb/Isaac-Lab-3-and-add-SO-101-leader-teleoperation) (see `ACT_M3.md`). > **This is an early research checkpoint.** It completes the task in roughly one episode in > four, and M3 masking showed no measurable benefit over plain ACT in the evaluation below. ## Model | | | |---|---| | Inputs | Three 320×240 RGB cameras (`top_rgb`, `left_rgb`, `right_rgb`) and the 14-value joint state | | Output | Chunk of 100 actions × 14 joint-position targets (6 joints + gripper per arm) | | Backbone | ImageNet-pretrained ResNet-18, shared across cameras | | Transformer | 4 encoder layers, 4 decoder layers, width 512, 8 heads, VAE latent 32 | | Size | 68M parameters | M3 is applied during training only; this checkpoint is an ordinary ACT checkpoint. - **Wrist-camera masking:** for 30% of training samples, both wrist cameras are hidden from attention together. The top camera is never hidden. - **Query masking:** each of the 100 action queries is hidden from the others with probability 0.1, and the visible ones are rescaled by 1/0.9. - The paper's language masking does not apply, since ACT has no language input. ## Training | | | |---|---| | Data | `kabilanKB/yam_put_pot`: 20 MimicGen episodes, 7,779 frames at 30 FPS, domain randomization on | | Steps / batch size | 20,000 / 32 (about 82 passes over the data) | | Optimizer | AdamW, learning rate 1e-5 | | Seed | 1000 (single run) | | Final training loss | 0.041 | ## Evaluation PutPotOnCooktop-v0 with `pot_000` / `cooktop_000`, plain scene (no domain randomization), 25 actions executed per prediction, 900-step limit, 50 episodes, 10 parallel environments. | Model | Full task succeeded | Pot lifted | |---|---|---| | ACT + M3 (this checkpoint) | 11 of 50 (22%) | 13 of 50 (26%) | | Plain ACT, same settings | 12 of 50 (24%) | 15 of 50 (30%) | - The two models are indistinguishable on this test. - The rates are slightly optimistic: the run stopped at the first 50 finished episodes, and successes finish sooner than timeouts. - Not tested: randomized or cluttered scenes, which is where the paper claims M3 helps. - The dataset is small (two demonstrations per pot and cooktop pair); more data is the most likely way to raise the success rate. ## Usage Requires the LeRobot v2.0 fork used by YAMLab (`RogerDAI1217/lerobot`, branch `lerobotv2.0`). Its `config.json` has no `type` key, so pass the config explicitly: ```python import draccus from huggingface_hub import snapshot_download from lerobot.common.policies.act.configuration_act import ACTConfig from lerobot.common.policies.act.modeling_act import ACTPolicy ckpt = snapshot_download("kabilanKB/yam_put_pot_act_m3") with draccus.config_type("json"): config = draccus.parse(ACTConfig, f"{ckpt}/config.json", args=[]) config.n_action_steps = 25 policy = ACTPolicy.from_pretrained(ckpt, config=config) ``` To run it in the simulator, download the checkpoint and use `scripts/eval/eval_act.py` from the code repository: ```bash D=yamlab_datasets/tasks_data/PutPotOnCooktop/objects python scripts/eval/eval_act.py --checkpoint --n_action_steps 25 \ --task PutPotOnCooktop-v0 --asset pot=$D/Pot/pot_000 --asset cooktop=$D/Cooktop/cooktop_000 \ --num_episodes 5 --enable_gripper_clamp --enable_cameras --viz kit ```