Instructions to use kabilanKB/yam_put_pot_act_m3 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LeRobot
How to use kabilanKB/yam_put_pot_act_m3 with LeRobot:
- Notebooks
- Google Colab
- Kaggle
File size: 4,136 Bytes
7511d13 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 | ---
license: apache-2.0
library_name: lerobot
pipeline_tag: robotics
tags:
- LeRobot
- act
- yamlab
- bimanual
- yam
- isaac-sim
- simulation
- modality-masking
datasets:
- kabilanKB/yam_put_pot
---
# ACT + M3 for bimanual YAM: PutPotOnCooktop
An [ACT](https://arxiv.org/abs/2304.13705) policy for the simulated bimanual YAM robot, trained
on [`kabilanKB/yam_put_pot`](https://huggingface.co/datasets/kabilanKB/yam_put_pot) with M3
modality masking, adapted from "Robust Bimanual Vision-Language-Action Models via Embarrassingly
Simple Modality Masking" ([arXiv 2608.22419](https://arxiv.org/abs/2608.22419)).
The task is PutPotOnCooktop in Isaac Sim 6 / Isaac Lab 3: both arms grasp a pot by its handles,
lift it and place it on a cooktop.
Code: [kabilankb/Isaac-Lab-3-and-add-SO-101-leader-teleoperation](https://github.com/kabilankb/Isaac-Lab-3-and-add-SO-101-leader-teleoperation)
(see `ACT_M3.md`).
> **This is an early research checkpoint.** It completes the task in roughly one episode in
> four, and M3 masking showed no measurable benefit over plain ACT in the evaluation below.
## Model
| | |
|---|---|
| Inputs | Three 320×240 RGB cameras (`top_rgb`, `left_rgb`, `right_rgb`) and the 14-value joint state |
| Output | Chunk of 100 actions × 14 joint-position targets (6 joints + gripper per arm) |
| Backbone | ImageNet-pretrained ResNet-18, shared across cameras |
| Transformer | 4 encoder layers, 4 decoder layers, width 512, 8 heads, VAE latent 32 |
| Size | 68M parameters |
M3 is applied during training only; this checkpoint is an ordinary ACT checkpoint.
- **Wrist-camera masking:** for 30% of training samples, both wrist cameras are hidden from
attention together. The top camera is never hidden.
- **Query masking:** each of the 100 action queries is hidden from the others with probability
0.1, and the visible ones are rescaled by 1/0.9.
- The paper's language masking does not apply, since ACT has no language input.
## Training
| | |
|---|---|
| Data | `kabilanKB/yam_put_pot`: 20 MimicGen episodes, 7,779 frames at 30 FPS, domain randomization on |
| Steps / batch size | 20,000 / 32 (about 82 passes over the data) |
| Optimizer | AdamW, learning rate 1e-5 |
| Seed | 1000 (single run) |
| Final training loss | 0.041 |
## Evaluation
PutPotOnCooktop-v0 with `pot_000` / `cooktop_000`, plain scene (no domain randomization),
25 actions executed per prediction, 900-step limit, 50 episodes, 10 parallel environments.
| Model | Full task succeeded | Pot lifted |
|---|---|---|
| ACT + M3 (this checkpoint) | 11 of 50 (22%) | 13 of 50 (26%) |
| Plain ACT, same settings | 12 of 50 (24%) | 15 of 50 (30%) |
- The two models are indistinguishable on this test.
- The rates are slightly optimistic: the run stopped at the first 50 finished episodes, and
successes finish sooner than timeouts.
- Not tested: randomized or cluttered scenes, which is where the paper claims M3 helps.
- The dataset is small (two demonstrations per pot and cooktop pair); more data is the most
likely way to raise the success rate.
## Usage
Requires the LeRobot v2.0 fork used by YAMLab (`RogerDAI1217/lerobot`, branch `lerobotv2.0`).
Its `config.json` has no `type` key, so pass the config explicitly:
```python
import draccus
from huggingface_hub import snapshot_download
from lerobot.common.policies.act.configuration_act import ACTConfig
from lerobot.common.policies.act.modeling_act import ACTPolicy
ckpt = snapshot_download("kabilanKB/yam_put_pot_act_m3")
with draccus.config_type("json"):
config = draccus.parse(ACTConfig, f"{ckpt}/config.json", args=[])
config.n_action_steps = 25
policy = ACTPolicy.from_pretrained(ckpt, config=config)
```
To run it in the simulator, download the checkpoint and use `scripts/eval/eval_act.py` from the
code repository:
```bash
D=yamlab_datasets/tasks_data/PutPotOnCooktop/objects
python scripts/eval/eval_act.py --checkpoint <downloaded folder> --n_action_steps 25 \
--task PutPotOnCooktop-v0 --asset pot=$D/Pot/pot_000 --asset cooktop=$D/Cooktop/cooktop_000 \
--num_episodes 5 --enable_gripper_clamp --enable_cameras --viz kit
```
|