File size: 4,136 Bytes
7511d13
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
---
license: apache-2.0
library_name: lerobot
pipeline_tag: robotics
tags:
  - LeRobot
  - act
  - yamlab
  - bimanual
  - yam
  - isaac-sim
  - simulation
  - modality-masking
datasets:
  - kabilanKB/yam_put_pot
---

# ACT + M3 for bimanual YAM: PutPotOnCooktop

An [ACT](https://arxiv.org/abs/2304.13705) policy for the simulated bimanual YAM robot, trained
on [`kabilanKB/yam_put_pot`](https://huggingface.co/datasets/kabilanKB/yam_put_pot) with M3
modality masking, adapted from "Robust Bimanual Vision-Language-Action Models via Embarrassingly
Simple Modality Masking" ([arXiv 2608.22419](https://arxiv.org/abs/2608.22419)).

The task is PutPotOnCooktop in Isaac Sim 6 / Isaac Lab 3: both arms grasp a pot by its handles,
lift it and place it on a cooktop.

Code: [kabilankb/Isaac-Lab-3-and-add-SO-101-leader-teleoperation](https://github.com/kabilankb/Isaac-Lab-3-and-add-SO-101-leader-teleoperation)
(see `ACT_M3.md`).

> **This is an early research checkpoint.** It completes the task in roughly one episode in
> four, and M3 masking showed no measurable benefit over plain ACT in the evaluation below.

## Model

| | |
|---|---|
| Inputs | Three 320×240 RGB cameras (`top_rgb`, `left_rgb`, `right_rgb`) and the 14-value joint state |
| Output | Chunk of 100 actions × 14 joint-position targets (6 joints + gripper per arm) |
| Backbone | ImageNet-pretrained ResNet-18, shared across cameras |
| Transformer | 4 encoder layers, 4 decoder layers, width 512, 8 heads, VAE latent 32 |
| Size | 68M parameters |

M3 is applied during training only; this checkpoint is an ordinary ACT checkpoint.

- **Wrist-camera masking:** for 30% of training samples, both wrist cameras are hidden from
  attention together. The top camera is never hidden.
- **Query masking:** each of the 100 action queries is hidden from the others with probability
  0.1, and the visible ones are rescaled by 1/0.9.
- The paper's language masking does not apply, since ACT has no language input.

## Training

| | |
|---|---|
| Data | `kabilanKB/yam_put_pot`: 20 MimicGen episodes, 7,779 frames at 30 FPS, domain randomization on |
| Steps / batch size | 20,000 / 32 (about 82 passes over the data) |
| Optimizer | AdamW, learning rate 1e-5 |
| Seed | 1000 (single run) |
| Final training loss | 0.041 |

## Evaluation

PutPotOnCooktop-v0 with `pot_000` / `cooktop_000`, plain scene (no domain randomization),
25 actions executed per prediction, 900-step limit, 50 episodes, 10 parallel environments.

| Model | Full task succeeded | Pot lifted |
|---|---|---|
| ACT + M3 (this checkpoint) | 11 of 50 (22%) | 13 of 50 (26%) |
| Plain ACT, same settings | 12 of 50 (24%) | 15 of 50 (30%) |

- The two models are indistinguishable on this test.
- The rates are slightly optimistic: the run stopped at the first 50 finished episodes, and
  successes finish sooner than timeouts.
- Not tested: randomized or cluttered scenes, which is where the paper claims M3 helps.
- The dataset is small (two demonstrations per pot and cooktop pair); more data is the most
  likely way to raise the success rate.

## Usage

Requires the LeRobot v2.0 fork used by YAMLab (`RogerDAI1217/lerobot`, branch `lerobotv2.0`).
Its `config.json` has no `type` key, so pass the config explicitly:

```python
import draccus
from huggingface_hub import snapshot_download
from lerobot.common.policies.act.configuration_act import ACTConfig
from lerobot.common.policies.act.modeling_act import ACTPolicy

ckpt = snapshot_download("kabilanKB/yam_put_pot_act_m3")
with draccus.config_type("json"):
    config = draccus.parse(ACTConfig, f"{ckpt}/config.json", args=[])
config.n_action_steps = 25
policy = ACTPolicy.from_pretrained(ckpt, config=config)
```

To run it in the simulator, download the checkpoint and use `scripts/eval/eval_act.py` from the
code repository:

```bash
D=yamlab_datasets/tasks_data/PutPotOnCooktop/objects
python scripts/eval/eval_act.py --checkpoint <downloaded folder> --n_action_steps 25 \
    --task PutPotOnCooktop-v0 --asset pot=$D/Pot/pot_000 --asset cooktop=$D/Cooktop/cooktop_000 \
    --num_episodes 5 --enable_gripper_clamp --enable_cameras --viz kit
```