File size: 4,109 Bytes
5d08972
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
# Action-layout check for gr00t_views datasets

Found 2026-09-20 while investigating a 0/160 eval. Use it on any gr00t_views dataset before
training on it.

## The bug

RoboCasa's simulator emits a 12-d action as

```
[eef_pos(3), eef_rot(3), gripper(1), base(3), torso(1), base_mode(1)]      robosuite order
```

A gr00t_views dataset's `meta/modality.json` declares

```
base_motion[0:4]  control_mode[4:5]  eef_pos[5:8]  eef_rot[8:11]  gripper_close[11:12]
```

which is a **different order**. A builder that copies simulator actions straight into the
parquet writes correct numbers under the wrong column names. Nothing errors, the loss
converges to a small value, and the resulting policy is useless: the slices land as

| modality.json reads | what is actually there |
|---|---|
| `base_motion[0:4]` | eef x, y, z and rot x |
| `control_mode[4:5]` | rot y |
| `eef_pos[5:8]` | rot z, **gripper**, 0 |
| `eef_rot[8:11]` | 0, 0, 0 |
| `gripper_close[11:12]` | base_mode, a constant |

So the model is trained to emit arm motion on the base-motion channel and a constant on the
gripper channel. At eval the robot drives its base away from the counter and never closes the
gripper. Measured: 0/160 on `pnpcountertocab_mimicgen8_exact160` for a checkpoint whose base
scored 15/160 on the same episodes.

## How the check works

No reference dataset needed. In this task the mobile base never moves, so the two layouts are
distinguishable by which dimensions are constant:

```
LeRobot     dims 0-3 all zero, dim 4 constant, dim 11 two-valued (the gripper)
robosuite   dims 0-5 continuous, dim 6 two-valued (the gripper), dims 7-10 zero, dim 11 constant
```

## Use

```bash
python check_action_layout.py --dataset <gr00t_views dataset> [--episodes 20]
```

Exit codes: `0` PASS (LeRobot), `1` FAIL (robosuite order), `2` not 12-d, `3` neither matched.

A `3` on a small sample can just mean the sampled episodes are degenerate -- re-run with a
larger `--episodes` before concluding anything.

```
$ python check_action_layout.py --dataset .../mimicgen_natural_256
  dim 6 gripper   -1.000 .. 1.000  uniq 2  <- two-valued
  dim11 base_mode -1.000 .. -1.000 uniq 1  <- constant
FAIL  action column is in RoboCasa/robosuite order but modality.json declares LeRobot order.
      Repair: actions = actions[:, [7, 8, 9, 10, 11, 0, 1, 2, 3, 4, 5, 6]]

$ python check_action_layout.py --dataset .../pickplace_target_human/PickPlaceCounterToCabinet
PASS  action column is in LeRobot order, matching modality.json.
```

## Fix

In the builder, reorder before writing the parquet:

```python
# robosuite [eef_pos3, eef_rot3, gripper, base3, torso, base_mode]
#   -> LeRobot [base3+torso, base_mode, eef_pos3, eef_rot3, gripper]
actions = all_actions[:, [7, 8, 9, 10, 11, 0, 1, 2, 3, 4, 5, 6]]
```

Rebuild the dataset rather than patching parquet in place, unless you have confirmed nothing
else already consumed it.

## Which builders are affected

| builder | source of `action` | affected |
|---|---|---|
| `baseline/mimicgen/gr00t_build/render_to_gr00t.py` (line 180) | `demo_grp["actions"]` from the MimicGen HDF5, simulator order | **yes** |
| `train_robocasa/scripts/dataset_build/build_vace_objwise_gr00t_dataset.py` | `pd.read_parquet(src_parquet)` from the original dataset, already LeRobot order | no |

The rule of thumb: a builder that **re-derives** actions from a simulator rollout needs the
reorder; one that **copies** rows from an existing LeRobot dataset does not.

Note that `observation.state` is not affected in either builder -- `extract_obs_state()`
assembles it field by field (`base_pos(3) + base_quat(4) + eef_pos(3) + eef_quat(4) +
gripper_qpos(2)`) in the declared order, so only `action` was ever passed through raw.

## Datasets checked

| dataset | result |
|---|---|
| `baseline/mimicgen/gr00t_views/mimicgen_natural_256` | FAIL |
| `baseline/mimicgen/gr00t_views/mimicgen_per_target_32_256eps` | FAIL (same builder) |
| `robocasa_full/pickplace_target_human/PickPlaceCounterToCabinet` | PASS |
| any VACE gr00t_views set | expected PASS -- run the checker to confirm on that machine |