|
Download tools/ACTION_LAYOUT.md from mlnha/pose6daug-scripts: direct link, hf CLI and curl.
- Browser
- Download file 4.11 kB
-
https://huggingface.co/mlnha/pose6daug-scripts/resolve/main/tools/ACTION_LAYOUT.md
- Command line
-
hf download hf://mlnha/pose6daug-scripts/tools/ACTION_LAYOUT.md
-
curl -L -o ACTION_LAYOUT.md https://huggingface.co/mlnha/pose6daug-scripts/resolve/main/tools/ACTION_LAYOUT.md
4.11 kB
| # Action-layout check for gr00t_views datasets | |
| Found 2026-09-20 while investigating a 0/160 eval. Use it on any gr00t_views dataset before | |
| training on it. | |
| ## The bug | |
| RoboCasa's simulator emits a 12-d action as | |
| ``` | |
| [eef_pos(3), eef_rot(3), gripper(1), base(3), torso(1), base_mode(1)] robosuite order | |
| ``` | |
| A gr00t_views dataset's `meta/modality.json` declares | |
| ``` | |
| base_motion[0:4] control_mode[4:5] eef_pos[5:8] eef_rot[8:11] gripper_close[11:12] | |
| ``` | |
| which is a **different order**. A builder that copies simulator actions straight into the | |
| parquet writes correct numbers under the wrong column names. Nothing errors, the loss | |
| converges to a small value, and the resulting policy is useless: the slices land as | |
| | modality.json reads | what is actually there | | |
| |---|---| | |
| | `base_motion[0:4]` | eef x, y, z and rot x | | |
| | `control_mode[4:5]` | rot y | | |
| | `eef_pos[5:8]` | rot z, **gripper**, 0 | | |
| | `eef_rot[8:11]` | 0, 0, 0 | | |
| | `gripper_close[11:12]` | base_mode, a constant | | |
| So the model is trained to emit arm motion on the base-motion channel and a constant on the | |
| gripper channel. At eval the robot drives its base away from the counter and never closes the | |
| gripper. Measured: 0/160 on `pnpcountertocab_mimicgen8_exact160` for a checkpoint whose base | |
| scored 15/160 on the same episodes. | |
| ## How the check works | |
| No reference dataset needed. In this task the mobile base never moves, so the two layouts are | |
| distinguishable by which dimensions are constant: | |
| ``` | |
| LeRobot dims 0-3 all zero, dim 4 constant, dim 11 two-valued (the gripper) | |
| robosuite dims 0-5 continuous, dim 6 two-valued (the gripper), dims 7-10 zero, dim 11 constant | |
| ``` | |
| ## Use | |
| ```bash | |
| python check_action_layout.py --dataset <gr00t_views dataset> [--episodes 20] | |
| ``` | |
| Exit codes: `0` PASS (LeRobot), `1` FAIL (robosuite order), `2` not 12-d, `3` neither matched. | |
| A `3` on a small sample can just mean the sampled episodes are degenerate -- re-run with a | |
| larger `--episodes` before concluding anything. | |
| ``` | |
| $ python check_action_layout.py --dataset .../mimicgen_natural_256 | |
| dim 6 gripper -1.000 .. 1.000 uniq 2 <- two-valued | |
| dim11 base_mode -1.000 .. -1.000 uniq 1 <- constant | |
| FAIL action column is in RoboCasa/robosuite order but modality.json declares LeRobot order. | |
| Repair: actions = actions[:, [7, 8, 9, 10, 11, 0, 1, 2, 3, 4, 5, 6]] | |
| $ python check_action_layout.py --dataset .../pickplace_target_human/PickPlaceCounterToCabinet | |
| PASS action column is in LeRobot order, matching modality.json. | |
| ``` | |
| ## Fix | |
| In the builder, reorder before writing the parquet: | |
| ```python | |
| # robosuite [eef_pos3, eef_rot3, gripper, base3, torso, base_mode] | |
| # -> LeRobot [base3+torso, base_mode, eef_pos3, eef_rot3, gripper] | |
| actions = all_actions[:, [7, 8, 9, 10, 11, 0, 1, 2, 3, 4, 5, 6]] | |
| ``` | |
| Rebuild the dataset rather than patching parquet in place, unless you have confirmed nothing | |
| else already consumed it. | |
| ## Which builders are affected | |
| | builder | source of `action` | affected | | |
| |---|---|---| | |
| | `baseline/mimicgen/gr00t_build/render_to_gr00t.py` (line 180) | `demo_grp["actions"]` from the MimicGen HDF5, simulator order | **yes** | | |
| | `train_robocasa/scripts/dataset_build/build_vace_objwise_gr00t_dataset.py` | `pd.read_parquet(src_parquet)` from the original dataset, already LeRobot order | no | | |
| The rule of thumb: a builder that **re-derives** actions from a simulator rollout needs the | |
| reorder; one that **copies** rows from an existing LeRobot dataset does not. | |
| Note that `observation.state` is not affected in either builder -- `extract_obs_state()` | |
| assembles it field by field (`base_pos(3) + base_quat(4) + eef_pos(3) + eef_quat(4) + | |
| gripper_qpos(2)`) in the declared order, so only `action` was ever passed through raw. | |
| ## Datasets checked | |
| | dataset | result | | |
| |---|---| | |
| | `baseline/mimicgen/gr00t_views/mimicgen_natural_256` | FAIL | | |
| | `baseline/mimicgen/gr00t_views/mimicgen_per_target_32_256eps` | FAIL (same builder) | | |
| | `robocasa_full/pickplace_target_human/PickPlaceCounterToCabinet` | PASS | | |
| | any VACE gr00t_views set | expected PASS -- run the checker to confirm on that machine | | |