# actaug — perturbed-object action re-earning for robocasa episodes Standalone handoff package. **No imports from any other project directory** — only installed libraries (robocasa / robosuite / robomimic / numpy / pandas / zmq / torch / imageio). Reference implementations it was distilled from (read-only): `/home/nvidia/jonghoon/idm_eval/scripts/gripper_state_replay_sandbox/replay/` (`gripper_state_replay.py` hybrid mode, `grip_perturb.py`). Physical location `/lp-dev/jonghoon/actaug` (symlinked at `/home/nvidia/jonghoon/actaug`). ## What it does For an arbitrary episode of a robocasa LeRobot export (per-episode `extras/episode_XXXXXX/{model.xml.gz, states.npz, ep_meta.json}` + 12-d actions in `data/chunk-XXX/episode_XXXXXX.parquet`): 1. **Reset** the MuJoCo sim to the recorded initial state — the episode's own compiled `model.xml` + flattened state via `EnvRobosuite.reset_to`, so the scene / layout / cameras are bit-identical to the source episode. 2. **Perturb** the manipulated object (`obj_joint0` free joint): yaw about the WORLD vertical + xy translation, then a short physics settle. Random draws: yaw ~ U[-rot_deg, rot_deg], translation direction uniform on the circle, radius ~ U[trans_min_m, trans_m]. 3. **Replay** the recorded actions until `grasp_t - pre_buffer` (`grasp_t` = first recorded gripper-close command). 4. **Policy handoff**: GR00T-N1.5 grasps (chunks of `exec_h=16` env steps, stop when the object is lifted > 3 cm or budget runs out), then the recorded transport actions resume at `grasp_t + resume_offset`. 5. **Sample N times** per perturbation. Each attempt is fully seeded: the obs dict carries `_policy_action_seed` (honored by myGR00T `policy.get_action`), so attempts are diverse AND reproducible. 6. **Filter on task success** (`env.is_success()["task"]`). Successful attempts are dumped under `accepted/` with executed `actions.npy` (parquet order), `states.npz` (16-d observation.state layout), and mp4s from the SAME three cameras as the dataset (`robot0_agentview_left/right`, `robot0_eye_in_hand`, 256×256) + a side-by-side `3cam.mp4` + `meta.json`. ## Two grasp-re-earning algorithms (pick via YAML `mode:`) **`hybrid` (unconstrained)** — GT actions to `grasp_t − pre_buffer`, then the policy commands the FULL action space (EEF pose + gripper through the OSC controller) until the object lifts, then recorded transport resumes. Handles translated objects (the policy can chase them). `configs/unconstrained.yaml`. **`constrained`** — the sim_action_aug algorithm (`sim_action_aug/code/rollout/r02_rollout.py`, `ws/ROLLOUT_REPORT.md` §3) ported to robocasa: every arm joint NOT listed in `constrained.policy_joints` (+ mobile base + torso) is PINNED to the recorded MuJoCo state path every physics substep (zero drift, the grip_state/grip_policy machinery). The policy owns ONLY the listed joints (default `[6, 7]`, i.e. wrist pitch + roll, like the sim_action_aug pipeline) and the gripper: - **each joint j in `policy_joints`**: `recorded + servo` — the policy's delta rotation projected onto joint j's world axis, capped `step_cap` rad/step and `servo_caps[j]` rad total (defaults: 6→0.30, 7→0.60); - **joint 7** (when listed) additionally: `recorded + dj7_init + servo` — `dj7_init` is a searched grid (`constrained.dj7_init_deg`, slewed in over ≤48 frames instead of teleported, per r02 phase R), with the (−π/2, π/2] wrap (2-jaw symmetry) and `|dj7_init + servo| ≤ j7_total_cap`; - **the gripper**: `action.gripper_close` in one of three interpretations — `bit` (binarized, robocasa-native), `threshold` (latch full close once ≥ `grip_tau`; this is what unlocked ep 0009 in sim_action_aug), `continuous` (ctrl interpolated open→close), `recorded` (diagnostic: the recorded close bit, = grip_state). Transport = relative replay: everything pinned, dj7 frozen at hand-off, gripper held closed until the recorded release frame, then the recorded opening schedule (never re-closes), plus 30 settle frames. **rel-6D weld (on by default, `constrained.rel6d`)** — sim_action_aug's simfix, required here too: a pinned (infinite-impedance) arm carry slips out of the friction grasp (the old robocasa pin-sweep peaked at ~33% success; reproduced on ep0: the pinch engages, then the bottle squeezes out during the lift, with recorded gripper timing and no policy at all). So an INACTIVE `` is injected into the episode XML; it engages at the CURRENT relative pose once both jaw pads touch the object while it is lifted ≥ `weld_lift_m` under a close command, and it drops the moment the gripper is commanded open — the release is a genuine physical fall. Engage/release frames are logged per attempt (`weld_engage_t`, `weld_release_t`). Tuning matters: with a 15 mm engage gate + soft solref 0.02 the bottle sagged ~19 mm in the jaw, caught the cabinet-shelf edge on the way in, and fell out at release; `weld_lift_m: 0.004` + `weld_solref: "0.005 1"` (the defaults) fixed it — recorded-gripper baseline then reproduces the GT success exactly. `configs/constrained.yaml`. Caveat: the pinned arm cannot chase a TRANSLATED object — use it for yaw-dominant perturbations / small translations, and `hybrid` for the 3–7 cm displacement regime. Note: attempts count = `n_perturbs × len(dj7_init_deg) × n_samples` for constrained, `n_perturbs × n_samples` for the others. The wrist servo uses the policy's raw delta-rotation command (controller units) — direction is what matters; the per-step/total caps bound the magnitude exactly as in r02. ## YAML reference (every key) A YAML passed as `--config` sets argparse defaults; any explicit CLI flag overrides it. Top-level keys mirror `rollout.py` flags; the `constrained:` block is passed to `run_constrained_attempt` verbatim. ### Top-level keys (both algorithms) | key | default | meaning | |---|---|---| | `mode` | hybrid | `gt` (pure replay sanity) / `hybrid` / `policy` (end-to-end) / `constrained` | | `dataset` | PickPlaceCounterToCabinet export | robocasa LeRobot export root | | `episode` | — (required) | episode index | | `perturb` | "" | explicit `"dyaw_deg,dx_m,dy_m"`; overrides random draws | | `n_perturbs` | 0 | number of RANDOM perturbations (0 + no `perturb` = nominal) | | `rot_deg` / `rot_min_deg` | 60 / 0 | yaw magnitude ~ U[rot_min_deg, rot_deg], sign random | | `trans_m` / `trans_min_m` | 0.05 / 0 | translation radius ~ U[trans_min_m, trans_m], direction uniform | | `seed` | 0 | seeds BOTH the perturbation draws and the per-attempt policy seeds | | `n_samples` | 1 | **N**: policy samples per initial condition (rejection sampling) | | `pre_buffer` | 10 | **the buffer**: policy inference starts at `grasp_t − pre_buffer` frames (`grasp_t` = first recorded gripper-close). hybrid-mode flag; constrained mode reads it from its own block (below) | | `policy_budget` | 200 | hybrid only: max policy env-steps to achieve the lift | | `resume_offset` | 0 | hybrid only: recorded transport resumes at `grasp_t + resume_offset` | | `exec_h` | 16 | hybrid only: env steps executed per policy chunk (inference stride) | | `host` / `port` | 127.0.0.1 / 8801 | GR00T server address | | `save` | success | `success` (dump accepted only) / `all` (+ rejected/) / `none` | | `fps` | 20 | output video fps (= dataset rate; don't change unless re-timing) | | `out` | outputs/ep\_\ | output dir | ### `constrained:` block | key | default | meaning | |---|---|---| | `pre_buffer` (alias `start_off`) | 4 | **the buffer**: inference starts at `grasp_t − pre_buffer`; the dj7_init ramp is scheduled BEFORE this point so it never eats into the policy window | | `k_post` | 12 | policy keeps control until `grasp_t + k_post`, then hand-off | | `replan` | 8 | **inference stride**: control steps between policy queries; each query returns a 16-step chunk, of which the first `replan` steps are executed | | `policy_joints` | [6, 7] | **the joint constraint**: 1-indexed arm joints the policy may servo; everything else (joints, base, torso) stays pinned to the recording. Gripper is ALWAYS the policy's unless `gripper_mode: recorded` | | `step_cap` | 0.05 | **inference speed**: max servo change per 20 Hz control step (rad) for every policy joint; also paces the dj7_init ramp. 0.05 ≈ 1 rad/s (rate-matched to sim_action_aug) | | `servo_caps` | {6: 0.30, 7: 0.60} | total authority per joint (rad) — how FAR, not how fast | | `servo_cap_default` | 0.30 | cap for listed joints missing from `servo_caps` | | `j7_total_cap` | π | joint 7 only: bound on `dj7_init + servo` | | `dj7_init_deg` | [0] | searched initial wrist-roll offsets (deg); grid × `n_samples` = attempts per perturbation. Use `[0, 90, -90]` for 90°+ object rotations | | `gripper_mode` | threshold | `bit` / `threshold` / `continuous` / `recorded` (see algorithm section) | | `grip_tau` | 0.35 | threshold mode: latch full close once policy's close fraction ≥ τ (use ~0.25 for the base model) | | `embodiment` | robocasa | `robocasa` (finetuned head) / `oxe_droid` (BASE model via DROID head) | | `oxe_grip_flip` | false | oxe_droid: invert gripper polarity if the base model's convention is reversed | | `rel6d` | true | the grasp weld (mandatory for reliable carry — see weld section) | | `weld_lift_m` | 0.004 | weld engages once both pads touch + object lifted this much + close commanded | | `weld_solref` | "0.005 1" | weld stiffness injected into the XML | | `weld_open_frac` / `weld_open_floor` | 0.5 / 0.10 | weld releases when the close command drops below `max(frac·held, floor)` | | `settle_frames` | 30 | pinned hold after the last frame so the final pose is a rest pose | | `qvel_mode` | finitediff | qvel written for pinned joints each substep (`zero` = no injected velocity) | ## Code documentation (what each file does) ### `actaug_core.py` — the shared library (no CLI) Everything both algorithms need, importable, side-effect free: - **Episode I/O**: `load_episode(root, idx)` → dict with the compiled `model.xml`, the full flattened MuJoCo `states` (T×(1+nq+nv)), `ep_meta`, the 12-d parquet `actions`, the language `instruction` (tasks.jsonl), and `env_args`. `grasp_frame(actions)` = first recorded gripper-close. - **Env**: `make_env(env_args)` builds the robomimic `EnvRobosuite` (cached per env name — construction is minutes); `reset_episode(env, ep)` = `reset_to` with the episode's own XML + state 0. - **Perturbation**: `apply_object_perturbation(env, dyaw, dx, dy)` rotates the object's free joint about the world vertical and translates it in-sim; `sample_perturbation(rng, ...)` draws random ones; `settle(env)` lets physics rest after. - **Policy bridges**: `groot_obs` / `groot_obs_oxe` build the obs dict (3 camera renders + proprio + instruction + `_policy_action_seed`) for the finetuned robocasa head / the base model's DROID head; `policy_chunk_to_env` / `oxe_chunk_to_env` decode a returned action chunk into robosuite env order; `to_env_action` / `env_to_parquet_action` convert between parquet layout and env layout. - **PolicyClient**: minimal ZMQ REQ client for the myGR00T inference server (torch.save wire format). - **Predicates & output**: `obj_z`, `is_success`, `render3`, `groot_state_vec` (16-d observation.state), `write_mp4`. ### `rollout.py` — the CLI driver Parses CLI + `--config` YAML, loads the episode, builds the env, draws the perturbations, then loops `perturbation × dj7_grid × n_samples`: - `run_attempt(...)` implements `gt` / `policy` / `hybrid` (GT approach via `env.step` → `run_policy` chunks until lift → GT transport), recording every executed step; - for `mode: constrained` it injects the weld into the episode XML (`inject_weld`), resets + perturbs, and delegates to `constrained.run_constrained_attempt`; - `save_dump(...)` writes each kept attempt: `actions.npy` (parquet order), `states.npz`, `left/right/wrist.mp4` + `3cam.mp4`, `obj_track.npz` (executed-vs-recorded object positions, constrained only), `meta.json`; `results.csv` is re-flushed after every attempt. ### `constrained.py` — the constrained algorithm - `inject_weld(xml, solref)`: adds the inactive `` equality to the episode XML (pure text transform, idempotent). - `PinRig(env)`: all model addressing for one compiled scene — arm/base/torso joint qpos/dof addresses, finger actuators (+ open/close ctrl targets), object vs gripper geom sets, left/right pad geom sets, joint world axes, weld eq id. Methods: `pin_step` (one control frame: lerp the pinned joints across physics substeps while gripper + object integrate), `grip_ctrl`, `grip_contact`, `pad_contacts`, `weld_engage/release/relpose/active`. - `run_constrained_attempt(env, ep, client, dj7_init_deg, seed, cfg, record)`: the four phases — pinned approach (with the dj7_init ramp), policy window (per-joint servo + gripper, chunk every `replan` steps, per-call action seed), relative-replay transport (frozen offsets, recorded opening), settle — plus the weld engage/release bookkeeping and the result row. ### `compose_2x3.py` — comparison videos `compose_2x3.py `: row 1 = the ORIGINAL episode's three dataset videos, row 2 = the augmented rollout's same three views; the shorter row freezes on its last frame. ### Shell scripts - `launch_server.sh [gpu] [port] [model] [embodiment] [data_config]` — detached GR00T inference server (defaults = the finetuned 60k checkpoint; pass the HF snapshot + `oxe_droid oxe_droid` for the base model). - `run_sanity.sh [gpu] [port] [ep]` — GT replay sanity → hybrid nominal → 3 random perturbations × 3 samples. ### `configs/` `unconstrained.yaml` (hybrid preset), `constrained.yaml` (fully commented constrained preset), `asym_base_j6.yaml` / `asym_bigrot.yaml` (the base-model asymmetric-object example runs, standard and 90–150° rotation). ## Interpreters / environment - sim client: `/lp-dev/jonghoon/mimicgen_augment/envs/mimicgen/bin/python` with `MUJOCO_GL=egl MUJOCO_EGL_DEVICE_ID=` (osmesa is broken on this box). - policy server: `/data/nvidia/gripper_augmentator/conda-envs/mygr00t/bin/python` running `myGR00T/scripts/inference_service.py` (see `launch_server.sh`; `HF_HUB_OFFLINE=1` is mandatory). Default checkpoint: `/lp-dev/jonghoon/myGR00T_outputs/pnpcountertocab_all502_gbs64_wandb_60k_save5k_20260427_193901/checkpoint-60000` (GR00T-N1.5 finetuned on all 502 PnPCounterToCabinet episodes — the same dataset these rollouts perturb). - **BASE (non-finetuned) GR00T-N1.5**: the base checkpoint has NO robocasa embodiment head (only gr1 / oxe_droid / agibot_genie1), so it is served via its **oxe_droid (DROID) head** with a best-effort obs/action mapping (`groot_obs_oxe` / `oxe_chunk_to_env` in actaug_core, same approach as the old gripper_state_replay code): `bash launch_server.sh 8802 oxe_droid oxe_droid` and set `constrained.embodiment: oxe_droid` (+ optionally lower `grip_tau` to ~0.25 — the base policy's close commands are weaker). Expect lower sample efficiency; that's what `n_samples` is for. - default dataset: `/lp-dev/jonghoon/robocasa_full/pickplace_target_human/PickPlaceCounterToCabinet` (502 eps, PandaOmron, HYBRID_MOBILE_BASE, fps 20). ## Usage ```bash # one-time server (GPU 5, port 8801) bash launch_server.sh 5 8801 SIMPY=/lp-dev/jonghoon/mimicgen_augment/envs/mimicgen/bin/python export MUJOCO_GL=egl MUJOCO_EGL_DEVICE_ID=2 # sanity: unperturbed GT replay must succeed $SIMPY rollout.py --episode 0 --mode gt # the pipeline: 5 random perturbations x 8 policy samples, keep successes $SIMPY rollout.py --episode 0 --mode hybrid \ --n_perturbs 5 --n_samples 8 --rot_deg 60 --trans_m 0.05 \ --pre_buffer 10 --policy_budget 200 --port 8801 --seed 7 # explicit perturbation (yaw +45deg, +3cm x, -2cm y) $SIMPY rollout.py --episode 0 --mode hybrid --perturb "45,0.03,-0.02" \ --n_samples 8 --port 8801 # constrained algorithm, all knobs from YAML (CLI flags override YAML values) $SIMPY rollout.py --config configs/constrained.yaml --episode 0 # or everything at once bash run_sanity.sh 2 8801 0 ``` Outputs: `outputs//results.csv` (one row per attempt: perturbation, seed, peak lift, grasped, success, wall time) + `accepted/pertXX_sYY/` dumps. `--save all` also dumps failures under `rejected/`. ## Controlling how FAST the policy moves (inference speed) The policy-driven motion speed is set entirely by per-control-step caps in the `constrained:` block — the sim always runs at the dataset's 20 Hz, so these caps ARE the deg/s of the wrist during the inference window: | key | meaning | default | rad/s @20 Hz | |---|---|---|---| | `step_cap` | max servo change per control step, every policy joint; also paces the `dj7_init` ramp | 0.05 | 1.0 (≈57°/s) | | `servo_caps: {j: cap}` | TOTAL authority per joint (how far, not how fast) | 6→0.30, 7→0.60 | — | | `replan` | control steps between policy queries (reaction latency, not speed) | 8 | — | History: the port initially reused r02's 0.10 rad/step, but r02 ran at 9.2 Hz — at robocasa's 20 Hz that is ~2 rad/s and the wrist visibly whips. 0.05 is rate-matched to the original. Raise/lower `step_cap` in the YAML to make the policy's wrist motion faster/slower; the gripper close speed is the actuator's own dynamics (same as GT) and is not affected. For `hybrid` mode the analogous knob is `exec_h` (env steps executed per chunk) — the policy's actions are executed at the recorded control rate either way. ## Knobs that matter - `--pre_buffer` — how early the policy takes over. 10 works for ≤5 cm translations; increase (or use `--mode policy`) for larger displacements, since the recorded approach aims at the ORIGINAL object location. - `--policy_budget` — max policy env-steps to achieve the lift (default 200). - `--resume_offset` — where recorded transport resumes. 0 = at the recorded grasp frame. The transport is delta-eef, so it carries the object from wherever the policy lifted it. - Perturbation ranges: prior sweeps (`grip_perturb.py`) used yaw up to ±180°, translations ≤ 7 cm, dz ≤ 2 cm. Large yaws are fine for rotationally symmetric objects; translations > ~7 cm start leaving the reachable counter region. ## Gotchas - Env construction takes ~2–4 min (kitchen scene compile) — it is cached per process; batch many attempts/episodes per process launch. - `EnvRobosuite.reset_to` internally remaps the export's absolute asset paths (`/root/robocasa/...`, `/opt/conda/...`) to the local robocasa/robosuite installs — only the mimicgen env's robocasa checkout has all assets; the `robocasa_calib` env historically had a geom mismatch. Use the mimicgen python. - GPU 4 is often another user's — check `nvidia-smi` before picking GPUs. - `obj_pos` observable + `obj_joint0` joint exist in all robocasa kitchen envs; success predicate is the env's own `is_success()["task"]`.