Download actaug/code/README.md from Ronaldo-GOAT/transfer: direct link, hf CLI and curl.
- Browser
- Download file 19.2 kB
-
https://huggingface.co/Ronaldo-GOAT/transfer/resolve/main/actaug/code/README.md
- Command line
-
hf download hf://Ronaldo-GOAT/transfer/actaug/code/README.md
-
curl -L -o README.md https://huggingface.co/Ronaldo-GOAT/transfer/resolve/main/actaug/code/README.md
actaug β perturbed-object action re-earning for robocasa episodes
Standalone handoff package. No imports from any other project directory β only
installed libraries (robocasa / robosuite / robomimic / numpy / pandas / zmq /
torch / imageio). Reference implementations it was distilled from (read-only):
/home/nvidia/jonghoon/idm_eval/scripts/gripper_state_replay_sandbox/replay/
(gripper_state_replay.py hybrid mode, grip_perturb.py).
Physical location /lp-dev/jonghoon/actaug (symlinked at
/home/nvidia/jonghoon/actaug).
What it does
For an arbitrary episode of a robocasa LeRobot export (per-episode
extras/episode_XXXXXX/{model.xml.gz, states.npz, ep_meta.json} + 12-d actions in
data/chunk-XXX/episode_XXXXXX.parquet):
- Reset the MuJoCo sim to the recorded initial state β the episode's own
compiled
model.xml+ flattened state viaEnvRobosuite.reset_to, so the scene / layout / cameras are bit-identical to the source episode. - Perturb the manipulated object (
obj_joint0free joint): yaw about the WORLD vertical + xy translation, then a short physics settle. Random draws: yaw ~ U[-rot_deg, rot_deg], translation direction uniform on the circle, radius ~ U[trans_min_m, trans_m]. - Replay the recorded actions until
grasp_t - pre_buffer(grasp_t= first recorded gripper-close command). - Policy handoff: GR00T-N1.5 grasps (chunks of
exec_h=16env steps, stop when the object is lifted > 3 cm or budget runs out), then the recorded transport actions resume atgrasp_t + resume_offset. - Sample N times per perturbation. Each attempt is fully seeded: the obs
dict carries
_policy_action_seed(honored by myGR00Tpolicy.get_action), so attempts are diverse AND reproducible. - Filter on task success (
env.is_success()["task"]). Successful attempts are dumped underaccepted/with executedactions.npy(parquet order),states.npz(16-d observation.state layout), and mp4s from the SAME three cameras as the dataset (robot0_agentview_left/right,robot0_eye_in_hand, 256Γ256) + a side-by-side3cam.mp4+meta.json.
Two grasp-re-earning algorithms (pick via YAML mode:)
hybrid (unconstrained) β GT actions to grasp_t β pre_buffer, then the
policy commands the FULL action space (EEF pose + gripper through the OSC
controller) until the object lifts, then recorded transport resumes. Handles
translated objects (the policy can chase them). configs/unconstrained.yaml.
constrained β the sim_action_aug algorithm
(sim_action_aug/code/rollout/r02_rollout.py, ws/ROLLOUT_REPORT.md Β§3)
ported to robocasa: every arm joint NOT listed in constrained.policy_joints
(+ mobile base + torso) is PINNED to the recorded MuJoCo state path every
physics substep (zero drift, the grip_state/grip_policy machinery). The policy
owns ONLY the listed joints (default [6, 7], i.e. wrist pitch + roll, like
the sim_action_aug pipeline) and the gripper:
- each joint j in
policy_joints:recorded + servoβ the policy's delta rotation projected onto joint j's world axis, cappedstep_caprad/step andservo_caps[j]rad total (defaults: 6β0.30, 7β0.60); - joint 7 (when listed) additionally:
recorded + dj7_init + servoβdj7_initis a searched grid (constrained.dj7_init_deg, slewed in over β€48 frames instead of teleported, per r02 phase R), with the (βΟ/2, Ο/2] wrap (2-jaw symmetry) and|dj7_init + servo| β€ j7_total_cap; - the gripper:
action.gripper_closein one of three interpretations βbit(binarized, robocasa-native),threshold(latch full close once β₯grip_tau; this is what unlocked ep 0009 in sim_action_aug),continuous(ctrl interpolated openβclose),recorded(diagnostic: the recorded close bit, = grip_state).
Transport = relative replay: everything pinned, dj7 frozen at hand-off, gripper held closed until the recorded release frame, then the recorded opening schedule (never re-closes), plus 30 settle frames.
rel-6D weld (on by default, constrained.rel6d) β sim_action_aug's simfix,
required here too: a pinned (infinite-impedance) arm carry slips out of the
friction grasp (the old robocasa pin-sweep peaked at ~33% success; reproduced
on ep0: the pinch engages, then the bottle squeezes out during the lift, with
recorded gripper timing and no policy at all). So an INACTIVE
<weld body1="gripper0_right_eef" body2="<obj root>"/> is injected into the
episode XML; it engages at the CURRENT relative pose once both jaw pads touch
the object while it is lifted β₯ weld_lift_m under a close command, and it
drops the moment the gripper is commanded open β the release is a genuine
physical fall. Engage/release frames are logged per attempt (weld_engage_t,
weld_release_t). Tuning matters: with a 15 mm engage gate + soft solref 0.02
the bottle sagged ~19 mm in the jaw, caught the cabinet-shelf edge on the way
in, and fell out at release; weld_lift_m: 0.004 + weld_solref: "0.005 1"
(the defaults) fixed it β recorded-gripper baseline then reproduces the GT
success exactly.
configs/constrained.yaml. Caveat: the pinned arm cannot chase a TRANSLATED
object β use it for yaw-dominant perturbations / small translations, and
hybrid for the 3β7 cm displacement regime.
Note: attempts count = n_perturbs Γ len(dj7_init_deg) Γ n_samples for
constrained, n_perturbs Γ n_samples for the others. The wrist servo uses the
policy's raw delta-rotation command (controller units) β direction is what
matters; the per-step/total caps bound the magnitude exactly as in r02.
YAML reference (every key)
A YAML passed as --config sets argparse defaults; any explicit CLI flag
overrides it. Top-level keys mirror rollout.py flags; the constrained:
block is passed to run_constrained_attempt verbatim.
Top-level keys (both algorithms)
| key | default | meaning |
|---|---|---|
mode |
hybrid | gt (pure replay sanity) / hybrid / policy (end-to-end) / constrained |
dataset |
PickPlaceCounterToCabinet export | robocasa LeRobot export root |
episode |
β (required) | episode index |
perturb |
"" | explicit "dyaw_deg,dx_m,dy_m"; overrides random draws |
n_perturbs |
0 | number of RANDOM perturbations (0 + no perturb = nominal) |
rot_deg / rot_min_deg |
60 / 0 | yaw magnitude ~ U[rot_min_deg, rot_deg], sign random |
trans_m / trans_min_m |
0.05 / 0 | translation radius ~ U[trans_min_m, trans_m], direction uniform |
seed |
0 | seeds BOTH the perturbation draws and the per-attempt policy seeds |
n_samples |
1 | N: policy samples per initial condition (rejection sampling) |
pre_buffer |
10 | the buffer: policy inference starts at grasp_t β pre_buffer frames (grasp_t = first recorded gripper-close). hybrid-mode flag; constrained mode reads it from its own block (below) |
policy_budget |
200 | hybrid only: max policy env-steps to achieve the lift |
resume_offset |
0 | hybrid only: recorded transport resumes at grasp_t + resume_offset |
exec_h |
16 | hybrid only: env steps executed per policy chunk (inference stride) |
host / port |
127.0.0.1 / 8801 | GR00T server address |
save |
success | success (dump accepted only) / all (+ rejected/) / none |
fps |
20 | output video fps (= dataset rate; don't change unless re-timing) |
out |
outputs/ep<N>_<mode> | output dir |
constrained: block
| key | default | meaning |
|---|---|---|
pre_buffer (alias start_off) |
4 | the buffer: inference starts at grasp_t β pre_buffer; the dj7_init ramp is scheduled BEFORE this point so it never eats into the policy window |
k_post |
12 | policy keeps control until grasp_t + k_post, then hand-off |
replan |
8 | inference stride: control steps between policy queries; each query returns a 16-step chunk, of which the first replan steps are executed |
policy_joints |
[6, 7] | the joint constraint: 1-indexed arm joints the policy may servo; everything else (joints, base, torso) stays pinned to the recording. Gripper is ALWAYS the policy's unless gripper_mode: recorded |
step_cap |
0.05 | inference speed: max servo change per 20 Hz control step (rad) for every policy joint; also paces the dj7_init ramp. 0.05 β 1 rad/s (rate-matched to sim_action_aug) |
servo_caps |
{6: 0.30, 7: 0.60} | total authority per joint (rad) β how FAR, not how fast |
servo_cap_default |
0.30 | cap for listed joints missing from servo_caps |
j7_total_cap |
Ο | joint 7 only: bound on dj7_init + servo |
dj7_init_deg |
[0] | searched initial wrist-roll offsets (deg); grid Γ n_samples = attempts per perturbation. Use [0, 90, -90] for 90Β°+ object rotations |
gripper_mode |
threshold | bit / threshold / continuous / recorded (see algorithm section) |
grip_tau |
0.35 | threshold mode: latch full close once policy's close fraction β₯ Ο (use ~0.25 for the base model) |
embodiment |
robocasa | robocasa (finetuned head) / oxe_droid (BASE model via DROID head) |
oxe_grip_flip |
false | oxe_droid: invert gripper polarity if the base model's convention is reversed |
rel6d |
true | the grasp weld (mandatory for reliable carry β see weld section) |
weld_lift_m |
0.004 | weld engages once both pads touch + object lifted this much + close commanded |
weld_solref |
"0.005 1" | weld stiffness injected into the XML |
weld_open_frac / weld_open_floor |
0.5 / 0.10 | weld releases when the close command drops below max(fracΒ·held, floor) |
settle_frames |
30 | pinned hold after the last frame so the final pose is a rest pose |
qvel_mode |
finitediff | qvel written for pinned joints each substep (zero = no injected velocity) |
Code documentation (what each file does)
actaug_core.py β the shared library (no CLI)
Everything both algorithms need, importable, side-effect free:
- Episode I/O:
load_episode(root, idx)β dict with the compiledmodel.xml, the full flattened MuJoCostates(TΓ(1+nq+nv)),ep_meta, the 12-d parquetactions, the languageinstruction(tasks.jsonl), andenv_args.grasp_frame(actions)= first recorded gripper-close. - Env:
make_env(env_args)builds the robomimicEnvRobosuite(cached per env name β construction is minutes);reset_episode(env, ep)=reset_towith the episode's own XML + state 0. - Perturbation:
apply_object_perturbation(env, dyaw, dx, dy)rotates the object's free joint about the world vertical and translates it in-sim;sample_perturbation(rng, ...)draws random ones;settle(env)lets physics rest after. - Policy bridges:
groot_obs/groot_obs_oxebuild the obs dict (3 camera renders + proprio + instruction +_policy_action_seed) for the finetuned robocasa head / the base model's DROID head;policy_chunk_to_env/oxe_chunk_to_envdecode a returned action chunk into robosuite env order;to_env_action/env_to_parquet_actionconvert between parquet layout and env layout. - PolicyClient: minimal ZMQ REQ client for the myGR00T inference server (torch.save wire format).
- Predicates & output:
obj_z,is_success,render3,groot_state_vec(16-d observation.state),write_mp4.
rollout.py β the CLI driver
Parses CLI + --config YAML, loads the episode, builds the env, draws the
perturbations, then loops perturbation Γ dj7_grid Γ n_samples:
run_attempt(...)implementsgt/policy/hybrid(GT approach viaenv.stepβrun_policychunks until lift β GT transport), recording every executed step;- for
mode: constrainedit injects the weld into the episode XML (inject_weld), resets + perturbs, and delegates toconstrained.run_constrained_attempt; save_dump(...)writes each kept attempt:actions.npy(parquet order),states.npz,left/right/wrist.mp4+3cam.mp4,obj_track.npz(executed-vs-recorded object positions, constrained only),meta.json;results.csvis re-flushed after every attempt.
constrained.py β the constrained algorithm
inject_weld(xml, solref): adds the inactive<weld body1="gripper0_right_eef" body2="<object root>"/>equality to the episode XML (pure text transform, idempotent).PinRig(env): all model addressing for one compiled scene β arm/base/torso joint qpos/dof addresses, finger actuators (+ open/close ctrl targets), object vs gripper geom sets, left/right pad geom sets, joint world axes, weld eq id. Methods:pin_step(one control frame: lerp the pinned joints across physics substeps while gripper + object integrate),grip_ctrl,grip_contact,pad_contacts,weld_engage/release/relpose/active.run_constrained_attempt(env, ep, client, dj7_init_deg, seed, cfg, record): the four phases β pinned approach (with the dj7_init ramp), policy window (per-joint servo + gripper, chunk everyreplansteps, per-call action seed), relative-replay transport (frozen offsets, recorded opening), settle β plus the weld engage/release bookkeeping and the result row.
compose_2x3.py β comparison videos
compose_2x3.py <episode> <dump_dir> <out.mp4>: row 1 = the ORIGINAL
episode's three dataset videos, row 2 = the augmented rollout's same three
views; the shorter row freezes on its last frame.
Shell scripts
launch_server.sh [gpu] [port] [model] [embodiment] [data_config]β detached GR00T inference server (defaults = the finetuned 60k checkpoint; pass the HF snapshot +oxe_droid oxe_droidfor the base model).run_sanity.sh [gpu] [port] [ep]β GT replay sanity β hybrid nominal β 3 random perturbations Γ 3 samples.
configs/
unconstrained.yaml (hybrid preset), constrained.yaml (fully commented
constrained preset), asym_base_j6.yaml / asym_bigrot.yaml (the base-model
asymmetric-object example runs, standard and 90β150Β° rotation).
Interpreters / environment
- sim client:
/lp-dev/jonghoon/mimicgen_augment/envs/mimicgen/bin/pythonwithMUJOCO_GL=egl MUJOCO_EGL_DEVICE_ID=<gpu>(osmesa is broken on this box). - policy server:
/data/nvidia/gripper_augmentator/conda-envs/mygr00t/bin/pythonrunningmyGR00T/scripts/inference_service.py(seelaunch_server.sh;HF_HUB_OFFLINE=1is mandatory). Default checkpoint:/lp-dev/jonghoon/myGR00T_outputs/pnpcountertocab_all502_gbs64_wandb_60k_save5k_20260427_193901/checkpoint-60000(GR00T-N1.5 finetuned on all 502 PnPCounterToCabinet episodes β the same dataset these rollouts perturb). - BASE (non-finetuned) GR00T-N1.5: the base checkpoint has NO robocasa
embodiment head (only gr1 / oxe_droid / agibot_genie1), so it is served via
its oxe_droid (DROID) head with a best-effort obs/action mapping
(
groot_obs_oxe/oxe_chunk_to_envin actaug_core, same approach as the old gripper_state_replay code):bash launch_server.sh <gpu> 8802 <hf-snapshot-dir> oxe_droid oxe_droidand setconstrained.embodiment: oxe_droid(+ optionally lowergrip_tauto ~0.25 β the base policy's close commands are weaker). Expect lower sample efficiency; that's whatn_samplesis for. - default dataset:
/lp-dev/jonghoon/robocasa_full/pickplace_target_human/PickPlaceCounterToCabinet(502 eps, PandaOmron, HYBRID_MOBILE_BASE, fps 20).
Usage
# one-time server (GPU 5, port 8801)
bash launch_server.sh 5 8801
SIMPY=/lp-dev/jonghoon/mimicgen_augment/envs/mimicgen/bin/python
export MUJOCO_GL=egl MUJOCO_EGL_DEVICE_ID=2
# sanity: unperturbed GT replay must succeed
$SIMPY rollout.py --episode 0 --mode gt
# the pipeline: 5 random perturbations x 8 policy samples, keep successes
$SIMPY rollout.py --episode 0 --mode hybrid \
--n_perturbs 5 --n_samples 8 --rot_deg 60 --trans_m 0.05 \
--pre_buffer 10 --policy_budget 200 --port 8801 --seed 7
# explicit perturbation (yaw +45deg, +3cm x, -2cm y)
$SIMPY rollout.py --episode 0 --mode hybrid --perturb "45,0.03,-0.02" \
--n_samples 8 --port 8801
# constrained algorithm, all knobs from YAML (CLI flags override YAML values)
$SIMPY rollout.py --config configs/constrained.yaml --episode 0
# or everything at once
bash run_sanity.sh 2 8801 0
Outputs: outputs/<run>/results.csv (one row per attempt: perturbation, seed,
peak lift, grasped, success, wall time) + accepted/pertXX_sYY/ dumps.
--save all also dumps failures under rejected/.
Controlling how FAST the policy moves (inference speed)
The policy-driven motion speed is set entirely by per-control-step caps in the
constrained: block β the sim always runs at the dataset's 20 Hz, so these
caps ARE the deg/s of the wrist during the inference window:
| key | meaning | default | rad/s @20 Hz |
|---|---|---|---|
step_cap |
max servo change per control step, every policy joint; also paces the dj7_init ramp |
0.05 | 1.0 (β57Β°/s) |
servo_caps: {j: cap} |
TOTAL authority per joint (how far, not how fast) | 6β0.30, 7β0.60 | β |
replan |
control steps between policy queries (reaction latency, not speed) | 8 | β |
History: the port initially reused r02's 0.10 rad/step, but r02 ran at 9.2 Hz β
at robocasa's 20 Hz that is ~2 rad/s and the wrist visibly whips. 0.05 is
rate-matched to the original. Raise/lower step_cap in the YAML to make the
policy's wrist motion faster/slower; the gripper close speed is the actuator's
own dynamics (same as GT) and is not affected.
For hybrid mode the analogous knob is exec_h (env steps executed per chunk)
β the policy's actions are executed at the recorded control rate either way.
Knobs that matter
--pre_bufferβ how early the policy takes over. 10 works for β€5 cm translations; increase (or use--mode policy) for larger displacements, since the recorded approach aims at the ORIGINAL object location.--policy_budgetβ max policy env-steps to achieve the lift (default 200).--resume_offsetβ where recorded transport resumes. 0 = at the recorded grasp frame. The transport is delta-eef, so it carries the object from wherever the policy lifted it.- Perturbation ranges: prior sweeps (
grip_perturb.py) used yaw up to Β±180Β°, translations β€ 7 cm, dz β€ 2 cm. Large yaws are fine for rotationally symmetric objects; translations > ~7 cm start leaving the reachable counter region.
Gotchas
- Env construction takes ~2β4 min (kitchen scene compile) β it is cached per process; batch many attempts/episodes per process launch.
EnvRobosuite.reset_tointernally remaps the export's absolute asset paths (/root/robocasa/...,/opt/conda/...) to the local robocasa/robosuite installs β only the mimicgen env's robocasa checkout has all assets; therobocasa_calibenv historically had a geom mismatch. Use the mimicgen python.- GPU 4 is often another user's β check
nvidia-smibefore picking GPUs. obj_posobservable +obj_joint0joint exist in all robocasa kitchen envs; success predicate is the env's ownis_success()["task"].