YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
actaug GR00T β self-contained train + eval bundle
Everything needed to train GR00T-N1.5 on the action-augmentation RoboCasa datasets
(actaug-prio-1024 full + its 512 subset) from the 60k base checkpoint, and to
evaluate the resulting checkpoints on the 160-episode exact-replay benchmark.
This repo is self-contained: it ships the training code (with our save/stop knobs), the eval stack (with the porting fixes), the base checkpoint, the subset builder, and this protocol doc. Large external artifacts (the training datasets, the eval replay bank) are on HF and fetched by the setup steps below.
What we're doing (and why)
Goal. Measure how much the action-augmentation data helps GR00T on RoboCasa PickPlaceCounterToCabinet, and how that depends on dataset size and training length.
- We start from the 60k base GR00T checkpoint (already trained on RoboCasa) and fine-tune
it on the action-augmented data (
actaug-prio-1024). - We run two data conditions in parallel to isolate the effect of dataset size:
- full = all 1024 augmented episodes,
- 512 = a 512-episode subset (first 64 of each of the 8 objects).
- We keep the learning-rate schedule identical to a full 30k-step run (warmup β1500, cosine over 30k) but stop early at step 10k. This way every checkpoint's LR is what it would be in the real 30k run β so checkpoints at 4k/5k/β¦/10k are directly comparable to that schedule instead of being from a schedule that was artificially compressed into 10k.
- We save checkpoints along the way (every 500 steps from 4k to 10k; the 512 run also saves 2k) so we can see the learning curve β success rate as a function of training step β not just a single final number.
- We evaluate every saved checkpoint on a fixed 160-episode exact-replay benchmark and compare the two conditions' curves against the base-60k baseline (11/160).
Why a queue? Training emits checkpoints over time (4k appears first, then 4.5k, 5k, β¦). Each eval takes ~2β3 h on one GPU. If we waited for training to finish and then evaluated 13β14 checkpoints Γ 2 conditions serially, it would take days. Instead we run a small pool of eval workers that start evaluating each checkpoint the moment it lands, in parallel with training and with each other β so evals finish almost as fast as the checkpoints arrive.
The FCFS eval queue β logic
A pool of 2β3 eval workers (each = 1 GPU running the 4-server/4-client 160-ep eval of one checkpoint at a time) shares one work queue across both training runs:
- Discover. Scan both runs' output dirs for
checkpoint-N/that are complete (havemodel.safetensors.index.json+trainer_state.jsonβ i.e. fully flushed, not mid-write). - Filter. Drop checkpoints already evaluated or currently claimed by another worker.
- Order (priority, not strict arrival): evaluate integer-thousand checkpoints first (4k, 5k, 6k, β¦, 10k), then the half-steps (4.5k, 5.5k, β¦). Within each tier, earliest step first. Rationale: the round-thousand points are the ones we report the curve on, so we want them ASAP.
- Claim atomically. A worker grabs the next checkpoint with an atomic marker (e.g.
mkdira<ckpt>.claim) so two workers never eval the same one. - Evaluate. Run
eval_one_checkpoint_local.shon it β writessummary.json(all-160 success rate + per-episode stage flags) into that checkpoint's eval dir. - Mark done, loop. If no checkpoint is pending, sleep and poll; a worker exits only when both training runs have finished and the queue is empty. It waits for arriving checkpoints β it does not need training to be done to start.
"Checkpoint-wise FCFS, not run-wise" means the two training runs' checkpoints are interleaved into one queue and taken as they become ready β a worker takes whichever eligible checkpoint is next by the ordering above, regardless of which run produced it.
0. TL;DR
# 1. env: mygr00t (train + serve), robocasa (eval client) β see Β§1
# 2. base checkpoint: base60k/ (shipped here) == mlnha/gr00t-n15-robocasa-base60k
# 3. datasets: Ronaldo-GOAT/actaug-prio-1024 (full 1024); build the 512 subset with scripts/build_512_subset.py
# 4. TRAIN (2 GPU, premium, 30k-schedule but STOP at 10k, save 4k->10k every 500):
sbatch -p sjw_alinlab_premium --wckey=project-short-name:others --export=ALL scripts/train_actaug.sbatch
# 5. EVAL each checkpoint (1 GPU, 4 workers, exact160 bank, POLICY_SEED=12345):
MODEL_PATH=<ckpt> ROOT=<out> bash code/eval/eval_one_checkpoint_local.sh
1. Environments
Two conda envs (GR00T serving and the RoboCasa sim client have incompatible deps):
mygr00tβ trains and serves the GR00T policy.PYTHONPATH=code/myGR00T. Has torch+CUDA+gr00t.robocasaβ runs the eval client (drives the RoboCasa sim). Hasrobocasa==1.0.0 + a vendoredrobosuitewithload_model_on_init. The eval client also importsgr00t.eval.wrappers, socode/myGR00Tmust be on itsPYTHONPATH(the runner sets this).
RoboCasa assets must be writable (robosuite writes a temp processed XML into each object's asset
dir). If the assets tree is a read-only mount, make a writable symlink farm
(cp -as <ro-assets>/. <writable>/) and point robocasa's models/assets symlink at it.
2. Base checkpoint
base60k/ == mlnha/gr00t-n15-robocasa-base60k (GR00T-N1.5, RoboCasa PnPCounterToCab, global
step 60000). All fine-tunes start from here (--base-model-path base60k).
3. Datasets
- Full:
Ronaldo-GOAT/actaug-prio-1024(HF dataset) β 1024 eps = 128 per object Γ 8 GR00T objects (donut_5, Jar023, MeasuringCup009, SoapDispenser010, steak_8, SyrupBottle005, teapot_6, teapot_7), 3 views (256Γ256), 20 fps.meta/modality.jsonis already correct forsingle_panda_gripper(video keysleft_view/right_view/wrist_viewβobservation.images.robot0_agentview_{left,right}/robot0_eye_in_hand; annotationhuman.action.task_description;rotation_type: quaternionon base_rotation + eef_rotation_relative). No modality fix needed. - 512 subset: built locally from the full set's
index/actaug_prio_512/episode_index.jsonl(first 64/object, "sample A";actaug_prio_512bis the disjoint last-64/object "sample B"):
This makes a lightweight VIEW (symlinks data/+videos/, subset meta/episodes.jsonl) β no parquet rewrite.python scripts/build_512_subset.py --full <actaug_prio_1024> --out <actaug_prio_512> --index-name actaug_prio_512
4. Training protocol
Trainer: code/myGR00T/scripts/gr00t_finetune.py (env mygr00t, PYTHONPATH=code/myGR00T, cwd there).
Config = the "30k config, terminate at 10k":
| flag | value | why |
|---|---|---|
--data-config |
single_panda_gripper |
matches base-60k embodiment |
--embodiment-tag |
new_embodiment |
|
--backbone-model-type / --backbone-select-layer |
eagle / 12 |
|
--base-model-path |
base60k |
warm start |
--num-gpus / --batch-size |
2 / 32 |
global batch 64 (32Γ2, DDP). VRAM is set by per-device batch (bs32 fits comfortably; literal bs64 on 1 GPU β 70 GB, bs16 β 26 GB). |
--max-steps |
30000 |
schedule horizon β warmup=0.05Γ30000β1500, cosine over 30k, so at step N the LR matches a full 30k run (NOT a run that decays to ~0 by N). |
--stop-at |
10000 |
terminate here while keeping the 30k schedule (custom SaveAtStepsCallback). |
--save-at |
4000,4500,β¦,10000 (full); 2000,4000,4500,β¦,10000 (512) |
explicit save steps; periodic saving disabled. |
Two knobs we added to the stock trainer (--save-at, --stop-at) via SaveAtStepsCallback
β saves exactly at listed steps and stops at stop_at without shortening the LR schedule.
Submit (per Β§0). The sbatch (scripts/train_actaug.sbatch) waits for the dataset to be fully present,
then trains; on preemption it auto-detects the newest checkpoint and passes --resume (the
save-at checkpoints are full β optimizer+scheduler+trainer_state β so the cosine continues). Use
sjw_alinlab_premium (non-preemptible) to avoid requeue churn.
Cluster submit-filter rules (this cluster): pass --wckey=project-short-name:others; do NOT pass
--cpus-per-task; export MODEL_OUTPUT_DIR starting with /rlwrld-unified-checkpoints/jonghoon/;
comma-valued env vars (like --save-at) must go through --export=ALL on an exported shell var, NOT
inline in --export=ALL,SAVE_AT=... (commas are the --export delimiter β truncation).
5. Eval protocol
Exact-state replay on PickPlaceCounterToCabinet, 160 episodes = 8 objects Γ 20, restoring
the exact recorded MuJoCo scene per episode. Bank:
Ronaldo-GOAT/transfer :: actaug/eval/table2_exact160/bank_pnpcountertocab_mimicgen8_exact160_replay_policyseed12345_20260506
(its 8 objects are exactly this dataset's training objects).
Per checkpoint (1 GPU, 4 servers + 4 clients, ~2β3 h):
MODEL_PATH=<step_dir> ROOT=<eval_out> CUDA_DEV=0 \
bash code/eval/eval_one_checkpoint_local.sh
Fixed protocol: POLICY_SEED=12345 (server RNG + per-step policy_seed+ep_idx*1000003+step),
SEED_BASE=42, ACTION_HORIZON=16, N_EPISODES=160. 4 workers = length of GPUS list
(GPUS=0,0,0,0 packs 4 on one GPU). Writes summary.json + per-episode stage flags
(grasped/grasped_strict/lifted/in_cab). Baseline: base-60k = 11/160 (6.9%) on this set.
FCFS eval queue (evaluate checkpoints as training produces them, concurrently): run 2β3 of the
above as independent single-GPU jobs, each claiming the next un-evaluated checkpoint, integer-k
first (4k,5k,6k,β¦ before 4.5k,5.5k,β¦). POLICY_SEED is seed-sensitive β don't over-read small
single-seed differences.
Eval porting fixes (why the eval code has local edits vs the 9/22 verified version)
The verified eval was captured on the original NVIDIA cluster; the replay bank bakes absolute
/lp-dev/... asset paths. To run elsewhere we:
- repoint the bank's
model_xml_gz/state_npz/ep_meta_pickleto the local bank, - call robocasa's
edit_model_xml()on the loaded scene XML (repaths recorded assets to the local install β mirrors robocasa's ownplayback_dataset.reset_to), - gate
--generative_textures(GENERATIVE_TEXTURES=0) β it only affects the thrown-away first reset; the real scene comes from the recorded XML, so results are unaffected. These are porting-only; the replayed scene and policy I/O are unchanged. GPU architecture is a controlled variable β compare checkpoints only on the same GPU model.
6. What we are running on this server (2026-09-23)
- Train full-1024 and 512 from base-60k, 30k-schedule, stop at 10k, save 4kβ10k every 500
(512 also saves 2k), 2 GPU each,
sjw_alinlab_premium. - Eval every saved checkpoint on the exact160 bank, FCFS, integer-k first, 2β3 concurrent 1-GPU workers.
7. Contents
README.md this protocol
base60k/ the 60k base checkpoint (== mlnha/gr00t-n15-robocasa-base60k)
code/myGR00T/ training + serving code (adds --save-at / --stop-at)
code/eval/ eval stack (exact-replay runner + client, with porting fixes)
eval_groot15_exact_replay.sh multi-worker eval launcher (server + clients)
eval_robocasa_replay_state_grasp.py robocasa replay client (+ stage flags)
eval_one_checkpoint_local.sh local wrapper: 1 GPU / 4 workers, all env overrides set
EVALUATION.md upstream eval notes
scripts/train_actaug.sbatch 2-GPU training job (dataset guard + resume-on-preempt)
scripts/build_512_subset.py build the 512 view from the full dataset's index list