mp_yam_code / doc /generation.md
yqi19's picture
docs split into doc/; scale-up generator; per-task frames and episode videos
000b008 verified
|
Raw
History Blame Contribute Delete
2.67 kB
# Scaling up to a dataset
`scripts/generate.py` runs many randomized episodes per task and keeps the ones that succeeded.
Each episode gets a fresh seed, so the scene randomization differs every time; the scripted
solver plays it and the task's own success conditions decide whether it counts. That filtering
is what makes this a demonstration generator rather than a batch runner.
```bash
python scripts/generate.py --episodes 20 # every registered task
python scripts/generate.py --tasks passing --episodes 100 # only tasks known to work
python scripts/generate.py --tasks grape_box,stack_blocks --episodes 50
python scripts/generate.py --episodes 20 --resume # skip episodes already done
```
## What it writes
```
data/
<task>/ep<seed>.mp4 video β€” successes only, unless --keep-failures
<task>/ep<seed>.json result, per-condition PASS/FAIL, and the episode's actual placements
manifest.json running totals and success rate per task
```
The per-episode JSON records the **randomized placements the environment actually used**, not the
nominal ones, so an episode can be replayed exactly with `--seed`.
## Options worth knowing
| flag | why |
|---|---|
| `--tasks passing` | reads `outputs/tasks/v3_*.log` and selects only tasks whose last development run succeeded β€” spending 100 seeds on a task that cannot complete is wasted GPU time |
| `--start-seed` | defaults to 100, deliberately clear of the 0–9 seeds used during development, so generated data never overlaps the episodes used to tune the task |
| `--resume` | skips episodes whose JSON already exists; the script is safe to interrupt |
| `--keep-failures` | keeps failed videos for debugging instead of deleting them |
## Cost
One episode is one simulator launch β€” expect roughly 40–90 s each, so this is long-running by
nature. It prints one line per episode and writes `manifest.json` after every episode, so
progress survives an interruption.
Success rates vary by task and are worth checking before committing to a large run: a task that
passes at seed 0 may be sensitive to the jitter. The manifest gives you exactly that number.
## A caution
The filter is only as honest as the success conditions. Three checks in this suite were caught
reporting success for episodes that did nothing (see [`../DIAGNOSTICS.md`](../DIAGNOSTICS.md)) β€”
a pour whose grasp had failed, and two joints that "moved" under gravity alone. Before generating
at scale on a task you have just written, render a few contact sheets and confirm the numeric
check agrees with what you see; otherwise you will mass-produce convincing-looking failures.