wall05-contact2: distribution of the first 1000, the gap check, and the last 1000The plan was: generate the first 1000 episodes, measure how their starts
are distributed, find any under-represented region, and if one exists,
choose the last 1000 seeds to fill it. This report records what was found
at each step. Conclusion up front: the first 1000 are a uniform draw to
within sampling noise on every axis tested, so no correction was applied;
the last 1000 were run on the seeds as originally drawn, and they match
the first 1000. Full per-bin tables: ds_distribution_at_1000.md.
The seeds are random.Random("pusht-wall05-contact2").sample(range(100_000,
1_000_000), 3000), and each seed produces a gym-pusht reset: block
position integer-uniform on [100, 400)², block angle uniform on [−π, π),
pusher uniform on [50, 450)². The goal is fixed. So the target
distribution is not a guess - it is exactly computable, and that is what
the first 1000 were tested against:
| quantity | analytic expectation |
|---|---|
| signed angle error (goal − block) | uniform, 1/24 per 15° bin |
| goal direction in the block frame | uniform, 1/24 per 15° bin (independent of position) |
| angle band | aligned 5.6 %, mixed 44.4 %, flipped 38.9 %, inverted 11.1 % |
| position band (CoG to goal CoG) | near 5.6 %, moderate 29.3 %, far 65.1 % |
| near a wall (< 115 px) | 24.8 % |
Tests: chi-square goodness-of-fit per histogram, plus a one-sided binomial test per bin for under-representation, Bonferroni-corrected across the bins of each histogram.
| histogram | bins | chi-square p | any bin under-represented? |
|---|---|---|---|
| signed angle error | 24 | 0.20 | no (smallest one-sided p 0.003 vs threshold 0.002) |
| |angle error| | 12 | 0.30 | no |
| position error | 10 | 0.77 | no |
| goal direction in block frame | 24 | 0.78 | no |
| wall distance | 8 | 0.78 | no |
| angle band | 4 | 0.16 | no |
| position band | 3 | 0.73 | no |
| near-wall share | 2 | 0.48 | no |
| angle band × position band | 12 | 0.55 | no |
| situation class (16 labels) | 16 | 0.37 | no |
Observed vs expected on the bands: aligned 4.8 % (5.6), mixed 46.9 % (44.4), flipped 36.2 % (38.9), inverted 12.1 % (11.1); near a wall 23.8 % (24.8). The one bin that looks low - signed angle error in [120°, 135°), 25 observed vs 41.7 expected, one-sided p = 0.003 - is inside the Bonferroni threshold for 24 bins (0.0021) and has no neighbour bin low with it; it is what one expects to see once in ~116 bins.
Data integrity on the same 1000: every episode has 200 logged steps, at least one plan, a non-empty video, finite record fields, a seed from the planned list, and no seed repeats.

None beyond sampling noise. Every histogram is consistent with the analytic reset distribution; no band, cell or situation label is under-represented after correction. The sampler does what it says.
The correction step therefore had nothing to correct. It was prepared
anyway: a rebalancing routine that would draw candidate seeds from a
disjoint range (1 000 000-10 000 000), instantiate only the start of
each (make_env("pymunk-wall", seed).get_state(), no rollout), classify
it, and pick 1000 that bring the 3000 closest to the target. It was not
invoked, because invoking it on a distribution already at target would
have replaced a uniform draw with a hand-shaped one and made the dataset
less faithful to the published reset, not more.
Run on seeds 2000-2999 of the original list. Compared with the first 1000:
| first 1000 | middle 1000 | last 1000 | |
|---|---|---|---|
| near a wall | 23.8 % | 25.5 % | 22.9 % |
| |angle error| uniform, chi-square p | 0.28 | 0.73 | 0.75 |
| solved | 8.9 % | 11.9 % | 10.9 % |
| final coverage mean | 0.627 | 0.621 | 0.632 |
Two-sample KS, first 1000 vs last 1000: start angle error D = 0.030, p = 0.76; start position error D = 0.030, p = 0.76; final coverage D = 0.049, p = 0.18. The thirds are draws from the same distribution.
Acceptance check against the 50-seed eval of the same engine and brain: solved 10.6 % (eval 12.0 %), coverage mean 0.627 (0.584), median 0.678 (0.623), bouts 3.69 (3.72), plans 4.29 (4.34); KS on final coverage p = 0.25; situation mix mixed 45.7 / flipped 37.7 / inverted 11.6 / aligned 4.6 / already-nearly-solved 0.3 %; 0 artifact problems.
What the dataset is therefore good for, and what it is not: it is an unbiased sample of the published reset distribution on the friction engine with the rule brain - the right thing for training or for measuring a policy against the same starts the published benchmark uses. It is not balanced by difficulty: 11.6 % of starts are inverted and the brain solves none of them, so a model trained on it sees ~350 inverted failures and few inverted successes. A difficulty-balanced companion set would be a different sampler, stated as such, not a correction of this one.