Dataset wall05-contact2: distribution of the first 1000, the gap check, and the last 1000

The plan was: generate the first 1000 episodes, measure how their starts are distributed, find any under-represented region, and if one exists, choose the last 1000 seeds to fill it. This report records what was found at each step. Conclusion up front: the first 1000 are a uniform draw to within sampling noise on every axis tested, so no correction was applied; the last 1000 were run on the seeds as originally drawn, and they match the first 1000. Full per-bin tables: ds_distribution_at_1000.md.

1. What "correctly distributed" means here

The seeds are random.Random("pusht-wall05-contact2").sample(range(100_000, 1_000_000), 3000), and each seed produces a gym-pusht reset: block position integer-uniform on [100, 400)², block angle uniform on [−π, π), pusher uniform on [50, 450)². The goal is fixed. So the target distribution is not a guess - it is exactly computable, and that is what the first 1000 were tested against:

quantity analytic expectation
signed angle error (goal − block) uniform, 1/24 per 15° bin
goal direction in the block frame uniform, 1/24 per 15° bin (independent of position)
angle band aligned 5.6 %, mixed 44.4 %, flipped 38.9 %, inverted 11.1 %
position band (CoG to goal CoG) near 5.6 %, moderate 29.3 %, far 65.1 %
near a wall (< 115 px) 24.8 %

Tests: chi-square goodness-of-fit per histogram, plus a one-sided binomial test per bin for under-representation, Bonferroni-corrected across the bins of each histogram.

2. The first 1000

histogram bins chi-square p any bin under-represented?
signed angle error 24 0.20 no (smallest one-sided p 0.003 vs threshold 0.002)
|angle error| 12 0.30 no
position error 10 0.77 no
goal direction in block frame 24 0.78 no
wall distance 8 0.78 no
angle band 4 0.16 no
position band 3 0.73 no
near-wall share 2 0.48 no
angle band × position band 12 0.55 no
situation class (16 labels) 16 0.37 no

Observed vs expected on the bands: aligned 4.8 % (5.6), mixed 46.9 % (44.4), flipped 36.2 % (38.9), inverted 12.1 % (11.1); near a wall 23.8 % (24.8). The one bin that looks low - signed angle error in [120°, 135°), 25 observed vs 41.7 expected, one-sided p = 0.003 - is inside the Bonferroni threshold for 24 bins (0.0021) and has no neighbour bin low with it; it is what one expects to see once in ~116 bins.

Data integrity on the same 1000: every episode has 200 logged steps, at least one plan, a non-empty video, finite record fields, a seed from the planned list, and no seed repeats.

3. The gap

None beyond sampling noise. Every histogram is consistent with the analytic reset distribution; no band, cell or situation label is under-represented after correction. The sampler does what it says.

The correction step therefore had nothing to correct. It was prepared anyway: a rebalancing routine that would draw candidate seeds from a disjoint range (1 000 000-10 000 000), instantiate only the start of each (make_env("pymunk-wall", seed).get_state(), no rollout), classify it, and pick 1000 that bring the 3000 closest to the target. It was not invoked, because invoking it on a distribution already at target would have replaced a uniform draw with a hand-shaped one and made the dataset less faithful to the published reset, not more.

4. The last 1000

Run on seeds 2000-2999 of the original list. Compared with the first 1000:

first 1000 middle 1000 last 1000
near a wall 23.8 % 25.5 % 22.9 %
|angle error| uniform, chi-square p 0.28 0.73 0.75
solved 8.9 % 11.9 % 10.9 %
final coverage mean 0.627 0.621 0.632

Two-sample KS, first 1000 vs last 1000: start angle error D = 0.030, p = 0.76; start position error D = 0.030, p = 0.76; final coverage D = 0.049, p = 0.18. The thirds are draws from the same distribution.

5. The whole 3000

Acceptance check against the 50-seed eval of the same engine and brain: solved 10.6 % (eval 12.0 %), coverage mean 0.627 (0.584), median 0.678 (0.623), bouts 3.69 (3.72), plans 4.29 (4.34); KS on final coverage p = 0.25; situation mix mixed 45.7 / flipped 37.7 / inverted 11.6 / aligned 4.6 / already-nearly-solved 0.3 %; 0 artifact problems.

What the dataset is therefore good for, and what it is not: it is an unbiased sample of the published reset distribution on the friction engine with the rule brain - the right thing for training or for measuring a policy against the same starts the published benchmark uses. It is not balanced by difficulty: 11.6 % of starts are inverted and the brain solves none of them, so a model trained on it sees ~350 inverted failures and few inverted successes. A difficulty-balanced companion set would be a different sampler, stated as such, not a correction of this one.