Dataset definition: pusht-simulator-wall05-contact2

The specification every published pusht-simulator dataset follows, fixed here before the first 100 episodes of the pymunk-wall dataset were generated. It is the LeWM-compatible layout the two reference datasets already use (scripts/publish_hf.py), plus sharding for size and the provenance fields the friction study introduced. A dataset that deviates from this file says so in its README.

Identity

field value
repo desertmouse/pusht-simulator-wall05-contact2 (moved to the team org when it exists)
engine pymunk-wall: the published gym-pusht environment (pixels, kinematic PD pusher k_p=100, k_v=20, mass 1, moment 3000, damping=0, frictionless walls) plus Coulomb friction 0.5 between the pusher and the T (contact_friction, set on both shapes; pymunk's pair rule is the geometric mean)
brain Brain-contact2 + trim: rule planner with the friction-study corrections (ContactLoopPlanner, planner: contact2) and the executor's near-goal trim release. Deterministic; no model call; model: null in every plan record
horizon 200 control steps at 10 Hz; episodes are not cut short on success (the executor backs off and holds)
goal the published fixed goal: block centre (256, 256), angle π/4; the pusher's goal position follows the goal convention below
seeds drawn once, deterministically: random.Random("pusht-wall05-contact2").sample(range(100_000, 1_000_000), 3000); the first 100 are the sample, positions 100–2999 the extension. Disjoint from every eval and study batch (those are < 100 000)
target size 100 (sample, reviewed) → 3000

Per episode (meta/episodes.parquet, one row)

The full episode-store record (33 columns: seed, situation labels, start and final errors, coverage, first-success and solved step, bouts, speeds, phase shares, planner, git commit, render style, ...) plus:

Per step (data/<config>/frames-NNNNN.parquet, one row per control step)

Row k is the observation before action k (video frame k; frame 0 is the reset state). reward_*, next.*, terminated_*, n_contacts describe the state after action k.

column dtype / shape meaning
episode_index, frame_index, timestamp, index int, int, float (s), int (global, gapless)
pixels PNG bytes 512×512×3 in native512, 224×224×3 (cv2.INTER_AREA, LeWM's size) in lewm224; decoded from the episode mp4
state float[7] [agent_x, agent_y, block_x, block_y, angle % 2π, agent_vx, agent_vy]
proprio float[4] agent position and velocity
observation.state float[5] LeRobot's 5-D state (raw angle), kept for LeRobot readers
pos_agent, vel_agent, block_pose float[2], float[2], float[3] LeWM's info keys
action float[2] relative, LeWM's clip((target − agent) / 100, −1, 1)
action_abs float[2] the absolute pusher target in px the brain actually issued
reward_lewm float −‖goal_state − state_after‖ over 7 dims
reward_coverage float clip(coverage / 0.95, 0, 1), LeRobot style
next.coverage, next.gap_px float after the step
terminated_lewm bool ‖goal[:4] − state[:4]‖ < 20 px and wrapped angle diff < π/9
terminated_ours bool coverage ≥ 0.95 and pusher gap ≥ 10 px (the workbench's success)
truncated bool last step
n_contacts int geometric proxy: 1 if post-step gap ≤ 1 px (the logs do not hold gym-pusht's per-substep collision count)
phase, plan_index, contact, gap_px str, int, bool, float executor phase, which plan was active, in contact, pusher-to-outline gap

Velocities are exact, not estimated. The pusher is kinematic under the published PD law, so agent_vx, agent_vy come from replaying that law from the recorded initial position and actions; the export asserts the replayed positions match the logged ones to < 1e-3 px on every step of every episode and aborts otherwise.

Sharding. 100 episodes (20 000 rows) per parquet file, written as episodes are processed; nothing is held in memory across shards. HF's config globs data/native512/*.parquet and data/lewm224/*.parquet read all shards as one table.

Per plan (meta/plans.parquet)

Every plan the brain made: episode_index, plan_index, the plan (edge, lever_px, orbit, release_alignment_deg, intent, reasoning - which for contact2 includes the online gain and trim notes), fallback, latency_s, and outcome.* (angle and position error before/after, coverage before/after, bout frames, touched, contact turn/travel, why it released).

Narratives (meta/narratives.parquet)

The generated auto narrative per episode (situation, plan sequence, outcome, what the last bout did), built from measured facts only. No model narratives.

Videos (videos/episode_NNNNNN.mp4)

512×512, 10 fps, 201 frames (reset + 200 steps), H.264 yuv420p, rendered in pygame-parity style; no target cross (the commanded target is in action_abs).

meta/info.json

codebase, source_commit, tag, backend, planner, brain ("Brain-contact2 + trim"), contact_friction, fps, total_episodes, total_frames, solved_episodes, seed_rule (the line above), features (every column with dtype/shape/units/notes), episode_features, configs, goals, goal_convention, success_criterion, lewm_success_criterion, velocity_source, pd_law, render_note (OpenCV frames match the published pygame frames on 99.7 % of pixels, not bit-identical), frame_convention, shards.

Acceptance before extending 100 → 3000

Run scripts/contact_study.py-style checks on the 100 and compare with the 50-seed eval of the same engine + brain: solved rate, coverage distribution, bouts per episode, situation mix, and release-reason mix within the ranges seen there; the PD-replay assertion passes; every episode has 201 video frames, a goal image at both sizes, ≥ 1 plan, a narrative; no NaN; indices gapless. The sample is uploaded to HF before review so nothing waits on it.