pusht-simulator / docs /html /contact_friction_study.html
desertmouse's picture
pages
473cc8a verified
Raw History Blame Contribute Delete
18.9 kB
<!doctype html><html><head><meta charset='utf-8'><title>Contact-friction study: what μ = 0.5 between pusher and T does to the brain</title><style>
body{max-width:980px;margin:2rem auto;padding:0 1.25rem;font:15px/1.55 -apple-system,Segoe UI,Roboto,sans-serif;color:#1c2330;background:#fafbfc}
h1{font-size:1.7rem;border-bottom:2px solid #d8dee8;padding-bottom:.4rem}h2{margin-top:2.2rem;border-bottom:1px solid #e3e8ef;padding-bottom:.25rem}
table{border-collapse:collapse;margin:1rem 0;font-size:13.5px;font-variant-numeric:tabular-nums}th,td{border:1px solid #d8dee8;padding:.35rem .6rem;text-align:left}
th{background:#eef2f7}td:not(:first-child){text-align:right}code{background:#eef2f7;padding:.1rem .3rem;border-radius:3px;font-size:.92em}
pre{background:#1c2330;color:#e6edf5;padding:.8rem 1rem;border-radius:6px;overflow:auto}pre code{background:none;color:inherit}
img{max-width:100%;border:1px solid #d8dee8;border-radius:4px;margin:.5rem 0}blockquote{border-left:3px solid #6aa9ff;margin:0;padding:.2rem 1rem;color:#4b5563}
nav{font-size:13px;margin-bottom:1.5rem}nav a{margin-right:1rem}
</style></head><body><nav><a href="index.html">index</a> <a href="brain3_vs_3a_events.html">brain3 vs 3a events</a> <a href="contact_friction_study.html">contact friction study</a> <a href="convergence_study.html">convergence study</a> <a href="dataset_definition.html">dataset definition</a> <a href="ds08_distribution_at_1000.html">ds08 distribution at 1000</a> <a href="ds_distribution_at_1000.html">ds distribution at 1000</a> <a href="ds_distribution_report.html">ds distribution report</a> <a href="ds_mu03_brain3b_distribution.html">ds mu03 brain3b distribution</a> <a href="episode_store.html">episode store</a> <a href="human_pushing_analysis.html">human pushing analysis</a> <a href="lewm_parity.html">lewm parity</a> <a href="reference_parity.html">reference parity</a> <a href="session_fidelity_and_behaviour.html">session fidelity and behaviour</a></nav><h1 id="contact-friction-study-what-05-between-pusher-and-t-does-to-the-brain">Contact-friction study: what μ = 0.5 between pusher and T does to the brain</h1>
<p>All runs in this study are on <strong>pymunk</strong>. Nothing here uses SuperDex. The
two engines compared are:</p>
<ul>
<li><strong>B</strong> <code>pymunk-reference</code> - the published LeWM/gym-pusht environment, bit-exact, frictionless (<code>Shape.friction</code> never assigned, so 0).</li>
<li><strong>A</strong> <code>pymunk-wall</code> - the same environment with one change: <code>contact_friction = 0.5</code> on the pusher circle and the two T polygons (<code>src/pusht_sim/envs/pymunk_wall_env.py</code>). Walls, <code>damping = 0</code>, mass, moment, seeds - unchanged; μ = 0 reproduces B bit for bit (<code>tests/test_env_pymunk_wall.py</code>).</li>
</ul>
<p>Same 25 seeds on both (the <code>batch_pymunk_ref</code> eval seeds, also published as
<code>desertmouse/pusht-simulator-pymunk-reference</code>): <code>44177 18033 67870 16370
57580 75204 98135 87949 81400 31775 75263 64785 52463 11980 68814 35669 2103
39679 2228 45997 32063 81358 79092 89650 91922</code>. Same goal (256, 256, π/4),
same 200-step horizon, same brain (<code>rule</code> planner + the measured-behaviour
executor). The planner is deterministic on the start state, so the <strong>first
plan is identical on all 25 seed pairs</strong>; its delivery is the clean
per-engine measurement. Everything below is produced by
<code>scripts/contact_study.py --ref B_ref_mu0=batch_pymunk_ref --tag A_wall_mu05=study_A_wall05</code>.</p>
<h2 id="stage-a-25-episodes-on-pymunk-wall-with-the-current-brain">Stage A - 25 episodes on pymunk-wall with the current brain</h2>
<p>Store tag <code>study_A_wall05</code>; videos <code>docs/videos/study_A_wall05/</code>.</p>
<h2 id="stage-b-paired-comparison-against-the-25-pymunk-reference-episodes">Stage B - paired comparison against the 25 pymunk-reference episodes</h2>
<h3 id="outcomes">Outcomes</h3>
<table>
<thead>
<tr>
<th>metric</th>
<th>B ref μ=0</th>
<th>A wall μ=0.5</th>
</tr>
</thead>
<tbody>
<tr>
<td>solved (cov ≥ 0.95 and pusher clear)</td>
<td>2</td>
<td>3</td>
</tr>
<tr>
<td>final coverage mean</td>
<td>0.613</td>
<td>0.582</td>
</tr>
<tr>
<td>final coverage median</td>
<td>0.713</td>
<td>0.469</td>
</tr>
<tr>
<td>final coverage IQR</td>
<td>0.375–0.918</td>
<td>0.358–0.918</td>
</tr>
<tr>
<td>final angle error mean</td>
<td>20.2°</td>
<td>27.6°</td>
</tr>
<tr>
<td>final position error mean</td>
<td>17.6 px</td>
<td>23.7 px</td>
</tr>
<tr>
<td>contact bouts per episode</td>
<td>3.6</td>
<td>4.0</td>
</tr>
</tbody>
</table>
<p>Headline: one more solve, a worse median. Friction did not shift all
episodes one way; it split them. Per seed (<code>coverage_per_seed.png</code>):</p>
<ul>
<li><strong>Improved, three of them to solved:</strong> 32063 (0.746 → 0.969), 44177 (0.918 → 0.969), 67870 (0.834 → 0.971), 64785 (0.795 → 0.897). All "mixed, far" starts whose plans were translate-first.</li>
<li><strong>Collapsed:</strong> 98135 (0.390 → 0.047), 87949 (0.815 → 0.389), 16370 (0.713 → 0.492), 18033 (0.612 → 0.440). Flipped or near-wall starts.</li>
<li><strong>Unchanged (±0.05):</strong> the other 17, including every inverted start the brain fails on either engine.</li>
</ul>
<p><img alt="" src="../figures/contact_study/coverage_per_seed.png" />
<img alt="" src="../figures/contact_study/coverage_distribution.png" />
<img alt="" src="../figures/contact_study/coverage_trajectory.png" /></p>
<h3 id="what-friction-changed-at-the-stroke-level">What friction changed at the stroke level</h3>
<p>Measured over every contact frame of every episode (step logs, wrapped angles):</p>
<table>
<thead>
<tr>
<th>PUSH phase, in contact</th>
<th>B μ=0</th>
<th>A μ=0.5</th>
<th>change</th>
</tr>
</thead>
<tbody>
<tr>
<td>contact point drift along the edge, px/frame (rotate plans)</td>
<td>0.84</td>
<td>0.55</td>
<td><strong>−35 %</strong></td>
</tr>
<tr>
<td>contact point drift along the edge, px/frame (translate plans)</td>
<td>0.64</td>
<td>0.37</td>
<td><strong>−42 %</strong></td>
</tr>
<tr>
<td>block rotation per px of block travel (rotate plans)</td>
<td>0.714 °/px</td>
<td>0.617 °/px</td>
<td><strong>−14 %</strong></td>
</tr>
<tr>
<td>block rotation per px of block travel (translate plans)</td>
<td>0.396 °/px</td>
<td>0.352 °/px</td>
<td>−11 %</td>
</tr>
</tbody>
</table>
<p>First bout, identical plan on identical state, 25 pairs:</p>
<table>
<thead>
<tr>
<th>intent</th>
<th>pairs</th>
<th>B delivered</th>
<th>A delivered</th>
<th>A / B median (IQR)</th>
</tr>
</thead>
<tbody>
<tr>
<td>rotate</td>
<td>9</td>
<td>14.4°</td>
<td>12.4°</td>
<td><strong>0.85</strong> (0.75–1.02)</td>
</tr>
<tr>
<td>translate</td>
<td>14</td>
<td>58.7 px</td>
<td>64.0 px</td>
<td>1.05 (1.01–1.20)</td>
</tr>
</tbody>
</table>
<p>So the earlier expectation - "friction makes every stroke over-deliver" -
was <strong>wrong</strong>. The measured effect is the opposite for rotation:</p>
<ol>
<li><strong>The pusher sticks.</strong> With a friction cone of 26.6° the contact point stops migrating along the edge (drift down 35–42 %). A frictionless push lets the block swing under the pusher until the pushed edge faces the push direction; the pusher slides along the edge as it does so. Friction resists that relative slide, and the resisting tangential force has a moment about the CoG that <strong>opposes the swing</strong>.</li>
<li><strong>Pushes become straighter.</strong> Rotation per px of travel drops 11–14 %; translation is delivered in full or slightly more (1.05×). Rotate plans under-deliver (0.85×). Translate plans also rotate the block less as a side effect (angle change during translate bouts −4.3° vs −7.2°), which is exactly why the translate-first "mixed, far" seeds got <em>better</em>: their translations stopped disturbing the angle.</li>
</ol>
<h3 id="why-the-collapsed-seeds-collapsed-the-alignment-stall-loop">Why the collapsed seeds collapsed: the alignment stall loop</h3>
<p>The executor releases a translating bout when the push direction (the
pushed edge's inward normal) is more than <code>release_alignment_deg</code> off the
CoG→goal direction (<code>executor.py</code>, reason <code>"alignment"</code>). On the frictionless
engine this is the human criterion and it works because the block
self-aligns: a few frames of pushing swing the edge toward the goal
direction and the bout continues. With friction the swing is damped, the
edge does not come round, and the release fires at the earliest permitted
frame - every time, on the same plan, because the planner is deterministic
on a state that has barely moved.</p>
<p>Seed 98135, plans and outcomes:</p>
<pre><code>B: tra bar_end_R -&gt; alignment ang -134-&gt;-132 pos 147-&gt;146
tra bar_end_R -&gt; alignment ang -132-&gt;-104 pos 146-&gt;107 &lt;- the block swung, bout delivered
rot bar_end_R -&gt; bout_length ang -104-&gt;-45 pos 107-&gt;7
A: tra bar_end_R -&gt; alignment ang -134-&gt;-132 pos 147-&gt;146
tra bar_end_R -&gt; alignment ang -132-&gt;-131 pos 146-&gt;144
... nine identical bouts, 1-2 degrees and 1-2 px each ...
tra bar_top -&gt; horizon
</code></pre>
<p>Bout-outcome counts over the 25 episodes: <code>alignment</code> 56 → 63, <code>placed</code> 7 →
10; translate bouts shortened from 19 to 16 frames. The other collapsed
seeds (87949, 16370) show the same pattern with a rotate plan chosen where
the frictionless run's translate had self-aligned.</p>
<h3 id="what-b-says-the-next-brain-must-do">What B says the next brain must do</h3>
<ol>
<li><strong>Expect 0.85× on rotation and ~1.0× on translation</strong> at μ = 0.5, i.e. ask for ~17 % more lever on rotate plans. The gain is small; it is not the main effect.</li>
<li><strong>Stop relying on self-alignment.</strong> A translating bout that is released on <code>alignment</code> having delivered nearly nothing must not be re-issued as the same plan. Either the executor keeps pushing (the block will translate straight - the thing friction is good at) and lets the planner fix the angle afterwards, or the planner treats such a release as "this edge is stalled" and changes edge or intent.</li>
<li><strong>Use the straighter push.</strong> Since translation no longer rotates the block as a side effect, translate-then-rotate sequences are cleaner under friction than they were - the three new solves show it.</li>
</ol>
<p>Per-bout data: <code>figures/contact_study/bout_delivery.png</code>. Raw JSON of every
number above: <code>scripts/contact_study.py --json</code>.</p>
<h2 id="stage-c-brain-contact1">Stage C - Brain-contact1</h2>
<p><code>contact1</code> planner (<code>ContactRulePlanner</code>, <code>src/pusht_sim/brain/planner.py</code>): the
rule planner with three changes, each traced to stage B: rotate levers x1/0.85;
an edge whose last bout stalled (&lt; 5 px and &lt; 3 deg) is skipped once and two
stalls in a row flip the intent; translating bouts release at 90 deg instead
of 60. Store tag <code>study_C_contact1</code>.</p>
<table>
<thead>
<tr>
<th>metric</th>
<th>B ref μ=0</th>
<th>A wall μ=0.5</th>
<th>C contact1</th>
</tr>
</thead>
<tbody>
<tr>
<td>solved</td>
<td>2</td>
<td>3</td>
<td>1</td>
</tr>
<tr>
<td>coverage mean / median</td>
<td>0.613 / 0.713</td>
<td>0.582 / 0.469</td>
<td><strong>0.640 / 0.725</strong></td>
</tr>
<tr>
<td>final angle / position error</td>
<td>20.2° / 17.6 px</td>
<td>27.6° / 23.7 px</td>
<td><strong>16.8° / 12.8 px</strong></td>
</tr>
</tbody>
</table>
<p>First-bout delivery against the reference: rotate 1.06x (was 0.85x), so the
lever gain is right. The stall loops are gone: 98135 0.047 → 0.810, 89650
0.000 → 0.725, 35669 0.092 → 0.390. But the 90 deg release runs translating
bouts past the goal when close - <code>overshoot</code> releases 1 → 18 - and three
seeds that A solved now end in finish loops at 0-3 px with 4-5 deg of angle
(32063, 44177, 67870): the finish pushes through the CoG and cannot turn the
block, and the lever rule is zero below 5 deg, so nothing in the brain can
close a 3-5 deg residual. Aggregate best of the three; solves worst.</p>
<h2 id="stage-d-brain-contact2-two-closed-loop-passes">Stage D - Brain-contact2, two closed-loop passes</h2>
<p><code>contact2</code> (<code>ContactLoopPlanner</code>): contact1 plus (1) an online rotation gain -
delivered angle over the executor's own model (0.56 deg/s per px of lever x
bout seconds), averaged over the episode's rotate bouts, prior 0.85, scaling
the next lever within [0.5, 2]; (2) the 90 deg release only beyond 60 px from
the goal; (3) inside the finish box with &gt; 2 deg left, a small-lever rotate
(4 px/deg, max 20 px) instead of a finish.</p>
<p><strong>D1</strong> (<code>study_D_contact2</code>): 3 solved, mean 0.571. The small-lever trim solved
81358 and 11980 cleanly (3.7 → 2.0 deg, 2.6 → 1.4 deg, no position lost) and
destroyed 32063, 44177, 52463, 57580 (−0.45 to −0.58 each): a rotating bout
near the goal had no angle-reached release and let go only on 25 px of drift.</p>
<p><strong>D2</strong> (<code>study_D2_contact2_trim</code>): one executor rule added - a rotate bout that
starts inside the alignment floor is a <em>trim</em> and releases at the finish
angle tolerance or 6 px of drift (<code>TRIM_DRIFT_PX</code>, <code>executor.py</code>).</p>
<table>
<thead>
<tr>
<th>metric</th>
<th>B</th>
<th>A</th>
<th>C</th>
<th>D1</th>
<th><strong>D2</strong></th>
</tr>
</thead>
<tbody>
<tr>
<td>solved</td>
<td>2</td>
<td>3</td>
<td>1</td>
<td>3</td>
<td><strong>4</strong></td>
</tr>
<tr>
<td>coverage mean</td>
<td>0.613</td>
<td>0.582</td>
<td>0.640</td>
<td>0.571</td>
<td><strong>0.653</strong></td>
</tr>
<tr>
<td>coverage median</td>
<td>0.713</td>
<td>0.469</td>
<td>0.725</td>
<td>0.458</td>
<td>0.712</td>
</tr>
<tr>
<td>final position error</td>
<td>17.6</td>
<td>23.7</td>
<td>12.8</td>
<td>16.3</td>
<td><strong>9.8 px</strong></td>
</tr>
<tr>
<td>final angle error</td>
<td>20.2</td>
<td>27.6</td>
<td>16.8</td>
<td><strong>16.1</strong></td>
<td>23.1°</td>
</tr>
</tbody>
</table>
<p>The four wrecked seeds recovered (+0.36 to +0.48 each); solved: 79092, 81358,
39679, 11980. Two rotation-heavy seeds fell against D1 (75204 −0.25, 18033
−0.22).</p>
<h3 id="what-the-closed-loop-actually-measured">What the closed loop actually measured</h3>
<p>The online gain estimates came out at median <strong>0.59</strong> (IQR 0.54-0.62, 33
bouts), pinning the lever scale at its 2.0 clamp for most of every episode.
That is not a friction measurement: rotate bouts release on alignment or
drift long before the lever's predicted rotation is delivered, so
"delivered / predicted" mostly measures bout truncation. A stroke-level loop
needs a prediction that includes the release rule, or a signal taken while
the bout is still in contact. Open.</p>
<h3 id="standing-of-the-brains-at-05-same-25-seeds">Standing of the brains at μ = 0.5, same 25 seeds</h3>
<p>Best available: <strong>Brain-contact2 + trim (D2)</strong> - most solves, highest mean
coverage, lowest final position error. Its remaining losses are the
inverted/flipped starts every brain fails on either engine (11 seeds below
0.5 in every column), and the gain saturation above. Pages:
<code>docs/videos/contact-study/index.html</code> (all stages, every seed);
<code>scripts/contact_study.py</code>, <code>scripts/contact_study_page.py</code>.</p>
<h2 id="stage-e-contact3-and-the-decision">Stage E - contact3, and the decision</h2>
<p><code>contact3</code> (<code>ContactFrameLoopPlanner</code>) replaces the per-bout gain with one
measured <strong>in contact, per frame</strong>: the executor accumulates block rotation
and CoG travel over the frames it touches the block (<code>contact_turn_deg</code>,
<code>contact_travel_px</code> in every bout outcome), and the ratio against the
reference engine's 0.714 deg/px is the friction effect alone. Clamp [0.7, 1.4].
Store tag <code>study_E_contact3</code>.</p>
<p>The measurement is now right: median gain <strong>0.91</strong> (IQR 0.77-0.99, 36
bouts), consistent with stage B's 0.85. The outcome is not: 3 solved (D2: 4),
mean coverage 0.631 (D2: 0.653), 75204/18033 not recovered, 39679's solve
lost. A 10 % lever change does not move outcomes because rotation delivery
on this engine is decided by where a bout releases, not by lever magnitude.</p>
<p><strong>Decision (acceptance set before the run: adopt only if &gt;= D2 on solves and
mean coverage):</strong> <code>pymunk-wall</code>'s default brain is <strong>Brain-contact2 + trim</strong>
(<code>DEFAULT_PLANNER_BY_BACKEND</code> in <code>brain/planner.py</code>; the tab, the CLI and
<code>run_brain_episode</code> follow it; <code>--planner</code> overrides). <code>pymunk-reference</code>
keeps the rule brain. <code>contact3</code> stays registered as the correct friction
measurement for whoever builds the next release rule.</p>
<h2 id="50-seed-eval-each-engine-with-its-own-default-brain">50-seed eval: each engine with its own default brain</h2>
<p>50 fresh seeds (disjoint from every earlier batch), same goal, 200 steps.
<code>pymunk-reference</code> with the rule brain against <code>pymunk-wall</code> (μ = 0.5)
with Brain-contact2 + trim. Tags <code>eval50_pymunk_ref</code>, <code>eval50_pymunk_wall</code>;
page <code>docs/videos/eval50/</code>. This compares engine+brain pairs; the
friction-only comparison is the 25-seed study above.</p>
<table>
<thead>
<tr>
<th>metric</th>
<th>REF rule</th>
<th>WALL contact2</th>
</tr>
</thead>
<tbody>
<tr>
<td>solved</td>
<td>5</td>
<td>6</td>
</tr>
<tr>
<td>coverage mean / median</td>
<td>0.591 / 0.719</td>
<td>0.584 / 0.623</td>
</tr>
<tr>
<td>final position error</td>
<td>24.6 px</td>
<td>20.9 px</td>
</tr>
<tr>
<td>final angle error</td>
<td>19.7°</td>
<td>18.8°</td>
</tr>
</tbody>
</table>
<p>Paired per seed: wall − ref mean −0.008, median +0.004; better on 18 seeds,
worse on 16, within ±0.05 on 16; Wilcoxon p = 0.80. <strong>No difference at the
population level.</strong> The two batches solve <em>different</em> seeds (no overlap:
ref 10560, 22402, 73939, 93128, 97593; wall 3075, 33340, 45081, 67306,
91328, 95399) and the eight largest swings are ±0.5-0.66 in both directions -
the brains are deterministic and each has starts it handles and starts it
does not. By start type: flipped n=19 0.530 vs 0.506, mixed n=20 0.711 vs
0.696, aligned n=6 0.608 vs 0.669, inverted n=5 0.33 vs 0.33 (0 solved on
either).</p>
<p>Reading: at μ = 0.5 with its own brain, pymunk-wall performs as the
published frictionless environment does with the published brain - the
friction study's adaptations bought back exactly what friction cost, no
more. The ceiling on both is the planner (inverted and most flipped starts),
not the contact model.</p></body></html>