Spaces:
Running
Running
Download docs/html/contact_friction_study.html from desertmouse/pusht-simulator: direct link, hf CLI and curl.
- Browser
- Download file 18.9 kB
-
https://huggingface.co/spaces/desertmouse/pusht-simulator/resolve/main/docs/html/contact_friction_study.html
- Command line
-
hf download hf://spaces/desertmouse/pusht-simulator/docs/html/contact_friction_study.html
-
curl -L -o contact_friction_study.html https://huggingface.co/spaces/desertmouse/pusht-simulator/resolve/main/docs/html/contact_friction_study.html
18.9 kB
| <html><head><meta charset='utf-8'><title>Contact-friction study: what μ = 0.5 between pusher and T does to the brain</title><style> | |
| body{max-width:980px;margin:2rem auto;padding:0 1.25rem;font:15px/1.55 -apple-system,Segoe UI,Roboto,sans-serif;color:#1c2330;background:#fafbfc} | |
| h1{font-size:1.7rem;border-bottom:2px solid #d8dee8;padding-bottom:.4rem}h2{margin-top:2.2rem;border-bottom:1px solid #e3e8ef;padding-bottom:.25rem} | |
| table{border-collapse:collapse;margin:1rem 0;font-size:13.5px;font-variant-numeric:tabular-nums}th,td{border:1px solid #d8dee8;padding:.35rem .6rem;text-align:left} | |
| th{background:#eef2f7}td:not(:first-child){text-align:right}code{background:#eef2f7;padding:.1rem .3rem;border-radius:3px;font-size:.92em} | |
| pre{background:#1c2330;color:#e6edf5;padding:.8rem 1rem;border-radius:6px;overflow:auto}pre code{background:none;color:inherit} | |
| img{max-width:100%;border:1px solid #d8dee8;border-radius:4px;margin:.5rem 0}blockquote{border-left:3px solid #6aa9ff;margin:0;padding:.2rem 1rem;color:#4b5563} | |
| nav{font-size:13px;margin-bottom:1.5rem}nav a{margin-right:1rem} | |
| </style></head><body><nav><a href="index.html">index</a> <a href="brain3_vs_3a_events.html">brain3 vs 3a events</a> <a href="contact_friction_study.html">contact friction study</a> <a href="convergence_study.html">convergence study</a> <a href="dataset_definition.html">dataset definition</a> <a href="ds08_distribution_at_1000.html">ds08 distribution at 1000</a> <a href="ds_distribution_at_1000.html">ds distribution at 1000</a> <a href="ds_distribution_report.html">ds distribution report</a> <a href="ds_mu03_brain3b_distribution.html">ds mu03 brain3b distribution</a> <a href="episode_store.html">episode store</a> <a href="human_pushing_analysis.html">human pushing analysis</a> <a href="lewm_parity.html">lewm parity</a> <a href="reference_parity.html">reference parity</a> <a href="session_fidelity_and_behaviour.html">session fidelity and behaviour</a></nav><h1 id="contact-friction-study-what-05-between-pusher-and-t-does-to-the-brain">Contact-friction study: what μ = 0.5 between pusher and T does to the brain</h1> | |
| <p>All runs in this study are on <strong>pymunk</strong>. Nothing here uses SuperDex. The | |
| two engines compared are:</p> | |
| <ul> | |
| <li><strong>B</strong> <code>pymunk-reference</code> - the published LeWM/gym-pusht environment, bit-exact, frictionless (<code>Shape.friction</code> never assigned, so 0).</li> | |
| <li><strong>A</strong> <code>pymunk-wall</code> - the same environment with one change: <code>contact_friction = 0.5</code> on the pusher circle and the two T polygons (<code>src/pusht_sim/envs/pymunk_wall_env.py</code>). Walls, <code>damping = 0</code>, mass, moment, seeds - unchanged; μ = 0 reproduces B bit for bit (<code>tests/test_env_pymunk_wall.py</code>).</li> | |
| </ul> | |
| <p>Same 25 seeds on both (the <code>batch_pymunk_ref</code> eval seeds, also published as | |
| <code>desertmouse/pusht-simulator-pymunk-reference</code>): <code>44177 18033 67870 16370 | |
| 57580 75204 98135 87949 81400 31775 75263 64785 52463 11980 68814 35669 2103 | |
| 39679 2228 45997 32063 81358 79092 89650 91922</code>. Same goal (256, 256, π/4), | |
| same 200-step horizon, same brain (<code>rule</code> planner + the measured-behaviour | |
| executor). The planner is deterministic on the start state, so the <strong>first | |
| plan is identical on all 25 seed pairs</strong>; its delivery is the clean | |
| per-engine measurement. Everything below is produced by | |
| <code>scripts/contact_study.py --ref B_ref_mu0=batch_pymunk_ref --tag A_wall_mu05=study_A_wall05</code>.</p> | |
| <h2 id="stage-a-25-episodes-on-pymunk-wall-with-the-current-brain">Stage A - 25 episodes on pymunk-wall with the current brain</h2> | |
| <p>Store tag <code>study_A_wall05</code>; videos <code>docs/videos/study_A_wall05/</code>.</p> | |
| <h2 id="stage-b-paired-comparison-against-the-25-pymunk-reference-episodes">Stage B - paired comparison against the 25 pymunk-reference episodes</h2> | |
| <h3 id="outcomes">Outcomes</h3> | |
| <table> | |
| <thead> | |
| <tr> | |
| <th>metric</th> | |
| <th>B ref μ=0</th> | |
| <th>A wall μ=0.5</th> | |
| </tr> | |
| </thead> | |
| <tbody> | |
| <tr> | |
| <td>solved (cov ≥ 0.95 and pusher clear)</td> | |
| <td>2</td> | |
| <td>3</td> | |
| </tr> | |
| <tr> | |
| <td>final coverage mean</td> | |
| <td>0.613</td> | |
| <td>0.582</td> | |
| </tr> | |
| <tr> | |
| <td>final coverage median</td> | |
| <td>0.713</td> | |
| <td>0.469</td> | |
| </tr> | |
| <tr> | |
| <td>final coverage IQR</td> | |
| <td>0.375–0.918</td> | |
| <td>0.358–0.918</td> | |
| </tr> | |
| <tr> | |
| <td>final angle error mean</td> | |
| <td>20.2°</td> | |
| <td>27.6°</td> | |
| </tr> | |
| <tr> | |
| <td>final position error mean</td> | |
| <td>17.6 px</td> | |
| <td>23.7 px</td> | |
| </tr> | |
| <tr> | |
| <td>contact bouts per episode</td> | |
| <td>3.6</td> | |
| <td>4.0</td> | |
| </tr> | |
| </tbody> | |
| </table> | |
| <p>Headline: one more solve, a worse median. Friction did not shift all | |
| episodes one way; it split them. Per seed (<code>coverage_per_seed.png</code>):</p> | |
| <ul> | |
| <li><strong>Improved, three of them to solved:</strong> 32063 (0.746 → 0.969), 44177 (0.918 → 0.969), 67870 (0.834 → 0.971), 64785 (0.795 → 0.897). All "mixed, far" starts whose plans were translate-first.</li> | |
| <li><strong>Collapsed:</strong> 98135 (0.390 → 0.047), 87949 (0.815 → 0.389), 16370 (0.713 → 0.492), 18033 (0.612 → 0.440). Flipped or near-wall starts.</li> | |
| <li><strong>Unchanged (±0.05):</strong> the other 17, including every inverted start the brain fails on either engine.</li> | |
| </ul> | |
| <p><img alt="" src="../figures/contact_study/coverage_per_seed.png" /> | |
| <img alt="" src="../figures/contact_study/coverage_distribution.png" /> | |
| <img alt="" src="../figures/contact_study/coverage_trajectory.png" /></p> | |
| <h3 id="what-friction-changed-at-the-stroke-level">What friction changed at the stroke level</h3> | |
| <p>Measured over every contact frame of every episode (step logs, wrapped angles):</p> | |
| <table> | |
| <thead> | |
| <tr> | |
| <th>PUSH phase, in contact</th> | |
| <th>B μ=0</th> | |
| <th>A μ=0.5</th> | |
| <th>change</th> | |
| </tr> | |
| </thead> | |
| <tbody> | |
| <tr> | |
| <td>contact point drift along the edge, px/frame (rotate plans)</td> | |
| <td>0.84</td> | |
| <td>0.55</td> | |
| <td><strong>−35 %</strong></td> | |
| </tr> | |
| <tr> | |
| <td>contact point drift along the edge, px/frame (translate plans)</td> | |
| <td>0.64</td> | |
| <td>0.37</td> | |
| <td><strong>−42 %</strong></td> | |
| </tr> | |
| <tr> | |
| <td>block rotation per px of block travel (rotate plans)</td> | |
| <td>0.714 °/px</td> | |
| <td>0.617 °/px</td> | |
| <td><strong>−14 %</strong></td> | |
| </tr> | |
| <tr> | |
| <td>block rotation per px of block travel (translate plans)</td> | |
| <td>0.396 °/px</td> | |
| <td>0.352 °/px</td> | |
| <td>−11 %</td> | |
| </tr> | |
| </tbody> | |
| </table> | |
| <p>First bout, identical plan on identical state, 25 pairs:</p> | |
| <table> | |
| <thead> | |
| <tr> | |
| <th>intent</th> | |
| <th>pairs</th> | |
| <th>B delivered</th> | |
| <th>A delivered</th> | |
| <th>A / B median (IQR)</th> | |
| </tr> | |
| </thead> | |
| <tbody> | |
| <tr> | |
| <td>rotate</td> | |
| <td>9</td> | |
| <td>14.4°</td> | |
| <td>12.4°</td> | |
| <td><strong>0.85</strong> (0.75–1.02)</td> | |
| </tr> | |
| <tr> | |
| <td>translate</td> | |
| <td>14</td> | |
| <td>58.7 px</td> | |
| <td>64.0 px</td> | |
| <td>1.05 (1.01–1.20)</td> | |
| </tr> | |
| </tbody> | |
| </table> | |
| <p>So the earlier expectation - "friction makes every stroke over-deliver" - | |
| was <strong>wrong</strong>. The measured effect is the opposite for rotation:</p> | |
| <ol> | |
| <li><strong>The pusher sticks.</strong> With a friction cone of 26.6° the contact point stops migrating along the edge (drift down 35–42 %). A frictionless push lets the block swing under the pusher until the pushed edge faces the push direction; the pusher slides along the edge as it does so. Friction resists that relative slide, and the resisting tangential force has a moment about the CoG that <strong>opposes the swing</strong>.</li> | |
| <li><strong>Pushes become straighter.</strong> Rotation per px of travel drops 11–14 %; translation is delivered in full or slightly more (1.05×). Rotate plans under-deliver (0.85×). Translate plans also rotate the block less as a side effect (angle change during translate bouts −4.3° vs −7.2°), which is exactly why the translate-first "mixed, far" seeds got <em>better</em>: their translations stopped disturbing the angle.</li> | |
| </ol> | |
| <h3 id="why-the-collapsed-seeds-collapsed-the-alignment-stall-loop">Why the collapsed seeds collapsed: the alignment stall loop</h3> | |
| <p>The executor releases a translating bout when the push direction (the | |
| pushed edge's inward normal) is more than <code>release_alignment_deg</code> off the | |
| CoG→goal direction (<code>executor.py</code>, reason <code>"alignment"</code>). On the frictionless | |
| engine this is the human criterion and it works because the block | |
| self-aligns: a few frames of pushing swing the edge toward the goal | |
| direction and the bout continues. With friction the swing is damped, the | |
| edge does not come round, and the release fires at the earliest permitted | |
| frame - every time, on the same plan, because the planner is deterministic | |
| on a state that has barely moved.</p> | |
| <p>Seed 98135, plans and outcomes:</p> | |
| <pre><code>B: tra bar_end_R -> alignment ang -134->-132 pos 147->146 | |
| tra bar_end_R -> alignment ang -132->-104 pos 146->107 <- the block swung, bout delivered | |
| rot bar_end_R -> bout_length ang -104->-45 pos 107->7 | |
| A: tra bar_end_R -> alignment ang -134->-132 pos 147->146 | |
| tra bar_end_R -> alignment ang -132->-131 pos 146->144 | |
| ... nine identical bouts, 1-2 degrees and 1-2 px each ... | |
| tra bar_top -> horizon | |
| </code></pre> | |
| <p>Bout-outcome counts over the 25 episodes: <code>alignment</code> 56 → 63, <code>placed</code> 7 → | |
| 10; translate bouts shortened from 19 to 16 frames. The other collapsed | |
| seeds (87949, 16370) show the same pattern with a rotate plan chosen where | |
| the frictionless run's translate had self-aligned.</p> | |
| <h3 id="what-b-says-the-next-brain-must-do">What B says the next brain must do</h3> | |
| <ol> | |
| <li><strong>Expect 0.85× on rotation and ~1.0× on translation</strong> at μ = 0.5, i.e. ask for ~17 % more lever on rotate plans. The gain is small; it is not the main effect.</li> | |
| <li><strong>Stop relying on self-alignment.</strong> A translating bout that is released on <code>alignment</code> having delivered nearly nothing must not be re-issued as the same plan. Either the executor keeps pushing (the block will translate straight - the thing friction is good at) and lets the planner fix the angle afterwards, or the planner treats such a release as "this edge is stalled" and changes edge or intent.</li> | |
| <li><strong>Use the straighter push.</strong> Since translation no longer rotates the block as a side effect, translate-then-rotate sequences are cleaner under friction than they were - the three new solves show it.</li> | |
| </ol> | |
| <p>Per-bout data: <code>figures/contact_study/bout_delivery.png</code>. Raw JSON of every | |
| number above: <code>scripts/contact_study.py --json</code>.</p> | |
| <h2 id="stage-c-brain-contact1">Stage C - Brain-contact1</h2> | |
| <p><code>contact1</code> planner (<code>ContactRulePlanner</code>, <code>src/pusht_sim/brain/planner.py</code>): the | |
| rule planner with three changes, each traced to stage B: rotate levers x1/0.85; | |
| an edge whose last bout stalled (< 5 px and < 3 deg) is skipped once and two | |
| stalls in a row flip the intent; translating bouts release at 90 deg instead | |
| of 60. Store tag <code>study_C_contact1</code>.</p> | |
| <table> | |
| <thead> | |
| <tr> | |
| <th>metric</th> | |
| <th>B ref μ=0</th> | |
| <th>A wall μ=0.5</th> | |
| <th>C contact1</th> | |
| </tr> | |
| </thead> | |
| <tbody> | |
| <tr> | |
| <td>solved</td> | |
| <td>2</td> | |
| <td>3</td> | |
| <td>1</td> | |
| </tr> | |
| <tr> | |
| <td>coverage mean / median</td> | |
| <td>0.613 / 0.713</td> | |
| <td>0.582 / 0.469</td> | |
| <td><strong>0.640 / 0.725</strong></td> | |
| </tr> | |
| <tr> | |
| <td>final angle / position error</td> | |
| <td>20.2° / 17.6 px</td> | |
| <td>27.6° / 23.7 px</td> | |
| <td><strong>16.8° / 12.8 px</strong></td> | |
| </tr> | |
| </tbody> | |
| </table> | |
| <p>First-bout delivery against the reference: rotate 1.06x (was 0.85x), so the | |
| lever gain is right. The stall loops are gone: 98135 0.047 → 0.810, 89650 | |
| 0.000 → 0.725, 35669 0.092 → 0.390. But the 90 deg release runs translating | |
| bouts past the goal when close - <code>overshoot</code> releases 1 → 18 - and three | |
| seeds that A solved now end in finish loops at 0-3 px with 4-5 deg of angle | |
| (32063, 44177, 67870): the finish pushes through the CoG and cannot turn the | |
| block, and the lever rule is zero below 5 deg, so nothing in the brain can | |
| close a 3-5 deg residual. Aggregate best of the three; solves worst.</p> | |
| <h2 id="stage-d-brain-contact2-two-closed-loop-passes">Stage D - Brain-contact2, two closed-loop passes</h2> | |
| <p><code>contact2</code> (<code>ContactLoopPlanner</code>): contact1 plus (1) an online rotation gain - | |
| delivered angle over the executor's own model (0.56 deg/s per px of lever x | |
| bout seconds), averaged over the episode's rotate bouts, prior 0.85, scaling | |
| the next lever within [0.5, 2]; (2) the 90 deg release only beyond 60 px from | |
| the goal; (3) inside the finish box with > 2 deg left, a small-lever rotate | |
| (4 px/deg, max 20 px) instead of a finish.</p> | |
| <p><strong>D1</strong> (<code>study_D_contact2</code>): 3 solved, mean 0.571. The small-lever trim solved | |
| 81358 and 11980 cleanly (3.7 → 2.0 deg, 2.6 → 1.4 deg, no position lost) and | |
| destroyed 32063, 44177, 52463, 57580 (−0.45 to −0.58 each): a rotating bout | |
| near the goal had no angle-reached release and let go only on 25 px of drift.</p> | |
| <p><strong>D2</strong> (<code>study_D2_contact2_trim</code>): one executor rule added - a rotate bout that | |
| starts inside the alignment floor is a <em>trim</em> and releases at the finish | |
| angle tolerance or 6 px of drift (<code>TRIM_DRIFT_PX</code>, <code>executor.py</code>).</p> | |
| <table> | |
| <thead> | |
| <tr> | |
| <th>metric</th> | |
| <th>B</th> | |
| <th>A</th> | |
| <th>C</th> | |
| <th>D1</th> | |
| <th><strong>D2</strong></th> | |
| </tr> | |
| </thead> | |
| <tbody> | |
| <tr> | |
| <td>solved</td> | |
| <td>2</td> | |
| <td>3</td> | |
| <td>1</td> | |
| <td>3</td> | |
| <td><strong>4</strong></td> | |
| </tr> | |
| <tr> | |
| <td>coverage mean</td> | |
| <td>0.613</td> | |
| <td>0.582</td> | |
| <td>0.640</td> | |
| <td>0.571</td> | |
| <td><strong>0.653</strong></td> | |
| </tr> | |
| <tr> | |
| <td>coverage median</td> | |
| <td>0.713</td> | |
| <td>0.469</td> | |
| <td>0.725</td> | |
| <td>0.458</td> | |
| <td>0.712</td> | |
| </tr> | |
| <tr> | |
| <td>final position error</td> | |
| <td>17.6</td> | |
| <td>23.7</td> | |
| <td>12.8</td> | |
| <td>16.3</td> | |
| <td><strong>9.8 px</strong></td> | |
| </tr> | |
| <tr> | |
| <td>final angle error</td> | |
| <td>20.2</td> | |
| <td>27.6</td> | |
| <td>16.8</td> | |
| <td><strong>16.1</strong></td> | |
| <td>23.1°</td> | |
| </tr> | |
| </tbody> | |
| </table> | |
| <p>The four wrecked seeds recovered (+0.36 to +0.48 each); solved: 79092, 81358, | |
| 39679, 11980. Two rotation-heavy seeds fell against D1 (75204 −0.25, 18033 | |
| −0.22).</p> | |
| <h3 id="what-the-closed-loop-actually-measured">What the closed loop actually measured</h3> | |
| <p>The online gain estimates came out at median <strong>0.59</strong> (IQR 0.54-0.62, 33 | |
| bouts), pinning the lever scale at its 2.0 clamp for most of every episode. | |
| That is not a friction measurement: rotate bouts release on alignment or | |
| drift long before the lever's predicted rotation is delivered, so | |
| "delivered / predicted" mostly measures bout truncation. A stroke-level loop | |
| needs a prediction that includes the release rule, or a signal taken while | |
| the bout is still in contact. Open.</p> | |
| <h3 id="standing-of-the-brains-at-05-same-25-seeds">Standing of the brains at μ = 0.5, same 25 seeds</h3> | |
| <p>Best available: <strong>Brain-contact2 + trim (D2)</strong> - most solves, highest mean | |
| coverage, lowest final position error. Its remaining losses are the | |
| inverted/flipped starts every brain fails on either engine (11 seeds below | |
| 0.5 in every column), and the gain saturation above. Pages: | |
| <code>docs/videos/contact-study/index.html</code> (all stages, every seed); | |
| <code>scripts/contact_study.py</code>, <code>scripts/contact_study_page.py</code>.</p> | |
| <h2 id="stage-e-contact3-and-the-decision">Stage E - contact3, and the decision</h2> | |
| <p><code>contact3</code> (<code>ContactFrameLoopPlanner</code>) replaces the per-bout gain with one | |
| measured <strong>in contact, per frame</strong>: the executor accumulates block rotation | |
| and CoG travel over the frames it touches the block (<code>contact_turn_deg</code>, | |
| <code>contact_travel_px</code> in every bout outcome), and the ratio against the | |
| reference engine's 0.714 deg/px is the friction effect alone. Clamp [0.7, 1.4]. | |
| Store tag <code>study_E_contact3</code>.</p> | |
| <p>The measurement is now right: median gain <strong>0.91</strong> (IQR 0.77-0.99, 36 | |
| bouts), consistent with stage B's 0.85. The outcome is not: 3 solved (D2: 4), | |
| mean coverage 0.631 (D2: 0.653), 75204/18033 not recovered, 39679's solve | |
| lost. A 10 % lever change does not move outcomes because rotation delivery | |
| on this engine is decided by where a bout releases, not by lever magnitude.</p> | |
| <p><strong>Decision (acceptance set before the run: adopt only if >= D2 on solves and | |
| mean coverage):</strong> <code>pymunk-wall</code>'s default brain is <strong>Brain-contact2 + trim</strong> | |
| (<code>DEFAULT_PLANNER_BY_BACKEND</code> in <code>brain/planner.py</code>; the tab, the CLI and | |
| <code>run_brain_episode</code> follow it; <code>--planner</code> overrides). <code>pymunk-reference</code> | |
| keeps the rule brain. <code>contact3</code> stays registered as the correct friction | |
| measurement for whoever builds the next release rule.</p> | |
| <h2 id="50-seed-eval-each-engine-with-its-own-default-brain">50-seed eval: each engine with its own default brain</h2> | |
| <p>50 fresh seeds (disjoint from every earlier batch), same goal, 200 steps. | |
| <code>pymunk-reference</code> with the rule brain against <code>pymunk-wall</code> (μ = 0.5) | |
| with Brain-contact2 + trim. Tags <code>eval50_pymunk_ref</code>, <code>eval50_pymunk_wall</code>; | |
| page <code>docs/videos/eval50/</code>. This compares engine+brain pairs; the | |
| friction-only comparison is the 25-seed study above.</p> | |
| <table> | |
| <thead> | |
| <tr> | |
| <th>metric</th> | |
| <th>REF rule</th> | |
| <th>WALL contact2</th> | |
| </tr> | |
| </thead> | |
| <tbody> | |
| <tr> | |
| <td>solved</td> | |
| <td>5</td> | |
| <td>6</td> | |
| </tr> | |
| <tr> | |
| <td>coverage mean / median</td> | |
| <td>0.591 / 0.719</td> | |
| <td>0.584 / 0.623</td> | |
| </tr> | |
| <tr> | |
| <td>final position error</td> | |
| <td>24.6 px</td> | |
| <td>20.9 px</td> | |
| </tr> | |
| <tr> | |
| <td>final angle error</td> | |
| <td>19.7°</td> | |
| <td>18.8°</td> | |
| </tr> | |
| </tbody> | |
| </table> | |
| <p>Paired per seed: wall − ref mean −0.008, median +0.004; better on 18 seeds, | |
| worse on 16, within ±0.05 on 16; Wilcoxon p = 0.80. <strong>No difference at the | |
| population level.</strong> The two batches solve <em>different</em> seeds (no overlap: | |
| ref 10560, 22402, 73939, 93128, 97593; wall 3075, 33340, 45081, 67306, | |
| 91328, 95399) and the eight largest swings are ±0.5-0.66 in both directions - | |
| the brains are deterministic and each has starts it handles and starts it | |
| does not. By start type: flipped n=19 0.530 vs 0.506, mixed n=20 0.711 vs | |
| 0.696, aligned n=6 0.608 vs 0.669, inverted n=5 0.33 vs 0.33 (0 solved on | |
| either).</p> | |
| <p>Reading: at μ = 0.5 with its own brain, pymunk-wall performs as the | |
| published frictionless environment does with the published brain - the | |
| friction study's adaptations bought back exactly what friction cost, no | |
| more. The ceiling on both is the planner (inverted and most flipped starts), | |
| not the contact model.</p></body></html> |