--- license: mit library_name: pytorch tags: - tactile - contact-estimation - pressure-estimation - egocentric - touchanything --- # c2p — Coarse-to-Fine Contact → Pressure (Ego2Contact) Bimanual tactile estimation from egocentric video, built on **TouchAnything + EgoTouch**. A decoupled two-stage design: **Stage-1** trains a standalone contact head (per-cell contact probability `p_c` on the 21×21 grid, both hands) from a frozen **DINOv2** view + **WiLoR** hand features; **Stage-2/3** add pressure *magnitude* with a flow-matching GridDiT that consumes the Stage-1 contact map. This repo holds the checkpoints + full test-eval metrics for every stage. ## Model families | dir | model | one-line | |---|---|---| | `stage1_contact_head/` | **Stage-1 head** | contact prob only, no magnitude (soft-BCE + Dice + part-CoP) | | `stage2a/` | **Stage-2a** | frozen head (detached) → DiT-from-scratch, p_c as zero-init conditioning | | `stage2_gate/` | **Stage-2 GATE** | frozen head `p_c` **soft-gates** the DiT output (floor 0.3), gate trained in-loop | | `stage3/` | **Stage-3 JOINT** | init from GATE, **unfreeze** head + let the gated-final gradient co-train it (0.1× LR) | Each `stageX/` has `best_model.pth` (+ `last_model.pth` for stage2a/stage3), and `eval/stageX/test_{seen,unseen}_metrics.json`. The two reference baselines' full evals (same protocol) are in `eval/v2_dit/` and `eval/ta_base/` (`test_{seen,unseen}_metrics.json`), for direct comparison with the c2p line. ## Test eval — grid space, threshold @0.05, per-seq pooled (179 seen / 85 unseen) Contact metrics for the pressure models are on the **final pressure** map; Stage-1 scores `p_c>0.5` vs `gt>0.05`. | model | cIoU s | cIoU u | T.Acc s | T.Acc u | vIoU s | vIoU u | MAE↓ s | MAE↓ u | CoP↓ s | CoP↓ u | |---|---|---|---|---|---|---|---|---|---|---| | Stage-1 head | 0.5117 | 0.4781 | 0.891 | 0.856 | — | — | — | — | 0.184 | 0.166 | | Stage-2a | 0.5428 | 0.4647 | 0.889 | 0.835 | 0.442 | 0.315 | 0.042 | 0.061 | — | — | | **Stage-2 GATE** | **0.5585** | **0.4739** | 0.898 | 0.838 | 0.479 | 0.325 | 0.039 | 0.060 | 0.164 | 0.155 | | Stage-3 JOINT | 0.5680 | 0.4631 | **0.903** | 0.837 | 0.493 | 0.322 | **0.038** | 0.059 | 0.149 | 0.144 | | _v2_dit (single-stage bar)_ | _0.5735_ | _0.4812_ | _0.902_ | _0.854_ | _0.496_ | _0.342_ | _0.038_ | _0.059_ | _—_ | _—_ | | _ta_base (regression bar)_ | _0.5064_ | _0.4720_ | _0.890_ | _0.838_ | _0.425_ | _0.345_ | _0.049_ | _0.064_ | _—_ | _—_ | ## Findings - **Stage-2a** (conditioning) improves *seen* but its from-scratch DiT overfits seen magnitude and **degrades unseen** (0.4647, below the head) — the DiT re-learns localization from scratch. - **Stage-2 GATE** (learned in-loop gate) is the **best-balanced two-stage model**: it fixes Stage-2a's OOD collapse (unseen +.009 *and* seen +.016 vs 2a) and beats the whole contact-prior family + a post-hoc late-fusion gate. It trails single-stage v2_dit but ties it on unseen. - **Stage-3 JOINT** (unfreeze + co-train the head) gives the **best in-distribution** numbers (seen cIoU 0.568, edging v2_dit on seen T.Acc/MAE) and lifts the head's own contact cIoU 0.518→0.55, **but sacrifices OOD** — unseen drops to the family's lowest (0.4631). A clean demonstration that freezing the contact prior protects generalization. - **Cross-model failure analysis** (per-task cIoU, GATE vs v2_dit vs ta_base): all three heads bottom out on the *same* tasks — `press_multimeter_button`, `move_toolbox`, `fold_clothes_and_bag`, `shop_for_fruit`, `organize_kitchen_items`, `toss_and_catch_tennis_ball` — in nearly the same rank order, so the long left tail is **task-intrinsic** (small / occluded / dynamic / OOD contact), not model-specific. The one sharp exception: on OOD `move_toolbox` the **GATE is uniquely worst** (0.326 vs v2_dit 0.361 / ta_base 0.360) — the frozen head under-localizes the occluded heavy-grip and the multiplicative gate *enforces* it, while the un-gated v2_dit and the area-over-predicting regression baseline both escape it. The concrete cost of the freeze. **Takeaway:** the two-stage contact-prior closes most of the gap to a well-optimized single-stage flow model but does not overtake it; among two-stage variants, **freeze-the-head (GATE)** is the right default for OOD, while **co-training (Stage-3)** trades OOD robustness for in-distribution accuracy. Demo videos (GT vs predicted pressure): [qqyang/ego2contact-demo-videos](https://huggingface.co/qqyang/ego2contact-demo-videos).