c2p β€” Coarse-to-Fine Contact β†’ Pressure (Ego2Contact)

Bimanual tactile estimation from egocentric video, built on TouchAnything + EgoTouch. A decoupled two-stage design: Stage-1 trains a standalone contact head (per-cell contact probability p_c on the 21Γ—21 grid, both hands) from a frozen DINOv2 view + WiLoR hand features; Stage-2/3 add pressure magnitude with a flow-matching GridDiT that consumes the Stage-1 contact map. This repo holds the checkpoints + full test-eval metrics for every stage.

Model families

dir model one-line
stage1_contact_head/ Stage-1 head contact prob only, no magnitude (soft-BCE + Dice + part-CoP)
stage2a/ Stage-2a frozen head (detached) β†’ DiT-from-scratch, p_c as zero-init conditioning
stage2_gate/ Stage-2 GATE frozen head p_c soft-gates the DiT output (floor 0.3), gate trained in-loop
stage3/ Stage-3 JOINT init from GATE, unfreeze head + let the gated-final gradient co-train it (0.1Γ— LR)

Each stageX/ has best_model.pth (+ last_model.pth for stage2a/stage3), and eval/stageX/test_{seen,unseen}_metrics.json. The two reference baselines' full evals (same protocol) are in eval/v2_dit/ and eval/ta_base/ (test_{seen,unseen}_metrics.json), for direct comparison with the c2p line.

Test eval β€” grid space, threshold @0.05, per-seq pooled (179 seen / 85 unseen)

Contact metrics for the pressure models are on the final pressure map; Stage-1 scores p_c>0.5 vs gt>0.05.

model cIoU s cIoU u T.Acc s T.Acc u vIoU s vIoU u MAE↓ s MAE↓ u CoP↓ s CoP↓ u
Stage-1 head 0.5117 0.4781 0.891 0.856 β€” β€” β€” β€” 0.184 0.166
Stage-2a 0.5428 0.4647 0.889 0.835 0.442 0.315 0.042 0.061 β€” β€”
Stage-2 GATE 0.5585 0.4739 0.898 0.838 0.479 0.325 0.039 0.060 0.164 0.155
Stage-3 JOINT 0.5680 0.4631 0.903 0.837 0.493 0.322 0.038 0.059 0.149 0.144
v2_dit (single-stage bar) 0.5735 0.4812 0.902 0.854 0.496 0.342 0.038 0.059 β€” β€”
ta_base (regression bar) 0.5064 0.4720 0.890 0.838 0.425 0.345 0.049 0.064 β€” β€”

Findings

  • Stage-2a (conditioning) improves seen but its from-scratch DiT overfits seen magnitude and degrades unseen (0.4647, below the head) β€” the DiT re-learns localization from scratch.
  • Stage-2 GATE (learned in-loop gate) is the best-balanced two-stage model: it fixes Stage-2a's OOD collapse (unseen +.009 and seen +.016 vs 2a) and beats the whole contact-prior family + a post-hoc late-fusion gate. It trails single-stage v2_dit but ties it on unseen.
  • Stage-3 JOINT (unfreeze + co-train the head) gives the best in-distribution numbers (seen cIoU 0.568, edging v2_dit on seen T.Acc/MAE) and lifts the head's own contact cIoU 0.518β†’0.55, but sacrifices OOD β€” unseen drops to the family's lowest (0.4631). A clean demonstration that freezing the contact prior protects generalization.
  • Cross-model failure analysis (per-task cIoU, GATE vs v2_dit vs ta_base): all three heads bottom out on the same tasks β€” press_multimeter_button, move_toolbox, fold_clothes_and_bag, shop_for_fruit, organize_kitchen_items, toss_and_catch_tennis_ball β€” in nearly the same rank order, so the long left tail is task-intrinsic (small / occluded / dynamic / OOD contact), not model-specific. The one sharp exception: on OOD move_toolbox the GATE is uniquely worst (0.326 vs v2_dit 0.361 / ta_base 0.360) β€” the frozen head under-localizes the occluded heavy-grip and the multiplicative gate enforces it, while the un-gated v2_dit and the area-over-predicting regression baseline both escape it. The concrete cost of the freeze.

Takeaway: the two-stage contact-prior closes most of the gap to a well-optimized single-stage flow model but does not overtake it; among two-stage variants, freeze-the-head (GATE) is the right default for OOD, while co-training (Stage-3) trades OOD robustness for in-distribution accuracy.

Demo videos (GT vs predicted pressure): qqyang/ego2contact-demo-videos.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support