Millimetre object pose from frozen DINOv3 β€” MimicGen threading_d0

Decode metric 3D object position from vision on MimicGen threading_d0, and predict it forward with an action-conditioned world model. Task requirement for reliable insertion is ~4 mm; offset (needle βˆ’ tripod) is the decision metric, because relative pose is what the controller consumes.

All numbers are mean Euclidean distance in millimetres on held-out episodes β€” mean_i β€–p_i βˆ’ t_iβ€–β‚‚, the average 3D miss distance, not RMSE. Camera geometry: 45Β° fovy, 0.762 m to the target plane, 2.82 mm per pixel at 224 px; one patch-16 token spans β‰ˆ45 mm.

Headline

representation needle tripod offset
frozen DINOv3-S/16 + soft-argmax readout 2.1 5.5 7.5
SegHead sup224 (dense seg/depth/RGB supervision) 3.5 5.7 9.6
GeoHead (uses GT pose labels) 2.86 5.60 7.39
ResNet-18 from scratch (0.68M) 2.33 6.31 7.65

The best result uses a completely frozen, stock DINOv3-S/16 with no object supervision of any kind β€” it beats a head trained with ground-truth pose labels.

offset is met on successful trajectories

The pooled number is dominated by failure rollouts (48% of the val set). Per split, one readout:

split needle tripod offset vs 4 mm
expert demos 2.15 0.95 2.50 met
policy rollouts (success) 2.78 1.59 3.61 met
policy rollouts (failure) 5.52 14.19 16.81 not met

The entire deficit is tripod localisation once the tripod has been displaced β€” at which point the threading attempt has already failed.

Two findings that dominate everything else

1. The readout, not the backbone. On identical frozen features, a linear ridge reads 15.6 mm and a learned soft-argmax readout reads 2.1 mm β€” a 7.4Γ— gap. Encoder choice is irrelevant by comparison: DINOv2-small 3.4, DINOv3-S/16 3.3, VGGT-1B 4.5 mm (matched protocol, 57Γ— parameter range). Finetuning the backbone makes things worse (~50%, with feature-spread collapse).

2. A starved readout manufactures ties. At n=9,600 / 4,000 steps every representation reads ~3.3 mm and looks equivalent. At n=60,000 / 25,000 steps:

representation values/sample starved saturated gain
raw frozen DINOv3 602,112 3.3 2.1 36%
SegHead sup224 200,704 3.6 3.5 3%

SegHead saturates; raw DINOv3 does not. SegHead compresses 602k values into 200k against a seg/depth/RGB objective and discards what a strong readout needs. Its geometry is fully linearly accessible (ridge 2.9 vs head 3.5 β€” no gap) and therefore capped; DINOv3's is richer but non-linear.

Always verify the probe is saturated before concluding two representations are equivalent.

SegHead is a trade, not an upgrade

goal use numbers
best pose raw frozen DINOv3 2.1 vs SegHead 3.5
a linear readout SegHead 2.9 vs raw 15.6
a world model SegHead 0.207 vs raw 0.239

SegHead is supervised only on {segmentation, depth, RGB} + proprioception β€” no object pose β€” so it is trainable from signals available on real hardware.

World model

Action-conditioned dynamics over the frozen feature map, 8-step horizon, expert split:

features imagined copy baseline readout floor ratio
SegHead sup224 4.82 mm 23.03 1.35 0.207
SegHead sup28 4.85 mm 22.93 1.28 0.213
raw frozen DINOv3 5.67 mm 23.81 2.70 0.239

4.8Γ— better than assuming nothing moves. For reference, a DINOv2 DINO-WM reaches 0.45, and a world model on pose-supervised GeoHead features reaches 43Γ— worse than copy β€” its features had collapsed to effective rank 2.4 with 90% dead channels.

Contents

  • readout_raw_dinov3/ β€” the 2.1 mm readout (the headline result). Backbone is stock frozen facebook/dinov3-vits16-pretrain-lvd1689m; only this readout is trained.
  • seghead_sup224/ β€” SegHead, label-free dense supervision at native 224Β² resolution
  • wm_sup224/ β€” world model on those features (best dynamics result) + its frozen measuring readout
  • measurements/ β€” probe logs

Reproducing

Feed both cameras at 224 px through frozen DINOv3-S/16, take hidden_states[3,6,9,12], drop the first 5 tokens (CLS + 4 registers) to get a 14Γ—14Γ—384 grid per view, then apply the readout: 1Γ—1 conv β†’ 16 heatmaps β†’ spatial softmax β†’ soft-argmax β†’ MLP with z-scored proprioception. Sub-cell precision comes from the soft-argmax expectation: the grid is 45 mm per token and the readout resolves 2.1 mm, ~21Γ— sub-token.

Caveats

  • Single seed per configuration; measured spread across 3 seeds is Οƒ β‰ˆ 0.15 mm.
  • Only an 8-step horizon was measured for the world model; MPC would use shorter horizons.
  • One scene, fixed cameras, one task β€” this says nothing about transfer, which is where a foundation model would be expected to earn its keep.
Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading