Millimetre object pose from frozen DINOv3 β MimicGen threading_d0
Decode metric 3D object position from vision on MimicGen threading_d0, and predict it forward with
an action-conditioned world model. Task requirement for reliable insertion is ~4 mm; offset
(needle β tripod) is the decision metric, because relative pose is what the controller consumes.
All numbers are mean Euclidean distance in millimetres on held-out episodes β mean_i βp_i β t_iββ,
the average 3D miss distance, not RMSE. Camera geometry: 45Β° fovy, 0.762 m to the target plane,
2.82 mm per pixel at 224 px; one patch-16 token spans β45 mm.
Headline
| representation | needle | tripod | offset |
|---|---|---|---|
| frozen DINOv3-S/16 + soft-argmax readout | 2.1 | 5.5 | 7.5 |
SegHead sup224 (dense seg/depth/RGB supervision) |
3.5 | 5.7 | 9.6 |
| GeoHead (uses GT pose labels) | 2.86 | 5.60 | 7.39 |
| ResNet-18 from scratch (0.68M) | 2.33 | 6.31 | 7.65 |
The best result uses a completely frozen, stock DINOv3-S/16 with no object supervision of any kind β it beats a head trained with ground-truth pose labels.
offset is met on successful trajectories
The pooled number is dominated by failure rollouts (48% of the val set). Per split, one readout:
| split | needle | tripod | offset | vs 4 mm |
|---|---|---|---|---|
| expert demos | 2.15 | 0.95 | 2.50 | met |
| policy rollouts (success) | 2.78 | 1.59 | 3.61 | met |
| policy rollouts (failure) | 5.52 | 14.19 | 16.81 | not met |
The entire deficit is tripod localisation once the tripod has been displaced β at which point the threading attempt has already failed.
Two findings that dominate everything else
1. The readout, not the backbone. On identical frozen features, a linear ridge reads 15.6 mm and a learned soft-argmax readout reads 2.1 mm β a 7.4Γ gap. Encoder choice is irrelevant by comparison: DINOv2-small 3.4, DINOv3-S/16 3.3, VGGT-1B 4.5 mm (matched protocol, 57Γ parameter range). Finetuning the backbone makes things worse (~50%, with feature-spread collapse).
2. A starved readout manufactures ties. At n=9,600 / 4,000 steps every representation reads ~3.3 mm and looks equivalent. At n=60,000 / 25,000 steps:
| representation | values/sample | starved | saturated | gain |
|---|---|---|---|---|
| raw frozen DINOv3 | 602,112 | 3.3 | 2.1 | 36% |
SegHead sup224 |
200,704 | 3.6 | 3.5 | 3% |
SegHead saturates; raw DINOv3 does not. SegHead compresses 602k values into 200k against a seg/depth/RGB objective and discards what a strong readout needs. Its geometry is fully linearly accessible (ridge 2.9 vs head 3.5 β no gap) and therefore capped; DINOv3's is richer but non-linear.
Always verify the probe is saturated before concluding two representations are equivalent.
SegHead is a trade, not an upgrade
| goal | use | numbers |
|---|---|---|
| best pose | raw frozen DINOv3 | 2.1 vs SegHead 3.5 |
| a linear readout | SegHead | 2.9 vs raw 15.6 |
| a world model | SegHead | 0.207 vs raw 0.239 |
SegHead is supervised only on {segmentation, depth, RGB} + proprioception β no object pose β so it is trainable from signals available on real hardware.
World model
Action-conditioned dynamics over the frozen feature map, 8-step horizon, expert split:
| features | imagined | copy baseline | readout floor | ratio |
|---|---|---|---|---|
SegHead sup224 |
4.82 mm | 23.03 | 1.35 | 0.207 |
SegHead sup28 |
4.85 mm | 22.93 | 1.28 | 0.213 |
| raw frozen DINOv3 | 5.67 mm | 23.81 | 2.70 | 0.239 |
4.8Γ better than assuming nothing moves. For reference, a DINOv2 DINO-WM reaches 0.45, and a world model on pose-supervised GeoHead features reaches 43Γ worse than copy β its features had collapsed to effective rank 2.4 with 90% dead channels.
Contents
readout_raw_dinov3/β the 2.1 mm readout (the headline result). Backbone is stock frozenfacebook/dinov3-vits16-pretrain-lvd1689m; only this readout is trained.seghead_sup224/β SegHead, label-free dense supervision at native 224Β² resolutionwm_sup224/β world model on those features (best dynamics result) + its frozen measuring readoutmeasurements/β probe logs
Reproducing
Feed both cameras at 224 px through frozen DINOv3-S/16, take hidden_states[3,6,9,12], drop the
first 5 tokens (CLS + 4 registers) to get a 14Γ14Γ384 grid per view, then apply the readout:
1Γ1 conv β 16 heatmaps β spatial softmax β soft-argmax β MLP with z-scored proprioception.
Sub-cell precision comes from the soft-argmax expectation: the grid is 45 mm per token and the
readout resolves 2.1 mm, ~21Γ sub-token.
Caveats
- Single seed per configuration; measured spread across 3 seeds is Ο β 0.15 mm.
- Only an 8-step horizon was measured for the world model; MPC would use shorter horizons.
- One scene, fixed cameras, one task β this says nothing about transfer, which is where a foundation model would be expected to earn its keep.