YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

FARM UF850 Patch Policy β€” task_2 "stack the three cubes" (L1 head)

Reimplementation of Patch Policy (arXiv 2607.18236) as a single-task specialist for the FARM UF850 arm: frozen DINOv2-S/14 dense patch tokens (all 256/camera, base + wrist, no pooling) -> 8-layer / d=512 transformer (26.8M trainable) -> L1-regressed 50-step action chunk. No language conditioning (single task). Trained on 84 episodes, 10 held out.

Training cost: ~0.5 GPU-hours (30k steps, batch 64, one H200, ~24 it/s β€” episode subset fully RAM-preloaded).

Results (offline, held-out episodes β€” triage proxy ONLY, no arm rollouts yet)

variant held-out chunk-MAE (rad) first-action MAE (rad)
L1 head (this ckpt) 0.095 0.033 (~1.9Β°/joint)
diffusion head (MLP denoiser) 0.319 0.220

Open-loop prediction accuracy is a screen, not a verdict β€” closed-loop rollouts on the arm are the only real test, and this model has not had them.

Inference

from infer_patch_policy import PatchPolicyRunner
r = PatchPolicyRunner("patch_policy_task2_l1.pt")          # downloads DINOv2-S via torch.hub
chunk = r.predict(base_rgb, wrist_rgb, state7)             # (50, 7)
  • Inputs: two RGB frames (any resolution; resized to 224x224) + 7-d proprio (6 joint positions rad + gripper). State is normalized internally (stats stored in ckpt).
  • Output: 50-step chunk in the FARM dataset action space β€” 6 relative joint-position deltas (rad) + 1 gripper position [0,1], recorded at 30 fps. Execute like the dataset semantics (apply deltas to current joint positions; gripper is absolute). Re-plan by re-querying at chunk end or earlier.
  • Deps: torch >= 2.x, opencv, numpy, internet on first run (torch.hub DINOv2 weights, ~84MB).

Files

  • patch_policy_task2_l1.pt β€” EMA weights (incl. frozen ViT), config, normalization stats.
  • train_patch_policy.py β€” full self-contained training script (data loader for lerobot-v2 FARM episodes, both heads, eval). Repro: see docstring; ~31 min on one H200.
  • infer_patch_policy.py β€” the runner above.

Caveats

Single-task specialist (no language, no multi-task); T=1 observation (no temporal context); diffusion variant underperforms here (simplified MLP denoiser, not the paper's DP head); offline-validated only.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support