UniPose9D: Universal Category-Agnostic Object Pose Estimation

UniPose9D is a category-agnostic foundation model for 9D object pose estimation. Given an RGB-D observation (or an RGB image with predicted depth) and an instance mask, it predicts rotation, translation and metric size without category labels, CAD models, mean-shape priors or reference views.

Files

File Description
last.ckpt UniPose9D checkpoint (PyTorch Lightning)
config.yaml Training configuration for the checkpoint

Usage

Download the checkpoint into the checkpoints/ folder of the code repository, then run inference:

pip install -U "huggingface_hub[cli]"
hf download qq456cvb/UniPose9D last.ckpt config.yaml --local-dir checkpoints

python infer/unipose9d_inference.py \
  --rgb examples/desktop_scene/example.jpg \
  --sam2-point 540 430 1

Masks come from SAM2 (facebook/sam2.1-hiera-large), and missing depth and intrinsics are estimated with MoGe (Ruicheng/moge-2-vitl-normal). Both are downloaded automatically. See the code repository for all options, including your own depth maps and camera intrinsics.

Citation

@article{you2026unipose9d,
  title={UniPose9D: Universal Category-Agnostic Object Pose Estimation},
  author={You, Yang and Du, Yi and Harrison, Cole and Guibas, Leonidas},
  journal={arXiv preprint arXiv:2607.09985},
  year={2026}
}
Downloads last month
9
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Space using qq456cvb/UniPose9D 1

Collection including qq456cvb/UniPose9D

Paper for qq456cvb/UniPose9D