Espresso-3D — unofficial re-implementation weights

Weights for an unofficial re-implementation of Espresso-3D from Affogato: Open-Vocabulary Affordance Grounding with Automated Data Generation at Scale (Lee et al.). Code: github.com/LHC0312/espresso-3d-reimplementation.

Not affiliated with the Affogato authors and not the official Espresso-3D weights. The official code and weights were not public when this was trained (Sep 2026); the model was rebuilt from the paper.

Given a 3D point cloud and a natural-language query ("Point to the part you would hold to carry it."), the model outputs a per-point affordance heatmap in [0, 1].

Files

File Training
espresso3d_affogato_pretrain.safetensors Affogato-750K pretraining, 185 min on one RTX PRO 6000 (3.2 epochs, 29,689 steps)
espresso3d_laso_seen_best.safetensors + LASO Seen finetuning, best LASO-val aIoU (epoch 13)
espresso3d_laso_seen_final.safetensors + LASO Seen finetuning, last (epoch 34)
espresso3d_laso_unseen_best.safetensors + LASO Unseen finetuning, best LASO-val aIoU (epoch 9)
espresso3d_laso_unseen_final.safetensors + LASO Unseen finetuning, last (epoch 44)

Each file is fp16 (341M parameters, 683 MB): the heatmap decoder (7.5M) and the finetuned Recap-CLIP text tower. The frozen PartField encoder is not included; the code downloads the official checkpoint from mikaelaangel/partfield-ckpt.

Usage

git clone https://github.com/LHC0312/espresso-3d-reimplementation && cd espresso-3d-reimplementation
pip install -r requirements.txt
python inference.py --points object.npy --query "Point to the part you would press to turn it on." \
  --weights LHC0312/espresso-3d-reimplementation:espresso3d_laso_seen_best.safetensors
from espresso import load_espresso
model = load_espresso("LHC0312/espresso-3d-reimplementation:espresso3d_affogato_pretrain.safetensors").cuda().eval()

Results (LASO test, 2,416 pairs)

aIoU ↑ AUC ↑ SIM ↑ MAE ↓
Seen · paper (Espresso-3D + Affogato pretrain) 21.9 85.9 63.7 0.116
Seen · laso_seen_best 19.5 84.1 60.2 0.099
Seen · laso_seen_final 18.9 82.7 59.0 0.104
Unseen · paper (Espresso-3D + Affogato pretrain) 20.8 82.9 61.4 0.122
Unseen · laso_unseen_best 17.6 81.7 59.1 0.104
Unseen · laso_unseen_final 16.2 77.7 54.0 0.116
Zero-shot · affogato_pretrain (no LASO training) 7.2 63.4 41.5 0.157

On our own Affogato validation split (1,050 objects / 5,250 pairs), affogato_pretrain reaches aIoU 13.98, AUC 77.0, SIM 48.4, MAE 0.116. The LASO Unseen training split is rebuilt from LASO supplementary Table 6 (11,558 training pairs, matching the paper). See the GitHub README for the full protocol.

Limitations

  • Pretrained for 3.2 epochs versus 50 in the paper; LASO scores are 2–3 aIoU points below the paper.
  • Decoder depth, width and projections are not specified in the paper and were chosen for this implementation.
  • Longer LASO finetuning overfits the seen (object, affordance) pairs: on held-out Unseen pairs, laso_unseen_best gets 12.1 aIoU and laso_unseen_final 7.0.
  • Trained on synthetic Objaverse objects (Affogato) and PartNet shapes (LASO); no scene-level context.

License

CC BY-NC 4.0. The weights contain a finetuned copy of the Recap-CLIP text tower (UCSC-VLAA, CC BY 4.0) and require the PartField encoder, whose code and weights are for non-commercial research and educational use only (NVIDIA License). The training data are Affogato-750K and LASO; follow their terms.

Citation

@article{lee2025affogato,
  title   = {Affogato: Open-Vocabulary Affordance Grounding with Automated Data Generation at Scale},
  author  = {Lee, Junha and Park, Eunha and Park, Chunghyun and Kang, Dahyun and Cho, Minsu},
  journal = {arXiv preprint arXiv:2506.12009},
  year    = {2025}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for LHC0312/espresso-3d-reimplementation

Finetuned
(1)
this model

Dataset used to train LHC0312/espresso-3d-reimplementation

Paper for LHC0312/espresso-3d-reimplementation