Espresso-3D — unofficial re-implementation weights
Weights for an unofficial re-implementation of Espresso-3D from Affogato: Open-Vocabulary Affordance Grounding with Automated Data Generation at Scale (Lee et al.). Code: github.com/LHC0312/espresso-3d-reimplementation.
Not affiliated with the Affogato authors and not the official Espresso-3D weights. The official code and weights were not public when this was trained (Sep 2026); the model was rebuilt from the paper.
Given a 3D point cloud and a natural-language query ("Point to the part you would hold to carry it."), the model outputs a per-point affordance heatmap in [0, 1].
Files
| File | Training |
|---|---|
espresso3d_affogato_pretrain.safetensors |
Affogato-750K pretraining, 185 min on one RTX PRO 6000 (3.2 epochs, 29,689 steps) |
espresso3d_laso_seen_best.safetensors |
+ LASO Seen finetuning, best LASO-val aIoU (epoch 13) |
espresso3d_laso_seen_final.safetensors |
+ LASO Seen finetuning, last (epoch 34) |
espresso3d_laso_unseen_best.safetensors |
+ LASO Unseen finetuning, best LASO-val aIoU (epoch 9) |
espresso3d_laso_unseen_final.safetensors |
+ LASO Unseen finetuning, last (epoch 44) |
Each file is fp16 (341M parameters, 683 MB): the heatmap decoder (7.5M) and the finetuned Recap-CLIP text tower. The frozen PartField encoder is not included; the code downloads the official checkpoint from mikaelaangel/partfield-ckpt.
Usage
git clone https://github.com/LHC0312/espresso-3d-reimplementation && cd espresso-3d-reimplementation
pip install -r requirements.txt
python inference.py --points object.npy --query "Point to the part you would press to turn it on." \
--weights LHC0312/espresso-3d-reimplementation:espresso3d_laso_seen_best.safetensors
from espresso import load_espresso
model = load_espresso("LHC0312/espresso-3d-reimplementation:espresso3d_affogato_pretrain.safetensors").cuda().eval()
Results (LASO test, 2,416 pairs)
| aIoU ↑ | AUC ↑ | SIM ↑ | MAE ↓ | |
|---|---|---|---|---|
| Seen · paper (Espresso-3D + Affogato pretrain) | 21.9 | 85.9 | 63.7 | 0.116 |
Seen · laso_seen_best |
19.5 | 84.1 | 60.2 | 0.099 |
Seen · laso_seen_final |
18.9 | 82.7 | 59.0 | 0.104 |
| Unseen · paper (Espresso-3D + Affogato pretrain) | 20.8 | 82.9 | 61.4 | 0.122 |
Unseen · laso_unseen_best |
17.6 | 81.7 | 59.1 | 0.104 |
Unseen · laso_unseen_final |
16.2 | 77.7 | 54.0 | 0.116 |
Zero-shot · affogato_pretrain (no LASO training) |
7.2 | 63.4 | 41.5 | 0.157 |
On our own Affogato validation split (1,050 objects / 5,250 pairs), affogato_pretrain reaches aIoU 13.98,
AUC 77.0, SIM 48.4, MAE 0.116. The LASO Unseen training split is rebuilt from LASO supplementary Table 6
(11,558 training pairs, matching the paper). See the GitHub README for the full protocol.
Limitations
- Pretrained for 3.2 epochs versus 50 in the paper; LASO scores are 2–3 aIoU points below the paper.
- Decoder depth, width and projections are not specified in the paper and were chosen for this implementation.
- Longer LASO finetuning overfits the seen (object, affordance) pairs: on held-out Unseen pairs,
laso_unseen_bestgets 12.1 aIoU andlaso_unseen_final7.0. - Trained on synthetic Objaverse objects (Affogato) and PartNet shapes (LASO); no scene-level context.
License
CC BY-NC 4.0. The weights contain a finetuned copy of the Recap-CLIP text tower (UCSC-VLAA, CC BY 4.0) and require the PartField encoder, whose code and weights are for non-commercial research and educational use only (NVIDIA License). The training data are Affogato-750K and LASO; follow their terms.
Citation
@article{lee2025affogato,
title = {Affogato: Open-Vocabulary Affordance Grounding with Automated Data Generation at Scale},
author = {Lee, Junha and Park, Eunha and Park, Chunghyun and Kang, Dahyun and Cho, Minsu},
journal = {arXiv preprint arXiv:2506.12009},
year = {2025}
}
Model tree for LHC0312/espresso-3d-reimplementation
Base model
UCSC-VLAA/ViT-L-16-HTxt-Recap-CLIP