Argus-Lite
Argus-Lite is Argus on a frozen
EUPE-ViT-S backbone (Zhu et al.,
Meta FAIR, arXiv:2603.22387): 38M
parameters in total, of which 16.5M are the heads. It shares Argus's code and
API; config.json selects the backbone width and the head variants.
| Task | Head | Output | Trained on | Result |
|---|---|---|---|---|
| Classification | Linear softmax on the L2-normalized CLS token | 1,000 ImageNet classes | ImageNet-1k | 79.13% top-1, 95.53% top-5 (val) |
| Segmentation | BatchNorm and 1×1 conv | 150 ADE20K classes | ADE20K | 41.9 mIoU (val) |
| Depth | DPT decoder on blocks 2, 5, 8 and 11, 256 log-spaced bins over 0.001–10 m | Metric depth | NYU Depth V2 | RMSE 0.321, abs rel 0.103, δ<1.25 0.898 (654-image Eigen test split, Eigen crop, evaluate.py) |
| Detection | Cofiber pyramid with a text-aligned cosine classifier, FCOS decoding | 80 COCO classes | COCO 2017 | 27.3 mAP, 49.6 AP@50 (val2017) |
| Correspondence | None: cosine matching of patch features | Patch matches | — | — |
Class-agnostic AR@100 across 20 RF100-VL domains is 0.266, with the
per-domain results in rf100vl_results.json; cls_val_eval.json and
coco_val_eval.json hold the classification and detection breakdowns.
Usage
from PIL import Image
from transformers import AutoModel
model = AutoModel.from_pretrained("phanerozoic/argus-lite", trust_remote_code=True)
image = Image.open("image.jpg").convert("RGB")
model.classify(image, top_k=5) # [{"class_id", "class_name", "score"}, ...]
model.segment(image) # [H, W] ADE20K class indices
model.depth(image) # [H, W] depth in meters
model.detect(image, score_thresh=0.3) # [{"box", "score", "label", "class_name"}, ...]
model.embed(image) # [384] L2-normalized CLS embedding
model.perceive(image) # classification, segmentation and depth
model.correspond(image, other_image) # {"matches", "scores", "grid"}
Each method also takes a list of images and returns a list of results. Images
are resized to the resolution of the head: 224 for classification and
embedding, 512 for segmentation, 416 for depth and 768 for detection.
Segmentation and depth maps come back at each image's own height and width,
and boxes in its pixel coordinates. segment() and depth() return
per-pixel confidence and depth standard deviation with
return_confidence=True. The depth head is trained on indoor scenes and
predicts at most 10 m. quantize_int8(), compile(), export_onnx() and
evaluate.py work as in Argus.
License
The EUPE-ViT-S weights are released under the FAIR Noncommercial Research License. The checkpoint contains them and is released under the same license.
Citation
@misc{zhu2026eupe,
title={Efficient Universal Perception Encoder},
author={Zhu, Chenchen and Suri, Saksham and Jose, Cijo and Oquab, Maxime and Szafraniec, Marc and Wen, Wei and Xiong, Yunyang and Labatut, Patrick and Bojanowski, Piotr and Krishnamoorthi, Raghuraman and Chandra, Vikas},
year={2026},
eprint={2603.22387},
archivePrefix={arXiv},
primaryClass={cs.CV}
}
- Downloads last month
- 21
Model tree for phanerozoic/argus-lite
Base model
facebook/EUPE-ViT-S