Argus-Lite

Argus-Lite is Argus on a frozen EUPE-ViT-S backbone (Zhu et al., Meta FAIR, arXiv:2603.22387): 38M parameters in total, of which 16.5M are the heads. It shares Argus's code and API; config.json selects the backbone width and the head variants.

Task Head Output Trained on Result
Classification Linear softmax on the L2-normalized CLS token 1,000 ImageNet classes ImageNet-1k 79.13% top-1, 95.53% top-5 (val)
Segmentation BatchNorm and 1×1 conv 150 ADE20K classes ADE20K 41.9 mIoU (val)
Depth DPT decoder on blocks 2, 5, 8 and 11, 256 log-spaced bins over 0.001–10 m Metric depth NYU Depth V2 RMSE 0.321, abs rel 0.103, δ<1.25 0.898 (654-image Eigen test split, Eigen crop, evaluate.py)
Detection Cofiber pyramid with a text-aligned cosine classifier, FCOS decoding 80 COCO classes COCO 2017 27.3 mAP, 49.6 AP@50 (val2017)
Correspondence None: cosine matching of patch features Patch matches — —

Class-agnostic AR@100 across 20 RF100-VL domains is 0.266, with the per-domain results in rf100vl_results.json; cls_val_eval.json and coco_val_eval.json hold the classification and detection breakdowns.

Usage

from PIL import Image
from transformers import AutoModel

model = AutoModel.from_pretrained("phanerozoic/argus-lite", trust_remote_code=True)
image = Image.open("image.jpg").convert("RGB")

model.classify(image, top_k=5)          # [{"class_id", "class_name", "score"}, ...]
model.segment(image)                    # [H, W] ADE20K class indices
model.depth(image)                      # [H, W] depth in meters
model.detect(image, score_thresh=0.3)   # [{"box", "score", "label", "class_name"}, ...]
model.embed(image)                      # [384] L2-normalized CLS embedding
model.perceive(image)                   # classification, segmentation and depth
model.correspond(image, other_image)    # {"matches", "scores", "grid"}

Each method also takes a list of images and returns a list of results. Images are resized to the resolution of the head: 224 for classification and embedding, 512 for segmentation, 416 for depth and 768 for detection. Segmentation and depth maps come back at each image's own height and width, and boxes in its pixel coordinates. segment() and depth() return per-pixel confidence and depth standard deviation with return_confidence=True. The depth head is trained on indoor scenes and predicts at most 10 m. quantize_int8(), compile(), export_onnx() and evaluate.py work as in Argus.

License

The EUPE-ViT-S weights are released under the FAIR Noncommercial Research License. The checkpoint contains them and is released under the same license.

Citation

@misc{zhu2026eupe,
  title={Efficient Universal Perception Encoder},
  author={Zhu, Chenchen and Suri, Saksham and Jose, Cijo and Oquab, Maxime and Szafraniec, Marc and Wen, Wei and Xiong, Yunyang and Labatut, Patrick and Bojanowski, Piotr and Krishnamoorthi, Raghuraman and Chandra, Vikas},
  year={2026},
  eprint={2603.22387},
  archivePrefix={arXiv},
  primaryClass={cs.CV}
}
Downloads last month
21
Safetensors
Model size
38.1M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for phanerozoic/argus-lite

Finetuned
(1)
this model

Datasets used to train phanerozoic/argus-lite

Space using phanerozoic/argus-lite 1

Paper for phanerozoic/argus-lite