PyTorch
llama

siglip2_logo

SigLIP 2 is a multilingual vision–language encoder that extends the original SigLIP training objective with captioning, self-distillation, masked prediction, and improved data curation for stronger semantic understanding, localization, and dense visual representations.

Original paper: SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features, Michael Tschannen, et al., 2025

SigLIP2-Base-Patch16-224

This model uses the SigLIP 2 Base-Patch16-224x224 variant, based on a ViT-Base vision encoder with 16×16 image patches and approximately 86M parameters. It is well suited for zero-shot image classification, image–text retrieval, and as a vision encoder for VLMs and downstream vision tasks.

Model Configuration:

Model Device Model Link
SigLIP2-Base-Patch16 Image Encoder N1-655 Model_Link
SigLIP2-Base-Patch16 Text Encoder N1-655 Model_Link
SigLIP2-Base-Patch16 Image Encoder X7 Model_Link
SigLIP2-Base-Patch16 Text Encoder X7 Model_Link
SigLIP2-Base-Patch16 Image Encoder CV7 Model_Link
SigLIP2-Base-Patch16 Text Encoder CV7 Model_Link
SigLIP2-Base-Patch16 Image Encoder CV72 Model_Link
SigLIP2-Base-Patch16 Text Encoder CV72 Model_Link
SigLIP2-Base-Patch16 Image Encoder CV75 Model_Link
SigLIP2-Base-Patch16 Text Encoder CV75 Model_Link
Downloads last month
50
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for Ambarella/SigLIP2