SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
Paper • 2502.14786 • Published • 169
SigLIP 2 is a multilingual vision–language encoder that extends the original SigLIP training objective with captioning, self-distillation, masked prediction, and improved data curation for stronger semantic understanding, localization, and dense visual representations.
This model uses the SigLIP 2 Base-Patch16-224x224 variant, based on a ViT-Base vision encoder with 16×16 image patches and approximately 86M parameters. It is well suited for zero-shot image classification, image–text retrieval, and as a vision encoder for VLMs and downstream vision tasks.
Model Configuration:
| Model | Device | Model Link |
|---|---|---|
| SigLIP2-Base-Patch16 Image Encoder | N1-655 | Model_Link |
| SigLIP2-Base-Patch16 Text Encoder | N1-655 | Model_Link |
| SigLIP2-Base-Patch16 Image Encoder | X7 | Model_Link |
| SigLIP2-Base-Patch16 Text Encoder | X7 | Model_Link |
| SigLIP2-Base-Patch16 Image Encoder | CV7 | Model_Link |
| SigLIP2-Base-Patch16 Text Encoder | CV7 | Model_Link |
| SigLIP2-Base-Patch16 Image Encoder | CV72 | Model_Link |
| SigLIP2-Base-Patch16 Text Encoder | CV72 | Model_Link |
| SigLIP2-Base-Patch16 Image Encoder | CV75 | Model_Link |
| SigLIP2-Base-Patch16 Text Encoder | CV75 | Model_Link |