|
Download README.md from Ambarella/SigLIP2: direct link, hf CLI and curl.
- Browser
- Download file 2.95 kB
-
https://huggingface.co/Ambarella/SigLIP2/resolve/main/README.md
- Command line
-
hf download hf://Ambarella/SigLIP2/README.md
-
curl -L -o README.md https://huggingface.co/Ambarella/SigLIP2/resolve/main/README.md
2.95 kB
| library_name: pytorch | |
|  | |
| **SigLIP 2** is a multilingual vision–language encoder that extends the original SigLIP training objective with captioning, self-distillation, masked prediction, and improved data curation for stronger semantic understanding, localization, and dense visual representations. | |
| Original paper: [SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features, Michael Tschannen, et al., 2025](https://arxiv.org/abs/2502.14786) | |
| # SigLIP2-Base-Patch16-224 | |
| This model uses the **SigLIP 2 Base-Patch16-224x224** variant, based on a ViT-Base vision encoder with **16×16 image patches** and approximately **86M parameters**. It is well suited for zero-shot image classification, image–text retrieval, and as a vision encoder for VLMs and downstream vision tasks. | |
| Model Configuration: | |
| - Reference implementation: [google/siglip2-base-patch16-224](https://huggingface.co/google/siglip2-base-patch16-224) | |
| - Original Weight: [SigLIP2-Base-Patch16](https://huggingface.co/google/siglip2-base-patch16-224) | |
| - Dataset: [ImageNet](https://image-net.org) | |
| - Resolution: 3x224x224 | |
| - Support Cooper version: | |
| - Cooper SDK: [2.5.4] | |
| - Cooper Foundry: [2.3] | |
| | Model | Device | Model Link | | |
| | :-----: | :-----: | :-----: | | |
| | SigLIP2-Base-Patch16 Image Encoder | N1-655 | [Model_Link](https://huggingface.co/Ambarella/SigLIP2/blob/main/n1-655_siglip2_base_patch16_image_encoder_act16.bin) | | |
| | SigLIP2-Base-Patch16 Text Encoder | N1-655 | [Model_Link](https://huggingface.co/Ambarella/SigLIP2/blob/main/n1-655_siglip2_base_patch16_text_encoder_act16.bin) | | |
| | SigLIP2-Base-Patch16 Image Encoder | X7 | [Model_Link](https://huggingface.co/Ambarella/SigLIP2/blob/main/x7_siglip2_base_patch16_image_encoder_act16.bin) | | |
| | SigLIP2-Base-Patch16 Text Encoder | X7 | [Model_Link](https://huggingface.co/Ambarella/SigLIP2/blob/main/x7_siglip2_base_patch16_text_encoder_act16.bin) | | |
| | SigLIP2-Base-Patch16 Image Encoder | CV7 | [Model_Link](https://huggingface.co/Ambarella/SigLIP2/blob/main/cv7_siglip2_base_patch16_image_encoder_act16.bin) | | |
| | SigLIP2-Base-Patch16 Text Encoder | CV7 | [Model_Link](https://huggingface.co/Ambarella/SigLIP2/blob/main/cv7_siglip2_base_patch16_text_encoder_act16.bin) | | |
| | SigLIP2-Base-Patch16 Image Encoder | CV72 | [Model_Link](https://huggingface.co/Ambarella/SigLIP2/blob/main/cv72_siglip2_base_patch16_image_encoder_act16.bin) | | |
| | SigLIP2-Base-Patch16 Text Encoder | CV72 | [Model_Link](https://huggingface.co/Ambarella/SigLIP2/blob/main/cv72_siglip2_base_patch16_text_encoder_act16.bin) | | |
| | SigLIP2-Base-Patch16 Image Encoder | CV75 | [Model_Link](https://huggingface.co/Ambarella/SigLIP2/blob/main/cv75_siglip2_base_patch16_image_encoder_act16.bin) | | |
| | SigLIP2-Base-Patch16 Text Encoder | CV75 | [Model_Link](https://huggingface.co/Ambarella/SigLIP2/blob/main/cv75_siglip2_base_patch16_text_encoder_act16.bin) | | |