SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
Paper • 2502.14786 • Published • 169
ONNX export of Google's SigLIP2 ViT-B/16 (WebLI, 224 px), from
timm/ViT-B-16-SigLIP2 at revision eee10eff6dd8cabae2d7f379d4e8cfcd352030aa, used by
Exposure to search photos by what they show. Apache 2.0, like the original weights.
visual/model.onnx: image float32 [1, 3, 224, 224], RGB squashed to 224×224 (bicubic), scaled to [-1, 1] → embedding [1, 768], L2-normalized.textual/model.onnx: text int32 [1, 64] → embedding [1, 768], L2-normalized. Lowercase the query, drop punctuation, collapse spaces,
tokenize with textual/tokenizer.json (appends <eos>), pad with <pad> to 64. The token embedding table is int8 (cosine to fp32 ≥ 0.999); everything else is fp32.Made with server/scripts/export-model.py.
| File | Bytes | SHA-256 |
|---|---|---|
visual/model.onnx |
371,674,653 | c0cd20ed79d53c5a3fa68fed5ff28b601c78bdde58dd70b24d0dd9595a2dc3dd |
textual/model.onnx |
539,694,825 | c856f2b3a5b71b04bf7cb68e93cb64914a79e47fec1e5648b4b4e2cad1226b36 |
textual/tokenizer.json |
34,362,885 | 220c63d496e0c14e63eb656c91e0215e926202e4c74b1f089e09f1920d779b04 |
textual/tokenizer_config.json |
46,382 | 513e778583a9ac0b74ae47784216f9d654d5132eee62d777597f5e842701323c |
Base model
timm/ViT-B-16-SigLIP2