siglip2-base-patch16-384 text model — ONNX
Only the text model of upshift/siglip2-base-patch16-384-onnx, the ONNX export of
google/siglip2-base-patch16-384 (revision f775b65, opset 18), for when
you only need text embeddings, for example to embed search queries against image embeddings computed
elsewhere. The files are identical to the ones in the full repo, which also has the vision model.
| File | Precision | Size | Inputs → output |
|---|---|---|---|
text_model.onnx + text_model.onnx_data |
fp32 | 1.13 GB | input_ids int64 (B, 64) → text_embeds float32 (B, 768) |
text_model_fp16.onnx + text_model_fp16.onnx_data |
fp16 | 565 MB | same as fp32 |
B (batch) is dynamic. Embeddings are not normalised. config.json and the tokenizer files are copied
from the original repo; export_meta.json holds logit_scale (already exponentiated) and logit_bias
for scoring against image embeddings.
Parity
Compared with the PyTorch text model on onnxruntime (CPU):
| Precision | Inputs | Min cosine | Max abs diff | Threshold | Result |
|---|---|---|---|---|---|
| fp32 | 6 | 1.0000000 | 1.0e-05 | 0.9999 | passed |
| fp16 | 6 | 0.9999989 | 1.7e-02 | 0.999 | passed |
Inputs
Lowercase the text, tokenise with the included Gemma tokenizer.json, append <eos> and pad with
<pad> (id 0) to exactly 64 tokens. The model pools the last token, so the padding is required. In
Python, use Siglip2Tokenizer (transformers 5.x AutoTokenizer returns a GemmaTokenizer that does not
lowercase).
Scoring against an image embedding from the vision model of
upshift/siglip2-base-patch16-384-onnx:
p = sigmoid(logit_scale · cos(image, text) + logit_bias), per label.
Usage with onnxruntime-web
import * as ort from 'onnxruntime-web/webgpu';
const base = 'https://huggingface.co/upshift/siglip2-base-patch16-384-text-onnx/resolve/main/';
const text = await ort.InferenceSession.create(base + 'text_model.onnx', {
executionProviders: ['webgpu', 'wasm'],
externalData: [{ path: 'text_model.onnx_data', data: base + 'text_model.onnx_data' }],
});
const { text_embeds } = await text.run({
input_ids: new ort.Tensor('int64', inputIds, [1, 64]),
});
Limitations
- Image embeddings need the vision model from upshift/siglip2-base-patch16-384-onnx, from the same checkpoint.
License
Apache-2.0, as the original model. Original work by Google; see the SigLIP 2 paper and the original model card.
- Downloads last month
- -
Model tree for upshift/siglip2-base-patch16-384-text-onnx
Base model
google/siglip2-base-patch16-384