SigLIP2 Vision Tower β Karume
What is this
An image embedding distribution: the vision tower of SigLIP2, converted into the
WebGPU inference runtime Karume's container format (a .krm part sequence whose
first part carries the graph and model descriptors). Runs as-is in the browser and in
Deno.
- One graph, one call: pixels in,
pooler_outputout β the pooled[1, 768]vector from the attention-pooling (MAP) head. - Preprocessing is included. The pipeline resizes to 224 Γ 224, rescales and normalizes with the
constants below, so callers hand over raw RGB8 pixels. Decoding PNG / JPEG is not
part of this β use
createImageBitmapin the browser, or any decoder in Deno. - The text tower is not here. Zero-shot classification and image/text similarity need it (plus the logit scale and bias that go with it), so this repository cannot do them; what it does is turn an image into a vector.
- Not readable by transformers (it's a different container with an embedded graph); the reader is a pipeline that implements
siglip2/1. - Exporter used for the conversion:
karume/0.13.0. The distribution manifest iskarume.json(karume/5).
Base weights and attribution
Converted into the container format β the original checkpoints are not distributed here.
base: google/siglip2-base-patch16-224, licensed apache-2.0 (as of retrieval; full text β a verbatim copy is inLICENSE.md).so400m: google/siglip2-so400m-patch14-384, licensed apache-2.0 (as of retrieval; full text β a verbatim copy is inLICENSE.md).- Changes made here (also listed in
NOTICE.md, per Apache 2.0 Β§4(b)): conversion into the Karume container format, vision tower only. No retraining, no fine-tuning and no quantization β the weights are the source checkpoint's own f32 values. Two rewrites were needed to export the graph: the patch embedding's padding and the position embedding lookup were folded into equivalent operations (bit-exact), and the pooling head's attention was rewritten with explicit q/k/v projections, which is equivalent up to floating-point rounding (7.75e-07 to 2.38e-06 measured on the pooled vector, whose L2 norm is around 13).
Models
| Model | Pipeline | Quants | Default quant |
|---|---|---|---|
base (default) |
siglip2/1 |
f32 |
f32 |
so400m |
siglip2/1 |
f32 |
f32 |
model selects one of these; omitted, it is base. quant defaults to that model's own default quant.
Usage
import { Siglip2Pipeline } from "jsr:@karume/models";
await using pipeline = await Siglip2Pipeline.fromPretrained({
repo: "hdae/karume-siglip2",
// Pin a commit for reproducible builds β without it you track `main`, and a future
// repo update (renamed files, new manifest format) may break your app.
// Copy the full hash from this repo's "Files and versions" tab:
// revision: "<full commit sha>",
}, {
// model: "base", // default β available: base / so400m
// quant: "f32", // default β available: f32
});
// RGB8, row-major, 3 bytes per pixel. Decoding is the caller's job.
const embedding = await pipeline.embed({ data: pixels, width, height });
// pooler_output is not L2-normalized β normalize it yourself for cosine similarity:
const norm = Math.sqrt(embedding.reduce((sum, value) => sum + value * value, 0));
const unit = embedding.map((value) => value / norm);
embed() returns the pooled vector as an f32 array, exactly as the graph produces it.
It keeps one GPU session alive for the lifetime of the pipeline, so embedding many images
uploads the weights once; concurrent calls are queued rather than run side by side.
Weights are fetched once and cached (verified against karume.json's size / sha256).
Model: base
Quants
| Quant | What it is | Download | Weights | Compute |
|---|---|---|---|---|
f32 (default) |
β | 354 MiB | vision = f32 |
β |
If no quant is given, it runs as f32 (this model's recommended default).
Per-file size and sha256 live in karume.json β verify against that at the fetch layer.
Dtype labels use the runtime's storage dtype vocabulary (f16 / i8 / i4 / i2), not the fp16 spelling common elsewhere in the ecosystem.
Weights ship as Karume container files (.krm), split into numbered parts; a part is fetched and verified on its own.
Input and output
Derived from the checkpoint's own preprocessor_config.json and the exported graph, and
checked against each other when this repository was assembled.
- input: RGB8 pixels, resized to 224 Γ 224 (bilinear, antialiased). The aspect ratio is not preserved β there is no crop and no padding, matching the upstream processor.
- normalization:
(pixel / 255 - mean) / std, mean 0.5 / 0.5 / 0.5, std 0.5 / 0.5 / 0.5 - output:
pooler_output, 768 f32 values, not L2-normalized
Model: so400m
Quants
| Quant | What it is | Download | Weights | Compute |
|---|---|---|---|---|
f32 (default) |
β | 1.60 GiB | vision = f32 |
β |
If no quant is given, it runs as f32 (this model's recommended default).
Per-file size and sha256 live in karume.json β verify against that at the fetch layer.
Dtype labels use the runtime's storage dtype vocabulary (f16 / i8 / i4 / i2), not the fp16 spelling common elsewhere in the ecosystem.
Weights ship as Karume container files (.krm), split into numbered parts; a part is fetched and verified on its own.
Input and output
Derived from the checkpoint's own preprocessor_config.json and the exported graph, and
checked against each other when this repository was assembled.
- input: RGB8 pixels, resized to 384 Γ 384 (bilinear, antialiased). The aspect ratio is not preserved β there is no crop and no padding, matching the upstream processor.
- normalization:
(pixel / 255 - mean) / std, mean 0.5 / 0.5 / 0.5, std 0.5 / 0.5 / 0.5 - output:
pooler_output, 1152 f32 values, not L2-normalized
Model tree for hdae/karume-siglip2
Base model
google/siglip2-base-patch16-224