SigLIP2 Vision Tower β€” Karume

What is this

An image embedding distribution: the vision tower of SigLIP2, converted into the WebGPU inference runtime Karume's container format (a .krm part sequence whose first part carries the graph and model descriptors). Runs as-is in the browser and in Deno.

  • One graph, one call: pixels in, pooler_output out β€” the pooled [1, 768] vector from the attention-pooling (MAP) head.
  • Preprocessing is included. The pipeline resizes to 224 Γ— 224, rescales and normalizes with the constants below, so callers hand over raw RGB8 pixels. Decoding PNG / JPEG is not part of this β€” use createImageBitmap in the browser, or any decoder in Deno.
  • The text tower is not here. Zero-shot classification and image/text similarity need it (plus the logit scale and bias that go with it), so this repository cannot do them; what it does is turn an image into a vector.
  • Not readable by transformers (it's a different container with an embedded graph); the reader is a pipeline that implements siglip2/1.
  • Exporter used for the conversion: karume/0.13.0. The distribution manifest is karume.json (karume/5).

Base weights and attribution

Converted into the container format β€” the original checkpoints are not distributed here.

  • base: google/siglip2-base-patch16-224, licensed apache-2.0 (as of retrieval; full text β€” a verbatim copy is in LICENSE.md).
  • so400m: google/siglip2-so400m-patch14-384, licensed apache-2.0 (as of retrieval; full text β€” a verbatim copy is in LICENSE.md).
  • Changes made here (also listed in NOTICE.md, per Apache 2.0 Β§4(b)): conversion into the Karume container format, vision tower only. No retraining, no fine-tuning and no quantization β€” the weights are the source checkpoint's own f32 values. Two rewrites were needed to export the graph: the patch embedding's padding and the position embedding lookup were folded into equivalent operations (bit-exact), and the pooling head's attention was rewritten with explicit q/k/v projections, which is equivalent up to floating-point rounding (7.75e-07 to 2.38e-06 measured on the pooled vector, whose L2 norm is around 13).

Models

Model Pipeline Quants Default quant
base (default) siglip2/1 f32 f32
so400m siglip2/1 f32 f32

model selects one of these; omitted, it is base. quant defaults to that model's own default quant.

Usage

import { Siglip2Pipeline } from "jsr:@karume/models";

await using pipeline = await Siglip2Pipeline.fromPretrained({
  repo: "hdae/karume-siglip2",
  // Pin a commit for reproducible builds β€” without it you track `main`, and a future
  // repo update (renamed files, new manifest format) may break your app.
  // Copy the full hash from this repo's "Files and versions" tab:
  // revision: "<full commit sha>",
}, {
  // model: "base", // default β€” available: base / so400m
  // quant: "f32", // default β€” available: f32
});

// RGB8, row-major, 3 bytes per pixel. Decoding is the caller's job.
const embedding = await pipeline.embed({ data: pixels, width, height });

// pooler_output is not L2-normalized β€” normalize it yourself for cosine similarity:
const norm = Math.sqrt(embedding.reduce((sum, value) => sum + value * value, 0));
const unit = embedding.map((value) => value / norm);

embed() returns the pooled vector as an f32 array, exactly as the graph produces it. It keeps one GPU session alive for the lifetime of the pipeline, so embedding many images uploads the weights once; concurrent calls are queued rather than run side by side. Weights are fetched once and cached (verified against karume.json's size / sha256).

Model: base

Quants

Quant What it is Download Weights Compute
f32 (default) β€” 354 MiB vision = f32 β€”

If no quant is given, it runs as f32 (this model's recommended default). Per-file size and sha256 live in karume.json β€” verify against that at the fetch layer. Dtype labels use the runtime's storage dtype vocabulary (f16 / i8 / i4 / i2), not the fp16 spelling common elsewhere in the ecosystem. Weights ship as Karume container files (.krm), split into numbered parts; a part is fetched and verified on its own.

Input and output

Derived from the checkpoint's own preprocessor_config.json and the exported graph, and checked against each other when this repository was assembled.

  • input: RGB8 pixels, resized to 224 Γ— 224 (bilinear, antialiased). The aspect ratio is not preserved β€” there is no crop and no padding, matching the upstream processor.
  • normalization: (pixel / 255 - mean) / std, mean 0.5 / 0.5 / 0.5, std 0.5 / 0.5 / 0.5
  • output: pooler_output, 768 f32 values, not L2-normalized

Model: so400m

Quants

Quant What it is Download Weights Compute
f32 (default) β€” 1.60 GiB vision = f32 β€”

If no quant is given, it runs as f32 (this model's recommended default). Per-file size and sha256 live in karume.json β€” verify against that at the fetch layer. Dtype labels use the runtime's storage dtype vocabulary (f16 / i8 / i4 / i2), not the fp16 spelling common elsewhere in the ecosystem. Weights ship as Karume container files (.krm), split into numbered parts; a part is fetched and verified on its own.

Input and output

Derived from the checkpoint's own preprocessor_config.json and the exported graph, and checked against each other when this repository was assembled.

  • input: RGB8 pixels, resized to 384 Γ— 384 (bilinear, antialiased). The aspect ratio is not preserved β€” there is no crop and no padding, matching the upstream processor.
  • normalization: (pixel / 255 - mean) / std, mean 0.5 / 0.5 / 0.5, std 0.5 / 0.5 / 0.5
  • output: pooler_output, 1152 f32 values, not L2-normalized
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for hdae/karume-siglip2

Finetuned
(135)
this model