EmbeddingGemma 2 image and text embeddings on the AMD NPU

An IRON export of google/embeddinggemma-2 for AMD Ryzen AI NPUs: the compiled NPU kernels (.xclbin + instruction streams) and the packed weights the taconite-embeddinggemma2 Rust runtime replays, over XRT or directly over the amdxdna driver.

The kernels are compiled for NPU2 (AIE2P: Strix Point, Strix Halo, Krackan) and will not load on NPU1 (Phoenix, Hawk Point).

Download

hf download brishen/iron-embeddinggemma2-npu2 --local-dir embeddinggemma2
# or, from an IRON checkout:
python scripts/hf_models.py download embeddinggemma2 --repo brishen/iron-embeddinggemma2-npu2 --out embeddinggemma2

Images and texts (up to 2048 tokens, with the model's task prompts) in, the unit 768-d embeddings sentence-transformers computes out, in the model's one space; the transformers' projections and the vision attention run on the NPU (~1.1 s an image on a Ryzen AI 9 HX 370). Against float32 sentence-transformers: image preprocessing bit-exact and text tokenization id for id; embedding cosine >= 0.999 on the bundle's reference images and texts, which embeddinggemma2 check repeats. Text needs taconite-embeddinggemma2 >= 0.2.0 (0.1.0 runs this bundle's image path).

Usage

cargo install taconite-embeddinggemma2             # over XRT
# or, with no XRT at all (straight to the amdxdna driver):
cargo install taconite-embeddinggemma2 --no-default-features --features cli,direct
embeddinggemma2 embed embeddinggemma2 cats.jpg car.png [--dim 256] -o emb.f32
embeddinggemma2 embed embeddinggemma2 cats.jpg --text "two cats sleeping" --prompt SearchQuery
embeddinggemma2 check embeddinggemma2             # every stage against the float32 references

See taconite-embeddinggemma2 (docs, source) for the full API. The bundle is exported by IRON's iron/applications/embeddinggemma2.

Provenance

  • Upstream model: google/embeddinggemma-2 (license: apache-2.0; its terms apply to these weights)
  • IRON commit: 0229cc3
  • Uploaded: 2026-10-06
  • Files: 142, 1.0 GB
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for brishen/iron-embeddinggemma2-npu2

Finetuned
(34)
this model