Instructions to use Nurymanau/EmbeddingGemma-2-MLX-Swift with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use Nurymanau/EmbeddingGemma-2-MLX-Swift with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download Nurymanau/EmbeddingGemma-2-MLX-Swift --local-dir EmbeddingGemma-2-MLX-Swift
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
EmbeddingGemma 2 — MLX Swift weights
Ready-to-use modular weights for EmbeddingGemma2Swift, a native Swift runtime built on MLX for Mac and iPhone. Google DeepMind created the original model. This release adapts and packages it; it is not a newly trained model, a chat model, or a new quantization algorithm.
Files
| Folder | Contents | Weight bytes |
|---|---|---|
text-q8 |
Shared text encoder/tokenizer, affine8-bit/group64 | 288118885 |
vision-q8 |
Image tower + bridge, affine8-bit/group64 | 193095171 |
audio-bf16 |
Optional unchanged audio tower + bridge, BF16 | 611314600 |
Text is required for all modalities. Text+vision weights total481214056bytes
(~481MB); tokenizer/config files add about32MB. Audio is an optional separate
download. Read exact file sizes and SHA256 in manifest.json.
Use the linked Swift runtime. The packed layout has versioned
quantization.json metadata and is not a drop-in Python mlx-vlm checkpoint.
FP16 activations are unsupported. The default working activation dtype is BF16;
output vectors are normalized Float32, with128/256/512/768 dimensions.
Quick start
git clone https://github.com/Obscyra-app/EmbeddingGemma2Swift.git
cd EmbeddingGemma2Swift
python3 scripts/download_release.py --output models/release
xcodebuild -scheme eg2-embed -configuration Release \
-destination 'platform=macOS,arch=arm64' -derivedDataPath .build-xcode build
Create request.json containing:
{"query":"How can I check a downloaded file?","documents":["Compare its SHA-256 checksum with the source.","Tomorrow will be sunny."]}
.build-xcode/Build/Products/Release/eg2-embed models/release/text-q8 request.json
The helper downloads a pinned revision and verifies SHA256. Add--audio to
download the audio tower. Swift inference runs locally without Python or
network requests. Downloading weights is an explicit separate setup step.
For Swift package integration, image/audio examples and the iOS demo, see the source README.
Provenance and validation
Original: google/embeddinggemma-2,
revision914f7f89142e33e77833254d9c9b90c3cef7303b.
Original safetensors SHA256197a32965d4b1105faf060417baa899e193fb73cd401f42ec9295234d5553d79.
MLX Swift0.32.3 and Swift Transformers1.3.0. No training or calibration was performed.
All compatible2D text/vision weight matrices were quantized with standard MLX;
norms/scalars/3Dvision position table remain unquantized.
Validated on MacBook Air M3/16GiB and physical iPhone16ProMax/iOS27.0. Independent CPUFP32 and NumPy-unpack audits check implementation separately from quantization loss.11 Swift tests and input validation checks passed. The selected q8 profile passed a fixed dev drift gate; q4 was rejected despite correct top1 on the simple dev set. Synthetic bilingual dev24/24 and separate holdout16/16 text matches were retained by q8. Three held-out synthetic images retained6/6 RU/ENtext→image matches. These are small synthetic checks, not broad real-world quality benchmarks. Speedup over BF16 has not been established.
See validation details and machine-readable summaries.
Limitations
- Audio is experimental for retrieval quality: a tiny Russian-description smoke test achieved only1/4 matches in each direction, despite numerical parity.
- Images are processed one at a time. Transparency is rejected; exact JPEG/HEIC/color-profile decoder parity has not been established.
- Audio input: mono16kHz PCM16 WAV or normalized Float samples, up to30seconds; full neural inference at30seconds remains unvalidated (frontend was checked).
- No video, interleaved multimodal documents, image generation, chat, ASR or TTS.
- Older devices/OS, sustained thermals, battery use and broad retrieval remain unmeasured.
Weights are Apache-2.0; seeLICENSEandNOTICE. Runtime source has its own MIT
license and bundled third-party notices. No affiliation, endorsement, or claim
of first MLX/iOS support is implied.
Quantized
Model tree for Nurymanau/EmbeddingGemma-2-MLX-Swift
Base model
google/embeddinggemma-2