--- license: apache-2.0 base_model: Qwen/Qwen3-Embedding-0.6B tags: - text-embeddings - sentence-similarity - retrieval - core-ai - apple-silicon - macos library_name: swift-text-embedder --- # embed-qwen3 — Qwen3-Embedding-0.6B for Core AI (macOS 27) Alibaba's [Qwen3-Embedding-0.6B](https://huggingface.co/Qwen/Qwen3-Embedding-0.6B) (Apache 2.0) exported to an Apple Core AI `.aimodel` for on-device text embeddings and semantic search. Read by [swift-text-embedder](https://github.com/arraypress/swift-text-embedder) and the [`embed`](https://github.com/arraypress/swift-embed-cli) command-line tool. ## Files | Path | What | |---|---| | `embed-qwen3-0.6b-float32.aimodel/` | Entry points `t128` and `t512`: token ids `[1, L]` + attention mask `[1, L]` (int32, LEFT-padded with `<|endoftext|>` = 151643) → the last position's hidden state `[1, 1024]`. Float32, 2.4 GB. | | `embed-support/vocab.json`, `merges.txt` | The checkpoint's own Qwen2 byte-level BPE files. | | `embed-support/config.json` | The pad / end-of-text id, the exported lengths, the dimension. | The host tokenises (NFC, Qwen2's pre-tokeniser, end-of-text appended), pads on the left, picks the shortest length that holds the text, and L2-normalises the output. Queries are embedded as `Instruct: \nQuery: `, documents as they are, per the model card. ## Fidelity Against `transformers` 5.17 on the same inputs: tokenizer ids and masks identical over 13 texts at both lengths (Unicode, code, newlines included); embeddings 125–126 dB PSNR, cosine ≥ 0.999999. The fixed-length left padding changes an embedding by at most 3e-7 against the unpadded model. ## Use ```sh hf download arraypress/embed-qwen3 --local-dir models embed model install models embed index ~/Notes && embed search "how do I rotate the API key" ``` Requires macOS 27 (Core AI) on Apple silicon. Converted with coreai-torch 0.4.2 / torch 2.13.