qwen3-embedding-0.6b
A Core AI bundle of Qwen/Qwen3-Embedding-0.6B for the Aether SDK (iOS and macOS 27+).
- Source:
Qwen/Qwen3-Embedding-0.6Bat revision97b0c614be4d77ee51c0cef4e5f07c00f9eb65b3, licence Apache-2.0. The licence is included asLICENSE.Qwen/Qwen3-Embedding-0.6Bdeclares apache-2.0 in its model card but ships no licence file;LICENSEis the canonical text from apache.org. - Changes from the source: converted from PyTorch to Core AI (
.aimodel) by Aether forge (recipeqwen3-embedding-0.6b@2). Weights are int8-linear-perblock32 (8-bit weights). The tokenizer files are the source's own. - Output: one vector per text (
embeddings), for search and similarity; see Embeddings below.
Variants
| Variant | Platform | Arch | Compute | Compiled | Assets | Download |
|---|---|---|---|---|---|---|
macos-any-gpu |
macos | any | gpu | no (specialized on first load) | qwen3_embedding_0_6b.aimodel 633.5 MB |
649.3 MB |
ios-any-gpu |
ios | any | gpu | no (specialized on first load) | qwen3_embedding_0_6b.aimodel 633.5 MB |
649.3 MB |
Embeddings
- Width: 1024 dimensions; Matryoshka: any prefix from 32 to 1024 dimensions (
EmbedOptions(dimensions:)), re-normalized. - Pooling: the last token's hidden state (
lastToken); vectors are unit length (normalized in the graph). - Token limit: 8192 tokens per text, the appended tokens included; longer texts are refused (AE-EMB-001).
- Appended tokens: none beyond the tokenizer's own.
- Queries are rendered as
Instruct: {instruction}\nQuery:{text}with the default instructionGiven a web search query, retrieve relevant passages that answer the query; documents as{text}.
Verification
Every row is a record in verification/ about exactly these bytes (matched by bundle digest). Reference rows
are strict T2 passes of the unquantized export on the same fixture, in verification/reference/.
| Variant | Tier | Result | Detail | Device | OS build | Compute | Record |
|---|---|---|---|---|---|---|---|
ios-any-gpu |
T2 | pass | 14/14 cases; profile quantized-8bit; cosine ≥ 0.999; fixture d7defd0591ac7299 |
iPhone18,2 | 24A446 | target | f09227c4 |
macos-any-gpu |
T2 | pass | 14/14 cases; 14/14 cases; profile quantized-8bit; cosine ≥ 0.999; fixture d7defd0591ac7299 |
Mac17,6 | 26A434 | target | 65fc8aa7 |
| unquantized reference (not published) | T2 | pass | 14/14 cases; 14/14 cases; profile strict; cosine ≥ 0.999; fixture d7defd0591ac7299 |
Mac17,6 | 26A434 | target | e1539410 |
The quantized profile also requires: T2 strict on the unquantized reference export (met by the reference row).
Use
aether embed qwen3-embedding-0.6b --query "What is the capital of China?" --text "The capital of China is Beijing."
import Aether
let aether = try Aether()
let embedder = try await aether.embedder("qwen3-embedding-0.6b")
let vectors = try await embedder.embed([.query("What is the capital of China?"), .document("The capital of China is Beijing.")])
print(Embedding.cosine(vectors[0], vectors[1]))