qwen3-embedding-0.6b

A Core AI bundle of Qwen/Qwen3-Embedding-0.6B for the Aether SDK (iOS and macOS 27+).

  • Source: Qwen/Qwen3-Embedding-0.6B at revision 97b0c614be4d77ee51c0cef4e5f07c00f9eb65b3, licence Apache-2.0. The licence is included as LICENSE. Qwen/Qwen3-Embedding-0.6B declares apache-2.0 in its model card but ships no licence file; LICENSE is the canonical text from apache.org.
  • Changes from the source: converted from PyTorch to Core AI (.aimodel) by Aether forge (recipe qwen3-embedding-0.6b@2). Weights are int8-linear-perblock32 (8-bit weights). The tokenizer files are the source's own.
  • Output: one vector per text (embeddings), for search and similarity; see Embeddings below.

Variants

Variant Platform Arch Compute Compiled Assets Download
macos-any-gpu macos any gpu no (specialized on first load) qwen3_embedding_0_6b.aimodel 633.5 MB 649.3 MB
ios-any-gpu ios any gpu no (specialized on first load) qwen3_embedding_0_6b.aimodel 633.5 MB 649.3 MB

Embeddings

  • Width: 1024 dimensions; Matryoshka: any prefix from 32 to 1024 dimensions (EmbedOptions(dimensions:)), re-normalized.
  • Pooling: the last token's hidden state (lastToken); vectors are unit length (normalized in the graph).
  • Token limit: 8192 tokens per text, the appended tokens included; longer texts are refused (AE-EMB-001).
  • Appended tokens: none beyond the tokenizer's own.
  • Queries are rendered as Instruct: {instruction}\nQuery:{text} with the default instruction Given a web search query, retrieve relevant passages that answer the query; documents as {text}.

Verification

Every row is a record in verification/ about exactly these bytes (matched by bundle digest). Reference rows are strict T2 passes of the unquantized export on the same fixture, in verification/reference/.

Variant Tier Result Detail Device OS build Compute Record
ios-any-gpu T2 pass 14/14 cases; profile quantized-8bit; cosine ≥ 0.999; fixture d7defd0591ac7299 iPhone18,2 24A446 target f09227c4
macos-any-gpu T2 pass 14/14 cases; 14/14 cases; profile quantized-8bit; cosine ≥ 0.999; fixture d7defd0591ac7299 Mac17,6 26A434 target 65fc8aa7
unquantized reference (not published) T2 pass 14/14 cases; 14/14 cases; profile strict; cosine ≥ 0.999; fixture d7defd0591ac7299 Mac17,6 26A434 target e1539410

The quantized profile also requires: T2 strict on the unquantized reference export (met by the reference row).

Use

aether embed qwen3-embedding-0.6b --query "What is the capital of China?" --text "The capital of China is Beijing."
import Aether

let aether = try Aether()
let embedder = try await aether.embedder("qwen3-embedding-0.6b")
let vectors = try await embedder.embed([.query("What is the capital of China?"), .document("The capital of China is Beijing.")])
print(Embedding.cosine(vectors[0], vectors[1]))
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for aether-models/qwen3-embedding-0.6b

Finetuned
(554)
this model