Add Qwen3-Embedding-0.6B
Browse filesAdds **Qwen/Qwen3-Embedding-0.6B** as a built-in model, usable as `ts/Qwen3-Embedding-0.6B`.
> **Blocked on typesense/typesense#3039.** This model needs the `qwen` byte level BPE tokenizer added there; until that ships, `model_type: "qwen"` is rejected by `validate_and_init_local_model`.
Qwen3-Embedding is multilingual across 100+ languages at 1024 dims, and currently sits at the top of the MTEB multilingual leaderboard among openly licensed models. Apache-2.0.
### Files
| File | Notes |
|---|---|
| `Qwen3-Embedding-0.6B/model.onnx` | exported from `Qwen/Qwen3-Embedding-0.6B` |
| `Qwen3-Embedding-0.6B/model.onnx_data` | external weights, 2.38 GB |
| `Qwen3-Embedding-0.6B/tokenizer.json` | the Hugging Face tokenizer, unmodified — it carries both the vocabulary and the merges, so it fits `vocab_file_name` |
| `Qwen3-Embedding-0.6B/config.json` | `model_type: qwen`, with the model's query instruction as `query_prefix` |
`.gitattributes` gains LFS rules for the weights and for `tokenizer.json`, which is 11 MB.
### About the export
Qwen pools on the **last non-padded token**, not the mean, and then L2 normalises. Since `TextEmbedder::mean_pooling` is not configurable, the graph does that pooling itself and broadcasts the result across the sequence axis, so mean pooling over identical rows returns it unchanged. Same approach as `bge-m3` in #17 and `embeddinggemma-300m` in #19.
One detail specific to this model: Typesense **right pads** its batches (`batch_encode`), while Qwen's own tokenizer left pads. The graph therefore derives the last token index from the attention mask (`sum(mask) - 1`) rather than assuming position -1. A left-padding assumption would silently read a pad token for every input in a batch except the longest.
As with the other exports, `torch.onnx.export` emits IR version 10 and the pinned onnxruntime (`rel-1.14.1`) accepts at most IR 8, so `ir_version` has to be set back to 8. Opset 18 is fine.
### Testing
Built `typesense-server` from typesense/typesense#3039 and ran it with these exact files in `<data-dir>/models/`:
- collection creates, `num_dim` resolves to 1024
- embeddings match `sentence-transformers` at **cosine 1.000000** on six documents spanning Latin, Chinese, German, French and Russian text
- the same holds for a right padded batch of six, which is what exercises the last token indexing
- search resolves correctly: "What is the capital of Japan?" ranks 东京是日本的首都。first at 0.279, "Wo ist der Hund?" the German sentence at 0.483, "programozási nyelv" the English Python sentence at 0.541, "neural networks need data" the machine learning sentence at 0.422
### Two things maintainers should know
**The query prefix cannot be reproduced exactly.** Qwen's documented query format ends `...\nQuery:` with no trailing space, but `get_query_prefix` appends one unconditionally, so the model sees `Query: <text>` rather than `Query:<text>`. Measured cost: cosine 0.986 to 0.9998 against the canonical form, and top-1 ranking was unchanged on all four queries I tried. I have written the prefix without a trailing space, which is the closest the current API allows. Making that space conditional would let this be exact.
**This model is slower than the others here.** It is 0.6B parameters, so on CPU a single document costs roughly 0.2s at 128 tokens and 7s at 2048, with peak RSS of 1.7 GB and 2.6 GB. The tokenizer in #3039 caps input at 2048 tokens for that reason, even though the model itself accepts 32k. Worth a note in the README next to it, since the existing entries are all far smaller.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
https://claude.ai/code/session_01A9ydwqcWZWFG2bLUnDLt8E
|
@@ -33,3 +33,5 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
|
| 33 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
| 35 |
multilingual-e5-large/model.onnx_data filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
|
|
|
| 33 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
| 35 |
multilingual-e5-large/model.onnx_data filter=lfs diff=lfs merge=lfs -text
|
| 36 |
+
Qwen3-Embedding-0.6B/model.onnx_data filter=lfs diff=lfs merge=lfs -text
|
| 37 |
+
Qwen3-Embedding-0.6B/tokenizer.json filter=lfs diff=lfs merge=lfs -text
|
|
@@ -0,0 +1,8 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"model_md5": "d87d9a4b9f7ef88ed839e7afaec4e29b",
|
| 3 |
+
"vocab_file_name": "tokenizer.json",
|
| 4 |
+
"vocab_md5": "665a9bd72da35d9c577a5e85df887543",
|
| 5 |
+
"data_md5": "edde49f77d44f5726c5b5b791c15f761",
|
| 6 |
+
"model_type": "qwen",
|
| 7 |
+
"query_prefix": "Instruct: Given a web search query, retrieve relevant passages that answer the query\nQuery:"
|
| 8 |
+
}
|
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:16065585c6073f8574dee1c3ada34984e04c5529bf42924d7dc9d8586cf59dd3
|
| 3 |
+
size 4752658
|
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:7c7569e58783ee0ad8c5fb797d7944aa4f5928af53fb4c1f626f71885af22969
|
| 3 |
+
size 2383077376
|
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:def76fb086971c7867b829c23a26261e38d9d74e02139253b38aeb9df8b4b50a
|
| 3 |
+
size 11423705
|