honzajara commited on
Commit
82d4dcd
·
verified ·
1 Parent(s): 30e1d21

Add Qwen3-Embedding-0.6B

Browse files

Adds **Qwen/Qwen3-Embedding-0.6B** as a built-in model, usable as `ts/Qwen3-Embedding-0.6B`.

> **Blocked on typesense/typesense#3039.** This model needs the `qwen` byte level BPE tokenizer added there; until that ships, `model_type: "qwen"` is rejected by `validate_and_init_local_model`.

Qwen3-Embedding is multilingual across 100+ languages at 1024 dims, and currently sits at the top of the MTEB multilingual leaderboard among openly licensed models. Apache-2.0.

### Files

| File | Notes |
|---|---|
| `Qwen3-Embedding-0.6B/model.onnx` | exported from `Qwen/Qwen3-Embedding-0.6B` |
| `Qwen3-Embedding-0.6B/model.onnx_data` | external weights, 2.38 GB |
| `Qwen3-Embedding-0.6B/tokenizer.json` | the Hugging Face tokenizer, unmodified — it carries both the vocabulary and the merges, so it fits `vocab_file_name` |
| `Qwen3-Embedding-0.6B/config.json` | `model_type: qwen`, with the model's query instruction as `query_prefix` |

`.gitattributes` gains LFS rules for the weights and for `tokenizer.json`, which is 11 MB.

### About the export

Qwen pools on the **last non-padded token**, not the mean, and then L2 normalises. Since `TextEmbedder::mean_pooling` is not configurable, the graph does that pooling itself and broadcasts the result across the sequence axis, so mean pooling over identical rows returns it unchanged. Same approach as `bge-m3` in #17 and `embeddinggemma-300m` in #19.

One detail specific to this model: Typesense **right pads** its batches (`batch_encode`), while Qwen's own tokenizer left pads. The graph therefore derives the last token index from the attention mask (`sum(mask) - 1`) rather than assuming position -1. A left-padding assumption would silently read a pad token for every input in a batch except the longest.

As with the other exports, `torch.onnx.export` emits IR version 10 and the pinned onnxruntime (`rel-1.14.1`) accepts at most IR 8, so `ir_version` has to be set back to 8. Opset 18 is fine.

### Testing

Built `typesense-server` from typesense/typesense#3039 and ran it with these exact files in `<data-dir>/models/`:

- collection creates, `num_dim` resolves to 1024
- embeddings match `sentence-transformers` at **cosine 1.000000** on six documents spanning Latin, Chinese, German, French and Russian text
- the same holds for a right padded batch of six, which is what exercises the last token indexing
- search resolves correctly: "What is the capital of Japan?" ranks 东京是日本的首都。first at 0.279, "Wo ist der Hund?" the German sentence at 0.483, "programozási nyelv" the English Python sentence at 0.541, "neural networks need data" the machine learning sentence at 0.422

### Two things maintainers should know

**The query prefix cannot be reproduced exactly.** Qwen's documented query format ends `...\nQuery:` with no trailing space, but `get_query_prefix` appends one unconditionally, so the model sees `Query: <text>` rather than `Query:<text>`. Measured cost: cosine 0.986 to 0.9998 against the canonical form, and top-1 ranking was unchanged on all four queries I tried. I have written the prefix without a trailing space, which is the closest the current API allows. Making that space conditional would let this be exact.

**This model is slower than the others here.** It is 0.6B parameters, so on CPU a single document costs roughly 0.2s at 128 tokens and 7s at 2048, with peak RSS of 1.7 GB and 2.6 GB. The tokenizer in #3039 caps input at 2048 tokens for that reason, even though the model itself accepts 32k. Worth a note in the README next to it, since the existing entries are all far smaller.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01A9ydwqcWZWFG2bLUnDLt8E

.gitattributes CHANGED
@@ -33,3 +33,5 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zst filter=lfs diff=lfs merge=lfs -text
34
  *tfevents* filter=lfs diff=lfs merge=lfs -text
35
  multilingual-e5-large/model.onnx_data filter=lfs diff=lfs merge=lfs -text
 
 
 
33
  *.zst filter=lfs diff=lfs merge=lfs -text
34
  *tfevents* filter=lfs diff=lfs merge=lfs -text
35
  multilingual-e5-large/model.onnx_data filter=lfs diff=lfs merge=lfs -text
36
+ Qwen3-Embedding-0.6B/model.onnx_data filter=lfs diff=lfs merge=lfs -text
37
+ Qwen3-Embedding-0.6B/tokenizer.json filter=lfs diff=lfs merge=lfs -text
Qwen3-Embedding-0.6B/config.json ADDED
@@ -0,0 +1,8 @@
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "model_md5": "d87d9a4b9f7ef88ed839e7afaec4e29b",
3
+ "vocab_file_name": "tokenizer.json",
4
+ "vocab_md5": "665a9bd72da35d9c577a5e85df887543",
5
+ "data_md5": "edde49f77d44f5726c5b5b791c15f761",
6
+ "model_type": "qwen",
7
+ "query_prefix": "Instruct: Given a web search query, retrieve relevant passages that answer the query\nQuery:"
8
+ }
Qwen3-Embedding-0.6B/model.onnx ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:16065585c6073f8574dee1c3ada34984e04c5529bf42924d7dc9d8586cf59dd3
3
+ size 4752658
Qwen3-Embedding-0.6B/model.onnx_data ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:7c7569e58783ee0ad8c5fb797d7944aa4f5928af53fb4c1f626f71885af22969
3
+ size 2383077376
Qwen3-Embedding-0.6B/tokenizer.json ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:def76fb086971c7867b829c23a26261e38d9d74e02139253b38aeb9df8b4b50a
3
+ size 11423705