Qwen3-Reranker-0.6B — Core ML int8 (Apple Neural Engine, compiled)
English
Qwen/Qwen3-Reranker-0.6B converted to Core ML:
- Neural Engine. The model runs on the Apple Neural Engine (2,566 of 2,573 ops; the other 7 are the token-embedding lookup, on the CPU).
- int8 weights. Activations stay fp16.
- Precompiled. It ships as a compiled
.mlmodelc. - Fixed lengths. Nine functions have fixed lengths from 128 to 768 tokens and share one set of weights.
It scores a (query, document) pair the way the model card does: the logits of "no" and "yes" at the last position,
and the relevance score is sigmoid(yes − no).
Files
| File | What | Size |
|---|---|---|
Qwen3Reranker.mlmodelc |
int8 weights, functions L128, L160, L192, L224, L256, L320, L384, L512, L768 |
606 MB |
tokenizer.json |
The original Qwen3 tokenizer, unchanged | 11 MB |
config.json |
Prompt strings, token ids, functions, input and output layout | |
SHA256SUMS |
Checksums of every file |
Sizes are in MB = 10⁶ bytes.
Inputs and output
Every function L{N} takes:
| Name | Type | Shape | Content |
|---|---|---|---|
input_ids |
int32 | [1, N] | Token ids, left-padded with 151643 (<|endoftext|>) |
mask |
fp16 | [1, N] | 0 for padding, 1 for tokens |
It returns logits, fp32 [1, 2], in the order [no, yes]. Relevance = 1 / (1 + exp(no − yes)), which is the
same as the model card's softmax over the "no" and "yes" logits.
Use
Build three strings (
\nis a newline). The prefix and suffix are fixed and are also inconfig.json:prefix: <|im_start|>system\nJudge whether the Document meets the requirements based on the Query and the Instruct provided. Note that the answer can only be "yes" or "no".<|im_end|>\n<|im_start|>user\n body: <Instruct>: {instruction}\n<Query>: {query}\n<Document>: {document} suffix: <|im_end|>\n<|im_start|>assistant\n<think>\n\n</think>\n\nIn the body, replace
{instruction},{query}and{document}with your instruction, query and document. The default instruction isGiven a web search query, retrieve relevant passages that answer the query.Tokenize each string with
tokenizer.json, without adding special tokens. The prefix is 39 tokens and the suffix is 9.Cut the body's tokens so that prefix + body + suffix ≤ 768. Join the three lists; the result has n tokens.
Pick the smallest N ≥ n from {128, 160, 192, 224, 256, 320, 384, 512, 768}, left-pad to N, and call function
L{N}.
import numpy as np, coremltools as ct
from tokenizers import Tokenizer
tok = Tokenizer.from_file("tokenizer.json")
enc = lambda s: tok.encode(s, add_special_tokens=False).ids
PREFIX = ("<|im_start|>system\nJudge whether the Document meets the requirements based on the Query and the Instruct "
"provided. Note that the answer can only be \"yes\" or \"no\".<|im_end|>\n<|im_start|>user\n")
SUFFIX = "<|im_end|>\n<|im_start|>assistant\n<think>\n\n</think>\n\n"
INSTRUCT = "Given a web search query, retrieve relevant passages that answer the query"
SIZES = [128, 160, 192, 224, 256, 320, 384, 512, 768]
models = {}
def score(query, doc):
pre, suf = enc(PREFIX), enc(SUFFIX)
body = enc(f"<Instruct>: {INSTRUCT}\n<Query>: {query}\n<Document>: {doc}")[:768 - len(pre) - len(suf)]
ids = pre + body + suf; n = len(ids); N = next(s for s in SIZES if s >= n)
if N not in models:
models[N] = ct.models.CompiledMLModel("Qwen3Reranker.mlmodelc",
compute_units=ct.ComputeUnit.CPU_AND_NE, function_name=f"L{N}")
feed = {"input_ids": np.array([[151643] * (N - n) + ids], np.int32),
"mask": np.array([[0] * (N - n) + [1] * n], np.float16)}
no, yes = models[N].predict(feed)["logits"].reshape(-1)
return float(1 / (1 + np.exp(no - yes)))
print(score("周末去哪里爬山", "香山公园秋季红叶最好看,周末游客较多,建议早上出发。"))
import CoreML
let config = MLModelConfiguration()
config.computeUnits = .cpuAndNeuralEngine
config.functionName = "L256" // smallest bucket that fits
let model = try MLModel(contentsOf: rerankerURL, configuration: config) // Qwen3Reranker.mlmodelc
let ids = try MLMultiArray(shape: [1, 256], dataType: .int32) // left-padded with 151643
let mask = try MLMultiArray(shape: [1, 256], dataType: .float16) // 0 = padding, 1 = token
// … fill ids and mask …
let out = try model.prediction(from: MLDictionaryFeatureProvider(dictionary: ["input_ids": ids, "mask": mask]))
let logits = out.featureValue(for: "logits")!.multiArrayValue!
let relevance = 1 / (1 + exp(logits[0].doubleValue - logits[1].doubleValue)) // [no, yes]
Load each function once and keep it. Loading a function is slow; scoring is fast.
How it was converted
We wrote the Qwen3 forward pass by hand in PyTorch so it can be traced with fixed shapes:
- Fixed length per function. Inputs are left-padded and use a causal mask plus a padding mask. The RoPE tables are precomputed.
- RMSNorm safe for fp16. On our test pairs, Qwen3's hidden states reach values above 8,000, and squaring them inside RMSNorm overflows fp16. Each RMSNorm first divides the row by its largest absolute value, which leaves the result mathematically unchanged. (With plain RMSNorm, our conversion of Qwen3-Embedding-0.6B, which has the same backbone, reached only cosine 0.90 to the original on the Neural Engine.)
- MLP scaled for fp16. The MLP's gate × up product is scaled by 1/64 before
down_proj, and the result is scaled back by 64. - Two-row LM head. Only the two LM-head rows that are needed ("no" 2152, "yes" 9693) are kept, so each call returns 2 logits instead of 151,669.
Then, for each length:
- trace with
torch.jit.trace; - convert with coremltools 9 to an ML Program (fp16 compute, macOS 15 target);
- quantize to int8 with
linear_quantize_weights(linear symmetric, per output channel, weights with more than 2048 elements); - merge the nine functions with
ct.utils.save_multifunction, which stores the weights once; - compile with
ct.utils.compile_model.
Validation
Tests ran on one MacBook Pro (M3 Max, macOS 26.7) with the Neural Engine. For each query we took EmbeddingGemma 2's top 30 results and reranked them with this model, using the default instruction above. The reference is the original model in PyTorch (fp16, GPU), called as in its model card.
| Test set | EmbeddingGemma 2 alone | + PyTorch reranker | + this model (int8, ANE) | Rank correlation with PyTorch |
|---|---|---|---|---|
| Simulated personal knowledge base, 113 queries | 93.5 | 97.5 | 96.9 | 0.992 |
| MLQA, Chinese question → English passage, 100 queries | 60.8 | 69.0 | 68.9 | 0.993 |
| MLQA, English question → Chinese passage, 100 queries | 63.3 | 78.9 | 78.6 | 0.995 |
All scores are nDCG@10. The simulated knowledge base has 2,829 chunks of notes, Wikipedia pages, open-source documentation and code, mixed Chinese and English, with queries written by an LLM. Rank correlation is the mean Spearman correlation over each query's 30 candidates.
Other checks. An fp16 build of the same conversion matched PyTorch on three more sets:
| Test set | PyTorch reranker | fp16 build (ANE) |
|---|---|---|
| T2Retrieval (Chinese) | 94.2 | 94.3 |
| DuRetrieval (Chinese) | 93.7 | 93.8 |
| SciFact (English) | 73.0 | 71.1 |
The instruction matters. On SciFact, reranking with the default web-search instruction lowered nDCG@10 from 82.2 to 73.0, because SciFact's queries are scientific claims, not questions. Write an instruction that matches your task.
Speed (M3 Max, Neural Engine, one pair per call, median):
| Function | L128 | L160 | L192 | L224 | L256 | L320 | L384 | L512 | L768 |
|---|---|---|---|---|---|---|---|---|---|
| ms per pair | 14.5 | 20.2 | 22.7 | 25.8 | 31.1 | 44.9 | 54.0 | 83.5 | 145 |
Reranking the top 10 takes about 0.3–0.4 s per query on these test sets; the top 30 takes 0.8–1.2 s. int8 is barely faster than fp16 (at most 6%), but it is half the size. In a test with fp16 weights, scoring 4 pairs as one batch was slower per pair than scoring them one at a time.
Limitations
- Tested on one machine only: M3 Max with macOS 26.7. The model needs macOS 15 / iOS 18 or later, because it is a multifunction model.
- First load is slow. The first load on a device compiles each function for the Neural Engine. On the M3 Max this took 13–30 s per function, 2–3 minutes for all nine. Core ML caches the result. Apps should warm the functions up in the background.
- 768-token limit. Pairs longer than 768 tokens must be truncated. The original model accepts much longer inputs.
- Small test sets. The test sets are small and some queries are LLM-written.
License and attribution
- License. Apache 2.0, the same as Qwen3-Reranker-0.6B.
- Qwen. The model is by the Qwen team, Alibaba Cloud.
- This repo. We did the Core ML conversion, the int8 quantization and the tests. This repo is not affiliated with the Qwen team.
中文
本仓库是 Qwen/Qwen3-Reranker-0.6B 的 Core ML 版本:
- 神经网络引擎:模型在 Apple 神经网络引擎(ANE)上运行(2,573 个算子中有 2,566 个在 ANE 上,剩下 7 个是查词表那一步,在 CPU 上)。
- int8 权重:计算仍用 fp16。
- 预先编译:以编译好的
.mlmodelc发布。 - 固定长度:有 9 个长度固定的函数,从 128 到 768 个 token,共用一份权重。
打分方式和官方模型卡相同:取最后一个位置上“no”和“yes”两个词的 logit,相关度 = sigmoid(yes − no)。
文件
| 文件 | 内容 | 大小 |
|---|---|---|
Qwen3Reranker.mlmodelc |
int8 权重;函数 L128、L160、L192、L224、L256、L320、L384、L512、L768 |
606 MB |
tokenizer.json |
Qwen3 原版分词器,未改动 | 11 MB |
config.json |
提示词、特殊 token、函数列表、输入输出格式 | |
SHA256SUMS |
所有文件的校验和 |
本页的 MB 指 10⁶ 字节。
输入和输出
每个函数 L{N} 的输入:
| 名称 | 类型 | 形状 | 内容 |
|---|---|---|---|
input_ids |
int32 | [1, N] | token id,在左侧补齐,补齐值为 151643(<|endoftext|>) |
mask |
fp16 | [1, N] | 补齐的位置为 0,有 token 的位置为 1 |
输出是 logits,fp32 [1, 2],顺序是 **[no, yes]**。相关度 = 1 / (1 + exp(no − yes)),
和模型卡里对 no、yes 两个 logit 做 softmax 的结果相同。
用法
拼三段文本(
\n表示换行)。前缀和后缀是固定的,config.json里也有:prefix: <|im_start|>system\nJudge whether the Document meets the requirements based on the Query and the Instruct provided. Note that the answer can only be "yes" or "no".<|im_end|>\n<|im_start|>user\n body: <Instruct>: {instruction}\n<Query>: {query}\n<Document>: {document} suffix: <|im_end|>\n<|im_start|>assistant\n<think>\n\n</think>\n\n把 body 里的
{instruction}、{query}、{document}分别换成指令、查询和文档。 默认指令是Given a web search query, retrieve relevant passages that answer the query。分词。 三段分别用
tokenizer.json分词,不加特殊 token。前缀 39 个 token,后缀 9 个。截断。 截短中间一段,使前缀 + 中间 + 后缀 ≤ 768 个 token,再把三段拼起来,共 n 个 token。
运行。 在 {128, 160, 192, 224, 256, 320, 384, 512, 768} 中选出不小于 n 的最小值 N,在左侧补齐到 N, 然后调用函数
L{N}。
import numpy as np, coremltools as ct
from tokenizers import Tokenizer
tok = Tokenizer.from_file("tokenizer.json")
enc = lambda s: tok.encode(s, add_special_tokens=False).ids
PREFIX = ("<|im_start|>system\nJudge whether the Document meets the requirements based on the Query and the Instruct "
"provided. Note that the answer can only be \"yes\" or \"no\".<|im_end|>\n<|im_start|>user\n")
SUFFIX = "<|im_end|>\n<|im_start|>assistant\n<think>\n\n</think>\n\n"
INSTRUCT = "Given a web search query, retrieve relevant passages that answer the query"
SIZES = [128, 160, 192, 224, 256, 320, 384, 512, 768]
models = {}
def score(query, doc):
pre, suf = enc(PREFIX), enc(SUFFIX)
body = enc(f"<Instruct>: {INSTRUCT}\n<Query>: {query}\n<Document>: {doc}")[:768 - len(pre) - len(suf)]
ids = pre + body + suf; n = len(ids); N = next(s for s in SIZES if s >= n)
if N not in models:
models[N] = ct.models.CompiledMLModel("Qwen3Reranker.mlmodelc",
compute_units=ct.ComputeUnit.CPU_AND_NE, function_name=f"L{N}")
feed = {"input_ids": np.array([[151643] * (N - n) + ids], np.int32),
"mask": np.array([[0] * (N - n) + [1] * n], np.float16)}
no, yes = models[N].predict(feed)["logits"].reshape(-1)
return float(1 / (1 + np.exp(no - yes)))
print(score("周末去哪里爬山", "香山公园秋季红叶最好看,周末游客较多,建议早上出发。"))
import CoreML
let config = MLModelConfiguration()
config.computeUnits = .cpuAndNeuralEngine
config.functionName = "L256" // 能放下这一对的最小长度档
let model = try MLModel(contentsOf: rerankerURL, configuration: config) // Qwen3Reranker.mlmodelc
let ids = try MLMultiArray(shape: [1, 256], dataType: .int32) // 在左侧用 151643 补齐
let mask = try MLMultiArray(shape: [1, 256], dataType: .float16) // 0 = 补齐,1 = token
// …… 填入 ids 和 mask ……
let out = try model.prediction(from: MLDictionaryFeatureProvider(dictionary: ["input_ids": ids, "mask": mask]))
let logits = out.featureValue(for: "logits")!.multiArrayValue!
let relevance = 1 / (1 + exp(logits[0].doubleValue - logits[1].doubleValue)) // 顺序是 [no, yes]
每个函数加载一次后保留复用:加载慢,打分快。
转换方法
我们用 PyTorch 手写了 Qwen3 的前向计算,使它能以固定形状被追踪:
- 每个函数长度固定:输入在左侧补齐,使用因果遮罩和补齐遮罩;RoPE 表预先算好。
- 防 fp16 溢出的 RMSNorm:在我们的测试样本上,Qwen3 的隐藏状态有超过 8,000 的离群值,在 RMSNorm 里平方后会超出 fp16 的范围。 每个 RMSNorm 先把这一行除以它的最大绝对值,数学上结果不变。(我们转换同一骨干的 Qwen3-Embedding-0.6B 时, 用普通 RMSNorm 在神经网络引擎上和原版的余弦相似度只有 0.90。)
- 防溢出的 MLP:MLP 中 gate × up 的乘积先乘以 1/64,经过
down_proj之后再乘回 64。 - 只保留两行输出层:只保留需要的两行(“no” 2152、“yes” 9693),每次调用只输出 2 个 logit,而不是 151,669 个。
然后对每个长度:
- 用
torch.jit.trace追踪; - 用 coremltools 9 转成 ML Program(fp16 计算,最低 macOS 15);
- 用
linear_quantize_weights压成 int8(线性对称,按输出通道;只压缩元素数超过 2048 的权重); - 用
ct.utils.save_multifunction把 9 个函数合并成一个模型,权重只存一份; - 用
ct.utils.compile_model编译。
验证
在一台 MacBook Pro(M3 Max,macOS 26.7)上测试,模型跑在神经网络引擎上。对每条查询,先取 EmbeddingGemma 2 检索出的前 30 条结果,再用本模型重排,指令用上面的默认指令。 参照是 PyTorch 原版模型(fp16,GPU),调用方式和模型卡相同。
| 测试集 | 只用 EmbeddingGemma 2 | + PyTorch 原版重排 | + 本模型(int8,ANE) | 和 PyTorch 排序的相关系数 |
|---|---|---|---|---|
| 模拟个人知识库,113 条查询 | 93.5 | 97.5 | 96.9 | 0.992 |
| MLQA,中文问题找英文段落,100 条查询 | 60.8 | 69.0 | 68.9 | 0.993 |
| MLQA,英文问题找中文段落,100 条查询 | 63.3 | 78.9 | 78.6 | 0.995 |
- 指标:都是 nDCG@10。
- 模拟个人知识库:2,829 个文本块,内容包括笔记、维基百科页面、开源文档和代码,中英混合;查询由大模型编写。
- 相关系数:对每条查询的 30 个候选计算 Spearman 相关系数,再取平均。
其他检查:同一转换方法的 fp16 版本在另外三个测试集上也和 PyTorch 一致:
| 测试集 | PyTorch 原版重排 | fp16 版本(ANE) |
|---|---|---|
| T2Retrieval(中文) | 94.2 | 94.3 |
| DuRetrieval(中文) | 93.7 | 93.8 |
| SciFact(英文) | 73.0 | 71.1 |
指令很重要:在 SciFact 上,用默认的网页搜索指令重排后,nDCG@10 从 82.2 降到了 73.0。 原因是 SciFact 的查询是科学论断,不是问题。请按自己的任务写指令。
速度(M3 Max,神经网络引擎,每次调用打一对分,取中位数):
| 函数 | L128 | L160 | L192 | L224 | L256 | L320 | L384 | L512 | L768 |
|---|---|---|---|---|---|---|---|---|---|
| 每对耗时(ms) | 14.5 | 20.2 | 22.7 | 25.8 | 31.1 | 44.9 | 54.0 | 83.5 | 145 |
- 每条查询的耗时:在这些测试集上,重排前 10 条约 0.3–0.4 秒,前 30 条约 0.8–1.2 秒。
- int8 和 fp16 对比:int8 只比 fp16 快一点(最多 6%),但体积减半。
- 批量打分:在 fp16 版本上测过,把 4 对放在一个批次里打分,平均每对反而比逐对打分更慢。
局限
- 只在一台机器上测过:M3 Max,macOS 26.7。模型是多函数模型,需要 macOS 15 / iOS 18 或更高版本。
- 第一次加载很慢:在每台设备上第一次加载时,系统要把每个函数编译给神经网络引擎。M3 Max 上每个函数需要 13–30 秒, 9 个函数共 2–3 分钟。Core ML 会缓存编译结果。建议 App 在后台预热各个函数。
- 最长 768 个 token:超过 768 个 token 的输入必须截断,而原版模型支持长得多的输入。
- 测试集很小:部分查询由大模型编写。
许可证和来源
- 许可证:Apache 2.0,和 Qwen3-Reranker-0.6B 相同。
- Qwen:模型由阿里云通义千问团队发布。
- 本仓库:我们做了 Core ML 转换、int8 量化和测试。本仓库与通义千问团队没有关联。
- Downloads last month
- 14