license: mit
language:
- code
- multilingual
tags:
- code
- code-search
- code-retrieval
- embeddings
- feature-extraction
- sentence-similarity
- knowledge-distillation
- quantized
- int8
pipeline_tag: feature-extraction
base_model:
- intfloat/multilingual-e5-base
- Alibaba-NLP/gte-modernbert-base
- Qwen/Qwen3-Reranker-4B
datasets:
- CoIR-Retrieval/cosqa
- CoIR-Retrieval/codesearchnet
- CoIR-Retrieval/stackoverflow-qa
code-daemon-embed-v1
A 46.8M-parameter, 4-layer code-embedding model that maps short code units β function and method bodies, signatures, docstrings, symbol names β and short natural-language queries into a shared 768-dim space. INT8, 128-token window, ~10.7k embeddings/sec on a laptop RTX 5060.
It is built for one job: embedding a whole repository fast enough to re-index it on every commit, so that a coding agent can run semantic search on every question it is asked. Every trade-off below follows from that β the depth, the 22.7k vocabulary, the 128-token cap, the INT8 weights, the four length-bucketed engines.
# the whole API surface β pooled AND L2-normalized inside the graph
vec = session.run(None, {"input_ids": ids, "attention_mask": mask})[0] # [B, 768], ready
Queries and documents are embedded the same way β no query: / passage: prefix.
1. What it is for, and why not a standard embedder
The target workload
- A repository is indexed: every function, method, type and doc chunk becomes one short text and one vector. A 700k-entity C++ codebase embeds in ~81 seconds on one consumer GPU.
- Someone asks something short and keyword-shaped β "git watcher head change reindex", "acquire database lock for project hash", "where does the daemon start".
- The vector channel runs next to a lexical (BM25-style) channel and the two are fused.
Point 3 matters: this model was trained to be the dense half of a hybrid retriever, not to win alone. Its training queries imitate real captured agent traffic β short keyword bags, behaviour descriptions, identifier fragments β rather than docstring paraphrases.
Design choices vs a typical general-purpose text embedder
| typical embedder (e5 / bge / gte class) | code-daemon-embed-v1 | |
|---|---|---|
| Parameters | 110M β 7B | 46.8M |
| Vocabulary | 30k β 250k, general text | 22.7k, pruned to code + code-adjacent English |
| Max sequence | 512 β 8192 tokens | 128 (hard cap, by design) |
| Query format | often needs an instruction prefix | none β symmetric |
| Pooling | you implement it | inside the graph β output is already [B, 768] |
| Weights | FP32 / FP16 | INT8 from quantization-aware training |
| Top-rank precision | needs a separate cross-encoder pass | cross-encoder ranking distilled into the vectors |
The built-in reranker
A bi-encoder compresses each document into a vector before it sees the query, so it cannot
do the pairwise comparison a cross-encoder reranker does. This model closes part of that gap
at training time instead of inference time: Qwen/Qwen3-Reranker-4B scored every
(query, candidate) pair over mined hard negatives, and the student was trained with a
listwise-KL objective to reproduce the teacher's ranking distribution β not merely
"positive above negative", but how much each near-miss should trail.
Practical consequence: you probably do not need a runtime reranker on top of it. In the production system it was built for, switching one on measured net-negative β the distillation had already captured the useful part, at zero inference cost.
Strong / weak
Strong
- Repository search with short NL, keyword and identifier queries.
- Natural language β code at short lengths.
- The dense channel of a hybrid retriever.
- Throughput-bound work: bulk re-index, index-on-save, re-embed per commit.
Weak / out of scope
- Long documents. Hard 128-token cap. Not a long-context retriever (see the window recipe in Β§4.6 if you must handle longer text).
- Benchmark-style long problem statements, multi-turn dialogue, codeβcode translation.
- General English prose (medical / financial / news) β the pruned vocabulary trades that away deliberately.
2. Architecture β what was kept, cut and added
Specification
| Base | intfloat/multilingual-e5-base (XLM-RoBERTa encoder, 12L, 278M) |
| Encoder layers | 4 |
| Hidden size | 768 |
| Attention heads | 12 |
| FFN | 3072, GELU |
| LayerNorm eps | 1e-5 |
| Vocabulary | 22,739 SentencePiece unigram pieces, byte fallback |
| Position embeddings | 514 allocated (learned, absolute) |
| Max sequence used | 128 tokens |
| Output | 768-dim, mask-mean-pooled AND L2-normalized inside the graph β use it as it comes |
| Parameters | 46.80M = 17.46M embedding table + 4 Γ ~7.3M encoder + positions |
| Weights | INT8, quantization-aware training, Q/DQ nodes carry the trained scales |
| ONNX | opset 19; inputs input_ids, attention_mask, both int64 [batch, seq] |
| Special token ids | pad=0, unk=1, bos=2, eos=3 (raw-SentencePiece indexing) |
Removed from the base model
- 8 of 12 encoder layers β strided-truncated 12 β 8 β 6 β 4, healed after each cut.
- 227k of 250k vocabulary pieces: 250,002 β 22,739, and the embedding table with them (192M β 17.5M parameters β the single largest saving in the model).
- The E5 instruction prefixes. No
query:/passage:asymmetry; both sides are encoded identically. - The pooler head, the MLM head, token-type inputs. The graph has exactly two inputs.
- FP32 weights β replaced by INT8 with trained scales.
Added
- Mask-aware mean pooling and L2 normalization fused into the ONNX graph. The model
returns unit-norm
[B, 768], not[B, seq, 768]β there is no pooling code to get wrong, nolast_hidden_statecopy, and no way to accidentally compare unnormalized vectors. - Q/DQ nodes with QAT-trained scales, so a TensorRT/OpenVINO build is INT8 end-to-end without a calibration pass.
- Ranking knowledge from a 4B cross-encoder (the built-in reranker above).
- Four length-bucketed engine builds with a dynamic sequence dimension (Β§3, Β§4.4).
Training recipe, in brief
- Strided depth truncation of the e5 backbone, vocabulary pruning, re-indexing to the raw-SentencePiece id convention.
- Dense representation distillation from
Alibaba-NLP/gte-modernbert-base(MSE against the teacher's embeddings) over real repository entities. - Listwise-KL ranking distillation from
Qwen/Qwen3-Reranker-4Bover mined hard negatives. - Diversity regularization β ~26% of the grouped training stream is out-of-domain code-retrieval pairs (CoIR-derived). Measured sweet spot: 0% loses deep recall, 50% dilutes the top.
- Self-distillation QAT to INT8, exported with the scales learned during training.
If you re-train, cut or re-quantize it
Two results that will cost you a training run if you rediscover them yourself:
- β Do not run PTQ/calibration over the shipped INT8 graph. It overwrites the trained
scales with fitted ones and measurably degrades the model (cosine .93 against the correct
export).
model_int8qdt.onnxis the artifact β feed it to your builder as-is. Plain post-training quantization of this body was also tried on its own and is quality-broken: engine hit@1 fell .200 β .133, which is why the shipped weights come from QAT. - β Heal a depth cut on documents and queries. Re-distilling the 4-layer student against document texts only cost β29.8% hit@1; the same initialization and budget against documents and queries cost β3.9% on the same 2,441-query gate. Retrieval lives in the relation between the two sides β a document-only anchor lets the query side drift.
3. Performance
Measured on one RTX 5060 Laptop GPU (sm_120, 8 GB), TensorRT INT8, pinned TDP,
clocks.sm 2,647 MHz median at 114 W.
End-to-end, inside the production indexer
The whole path: build the serve text β tokenize on host threads β IPC to the worker β GPU forward β pool β collect.
| corpus | vectors | wall | pipeline emb/s | GPU-only emb/s |
|---|---|---|---|---|
| mysql-server (sustained, 54 batches) | 868,795 | 80.96 s | 10,731 | 11,712 |
pipeline emb/s = vectors Γ· wall clock, everything included. GPU-only = vectors Γ· summed
inference time; it is higher because host work overlaps with the GPU.
Solo engine profiles β cost per sequence length
Four engines, one per length bucket, each with a dynamic sequence dimension so a batch can be dispatched at its own longest text instead of the bucket ceiling. One engine alone, dispatch only, 200 iterations:
| bucket | profile (batch Γ seq) | seq | Β΅s/text | texts/s |
|---|---|---|---|---|
| s | 96 Γ 8β¦48, opt 48 | 8 | 13.78 | 72,587 |
| 16 | 19.73 | 50,681 | ||
| 24 | 26.86 | 37,228 | ||
| 32 | 33.73 | 29,651 | ||
| 40 | 42.56 | 23,495 | ||
| 48 | 50.54 | 19,785 | ||
| m | 128 Γ 56β¦64, opt 56 | 56 | 63.08 | 15,853 |
| 64 | 76.52 | 13,068 | ||
| l | 128 Γ 72β¦80, opt 72 | 72 | 88.48 | 11,302 |
| 80 | 100.87 | 9,913 | ||
| xl | 256 Γ 88β¦128, opt 96 | 88 | 117.14 | 8,537 |
| 96 | 127.95 | 7,816 | ||
| 104 | 142.31 | 7,027 | ||
| 112 | 155.74 | 6,421 | ||
| 120 | 169.28 | 5,907 | ||
| 128 | 183.03 | 5,464 |
Three things to take from that table when you build your own serving path:
- Cost is per token, not per text, and it is superlinear: solo token throughput falls from 1.01M tok/s at seq 48 to 0.71M at seq 128. Batch size is not the lever β a 256Γ64 engine measured 5.8% slower per text than 128Γ64, because 96Γ48 already saturates the SMs.
- Padding is the lever. A batch pays for its longest member, so sorting texts by length before batching β and dispatching each batch at its own length β was worth +10.2% end-to-end. Round lengths up to a multiple of 8: a non-multiple is slower than a longer multiple (seq 70 costs more than seq 72).
- Running the four engines concurrently does not add throughput on one GPU. In parallel they measured slower than serially (1,487 ms vs 1,423 ms for the same work). The lanes exist to keep the GPU fed while the host works, not to multiply throughput.
Without a discrete GPU β OpenVINO on CPU, integrated GPU and NPU
Same model, no NVIDIA card involved. Measured on an Intel Core Ultra 9 275HX β its CPU cores, its integrated GPU, and its AI Boost NPU β with OpenVINO 2026.3. An OpenVINO IR is reshaped to one static shape at build time (unlike the TensorRT engines, whose sequence dimension is a range), so each bucket is a single number per device:
| device | precision | bucket | shape (batch Γ seq) | Β΅s/text | texts/s |
|---|---|---|---|---|---|
| CPU | INT8 | s | 64 Γ 48 | 1,424 | 702 |
| m | 64 Γ 64 | 1,965 | 509 | ||
| l | 64 Γ 80 | 2,485 | 402 | ||
| xl | 64 Γ 128 | 4,224 | 237 | ||
| iGPU | INT8 | s | 64 Γ 48 | 916 | 1,092 |
| m | 64 Γ 64 | 1,181 | 847 | ||
| l | 64 Γ 80 | 1,476 | 678 | ||
| xl | 64 Γ 128 | 2,381 | 420 | ||
| NPU | INT4 | s | 16 Γ 48 | 1,770 | 565 |
| m | 16 Γ 64 | 2,326 | 430 | ||
| l | 16 Γ 80 | 3,355 | 298 | ||
| xl | 16 Γ 128 | 5,911 | 169 |
All three at once. These are solo numbers β one device, nothing else running. The three
share one LPDDR5 controller, so they do not simply add up, but they come close: measured in a
single shared window at 64 Γ 64, CPU 416 + iGPU 750 + NPU 377 = 1,543 against 527 + 822 + 423 = 1,772 solo, i.e. 87 % of the sum. The noisiest neighbour is also the fastest device
(the iGPU costs the others ~13 %, the CPU ~7 %), so there is nobody worth switching off.
End to end that lands at ~1,050 texts/s on a real 20k-entity repository β the daemon runs all three workers concurrently and that figure includes tokenization, IPC and vector writes, not just inference.
Reference point: the same buckets on the RTX 5060 run at 19,785 / 13,068 / 9,913 / 5,464 texts/s, so the integrated GPU is 15β18Γ slower than the discrete one. That still leaves a 100k-entity repository indexed in a few minutes on a machine with no dedicated GPU at all.
Three notes if you deploy this path:
- INT8 on CPU and iGPU, INT4 only on the NPU. At 128 Γ 64 the iGPU does 859 texts/s in INT8, 584 in FP16 β and 584 in INT4, exactly the FP16 number. INT4 here is weight-only compression: the weights decompress to fp16 before the GEMM, so it is the same GEMM and the 4 bits buy file size, not speed. Only INT8, which the kernels execute natively, moves the number. The NPU is the exception because it requires INT4 with static shapes.
- Feed it the Q/DQ graph, not a calibrated one. OpenVINO keeps the QAT scales as FakeQuantize; re-fitting them with a post-training pass measured hit@1 .200 β .133 on this model. See Β§4.5 for the two-line recipe (and the transformation that keeps the IR at 96 MB instead of 179 MB).
- The NPU pays for batch, not for tokens. Its IRs are baked at batch 16 against 96β256 elsewhere, which is why its per-text cost is the flattest across buckets and its per-batch latency the lowest (28β95 ms). It suits interactive single-query embedding better than bulk indexing.
Retrieval quality
80 captured real agent queries over a live indexed codebase, multi-positive, file-level, through a production hybrid retriever (dense + lexical, runtime reranker off):
| metric | value |
|---|---|
| hit@1 | 0.46 |
| hit@3 | 0.74 |
| hit@5 | 0.78 |
| hit@10 | 0.80 |
| mrr@10 | 0.59 |
| ndcg@10 | 0.61 |
This is the metric the model is optimized for. On a broader 2,441-query gate across four unrelated repositories, the 4-layer model trades about 4% relative hit@1 against its 6-layer parent for roughly 1.5Γ the engine speed.
CoIR β the out-of-domain reference
Run on this exact artifact β model_int8qdt.onnx, full corpora (2.3M documents across the
six tasks), NDCG@10, the same raw-SentencePiece tokenization the daemon uses:
| CoIR task | NDCG@10 | Pattern |
|---|---|---|
| synthetic-text2sql | 51.26 | NL β SQL |
| stackoverflow-qa | 41.63 | short question β code |
| codesearchnet (6-lang avg) | 39.02 | docstring / NL β code |
| codefeedback-st | 38.55 | NL instruction β code |
| codesearchnet-ccr (6-lang avg) | 36.07 | code β related code |
| cosqa | 17.50 | NL question β code (noisy / hard) |
| Average | 37.34 |
Per language β codesearchnet (NLβcode): python 68.30, go 46.15, java 34.64, ruby 28.76, js 28.33, php 27.95. codesearchnet-ccr (codeβcode): js 43.36, ruby 42.88, java 36.82, php 31.59, go 31.06, python 30.68.
Read this as a lower bound, not as the headline. CoIR queries are mostly docstrings and long problem statements β the opposite of what this model was tuned for, and the spread proves it: 68.30 on Python docstringβcode against 17.50 on cosqa's noisy questionβcode. Four of CoIR's ten tasks (codeβcode translation, multi-turn dialogue, long problem statements) exceed the 128-token scope and are not shown. ~2k CoIR-derived rows were present in training as diversity regularization; overlap with the test splits was not audited.
4. Using the model
4.1 What is in this repository
| File | What it is |
|---|---|
model_int8qdt.onnx |
The source of truth. INT8 Q/DQ graph with QAT-trained scales β what the TensorRT and OpenVINO INT8 engines are built from. |
model.onnx |
The FP32 twin of the same weights, for lanes that cannot read Q/DQ and for fine-tuning. Derived from the file above by dropping its Q/DQ pairs, so the weights are identical (cosine 0.98 between them is the quantization, not a different model). |
sentencepiece.bpe.model |
The tokenizer. Raw-SP ids: pad=0, unk=1, bos=2, eos=3. Its 22,739 pieces match the model's embedding table row for row. |
tokenizer_config.json, config.json |
HF-side metadata. |
code-daemon-embed-v1-{s,m,l,xl}_{win_x64,linux_x64}_trt11.0_sm_{75,86,89,120}.engine |
Prebuilt TensorRT engines. |
code-daemon-embed-v1-{s,m,l,xl}_ov2026.3_{cpu_int8,igpu_lnl_int8,igpu_arc_int4,npu_int4}_b*_s*.{xml,bin} |
Prebuilt OpenVINO models (Intel CPU / iGPU / NPU). |
code-daemon-embed-v1_{win_x64,linux_x64}_tvm0.25_vulkan.{dll,so} |
TVM Vulkan module (vendor-neutral GPU fallback). |
model_gpu_mlx0.22/ |
MLX weights for Apple Silicon. |
Compiled artifacts are uploaded progressively per architecture. The ONNX + tokenizer are the always-present source of truth; if the engine you want is missing, build it β Β§4.4.
4.2 Taking a prebuilt engine
The filenames are a loading contract, not decoration. A serialized TensorRT plan is keyed on
{GPU architecture Γ OS Γ TensorRT version} and deserializeEngine has no compatibility
fallback, so pick all four coordinates exactly:
code-daemon-embed-v1-m_win_x64_trt11.0_sm_120.engine
β β β βββ GPU arch: sm_75 Turing Β· sm_86 Ampere (RTX 30xx, A-series)
β β β sm_89 Ada (RTX 40xx, L4) Β· sm_120 Blackwell (RTX 50xx)
β β βββ TensorRT 11.0 β not interchangeable with 10.x
β βββ OS/ABI
βββ length bucket: s | m | l | xl
Bucket shapes (batch Γ seq): s 96Γ48, m 128Γ64, l 128Γ80, xl 256Γ128. Loading all four costs ~400 MB of VRAM; if you only want one, take m β it covers the fattest part of a typical corpus (25β40% of texts land in 49β¦64 tokens).
Route each text to the first bucket whose sequence ceiling fits its token count, and pad the batch to that bucket's shape. Padding is masked in attention, so a padded batch and an unpadded one give the same vector (verified at cosine 1.000000) β provided the mask is right.
4.3 Running the ONNX directly
Works anywhere ONNX Runtime does (CPU, CUDA, DirectML) with no build step:
import onnxruntime as ort, sentencepiece as spm, numpy as np
sp = spm.SentencePieceProcessor(model_file="sentencepiece.bpe.model") # pad=0 unk=1 bos=2 eos=3
sess = ort.InferenceSession("model_int8qdt.onnx", providers=["CPUExecutionProvider"])
def embed(texts, max_len=128):
ids = [[2, *sp.encode(t)[: max_len - 2], 3] for t in texts] # bos β¦ eos
L = max(len(x) for x in ids)
inp = np.array([x + [0] * (L - len(x)) for x in ids], dtype=np.int64) # pad=0
mask = (inp != 0).astype(np.int64)
# pooled and unit-norm already β no post-processing
return sess.run(None, {"input_ids": inp, "attention_mask": mask})[0] # [B, 768]
D = embed(["function acquireLock in src/db.zig: zig\npath: src db"])
Q = embed(["acquire database lock"])
print(Q @ D.T) # inner product IS cosine here
Two notes that cost real debugging time:
- Tokenize with SentencePiece, not a greedy-BPE merge loop. The vocabulary is a unigram model; scoring it with pair-merge BPE produces a different segmentation than training used, and that train/serve skew silently costs retrieval quality.
- Pad with id 0 and mask 0. Both inputs are
int64; feeding int32 buffers to a hand-written runtime is the classic "second half of every batch is garbage" bug.
4.4 Building your own TensorRT engines
TensorRT 11 reads precision entirely from the ONNX Q/DQ nodes β the per-precision flags
(--int8, --fp16, --calib) were removed. Feed it the Q/DQ graph and pass
--stronglyTyped. Feed it an unquantized graph instead and you get an FP16 engine that builds
cleanly, loads cleanly, is twice the size and a fraction of the speed β the size is the only
visible symptom, so check it.
Rectangular engine for one bucket (here m, 128Γ64):
trtexec --onnx=model_int8qdt.onnx \
--saveEngine=code-daemon-embed-v1-m.engine \
--stronglyTyped \
--builderOptimizationLevel=5 \
--timingCacheFile=timing.cache \
--minShapes=input_ids:1x1,attention_mask:1x1 \
--optShapes=input_ids:128x64,attention_mask:128x64 \
--maxShapes=input_ids:128x64,attention_mask:128x64
Dynamic-sequence engines (the shipped configuration, +10.2% end-to-end) keep the batch
dimension rectangular and let seq range. Use these profiles:
| bucket | batch | seq min | seq opt | seq max |
|---|---|---|---|---|
| s | 96 | 8 | 48 | 48 |
| m | 128 | 56 | 56 | 64 |
| l | 128 | 72 | 72 | 80 |
| xl | 256 | 88 | 96 | 128 |
Three rules behind those numbers:
- One dynamic engine per bucket,
optat that bucket's mean length β never one global dynamic engine. The tax is paid for being far from the opt point, not for being dynamic: a single wide (16β64β128) engine costs +0.8% at seq 64 but +15.4% at seq 80. - Sequence length must be a multiple of 8 at dispatch (seq 70 costs more than seq 72).
- Sort texts by length within a bucket and cut at batch boundaries before dispatch. Without the sort the dynamic engines buy ~0%: any 128-text batch drawn from the 49β64 range almost surely contains a 64 and dispatches at 64 anyway.
A dynamic build reports ~69 layers against a rectangular build's ~54 β a wider shape range costs fusions, and that is exactly where the small tax at the ceiling lives.
4.5 OpenVINO, TVM, MLX
OpenVINO reads the Q/DQ graph directly β the trained scales survive as FakeQuantize, so there is no calibration pass to run:
import openvino as ov
from openvino._offline_transformations import compress_quantize_weights_transformation
m = ov.Core().read_model("model_int8qdt.onnx")
compress_quantize_weights_transformation(m) # folds weights to i8: 179 MB -> 96 MB
m.reshape({"input_ids": [128, 64], "attention_mask": [128, 64]}) # one bucket, static
ov.save_model(m, "code-daemon-embed-v1-m_cpu.xml")
That transformation is not optional bookkeeping: read_model leaves every weight as f32 behind
a FakeQuantize, and the IR comes out nearly twice the size for the same arithmetic. Do not
reach for NNCF post-training quantization here β it would replace the trained scales with
fitted ones, which is the failure the warning in Β§2 describes.
The NPU artifacts are the exception: they start from the FP32 twin and apply INT4 weight compression at batch 16, trading accuracy for size on purpose. TVM Vulkan modules and the MLX weights are also built from the FP32 twin, per bucket.
4.6 Feeding it well
Documents. The model was trained on a compact, front-loaded "serve text": semantics first, identifiers after, everything inside the 128-token budget. Reproducing that shape on your own corpus is worth more than any inference tuning:
{type} {name} in {file}: {lang}
path: {directory tokens, space-separated}
[async] [exported] [test]
sig: ({params}) -> {return type}
{doc comment, first ~200 chars}
{one-to-two-sentence description}
for example:
function acquireProjectLock in src/storage/multi_db.zig: zig
path: src storage multi db
[exported]
sig: (allocator: Allocator, project_hash: []const u8) -> !Lock
Acquires the exclusive SQLite lock for one project.
Cap the whole text around 768 characters; the raw function body is deliberately not part of it (it belongs in the lexical channel, where it measurably helps, not in the vector).
Queries. Feed them raw, no prefix, no template. The model is tuned for short keyword bags and behaviour descriptions.
Text longer than 128 tokens. Split into overlapping windows β window 128, stride 96 β embed each, then mean-pool the window vectors and L2-renormalize. That is what the production indexer does for long doc chunks.
Retrieval. The vectors come out unit-norm, so inner product is cosine β no normalization step of your own. They are dense and 768-dimensional; an IVF/HNSW index over them behaves normally.
License & training data
Released under the MIT license.
The backbone (intfloat/multilingual-e5-base) is MIT; the teachers (gte-modernbert-base,
Qwen3-Reranker-4B) are Apache-2.0. As is standard practice for distilled embedding models,
the weights are released under MIT. Training corpus, for transparency:
| Data | License note |
|---|---|
| Repository entity texts + synthetic search queries over an MIT-licensed codebase | MIT |
| CoIR-derived code-retrieval pairs (~26% diversity stream) | per-subset CoIR licenses (research benchmark) |
| Teacher embeddings / scores (gte-modernbert, Qwen3-Reranker-4B) | Apache-2.0 teachers |
Attribution
Backbone: intfloat/multilingual-e5-base (MIT). Dense teacher: Alibaba-NLP/gte-modernbert-base (Apache-2.0). Ranking teacher: Qwen/Qwen3-Reranker-4B (Apache-2.0).