code-daemon-enrich-v1

A distilled Qwen3.5-0.8B worker that writes the short, structured labels in the UltraCode code-intelligence pipeline: RAPTOR L0/L1 cluster labels, community labels, and link-selection picks. It replaces a much larger teacher on exactly the high-volume, prefill-bound, short-output stages where a sub-billion-parameter model is enough — running them on the daemon's dedicated .enrich worker at a fraction of the main LLM's per-call prefill cost.

This is a purpose-built pipeline component, not a general assistant. It only does the four label tasks below; outside that distribution its behaviour is undefined.

⚠ The weights under this id changed on 2026-09-16

This card describes the Qwen3.5-0.8B SeqKD student. Until 2026-09-16 the same id served a Qwen3-0.6B distilled from Qwen2.5-7B (2026-07-02); that model is retired and its file and card are in this repository's git history. Do not mix figures across the two — everything here was measured on the 0.8B.

What it is

  • Base: Qwen/Qwen3.5-0.8B — 24 layers, ChatML. Hybrid attention: 18 of the 24 layers are gated DeltaNet, 6 are attention. That is a deployment fact, not trivia — see Memory below.
  • Teacher: Qwen3.8-27B, answering the same prompts the daemon sends in production.
  • Method: sequence-level knowledge distillation (SeqKD), merged into the base and exported to GGUF without the base's multi-token-prediction block.
  • Format: code-daemon-enrich-v1-Q8_0.gguf (774 MB, Q8_0, llama.cpp ≥ b10809). Q8_0 keeps a small model's logits crisp for short noun-phrase / single-pick outputs.
  • Thinking: the model is trained with an empty think block closing every prompt (.on_suppress_empty). Send it that way or the outputs drift.

What it does — the four label tasks

  • SYS_RAPTOR_LABEL_L0 / L1 — a cluster label, "<topic>: name1, name2, name3". → "DI service provider construction: BuildServiceProvider, GetService, ServiceProviderOptions"
  • SYS_COMMUNITY_LABEL — a graph-community label, a 2-to-5-word noun phrase, nothing else. → "Go standard library packages"
  • SYS_LINK_SELECT — candidate ids to link, c<N> picks or none.

Evaluation

215 held-out rows for the trained buckets and all 41 LINK_SELECT records (never trained on), scored as the GGUF the daemon actually loads, through llama-server on the daemon's own llama.cpp. Two comparators: the untouched Qwen3.5-0.8B (what you get without the distillation) and the retired 0.6B (what this replaced).

Q8_0 GGUF trained buckets rougeL exact LINK_SELECT exact (41)
this model 0.493 0.093 18 / 41
stock Qwen3.5-0.8B, same converter 0.193 0.005 14 / 41
the retired 0.6B 0.220 0.005 11 / 41

Paired, on the trained buckets: +0.300 over the stock base (95 % CI [+0.257, +0.343], p 0.0003, 158 wins / 22 losses) and +0.273 over the 0.6B. LINK_SELECT was not trained and did not regress — 4 discordant rows, all its way, which 41 records cannot call a gain.

What the score does not say: three failures it does not have

rougeL orders the models; it does not say what separates them. These three properties need no reference at all — they are properties of the answer alone, so they can be read in production:

this model the retired 0.6B stock 0.8B teacher
community labels answered as a bare identifier list 0.000 0.382 0.009 0.000
L0 in the contract's <phrase>: <members> shape 1.000 0.667 0.598 1.000
L1 answer length, median words 6 33 33 7
LINK_SELECT answered none, of 41 6 0 2 24
  • The 0.6B listed the input instead of naming it: "errors, fmt, bytes, log" where the teacher writes "Go standard library packages". A community label is a cluster's name in search and in generated docs, so a list of four members answers nothing.
  • It did not fold a level: at L1 it returned its children's labels concatenated — 33 words against the teacher's 7 — which is the one thing that tier exists to avoid.
  • It could not decline a link: the teacher answers none to 24 of 41 link questions; the 0.6B never did, and picked the same option in 30 of 41. Every ambiguous unit became an edge.

Where this model is worse

It loses 24 of the 102 L0 rows, and the losses have one shape: it paraphrases what should be copied. "Retry with delay utility: AttemptWithDelay, iter, dur" comes back as "Async function retrying: AttemptWithDelay, iterations, elapsed" — iterations and elapsed are not in the input. Where the right answer is a verbatim list of identifiers, the 0.6B's copying wins. 73 wins to 24 is worth the trade, but if your use needs exact identifier echo, measure that first.

Speed

Measured live inside the daemon (laptop RTX 5060 8 GB, CUDA, Q8_0, n_ctx=8192, other workers on the same GPU), per batch from the worker's own profile lines, against the retired 0.6B measured the same day on the same machine:

regime this model the retired 0.6B
decode, single stream 253 tok/s 305 tok/s
decode, 8 concurrent label slots 794 tok/s 1 083 tok/s

Decode is 17 % slower single-stream and 27 % batched — the honest cost of a larger, hybrid network. It matters far less than it looks, because these stages are prefill-bound. The same two runs, measured end to end over the whole label stage with prompt tokens counted:

label stage, end to end this model the retired 0.6B
throughput 5 858 tok/s 6 265 tok/s
the run behind it 48 calls, 12 258 in + 606 out, 2.2 s 27 calls, 61 802 in + 4 796 out, 10.6 s

A 27 % slower decode costs 6.5 % of the stage. Plan capacity from this table, not from the one above. (The two arms ran on different corpora in different phases of the same daemon, so read the ratios, not the third digit.)

Memory

2 281 MB resident at n_ctx=8192 with the default 28 sequences — measured from the per-process GPU counter, not estimated. The breakdown matters because the hybrid architecture spends it differently from a dense model of the same size:

buffer MiB
weights 763.78
recurrent state (28 sequences) 539.44
KV cache (6 attention layers) 96.00
compute (+ host) 101.52 + 12.49
output 26.52

The recurrent state is per sequence — 19.27 MiB each — so it scales with n_seq_max, not with context length, and it is the one line a dense model does not have. Size the deployment from 2 281 MB, not from the 774 MB file; halve the sequences and you get ~270 MiB back.

Usage (llama.cpp)

# ChatML, ONE user turn, no system turn, and an empty think block before the answer.
llama-cli -m code-daemon-enrich-v1-Q8_0.gguf -c 8192 \
  -p '<|im_start|>user
Write ONE line — 2 to 5 plain-English words — labelling this group. No quotes, no explanation.

Group members:
- parseArgs
- Command
- Usage
<|im_end|>
<|im_start|>assistant
<think>

</think>

'

Greedy decoding (temperature 0) is recommended — the outputs are factual labels.

License & attribution

Apache-2.0, matching the Qwen/Qwen3.5-0.8B base. Not legal advice — check the base and teacher model cards before redistributing. Base and teacher © the Qwen team; please also honour their cards.

Downloads last month
45
GGUF
Model size
0.8B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for faxenoff/code-daemon-enrich-v1

Quantized
(305)
this model