Instructions to use faxenoff/code-daemon-enrich-v1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use faxenoff/code-daemon-enrich-v1 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf faxenoff/code-daemon-enrich-v1:Q8_0 # Run inference directly in the terminal: llama cli -hf faxenoff/code-daemon-enrich-v1:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf faxenoff/code-daemon-enrich-v1:Q8_0 # Run inference directly in the terminal: llama cli -hf faxenoff/code-daemon-enrich-v1:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf faxenoff/code-daemon-enrich-v1:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf faxenoff/code-daemon-enrich-v1:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf faxenoff/code-daemon-enrich-v1:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf faxenoff/code-daemon-enrich-v1:Q8_0
Use Docker
docker model run hf.co/faxenoff/code-daemon-enrich-v1:Q8_0
- LM Studio
- Jan
- vLLM
How to use faxenoff/code-daemon-enrich-v1 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "faxenoff/code-daemon-enrich-v1" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "faxenoff/code-daemon-enrich-v1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/faxenoff/code-daemon-enrich-v1:Q8_0
- Ollama
How to use faxenoff/code-daemon-enrich-v1 with Ollama:
ollama run hf.co/faxenoff/code-daemon-enrich-v1:Q8_0
- Unsloth Desktop
- Pi
How to use faxenoff/code-daemon-enrich-v1 with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf faxenoff/code-daemon-enrich-v1:Q8_0
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "faxenoff/code-daemon-enrich-v1:Q8_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use faxenoff/code-daemon-enrich-v1 with Docker Model Runner:
docker model run hf.co/faxenoff/code-daemon-enrich-v1:Q8_0
- Lemonade
How to use faxenoff/code-daemon-enrich-v1 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull faxenoff/code-daemon-enrich-v1:Q8_0
Run and chat with the model
lemonade run user.code-daemon-enrich-v1-Q8_0
List all available models
lemonade list
- Hermes Agent
How to use faxenoff/code-daemon-enrich-v1 with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf faxenoff/code-daemon-enrich-v1:Q8_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default faxenoff/code-daemon-enrich-v1:Q8_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use faxenoff/code-daemon-enrich-v1 with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf faxenoff/code-daemon-enrich-v1:Q8_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "faxenoff/code-daemon-enrich-v1:Q8_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
code-daemon-enrich-v1
A distilled Qwen3.5-0.8B worker that writes the short, structured labels in the UltraCode
code-intelligence pipeline: RAPTOR L0/L1 cluster labels, community labels, and link-selection
picks. It replaces a much larger teacher on exactly the high-volume, prefill-bound,
short-output stages where a sub-billion-parameter model is enough — running them on the daemon's
dedicated .enrich worker at a fraction of the main LLM's per-call prefill cost.
This is a purpose-built pipeline component, not a general assistant. It only does the four label tasks below; outside that distribution its behaviour is undefined.
⚠ The weights under this id changed on 2026-09-16
This card describes the Qwen3.5-0.8B SeqKD student. Until 2026-09-16 the same id served a Qwen3-0.6B distilled from Qwen2.5-7B (2026-07-02); that model is retired and its file and card are in this repository's git history. Do not mix figures across the two — everything here was measured on the 0.8B.
What it is
- Base:
Qwen/Qwen3.5-0.8B— 24 layers, ChatML. Hybrid attention: 18 of the 24 layers are gated DeltaNet, 6 are attention. That is a deployment fact, not trivia — see Memory below. - Teacher:
Qwen3.8-27B, answering the same prompts the daemon sends in production. - Method: sequence-level knowledge distillation (SeqKD), merged into the base and exported to GGUF without the base's multi-token-prediction block.
- Format:
code-daemon-enrich-v1-Q8_0.gguf(774 MB, Q8_0, llama.cpp ≥ b10809). Q8_0 keeps a small model's logits crisp for short noun-phrase / single-pick outputs. - Thinking: the model is trained with an empty think block closing every prompt
(
.on_suppress_empty). Send it that way or the outputs drift.
What it does — the four label tasks
SYS_RAPTOR_LABEL_L0/L1— a cluster label,"<topic>: name1, name2, name3". →"DI service provider construction: BuildServiceProvider, GetService, ServiceProviderOptions"SYS_COMMUNITY_LABEL— a graph-community label, a 2-to-5-word noun phrase, nothing else. →"Go standard library packages"SYS_LINK_SELECT— candidate ids to link,c<N>picks ornone.
Evaluation
215 held-out rows for the trained buckets and all 41 LINK_SELECT records (never trained on),
scored as the GGUF the daemon actually loads, through llama-server on the daemon's own llama.cpp.
Two comparators: the untouched Qwen3.5-0.8B (what you get without the distillation) and the
retired 0.6B (what this replaced).
| Q8_0 GGUF | trained buckets rougeL | exact | LINK_SELECT exact (41) |
|---|---|---|---|
| this model | 0.493 | 0.093 | 18 / 41 |
| stock Qwen3.5-0.8B, same converter | 0.193 | 0.005 | 14 / 41 |
| the retired 0.6B | 0.220 | 0.005 | 11 / 41 |
Paired, on the trained buckets: +0.300 over the stock base (95 % CI [+0.257, +0.343],
p 0.0003, 158 wins / 22 losses) and +0.273 over the 0.6B. LINK_SELECT was not trained and did
not regress — 4 discordant rows, all its way, which 41 records cannot call a gain.
What the score does not say: three failures it does not have
rougeL orders the models; it does not say what separates them. These three properties need no reference at all — they are properties of the answer alone, so they can be read in production:
| this model | the retired 0.6B | stock 0.8B | teacher | |
|---|---|---|---|---|
| community labels answered as a bare identifier list | 0.000 | 0.382 | 0.009 | 0.000 |
L0 in the contract's <phrase>: <members> shape |
1.000 | 0.667 | 0.598 | 1.000 |
| L1 answer length, median words | 6 | 33 | 33 | 7 |
LINK_SELECT answered none, of 41 |
6 | 0 | 2 | 24 |
- The 0.6B listed the input instead of naming it:
"errors, fmt, bytes, log"where the teacher writes"Go standard library packages". A community label is a cluster's name in search and in generated docs, so a list of four members answers nothing. - It did not fold a level: at L1 it returned its children's labels concatenated — 33 words against the teacher's 7 — which is the one thing that tier exists to avoid.
- It could not decline a link: the teacher answers
noneto 24 of 41 link questions; the 0.6B never did, and picked the same option in 30 of 41. Every ambiguous unit became an edge.
Where this model is worse
It loses 24 of the 102 L0 rows, and the losses have one shape: it paraphrases what should be
copied. "Retry with delay utility: AttemptWithDelay, iter, dur" comes back as "Async function retrying: AttemptWithDelay, iterations, elapsed" — iterations and elapsed are not in the
input. Where the right answer is a verbatim list of identifiers, the 0.6B's copying wins. 73 wins
to 24 is worth the trade, but if your use needs exact identifier echo, measure that first.
Speed
Measured live inside the daemon (laptop RTX 5060 8 GB, CUDA, Q8_0, n_ctx=8192, other workers
on the same GPU), per batch from the worker's own profile lines, against the retired 0.6B measured
the same day on the same machine:
| regime | this model | the retired 0.6B |
|---|---|---|
| decode, single stream | 253 tok/s | 305 tok/s |
| decode, 8 concurrent label slots | 794 tok/s | 1 083 tok/s |
Decode is 17 % slower single-stream and 27 % batched — the honest cost of a larger, hybrid network. It matters far less than it looks, because these stages are prefill-bound. The same two runs, measured end to end over the whole label stage with prompt tokens counted:
| label stage, end to end | this model | the retired 0.6B |
|---|---|---|
| throughput | 5 858 tok/s | 6 265 tok/s |
| the run behind it | 48 calls, 12 258 in + 606 out, 2.2 s | 27 calls, 61 802 in + 4 796 out, 10.6 s |
A 27 % slower decode costs 6.5 % of the stage. Plan capacity from this table, not from the one above. (The two arms ran on different corpora in different phases of the same daemon, so read the ratios, not the third digit.)
Memory
2 281 MB resident at n_ctx=8192 with the default 28 sequences — measured from the
per-process GPU counter, not estimated. The breakdown matters because the hybrid architecture
spends it differently from a dense model of the same size:
| buffer | MiB |
|---|---|
| weights | 763.78 |
| recurrent state (28 sequences) | 539.44 |
| KV cache (6 attention layers) | 96.00 |
| compute (+ host) | 101.52 + 12.49 |
| output | 26.52 |
The recurrent state is per sequence — 19.27 MiB each — so it scales with n_seq_max, not with
context length, and it is the one line a dense model does not have. Size the deployment from
2 281 MB, not from the 774 MB file; halve the sequences and you get ~270 MiB back.
Usage (llama.cpp)
# ChatML, ONE user turn, no system turn, and an empty think block before the answer.
llama-cli -m code-daemon-enrich-v1-Q8_0.gguf -c 8192 \
-p '<|im_start|>user
Write ONE line — 2 to 5 plain-English words — labelling this group. No quotes, no explanation.
Group members:
- parseArgs
- Command
- Usage
<|im_end|>
<|im_start|>assistant
<think>
</think>
'
Greedy decoding (temperature 0) is recommended — the outputs are factual labels.
License & attribution
Apache-2.0, matching the Qwen/Qwen3.5-0.8B base.
Not legal advice — check the base and teacher model cards before redistributing. Base and teacher
© the Qwen team; please also honour their cards.
- Downloads last month
- 45
8-bit