Instructions to use devvexus/marlowe-22b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use devvexus/marlowe-22b with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf devvexus/marlowe-22b # Run inference directly in the terminal: llama cli -hf devvexus/marlowe-22b
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf devvexus/marlowe-22b # Run inference directly in the terminal: llama cli -hf devvexus/marlowe-22b
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf devvexus/marlowe-22b # Run inference directly in the terminal: ./llama-cli -hf devvexus/marlowe-22b
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf devvexus/marlowe-22b # Run inference directly in the terminal: ./build/bin/llama-cli -hf devvexus/marlowe-22b
Use Docker
docker model run hf.co/devvexus/marlowe-22b
- LM Studio
- Jan
- vLLM
How to use devvexus/marlowe-22b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "devvexus/marlowe-22b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "devvexus/marlowe-22b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/devvexus/marlowe-22b
- Ollama
How to use devvexus/marlowe-22b with Ollama:
ollama run hf.co/devvexus/marlowe-22b
- Unsloth Desktop
- Pi
How to use devvexus/marlowe-22b with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf devvexus/marlowe-22b
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "devvexus/marlowe-22b" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use devvexus/marlowe-22b with Docker Model Runner:
docker model run hf.co/devvexus/marlowe-22b
- Lemonade
How to use devvexus/marlowe-22b with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull devvexus/marlowe-22b
Run and chat with the model
lemonade run user.marlowe-22b-{{QUANT_TAG}}List all available models
lemonade list
- Hermes Agent
How to use devvexus/marlowe-22b with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf devvexus/marlowe-22b
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default devvexus/marlowe-22b
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use devvexus/marlowe-22b with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf devvexus/marlowe-22b
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "devvexus/marlowe-22b" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Marlowe-22B (R1.5)
A 27B-class reasoning model compressed to 22.3B so it runs entirely on one 16 GB consumer GPU — and measured, domain by domain, against the model it came from.
Marlowe-22B removes 12 layers from Qwen3.8-27B and heals the loss by distilling from the 27B itself. R1.5 keeps 94.8% of the facts its parent demonstrably knows, on 81.6% of the parent's text parameters (22.30 B against 27.32 B), with the whole model resident in 16 GB at Q4_K_M — no layer offload, no paging.
It is the second of five planned rounds. Three remain — R1.7, R1.8 and R2 — plus two acceleration variants; the roadmap at the bottom says what each is built to move.
| parameters | 22.30 B — 81.6% of the parent's 27.32 B text parameters (its 0.46 B vision tower is dropped) |
| layers | 52 — 36 linear-attention (DeltaNet) + 16 full-attention (parent: 64 = 48 + 16) |
| hidden / intermediate | 5120 / 17408 |
| KV heads × head dim | 4 × 256, on the 16 full-attention layers only |
| vocabulary | 248,320 |
| max position embeddings | 262,144 |
| modality | text only (the parent's vision tower is not included) |
| thinking | on by default; reasoning_effort xhigh / medium / low |
| this repository | GGUF Q4_K_M, 13.7 GB — marlowe-dusk-22b-r1.5.gguf (this release) and marlowe-dusk-22b-r1.gguf (previous round) |
Why compress this model in particular
The parent is a hybrid: most layers are linear-attention (DeltaNet), and only every fourth is full attention. Its KV cache is therefore tiny — 64 KB per token at f16, ~34 KB at q8_0 — so long contexts are cheap to hold. What it cannot do is fit in 16 GB at a bit width that leaves its answers intact. Marlowe-22B is the same architecture, twelve layers shorter, at the size where a 4-bit build and a long context fit on one card together.
Method
- Layer profiling. Each layer scored by the KL divergence its removal induces on held-out text — measured, not assumed from depth heuristics.
- Depth pruning. The twelve cheapest layers by that measure: indices
4, 5, 8, 9, 13, 14, 16, 17, 37, 38, 40, 41— all linear-attention. Every full-attention layer survived, as did the first four and the last ten layers. - Heal (R1). LoRA (r=32, α=64) trained against the parent's top-64 log-probs — forward KL against a real distribution, not hard labels — over 27.5 M tokens of a 35 M-token teacher cache (arXiv, open-web-math, AlgebraicStack, FineWeb-Edu, PG-19, plus 1,314 reasoning traces from the parent). Stopped deliberately when held-out KL flattened, rather than at budget exhaustion.
- Targeted repair (R1.5). Probing located specific facts the pruned model had lost — RMSNorm's mean-of-squares, RoPE's exponent, the GELU cubic constant, Adam's bias correction. A 1.5 M-token pass on documents carrying exactly those facts closed them.
The adapter is merged into the weights, and this repository ships the Q4_K_M GGUF built from them. The bf16 safetensors are not published yet.
Results
Knowledge retention against the parent
Facts are mined from documents hash-dropped against the heal corpus, and the parent is scored first: every item the parent itself fails is discarded — so only knowledge the 27B demonstrably has is counted, and the student is never charged for its teacher's gaps. Denominator: 3,859 parent-verified facts (6,002 mined items before gating). Each cell is the share of those facts the model also knows.
| domain | pruned, unhealed | R1.5 | recovered by healing |
|---|---|---|---|
| chemistry / biology | 90.5% | 96.0% | +5.5 |
| physics | 79.8% | 85.3% | +5.5 |
| engineering | 88.9% | 91.6% | +2.7 |
| statistics | 87.4% | 89.2% | +1.8 |
| mathematics | 93.4% | 94.8% | +1.4 |
| general | 96.3% | 97.4% | +1.1 |
| machine learning | 92.4% | 93.2% | +0.8 |
| computer science | 91.2% | 89.9% | −1.3 |
| overall | 93.3% | 94.8% | +1.5 |
Five domains sit above 91%; the program's standard is 90% in every domain, not on average, and physics, statistics and computer science are the remaining targets — see the roadmap.
Derivation, not recall
120 generated engineering problems, parameters sampled outside tutorial values and answers computed from the defining formula, so a memorised answer scores zero: RoPE frequencies, GELU, Adam updates, softmax, cosine schedules, KV-cache sizing, LoRA parameter counts. R1.5 answers 92 of 120.
The sharper number is the hardest 60 — the families the pruned model failed outright:
| correct | |
|---|---|
| pruned, unhealed | 10 / 60 |
| R1.5 | 32 / 60 |
| parent (27B) | 41 / 60 |
Healing recovered 71% of the gap pruning opened on the material it damaged most. All three models ran the identical protocol, greedy, at a 6K-token thinking budget; the parent was measured at IQ3_M, a heavier quantisation than this build's Q4_K_M, so the comparison does not flatter it. (Greedy is the wrong setting for daily use — see Sampling — but it removes sampling noise from a paired comparison.)
Arithmetic fidelity
Teacher-forced on held-out worked computations, the parent predicts the next computed digit with 67.9% top-1 accuracy; R1.5 reaches 64.1% — a 0.115-nat gap in log-probability. KV-cache and attention-scaling arithmetic is already at parent level; rotation- and trigonometry-heavy work (RoPE, cosine schedules) is where the remaining distance lives, and it is R1.7's first target.
How it is measured
The instruments matter as much as the weights here, and every one of them is built to make a flattering result hard to get:
- Parent-gated retention bank — the teacher is scored first and its own failures are dropped, so the headline cannot be inflated by items nobody knows. Sources are hash-dropped against the training corpus, and mined from documents rather than hand-written, so the bank cannot grade its author's homework.
- Generated derivation problems — computed answers, parameters outside tutorial ranges, plus a recorded recall trap per family: the exact value a model lands on when it quotes a remembered formula with the wrong exponent. Wrong answers say which shortcut was taken.
- Teacher-forced digit scoring — separates "can it do the arithmetic" from "can it run its own derivation", in five minutes and with no grading ambiguity.
- Position-resolved KL (in progress) — KL against the parent at every position of a full reasoning trace, up to 24K tokens, on the parent's traces and the student's own. Every heal so far optimised KL on 768-token windows; this measures whether that proxy holds at long horizons.
- Pre-registered decision rules. Each experiment's confirm/refute criteria are written down before the data exists, and results are reported against them either way. Several promising hypotheses have been refuted by their own tests and recorded as such.
Running it
llama.cpp — all 52 layers on the GPU, quantised KV:
llama-server -m marlowe-dusk-22b-r1.5.gguf -ngl 99 \
-c 32768 -fa on -ctk q8_0 -ctv q8_0 \
--temp 1.0 --top-p 0.95 --top-k 20 \
--reasoning-format deepseek
Get the file — marlowe-dusk-22b-r1.5.gguf is this release; marlowe-dusk-22b-r1.gguf is the previous
round, kept for comparison.
huggingface-cli download devvexus/marlowe-22b marlowe-dusk-22b-r1.5.gguf --local-dir .
The chat template ships inside the GGUF, so llama-server applies it for you — including the
thinking block, which --reasoning-format deepseek returns as reasoning_content.
Any runtime that speaks GGUF and supports the hybrid qwen3_5 architecture (DeltaNet linear
attention interleaved with full attention — the architecture family Qwen3.8-27B is built on) will
load it; build llama.cpp from a revision that includes that support.
Sampling — use temperature 1.0, top-p 0.95, top-k 20 (the parent's thinking preset), and do not decode greedily: at temperature 0, 24 of 59 traces that used their full thinking budget contained verbatim repetition loops. This is a property of the parent's thinking mode, inherited rather than introduced by pruning, and sampling is the fix for both models.
The chat template defaults to reasoning_effort="xhigh"; medium and low are supported, as is
enable_thinking=False.
Speed and memory. On a 16 GB RTX 4080 SUPER at Q4_K_M with every layer resident and q8_0 KV: ~36 tokens/s per stream with two concurrent streams. Weights are 13.74 GB; a 32K context adds roughly 1.1 GB.
Roadmap — three training rounds scheduled, plus two acceleration variants
| round | goal | status |
|---|---|---|
| R1 | heal the pruning loss broadly | done — 27.5 M tokens |
| R1.5 | close specific probed fact losses | this release |
| R1.7 | match the 27B on long-context, reasoning-heavy work; close the rotation/trig arithmetic gap | measurement built and running |
| R1.8 | professional depth across every major engineering field | scheduled |
| R2 | tool use and agentic execution on the strengthened base | scheduled |
From R1.7 onward, every round is gated on the full instrument set — retention, derivation, arithmetic fidelity and long-horizon KL — with checkpoints scored during training rather than only at the end, so a round that trades one capability for another is caught while it runs. That discipline came from a round whose in-loop metric improved while a capability it never measured regressed; the fix was to widen the instruments, and it is now a standing gate.
Two acceleration variants are planned on the same weights: -super (a multi-token-prediction
draft head, 1.3–1.7× on code) and -turbo (EAGLE-3, 3–4× on supporting runtimes). This release
is the dense build: mtp_num_hidden_layers is 0.
The same pipeline then ladders downward — an 18B distilled from this model's successor, and a 9B distilled into a native base — reusing the teacher caches and evaluation banks built here.
Scope of this release
- GGUF only, for now. This repository contains the Q4_K_M build and nothing else — no
safetensors, config or tokenizer files, so
transformers, vLLM and SGLang cannot load it from here. Use llama.cpp or another GGUF runtime. The bf16 weights are planned for a later upload. - Text only. Image and video inputs are not supported; the vision tower was dropped to spend the budget on reasoning.
- No public benchmark scores are claimed. The held-out benchmark subsets here are too small to publish, and none were used in training. Results above are from the project's own instruments, all of which are described in full so they can be argued with.
- Long-horizon parity against the parent is being measured now and will be reported with R1.7.
- It inherits the parent's biases, refusals and knowledge cutoff; no alignment training was added.
- It thinks at length. On multi-stage engineering calculations it will carry more intermediate precision than the answer needs — ask explicitly for the precision you want, and give it room.
Training data and attribution
Public data, subsampled: arXiv, open-web-math, AlgebraicStack (permissive-licensed subset of The Stack) and FineWeb-Edu — all ODC-By, whose attribution travels with this model — plus PG-19 (public domain). Reasoning traces were generated by the parent model itself. No benchmark test sets were included; the corpus was hash-checked against the evaluation banks.
License
Apache 2.0, inherited from Qwen3.8-27B. This is a derivative work of that model; the parent's terms apply to it and to anything derived from it.
- Downloads last month
- 805
We're not able to determine the quantization variants.
Model tree for devvexus/marlowe-22b
Base model
Qwen/Qwen3.8-27B