Marlowe-22B (R1.5)

A 27B-class reasoning model compressed to 22.3B so it runs entirely on one 16 GB consumer GPU — and measured, domain by domain, against the model it came from.

Marlowe-22B removes 12 layers from Qwen3.8-27B and heals the loss by distilling from the 27B itself. R1.5 keeps 94.8% of the facts its parent demonstrably knows, on 81.6% of the parent's text parameters (22.30 B against 27.32 B), with the whole model resident in 16 GB at Q4_K_M — no layer offload, no paging.

It is the second of five planned rounds. Three remain — R1.7, R1.8 and R2 — plus two acceleration variants; the roadmap at the bottom says what each is built to move.

parameters 22.30 B — 81.6% of the parent's 27.32 B text parameters (its 0.46 B vision tower is dropped)
layers 52 — 36 linear-attention (DeltaNet) + 16 full-attention (parent: 64 = 48 + 16)
hidden / intermediate 5120 / 17408
KV heads × head dim 4 × 256, on the 16 full-attention layers only
vocabulary 248,320
max position embeddings 262,144
modality text only (the parent's vision tower is not included)
thinking on by default; reasoning_effort xhigh / medium / low
this repository GGUF Q4_K_M, 13.7 GB — marlowe-dusk-22b-r1.5.gguf (this release) and marlowe-dusk-22b-r1.gguf (previous round)

Why compress this model in particular

The parent is a hybrid: most layers are linear-attention (DeltaNet), and only every fourth is full attention. Its KV cache is therefore tiny — 64 KB per token at f16, ~34 KB at q8_0 — so long contexts are cheap to hold. What it cannot do is fit in 16 GB at a bit width that leaves its answers intact. Marlowe-22B is the same architecture, twelve layers shorter, at the size where a 4-bit build and a long context fit on one card together.

Method

  1. Layer profiling. Each layer scored by the KL divergence its removal induces on held-out text — measured, not assumed from depth heuristics.
  2. Depth pruning. The twelve cheapest layers by that measure: indices 4, 5, 8, 9, 13, 14, 16, 17, 37, 38, 40, 41 — all linear-attention. Every full-attention layer survived, as did the first four and the last ten layers.
  3. Heal (R1). LoRA (r=32, α=64) trained against the parent's top-64 log-probs — forward KL against a real distribution, not hard labels — over 27.5 M tokens of a 35 M-token teacher cache (arXiv, open-web-math, AlgebraicStack, FineWeb-Edu, PG-19, plus 1,314 reasoning traces from the parent). Stopped deliberately when held-out KL flattened, rather than at budget exhaustion.
  4. Targeted repair (R1.5). Probing located specific facts the pruned model had lost — RMSNorm's mean-of-squares, RoPE's exponent, the GELU cubic constant, Adam's bias correction. A 1.5 M-token pass on documents carrying exactly those facts closed them.

The adapter is merged into the weights, and this repository ships the Q4_K_M GGUF built from them. The bf16 safetensors are not published yet.

Results

Knowledge retention against the parent

Facts are mined from documents hash-dropped against the heal corpus, and the parent is scored first: every item the parent itself fails is discarded — so only knowledge the 27B demonstrably has is counted, and the student is never charged for its teacher's gaps. Denominator: 3,859 parent-verified facts (6,002 mined items before gating). Each cell is the share of those facts the model also knows.

domain pruned, unhealed R1.5 recovered by healing
chemistry / biology 90.5% 96.0% +5.5
physics 79.8% 85.3% +5.5
engineering 88.9% 91.6% +2.7
statistics 87.4% 89.2% +1.8
mathematics 93.4% 94.8% +1.4
general 96.3% 97.4% +1.1
machine learning 92.4% 93.2% +0.8
computer science 91.2% 89.9% −1.3
overall 93.3% 94.8% +1.5

Five domains sit above 91%; the program's standard is 90% in every domain, not on average, and physics, statistics and computer science are the remaining targets — see the roadmap.

Derivation, not recall

120 generated engineering problems, parameters sampled outside tutorial values and answers computed from the defining formula, so a memorised answer scores zero: RoPE frequencies, GELU, Adam updates, softmax, cosine schedules, KV-cache sizing, LoRA parameter counts. R1.5 answers 92 of 120.

The sharper number is the hardest 60 — the families the pruned model failed outright:

correct
pruned, unhealed 10 / 60
R1.5 32 / 60
parent (27B) 41 / 60

Healing recovered 71% of the gap pruning opened on the material it damaged most. All three models ran the identical protocol, greedy, at a 6K-token thinking budget; the parent was measured at IQ3_M, a heavier quantisation than this build's Q4_K_M, so the comparison does not flatter it. (Greedy is the wrong setting for daily use — see Sampling — but it removes sampling noise from a paired comparison.)

Arithmetic fidelity

Teacher-forced on held-out worked computations, the parent predicts the next computed digit with 67.9% top-1 accuracy; R1.5 reaches 64.1% — a 0.115-nat gap in log-probability. KV-cache and attention-scaling arithmetic is already at parent level; rotation- and trigonometry-heavy work (RoPE, cosine schedules) is where the remaining distance lives, and it is R1.7's first target.

How it is measured

The instruments matter as much as the weights here, and every one of them is built to make a flattering result hard to get:

  • Parent-gated retention bank — the teacher is scored first and its own failures are dropped, so the headline cannot be inflated by items nobody knows. Sources are hash-dropped against the training corpus, and mined from documents rather than hand-written, so the bank cannot grade its author's homework.
  • Generated derivation problems — computed answers, parameters outside tutorial ranges, plus a recorded recall trap per family: the exact value a model lands on when it quotes a remembered formula with the wrong exponent. Wrong answers say which shortcut was taken.
  • Teacher-forced digit scoring — separates "can it do the arithmetic" from "can it run its own derivation", in five minutes and with no grading ambiguity.
  • Position-resolved KL (in progress) — KL against the parent at every position of a full reasoning trace, up to 24K tokens, on the parent's traces and the student's own. Every heal so far optimised KL on 768-token windows; this measures whether that proxy holds at long horizons.
  • Pre-registered decision rules. Each experiment's confirm/refute criteria are written down before the data exists, and results are reported against them either way. Several promising hypotheses have been refuted by their own tests and recorded as such.

Running it

llama.cpp — all 52 layers on the GPU, quantised KV:

llama-server -m marlowe-dusk-22b-r1.5.gguf -ngl 99 \
  -c 32768 -fa on -ctk q8_0 -ctv q8_0 \
  --temp 1.0 --top-p 0.95 --top-k 20 \
  --reasoning-format deepseek

Get the file — marlowe-dusk-22b-r1.5.gguf is this release; marlowe-dusk-22b-r1.gguf is the previous round, kept for comparison.

huggingface-cli download devvexus/marlowe-22b marlowe-dusk-22b-r1.5.gguf --local-dir .

The chat template ships inside the GGUF, so llama-server applies it for you — including the thinking block, which --reasoning-format deepseek returns as reasoning_content.

Any runtime that speaks GGUF and supports the hybrid qwen3_5 architecture (DeltaNet linear attention interleaved with full attention — the architecture family Qwen3.8-27B is built on) will load it; build llama.cpp from a revision that includes that support.

Sampling — use temperature 1.0, top-p 0.95, top-k 20 (the parent's thinking preset), and do not decode greedily: at temperature 0, 24 of 59 traces that used their full thinking budget contained verbatim repetition loops. This is a property of the parent's thinking mode, inherited rather than introduced by pruning, and sampling is the fix for both models.

The chat template defaults to reasoning_effort="xhigh"; medium and low are supported, as is enable_thinking=False.

Speed and memory. On a 16 GB RTX 4080 SUPER at Q4_K_M with every layer resident and q8_0 KV: ~36 tokens/s per stream with two concurrent streams. Weights are 13.74 GB; a 32K context adds roughly 1.1 GB.

Roadmap — three training rounds scheduled, plus two acceleration variants

round goal status
R1 heal the pruning loss broadly done — 27.5 M tokens
R1.5 close specific probed fact losses this release
R1.7 match the 27B on long-context, reasoning-heavy work; close the rotation/trig arithmetic gap measurement built and running
R1.8 professional depth across every major engineering field scheduled
R2 tool use and agentic execution on the strengthened base scheduled

From R1.7 onward, every round is gated on the full instrument set — retention, derivation, arithmetic fidelity and long-horizon KL — with checkpoints scored during training rather than only at the end, so a round that trades one capability for another is caught while it runs. That discipline came from a round whose in-loop metric improved while a capability it never measured regressed; the fix was to widen the instruments, and it is now a standing gate.

Two acceleration variants are planned on the same weights: -super (a multi-token-prediction draft head, 1.3–1.7× on code) and -turbo (EAGLE-3, 3–4× on supporting runtimes). This release is the dense build: mtp_num_hidden_layers is 0.

The same pipeline then ladders downward — an 18B distilled from this model's successor, and a 9B distilled into a native base — reusing the teacher caches and evaluation banks built here.

Scope of this release

  • GGUF only, for now. This repository contains the Q4_K_M build and nothing else — no safetensors, config or tokenizer files, so transformers, vLLM and SGLang cannot load it from here. Use llama.cpp or another GGUF runtime. The bf16 weights are planned for a later upload.
  • Text only. Image and video inputs are not supported; the vision tower was dropped to spend the budget on reasoning.
  • No public benchmark scores are claimed. The held-out benchmark subsets here are too small to publish, and none were used in training. Results above are from the project's own instruments, all of which are described in full so they can be argued with.
  • Long-horizon parity against the parent is being measured now and will be reported with R1.7.
  • It inherits the parent's biases, refusals and knowledge cutoff; no alignment training was added.
  • It thinks at length. On multi-stage engineering calculations it will carry more intermediate precision than the answer needs — ask explicitly for the precision you want, and give it room.

Training data and attribution

Public data, subsampled: arXiv, open-web-math, AlgebraicStack (permissive-licensed subset of The Stack) and FineWeb-Edu — all ODC-By, whose attribution travels with this model — plus PG-19 (public domain). Reasoning traces were generated by the parent model itself. No benchmark test sets were included; the corpus was hash-checked against the evaluation banks.

License

Apache 2.0, inherited from Qwen3.8-27B. This is a derivative work of that model; the parent's terms apply to it and to anything derived from it.

Downloads last month
805
GGUF
Model size
22B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for devvexus/marlowe-22b

Base model

Qwen/Qwen3.8-27B
Finetuned
(467)
this model