--- license: apache-2.0 base_model: Qwen/Qwen3.8-27B base_model_relation: finetune pipeline_tag: text-generation language: - en tags: - gguf - llama.cpp - qwen3_5 - pruning - depth-pruning - distillation - reasoning --- # Marlowe-22B (R1.5) **A 27B-class reasoning model compressed to 22.3B so it runs entirely on one 16 GB consumer GPU — and measured, domain by domain, against the model it came from.** Marlowe-22B removes 12 layers from **Qwen3.8-27B** and heals the loss by distilling from the 27B itself. R1.5 keeps **94.8% of the facts its parent demonstrably knows**, on **81.6% of the parent's text parameters** (22.30 B against 27.32 B), with the whole model resident in 16 GB at Q4_K_M — no layer offload, no paging. It is the second of five planned rounds. Three remain — R1.7, R1.8 and R2 — plus two acceleration variants; the roadmap at the bottom says what each is built to move. | | | |---|---| | parameters | 22.30 B — 81.6% of the parent's 27.32 B text parameters (its 0.46 B vision tower is dropped) | | layers | 52 — 36 linear-attention (DeltaNet) + 16 full-attention (parent: 64 = 48 + 16) | | hidden / intermediate | 5120 / 17408 | | KV heads × head dim | 4 × 256, on the 16 full-attention layers only | | vocabulary | 248,320 | | max position embeddings | 262,144 | | modality | text only (the parent's vision tower is not included) | | thinking | on by default; `reasoning_effort` xhigh / medium / low | | this repository | GGUF **Q4_K_M, 13.7 GB** — `marlowe-dusk-22b-r1.5.gguf` (this release) and `marlowe-dusk-22b-r1.gguf` (previous round) | ## Why compress this model in particular The parent is a hybrid: most layers are linear-attention (DeltaNet), and only every fourth is full attention. Its KV cache is therefore tiny — **64 KB per token at f16, ~34 KB at q8_0** — so long contexts are cheap to hold. What it cannot do is fit in 16 GB at a bit width that leaves its answers intact. Marlowe-22B is the same architecture, twelve layers shorter, at the size where a 4-bit build and a long context fit on one card together. ## Method 1. **Layer profiling.** Each layer scored by the KL divergence its removal induces on held-out text — measured, not assumed from depth heuristics. 2. **Depth pruning.** The twelve cheapest layers by that measure: indices `4, 5, 8, 9, 13, 14, 16, 17, 37, 38, 40, 41` — all linear-attention. Every full-attention layer survived, as did the first four and the last ten layers. 3. **Heal (R1).** LoRA (r=32, α=64) trained against the parent's **top-64 log-probs** — forward KL against a real distribution, not hard labels — over **27.5 M tokens** of a 35 M-token teacher cache (arXiv, open-web-math, AlgebraicStack, FineWeb-Edu, PG-19, plus 1,314 reasoning traces from the parent). Stopped deliberately when held-out KL flattened, rather than at budget exhaustion. 4. **Targeted repair (R1.5).** Probing located specific facts the pruned model had lost — RMSNorm's mean-of-squares, RoPE's exponent, the GELU cubic constant, Adam's bias correction. A 1.5 M-token pass on documents carrying exactly those facts closed them. The adapter is merged into the weights, and this repository ships the **Q4_K_M GGUF** built from them. The bf16 safetensors are not published yet. ## Results ### Knowledge retention against the parent Facts are mined from documents **hash-dropped against the heal corpus**, and **the parent is scored first: every item the parent itself fails is discarded** — so only knowledge the 27B demonstrably has is counted, and the student is never charged for its teacher's gaps. Denominator: **3,859 parent-verified facts** (6,002 mined items before gating). Each cell is the share of those facts the model also knows. | domain | pruned, unhealed | **R1.5** | recovered by healing | |---|---:|---:|---:| | chemistry / biology | 90.5% | **96.0%** | +5.5 | | physics | 79.8% | **85.3%** | +5.5 | | engineering | 88.9% | **91.6%** | +2.7 | | statistics | 87.4% | **89.2%** | +1.8 | | mathematics | 93.4% | **94.8%** | +1.4 | | general | 96.3% | **97.4%** | +1.1 | | machine learning | 92.4% | **93.2%** | +0.8 | | computer science | 91.2% | **89.9%** | −1.3 | | **overall** | **93.3%** | **94.8%** | **+1.5** | Five domains sit above 91%; the program's standard is **90% in every domain, not on average**, and physics, statistics and computer science are the remaining targets — see the roadmap. ### Derivation, not recall 120 generated engineering problems, parameters sampled outside tutorial values and answers computed from the defining formula, so a memorised answer scores zero: RoPE frequencies, GELU, Adam updates, softmax, cosine schedules, KV-cache sizing, LoRA parameter counts. **R1.5 answers 92 of 120.** The sharper number is the hardest 60 — the families the pruned model failed outright: | | correct | |---|---:| | pruned, unhealed | 10 / 60 | | **R1.5** | **32 / 60** | | parent (27B) | 41 / 60 | Healing recovered **71% of the gap** pruning opened on the material it damaged most. All three models ran the identical protocol, greedy, at a 6K-token thinking budget; the parent was measured at IQ3_M, a heavier quantisation than this build's Q4_K_M, so the comparison does not flatter it. (Greedy is the wrong setting for daily use — see Sampling — but it removes sampling noise from a paired comparison.) ### Arithmetic fidelity Teacher-forced on held-out worked computations, the parent predicts the next computed digit with **67.9%** top-1 accuracy; R1.5 reaches **64.1%** — a 0.115-nat gap in log-probability. KV-cache and attention-scaling arithmetic is already at parent level; rotation- and trigonometry-heavy work (RoPE, cosine schedules) is where the remaining distance lives, and it is R1.7's first target. ## How it is measured The instruments matter as much as the weights here, and every one of them is built to make a flattering result hard to get: - **Parent-gated retention bank** — the teacher is scored first and its own failures are dropped, so the headline cannot be inflated by items nobody knows. Sources are hash-dropped against the training corpus, and mined from documents rather than hand-written, so the bank cannot grade its author's homework. - **Generated derivation problems** — computed answers, parameters outside tutorial ranges, plus a recorded *recall trap* per family: the exact value a model lands on when it quotes a remembered formula with the wrong exponent. Wrong answers say *which* shortcut was taken. - **Teacher-forced digit scoring** — separates "can it do the arithmetic" from "can it run its own derivation", in five minutes and with no grading ambiguity. - **Position-resolved KL** (in progress) — KL against the parent at every position of a full reasoning trace, up to 24K tokens, on the parent's traces and the student's own. Every heal so far optimised KL on 768-token windows; this measures whether that proxy holds at long horizons. - **Pre-registered decision rules.** Each experiment's confirm/refute criteria are written down before the data exists, and results are reported against them either way. Several promising hypotheses have been refuted by their own tests and recorded as such. ## Running it **llama.cpp** — all 52 layers on the GPU, quantised KV: ```bash llama-server -m marlowe-dusk-22b-r1.5.gguf -ngl 99 \ -c 32768 -fa on -ctk q8_0 -ctv q8_0 \ --temp 1.0 --top-p 0.95 --top-k 20 \ --reasoning-format deepseek ``` **Get the file** — `marlowe-dusk-22b-r1.5.gguf` is this release; `marlowe-dusk-22b-r1.gguf` is the previous round, kept for comparison. ```bash huggingface-cli download devvexus/marlowe-22b marlowe-dusk-22b-r1.5.gguf --local-dir . ``` The chat template ships inside the GGUF, so `llama-server` applies it for you — including the thinking block, which `--reasoning-format deepseek` returns as `reasoning_content`. Any runtime that speaks GGUF and supports the hybrid `qwen3_5` architecture (DeltaNet linear attention interleaved with full attention — the architecture family Qwen3.8-27B is built on) will load it; build llama.cpp from a revision that includes that support. **Sampling — use temperature 1.0, top-p 0.95, top-k 20** (the parent's thinking preset), and do not decode greedily: at temperature 0, 24 of 59 traces that used their full thinking budget contained verbatim repetition loops. This is a property of the parent's thinking mode, inherited rather than introduced by pruning, and sampling is the fix for both models. The chat template defaults to `reasoning_effort="xhigh"`; `medium` and `low` are supported, as is `enable_thinking=False`. **Speed and memory.** On a 16 GB RTX 4080 SUPER at Q4_K_M with every layer resident and q8_0 KV: **~36 tokens/s per stream with two concurrent streams.** Weights are 13.74 GB; a 32K context adds roughly 1.1 GB. ## Roadmap — three training rounds scheduled, plus two acceleration variants | round | goal | status | |---|---|---| | R1 | heal the pruning loss broadly | done — 27.5 M tokens | | R1.5 | close specific probed fact losses | **this release** | | **R1.7** | match the 27B on long-context, reasoning-heavy work; close the rotation/trig arithmetic gap | measurement built and running | | **R1.8** | professional depth across every major engineering field | scheduled | | **R2** | tool use and agentic execution on the strengthened base | scheduled | From R1.7 onward, every round is gated on the full instrument set — retention, derivation, arithmetic fidelity and long-horizon KL — with checkpoints scored **during** training rather than only at the end, so a round that trades one capability for another is caught while it runs. That discipline came from a round whose in-loop metric improved while a capability it never measured regressed; the fix was to widen the instruments, and it is now a standing gate. Two acceleration variants are planned on the same weights: **-super** (a multi-token-prediction draft head, 1.3–1.7× on code) and **-turbo** (EAGLE-3, 3–4× on supporting runtimes). This release is the dense build: `mtp_num_hidden_layers` is 0. The same pipeline then ladders downward — an 18B distilled from this model's successor, and a 9B distilled into a native base — reusing the teacher caches and evaluation banks built here. ## Scope of this release - **GGUF only, for now.** This repository contains the Q4_K_M build and nothing else — no safetensors, config or tokenizer files, so `transformers`, vLLM and SGLang cannot load it from here. Use llama.cpp or another GGUF runtime. The bf16 weights are planned for a later upload. - **Text only.** Image and video inputs are not supported; the vision tower was dropped to spend the budget on reasoning. - **No public benchmark scores are claimed.** The held-out benchmark subsets here are too small to publish, and none were used in training. Results above are from the project's own instruments, all of which are described in full so they can be argued with. - **Long-horizon parity against the parent is being measured now** and will be reported with R1.7. - It inherits the parent's biases, refusals and knowledge cutoff; no alignment training was added. - It thinks at length. On multi-stage engineering calculations it will carry more intermediate precision than the answer needs — ask explicitly for the precision you want, and give it room. ## Training data and attribution Public data, subsampled: **arXiv, open-web-math, AlgebraicStack** (permissive-licensed subset of The Stack) and **FineWeb-Edu** — all **ODC-By**, whose attribution travels with this model — plus **PG-19** (public domain). Reasoning traces were generated by the parent model itself. No benchmark test sets were included; the corpus was hash-checked against the evaluation banks. ## License **Apache 2.0**, inherited from Qwen3.8-27B. This is a derivative work of that model; the parent's terms apply to it and to anything derived from it.