marlowe-22b / README.md
devvexus's picture
Update README.md
9ca16b3 verified
|
Raw History Blame Contribute Delete
12 kB
---
license: apache-2.0
base_model: Qwen/Qwen3.8-27B
base_model_relation: finetune
pipeline_tag: text-generation
language:
- en
tags:
- gguf
- llama.cpp
- qwen3_5
- pruning
- depth-pruning
- distillation
- reasoning
---
# Marlowe-22B (R1.5)
**A 27B-class reasoning model compressed to 22.3B so it runs entirely on one 16 GB consumer GPU β€”
and measured, domain by domain, against the model it came from.**
Marlowe-22B removes 12 layers from **Qwen3.8-27B** and heals the loss by distilling from the 27B
itself. R1.5 keeps **94.8% of the facts its parent demonstrably knows**, on **81.6% of the parent's
text parameters** (22.30 B against 27.32 B), with the whole model resident in 16 GB at Q4_K_M β€” no
layer offload, no paging.
It is the second of five planned rounds. Three remain β€” R1.7, R1.8 and R2 β€” plus two acceleration
variants; the roadmap at the bottom says what each is built to move.
| | |
|---|---|
| parameters | 22.30 B β€” 81.6% of the parent's 27.32 B text parameters (its 0.46 B vision tower is dropped) |
| layers | 52 β€” 36 linear-attention (DeltaNet) + 16 full-attention (parent: 64 = 48 + 16) |
| hidden / intermediate | 5120 / 17408 |
| KV heads Γ— head dim | 4 Γ— 256, on the 16 full-attention layers only |
| vocabulary | 248,320 |
| max position embeddings | 262,144 |
| modality | text only (the parent's vision tower is not included) |
| thinking | on by default; `reasoning_effort` xhigh / medium / low |
| this repository | GGUF **Q4_K_M, 13.7 GB** β€” `marlowe-dusk-22b-r1.5.gguf` (this release) and `marlowe-dusk-22b-r1.gguf` (previous round) |
## Why compress this model in particular
The parent is a hybrid: most layers are linear-attention (DeltaNet), and only every fourth is full
attention. Its KV cache is therefore tiny β€” **64 KB per token at f16, ~34 KB at q8_0** β€” so long
contexts are cheap to hold. What it cannot do is fit in 16 GB at a bit width that leaves its
answers intact. Marlowe-22B is the same architecture, twelve layers shorter, at the size where a
4-bit build and a long context fit on one card together.
## Method
1. **Layer profiling.** Each layer scored by the KL divergence its removal induces on held-out
text β€” measured, not assumed from depth heuristics.
2. **Depth pruning.** The twelve cheapest layers by that measure: indices
`4, 5, 8, 9, 13, 14, 16, 17, 37, 38, 40, 41` β€” all linear-attention. Every full-attention layer
survived, as did the first four and the last ten layers.
3. **Heal (R1).** LoRA (r=32, Ξ±=64) trained against the parent's **top-64 log-probs** β€” forward KL
against a real distribution, not hard labels β€” over **27.5 M tokens** of a 35 M-token teacher
cache (arXiv, open-web-math, AlgebraicStack, FineWeb-Edu, PG-19, plus 1,314 reasoning traces
from the parent). Stopped deliberately when held-out KL flattened, rather than at budget
exhaustion.
4. **Targeted repair (R1.5).** Probing located specific facts the pruned model had lost β€” RMSNorm's
mean-of-squares, RoPE's exponent, the GELU cubic constant, Adam's bias correction. A 1.5 M-token
pass on documents carrying exactly those facts closed them.
The adapter is merged into the weights, and this repository ships the **Q4_K_M GGUF** built from
them. The bf16 safetensors are not published yet.
## Results
### Knowledge retention against the parent
Facts are mined from documents **hash-dropped against the heal corpus**, and **the parent is scored
first: every item the parent itself fails is discarded** β€” so only knowledge the 27B demonstrably
has is counted, and the student is never charged for its teacher's gaps. Denominator: **3,859
parent-verified facts** (6,002 mined items before gating). Each cell is the share of those facts
the model also knows.
| domain | pruned, unhealed | **R1.5** | recovered by healing |
|---|---:|---:|---:|
| chemistry / biology | 90.5% | **96.0%** | +5.5 |
| physics | 79.8% | **85.3%** | +5.5 |
| engineering | 88.9% | **91.6%** | +2.7 |
| statistics | 87.4% | **89.2%** | +1.8 |
| mathematics | 93.4% | **94.8%** | +1.4 |
| general | 96.3% | **97.4%** | +1.1 |
| machine learning | 92.4% | **93.2%** | +0.8 |
| computer science | 91.2% | **89.9%** | βˆ’1.3 |
| **overall** | **93.3%** | **94.8%** | **+1.5** |
Five domains sit above 91%; the program's standard is **90% in every domain, not on average**, and
physics, statistics and computer science are the remaining targets β€” see the roadmap.
### Derivation, not recall
120 generated engineering problems, parameters sampled outside tutorial values and answers computed
from the defining formula, so a memorised answer scores zero: RoPE frequencies, GELU, Adam updates,
softmax, cosine schedules, KV-cache sizing, LoRA parameter counts. **R1.5 answers 92 of 120.**
The sharper number is the hardest 60 β€” the families the pruned model failed outright:
| | correct |
|---|---:|
| pruned, unhealed | 10 / 60 |
| **R1.5** | **32 / 60** |
| parent (27B) | 41 / 60 |
Healing recovered **71% of the gap** pruning opened on the material it damaged most. All three
models ran the identical protocol, greedy, at a 6K-token thinking budget; the parent was measured
at IQ3_M, a heavier quantisation than this build's Q4_K_M, so the comparison does not flatter it.
(Greedy is the wrong setting for daily use β€” see Sampling β€” but it removes sampling noise from a
paired comparison.)
### Arithmetic fidelity
Teacher-forced on held-out worked computations, the parent predicts the next computed digit with
**67.9%** top-1 accuracy; R1.5 reaches **64.1%** β€” a 0.115-nat gap in log-probability. KV-cache and
attention-scaling arithmetic is already at parent level; rotation- and trigonometry-heavy work
(RoPE, cosine schedules) is where the remaining distance lives, and it is R1.7's first target.
## How it is measured
The instruments matter as much as the weights here, and every one of them is built to make a
flattering result hard to get:
- **Parent-gated retention bank** β€” the teacher is scored first and its own failures are dropped,
so the headline cannot be inflated by items nobody knows. Sources are hash-dropped against the
training corpus, and mined from documents rather than hand-written, so the bank cannot grade its
author's homework.
- **Generated derivation problems** β€” computed answers, parameters outside tutorial ranges, plus a
recorded *recall trap* per family: the exact value a model lands on when it quotes a remembered
formula with the wrong exponent. Wrong answers say *which* shortcut was taken.
- **Teacher-forced digit scoring** β€” separates "can it do the arithmetic" from "can it run its own
derivation", in five minutes and with no grading ambiguity.
- **Position-resolved KL** (in progress) β€” KL against the parent at every position of a full
reasoning trace, up to 24K tokens, on the parent's traces and the student's own. Every heal so
far optimised KL on 768-token windows; this measures whether that proxy holds at long horizons.
- **Pre-registered decision rules.** Each experiment's confirm/refute criteria are written down
before the data exists, and results are reported against them either way. Several promising
hypotheses have been refuted by their own tests and recorded as such.
## Running it
**llama.cpp** β€” all 52 layers on the GPU, quantised KV:
```bash
llama-server -m marlowe-dusk-22b-r1.5.gguf -ngl 99 \
-c 32768 -fa on -ctk q8_0 -ctv q8_0 \
--temp 1.0 --top-p 0.95 --top-k 20 \
--reasoning-format deepseek
```
**Get the file** β€” `marlowe-dusk-22b-r1.5.gguf` is this release; `marlowe-dusk-22b-r1.gguf` is the previous
round, kept for comparison.
```bash
huggingface-cli download devvexus/marlowe-22b marlowe-dusk-22b-r1.5.gguf --local-dir .
```
The chat template ships inside the GGUF, so `llama-server` applies it for you β€” including the
thinking block, which `--reasoning-format deepseek` returns as `reasoning_content`.
Any runtime that speaks GGUF and supports the hybrid `qwen3_5` architecture (DeltaNet linear
attention interleaved with full attention β€” the architecture family Qwen3.8-27B is built on) will
load it; build llama.cpp from a revision that includes that support.
**Sampling β€” use temperature 1.0, top-p 0.95, top-k 20** (the parent's thinking preset), and do not
decode greedily: at temperature 0, 24 of 59 traces that used their full thinking budget contained
verbatim repetition loops. This is a property of the parent's thinking mode, inherited rather than
introduced by pruning, and sampling is the fix for both models.
The chat template defaults to `reasoning_effort="xhigh"`; `medium` and `low` are supported, as is
`enable_thinking=False`.
**Speed and memory.** On a 16 GB RTX 4080 SUPER at Q4_K_M with every layer resident and q8_0 KV:
**~36 tokens/s per stream with two concurrent streams.** Weights are 13.74 GB; a 32K context adds
roughly 1.1 GB.
## Roadmap β€” three training rounds scheduled, plus two acceleration variants
| round | goal | status |
|---|---|---|
| R1 | heal the pruning loss broadly | done β€” 27.5 M tokens |
| R1.5 | close specific probed fact losses | **this release** |
| **R1.7** | match the 27B on long-context, reasoning-heavy work; close the rotation/trig arithmetic gap | measurement built and running |
| **R1.8** | professional depth across every major engineering field | scheduled |
| **R2** | tool use and agentic execution on the strengthened base | scheduled |
From R1.7 onward, every round is gated on the full instrument set β€” retention, derivation,
arithmetic fidelity and long-horizon KL β€” with checkpoints scored **during** training rather than
only at the end, so a round that trades one capability for another is caught while it runs. That
discipline came from a round whose in-loop metric improved while a capability it never measured
regressed; the fix was to widen the instruments, and it is now a standing gate.
Two acceleration variants are planned on the same weights: **-super** (a multi-token-prediction
draft head, 1.3–1.7Γ— on code) and **-turbo** (EAGLE-3, 3–4Γ— on supporting runtimes). This release
is the dense build: `mtp_num_hidden_layers` is 0.
The same pipeline then ladders downward β€” an 18B distilled from this model's successor, and a 9B
distilled into a native base β€” reusing the teacher caches and evaluation banks built here.
## Scope of this release
- **GGUF only, for now.** This repository contains the Q4_K_M build and nothing else β€” no
safetensors, config or tokenizer files, so `transformers`, vLLM and SGLang cannot load it from
here. Use llama.cpp or another GGUF runtime. The bf16 weights are planned for a later upload.
- **Text only.** Image and video inputs are not supported; the vision tower was dropped to spend
the budget on reasoning.
- **No public benchmark scores are claimed.** The held-out benchmark subsets here are too small to
publish, and none were used in training. Results above are from the project's own instruments,
all of which are described in full so they can be argued with.
- **Long-horizon parity against the parent is being measured now** and will be reported with R1.7.
- It inherits the parent's biases, refusals and knowledge cutoff; no alignment training was added.
- It thinks at length. On multi-stage engineering calculations it will carry more intermediate
precision than the answer needs β€” ask explicitly for the precision you want, and give it room.
## Training data and attribution
Public data, subsampled: **arXiv, open-web-math, AlgebraicStack** (permissive-licensed subset of
The Stack) and **FineWeb-Edu** β€” all **ODC-By**, whose attribution travels with this model β€” plus
**PG-19** (public domain). Reasoning traces were generated by the parent model itself. No benchmark
test sets were included; the corpus was hash-checked against the evaluation banks.
## License
**Apache 2.0**, inherited from Qwen3.8-27B. This is a derivative work of that model; the parent's
terms apply to it and to anything derived from it.