File size: 12,021 Bytes
e866f76
 
5d32e00
9ca16b3
5d32e00
 
 
 
9ca16b3
 
5d32e00
 
 
 
 
5bcfcc0
5d32e00
 
 
 
 
 
 
9ca16b3
 
 
5d32e00
9ca16b3
 
5d32e00
 
 
9ca16b3
5d32e00
 
 
 
 
 
 
9ca16b3
5d32e00
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9ca16b3
 
5d32e00
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9ca16b3
 
 
 
 
5d32e00
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9ca16b3
5d32e00
 
 
 
 
9ca16b3
 
 
 
 
5d32e00
 
9ca16b3
 
 
 
 
 
 
5d32e00
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9ca16b3
 
 
 
 
5d32e00
 
 
 
 
 
 
 
 
 
9ca16b3
 
 
5d32e00
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
---
license: apache-2.0
base_model: Qwen/Qwen3.8-27B
base_model_relation: finetune
pipeline_tag: text-generation
language:
- en
tags:
- gguf
- llama.cpp
- qwen3_5
- pruning
- depth-pruning
- distillation
- reasoning
---

# Marlowe-22B (R1.5)

**A 27B-class reasoning model compressed to 22.3B so it runs entirely on one 16 GB consumer GPU β€”
and measured, domain by domain, against the model it came from.**

Marlowe-22B removes 12 layers from **Qwen3.8-27B** and heals the loss by distilling from the 27B
itself. R1.5 keeps **94.8% of the facts its parent demonstrably knows**, on **81.6% of the parent's
text parameters** (22.30 B against 27.32 B), with the whole model resident in 16 GB at Q4_K_M β€” no
layer offload, no paging.

It is the second of five planned rounds. Three remain β€” R1.7, R1.8 and R2 β€” plus two acceleration
variants; the roadmap at the bottom says what each is built to move.

| | |
|---|---|
| parameters | 22.30 B β€” 81.6% of the parent's 27.32 B text parameters (its 0.46 B vision tower is dropped) |
| layers | 52 β€” 36 linear-attention (DeltaNet) + 16 full-attention (parent: 64 = 48 + 16) |
| hidden / intermediate | 5120 / 17408 |
| KV heads Γ— head dim | 4 Γ— 256, on the 16 full-attention layers only |
| vocabulary | 248,320 |
| max position embeddings | 262,144 |
| modality | text only (the parent's vision tower is not included) |
| thinking | on by default; `reasoning_effort` xhigh / medium / low |
| this repository | GGUF **Q4_K_M, 13.7 GB** β€” `marlowe-dusk-22b-r1.5.gguf` (this release) and `marlowe-dusk-22b-r1.gguf` (previous round) |

## Why compress this model in particular

The parent is a hybrid: most layers are linear-attention (DeltaNet), and only every fourth is full
attention. Its KV cache is therefore tiny β€” **64 KB per token at f16, ~34 KB at q8_0** β€” so long
contexts are cheap to hold. What it cannot do is fit in 16 GB at a bit width that leaves its
answers intact. Marlowe-22B is the same architecture, twelve layers shorter, at the size where a
4-bit build and a long context fit on one card together.

## Method

1. **Layer profiling.** Each layer scored by the KL divergence its removal induces on held-out
   text β€” measured, not assumed from depth heuristics.
2. **Depth pruning.** The twelve cheapest layers by that measure: indices
   `4, 5, 8, 9, 13, 14, 16, 17, 37, 38, 40, 41` β€” all linear-attention. Every full-attention layer
   survived, as did the first four and the last ten layers.
3. **Heal (R1).** LoRA (r=32, Ξ±=64) trained against the parent's **top-64 log-probs** β€” forward KL
   against a real distribution, not hard labels β€” over **27.5 M tokens** of a 35 M-token teacher
   cache (arXiv, open-web-math, AlgebraicStack, FineWeb-Edu, PG-19, plus 1,314 reasoning traces
   from the parent). Stopped deliberately when held-out KL flattened, rather than at budget
   exhaustion.
4. **Targeted repair (R1.5).** Probing located specific facts the pruned model had lost β€” RMSNorm's
   mean-of-squares, RoPE's exponent, the GELU cubic constant, Adam's bias correction. A 1.5 M-token
   pass on documents carrying exactly those facts closed them.

The adapter is merged into the weights, and this repository ships the **Q4_K_M GGUF** built from
them. The bf16 safetensors are not published yet.

## Results

### Knowledge retention against the parent

Facts are mined from documents **hash-dropped against the heal corpus**, and **the parent is scored
first: every item the parent itself fails is discarded** β€” so only knowledge the 27B demonstrably
has is counted, and the student is never charged for its teacher's gaps. Denominator: **3,859
parent-verified facts** (6,002 mined items before gating). Each cell is the share of those facts
the model also knows.

| domain | pruned, unhealed | **R1.5** | recovered by healing |
|---|---:|---:|---:|
| chemistry / biology | 90.5% | **96.0%** | +5.5 |
| physics | 79.8% | **85.3%** | +5.5 |
| engineering | 88.9% | **91.6%** | +2.7 |
| statistics | 87.4% | **89.2%** | +1.8 |
| mathematics | 93.4% | **94.8%** | +1.4 |
| general | 96.3% | **97.4%** | +1.1 |
| machine learning | 92.4% | **93.2%** | +0.8 |
| computer science | 91.2% | **89.9%** | βˆ’1.3 |
| **overall** | **93.3%** | **94.8%** | **+1.5** |

Five domains sit above 91%; the program's standard is **90% in every domain, not on average**, and
physics, statistics and computer science are the remaining targets β€” see the roadmap.

### Derivation, not recall

120 generated engineering problems, parameters sampled outside tutorial values and answers computed
from the defining formula, so a memorised answer scores zero: RoPE frequencies, GELU, Adam updates,
softmax, cosine schedules, KV-cache sizing, LoRA parameter counts. **R1.5 answers 92 of 120.**

The sharper number is the hardest 60 β€” the families the pruned model failed outright:

| | correct |
|---|---:|
| pruned, unhealed | 10 / 60 |
| **R1.5** | **32 / 60** |
| parent (27B) | 41 / 60 |

Healing recovered **71% of the gap** pruning opened on the material it damaged most. All three
models ran the identical protocol, greedy, at a 6K-token thinking budget; the parent was measured
at IQ3_M, a heavier quantisation than this build's Q4_K_M, so the comparison does not flatter it.
(Greedy is the wrong setting for daily use β€” see Sampling β€” but it removes sampling noise from a
paired comparison.)

### Arithmetic fidelity

Teacher-forced on held-out worked computations, the parent predicts the next computed digit with
**67.9%** top-1 accuracy; R1.5 reaches **64.1%** β€” a 0.115-nat gap in log-probability. KV-cache and
attention-scaling arithmetic is already at parent level; rotation- and trigonometry-heavy work
(RoPE, cosine schedules) is where the remaining distance lives, and it is R1.7's first target.

## How it is measured

The instruments matter as much as the weights here, and every one of them is built to make a
flattering result hard to get:

- **Parent-gated retention bank** β€” the teacher is scored first and its own failures are dropped,
  so the headline cannot be inflated by items nobody knows. Sources are hash-dropped against the
  training corpus, and mined from documents rather than hand-written, so the bank cannot grade its
  author's homework.
- **Generated derivation problems** β€” computed answers, parameters outside tutorial ranges, plus a
  recorded *recall trap* per family: the exact value a model lands on when it quotes a remembered
  formula with the wrong exponent. Wrong answers say *which* shortcut was taken.
- **Teacher-forced digit scoring** β€” separates "can it do the arithmetic" from "can it run its own
  derivation", in five minutes and with no grading ambiguity.
- **Position-resolved KL** (in progress) β€” KL against the parent at every position of a full
  reasoning trace, up to 24K tokens, on the parent's traces and the student's own. Every heal so
  far optimised KL on 768-token windows; this measures whether that proxy holds at long horizons.
- **Pre-registered decision rules.** Each experiment's confirm/refute criteria are written down
  before the data exists, and results are reported against them either way. Several promising
  hypotheses have been refuted by their own tests and recorded as such.

## Running it

**llama.cpp** β€” all 52 layers on the GPU, quantised KV:

```bash
llama-server -m marlowe-dusk-22b-r1.5.gguf -ngl 99 \
  -c 32768 -fa on -ctk q8_0 -ctv q8_0 \
  --temp 1.0 --top-p 0.95 --top-k 20 \
  --reasoning-format deepseek
```

**Get the file** β€” `marlowe-dusk-22b-r1.5.gguf` is this release; `marlowe-dusk-22b-r1.gguf` is the previous
round, kept for comparison.

```bash
huggingface-cli download devvexus/marlowe-22b marlowe-dusk-22b-r1.5.gguf --local-dir .
```

The chat template ships inside the GGUF, so `llama-server` applies it for you β€” including the
thinking block, which `--reasoning-format deepseek` returns as `reasoning_content`.

Any runtime that speaks GGUF and supports the hybrid `qwen3_5` architecture (DeltaNet linear
attention interleaved with full attention β€” the architecture family Qwen3.8-27B is built on) will
load it; build llama.cpp from a revision that includes that support.

**Sampling β€” use temperature 1.0, top-p 0.95, top-k 20** (the parent's thinking preset), and do not
decode greedily: at temperature 0, 24 of 59 traces that used their full thinking budget contained
verbatim repetition loops. This is a property of the parent's thinking mode, inherited rather than
introduced by pruning, and sampling is the fix for both models.

The chat template defaults to `reasoning_effort="xhigh"`; `medium` and `low` are supported, as is
`enable_thinking=False`.

**Speed and memory.** On a 16 GB RTX 4080 SUPER at Q4_K_M with every layer resident and q8_0 KV:
**~36 tokens/s per stream with two concurrent streams.** Weights are 13.74 GB; a 32K context adds
roughly 1.1 GB.

## Roadmap β€” three training rounds scheduled, plus two acceleration variants

| round | goal | status |
|---|---|---|
| R1 | heal the pruning loss broadly | done β€” 27.5 M tokens |
| R1.5 | close specific probed fact losses | **this release** |
| **R1.7** | match the 27B on long-context, reasoning-heavy work; close the rotation/trig arithmetic gap | measurement built and running |
| **R1.8** | professional depth across every major engineering field | scheduled |
| **R2** | tool use and agentic execution on the strengthened base | scheduled |

From R1.7 onward, every round is gated on the full instrument set β€” retention, derivation,
arithmetic fidelity and long-horizon KL β€” with checkpoints scored **during** training rather than
only at the end, so a round that trades one capability for another is caught while it runs. That
discipline came from a round whose in-loop metric improved while a capability it never measured
regressed; the fix was to widen the instruments, and it is now a standing gate.

Two acceleration variants are planned on the same weights: **-super** (a multi-token-prediction
draft head, 1.3–1.7Γ— on code) and **-turbo** (EAGLE-3, 3–4Γ— on supporting runtimes). This release
is the dense build: `mtp_num_hidden_layers` is 0.

The same pipeline then ladders downward β€” an 18B distilled from this model's successor, and a 9B
distilled into a native base β€” reusing the teacher caches and evaluation banks built here.

## Scope of this release

- **GGUF only, for now.** This repository contains the Q4_K_M build and nothing else β€” no
  safetensors, config or tokenizer files, so `transformers`, vLLM and SGLang cannot load it from
  here. Use llama.cpp or another GGUF runtime. The bf16 weights are planned for a later upload.
- **Text only.** Image and video inputs are not supported; the vision tower was dropped to spend
  the budget on reasoning.
- **No public benchmark scores are claimed.** The held-out benchmark subsets here are too small to
  publish, and none were used in training. Results above are from the project's own instruments,
  all of which are described in full so they can be argued with.
- **Long-horizon parity against the parent is being measured now** and will be reported with R1.7.
- It inherits the parent's biases, refusals and knowledge cutoff; no alignment training was added.
- It thinks at length. On multi-stage engineering calculations it will carry more intermediate
  precision than the answer needs β€” ask explicitly for the precision you want, and give it room.

## Training data and attribution

Public data, subsampled: **arXiv, open-web-math, AlgebraicStack** (permissive-licensed subset of
The Stack) and **FineWeb-Edu** β€” all **ODC-By**, whose attribution travels with this model β€” plus
**PG-19** (public domain). Reasoning traces were generated by the parent model itself. No benchmark
test sets were included; the corpus was hash-checked against the evaluation banks.

## License

**Apache 2.0**, inherited from Qwen3.8-27B. This is a derivative work of that model; the parent's
terms apply to it and to anything derived from it.