devvexus commited on
Commit
9ca16b3
Β·
verified Β·
1 Parent(s): 04d76c5

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +38 -28
README.md CHANGED
@@ -1,11 +1,13 @@
1
  ---
2
  license: apache-2.0
3
  base_model: Qwen/Qwen3.8-27B
4
- library_name: transformers
5
  pipeline_tag: text-generation
6
  language:
7
  - en
8
  tags:
 
 
9
  - qwen3_5
10
  - pruning
11
  - depth-pruning
@@ -19,15 +21,16 @@ tags:
19
  and measured, domain by domain, against the model it came from.**
20
 
21
  Marlowe-22B removes 12 layers from **Qwen3.8-27B** and heals the loss by distilling from the 27B
22
- itself. R1.5 keeps **94.8% of the facts its parent demonstrably knows**, at **81% of the
23
- parameters**, with the whole model resident in 16 GB at Q4_K_M β€” no layer offload, no paging.
 
24
 
25
- This is round three of a planned seven-round program. R1.7, R1.8 and R2 are scheduled; the
26
- roadmap is at the bottom, along with what each round is built to move.
27
 
28
  | | |
29
  |---|---|
30
- | parameters | 22.30 B (parent: 27 B) |
31
  | layers | 52 β€” 36 linear-attention (DeltaNet) + 16 full-attention (parent: 64 = 48 + 16) |
32
  | hidden / intermediate | 5120 / 17408 |
33
  | KV heads Γ— head dim | 4 Γ— 256, on the 16 full-attention layers only |
@@ -35,7 +38,7 @@ roadmap is at the bottom, along with what each round is built to move.
35
  | max position embeddings | 262,144 |
36
  | modality | text only (the parent's vision tower is not included) |
37
  | thinking | on by default; `reasoning_effort` xhigh / medium / low |
38
- | release build | GGUF **Q4_K_M, 13.74 GB** |
39
 
40
  ## Why compress this model in particular
41
 
@@ -61,7 +64,8 @@ answers intact. Marlowe-22B is the same architecture, twelve layers shorter, at
61
  mean-of-squares, RoPE's exponent, the GELU cubic constant, Adam's bias correction. A 1.5 M-token
62
  pass on documents carrying exactly those facts closed them.
63
 
64
- The adapter is merged; these are plain weights.
 
65
 
66
  ## Results
67
 
@@ -102,7 +106,11 @@ The sharper number is the hardest 60 β€” the families the pruned model failed ou
102
  | **R1.5** | **32 / 60** |
103
  | parent (27B) | 41 / 60 |
104
 
105
- Healing recovered **71% of the gap** pruning opened on the material it damaged most.
 
 
 
 
106
 
107
  ### Arithmetic fidelity
108
 
@@ -137,29 +145,26 @@ flattering result hard to get:
137
  **llama.cpp** β€” all 52 layers on the GPU, quantised KV:
138
 
139
  ```bash
140
- llama-server -m marlowe-22b-q4_K_M.gguf -ngl 99 \
141
  -c 32768 -fa on -ctk q8_0 -ctv q8_0 \
142
  --temp 1.0 --top-p 0.95 --top-k 20 \
143
  --reasoning-format deepseek
144
  ```
145
 
146
- **transformers:**
147
-
148
- ```python
149
- from transformers import AutoModelForCausalLM, AutoTokenizer
150
-
151
- tok = AutoTokenizer.from_pretrained("Marlowe-22B-R1.5")
152
- model = AutoModelForCausalLM.from_pretrained("Marlowe-22B-R1.5", dtype="bfloat16",
153
- device_map="auto")
154
- msgs = [{"role": "user", "content": "Derive the RoPE inverse frequency for head_dim 144, "
155
- "base 21000, pair index 7."}]
156
- ids = tok.apply_chat_template(msgs, add_generation_prompt=True, return_tensors="pt",
157
- reasoning_effort="xhigh").to(model.device)
158
- out = model.generate(ids, max_new_tokens=8192, do_sample=True, temperature=1.0,
159
- top_p=0.95, top_k=20)
160
- print(tok.decode(out[0][ids.shape[-1]:]))
161
  ```
162
 
 
 
 
 
 
 
 
163
  **Sampling β€” use temperature 1.0, top-p 0.95, top-k 20** (the parent's thinking preset), and do not
164
  decode greedily: at temperature 0, 24 of 59 traces that used their full thinking budget contained
165
  verbatim repetition loops. This is a property of the parent's thinking mode, inherited rather than
@@ -182,9 +187,11 @@ roughly 1.1 GB.
182
  | **R1.8** | professional depth across every major engineering field | scheduled |
183
  | **R2** | tool use and agentic execution on the strengthened base | scheduled |
184
 
185
- Each round is gated on the full instrument set β€” retention, derivation, arithmetic fidelity and
186
- long-horizon KL β€” with checkpoints scored during training rather than only at the end, so a round
187
- that trades one capability for another is caught while it runs.
 
 
188
 
189
  Two acceleration variants are planned on the same weights: **-super** (a multi-token-prediction
190
  draft head, 1.3–1.7Γ— on code) and **-turbo** (EAGLE-3, 3–4Γ— on supporting runtimes). This release
@@ -195,6 +202,9 @@ distilled into a native base β€” reusing the teacher caches and evaluation banks
195
 
196
  ## Scope of this release
197
 
 
 
 
198
  - **Text only.** Image and video inputs are not supported; the vision tower was dropped to spend
199
  the budget on reasoning.
200
  - **No public benchmark scores are claimed.** The held-out benchmark subsets here are too small to
 
1
  ---
2
  license: apache-2.0
3
  base_model: Qwen/Qwen3.8-27B
4
+ base_model_relation: finetune
5
  pipeline_tag: text-generation
6
  language:
7
  - en
8
  tags:
9
+ - gguf
10
+ - llama.cpp
11
  - qwen3_5
12
  - pruning
13
  - depth-pruning
 
21
  and measured, domain by domain, against the model it came from.**
22
 
23
  Marlowe-22B removes 12 layers from **Qwen3.8-27B** and heals the loss by distilling from the 27B
24
+ itself. R1.5 keeps **94.8% of the facts its parent demonstrably knows**, on **81.6% of the parent's
25
+ text parameters** (22.30 B against 27.32 B), with the whole model resident in 16 GB at Q4_K_M β€” no
26
+ layer offload, no paging.
27
 
28
+ It is the second of five planned rounds. Three remain β€” R1.7, R1.8 and R2 β€” plus two acceleration
29
+ variants; the roadmap at the bottom says what each is built to move.
30
 
31
  | | |
32
  |---|---|
33
+ | parameters | 22.30 B β€” 81.6% of the parent's 27.32 B text parameters (its 0.46 B vision tower is dropped) |
34
  | layers | 52 β€” 36 linear-attention (DeltaNet) + 16 full-attention (parent: 64 = 48 + 16) |
35
  | hidden / intermediate | 5120 / 17408 |
36
  | KV heads Γ— head dim | 4 Γ— 256, on the 16 full-attention layers only |
 
38
  | max position embeddings | 262,144 |
39
  | modality | text only (the parent's vision tower is not included) |
40
  | thinking | on by default; `reasoning_effort` xhigh / medium / low |
41
+ | this repository | GGUF **Q4_K_M, 13.7 GB** β€” `marlowe-dusk-22b-r1.5.gguf` (this release) and `marlowe-dusk-22b-r1.gguf` (previous round) |
42
 
43
  ## Why compress this model in particular
44
 
 
64
  mean-of-squares, RoPE's exponent, the GELU cubic constant, Adam's bias correction. A 1.5 M-token
65
  pass on documents carrying exactly those facts closed them.
66
 
67
+ The adapter is merged into the weights, and this repository ships the **Q4_K_M GGUF** built from
68
+ them. The bf16 safetensors are not published yet.
69
 
70
  ## Results
71
 
 
106
  | **R1.5** | **32 / 60** |
107
  | parent (27B) | 41 / 60 |
108
 
109
+ Healing recovered **71% of the gap** pruning opened on the material it damaged most. All three
110
+ models ran the identical protocol, greedy, at a 6K-token thinking budget; the parent was measured
111
+ at IQ3_M, a heavier quantisation than this build's Q4_K_M, so the comparison does not flatter it.
112
+ (Greedy is the wrong setting for daily use β€” see Sampling β€” but it removes sampling noise from a
113
+ paired comparison.)
114
 
115
  ### Arithmetic fidelity
116
 
 
145
  **llama.cpp** β€” all 52 layers on the GPU, quantised KV:
146
 
147
  ```bash
148
+ llama-server -m marlowe-dusk-22b-r1.5.gguf -ngl 99 \
149
  -c 32768 -fa on -ctk q8_0 -ctv q8_0 \
150
  --temp 1.0 --top-p 0.95 --top-k 20 \
151
  --reasoning-format deepseek
152
  ```
153
 
154
+ **Get the file** β€” `marlowe-dusk-22b-r1.5.gguf` is this release; `marlowe-dusk-22b-r1.gguf` is the previous
155
+ round, kept for comparison.
156
+
157
+ ```bash
158
+ huggingface-cli download devvexus/marlowe-22b marlowe-dusk-22b-r1.5.gguf --local-dir .
 
 
 
 
 
 
 
 
 
 
159
  ```
160
 
161
+ The chat template ships inside the GGUF, so `llama-server` applies it for you β€” including the
162
+ thinking block, which `--reasoning-format deepseek` returns as `reasoning_content`.
163
+
164
+ Any runtime that speaks GGUF and supports the hybrid `qwen3_5` architecture (DeltaNet linear
165
+ attention interleaved with full attention β€” the architecture family Qwen3.8-27B is built on) will
166
+ load it; build llama.cpp from a revision that includes that support.
167
+
168
  **Sampling β€” use temperature 1.0, top-p 0.95, top-k 20** (the parent's thinking preset), and do not
169
  decode greedily: at temperature 0, 24 of 59 traces that used their full thinking budget contained
170
  verbatim repetition loops. This is a property of the parent's thinking mode, inherited rather than
 
187
  | **R1.8** | professional depth across every major engineering field | scheduled |
188
  | **R2** | tool use and agentic execution on the strengthened base | scheduled |
189
 
190
+ From R1.7 onward, every round is gated on the full instrument set β€” retention, derivation,
191
+ arithmetic fidelity and long-horizon KL β€” with checkpoints scored **during** training rather than
192
+ only at the end, so a round that trades one capability for another is caught while it runs. That
193
+ discipline came from a round whose in-loop metric improved while a capability it never measured
194
+ regressed; the fix was to widen the instruments, and it is now a standing gate.
195
 
196
  Two acceleration variants are planned on the same weights: **-super** (a multi-token-prediction
197
  draft head, 1.3–1.7Γ— on code) and **-turbo** (EAGLE-3, 3–4Γ— on supporting runtimes). This release
 
202
 
203
  ## Scope of this release
204
 
205
+ - **GGUF only, for now.** This repository contains the Q4_K_M build and nothing else β€” no
206
+ safetensors, config or tokenizer files, so `transformers`, vLLM and SGLang cannot load it from
207
+ here. Use llama.cpp or another GGUF runtime. The bf16 weights are planned for a later upload.
208
  - **Text only.** Image and video inputs are not supported; the vision tower was dropped to spend
209
  the budget on reasoning.
210
  - **No public benchmark scores are claimed.** The held-out benchmark subsets here are too small to