Text Generation
GGUF
English
llama.cpp
qwen3_5
pruning
depth-pruning
distillation
reasoning
conversational
Instructions to use devvexus/marlowe-22b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use devvexus/marlowe-22b with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf devvexus/marlowe-22b # Run inference directly in the terminal: llama cli -hf devvexus/marlowe-22b
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf devvexus/marlowe-22b # Run inference directly in the terminal: llama cli -hf devvexus/marlowe-22b
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf devvexus/marlowe-22b # Run inference directly in the terminal: ./llama-cli -hf devvexus/marlowe-22b
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf devvexus/marlowe-22b # Run inference directly in the terminal: ./build/bin/llama-cli -hf devvexus/marlowe-22b
Use Docker
docker model run hf.co/devvexus/marlowe-22b
- LM Studio
- Jan
- vLLM
How to use devvexus/marlowe-22b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "devvexus/marlowe-22b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "devvexus/marlowe-22b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/devvexus/marlowe-22b
- Ollama
How to use devvexus/marlowe-22b with Ollama:
ollama run hf.co/devvexus/marlowe-22b
- Unsloth Desktop
- Pi
How to use devvexus/marlowe-22b with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf devvexus/marlowe-22b
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "devvexus/marlowe-22b" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use devvexus/marlowe-22b with Docker Model Runner:
docker model run hf.co/devvexus/marlowe-22b
- Lemonade
How to use devvexus/marlowe-22b with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull devvexus/marlowe-22b
Run and chat with the model
lemonade run user.marlowe-22b-{{QUANT_TAG}}List all available models
lemonade list
- Hermes Agent
How to use devvexus/marlowe-22b with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf devvexus/marlowe-22b
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default devvexus/marlowe-22b
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use devvexus/marlowe-22b with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf devvexus/marlowe-22b
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "devvexus/marlowe-22b" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Update README.md
Browse files
README.md
CHANGED
|
@@ -1,11 +1,13 @@
|
|
| 1 |
---
|
| 2 |
license: apache-2.0
|
| 3 |
base_model: Qwen/Qwen3.8-27B
|
| 4 |
-
|
| 5 |
pipeline_tag: text-generation
|
| 6 |
language:
|
| 7 |
- en
|
| 8 |
tags:
|
|
|
|
|
|
|
| 9 |
- qwen3_5
|
| 10 |
- pruning
|
| 11 |
- depth-pruning
|
|
@@ -19,15 +21,16 @@ tags:
|
|
| 19 |
and measured, domain by domain, against the model it came from.**
|
| 20 |
|
| 21 |
Marlowe-22B removes 12 layers from **Qwen3.8-27B** and heals the loss by distilling from the 27B
|
| 22 |
-
itself. R1.5 keeps **94.8% of the facts its parent demonstrably knows**,
|
| 23 |
-
parameters**, with the whole model resident in 16 GB at Q4_K_M β no
|
|
|
|
| 24 |
|
| 25 |
-
|
| 26 |
-
|
| 27 |
|
| 28 |
| | |
|
| 29 |
|---|---|
|
| 30 |
-
| parameters | 22.30 B
|
| 31 |
| layers | 52 β 36 linear-attention (DeltaNet) + 16 full-attention (parent: 64 = 48 + 16) |
|
| 32 |
| hidden / intermediate | 5120 / 17408 |
|
| 33 |
| KV heads Γ head dim | 4 Γ 256, on the 16 full-attention layers only |
|
|
@@ -35,7 +38,7 @@ roadmap is at the bottom, along with what each round is built to move.
|
|
| 35 |
| max position embeddings | 262,144 |
|
| 36 |
| modality | text only (the parent's vision tower is not included) |
|
| 37 |
| thinking | on by default; `reasoning_effort` xhigh / medium / low |
|
| 38 |
-
|
|
| 39 |
|
| 40 |
## Why compress this model in particular
|
| 41 |
|
|
@@ -61,7 +64,8 @@ answers intact. Marlowe-22B is the same architecture, twelve layers shorter, at
|
|
| 61 |
mean-of-squares, RoPE's exponent, the GELU cubic constant, Adam's bias correction. A 1.5 M-token
|
| 62 |
pass on documents carrying exactly those facts closed them.
|
| 63 |
|
| 64 |
-
The adapter is merged
|
|
|
|
| 65 |
|
| 66 |
## Results
|
| 67 |
|
|
@@ -102,7 +106,11 @@ The sharper number is the hardest 60 β the families the pruned model failed ou
|
|
| 102 |
| **R1.5** | **32 / 60** |
|
| 103 |
| parent (27B) | 41 / 60 |
|
| 104 |
|
| 105 |
-
Healing recovered **71% of the gap** pruning opened on the material it damaged most.
|
|
|
|
|
|
|
|
|
|
|
|
|
| 106 |
|
| 107 |
### Arithmetic fidelity
|
| 108 |
|
|
@@ -137,29 +145,26 @@ flattering result hard to get:
|
|
| 137 |
**llama.cpp** β all 52 layers on the GPU, quantised KV:
|
| 138 |
|
| 139 |
```bash
|
| 140 |
-
llama-server -m marlowe-22b-
|
| 141 |
-c 32768 -fa on -ctk q8_0 -ctv q8_0 \
|
| 142 |
--temp 1.0 --top-p 0.95 --top-k 20 \
|
| 143 |
--reasoning-format deepseek
|
| 144 |
```
|
| 145 |
|
| 146 |
-
**
|
| 147 |
-
|
| 148 |
-
|
| 149 |
-
|
| 150 |
-
|
| 151 |
-
tok = AutoTokenizer.from_pretrained("Marlowe-22B-R1.5")
|
| 152 |
-
model = AutoModelForCausalLM.from_pretrained("Marlowe-22B-R1.5", dtype="bfloat16",
|
| 153 |
-
device_map="auto")
|
| 154 |
-
msgs = [{"role": "user", "content": "Derive the RoPE inverse frequency for head_dim 144, "
|
| 155 |
-
"base 21000, pair index 7."}]
|
| 156 |
-
ids = tok.apply_chat_template(msgs, add_generation_prompt=True, return_tensors="pt",
|
| 157 |
-
reasoning_effort="xhigh").to(model.device)
|
| 158 |
-
out = model.generate(ids, max_new_tokens=8192, do_sample=True, temperature=1.0,
|
| 159 |
-
top_p=0.95, top_k=20)
|
| 160 |
-
print(tok.decode(out[0][ids.shape[-1]:]))
|
| 161 |
```
|
| 162 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 163 |
**Sampling β use temperature 1.0, top-p 0.95, top-k 20** (the parent's thinking preset), and do not
|
| 164 |
decode greedily: at temperature 0, 24 of 59 traces that used their full thinking budget contained
|
| 165 |
verbatim repetition loops. This is a property of the parent's thinking mode, inherited rather than
|
|
@@ -182,9 +187,11 @@ roughly 1.1 GB.
|
|
| 182 |
| **R1.8** | professional depth across every major engineering field | scheduled |
|
| 183 |
| **R2** | tool use and agentic execution on the strengthened base | scheduled |
|
| 184 |
|
| 185 |
-
|
| 186 |
-
long-horizon KL β with checkpoints scored during training rather than
|
| 187 |
-
that trades one capability for another is caught while it runs.
|
|
|
|
|
|
|
| 188 |
|
| 189 |
Two acceleration variants are planned on the same weights: **-super** (a multi-token-prediction
|
| 190 |
draft head, 1.3β1.7Γ on code) and **-turbo** (EAGLE-3, 3β4Γ on supporting runtimes). This release
|
|
@@ -195,6 +202,9 @@ distilled into a native base β reusing the teacher caches and evaluation banks
|
|
| 195 |
|
| 196 |
## Scope of this release
|
| 197 |
|
|
|
|
|
|
|
|
|
|
| 198 |
- **Text only.** Image and video inputs are not supported; the vision tower was dropped to spend
|
| 199 |
the budget on reasoning.
|
| 200 |
- **No public benchmark scores are claimed.** The held-out benchmark subsets here are too small to
|
|
|
|
| 1 |
---
|
| 2 |
license: apache-2.0
|
| 3 |
base_model: Qwen/Qwen3.8-27B
|
| 4 |
+
base_model_relation: finetune
|
| 5 |
pipeline_tag: text-generation
|
| 6 |
language:
|
| 7 |
- en
|
| 8 |
tags:
|
| 9 |
+
- gguf
|
| 10 |
+
- llama.cpp
|
| 11 |
- qwen3_5
|
| 12 |
- pruning
|
| 13 |
- depth-pruning
|
|
|
|
| 21 |
and measured, domain by domain, against the model it came from.**
|
| 22 |
|
| 23 |
Marlowe-22B removes 12 layers from **Qwen3.8-27B** and heals the loss by distilling from the 27B
|
| 24 |
+
itself. R1.5 keeps **94.8% of the facts its parent demonstrably knows**, on **81.6% of the parent's
|
| 25 |
+
text parameters** (22.30 B against 27.32 B), with the whole model resident in 16 GB at Q4_K_M β no
|
| 26 |
+
layer offload, no paging.
|
| 27 |
|
| 28 |
+
It is the second of five planned rounds. Three remain β R1.7, R1.8 and R2 β plus two acceleration
|
| 29 |
+
variants; the roadmap at the bottom says what each is built to move.
|
| 30 |
|
| 31 |
| | |
|
| 32 |
|---|---|
|
| 33 |
+
| parameters | 22.30 B β 81.6% of the parent's 27.32 B text parameters (its 0.46 B vision tower is dropped) |
|
| 34 |
| layers | 52 β 36 linear-attention (DeltaNet) + 16 full-attention (parent: 64 = 48 + 16) |
|
| 35 |
| hidden / intermediate | 5120 / 17408 |
|
| 36 |
| KV heads Γ head dim | 4 Γ 256, on the 16 full-attention layers only |
|
|
|
|
| 38 |
| max position embeddings | 262,144 |
|
| 39 |
| modality | text only (the parent's vision tower is not included) |
|
| 40 |
| thinking | on by default; `reasoning_effort` xhigh / medium / low |
|
| 41 |
+
| this repository | GGUF **Q4_K_M, 13.7 GB** β `marlowe-dusk-22b-r1.5.gguf` (this release) and `marlowe-dusk-22b-r1.gguf` (previous round) |
|
| 42 |
|
| 43 |
## Why compress this model in particular
|
| 44 |
|
|
|
|
| 64 |
mean-of-squares, RoPE's exponent, the GELU cubic constant, Adam's bias correction. A 1.5 M-token
|
| 65 |
pass on documents carrying exactly those facts closed them.
|
| 66 |
|
| 67 |
+
The adapter is merged into the weights, and this repository ships the **Q4_K_M GGUF** built from
|
| 68 |
+
them. The bf16 safetensors are not published yet.
|
| 69 |
|
| 70 |
## Results
|
| 71 |
|
|
|
|
| 106 |
| **R1.5** | **32 / 60** |
|
| 107 |
| parent (27B) | 41 / 60 |
|
| 108 |
|
| 109 |
+
Healing recovered **71% of the gap** pruning opened on the material it damaged most. All three
|
| 110 |
+
models ran the identical protocol, greedy, at a 6K-token thinking budget; the parent was measured
|
| 111 |
+
at IQ3_M, a heavier quantisation than this build's Q4_K_M, so the comparison does not flatter it.
|
| 112 |
+
(Greedy is the wrong setting for daily use β see Sampling β but it removes sampling noise from a
|
| 113 |
+
paired comparison.)
|
| 114 |
|
| 115 |
### Arithmetic fidelity
|
| 116 |
|
|
|
|
| 145 |
**llama.cpp** β all 52 layers on the GPU, quantised KV:
|
| 146 |
|
| 147 |
```bash
|
| 148 |
+
llama-server -m marlowe-dusk-22b-r1.5.gguf -ngl 99 \
|
| 149 |
-c 32768 -fa on -ctk q8_0 -ctv q8_0 \
|
| 150 |
--temp 1.0 --top-p 0.95 --top-k 20 \
|
| 151 |
--reasoning-format deepseek
|
| 152 |
```
|
| 153 |
|
| 154 |
+
**Get the file** β `marlowe-dusk-22b-r1.5.gguf` is this release; `marlowe-dusk-22b-r1.gguf` is the previous
|
| 155 |
+
round, kept for comparison.
|
| 156 |
+
|
| 157 |
+
```bash
|
| 158 |
+
huggingface-cli download devvexus/marlowe-22b marlowe-dusk-22b-r1.5.gguf --local-dir .
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 159 |
```
|
| 160 |
|
| 161 |
+
The chat template ships inside the GGUF, so `llama-server` applies it for you β including the
|
| 162 |
+
thinking block, which `--reasoning-format deepseek` returns as `reasoning_content`.
|
| 163 |
+
|
| 164 |
+
Any runtime that speaks GGUF and supports the hybrid `qwen3_5` architecture (DeltaNet linear
|
| 165 |
+
attention interleaved with full attention β the architecture family Qwen3.8-27B is built on) will
|
| 166 |
+
load it; build llama.cpp from a revision that includes that support.
|
| 167 |
+
|
| 168 |
**Sampling β use temperature 1.0, top-p 0.95, top-k 20** (the parent's thinking preset), and do not
|
| 169 |
decode greedily: at temperature 0, 24 of 59 traces that used their full thinking budget contained
|
| 170 |
verbatim repetition loops. This is a property of the parent's thinking mode, inherited rather than
|
|
|
|
| 187 |
| **R1.8** | professional depth across every major engineering field | scheduled |
|
| 188 |
| **R2** | tool use and agentic execution on the strengthened base | scheduled |
|
| 189 |
|
| 190 |
+
From R1.7 onward, every round is gated on the full instrument set β retention, derivation,
|
| 191 |
+
arithmetic fidelity and long-horizon KL β with checkpoints scored **during** training rather than
|
| 192 |
+
only at the end, so a round that trades one capability for another is caught while it runs. That
|
| 193 |
+
discipline came from a round whose in-loop metric improved while a capability it never measured
|
| 194 |
+
regressed; the fix was to widen the instruments, and it is now a standing gate.
|
| 195 |
|
| 196 |
Two acceleration variants are planned on the same weights: **-super** (a multi-token-prediction
|
| 197 |
draft head, 1.3β1.7Γ on code) and **-turbo** (EAGLE-3, 3β4Γ on supporting runtimes). This release
|
|
|
|
| 202 |
|
| 203 |
## Scope of this release
|
| 204 |
|
| 205 |
+
- **GGUF only, for now.** This repository contains the Q4_K_M build and nothing else β no
|
| 206 |
+
safetensors, config or tokenizer files, so `transformers`, vLLM and SGLang cannot load it from
|
| 207 |
+
here. Use llama.cpp or another GGUF runtime. The bf16 weights are planned for a later upload.
|
| 208 |
- **Text only.** Image and video inputs are not supported; the vision tower was dropped to spend
|
| 209 |
the budget on reasoning.
|
| 210 |
- **No public benchmark scores are claimed.** The held-out benchmark subsets here are too small to
|