Text Generation
GGUF
English
llama.cpp
qwen3_5
pruning
depth-pruning
distillation
reasoning
conversational
Instructions to use devvexus/marlowe-22b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use devvexus/marlowe-22b with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf devvexus/marlowe-22b # Run inference directly in the terminal: llama cli -hf devvexus/marlowe-22b
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf devvexus/marlowe-22b # Run inference directly in the terminal: llama cli -hf devvexus/marlowe-22b
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf devvexus/marlowe-22b # Run inference directly in the terminal: ./llama-cli -hf devvexus/marlowe-22b
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf devvexus/marlowe-22b # Run inference directly in the terminal: ./build/bin/llama-cli -hf devvexus/marlowe-22b
Use Docker
docker model run hf.co/devvexus/marlowe-22b
- LM Studio
- Jan
- vLLM
How to use devvexus/marlowe-22b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "devvexus/marlowe-22b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "devvexus/marlowe-22b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/devvexus/marlowe-22b
- Ollama
How to use devvexus/marlowe-22b with Ollama:
ollama run hf.co/devvexus/marlowe-22b
- Unsloth Desktop
- Pi
How to use devvexus/marlowe-22b with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf devvexus/marlowe-22b
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "devvexus/marlowe-22b" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use devvexus/marlowe-22b with Docker Model Runner:
docker model run hf.co/devvexus/marlowe-22b
- Lemonade
How to use devvexus/marlowe-22b with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull devvexus/marlowe-22b
Run and chat with the model
lemonade run user.marlowe-22b-{{QUANT_TAG}}List all available models
lemonade list
- Hermes Agent
How to use devvexus/marlowe-22b with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf devvexus/marlowe-22b
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default devvexus/marlowe-22b
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use devvexus/marlowe-22b with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf devvexus/marlowe-22b
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "devvexus/marlowe-22b" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
|
Download README.md from devvexus/marlowe-22b: direct link, hf CLI and curl.
- Browser
- Download file 12 kB
-
https://huggingface.co/devvexus/marlowe-22b/resolve/main/README.md
- Command line
-
hf download hf://devvexus/marlowe-22b/README.md
-
curl -L -o README.md https://huggingface.co/devvexus/marlowe-22b/resolve/main/README.md
12 kB
| license: apache-2.0 | |
| base_model: Qwen/Qwen3.8-27B | |
| base_model_relation: finetune | |
| pipeline_tag: text-generation | |
| language: | |
| - en | |
| tags: | |
| - gguf | |
| - llama.cpp | |
| - qwen3_5 | |
| - pruning | |
| - depth-pruning | |
| - distillation | |
| - reasoning | |
| # Marlowe-22B (R1.5) | |
| **A 27B-class reasoning model compressed to 22.3B so it runs entirely on one 16 GB consumer GPU β | |
| and measured, domain by domain, against the model it came from.** | |
| Marlowe-22B removes 12 layers from **Qwen3.8-27B** and heals the loss by distilling from the 27B | |
| itself. R1.5 keeps **94.8% of the facts its parent demonstrably knows**, on **81.6% of the parent's | |
| text parameters** (22.30 B against 27.32 B), with the whole model resident in 16 GB at Q4_K_M β no | |
| layer offload, no paging. | |
| It is the second of five planned rounds. Three remain β R1.7, R1.8 and R2 β plus two acceleration | |
| variants; the roadmap at the bottom says what each is built to move. | |
| | | | | |
| |---|---| | |
| | parameters | 22.30 B β 81.6% of the parent's 27.32 B text parameters (its 0.46 B vision tower is dropped) | | |
| | layers | 52 β 36 linear-attention (DeltaNet) + 16 full-attention (parent: 64 = 48 + 16) | | |
| | hidden / intermediate | 5120 / 17408 | | |
| | KV heads Γ head dim | 4 Γ 256, on the 16 full-attention layers only | | |
| | vocabulary | 248,320 | | |
| | max position embeddings | 262,144 | | |
| | modality | text only (the parent's vision tower is not included) | | |
| | thinking | on by default; `reasoning_effort` xhigh / medium / low | | |
| | this repository | GGUF **Q4_K_M, 13.7 GB** β `marlowe-dusk-22b-r1.5.gguf` (this release) and `marlowe-dusk-22b-r1.gguf` (previous round) | | |
| ## Why compress this model in particular | |
| The parent is a hybrid: most layers are linear-attention (DeltaNet), and only every fourth is full | |
| attention. Its KV cache is therefore tiny β **64 KB per token at f16, ~34 KB at q8_0** β so long | |
| contexts are cheap to hold. What it cannot do is fit in 16 GB at a bit width that leaves its | |
| answers intact. Marlowe-22B is the same architecture, twelve layers shorter, at the size where a | |
| 4-bit build and a long context fit on one card together. | |
| ## Method | |
| 1. **Layer profiling.** Each layer scored by the KL divergence its removal induces on held-out | |
| text β measured, not assumed from depth heuristics. | |
| 2. **Depth pruning.** The twelve cheapest layers by that measure: indices | |
| `4, 5, 8, 9, 13, 14, 16, 17, 37, 38, 40, 41` β all linear-attention. Every full-attention layer | |
| survived, as did the first four and the last ten layers. | |
| 3. **Heal (R1).** LoRA (r=32, Ξ±=64) trained against the parent's **top-64 log-probs** β forward KL | |
| against a real distribution, not hard labels β over **27.5 M tokens** of a 35 M-token teacher | |
| cache (arXiv, open-web-math, AlgebraicStack, FineWeb-Edu, PG-19, plus 1,314 reasoning traces | |
| from the parent). Stopped deliberately when held-out KL flattened, rather than at budget | |
| exhaustion. | |
| 4. **Targeted repair (R1.5).** Probing located specific facts the pruned model had lost β RMSNorm's | |
| mean-of-squares, RoPE's exponent, the GELU cubic constant, Adam's bias correction. A 1.5 M-token | |
| pass on documents carrying exactly those facts closed them. | |
| The adapter is merged into the weights, and this repository ships the **Q4_K_M GGUF** built from | |
| them. The bf16 safetensors are not published yet. | |
| ## Results | |
| ### Knowledge retention against the parent | |
| Facts are mined from documents **hash-dropped against the heal corpus**, and **the parent is scored | |
| first: every item the parent itself fails is discarded** β so only knowledge the 27B demonstrably | |
| has is counted, and the student is never charged for its teacher's gaps. Denominator: **3,859 | |
| parent-verified facts** (6,002 mined items before gating). Each cell is the share of those facts | |
| the model also knows. | |
| | domain | pruned, unhealed | **R1.5** | recovered by healing | | |
| |---|---:|---:|---:| | |
| | chemistry / biology | 90.5% | **96.0%** | +5.5 | | |
| | physics | 79.8% | **85.3%** | +5.5 | | |
| | engineering | 88.9% | **91.6%** | +2.7 | | |
| | statistics | 87.4% | **89.2%** | +1.8 | | |
| | mathematics | 93.4% | **94.8%** | +1.4 | | |
| | general | 96.3% | **97.4%** | +1.1 | | |
| | machine learning | 92.4% | **93.2%** | +0.8 | | |
| | computer science | 91.2% | **89.9%** | β1.3 | | |
| | **overall** | **93.3%** | **94.8%** | **+1.5** | | |
| Five domains sit above 91%; the program's standard is **90% in every domain, not on average**, and | |
| physics, statistics and computer science are the remaining targets β see the roadmap. | |
| ### Derivation, not recall | |
| 120 generated engineering problems, parameters sampled outside tutorial values and answers computed | |
| from the defining formula, so a memorised answer scores zero: RoPE frequencies, GELU, Adam updates, | |
| softmax, cosine schedules, KV-cache sizing, LoRA parameter counts. **R1.5 answers 92 of 120.** | |
| The sharper number is the hardest 60 β the families the pruned model failed outright: | |
| | | correct | | |
| |---|---:| | |
| | pruned, unhealed | 10 / 60 | | |
| | **R1.5** | **32 / 60** | | |
| | parent (27B) | 41 / 60 | | |
| Healing recovered **71% of the gap** pruning opened on the material it damaged most. All three | |
| models ran the identical protocol, greedy, at a 6K-token thinking budget; the parent was measured | |
| at IQ3_M, a heavier quantisation than this build's Q4_K_M, so the comparison does not flatter it. | |
| (Greedy is the wrong setting for daily use β see Sampling β but it removes sampling noise from a | |
| paired comparison.) | |
| ### Arithmetic fidelity | |
| Teacher-forced on held-out worked computations, the parent predicts the next computed digit with | |
| **67.9%** top-1 accuracy; R1.5 reaches **64.1%** β a 0.115-nat gap in log-probability. KV-cache and | |
| attention-scaling arithmetic is already at parent level; rotation- and trigonometry-heavy work | |
| (RoPE, cosine schedules) is where the remaining distance lives, and it is R1.7's first target. | |
| ## How it is measured | |
| The instruments matter as much as the weights here, and every one of them is built to make a | |
| flattering result hard to get: | |
| - **Parent-gated retention bank** β the teacher is scored first and its own failures are dropped, | |
| so the headline cannot be inflated by items nobody knows. Sources are hash-dropped against the | |
| training corpus, and mined from documents rather than hand-written, so the bank cannot grade its | |
| author's homework. | |
| - **Generated derivation problems** β computed answers, parameters outside tutorial ranges, plus a | |
| recorded *recall trap* per family: the exact value a model lands on when it quotes a remembered | |
| formula with the wrong exponent. Wrong answers say *which* shortcut was taken. | |
| - **Teacher-forced digit scoring** β separates "can it do the arithmetic" from "can it run its own | |
| derivation", in five minutes and with no grading ambiguity. | |
| - **Position-resolved KL** (in progress) β KL against the parent at every position of a full | |
| reasoning trace, up to 24K tokens, on the parent's traces and the student's own. Every heal so | |
| far optimised KL on 768-token windows; this measures whether that proxy holds at long horizons. | |
| - **Pre-registered decision rules.** Each experiment's confirm/refute criteria are written down | |
| before the data exists, and results are reported against them either way. Several promising | |
| hypotheses have been refuted by their own tests and recorded as such. | |
| ## Running it | |
| **llama.cpp** β all 52 layers on the GPU, quantised KV: | |
| ```bash | |
| llama-server -m marlowe-dusk-22b-r1.5.gguf -ngl 99 \ | |
| -c 32768 -fa on -ctk q8_0 -ctv q8_0 \ | |
| --temp 1.0 --top-p 0.95 --top-k 20 \ | |
| --reasoning-format deepseek | |
| ``` | |
| **Get the file** β `marlowe-dusk-22b-r1.5.gguf` is this release; `marlowe-dusk-22b-r1.gguf` is the previous | |
| round, kept for comparison. | |
| ```bash | |
| huggingface-cli download devvexus/marlowe-22b marlowe-dusk-22b-r1.5.gguf --local-dir . | |
| ``` | |
| The chat template ships inside the GGUF, so `llama-server` applies it for you β including the | |
| thinking block, which `--reasoning-format deepseek` returns as `reasoning_content`. | |
| Any runtime that speaks GGUF and supports the hybrid `qwen3_5` architecture (DeltaNet linear | |
| attention interleaved with full attention β the architecture family Qwen3.8-27B is built on) will | |
| load it; build llama.cpp from a revision that includes that support. | |
| **Sampling β use temperature 1.0, top-p 0.95, top-k 20** (the parent's thinking preset), and do not | |
| decode greedily: at temperature 0, 24 of 59 traces that used their full thinking budget contained | |
| verbatim repetition loops. This is a property of the parent's thinking mode, inherited rather than | |
| introduced by pruning, and sampling is the fix for both models. | |
| The chat template defaults to `reasoning_effort="xhigh"`; `medium` and `low` are supported, as is | |
| `enable_thinking=False`. | |
| **Speed and memory.** On a 16 GB RTX 4080 SUPER at Q4_K_M with every layer resident and q8_0 KV: | |
| **~36 tokens/s per stream with two concurrent streams.** Weights are 13.74 GB; a 32K context adds | |
| roughly 1.1 GB. | |
| ## Roadmap β three training rounds scheduled, plus two acceleration variants | |
| | round | goal | status | | |
| |---|---|---| | |
| | R1 | heal the pruning loss broadly | done β 27.5 M tokens | | |
| | R1.5 | close specific probed fact losses | **this release** | | |
| | **R1.7** | match the 27B on long-context, reasoning-heavy work; close the rotation/trig arithmetic gap | measurement built and running | | |
| | **R1.8** | professional depth across every major engineering field | scheduled | | |
| | **R2** | tool use and agentic execution on the strengthened base | scheduled | | |
| From R1.7 onward, every round is gated on the full instrument set β retention, derivation, | |
| arithmetic fidelity and long-horizon KL β with checkpoints scored **during** training rather than | |
| only at the end, so a round that trades one capability for another is caught while it runs. That | |
| discipline came from a round whose in-loop metric improved while a capability it never measured | |
| regressed; the fix was to widen the instruments, and it is now a standing gate. | |
| Two acceleration variants are planned on the same weights: **-super** (a multi-token-prediction | |
| draft head, 1.3β1.7Γ on code) and **-turbo** (EAGLE-3, 3β4Γ on supporting runtimes). This release | |
| is the dense build: `mtp_num_hidden_layers` is 0. | |
| The same pipeline then ladders downward β an 18B distilled from this model's successor, and a 9B | |
| distilled into a native base β reusing the teacher caches and evaluation banks built here. | |
| ## Scope of this release | |
| - **GGUF only, for now.** This repository contains the Q4_K_M build and nothing else β no | |
| safetensors, config or tokenizer files, so `transformers`, vLLM and SGLang cannot load it from | |
| here. Use llama.cpp or another GGUF runtime. The bf16 weights are planned for a later upload. | |
| - **Text only.** Image and video inputs are not supported; the vision tower was dropped to spend | |
| the budget on reasoning. | |
| - **No public benchmark scores are claimed.** The held-out benchmark subsets here are too small to | |
| publish, and none were used in training. Results above are from the project's own instruments, | |
| all of which are described in full so they can be argued with. | |
| - **Long-horizon parity against the parent is being measured now** and will be reported with R1.7. | |
| - It inherits the parent's biases, refusals and knowledge cutoff; no alignment training was added. | |
| - It thinks at length. On multi-stage engineering calculations it will carry more intermediate | |
| precision than the answer needs β ask explicitly for the precision you want, and give it room. | |
| ## Training data and attribution | |
| Public data, subsampled: **arXiv, open-web-math, AlgebraicStack** (permissive-licensed subset of | |
| The Stack) and **FineWeb-Edu** β all **ODC-By**, whose attribution travels with this model β plus | |
| **PG-19** (public domain). Reasoning traces were generated by the parent model itself. No benchmark | |
| test sets were included; the corpus was hash-checked against the evaluation banks. | |
| ## License | |
| **Apache 2.0**, inherited from Qwen3.8-27B. This is a derivative work of that model; the parent's | |
| terms apply to it and to anything derived from it. | |