Qwen3.8-27B-Splash (mirror)
⚠️ Settings: read this before you run it
| Setting | Use this | Why |
|---|---|---|
temperature |
0.8 | What I settled on for Qwen3.8-27B after it looped at Qwen3.6's settings. |
top_p |
0.9 | Same tuning. Honored by the runtime. |
top_k |
32 | Honored. Must be an integer from 1 to 32. I'd tuned 50 on other runtimes, but here anything above 32, or "off", returns HTTP 400: Splash top-k must be an integer from 1 to 32. |
max_tokens |
cap it, ~20k | This is your only loop guard. The Splash runtime has no working repetition penalty, so a runaway turn only ends at the cap or the timeout. |
repeat_penalty, presence_penalty, frequency_penalty |
don't rely on them | The Splash runtime ignores all three. I verified this on the 3.6 build, and it's the same runtime. |
| Thinking | on by default | Short answers stay short: "say hi in three words" took 117 tokens here, against 2,815 on the 3.6 build. |
LM Studio per-model default (saved as ~/.lmstudio/.internal/user-concrete-model-default-config/incoai/Qwen3.8-27B-Splash.json, which LM Studio reads live):
{
"preset": "",
"operation": {
"fields": [
{ "key": "llm.prediction.temperature", "value": 0.8 },
{ "key": "llm.prediction.topPSampling", "value": { "checked": true, "value": 0.9 } },
{ "key": "llm.prediction.topKSampling", "value": 32 },
{ "key": "llm.prediction.contextOverflowPolicy", "value": "rollingWindow" }
]
},
"load": { "fields": [] }
}
A byte-for-byte mirror of incoai/Qwen3.8-27B-Splash, kept here as a pinned copy for my local agent stack. All credit for the packing, the DFlash 2 draft, and the Splash engine goes to Inco AI. Read their card first. This one only adds what I found running the Splash runtime.
Verified identical to upstream on 2026-09-23: the artifact_set_sha256 in manifest.json matches incoai's published manifest (b0ec3657…8da81).
What's in the package
It's not a Transformers or MLX checkpoint. It only loads in Splash, or in LM Studio's Splash runtime.
| Path | Contents |
|---|---|
target/ |
Qwen3.8-27B dense, 4-bit packed (splash-packed-q4) |
draft/ |
DFlash 2 draft model, 5 layers, proposes 7 tokens per step |
vision/ |
Vision encoder |
tokenizer/ |
Tokenizer and chat template |
manifest.json |
Artifact hashes, execution geometry, upstream revisions |
Upstream sources, from the manifest: target, tokenizer, and vision from mlx-community/Qwen3.8-27B-4bit @ 3e6447f, draft from incoai/Qwen3.8-27B-DFlash2 @ dedf8df. About 17.4 GB on disk.
Running it
LM Studio: it shows up as qwen3.8-27b-splash and reports model format yuzu.
Splash: upstream's command is splash serve --model incoai/Qwen3.8-27B-Splash. I haven't tried pointing Splash at this mirror's repo id, so use upstream's if you're going that route.
Measured
LM Studio on an M5 Max (128 GB), OpenAI-compatible endpoint, non-streaming, at the settings above. Timings include prefill:
| Prompt | Tokens | Time | tok/s |
|---|---|---|---|
| Say hi in three words | 117 | 1.4 s | 82 |
| Capital of France | 33 | 0.3 s | 100 |
| Haiku about rain | 208 | 2.1 s | 99 |
| 300-word story | 3,636 | 35.4 s | 103 |
So it runs at about 100 tok/s. The Qwen3.8-27B MTPLX build did about 24 tok/s in the same agent. It's not as fast as the 3.6 MoE (about 148 tok/s), but it's far less wordy in its reasoning, and in my agent that matters more.
Sampling notes (LM Studio's Splash runtime, 2026-09-23)
I tested these on the 3.6 sibling, not on this model. It's the same runtime, so I'd expect the same behavior, but that's unverified:
- Honored:
temperature,top_p,top_k, plus LM Studio's per-model config defaults. - Ignored:
repeat_penalty,presence_penalty,frequency_penalty.
With no working repetition penalty, the 3.6 build fell into a reasoning loop on its first evening (about 112k tokens of one repeated line, and no answer). If you run this behind an agent, cap max_tokens. Also, avoid greedy decoding, because Qwen's guidance is that it makes endless repetition more likely.
If you drive LM Studio from code
@lmstudio/sdk 1.5.0 hangs on llm.listLoaded() whenever a Splash model is loaded (it doesn't know the yuzu format). 2.0.0 fixes it.
License
Apache 2.0, same as upstream and the Qwen base model.