Qwen3.8-27B MXFP4 + DFlash2 β radiance container
amd/Qwen3.8-27B-Quark-AWQ-MXFP4 β
Qwen3.8-27B with AWQ-calibrated OCP-MXFP4 weights from AMD Quark β as a single .rad container for
the radiance inference engine (AMD RDNA4, ROCm), with its vision tower and the
z-lab/Qwen3.8-27B-DFlash2 block-diffusion
drafter merged in for speculative decoding.
Needs radiance 1.2.3 or newer. The drafter's MXFP4 linears and the rescored 2-bit head arrive in 1.2.3; older releases refuse this file. The previous version of this file (fp8 drafter) is in this repository's history.
| File | qwen3.8-27b-mxfp4.rad β 18.05 GiB |
| Weights | AMD's MXFP4 trunk exactly as released β e2m1 codes, an E8M0 exponent per 32 β served at 4-bit weights against E4M3 activations; the lm_head made block FP8 by the recipe |
| Speculator | DFlash2, its linears OCP-MXFP4 from its bf16 release; its vocabulary head as 2-bit codes, whose top 32 candidates the engine rescores against the model's own FP8 head. Drafts are verified, so they change the speed and never the output |
| Vision | the 27-block vision tower, bf16: images and video in chat requests |
| Context | 262,144 tokens trained; 200K tested |
Speed
Radeon AI PRO R9700 (gfx1201), one card (--tp 1) and two (--tp 2):
| this file, TP1 | this file, TP2 | previous file, TP2 | |
|---|---|---|---|
| one stream, all eight prompts (tok/s) | 122 | 204 | 182 |
| one stream, prose / code (tok/s) | 77 / 164 | 130 / 285 | 112 / 256 |
| four streams, total (tok/s) | 280 | 438 | 387 |
Greedy, 768 tokens each of eight prompts (prose, code, an explanation, JSON, Rust, German, two with reasoning), depth 7, --max-num-seqs 8 --max-model-len 200000 --kv-cache-dtype fp8; "previous file" is this repository's earlier version (fp8 drafter, unrescored 2-bit head) served by 1.2.2 on the same machine.
AMD reports the quality of these weights on its model card (GSM8K 5-shot, 99β102% of bf16); this file serves the same weights.
Serve
radiance --model qwen3.8-27b-mxfp4.rad --tp 2 --max-model-len 200000 --kv-cache-dtype fp8 \
--max-num-seqs 8 --host 0.0.0.0 --port 8000
It also fits one 32 GB card: --tp 1, the same flags otherwise.
The server speaks the OpenAI API (/v1/chat/completions, /v1/completions), with tool calls and
structured output, and image_url / video_url parts in chat messages. The drafter's depth is
chosen automatically (--num-speculative-tokens N states one, 0 turns speculation off).
How this file was made
rad-convert reads Quark's MXFP4 linears as they ship (e2m1 codes and E8M0 exponents, byte for byte), so the recipe names only what the release left in bf16:
rad-convert amd/Qwen3.8-27B-Quark-AWQ-MXFP4 --draft-model z-lab/Qwen3.8-27B-DFlash2 \
--recipe q38-27b-quark-mxfp4-df2.recipe -o qwen3.8-27b-mxfp4.rad
The recipe (q38-27b-quark-mxfp4-df2.recipe in this repository); everything it does not name is
the checkpoint's own:
blk.*.ssm_ab.weight cast dtype=bf16
output.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
dflash.draft_head.weight rtn codes=u2 zero=u8 group=128 scale=f16 rule=search
dflash.fc.weight rtn codes=fp4_e2m1 group=32 scale=e8m0
dflash.blk.*.attn_q.weight rtn codes=fp4_e2m1 group=32 scale=e8m0
dflash.blk.*.attn_k.weight rtn codes=fp4_e2m1 group=32 scale=e8m0
dflash.blk.*.attn_v.weight rtn codes=fp4_e2m1 group=32 scale=e8m0
dflash.blk.*.attn_output.weight rtn codes=fp4_e2m1 group=32 scale=e8m0
dflash.blk.*.ffn_gate_up.weight rtn codes=fp4_e2m1 group=32 scale=e8m0
dflash.blk.*.ffn_down.weight rtn codes=fp4_e2m1 group=32 scale=e8m0
The delta net's a/b projection is declared bf16 and keeps AMD's values that way.