Qwen3.8-27B Uncensored FP8 + DFlash2 β radiance container
orcarouter/Qwen3.8-27B-Uncensored-FP8
β Qwen3.8-27B-FP8 with its refusal direction abliterated β as a single .rad container for the
radiance inference engine (AMD RDNA4, ROCm), with its vision tower and the
z-lab/Qwen3.8-27B-DFlash2 block-diffusion
drafter merged in for speculative decoding.
This model's safety alignment has been removed: it answers requests the original model refuses. The source model's card describes how it was made, how it was evaluated and what it is for.
Needs radiance 1.2.3 or newer. The drafter's MXFP4 linears and the rescored 2-bit head arrive in 1.2.3; older releases refuse this file. The previous version of this file (fp8 drafter) is in this repository's history.
| File | qwen3.8-27b-uncensored-fp8.rad β 28.67 GiB |
| Weights | the uncensored FP8 checkpoint as it ships: E4M3 with a scale per 128Γ128 block; the lm_head made block FP8 by the recipe |
| Speculator | DFlash2, its linears OCP-MXFP4 from its bf16 release; its vocabulary head as 2-bit codes, whose top 32 candidates the engine rescores against the model's own FP8 head. It was trained on the original model and drafts for this one too: drafts are verified, so they change the speed and never the output |
| Vision | the 27-block vision tower, bf16: images and video in chat requests |
| Context | 262,144 tokens trained; 200K tested |
Speed
2Γ Radeon AI PRO R9700 (gfx1201), --tp 2:
| this file | previous file | |
|---|---|---|
| one stream, all eight prompts (tok/s) | 139 | 124 |
| one stream, prose / code (tok/s) | 86 / 177 | 77 / 161 |
| four streams, total (tok/s) | 331 | 292 |
Greedy, 768 tokens each of eight prompts (prose, code, an explanation, JSON, Rust, German, two with reasoning), depth 7, --max-num-seqs 8 --max-model-len 200000 --kv-cache-dtype fp8; "previous file" is this repository's earlier version (fp8 drafter, unrescored 2-bit head) served by 1.2.2 on the same machine.
Serve
radiance --model qwen3.8-27b-uncensored-fp8.rad --tp 2 --max-model-len 200000 --kv-cache-dtype fp8 \
--max-num-seqs 8 --host 0.0.0.0 --port 8000
The server speaks the OpenAI API (/v1/chat/completions, /v1/completions), with tool calls and
structured output, and image_url / video_url parts in chat
messages (PNG, JPEG, WebP, GIF, BMP, TIFF, AVIF; MP4, WebM, MKV, MPEG-TS, AVI with H.264, HEVC, VP8,
VP9 or AV1). The drafter's depth is chosen automatically (--num-speculative-tokens N states one,
0 turns speculation off).
Built and tested on 2Γ Radeon AI PRO R9700 (gfx1201).
How this file was made
rad-convert orcarouter/Qwen3.8-27B-Uncensored-FP8 --draft-model z-lab/Qwen3.8-27B-DFlash2 \
--recipe q38-27b-fp8-df2.recipe -o qwen3.8-27b-uncensored-fp8.rad
The recipe (q38-27b-fp8-df2.recipe in this repository); everything it does not name is the
checkpoint's own:
output.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
dflash.draft_head.weight rtn codes=u2 zero=u8 group=128 scale=f16 rule=search
dflash.fc.weight rtn codes=fp4_e2m1 group=32 scale=e8m0
dflash.blk.*.attn_q.weight rtn codes=fp4_e2m1 group=32 scale=e8m0
dflash.blk.*.attn_k.weight rtn codes=fp4_e2m1 group=32 scale=e8m0
dflash.blk.*.attn_v.weight rtn codes=fp4_e2m1 group=32 scale=e8m0
dflash.blk.*.attn_output.weight rtn codes=fp4_e2m1 group=32 scale=e8m0
dflash.blk.*.ffn_gate_up.weight rtn codes=fp4_e2m1 group=32 scale=e8m0
dflash.blk.*.ffn_down.weight rtn codes=fp4_e2m1 group=32 scale=e8m0
Model tree for StillDeadcode/qwen3.8-27b-uncensored-fp8
Base model
Qwen/Qwen3.8-27B