Qwen3.8-27B Uncensored FP8 + DFlash2 β€” radiance container

orcarouter/Qwen3.8-27B-Uncensored-FP8 β€” Qwen3.8-27B-FP8 with its refusal direction abliterated β€” as a single .rad container for the radiance inference engine (AMD RDNA4, ROCm), with its vision tower and the z-lab/Qwen3.8-27B-DFlash2 block-diffusion drafter merged in for speculative decoding.

This model's safety alignment has been removed: it answers requests the original model refuses. The source model's card describes how it was made, how it was evaluated and what it is for.

Needs radiance 1.2.3 or newer. The drafter's MXFP4 linears and the rescored 2-bit head arrive in 1.2.3; older releases refuse this file. The previous version of this file (fp8 drafter) is in this repository's history.

File qwen3.8-27b-uncensored-fp8.rad β€” 28.67 GiB
Weights the uncensored FP8 checkpoint as it ships: E4M3 with a scale per 128Γ—128 block; the lm_head made block FP8 by the recipe
Speculator DFlash2, its linears OCP-MXFP4 from its bf16 release; its vocabulary head as 2-bit codes, whose top 32 candidates the engine rescores against the model's own FP8 head. It was trained on the original model and drafts for this one too: drafts are verified, so they change the speed and never the output
Vision the 27-block vision tower, bf16: images and video in chat requests
Context 262,144 tokens trained; 200K tested

Speed

2Γ— Radeon AI PRO R9700 (gfx1201), --tp 2:

this file previous file
one stream, all eight prompts (tok/s) 139 124
one stream, prose / code (tok/s) 86 / 177 77 / 161
four streams, total (tok/s) 331 292

Greedy, 768 tokens each of eight prompts (prose, code, an explanation, JSON, Rust, German, two with reasoning), depth 7, --max-num-seqs 8 --max-model-len 200000 --kv-cache-dtype fp8; "previous file" is this repository's earlier version (fp8 drafter, unrescored 2-bit head) served by 1.2.2 on the same machine.

Serve

radiance --model qwen3.8-27b-uncensored-fp8.rad --tp 2 --max-model-len 200000 --kv-cache-dtype fp8 \
    --max-num-seqs 8 --host 0.0.0.0 --port 8000

The server speaks the OpenAI API (/v1/chat/completions, /v1/completions), with tool calls and structured output, and image_url / video_url parts in chat messages (PNG, JPEG, WebP, GIF, BMP, TIFF, AVIF; MP4, WebM, MKV, MPEG-TS, AVI with H.264, HEVC, VP8, VP9 or AV1). The drafter's depth is chosen automatically (--num-speculative-tokens N states one, 0 turns speculation off).

Built and tested on 2Γ— Radeon AI PRO R9700 (gfx1201).

How this file was made

rad-convert orcarouter/Qwen3.8-27B-Uncensored-FP8 --draft-model z-lab/Qwen3.8-27B-DFlash2 \
    --recipe q38-27b-fp8-df2.recipe -o qwen3.8-27b-uncensored-fp8.rad

The recipe (q38-27b-fp8-df2.recipe in this repository); everything it does not name is the checkpoint's own:

output.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
dflash.draft_head.weight rtn codes=u2 zero=u8 group=128 scale=f16 rule=search
dflash.fc.weight rtn codes=fp4_e2m1 group=32 scale=e8m0
dflash.blk.*.attn_q.weight rtn codes=fp4_e2m1 group=32 scale=e8m0
dflash.blk.*.attn_k.weight rtn codes=fp4_e2m1 group=32 scale=e8m0
dflash.blk.*.attn_v.weight rtn codes=fp4_e2m1 group=32 scale=e8m0
dflash.blk.*.attn_output.weight rtn codes=fp4_e2m1 group=32 scale=e8m0
dflash.blk.*.ffn_gate_up.weight rtn codes=fp4_e2m1 group=32 scale=e8m0
dflash.blk.*.ffn_down.weight rtn codes=fp4_e2m1 group=32 scale=e8m0
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for StillDeadcode/qwen3.8-27b-uncensored-fp8

Base model

Qwen/Qwen3.8-27B
Finetuned
(1)
this model