|
Download README.md from StillDeadcode/minicpm5-2b-fp8: direct link, hf CLI and curl.
- Browser
- Download file 2.76 kB
-
https://huggingface.co/StillDeadcode/minicpm5-2b-fp8/resolve/main/README.md
- Command line
-
hf download hf://StillDeadcode/minicpm5-2b-fp8/README.md
-
curl -L -o README.md https://huggingface.co/StillDeadcode/minicpm5-2b-fp8/resolve/main/README.md
2.76 kB
metadata
license: apache-2.0
base_model:
- openbmb/MiniCPM5-2B
- openbmb/MiniCPM5-2B-DSpark
pipeline_tag: text-generation
language:
- en
- zh
tags:
- radiance
- rocm
- amd
- rdna4
- fp8
- speculative-decoding
- dspark
MiniCPM5-2B FP8 + DSpark — radiance container
openbmb/MiniCPM5-2B as a single .rad container for
the radiance inference engine (AMD RDNA4, ROCm), with the
MiniCPM5-2B-DSpark block drafter merged in for
speculative decoding. A bf16 version is at
StillDeadcode/minicpm5-2b-bf16.
| File | minicpm5-2b-fp8.rad — 3.03 GiB |
| Weights | every linear, the lm_head included, FP8 E4M3 with a bf16 scale per 128×128 block (round to nearest); embeddings and norms bf16 |
| Speculator | DSpark (5 layers, block size 7), FP8; its vocabulary head as 2-bit codes |
| Context | 131,072 tokens (the checkpoint's trained length) |
| Modalities | text |
Serve
radiance --model minicpm5-2b-fp8.rad --tp 2 --max-model-len 131072 --kv-cache-dtype fp8 \
--max-num-seqs 8 --host 0.0.0.0 --port 8000
--tp 1 serves it on one card. The server speaks the OpenAI API (/v1/chat/completions, /v1/completions), with tool calls and
structured output; the drafter's
depth is chosen automatically (--num-speculative-tokens N states one, 0 turns speculation off).
Built and tested on Radeon AI PRO R9700 (gfx1201).
How this file was made
rad-convert openbmb/MiniCPM5-2B --draft-model openbmb/MiniCPM5-2B-DSpark \
--tokenizer openbmb/MiniCPM5-2B/tokenizer.json \
--recipe minicpm5-2b-dspark.recipe -o minicpm5-2b-fp8.rad
The recipe (minicpm5-2b-dspark.recipe in this repository); everything it does not name is
the checkpoint's own:
output.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
dspark.draft_head.weight rtn codes=u2 zero=u8 group=128 scale=f16
blk.*.attn_qkv.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
blk.*.attn_output.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
blk.*.ffn_gate_up.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
blk.*.ffn_down.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
dspark.fc.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
dspark.blk.*.attn_q.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
dspark.blk.*.attn_k.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
dspark.blk.*.attn_v.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
dspark.blk.*.attn_output.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
dspark.blk.*.ffn_gate_up.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
dspark.blk.*.ffn_down.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16