|
Download README.md from StillDeadcode/minicpm5-2b-fp8: direct link, hf CLI and curl.
- Browser
- Download file 2.76 kB
-
https://huggingface.co/StillDeadcode/minicpm5-2b-fp8/resolve/main/README.md
- Command line
-
hf download hf://StillDeadcode/minicpm5-2b-fp8/README.md
-
curl -L -o README.md https://huggingface.co/StillDeadcode/minicpm5-2b-fp8/resolve/main/README.md
2.76 kB
| license: apache-2.0 | |
| base_model: | |
| - openbmb/MiniCPM5-2B | |
| - openbmb/MiniCPM5-2B-DSpark | |
| pipeline_tag: text-generation | |
| language: | |
| - en | |
| - zh | |
| tags: | |
| - radiance | |
| - rocm | |
| - amd | |
| - rdna4 | |
| - fp8 | |
| - speculative-decoding | |
| - dspark | |
| # MiniCPM5-2B FP8 + DSpark — radiance container | |
| [openbmb/MiniCPM5-2B](https://huggingface.co/openbmb/MiniCPM5-2B) as a single `.rad` container for | |
| the **radiance** inference engine (AMD RDNA4, ROCm), with the | |
| [MiniCPM5-2B-DSpark](https://huggingface.co/openbmb/MiniCPM5-2B-DSpark) block drafter merged in for | |
| speculative decoding. A bf16 version is at | |
| [StillDeadcode/minicpm5-2b-bf16](https://huggingface.co/StillDeadcode/minicpm5-2b-bf16). | |
| | | | | |
| |---|---| | |
| | File | `minicpm5-2b-fp8.rad` — 3.03 GiB | | |
| | Weights | every linear, the lm_head included, FP8 E4M3 with a bf16 scale per 128×128 block (round to nearest); embeddings and norms bf16 | | |
| | Speculator | DSpark (5 layers, block size 7), FP8; its vocabulary head as 2-bit codes | | |
| | Context | 131,072 tokens (the checkpoint's trained length) | | |
| | Modalities | text | | |
| ## Serve | |
| ```sh | |
| radiance --model minicpm5-2b-fp8.rad --tp 2 --max-model-len 131072 --kv-cache-dtype fp8 \ | |
| --max-num-seqs 8 --host 0.0.0.0 --port 8000 | |
| ``` | |
| `--tp 1` serves it on one card. The server speaks the OpenAI API (`/v1/chat/completions`, `/v1/completions`), with tool calls and | |
| structured output; the drafter's | |
| depth is chosen automatically (`--num-speculative-tokens N` states one, `0` turns speculation off). | |
| Built and tested on Radeon AI PRO R9700 (gfx1201). | |
| ## How this file was made | |
| ```sh | |
| rad-convert openbmb/MiniCPM5-2B --draft-model openbmb/MiniCPM5-2B-DSpark \ | |
| --tokenizer openbmb/MiniCPM5-2B/tokenizer.json \ | |
| --recipe minicpm5-2b-dspark.recipe -o minicpm5-2b-fp8.rad | |
| ``` | |
| The recipe (`minicpm5-2b-dspark.recipe` in this repository); everything it does not name is | |
| the checkpoint's own: | |
| ``` | |
| output.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16 | |
| dspark.draft_head.weight rtn codes=u2 zero=u8 group=128 scale=f16 | |
| blk.*.attn_qkv.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16 | |
| blk.*.attn_output.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16 | |
| blk.*.ffn_gate_up.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16 | |
| blk.*.ffn_down.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16 | |
| dspark.fc.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16 | |
| dspark.blk.*.attn_q.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16 | |
| dspark.blk.*.attn_k.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16 | |
| dspark.blk.*.attn_v.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16 | |
| dspark.blk.*.attn_output.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16 | |
| dspark.blk.*.ffn_gate_up.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16 | |
| dspark.blk.*.ffn_down.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16 | |
| ``` | |