minicpm5-2b-bf16 / README.md
StillDeadcode's picture
Model card: the command and the recipe
ef5a007 verified
|
Raw History Blame Contribute Delete
2.64 kB
---
license: apache-2.0
base_model:
- openbmb/MiniCPM5-2B
- openbmb/MiniCPM5-2B-DSpark
pipeline_tag: text-generation
language:
- en
- zh
tags:
- radiance
- rocm
- amd
- rdna4
- bf16
- speculative-decoding
- dspark
---
# MiniCPM5-2B BF16 + DSpark — radiance container
[openbmb/MiniCPM5-2B](https://huggingface.co/openbmb/MiniCPM5-2B) in bf16, exactly as the checkpoint
ships it, as a single `.rad` container for the **radiance** inference engine (AMD RDNA4, ROCm), with
the [MiniCPM5-2B-DSpark](https://huggingface.co/openbmb/MiniCPM5-2B-DSpark) block drafter merged in
for speculative decoding. An FP8 version is at
[StillDeadcode/minicpm5-2b-fp8](https://huggingface.co/StillDeadcode/minicpm5-2b-fp8).
| | |
|---|---|
| File | `minicpm5-2b-bf16.rad` — 5.12 GiB |
| Weights | every weight of the model, the lm_head included, bf16 as the checkpoint ships it |
| Speculator | DSpark (5 layers, block size 7). Its weights are FP8 E4M3 per 128×128 block and its vocabulary head 2-bit codes, the one form the engine's DSpark block reads; the drafter only proposes, and the bf16 model verifies every token it emits |
| Context | 131,072 tokens (the checkpoint's trained length) |
| Modalities | text |
## Serve
```sh
radiance --model minicpm5-2b-bf16.rad --tp 2 --max-model-len 131072 --kv-cache-dtype fp8 \
--max-num-seqs 8 --host 0.0.0.0 --port 8000
```
`--tp 1` serves it on one card. The server speaks the OpenAI API (`/v1/chat/completions`, `/v1/completions`), with tool calls and
structured output; the drafter's
depth is chosen automatically (`--num-speculative-tokens N` states one, `0` turns speculation off).
Built and tested on Radeon AI PRO R9700 (gfx1201).
## How this file was made
```sh
rad-convert openbmb/MiniCPM5-2B --draft-model openbmb/MiniCPM5-2B-DSpark \
--tokenizer openbmb/MiniCPM5-2B/tokenizer.json \
--recipe minicpm5-2b-bf16-dspark.recipe -o minicpm5-2b-bf16.rad
```
The recipe (`minicpm5-2b-bf16-dspark.recipe` in this repository) quantises the drafter only;
everything it does not name is the checkpoint's bf16:
```
dspark.draft_head.weight rtn codes=u2 zero=u8 group=128 scale=f16
dspark.fc.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
dspark.blk.*.attn_q.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
dspark.blk.*.attn_k.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
dspark.blk.*.attn_v.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
dspark.blk.*.attn_output.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
dspark.blk.*.ffn_gate_up.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
dspark.blk.*.ffn_down.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
```