StillDeadcode commited on
Commit
7d5ba7a
·
verified ·
1 Parent(s): c13ae03

Model card: the exact command that made the file

Browse files
Files changed (1) hide show
  1. README.md +49 -6
README.md CHANGED
@@ -22,13 +22,14 @@ tags:
22
  [openbmb/MiniCPM5-2B](https://huggingface.co/openbmb/MiniCPM5-2B) as a single `.rad` container for
23
  the **radiance** inference engine (AMD RDNA4, ROCm), with the
24
  [MiniCPM5-2B-DSpark](https://huggingface.co/openbmb/MiniCPM5-2B-DSpark) block drafter merged in for
25
- speculative decoding.
 
26
 
27
  | | |
28
  |---|---|
29
  | File | `minicpm5-2b-fp8.rad` — 3.03 GiB |
30
- | Weights | every linear FP8 E4M3 with a bf16 scale per 128×128 block (round to nearest); embeddings and norms bf16 |
31
- | Speculator | DSpark (5 layers, block size 7), FP8; its vocabulary head at 2 bits |
32
  | Context | 131,072 tokens (the checkpoint's trained length) |
33
  | Modalities | text |
34
 
@@ -39,8 +40,50 @@ radiance --model minicpm5-2b-fp8.rad --tp 2 --max-model-len 131072 --kv-cache-dt
39
  --max-num-seqs 8 --host 0.0.0.0 --port 8000
40
  ```
41
 
42
- `--tp 1` serves it on one card. The server speaks the OpenAI API (`/v1/chat/completions`,
43
- `/v1/completions`) with tool calls and structured output; the drafter's depth is chosen
44
- automatically (`--num-speculative-tokens N` states one, `0` turns speculation off).
45
 
46
  Built and tested on Radeon AI PRO R9700 (gfx1201).
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
22
  [openbmb/MiniCPM5-2B](https://huggingface.co/openbmb/MiniCPM5-2B) as a single `.rad` container for
23
  the **radiance** inference engine (AMD RDNA4, ROCm), with the
24
  [MiniCPM5-2B-DSpark](https://huggingface.co/openbmb/MiniCPM5-2B-DSpark) block drafter merged in for
25
+ speculative decoding. A bf16 version is at
26
+ [StillDeadcode/minicpm5-2b-bf16](https://huggingface.co/StillDeadcode/minicpm5-2b-bf16).
27
 
28
  | | |
29
  |---|---|
30
  | File | `minicpm5-2b-fp8.rad` — 3.03 GiB |
31
+ | Weights | every linear, the lm_head included, FP8 E4M3 with a bf16 scale per 128×128 block (round to nearest); embeddings and norms bf16 |
32
+ | Speculator | DSpark (5 layers, block size 7), FP8; its vocabulary head as 2-bit codes |
33
  | Context | 131,072 tokens (the checkpoint's trained length) |
34
  | Modalities | text |
35
 
 
40
  --max-num-seqs 8 --host 0.0.0.0 --port 8000
41
  ```
42
 
43
+ `--tp 1` serves it on one card. The server speaks the OpenAI API (`/v1/chat/completions`, `/v1/completions`), with tool calls and
44
+ structured output; the drafter's
45
+ depth is chosen automatically (`--num-speculative-tokens N` states one, `0` turns speculation off).
46
 
47
  Built and tested on Radeon AI PRO R9700 (gfx1201).
48
+
49
+ ## How this file was made
50
+
51
+ The exact command, run on 2026-10-01 from a radiance build directory (`radiance_home` is that
52
+ build's plugin directory):
53
+
54
+ ```sh
55
+ ./bin/rad-convert ~/models/openbmb/MiniCPM5-2B -o ~/models/rad/minicpm5-2b-dspark.v2.rad \
56
+ --home radiance_home --tokenizer ~/models/openbmb/MiniCPM5-2B/tokenizer.json \
57
+ --draft-model ~/models/openbmb/MiniCPM5-2B-DSpark --recipe ../data/recipes/minicpm5-2b-dspark.recipe
58
+ ```
59
+
60
+ The published file is that output under this repository's name, with the local directory removed
61
+ from the `created by` string in its header; the weights are the converter's bytes. The same
62
+ conversion with the radiance image (`docker/build.sh` in the radiance repository, commit 93bbe22)
63
+ reproduces them byte for byte:
64
+
65
+ ```sh
66
+ docker run --rm -v ~/models:/models --entrypoint rad-convert radiance \
67
+ /models/openbmb/MiniCPM5-2B --draft-model /models/openbmb/MiniCPM5-2B-DSpark \
68
+ --tokenizer /models/openbmb/MiniCPM5-2B/tokenizer.json \
69
+ --recipe /opt/radiance/share/radiance/recipes/minicpm5-2b-dspark.recipe \
70
+ -o /models/minicpm5-2b-fp8.rad
71
+ ```
72
+
73
+ The recipe (`rad-info --recipe minicpm5-2b-fp8.rad`):
74
+
75
+ ```
76
+ output.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
77
+ dspark.draft_head.weight rtn codes=u2 zero=u8 group=128 scale=f16
78
+ blk.*.attn_qkv.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
79
+ blk.*.attn_output.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
80
+ blk.*.ffn_gate_up.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
81
+ blk.*.ffn_down.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
82
+ dspark.fc.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
83
+ dspark.blk.*.attn_q.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
84
+ dspark.blk.*.attn_k.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
85
+ dspark.blk.*.attn_v.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
86
+ dspark.blk.*.attn_output.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
87
+ dspark.blk.*.ffn_gate_up.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
88
+ dspark.blk.*.ffn_down.weight rtn codes=fp8_e4m3 block=128x128 scale=bf16
89
+ ```