FP8 checkpoint for vLLM, measured against bf16
#3
by fogf - opened
Hi, thanks for releasing decider under Apache 2.0.
We quantized decider-2b v11 for vLLM:
- llmtech/decider-2b-fp8: 2.4 GB (bf16 3.8 GB). Peak prefill 1.32x to 1.42x of bf16 on an RTX PRO 6000 Blackwell (32K to 1K tokens). Accuracy within 0.1 points of bf16 in-task and held-out on your regression set.
Method: your 95-task regression set rebuilt from public data (144,226 rows) and the 231 public JevBench items, bf16 and quantized both run through vLLM 0.29.0. Our bf16 run matches your eval_results.json within 0.0005 on accuracy and NLL.
Your tokenizer and chat template are unchanged; decider_config.json is yours with only the version and quantization fields changed, temperatures included. The checkpoint runs with decider.serve_vllm.
The model card has the full aggregate table, the largest per-task drops, speed at 1K/8K/32K and the quantization details. Tell us if you spot anything off.