FP8 and NVFP4 checkpoints for vLLM, measured against bf16
Hi, thanks for releasing decider under Apache 2.0.
We quantized decider-4b v2.1 for vLLM and published two checkpoints:
- llmtech/decider-4b-nvfp4: 3.3 GB (bf16 8.4 GB). Peak prefill 1.67x to 2.00x of bf16 on an RTX PRO 6000 Blackwell (32K to 1K tokens). Accuracy -0.6 in-task, -0.7 held-out on your regression set.
- llmtech/decider-4b-fp8: 4.9 GB. Peak prefill 1.33x to 1.45x of bf16. Accuracy -0.1 in-task, -0.1 held-out.
Method: your 95-task regression set rebuilt from public data (144,226 rows) and the 231 public JevBench items, bf16 and quantized both run through vLLM 0.29.0. Our bf16 run matches your eval_results.json within 0.0005 on accuracy and NLL.
Your tokenizer and chat template are unchanged; decider_config.json is yours with only the version and quantization fields changed, temperatures included. decider-4b-nvfp4 was also run through decider.serve_vllm; decider-4b-fp8 only through vLLM directly.
The model cards have the full aggregate table, the largest per-task drops, speed at 1K/8K/32K and the quantization details. Tell us if you spot anything off.