FP8 checkpoint for vLLM, measured against bf16

#2
by fogf - opened

Hi, thanks for releasing decider under Apache 2.0.

We quantized decider-0.8b v1 for vLLM:

  • llmtech/decider-0.8b-fp8: 1.0 GB (bf16 1.5 GB). Peak prefill 1.17x to 1.24x of bf16 on an RTX PRO 6000 Blackwell (32K to 1K tokens). Accuracy -0.1 in-task, 0.0 held-out on your regression set.

Method: your 95-task regression set rebuilt from public data (144,226 rows) and the 231 public JevBench items, bf16 and quantized both run through vLLM 0.29.0. You have not published numbers for the 0.8B on this protocol, so the reference is our bf16 run; the same harness matches your published 2B and 4B numbers within 0.0005.

Your tokenizer and chat template are unchanged; decider_config.json is yours with only the version and quantization fields changed. The checkpoint runs with decider.serve_vllm.

The model card has the full aggregate table, the largest per-task drops, speed at 1K/8K/32K and the quantization details. Tell us if you spot anything off.

Sign up or log in to comment