Subject: NVFP4 quant of decider-0.8b/2b/4b
Hello,
I run LLM Tech, an inference provider in the EU. We run inference on rented RTX PRO 6000 Blackwell GPUs and we are building NVFP4 quantized versions of decider-0.8b, decider-2b and decider-4b for inference on that hardware.
We plan to publish the quants on Hugging Face with attribution to your original decider models and to Qwen3.5-*-Base, under Apache 2.0, along with a comparison against your bf16 checkpoints on your own evaluation sets so the accuracy difference is visible, not just claimed. We also want to host these quants on OpenRouter as a provider.
Nothing here needs your permission under Apache 2.0, so this is a heads-up and an offer:
- We will run your regression set against our quant and share the numbers with you.
- If the results hold up, would you consider linking to our quants from your own model cards?
Happy to send you the checkpoints or numbers directly once we have them.
Thank you,
Artem Burei
LLM Tech
artem@llmtech.eu
+48 690 978 945