dots3-note-prev NVFP4

Temporary withdrawal — 2026-08-14: Do not download or serve the current weight shards. A full audit found incompatible global scales in 11,266 of 11,565 fused expert gate/up pairs. vLLM accepts only one global scale for each fused pair, so the published checkpoint can be numerically misloaded. A corrected conversion is in progress and will replace this warning only after structural, numerical, generation, tool-call, and serving checks pass.

hero

Mixed NVFP4 conversion of dots-studio/dots3-note-prev (Apache-2.0). The intended policy is NVFP4 routed/shared expert projections; MLA, DSA, MTP, vision, and audio remain FP8 or BF16. No expert is dropped or reordered.

The upstream benchmark images previously shown here are not evidence for this quantized checkpoint and have been removed from the card while correction is underway.

Downloads last month
-
Safetensors
Model size
152B params
Tensor type
BF16
·
F32
·
F8_E4M3
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Frosty40/dots3-note-prev-NVFP4

Quantized
(1)
this model