DeepSeek-v4-Flash

Pulsar's build of DeepSeek-v4-Flash, derived from deepseek-ai/DeepSeek-V4-Flash-Vision-Exp at revision 6821d6ad3681a4b137b066b76094fa82ebd0a380.

This is the unqualified member of the family: the source-precision build. Every routed expert is stored in cutlass_mxfp4, repacked from the upstream checkpoint's own FP4 bytes โ€” nothing is quantized down, and there is no iq2_xxs_mmq_k tensor anywhere. Variants that deviate from source precision carry a suffix (-IQ2, -Mix-โ€ฆ); the bare name always means this one.

At ~168 GB it does not fit a single GB10. It is the artifact for a two-Spark rig; the 2-bit build (DeepSeek-v4-Flash-IQ2, ~92 GB) is the single-Spark one.

Layout

Layout Used by Bytes
cutlass_mxfp4 main routed experts (43 layers ร— 256) and drafter experts (3 ร— 256) per-expert E2M1 data then a swizzled E8M0 group-32 SF tile
mxfp8_lt attention / dense / shared-expert projections de-interleaved E4M3 then a swizzled E8M0 group-32 scale
fp8_e4m3_soa_k one drafter head projection E8M0 scale plane [rows][cols/32] then an E4M3 payload plane
native norms, router, embeddings, hyper-connection params plain BF16 / F32 / I32 as declared in the header

Engine-layout tensors are stored as flat U8 blobs. Each shard's __metadata__ is self-describing:

  • pulsar.tensors โ€” per tensor: layout and dims_ne (the logical shape in ggml ne[] order, which is the HF shape reversed; recorded for every tensor so a reader never has to reason about orientation).
  • pulsar.experts โ€” per routed-expert projection: n_experts, expert_bytes, layout, contiguous, and the original stacked tensor name. The expert tensors of a projection are written back-to-back with no inter-expert padding, so a kernel can address them as base + xid * expert_bytes from a single pointer.
  • pulsar.kv_arch โ€” the 53 architecture keys the engine reads. pulsar.kv (primary shard only) additionally carries the tokenizer block.
  • format โ€” pt, and pulsar.alignment โ€” 32. Every tensor offset is aligned to 32 bytes, which the swizzled scale tiles depend on.

Sizes

component upstream this artifact delta
routed-expert payloads 157.437 GB 157.437 GB 0
everything else 10.374 GB 10.529 GB +155.2 MB
total 167.819 GB 167.979 GB +0.09%

The expert payloads are byte-count identical โ€” 157,437,394,944 B in both โ€” and a pure permutation: the upstream E2M1 nibble plane is copied verbatim and the upstream E8M0 group-32 plane is scattered into the CUTLASS SFB swizzle. The +155 MB is the dense scale plane: mxfp8_lt carries one E8M0 scale per 32 columns, where the upstream FP8 weight block is 128ร—128. That re-encode is lossless to within E4M3 rounding on 0.0013% of dense elements (max |ฮ”| 1.5e-5).

Loading

This runs on pulsar only. Stock transformers / vLLM cannot load it โ€” they have no decoder for the fused layouts, and U8 blobs have no meaning without the metadata above. config.json's quantization_config.quant_method is pulsar-safetensors precisely so a standard loader fails loudly instead of misreading the bytes.

Verification

  • 48 shards, 36,909 tensors, index closed both ways, gap-free, 32 B aligned, exact file lengths.
  • 138/138 routed-expert projections cutlass_mxfp4 and contiguous.
  • Zero shards contain the substring iq2.
  • 45/45 spot-checked expert blobs invert to the upstream E2M1+E8M0 bytes exactly.
  • Non-expert tensors byte-identical to the base container (48/48 shard digests).
  • Dense values vs upstream: 83,326 of 6,304,038,912 elements differ, max |ฮ”| 1.5e-5.

See BUILD-RECORD.md for the full provenance and gate list.

Downloads last month
382
Safetensors
Model size
166B params
Tensor type
BF16
ยท
F32
ยท
U8
ยท
I32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for ElytronAI/DeepSeek-v4-Flash

Quantized
(27)
this model