DeepSeek-v4-Flash
Pulsar's build of DeepSeek-v4-Flash, derived from
deepseek-ai/DeepSeek-V4-Flash-Vision-Exp at revision
6821d6ad3681a4b137b066b76094fa82ebd0a380.
This is the unqualified member of the family: the source-precision build.
Every routed expert is stored in cutlass_mxfp4, repacked from the upstream
checkpoint's own FP4 bytes โ nothing is quantized down, and there is no
iq2_xxs_mmq_k tensor anywhere. Variants that deviate from source precision
carry a suffix (-IQ2, -Mix-โฆ); the bare name always means this one.
At ~168 GB it does not fit a single GB10. It is the artifact for a two-Spark
rig; the 2-bit build (DeepSeek-v4-Flash-IQ2, ~92 GB) is the single-Spark one.
Layout
| Layout | Used by | Bytes |
|---|---|---|
cutlass_mxfp4 |
main routed experts (43 layers ร 256) and drafter experts (3 ร 256) | per-expert E2M1 data then a swizzled E8M0 group-32 SF tile |
mxfp8_lt |
attention / dense / shared-expert projections | de-interleaved E4M3 then a swizzled E8M0 group-32 scale |
fp8_e4m3_soa_k |
one drafter head projection | E8M0 scale plane [rows][cols/32] then an E4M3 payload plane |
native |
norms, router, embeddings, hyper-connection params | plain BF16 / F32 / I32 as declared in the header |
Engine-layout tensors are stored as flat U8 blobs. Each shard's
__metadata__ is self-describing:
pulsar.tensorsโ per tensor:layoutanddims_ne(the logical shape in ggmlne[]order, which is the HF shape reversed; recorded for every tensor so a reader never has to reason about orientation).pulsar.expertsโ per routed-expert projection:n_experts,expert_bytes,layout,contiguous, and the original stacked tensor name. The expert tensors of a projection are written back-to-back with no inter-expert padding, so a kernel can address them asbase + xid * expert_bytesfrom a single pointer.pulsar.kv_archโ the 53 architecture keys the engine reads.pulsar.kv(primary shard only) additionally carries the tokenizer block.formatโpt, andpulsar.alignmentโ32. Every tensor offset is aligned to 32 bytes, which the swizzled scale tiles depend on.
Sizes
| component | upstream | this artifact | delta |
|---|---|---|---|
| routed-expert payloads | 157.437 GB | 157.437 GB | 0 |
| everything else | 10.374 GB | 10.529 GB | +155.2 MB |
| total | 167.819 GB | 167.979 GB | +0.09% |
The expert payloads are byte-count identical โ 157,437,394,944 B in both โ and
a pure permutation: the upstream E2M1 nibble plane is copied verbatim and the
upstream E8M0 group-32 plane is scattered into the CUTLASS SFB swizzle. The
+155 MB is the dense scale plane: mxfp8_lt carries one E8M0 scale per 32
columns, where the upstream FP8 weight block is 128ร128. That re-encode is
lossless to within E4M3 rounding on 0.0013% of dense elements (max |ฮ| 1.5e-5).
Loading
This runs on pulsar only. Stock transformers / vLLM cannot load it โ they
have no decoder for the fused layouts, and U8 blobs have no meaning without
the metadata above. config.json's quantization_config.quant_method is
pulsar-safetensors precisely so a standard loader fails loudly instead of
misreading the bytes.
Verification
- 48 shards, 36,909 tensors, index closed both ways, gap-free, 32 B aligned, exact file lengths.
- 138/138 routed-expert projections
cutlass_mxfp4and contiguous. - Zero shards contain the substring
iq2. - 45/45 spot-checked expert blobs invert to the upstream E2M1+E8M0 bytes exactly.
- Non-expert tensors byte-identical to the base container (48/48 shard digests).
- Dense values vs upstream: 83,326 of 6,304,038,912 elements differ, max |ฮ| 1.5e-5.
See BUILD-RECORD.md for the full provenance and gate list.
- Downloads last month
- 382
Model tree for ElytronAI/DeepSeek-v4-Flash
Base model
deepseek-ai/DeepSeek-V4-Flash-Vision-Exp