TensorFold

DeepSeek-V4-Flash MTP for MLX

This is DeepSeek-V4-Flash's own multi-token-prediction (MTP) layer, converted for Apple Silicon. The mlx-community 4-bit conversion leaves it out. It drafts tokens for mlx-community/DeepSeek-V4-Flash-4bit on TensorFold's lane engine, which verifies every draft against the target.

  • Source: deepseek-ai/DeepSeek-V4-Flash at revision 60d8d70770c6776ff598c94bb586a859a38244f1. The layer is the mtp.0.* tensors of model-00046-of-00046.safetensors.
  • License: MIT, DeepSeek's (see LICENSE, copyright 2023 DeepSeek). The conversion adds no terms.
  • Files: model.safetensors (3.52 GB), config.json, LICENSE, SHA256SUMS.

Use

TensorFold 0.3.6.4 or later on a Mac with 256 GB:

tensorfold pull mlx-community/DeepSeek-V4-Flash-4bit TensorFold/DeepSeek-V4-Flash-MTP-MLX
tensorfold serve mlx-community/DeepSeek-V4-Flash-4bit --drafter TensorFold/DeepSeek-V4-Flash-MTP-MLX

The layer drafts up to three tokens a round, one pass each, from the target's streams at the last verified row. A drafted reply equals the same request sent with "draft": false. Drafts change speed, never tokens. TensorFold's default draft head is DeepSeek's DSpark module, TensorFold/DeepSeek-V4-Flash-DSpark-MLX (10.67 GB).

What the conversion does

  • The layer's dense projections, e_proj and h_proj included, are FP8 (E4M3) with an E8M0 scale per 128 x 128 block. They are dequantized to bf16, which is exact, then quantized to MLX affine 4-bit in groups of 64. That step is lossy, and it is the format the mlx-community target stores for its own dense projections.
  • The routed experts keep DeepSeek's mxfp4 bytes and E8M0 scales unchanged.
  • Norms (enorm, hnorm and the block's own), attention sinks, router weights and biases and hyper-connection tensors are copied as stored.
  • Block mtp.0 becomes mtp, and its tensors take the names the mlx-community target uses for its own layers (attn.wq_a, ffn.switch_mlp.gate_proj and so on). config.json names the head with "model_type": "deepseek_v4_mtp".

TensorFold's converter reproduces model.safetensors byte for byte from the source:

python -m tensorfold.families.deepseek_v4.convert mtp model-00046-of-00046.safetensors DeepSeek-V4-Flash-MTP-MLX

Measured

On an M3 Ultra (60-core GPU, 256 GB) with MLX 0.32.2, TensorFold's lane engine served four short chat prompts (a Fibonacci function, how a GPU runs matrix multiplication, primes, a translation), 64 generated tokens each. Decode tok/s with this drafter and with "draft": false; sampled is temperature 1.0 with seed 1234:

Prompt Greedy, drafted Greedy, no drafts Sampled, drafted Sampled, no drafts
1 60.9 46.5 53.6 45.8
2 53.0 46.4 48.1 45.7
3 64.2 46.6 59.5 45.8
4 58.7 46.4 60.7 45.7

Each drafted reply equals its "draft": false reply. For comparison, mlx-lm PR #1797's server decodes 23.0-23.5 tok/s greedy and 28.7-29.4 sampled on the same machine and weights with its defaults, measured on a chat and code fixture rather than these four prompts.

Downloads last month
66
Safetensors
Model size
1B params
Tensor type
F32
路
BF16
路
U32
路
U8
路
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for TensorFold/DeepSeek-V4-Flash-MTP-MLX

Quantized
(127)
this model