Instructions to use TensorFold/DeepSeek-V4-Flash-MTP-MLX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use TensorFold/DeepSeek-V4-Flash-MTP-MLX with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir DeepSeek-V4-Flash-MTP-MLX TensorFold/DeepSeek-V4-Flash-MTP-MLX
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
DeepSeek-V4-Flash MTP for MLX
This is DeepSeek-V4-Flash's own multi-token-prediction (MTP) layer, converted for Apple Silicon. The mlx-community 4-bit conversion leaves it out. It drafts tokens for mlx-community/DeepSeek-V4-Flash-4bit on TensorFold's lane engine, which verifies every draft against the target.
- Source: deepseek-ai/DeepSeek-V4-Flash at revision
60d8d70770c6776ff598c94bb586a859a38244f1. The layer is themtp.0.*tensors ofmodel-00046-of-00046.safetensors. - License: MIT, DeepSeek's (see
LICENSE, copyright 2023 DeepSeek). The conversion adds no terms. - Files:
model.safetensors(3.52 GB),config.json,LICENSE,SHA256SUMS.
Use
TensorFold 0.3.6.4 or later on a Mac with 256 GB:
tensorfold pull mlx-community/DeepSeek-V4-Flash-4bit TensorFold/DeepSeek-V4-Flash-MTP-MLX
tensorfold serve mlx-community/DeepSeek-V4-Flash-4bit --drafter TensorFold/DeepSeek-V4-Flash-MTP-MLX
The layer drafts up to three tokens a round, one pass each, from the target's streams at the last verified row. A
drafted reply equals the same request sent with "draft": false. Drafts change speed, never tokens. TensorFold's
default draft head is DeepSeek's DSpark module,
TensorFold/DeepSeek-V4-Flash-DSpark-MLX (10.67 GB).
What the conversion does
- The layer's dense projections,
e_projandh_projincluded, are FP8 (E4M3) with an E8M0 scale per 128 x 128 block. They are dequantized to bf16, which is exact, then quantized to MLX affine 4-bit in groups of 64. That step is lossy, and it is the format the mlx-community target stores for its own dense projections. - The routed experts keep DeepSeek's mxfp4 bytes and E8M0 scales unchanged.
- Norms (
enorm,hnormand the block's own), attention sinks, router weights and biases and hyper-connection tensors are copied as stored. - Block
mtp.0becomesmtp, and its tensors take the names the mlx-community target uses for its own layers (attn.wq_a,ffn.switch_mlp.gate_projand so on).config.jsonnames the head with"model_type": "deepseek_v4_mtp".
TensorFold's converter reproduces model.safetensors byte for byte from the source:
python -m tensorfold.families.deepseek_v4.convert mtp model-00046-of-00046.safetensors DeepSeek-V4-Flash-MTP-MLX
Measured
On an M3 Ultra (60-core GPU, 256 GB) with MLX 0.32.2, TensorFold's lane engine served four short chat prompts (a
Fibonacci function, how a GPU runs matrix multiplication, primes, a translation), 64 generated tokens each. Decode
tok/s with this drafter and with "draft": false; sampled is temperature 1.0 with seed 1234:
| Prompt | Greedy, drafted | Greedy, no drafts | Sampled, drafted | Sampled, no drafts |
|---|---|---|---|---|
| 1 | 60.9 | 46.5 | 53.6 | 45.8 |
| 2 | 53.0 | 46.4 | 48.1 | 45.7 |
| 3 | 64.2 | 46.6 | 59.5 | 45.8 |
| 4 | 58.7 | 46.4 | 60.7 | 45.7 |
Each drafted reply equals its "draft": false reply. For comparison, mlx-lm PR #1797's server decodes 23.0-23.5
tok/s greedy and 28.7-29.4 sampled on the same machine and weights with its defaults, measured on a chat and code
fixture rather than these four prompts.
- Downloads last month
- 66
Quantized
Model tree for TensorFold/DeepSeek-V4-Flash-MTP-MLX
Base model
deepseek-ai/DeepSeek-V4-Flash