Run DeepSeek-V4-Flash-0731 with TyloQuant MFQ
TyloQuant MFQ provides neuron-anchored mixed-format quantization and high-fidelity inference for large language models.
- This repository contains MFQ-quantized weights derived from the official DeepSeek-V4-Flash-0731 release.
- MFQ combines NINT, NVQ/NPQ and NEPQ formats with expert-wise precision allocation for MoE models.
- The released 77.519 GiB S tier records 0.313576 Mean KLD and 82.2913% same-top on the complete official-0731 WikiText-2 evaluation.
- See the project documentation for runtime, format and deployment instructions.
Full WikiText-2 evaluation against official DeepSeek-V4-Flash-0731 BF16 logits: ctx=512, 573 chunks and 146,115 scored tokens. Lower Mean KLD is better. Complete results and protocol.
DeepSeek-V4-Flash-0731
Introduction
DeepSeek-V4-Flash-0731 is the official release of DeepSeek-V4-Flash, superseding the preview version, with substantially enhanced agentic capabilities. It has the same model structure as DeepSeek-V4-Flash-DSpark, i.e. it comes with a speculative decoding module attached.
DeepSeek-V4-Flash-0731 outperforms DeepSeek-V4-Pro (Preview) on benchmarks listed below despite its far smaller activated parameter count, and is broadly competitive with the strongest proprietary models available.
| Benchmark | DeepSeek-V4-Flash-0731 | DeepSeek-V4-Flash (Preview) | DeepSeek-V4-Pro (Preview) | GLM-5.2 | Opus-4.8 |
|---|---|---|---|---|---|
| Terminal Bench 2.1 | 82.7 | 61.8 | 72.1 | 81.0 | 85.0 |
| NL2Repo | 54.2 | 39.4 | 38.5 | 48.9 | 69.7 |
| Cybergym | 76.7 | 38.7 | 52.7 | - | 83.1 |
| DeepSWE | 54.4 | 7.3 | 12.8 | 46.2 | 58.0 |
| Toolathlon-Verified | 70.3 | 49.7 | 55.9 | 59.9 | 76.2 |
| Agents' Last Exam | 25.2 | 15.8 | 16.5 | 23.8 | 25.7 |
| AutomationBench Public | 25.1 | 10.8 | 12.8 | 12.9 | 27.2 |
| DSBench-FullStack † | 68.7 | 37.0 | 41.8 | 61.8 | 71.6 |
| DSBench-Hard † | 59.6 | 25.8 | 31.1 | 54.5 | 71.7 |
Notes:
- For the Code Agent tasks among the public benchmarks above, DeepSeek-V4-Flash-0731 is evaluated with the minimal mode of DeepSeek Harness (to be released) as the agent framework, using the
maxreasoning effort level withtemperature = 1.0, top_p = 0.95. - † DSBench-FullStack is an internal full-stack development test set; DSBench-Hard is an internal test set of difficult coding-agent problems.
License
This repository and the model weights are licensed under the MIT License.
Citation
@misc{deepseekai2026deepseekv4,
title={DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence},
author={DeepSeek-AI},
year={2026},
}
Model tree for Tylogi/DeepSeek-V4-Flash-0731-EW-MFQ
Base model
deepseek-ai/DeepSeek-V4-Flash-0731