Holo4-27B-NVFP4-MTP

Community packaging by RubyWitch of H Company's Holo4-27B-NVFP4, with the compatible BF16 multi-token prediction (MTP) head from Qwen/Qwen3.8-27B included for speculative decoding in vLLM.

The original Holo4 target weights are unchanged. This repository adds the donor MTP head and the metadata needed to load it. No additional training or requantization was performed. The donor head comes from Qwen, rather than from a Holo4-specific MTP fine-tune. This is an unofficial community derivative, not an H Company or Qwen release.

Provenance and changes

Component Source Pinned revision
NVFP4 target weights, tokenizer, processor and chat template Hcompany/Holo4-27B-NVFP4 6a02909667ba91dcd45e01be13c4d49992112510
BF16 MTP head Qwen/Qwen3.8-27B 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0

The upstream NVFP4 checkpoint declares one MTP layer but does not contain its weights. This package:

  • Preserves both original Holo4 weight shards byte for byte.
  • Adds model-mtp.safetensors, containing exactly 15 mtp.* BF16 tensors extracted from the donor's final shard. Tensor data occupies 849,398,784 bytes; the complete shard is 849,400,424 bytes.
  • Excludes the donor's separate lm_head.weight. Target and draft use Holo4's embeddings and output head.
  • Adds the MTP tensors to model.safetensors.index.json and updates its total size.
  • Adds re:^mtp\. to the compressed-tensors quantization ignore list, keeping the draft head in BF16.
  • Keeps the target's 64-layer count unchanged; the MTP layer is separate.
  • Preserves the tokenizer, processor, generation configuration, chat template and original quantization recipe.

crystal-mtp-provenance.json records immutable source identities, source file hashes, per-tensor hashes and hashes of the modified configuration, index and new shard. SHA256SUMS covers this package's files. README-Hcompany.md preserves the original model card, including its attribution and links.

Running with vLLM

Tested on 2026-10-06 with a native NVIDIA GB10 workstation and vLLM source revision 52358e6e192aeae73dd3764046c98fc4b156ea83. Use a build supporting this architecture, compressed-tensors NVFP4 and Qwen MTP; older vLLM versions may lack the required functionality or flags.

The following reproduces the tested 128K context, two concurrent requests, three speculative tokens profile. Its memory settings are specific to the GB10 workstation and must be adjusted for other hardware and workloads.

vllm serve RubyWitch/Holo4-27B-NVFP4-MTP \
  --served-model-name Holo4-27B-NVFP4-MTP \
  --host 127.0.0.1 --port 30000 \
  --dtype bfloat16 \
  --quantization compressed-tensors \
  --max-model-len 131072 \
  --max-num-seqs 2 \
  --gpu-memory-utilization 0.55 \
  --kv-cache-memory-bytes 30064771072 \
  --kv-cache-dtype bfloat16 \
  --mamba-ssm-cache-dtype float32 \
  --enable-prefix-caching \
  --prefix-cache-retention-interval 800 \
  --mamba-cache-mode align \
  --enable-chunked-prefill \
  --max-num-batched-tokens 8192 \
  --async-scheduling \
  --speculative-config '{"method":"mtp","num_speculative_tokens":3}' \
  --chat-template-content-format openai \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_coder \
  --reasoning-parser qwen3 \
  --limit-mm-per-prompt '{"image":5,"video":0}' \
  --mm-processor-cache-type shm \
  --mm-processor-cache-gb 4

The repository includes one MTP layer. num_speculative_tokens: 3 reuses that head for three drafting steps; it does not require three separate heads.

On the tested GB10 runtime, vLLM automatically selected Marlin for the NVFP4 W4A16 layers, retained its usual backend selection for other FP8 layers, and used FlashAttention 2 for BF16 attention. No global Marlin override, eager-mode fallback or runtime source patch was required. This does not establish the best kernel choice on other GPUs.

The explicit KV-cache allocation is 28 GiB, not the total memory requirement. With this flag, gpu_memory_utilization is not a 55% total-allocation cap in the tested runtime. Weights, activations, Mamba state, processor caches, compilation and other runtime allocations need additional memory. The 800-token prefix-retention interval matches the block size selected for this particular MTP depth and cache precision; review it if changing those settings. The model configuration retains the upstream context limit, but this package was qualified at 131,072 tokens.

Keep sufficient disk space available for CUDA/FlashInfer compilation scratch. On the test workstation, compilation initially exhausted a quota-limited /tmp; selecting a larger writable TMPDIR resolved that startup failure.

Example text request:

curl http://127.0.0.1:30000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "Holo4-27B-NVFP4-MTP",
    "messages": [{"role": "user", "content": "Explain speculative decoding briefly."}],
    "max_tokens": 256,
    "chat_template_kwargs": {"enable_thinking": false}
  }'

See H Company's documentation for function calling and element localization.

Local qualification

Functional checks passed for text generation, greedy and sampled image recognition, reasoning output, automatic tool calls and tool-result continuation, constrained-JSON grounding, and prefix-cache reuse.

Two simultaneous requests each used 130,944 input tokens plus 128 output tokens, recovered their distinct test codes, and completed with zero cache preemptions. Peak KV-cache use was 83.27% of the 28 GiB allocation. That stress test took about 391 seconds overall; it is not a 128K-context decode-speed benchmark.

Short-context measurements on the shared GB10 workstation used 512 generated tokens per prompt. Sequential decode rates exclude time to first token; the concurrent aggregate includes full request time.

Profile Code tokens/s Explanatory text tokens/s Screenshot response tokens/s Two-request aggregate tokens/s
No MTP, earlier automatic cache budget 9.56 9.64 9.62 18.77
MTP 3, final fixed 28 GiB cache 20.88 16.80 18.66 37.94

These are individual local runs with different cache budgets, not repeated, isolated performance trials or a general speed guarantee. Comparing three and five speculative tokens at the same earlier cache budget favored three for mixed use: five slightly improved code but reduced explanatory-text throughput. Acceptance and speed depend on prompts, sampling, context length, hardware and runtime version.

No independent quality benchmark of this composite has been performed. Upstream Holo4 benchmark claims belong to the upstream release and should not be treated as new evaluations of this package. The compatibility checks above establish serving functionality, not comprehensive accuracy equivalence across workloads.

License and attribution

Holo4 was developed by H Company and builds on Qwen3.8-27B from Alibaba Cloud / the Qwen team. The NVFP4 target weights and related upstream files are redistributed with H Company's CC BY-NC 4.0 license in LICENSE. The Qwen MTP tensors originate from the Apache-2.0 licensed Qwen model; its license is retained in LICENSE-APACHE.

This community packaging does not remove Holo4's noncommercial restriction. The modifications are listed above; original attribution and license texts are retained. Consult the upstream model card for the model's intended uses and limitations.

Downloads last month
-
Safetensors
Model size
19B params
Tensor type
BF16
·
F8_E4M3
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for RubyWitch/Holo4-27B-NVFP4-MTP

Base model

Qwen/Qwen3.8-27B
Quantized
(10)
this model