Instructions to use 0xSero/Step-5-Preview-Spark with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Trellis
How to use 0xSero/Step-5-Preview-Spark with Trellis:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Step-5-Preview-Spark
StepFun Step-5-Preview with every routed expert and every dense body linear re-encoded as EXL3 trellis
(MUL1 codebook), sized to serve on four NVIDIA DGX Spark (GB10, 4 x 128 GB unified memory, tensor
parallel 4). Every routed expert is K4 (4.01 bpw); the
attention / shared-expert / dense-MLP body is EXL3 K8. A BF16 copy of the body is included so prefill-sized batches
run as dense GEMMs while decode uses the EXL3 K8 GEMV ("hybrid" body). Activations run in fp16 (--dtype float16),
matching the reference implementation's precision; see Quality. The MTP draft layers, routers, norms,
embeddings, head and the vision tower are the upstream BF16 tensors.
EXL3 experts and body need the
step5_sparkvLLM plugin shipped in the companion image (below). Stock vLLM and plaintransformerscannot run this checkpoint.
At a glance
| Base | SHSLab/Step-5-Preview-BF16 @ 361580e06231fd3d3ad9c72a38fd8957aa776b52 |
| Architecture | ~600B total / ~27B active. 92 layers, hidden 4096; MoE layers 3-90 with 352 routed experts (top-8, sigmoid router) + 1 shared expert; sliding-window attention (512) with full attention every 4th layer; 3 MTP predictor layers; perception-encoder vision tower |
| Routed experts | EXL3 MUL1: K4 for all 88 MoE layers (3-90); 4.01 bpw |
| Body linears (q/k/v/o, gate, shared expert, dense MLP) | EXL3 K8 (decode) + BF16 copy (prefill) |
| Everything else | upstream BF16 tensors |
| Size | ~342 GB (EXL3 experts 293.4 GB, EXL3 K8 body ~12.6 GB, BF16 tensors incl. BF16 body copy 35.8 GB) |
| Context | 262,144 |
| Inputs | text, image, video (as sampled frames through the image path). No audio path exists in the model |
| Tools / reasoning | yes / yes (step3p5 parsers) |
| MTP speculative decoding | yes, 2 draft tokens |
| License | StepFun Community License, carried from the base |
Quality
Offline full-vocabulary token-wise KL divergence against the native BF16 model (teacher) on a 64-window x 2,048-token panel that was never used for calibration, with window-block bootstrap 95 % CIs. Lower KLD and higher top-1 agreement are better.
How the quantized model was scored: the served runtime (vLLM, TP4, this plugin, fp16) with the EXL3 K8 body on every token (the decode-path body; production runs prefill-sized batches through the BF16 body copy instead), in eager mode, without MTP and without prefix caching, capturing the full-vocabulary log-probs of the prompt (prefill) positions. This is the closest offline match to the serving path, not the serving path itself: CUDA graphs, MTP drafting and cached decode steps were not part of the scored run.
Cross-check on the serving path: the same checkpoint served on a different machine (4 x RTX PRO 6000) with the release serving settings (hybrid body, CUDA graphs, MTP 2, FP16 KV) scored on the same 64-window panel gives KLD 0.0716 [0.0652, 0.0788] and top-1 93.34 % [92.79, 93.89], consistent with the number above (overlapping CIs; different machine and capture path, so not a measurement of the decode path itself). In that run the prompt positions pass through the BF16 prefill body (that is what the serving path does for prefill-sized batches), and a separate check of decode steps against prefill on the same text agreed on the top-1 token 98.4 % of the time (top-20-truncated comparison, not a full-vocabulary teacher KL).
Panel reuse: windows 0-7 of this panel were also used for the 8-window comparisons that chose between candidate configurations (late-layer expert bits, body precision, the activation-precision fix). The 56 windows never used for any decision (8-63) give KLD 0.0737 [0.0671, 0.0808] and top-1 93.38 % [92.82, 93.93], the same as the full panel within noise.
| Build | Mean KLD (nats) [95 % CI] | Median KLD | Top-1 agreement [95 % CI] | ΔNLL vs teacher (nats) |
|---|---|---|---|---|
| this release (4.01 bpw experts, EXL3 K8 body on every token = the served decode path, fp16) | 0.0735 [0.0674, 0.0800] | 0.0029 | 0.933 [0.928, 0.938] | +0.013 [0.008, 0.018] |
| 3.40 bpw experts (layers 37-90 K3), hybrid body, fp16 | 0.142 [0.131, 0.155] | 0.0065 | 0.905 [0.897, 0.913] | +0.046 |
| 3.28 bpw experts (layers 3-26 K4), hybrid body, bf16 activations | 0.151 [0.139, 0.163] | 0.0073 | 0.902 [0.895, 0.910] | +0.052 |
| 3.01 bpw experts (all K3), hybrid body, bf16 activations | 0.156 [0.143, 0.170] | — | 0.901 [0.892, 0.908] | +0.051 |
| 3.01 bpw experts, EXL3 body everywhere, bf16 activations | 0.191 [0.176, 0.206] | 0.0104 | 0.889 [0.880, 0.897] | +0.075 |
Noise floor: this teacher is very sensitive. Re-running the reference model with unchanged weights and only a different floating-point summation order inside the MoE layers already gives KLD 0.027 / top-1 96.1 % (8 windows), because a 352-expert, top-8 router flips near-tied expert choices under tiny perturbations and the flips compound through depth. Quantization is measured on top of that floor. Our target was top-1 >= 93 % and mean KLD ~0.07 (no formal tolerance was set on the KLD target). The point estimates meet it (top-1 93.30 %, KLD 0.0735), but the 95 % CIs are wide enough to include values on both sides: top-1 [92.75, 93.84] %, KLD [0.067, 0.080].
What mattered: (1) running activations in fp16 instead of bf16 (the reference precision; bf16 rounding of the first norm output changed the selected top-8 expert set for ~30 % of tokens in every MoE layer in an 8-layer comparison, fp16 5-10 %); (2) expert bits on the late layers: with everything else fixed (same reference runtime, 8 windows), moving layers 37-90 from K3 to K4 took KLD from 0.133 to 0.084; (3) the body and scoring path: the release replaces the K4 body with a K8 EXL3 body and is scored in the served runtime, which accounts for the rest of the change (0.142 for the 3.40 bpw build to 0.0735). The step from 0.142 to 0.0735 is therefore expert bits plus body/path changes, not expert bits alone. Calibration-recipe changes (more data, router-weighted Hessians) made no measurable difference at K4.
Serving on 4 x DGX Spark
Measured with the release image recipe on four Sparks, TP4 over the Sparks' 200 GbE RoCE fabric, FP16 activations and KV cache, MTP speculative decoding (2 draft tokens), CUDA graphs, up to 4 concurrent sequences:
| KV cache | 999,279 tokens (BF16/FP16 KV, 20 GB per GPU) |
| Max context | 262,144 |
| Prefill (8k) | release config over 4 server launches: cold first request 1,106-1,327 tok/s, warm requests 1,365-1,565 tok/s (5 of 8 warm runs >= 1,500) · 32k: 1,445 tok/s (one launch) |
| Code decode, 1 stream | 25.7-29.9 tok/s |
| Prose decode, 1 stream | 26.2-29.5 tok/s |
| MTP acceptance (per position) | 0.89 / 0.67 (mean accepted length 2.56) |
| Load time | ~11 min |
All decode runs used the model's default sampling and stopped naturally; no output caps. Smoke checks passed for text, tool calls, image and video input. 4 concurrent code streams: 66.5 tok/s aggregate (prose 40.3). Long-context recall (passphrase at 40 % depth of a 240,925-token prompt): pass.
Against our own speed target (75 tok/s single-stream decode, 1,500 tok/s prefill) this release does not meet the
decode target (26-30 tok/s single stream, well short of 75) and does not reliably meet the prefill target: 8k warm
requests range 1,365-1,565 tok/s across launches (cold first requests 1,106-1,327) and 32k measured 1,445. Decode speed
also varies between server launches of the same configuration (per-launch means ~75-81 vs ~90 ms/step observed) and
grows with context (81 ms/step at 1-2k generated tokens to ~102 at 20-30k). Per-token decode time is dominated by
the EXL3 expert kernels and the dense-body GEMV at TP4, plus the all-reduce between the four Sparks.
Run it
- Download on every Spark (same path on all four):
hf download 0xSero/Step-5-Preview-Spark --local-dir ~/models/Step-5-Preview-Spark
- Launch with the four-Spark launcher
(
scripts/launch.sh; it starts one container per Spark and serves the OpenAI-compatible API on the head at:8000). Key settings:
vllm serve /model --served-model-name step-5-preview-spark \
--tensor-parallel-size 4 --nnodes 4 --dtype float16 \
--max-model-len 262144 --max-num-seqs 4 --kv-cache-memory-bytes 20000000000 \
--speculative-config '{"method":"mtp","num_speculative_tokens":2}' \
--reasoning-parser step3p5 --tool-call-parser step3p5 --enable-auto-tool-choice
# env: STEP5_BODY_FORMAT=hybrid EXL3_INT8_GEMV=1 VLLM_USE_V2_MODEL_RUNNER=0 NCCL_PROTO=Simple
# VLLM_DISABLE_SHARED_EXPERTS_STREAM=1 VLLM_ENABLE_ROCE_ALLREDUCE=1
Files
| Path | What it is |
|---|---|
model-0000{1..3}.safetensors |
upstream BF16 tensors that are not re-encoded (embeddings, norms, routers, head, MTP layers, vision tower) |
exl3/experts/L03..L90.safetensors |
EXL3 routed-expert banks, one file per MoE layer |
exl3/body/*.safetensors |
EXL3 K8 body linears |
body-bf16-*.safetensors |
BF16 body linears used for prefill-sized batches |
config.json |
Step-5 config with quantization_config.quant_method = step5_exl3, body_format = hybrid |
tokenizer*.json, chat_template.jinja |
upstream, unchanged |
How it was made
- Calibration: routed inputs captured layer by layer from the BF16 model on a calibration panel disjoint from the evaluation panel. Each expert is encoded with its own routed-input Hessian plus a pooled prior, gate/up first, then down on the decoded intermediate.
- Calibration data: 640 windows x 2,048 tokens (1.31M tokens) for the late layers, 160-320 windows for the earlier builds; uniform K4 for all experts. A K3 / K4 split by single-layer sensitivity was tried first and lost 0.07 KLD.
- Dense body: K8 for decode (EXL3 GEMV), plus the BF16 body for prefill-sized batches.
- Body linears that share an input (q/k/v, gate/up) share one Hessian so they can run as one fused GEMV.
Provenance
| Source | SHSLab/Step-5-Preview-BF16 rev 361580e06231fd3d3ad9c72a38fd8957aa776b52 |
| Encoder / kernels | exllamav3 1.5.3 d3739fd393337b1ff4d6c2a342b12f0c87a9592f |
| Serving runtime | local-inference-lab vLLM r38 + step5_spark plugin |
| Image | ghcr.io/0xsero/step-5-preview-spark@sha256:4e28f849a414a483b44be50a09198d76cbc859096b354796dc25ee2159e05a5e (tag s1, GitHub-hosted arm64 build with signed provenance) |
License and credits
StepFun Community License, carried from StepFun's release; this repo adds no terms. Thanks to StepFun for the model, turboderp for ExLlamaV3 / EXL3 (MIT) — the trellis encoder and kernels used here — and Local Inference Lab and the vLLM community for the Spark serving stack. The same-runtime evaluation (decoded experts swapped into the reference model) and the uniform-K4 comparison that led to this allocation follow ShapleyMcg's TR3 quantization work, used with permission.
- Downloads last month
- -
Model tree for 0xSero/Step-5-Preview-Spark
Base model
SHSLab/Step-5-Preview-BF16