Step-5-Preview
Imatrix quants best-in-class for ik_llama.cpp (Step-5 fork)
NOTE: ik_llama.cpp is highly versatile and runs mainline llama.cpp quants such as existing GGUFs from bartowski, unsloth, mradermacher, Aes Sedai etc
IQ3_KS-mix & IQ4_KS-mix released ✅ — more quants to come
General Specs
Value
ArchitectureStep-5 (MoE, 352 experts × 8 active + 1 shared, sparse GQA indexer)
Params600.1B total / 27B active
Layers92 (3 dense + 89 MoE) + 3 MTP
Attention64 q heads / 4 kv, head_dim 192; sliding (512) × 3 → full, × 92
Context1,048,576
Vocab128,896
Imatrixstep-preview-A.imatrix, 400K tokens @ ctx 8192, ubergarm-v02 corpus, sparse-enabled runtime
Imatrix
Value
Filestep-preview-A.imatrix (713 MB, 977/977 tensor entries)
Corpusubergarm-imatrix-calibration-corpus-v02
Tokens400K (50 chunks x 8192 tokens)
Context8192
Runtimeik_llama.cpp with sparse indexer active (see flags below)
Coverage977 entries; gaps are benign (gate_shexp kept at iq6_k, indexer.k q8_0 anyway, MTP/embeddings at standard tiers)
The imatrix was collected with the sparse selection ACTIVE at deployment context, so activation statistics match how the model actually runs. Dense-mode calibration at ctx 8192 would measure the wrong distribution: with top-k 512 blocks the sparse path discards most keys at that length.
Download: step-preview-A.imatrix — reusable for any tier quant from the BF16 master.
About our inference code
The sparse GQA indexer implementation (SSMax, CSA block compression, sparse selection) ships with ik_llama.cpp and is being open-sourced. Everything was derived from the model's sparse_config metadata and validated empirically: exact needle retrieval, PPL A/B tests, and logit-equivalence gates against the dense baseline. The knobs below make every candidate switchable without re-quantizing: the quants themselves stay valid.
Seriously experimental: this is a from-scratch implementation, not a port of reference code. Multiple GPU configurations are untested, including -sm tensor and -sm graph (split modes). Known reports: -sm graph segfault and layer-parallel token duplication under dense multi-GPU configs (fixes pending; single-GPU and default split mode unaffected). The Quick Start command uses the default split mode, which is the only one exercised. Expect rough edges and please report what you find.
FlagMeaning
IK_SPARSE=1Activates the sparse indexer on full-attention layers. Selection masks the attention instead of attending to every cached key. Default 0 runs dense (telemetry only).
IK_CSA=2Block-compressed indexer keys via the learned z projection (csa_z_norm_type=none in sparse_config). Mode 1 pools cached keys by averaging; mode 0 is token-grain. The z-proj variant measured better PPL (-8.6% at 8k ctx) and exact needle retrieval.
IK_SSMAX=1Per-query-head logit scaling (ssmax_s weights, ln(n) factor). PPL-neutral at deployment context and part of the config that passed needle retrieval exactly.
IK_SEL_SOFT=0.5Soft-gate penalty (attention-logit units) applied to keys outside the indexer's kept set. Set this whenever IK_SPARSE=1 — default 0 is the legacy hard cliff (strict top-k masking), which starves long-context recall. 0.5 is the validated setting; larger values approach dense-equivalent attention.
IK_SEL_KEEP_PCTOptional kept-slot scaling override (default 0 = fixed 4096-slot floor, proven sufficient with the soft gate; 50 = keep half of context; 100 = effectively dense selection). Experimentation knob only.
✅ Fixed — sparse long-context recall (update): early builds hard-masked every key past the kept-slot budget (4096 slots); beyond roughly 4–8k context that cliff starved verbatim recall — facts and paths came back as fabricated replacements while short tokens survived. Root cause was the hard rank cliff in the selection mask. The fix is a constant soft gate (IK_SEL_SOFT=0.5): keys outside the kept set keep a small attention-logit penalty instead of −∞, so strongly-scoring keys the indexer under-ranked still attend. Validated prod-parity at IQ3-tier stress mix (worst case): needle recall 18/18 at every depth 4.4k–44k, fresh AND multi-turn session, fabrications 0, n=3. Requires building current source (commit dd4359b9 or newer — the Quick Start clone above already includes it). Update (commit 81ac5fc7): the indexer now runs fused (selection scored and masked in-kernel) — behavior validated identical to the previous path, but compute-buffer memory drops ~70% at prod settings (≈2.2 GiB vs ≈7.4 GiB at 16k ctx), so long contexts fit in far less VRAM (launch-checked at 180k ctx). The old workaround IK_SEL_KEEP_PCT=100 is obsolete; kept-slot scaling remains only as an override knob (default 0 = fixed 4096-slot floor, proven sufficient with the soft gate). Honest boundaries: recall validated up to 44k context (larger contexts launch — checked to 180k — but recall beyond 44k is untested, and the 1M max is untested); sparse selection is still mask-only — it shapes which keys attend but does not cut attention compute (expect a small decode tax vs dense, unchanged by the fused rewrite, until KV-row gather kernels land).
All flags are read once at process start and stay fixed for the run. The Quick Start exports the exact combination the quant was calibrated and validated with (IK_SEL_SOFT=0.5 included).
PPL Analysis
TODO: PPL / KLD evaluation vs base reference pending.
Quant Collection
✅ IQ4_KS-mix  310.1 GiB · 36 shards
TODO: PPL / KLD evaluation vs base reference pending.
MetricValue
PPL (quant)Pending
PPL (base ref)Pending
(PPL(Q)/PPL(base)) - 1Pending
KLDPending
Same top-pPending
Δp RMSPending
Quantization recipe
# Attention (all layers) blk\..*\.attn_q\.weight=q8_0 blk\..*\.attn_k\.weight=q8_0 blk\..*\.attn_v\.weight=q8_0 blk\..*\.attn_output\.weight=q8_0 blk\..*\.attn_gate\.weight=q8_0 # Indexer (full-attn layers) blk\..*\.indexer\.q\.weight=q8_0 blk\..*\.indexer\.k\.weight=q8_0 blk\..*\.indexer\.z\.weight=q8_0 blk\..*\.indexer\.w\.weight=q8_0 # Dense layers 0-2 + 91 blk\.(0|1|2|91)\.ffn_down\.weight=iq6_k blk\.(0|1|2|91)\.ffn_(gate|up)\.weight=iq5_ks blk\.(0|1|2|91)\.ffn_gate\.weight=iq5_ks # Shared experts blk\..*\.ffn_down_shexp\.weight=iq6_k blk\..*\.ffn_gate_shexp\.weight=iq6_k blk\..*\.ffn_up_shexp\.weight=iq6_k # Routed experts (uniform) # down one tier above up/gate blk\..*\.ffn_down_exps\.weight=iq4_k blk\..*\.ffn_(gate|up)_exps\.weight=iq4_ks # MTP / nextn tail # shared_head_head = draft lm_head # eh_proj = layer-input projector blk\..*\.nextn\.eh_proj\.weight=iq4_kss blk\..*\.nextn\.shared_head_head\.weight=iq5_ks # Embeddings / output token_embd\.weight=iq6_k output\.weight=iq6_k
✅ IQ3_KS-mix  256.3 GiB · 30 shards
TODO: PPL / KLD evaluation vs base reference pending.
MetricValue
PPL (quant)Pending
PPL (base ref)Pending
(PPL(Q)/PPL(base)) - 1Pending
KLDPending
Same top-pPending
Δp RMSPending
Quantization recipe
# Attention (all layers) blk\..*\.attn_q\.weight=q8_0 blk\..*\.attn_k\.weight=q8_0 blk\..*\.attn_v\.weight=q8_0 blk\..*\.attn_output\.weight=q8_0 blk\..*\.attn_gate\.weight=q8_0 # Indexer (full-attn layers) blk\..*\.indexer\.q\.weight=q8_0 blk\..*\.indexer\.k\.weight=q8_0 blk\..*\.indexer\.z\.weight=q8_0 blk\..*\.indexer\.w\.weight=q8_0 # Dense layers 0-2 + 91 blk\(0|1|2|91)\.ffn_down\.weight=iq6_k blk\(0|1|2|91)\.ffn_(gate|up)\.weight=iq5_ks # Shared experts blk\..*\.ffn_down_shexp\.weight=iq6_k blk\..*\.ffn_gate_shexp\.weight=iq5_ks blk\..*\.ffn_up_shexp\.weight=iq5_ks # Routed experts (uniform) # down one tier above up/gate blk\..*\.ffn_down_exps\.weight=iq4_ks blk\..*\.ffn_(gate|up)_exps\.weight=iq3_ks # MTP / nextn tail # shared_head_head = draft lm_head # eh_proj = layer-input projector blk\..*\.nextn\.eh_proj\.weight=iq4_kss blk\..*\.nextn\.shared_head_head\.weight=iq5_ks # Embeddings / output token_embd\.weight=iq6_k output\.weight=iq6_k
Making your own quants (conversion pipeline)
IMPORTANT: the convert_hf_to_gguf.py converter in ik_llama.cpp is BROKEN for this architecture — do not use it. The correct pipeline is:
# 1. Convert HF -> BF16 GGUF with MAINLINE llama.cpp's converter + the Step-5 patch
#    (patch + instructions by avar6: https://huggingface.co/avar6/Step5-Preview)
git clone https://github.com/ggerganov/llama.cpp
cd llama.cpp && git apply step5.patch   # patch from the avar6 repo above
python convert_hf_to_gguf.py /path/to/Step5-Preview-HF \
  --outfile step5-preview-bf16.gguf --outtype bf16
# 2. Indexer tensor surgery — the mainline converter output lacks the 184 sparse-indexer
#    tensors (1783 total after surgery). Insert them from the HF checkpoint:
#    script: https://huggingface.co/L-Alchemyst/Step-5-Preview-GGUF/blob/main/indexer_insert.py
python3 indexer_insert.py plan    # dry-run: parse + assert + print the bill
python3 indexer_insert.py apply   # perform the surgery (keeps a .bak)
python3 indexer_insert.py verify  # reparse + byte-compare vs HF
# (edit the CONFIG block at the top of the script first: HF checkpoint dir + GGUF shard path)
# 3. Imatrix + quantize with ik_llama.cpp (IQK types are Ik-only)
./build/bin/llama-imatrix -m step5-preview-bf16.gguf \
  -f calibration-corpus.txt -o step5.imatrix \
  --gpu-layers 12 --ctx-size 8192
./build/bin/llama-quantize --imatrix step5.imatrix \
  step5-preview-bf16.gguf step5-preview-IQ4_KS-mix.gguf IQ4_XS
Conversion credit: avar6 (mainline converter patch). The indexer surgery script (indexer_insert.py) is available in this repo's files — edit the CONFIG block at the top with your local paths before running.
Quick Start
Requires ik_llama.cpp (Step-5 fork). Standard llama.cpp will not run these quants.

# Clone and build git clone --branch step5-preview-dense https://github.com/NF09830453/ik_llama.cpp cd ik_llama.cpp cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_CUDA=ON cmake --build build --config Release -j $(nproc)

# Download (main model) pip install huggingface_hub hf download L-Alchemyst/Step-5-Preview-GGUF --repo-type model --include "IQ4_KS-mix/*" --local-dir ./step5-iq4ks

IQ3_KS-mix alternative: swap the include to IQ3_KS-mix/* and point -m at step5-preview-IQ3_KS-mix-00001-of-00030.gguf (30 shards) — same run flags work for both.

# Optional: vision module (mmproj), required for image input hf download L-Alchemyst/Step-5-Preview-GGUF --repo-type model --include "mmproj-step5-vision-q8_0.gguf" --local-dir ./step5-iq4ks

Hybrid CPU+GPU: adjust -ot to your RAM/VRAM split — -cram is another adjustable flag. The sparse indexer export below is the prod-validated config: IK_SEL_SOFT=0.5 is required alongside IK_SPARSE=1 (the flag's default 0 is the legacy strict top-k cliff that starves long-context recall). Dense alternative (IK_SPARSE=0) is slightly faster today — sparse attention is mask-only until fused gather kernels land — but sparse matches the imatrix calibration regime.

export IK_SPARSE=1 IK_CSA=2 IK_SSMAX=1 IK_SEL_SOFT=0.5 ./build/bin/llama-server
-m ./step5-iq4ks/step5-preview-IQ4_KS-mix-00001-of-00036.gguf
--alias Alchemyst/Step-5
-ngl 999
-ot "blk.(2[1-9]|[3-8][0-9]|90).(ffn_up|ffn_gate|ffn_down)_exps=CPU"
-ub 4096 -b 4096
--flash-attn on
--ctx-checkpoints 5
--swa-compress
-c 80000
-cuda "offload-batch-size=128"
--spec-type mtp:n_max=1,p_min=0.0
-muge
-gr -ger
--merge-qkv
-ctk q8_0 -ctv q8_0
-khad -vhad
--jinja
-rtr
--threads 96 --threads-batch 128
--mmproj ./step5-iq4ks/mmproj-step5-vision-q8_0.gguf
--no-mmproj-offload
--port 8080

The --no-mmproj-offload flag keeps the vision module in system RAM instead of VRAM, trading a little image-processing speed for more GPU headroom.
CPU-only: use -ctk q8_0 and prefix with numactl -N ${SOCKET} -m ${SOCKET}.
Credits
Model: stepfun-ai
Runtime: ik_llama.cpp (Step-5 fork)
Downloads last month
1,917
GGUF
Model size
601B params
Architecture
step35
Hardware compatibility
Log In to add your hardware
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support