Instructions to use L-Alchemyst/Step-5-Preview-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use L-Alchemyst/Step-5-Preview-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf L-Alchemyst/Step-5-Preview-GGUF:BF16 # Run inference directly in the terminal: llama cli -hf L-Alchemyst/Step-5-Preview-GGUF:BF16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf L-Alchemyst/Step-5-Preview-GGUF:BF16 # Run inference directly in the terminal: llama cli -hf L-Alchemyst/Step-5-Preview-GGUF:BF16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf L-Alchemyst/Step-5-Preview-GGUF:BF16 # Run inference directly in the terminal: ./llama-cli -hf L-Alchemyst/Step-5-Preview-GGUF:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf L-Alchemyst/Step-5-Preview-GGUF:BF16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf L-Alchemyst/Step-5-Preview-GGUF:BF16
Use Docker
docker model run hf.co/L-Alchemyst/Step-5-Preview-GGUF:BF16
- LM Studio
- Jan
- Ollama
How to use L-Alchemyst/Step-5-Preview-GGUF with Ollama:
ollama run hf.co/L-Alchemyst/Step-5-Preview-GGUF:BF16
- Unsloth Desktop
- Pi
How to use L-Alchemyst/Step-5-Preview-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf L-Alchemyst/Step-5-Preview-GGUF:BF16
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "L-Alchemyst/Step-5-Preview-GGUF:BF16" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use L-Alchemyst/Step-5-Preview-GGUF with Docker Model Runner:
docker model run hf.co/L-Alchemyst/Step-5-Preview-GGUF:BF16
- Lemonade
How to use L-Alchemyst/Step-5-Preview-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull L-Alchemyst/Step-5-Preview-GGUF:BF16
Run and chat with the model
lemonade run user.Step-5-Preview-GGUF-BF16
List all available models
lemonade list
- Hermes Agent
How to use L-Alchemyst/Step-5-Preview-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf L-Alchemyst/Step-5-Preview-GGUF:BF16
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default L-Alchemyst/Step-5-Preview-GGUF:BF16
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use L-Alchemyst/Step-5-Preview-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf L-Alchemyst/Step-5-Preview-GGUF:BF16
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "L-Alchemyst/Step-5-Preview-GGUF:BF16" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
| Value | |
|---|---|
| Architecture | Step-5 (MoE, 352 experts × 8 active + 1 shared, sparse GQA indexer) |
| Params | 600.1B total / 27B active |
| Layers | 92 (3 dense + 89 MoE) + 3 MTP |
| Attention | 64 q heads / 4 kv, head_dim 192; sliding (512) × 3 → full, × 92 |
| Context | 1,048,576 |
| Vocab | 128,896 |
| Imatrix | step-preview-A.imatrix, 400K tokens @ ctx 8192, ubergarm-v02 corpus, sparse-enabled runtime |
| Value | |
|---|---|
| File | step-preview-A.imatrix (713 MB, 977/977 tensor entries) |
| Corpus | ubergarm-imatrix-calibration-corpus-v02 |
| Tokens | 400K (50 chunks x 8192 tokens) |
| Context | 8192 |
| Runtime | ik_llama.cpp with sparse indexer active (see flags below) |
| Coverage | 977 entries; gaps are benign (gate_shexp kept at iq6_k, indexer.k q8_0 anyway, MTP/embeddings at standard tiers) |
-sm tensor and -sm graph (split modes). Known reports: -sm graph segfault and layer-parallel token duplication under dense multi-GPU configs (fixes pending; single-GPU and default split mode unaffected). The Quick Start command uses the default split mode, which is the only one exercised. Expect rough edges and please report what you find.| Flag | Meaning |
|---|---|
| IK_SPARSE=1 | Activates the sparse indexer on full-attention layers. Selection masks the attention instead of attending to every cached key. Default 0 runs dense (telemetry only). |
| IK_CSA=2 | Block-compressed indexer keys via the learned z projection (csa_z_norm_type=none in sparse_config). Mode 1 pools cached keys by averaging; mode 0 is token-grain. The z-proj variant measured better PPL (-8.6% at 8k ctx) and exact needle retrieval. |
| IK_SSMAX=1 | Per-query-head logit scaling (ssmax_s weights, ln(n) factor). PPL-neutral at deployment context and part of the config that passed needle retrieval exactly. |
| IK_SEL_SOFT=0.5 | Soft-gate penalty (attention-logit units) applied to keys outside the indexer's kept set. Set this whenever IK_SPARSE=1 — default 0 is the legacy hard cliff (strict top-k masking), which starves long-context recall. 0.5 is the validated setting; larger values approach dense-equivalent attention. |
| IK_SEL_KEEP_PCT | Optional kept-slot scaling override (default 0 = fixed 4096-slot floor, proven sufficient with the soft gate; 50 = keep half of context; 100 = effectively dense selection). Experimentation knob only. |
IK_SEL_SOFT=0.5): keys outside the kept set keep a small attention-logit penalty instead of −∞, so strongly-scoring keys the indexer under-ranked still attend. Validated prod-parity at IQ3-tier stress mix (worst case): needle recall 18/18 at every depth 4.4k–44k, fresh AND multi-turn session, fabrications 0, n=3. Requires building current source (commit dd4359b9 or newer — the Quick Start clone above already includes it). Update (commit 81ac5fc7): the indexer now runs fused (selection scored and masked in-kernel) — behavior validated identical to the previous path, but compute-buffer memory drops ~70% at prod settings (≈2.2 GiB vs ≈7.4 GiB at 16k ctx), so long contexts fit in far less VRAM (launch-checked at 180k ctx). The old workaround IK_SEL_KEEP_PCT=100 is obsolete; kept-slot scaling remains only as an override knob (default 0 = fixed 4096-slot floor, proven sufficient with the soft gate). Honest boundaries: recall validated up to 44k context (larger contexts launch — checked to 180k — but recall beyond 44k is untested, and the 1M max is untested); sparse selection is still mask-only — it shapes which keys attend but does not cut attention compute (expect a small decode tax vs dense, unchanged by the fused rewrite, until KV-row gather kernels land).IK_SEL_SOFT=0.5 included).| Metric | Value |
|---|---|
| PPL (quant) | Pending |
| PPL (base ref) | Pending |
| (PPL(Q)/PPL(base)) - 1 | Pending |
| KLD | Pending |
| Same top-p | Pending |
| Δp RMS | Pending |
Quantization recipe
| Metric | Value |
|---|---|
| PPL (quant) | Pending |
| PPL (base ref) | Pending |
| (PPL(Q)/PPL(base)) - 1 | Pending |
| KLD | Pending |
| Same top-p | Pending |
| Δp RMS | Pending |
Quantization recipe
convert_hf_to_gguf.py converter in ik_llama.cpp is BROKEN for this architecture — do not use it. The correct pipeline is:# 1. Convert HF -> BF16 GGUF with MAINLINE llama.cpp's converter + the Step-5 patch
# (patch + instructions by avar6: https://huggingface.co/avar6/Step5-Preview)
git clone https://github.com/ggerganov/llama.cpp
cd llama.cpp && git apply step5.patch # patch from the avar6 repo above
python convert_hf_to_gguf.py /path/to/Step5-Preview-HF \
--outfile step5-preview-bf16.gguf --outtype bf16
# 2. Indexer tensor surgery — the mainline converter output lacks the 184 sparse-indexer
# tensors (1783 total after surgery). Insert them from the HF checkpoint:
# script: https://huggingface.co/L-Alchemyst/Step-5-Preview-GGUF/blob/main/indexer_insert.py
python3 indexer_insert.py plan # dry-run: parse + assert + print the bill
python3 indexer_insert.py apply # perform the surgery (keeps a .bak)
python3 indexer_insert.py verify # reparse + byte-compare vs HF
# (edit the CONFIG block at the top of the script first: HF checkpoint dir + GGUF shard path)
# 3. Imatrix + quantize with ik_llama.cpp (IQK types are Ik-only)
./build/bin/llama-imatrix -m step5-preview-bf16.gguf \
-f calibration-corpus.txt -o step5.imatrix \
--gpu-layers 12 --ctx-size 8192
./build/bin/llama-quantize --imatrix step5.imatrix \
step5-preview-bf16.gguf step5-preview-IQ4_KS-mix.gguf IQ4_XS
indexer_insert.py) is available in this repo's files — edit the CONFIG block at the top with your local paths before running.# Clone and build git clone --branch step5-preview-dense https://github.com/NF09830453/ik_llama.cpp cd ik_llama.cpp cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_CUDA=ON cmake --build build --config Release -j $(nproc)
# Download (main model) pip install huggingface_hub hf download L-Alchemyst/Step-5-Preview-GGUF --repo-type model --include "IQ4_KS-mix/*" --local-dir ./step5-iq4ks
IQ3_KS-mix/* and point -m at step5-preview-IQ3_KS-mix-00001-of-00030.gguf (30 shards) — same run flags work for both.# Optional: vision module (mmproj), required for image input hf download L-Alchemyst/Step-5-Preview-GGUF --repo-type model --include "mmproj-step5-vision-q8_0.gguf" --local-dir ./step5-iq4ks
-ot to your RAM/VRAM split — -cram is another adjustable flag. The sparse indexer export below is the prod-validated config: IK_SEL_SOFT=0.5 is required alongside IK_SPARSE=1 (the flag's default 0 is the legacy strict top-k cliff that starves long-context recall). Dense alternative (IK_SPARSE=0) is slightly faster today — sparse attention is mask-only until fused gather kernels land — but sparse matches the imatrix calibration regime.export IK_SPARSE=1 IK_CSA=2 IK_SSMAX=1 IK_SEL_SOFT=0.5
./build/bin/llama-server
-m ./step5-iq4ks/step5-preview-IQ4_KS-mix-00001-of-00036.gguf
--alias Alchemyst/Step-5
-ngl 999
-ot "blk.(2[1-9]|[3-8][0-9]|90).(ffn_up|ffn_gate|ffn_down)_exps=CPU"
-ub 4096 -b 4096
--flash-attn on
--ctx-checkpoints 5
--swa-compress
-c 80000
-cuda "offload-batch-size=128"
--spec-type mtp:n_max=1,p_min=0.0
-muge
-gr -ger
--merge-qkv
-ctk q8_0 -ctv q8_0
-khad -vhad
--jinja
-rtr
--threads 96 --threads-batch 128
--mmproj ./step5-iq4ks/mmproj-step5-vision-q8_0.gguf
--no-mmproj-offload
--port 8080
--no-mmproj-offload flag keeps the vision module in system RAM instead of VRAM, trading a little image-processing speed for more GPU headroom.-ctk q8_0 and prefix with numactl -N ${SOCKET} -m ${SOCKET}.- Downloads last month
- 1,917