Instructions to use ZenithLLM/ZenAlta-Draft with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use ZenithLLM/ZenAlta-Draft with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf ZenithLLM/ZenAlta-Draft:Q4_K_M # Run inference directly in the terminal: llama cli -hf ZenithLLM/ZenAlta-Draft:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf ZenithLLM/ZenAlta-Draft:Q4_K_M # Run inference directly in the terminal: llama cli -hf ZenithLLM/ZenAlta-Draft:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf ZenithLLM/ZenAlta-Draft:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf ZenithLLM/ZenAlta-Draft:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf ZenithLLM/ZenAlta-Draft:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf ZenithLLM/ZenAlta-Draft:Q4_K_M
Use Docker
docker model run hf.co/ZenithLLM/ZenAlta-Draft:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use ZenithLLM/ZenAlta-Draft with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ZenithLLM/ZenAlta-Draft" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ZenithLLM/ZenAlta-Draft", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ZenithLLM/ZenAlta-Draft:Q4_K_M
- Ollama
How to use ZenithLLM/ZenAlta-Draft with Ollama:
ollama run hf.co/ZenithLLM/ZenAlta-Draft:Q4_K_M
- Unsloth Desktop
- Pi
How to use ZenithLLM/ZenAlta-Draft with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ZenithLLM/ZenAlta-Draft:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ZenithLLM/ZenAlta-Draft:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use ZenithLLM/ZenAlta-Draft with Docker Model Runner:
docker model run hf.co/ZenithLLM/ZenAlta-Draft:Q4_K_M
- Lemonade
How to use ZenithLLM/ZenAlta-Draft with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull ZenithLLM/ZenAlta-Draft:Q4_K_M
Run and chat with the model
lemonade run user.ZenAlta-Draft-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use ZenithLLM/ZenAlta-Draft with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ZenithLLM/ZenAlta-Draft:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ZenithLLM/ZenAlta-Draft:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use ZenithLLM/ZenAlta-Draft with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ZenithLLM/ZenAlta-Draft:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ZenithLLM/ZenAlta-Draft:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Download README.md from ZenithLLM/ZenAlta-Draft: direct link, hf CLI and curl.
- Browser
- Download file 3.3 kB
-
https://huggingface.co/ZenithLLM/ZenAlta-Draft/resolve/main/README.md
- Command line
-
hf download hf://ZenithLLM/ZenAlta-Draft/README.md
-
curl -L -o README.md https://huggingface.co/ZenithLLM/ZenAlta-Draft/resolve/main/README.md
language:
- en
license: llama3.2
base_model: ZenithLLM/ZenAlta-1-3B-Phase2
tags:
- llama
- llama3.2
- speculative-decoding
- draft-model
- gguf
- conversational
- roleplay
- mobile-llm
pipeline_tag: text-generation
β‘ Zen Alta 4-Layer Speculative Decoding Draft Model (~790M)
Zen Alta Draft is a ultra-lightweight, 4-layer speculative decoding companion model engineered by ZenithLLM. Sliced from the top of the 24-layer Zen Alta architecture, it shares the exact same 128,256 BPE vocabulary and embedding/LM head, enabling lossless 2Γ speculative inference acceleration in llama.cpp, vLLM, and mobile runtimes.
π― What is Speculative Decoding?
In standard autoregressive generation, deep models calculate every single token sequentially (e.g. 24 transformer layers per token).
With Zen Alta Draft:
- The Fast Draft (4 Layers, 790M): Quickly guesses 4β5 candidate tokens in parallel in just ~20β30ms.
- The Target Model (Zen Alta 24 Layers, 2.8B): Verifies all proposed tokens in a single parallel forward pass (~40ms).
- Result: Accepted tokens are committed simultaneously, achieving 40β50+ tokens/sec on mobile chips with 0% degradation in output quality or persona.
π¦ Model Specifications
| Parameter | Value |
|---|---|
| Base Architecture | Llama 3.2 (CausalLM) |
| Hidden Layers | 4 (vs 24 in Target model) |
| Hidden Dimension | 3072 |
| Intermediate Size | 8192 |
| Attention Heads | 24 query heads / 8 KV heads |
| Vocabulary Size | 128,256 (Identical to Llama 3.2 & Zen Alta) |
| Context Length | 131,072 tokens |
| RoPE Theta | 500,000.0 |
π Repository Contents
This consolidated repository contains both the raw PyTorch weights and the ready-to-run quantized GGUF:
| File | Size | Description |
|---|---|---|
model.safetensors |
1.52 GB | Unquantized FP16 PyTorch weights (4 layers) |
zen-alta-draft-q4_k_m.gguf |
545.72 MB | Quantized 4-bit medium GGUF for llama.cpp & mobile |
config.json |
< 1 KB | 4-layer model configuration |
tokenizer.json |
16.4 MB | Fast BPE tokenizer definition |
chat_template.jinja |
< 4 KB | Llama 3.2 conversational chat template |
π How to Run in llama.cpp
Speculative Decoding (Target + Draft Pairing)
Download the target model from ZenithLLM/ZenAlta-1-3B-Phase2-GGUF and the draft model from this repo:
# Speculative decoding command
./llama-cli \
-m ZenAlta-1-3B-Pruned.Q4_K_M.gguf \
-md zen-alta-draft-q4_k_m.gguf \
--draft-max 5 \
-p "<|start_header_id|>user<|end_header_id|>\n\nhey who are you?<|eot_id|><|start_header_id|>assistant<|end_header_id|>\n\n" \
-n 128
Standalone Inference (Fast Preview)
./llama-cli -m zen-alta-draft-q4_k_m.gguf -p "what is up" -n 64
π Related Models
- Target Model (LoRA Adapter): ZenithLLM/ZenAlta-1-3B-Phase2
- Target Model (GGUF Q4_K_M): ZenithLLM/ZenAlta-1-3B-Phase2-GGUF
- Base Pruned Model (24-Layer): ZenithLLM/ZenAlta-1-3B-Pruned