--- language: - en license: llama3.2 base_model: ZenithLLM/ZenAlta-1-3B-Phase2 tags: - llama - llama3.2 - speculative-decoding - draft-model - gguf - conversational - roleplay - mobile-llm pipeline_tag: text-generation --- # ⚡ Zen Alta 4-Layer Speculative Decoding Draft Model (~790M) **Zen Alta Draft** is a ultra-lightweight, 4-layer speculative decoding companion model engineered by **ZenithLLM**. Sliced from the top of the 24-layer **Zen Alta** architecture, it shares the exact same 128,256 BPE vocabulary and embedding/LM head, enabling **lossless 2× speculative inference acceleration** in `llama.cpp`, vLLM, and mobile runtimes. --- ## 🎯 What is Speculative Decoding? In standard autoregressive generation, deep models calculate every single token sequentially (e.g. 24 transformer layers per token). With **Zen Alta Draft**: 1. **The Fast Draft (4 Layers, 790M)**: Quickly guesses 4–5 candidate tokens in parallel in just ~20–30ms. 2. **The Target Model (Zen Alta 24 Layers, 2.8B)**: Verifies all proposed tokens in a single parallel forward pass (~40ms). 3. **Result**: Accepted tokens are committed simultaneously, achieving **40–50+ tokens/sec on mobile chips** with **0% degradation in output quality or persona**. --- ## 📦 Model Specifications | Parameter | Value | |---|---| | **Base Architecture** | Llama 3.2 (CausalLM) | | **Hidden Layers** | **4** (vs 24 in Target model) | | **Hidden Dimension** | 3072 | | **Intermediate Size** | 8192 | | **Attention Heads** | 24 query heads / 8 KV heads | | **Vocabulary Size** | 128,256 (Identical to Llama 3.2 & Zen Alta) | | **Context Length** | 131,072 tokens | | **RoPE Theta** | 500,000.0 | --- ## 📂 Repository Contents This consolidated repository contains both the raw PyTorch weights and the ready-to-run quantized GGUF: | File | Size | Description | |---|---|---| | `model.safetensors` | 1.52 GB | Unquantized FP16 PyTorch weights (4 layers) | | `zen-alta-draft-q4_k_m.gguf` | 545.72 MB | Quantized 4-bit medium GGUF for `llama.cpp` & mobile | | `config.json` | < 1 KB | 4-layer model configuration | | `tokenizer.json` | 16.4 MB | Fast BPE tokenizer definition | | `chat_template.jinja`| < 4 KB | Llama 3.2 conversational chat template | --- ## 🚀 How to Run in llama.cpp ### Speculative Decoding (Target + Draft Pairing) Download the target model from [ZenithLLM/ZenAlta-1-3B-Phase2-GGUF](https://huggingface.co/ZenithLLM/ZenAlta-1-3B-Phase2-GGUF) and the draft model from this repo: ```bash # Speculative decoding command ./llama-cli \ -m ZenAlta-1-3B-Pruned.Q4_K_M.gguf \ -md zen-alta-draft-q4_k_m.gguf \ --draft-max 5 \ -p "<|start_header_id|>user<|end_header_id|>\n\nhey who are you?<|eot_id|><|start_header_id|>assistant<|end_header_id|>\n\n" \ -n 128 ``` ### Standalone Inference (Fast Preview) ```bash ./llama-cli -m zen-alta-draft-q4_k_m.gguf -p "what is up" -n 64 ``` --- ## 🔗 Related Models - **Target Model (LoRA Adapter)**: [ZenithLLM/ZenAlta-1-3B-Phase2](https://huggingface.co/ZenithLLM/ZenAlta-1-3B-Phase2) - **Target Model (GGUF Q4_K_M)**: [ZenithLLM/ZenAlta-1-3B-Phase2-GGUF](https://huggingface.co/ZenithLLM/ZenAlta-1-3B-Phase2-GGUF) - **Base Pruned Model (24-Layer)**: [ZenithLLM/ZenAlta-1-3B-Pruned](https://huggingface.co/ZenithLLM/ZenAlta-1-3B-Pruned)