ZenAlta-Draft / README.md
ZenithLLM's picture
Update comprehensive Model Card with speculative decoding specs
7136a63 verified
|
Raw History Blame Contribute Delete
3.3 kB
---
language:
- en
license: llama3.2
base_model: ZenithLLM/ZenAlta-1-3B-Phase2
tags:
- llama
- llama3.2
- speculative-decoding
- draft-model
- gguf
- conversational
- roleplay
- mobile-llm
pipeline_tag: text-generation
---
# ⚑ Zen Alta 4-Layer Speculative Decoding Draft Model (~790M)
**Zen Alta Draft** is a ultra-lightweight, 4-layer speculative decoding companion model engineered by **ZenithLLM**. Sliced from the top of the 24-layer **Zen Alta** architecture, it shares the exact same 128,256 BPE vocabulary and embedding/LM head, enabling **lossless 2Γ— speculative inference acceleration** in `llama.cpp`, vLLM, and mobile runtimes.
---
## 🎯 What is Speculative Decoding?
In standard autoregressive generation, deep models calculate every single token sequentially (e.g. 24 transformer layers per token).
With **Zen Alta Draft**:
1. **The Fast Draft (4 Layers, 790M)**: Quickly guesses 4–5 candidate tokens in parallel in just ~20–30ms.
2. **The Target Model (Zen Alta 24 Layers, 2.8B)**: Verifies all proposed tokens in a single parallel forward pass (~40ms).
3. **Result**: Accepted tokens are committed simultaneously, achieving **40–50+ tokens/sec on mobile chips** with **0% degradation in output quality or persona**.
---
## πŸ“¦ Model Specifications
| Parameter | Value |
|---|---|
| **Base Architecture** | Llama 3.2 (CausalLM) |
| **Hidden Layers** | **4** (vs 24 in Target model) |
| **Hidden Dimension** | 3072 |
| **Intermediate Size** | 8192 |
| **Attention Heads** | 24 query heads / 8 KV heads |
| **Vocabulary Size** | 128,256 (Identical to Llama 3.2 & Zen Alta) |
| **Context Length** | 131,072 tokens |
| **RoPE Theta** | 500,000.0 |
---
## πŸ“‚ Repository Contents
This consolidated repository contains both the raw PyTorch weights and the ready-to-run quantized GGUF:
| File | Size | Description |
|---|---|---|
| `model.safetensors` | 1.52 GB | Unquantized FP16 PyTorch weights (4 layers) |
| `zen-alta-draft-q4_k_m.gguf` | 545.72 MB | Quantized 4-bit medium GGUF for `llama.cpp` & mobile |
| `config.json` | < 1 KB | 4-layer model configuration |
| `tokenizer.json` | 16.4 MB | Fast BPE tokenizer definition |
| `chat_template.jinja`| < 4 KB | Llama 3.2 conversational chat template |
---
## πŸš€ How to Run in llama.cpp
### Speculative Decoding (Target + Draft Pairing)
Download the target model from [ZenithLLM/ZenAlta-1-3B-Phase2-GGUF](https://huggingface.co/ZenithLLM/ZenAlta-1-3B-Phase2-GGUF) and the draft model from this repo:
```bash
# Speculative decoding command
./llama-cli \
-m ZenAlta-1-3B-Pruned.Q4_K_M.gguf \
-md zen-alta-draft-q4_k_m.gguf \
--draft-max 5 \
-p "<|start_header_id|>user<|end_header_id|>\n\nhey who are you?<|eot_id|><|start_header_id|>assistant<|end_header_id|>\n\n" \
-n 128
```
### Standalone Inference (Fast Preview)
```bash
./llama-cli -m zen-alta-draft-q4_k_m.gguf -p "what is up" -n 64
```
---
## πŸ”— Related Models
- **Target Model (LoRA Adapter)**: [ZenithLLM/ZenAlta-1-3B-Phase2](https://huggingface.co/ZenithLLM/ZenAlta-1-3B-Phase2)
- **Target Model (GGUF Q4_K_M)**: [ZenithLLM/ZenAlta-1-3B-Phase2-GGUF](https://huggingface.co/ZenithLLM/ZenAlta-1-3B-Phase2-GGUF)
- **Base Pruned Model (24-Layer)**: [ZenithLLM/ZenAlta-1-3B-Pruned](https://huggingface.co/ZenithLLM/ZenAlta-1-3B-Pruned)