File size: 3,300 Bytes
b2298c7
7136a63
 
 
 
 
 
 
 
 
 
 
 
 
 
b2298c7
 
7136a63
b2298c7
7136a63
b2298c7
7136a63
b2298c7
7136a63
b2298c7
7136a63
b2298c7
7136a63
 
 
 
b2298c7
7136a63
b2298c7
7136a63
b2298c7
7136a63
 
 
 
 
 
 
 
 
 
b2298c7
7136a63
b2298c7
7136a63
b2298c7
7136a63
b2298c7
7136a63
 
 
 
 
 
 
b2298c7
7136a63
b2298c7
7136a63
b2298c7
7136a63
b2298c7
7136a63
b2298c7
7136a63
 
 
 
 
 
 
 
 
b2298c7
7136a63
 
 
 
b2298c7
7136a63
b2298c7
7136a63
b2298c7
7136a63
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
---
language:
- en
license: llama3.2
base_model: ZenithLLM/ZenAlta-1-3B-Phase2
tags:
- llama
- llama3.2
- speculative-decoding
- draft-model
- gguf
- conversational
- roleplay
- mobile-llm
pipeline_tag: text-generation
---

# ⚑ Zen Alta 4-Layer Speculative Decoding Draft Model (~790M)

**Zen Alta Draft** is a ultra-lightweight, 4-layer speculative decoding companion model engineered by **ZenithLLM**. Sliced from the top of the 24-layer **Zen Alta** architecture, it shares the exact same 128,256 BPE vocabulary and embedding/LM head, enabling **lossless 2Γ— speculative inference acceleration** in `llama.cpp`, vLLM, and mobile runtimes.

---

## 🎯 What is Speculative Decoding?

In standard autoregressive generation, deep models calculate every single token sequentially (e.g. 24 transformer layers per token). 

With **Zen Alta Draft**:
1. **The Fast Draft (4 Layers, 790M)**: Quickly guesses 4–5 candidate tokens in parallel in just ~20–30ms.
2. **The Target Model (Zen Alta 24 Layers, 2.8B)**: Verifies all proposed tokens in a single parallel forward pass (~40ms).
3. **Result**: Accepted tokens are committed simultaneously, achieving **40–50+ tokens/sec on mobile chips** with **0% degradation in output quality or persona**.

---

## πŸ“¦ Model Specifications

| Parameter | Value |
|---|---|
| **Base Architecture** | Llama 3.2 (CausalLM) |
| **Hidden Layers** | **4** (vs 24 in Target model) |
| **Hidden Dimension** | 3072 |
| **Intermediate Size** | 8192 |
| **Attention Heads** | 24 query heads / 8 KV heads |
| **Vocabulary Size** | 128,256 (Identical to Llama 3.2 & Zen Alta) |
| **Context Length** | 131,072 tokens |
| **RoPE Theta** | 500,000.0 |

---

## πŸ“‚ Repository Contents

This consolidated repository contains both the raw PyTorch weights and the ready-to-run quantized GGUF:

| File | Size | Description |
|---|---|---|
| `model.safetensors` | 1.52 GB | Unquantized FP16 PyTorch weights (4 layers) |
| `zen-alta-draft-q4_k_m.gguf` | 545.72 MB | Quantized 4-bit medium GGUF for `llama.cpp` & mobile |
| `config.json` | < 1 KB | 4-layer model configuration |
| `tokenizer.json` | 16.4 MB | Fast BPE tokenizer definition |
| `chat_template.jinja`| < 4 KB | Llama 3.2 conversational chat template |

---

## πŸš€ How to Run in llama.cpp

### Speculative Decoding (Target + Draft Pairing)

Download the target model from [ZenithLLM/ZenAlta-1-3B-Phase2-GGUF](https://huggingface.co/ZenithLLM/ZenAlta-1-3B-Phase2-GGUF) and the draft model from this repo:

```bash
# Speculative decoding command
./llama-cli \
  -m ZenAlta-1-3B-Pruned.Q4_K_M.gguf \
  -md zen-alta-draft-q4_k_m.gguf \
  --draft-max 5 \
  -p "<|start_header_id|>user<|end_header_id|>\n\nhey who are you?<|eot_id|><|start_header_id|>assistant<|end_header_id|>\n\n" \
  -n 128
```

### Standalone Inference (Fast Preview)
```bash
./llama-cli -m zen-alta-draft-q4_k_m.gguf -p "what is up" -n 64
```

---

## πŸ”— Related Models

- **Target Model (LoRA Adapter)**: [ZenithLLM/ZenAlta-1-3B-Phase2](https://huggingface.co/ZenithLLM/ZenAlta-1-3B-Phase2)
- **Target Model (GGUF Q4_K_M)**: [ZenithLLM/ZenAlta-1-3B-Phase2-GGUF](https://huggingface.co/ZenithLLM/ZenAlta-1-3B-Phase2-GGUF)
- **Base Pruned Model (24-Layer)**: [ZenithLLM/ZenAlta-1-3B-Pruned](https://huggingface.co/ZenithLLM/ZenAlta-1-3B-Pruned)