ZenAlta-Draft / README.md
ZenithLLM's picture
Update comprehensive Model Card with speculative decoding specs
7136a63 verified
|
Raw History Blame Contribute Delete
3.3 kB
metadata
language:
  - en
license: llama3.2
base_model: ZenithLLM/ZenAlta-1-3B-Phase2
tags:
  - llama
  - llama3.2
  - speculative-decoding
  - draft-model
  - gguf
  - conversational
  - roleplay
  - mobile-llm
pipeline_tag: text-generation

⚑ Zen Alta 4-Layer Speculative Decoding Draft Model (~790M)

Zen Alta Draft is a ultra-lightweight, 4-layer speculative decoding companion model engineered by ZenithLLM. Sliced from the top of the 24-layer Zen Alta architecture, it shares the exact same 128,256 BPE vocabulary and embedding/LM head, enabling lossless 2Γ— speculative inference acceleration in llama.cpp, vLLM, and mobile runtimes.


🎯 What is Speculative Decoding?

In standard autoregressive generation, deep models calculate every single token sequentially (e.g. 24 transformer layers per token).

With Zen Alta Draft:

  1. The Fast Draft (4 Layers, 790M): Quickly guesses 4–5 candidate tokens in parallel in just ~20–30ms.
  2. The Target Model (Zen Alta 24 Layers, 2.8B): Verifies all proposed tokens in a single parallel forward pass (~40ms).
  3. Result: Accepted tokens are committed simultaneously, achieving 40–50+ tokens/sec on mobile chips with 0% degradation in output quality or persona.

πŸ“¦ Model Specifications

Parameter Value
Base Architecture Llama 3.2 (CausalLM)
Hidden Layers 4 (vs 24 in Target model)
Hidden Dimension 3072
Intermediate Size 8192
Attention Heads 24 query heads / 8 KV heads
Vocabulary Size 128,256 (Identical to Llama 3.2 & Zen Alta)
Context Length 131,072 tokens
RoPE Theta 500,000.0

πŸ“‚ Repository Contents

This consolidated repository contains both the raw PyTorch weights and the ready-to-run quantized GGUF:

File Size Description
model.safetensors 1.52 GB Unquantized FP16 PyTorch weights (4 layers)
zen-alta-draft-q4_k_m.gguf 545.72 MB Quantized 4-bit medium GGUF for llama.cpp & mobile
config.json < 1 KB 4-layer model configuration
tokenizer.json 16.4 MB Fast BPE tokenizer definition
chat_template.jinja < 4 KB Llama 3.2 conversational chat template

πŸš€ How to Run in llama.cpp

Speculative Decoding (Target + Draft Pairing)

Download the target model from ZenithLLM/ZenAlta-1-3B-Phase2-GGUF and the draft model from this repo:

# Speculative decoding command
./llama-cli \
  -m ZenAlta-1-3B-Pruned.Q4_K_M.gguf \
  -md zen-alta-draft-q4_k_m.gguf \
  --draft-max 5 \
  -p "<|start_header_id|>user<|end_header_id|>\n\nhey who are you?<|eot_id|><|start_header_id|>assistant<|end_header_id|>\n\n" \
  -n 128

Standalone Inference (Fast Preview)

./llama-cli -m zen-alta-draft-q4_k_m.gguf -p "what is up" -n 64

πŸ”— Related Models