โšก Zen Alta Draft (GGUF Quantized)

This repository provides the quantized Q4_K_M GGUF binary (545.72 MB) for the Zen Alta 4-Layer Speculative Decoding Draft Model.

Consolidated Repository: The full model with both FP16 Safetensors weights and GGUF quantization is available at ZenithLLM/ZenAlta-Draft.


๐Ÿš€ Speculative Decoding Usage in llama.cpp

Pair this draft model with the target model ZenithLLM/ZenAlta-1-3B-Phase2-GGUF:

./llama-cli \
  -m ZenAlta-1-3B-Pruned.Q4_K_M.gguf \
  -md zen-alta-draft-q4_k_m.gguf \
  --draft-max 5 \
  -p "<|start_header_id|>user<|end_header_id|>\n\nhey what are you doing<|eot_id|><|start_header_id|>assistant<|end_header_id|>\n\n" \
  -n 128

๐Ÿ“ฆ File Information

  • File: zen-alta-draft-q4_k_m.gguf
  • Size: 545.72 MB
  • Quantization: Q4_K_M (5.67 bits per weight)
  • Vocabulary: 128,256 BPE (Identical to Llama 3.2 3B & Zen Alta)
Downloads last month
62
GGUF
Model size
0.8B params
Architecture
llama
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for ZenithLLM/zen-alta-draft-gguf

Quantized
(3)
this model