# Quantization Methodology ## Overview Sixpert K1 is released in Q4_K_M GGUF format. This document details the quantization methodology, quality benchmarks, and guidance for users selecting quantization levels. ## What is Q4_K_M? Q4_K_M is a 4-bit K-quantization method that provides: - **4-bit weights** with block-wise quantization - **Per-block scales** for fine-grained accuracy - **K-quant optimization** that preserves important weight groups - **Medium quality tier** balancing speed and accuracy ## Quantization Comparison | Method | Bits | File Size | Quality | Speed | |---|---|---|---|---| | FP16 (original) | 16 | ~17.4 GB | Maximum | Slowest | | Q8_0 | 8 | ~9.0 GB | Near-lossless | Fast | | Q6_K | 6 | ~6.8 GB | Excellent | Very Fast | | Q5_K_M | 5 | ~5.8 GB | Great | Very Fast | | **Q4_K_M** | **4** | **~5.0 GB** | **Good** | **Fastest** | | Q4_0 | 4 | ~4.6 GB | Acceptable | Fast | | Q3_K_M | 3 | ~3.8 GB | Lower | Fast | ## Quality Retention Benchmarks comparing Q4_K_M to FP16 baseline: | Benchmark | FP16 Score | Q4_K_M Score | Retention | |---|---|---|---| | MMLU | 72.1 | 70.8 | 98.2% | | HumanEval | 68.4 | 66.1 | 96.6% | | GSM8K | 82.3 | 80.5 | 97.8% | | TruthfulQA | 61.2 | 59.8 | 97.7% | | MATH | 54.7 | 52.9 | 96.7% | ## Conversion Commands To convert to other quantization levels: ```bash # Install llama.cpp git clone https://github.com/ggerganov/llama.cpp cd llama.cpp && make # Quantize to Q8_0 ./llama-quantize SixpertK1.gguf SixpertK1-Q8_0.gguf Q8_0 # Quantize to Q6_K ./llama-quantize SixpertK1.gguf SixpertK1-Q6_K.gguf Q6_K # Quantize to Q5_K_M ./llama-quantize SixpertK1.gguf SixpertK1-Q5_K_M.gguf Q5_K_M ``` ## GGUF Format Details The GGUF (GPT-Generated Unified Format) specification used: - **Version**: 3 - **Metadata**: Includes model architecture, tokenizer, and training info - **Alignment**: 512-byte aligned for mmap compatibility - **Metadata KV**: Contains all model hyperparameters ## Recommendations | Hardware | Recommended Quant | |---|---| | Apple M1/M2 (8GB) | Q4_K_M (this release) | | Apple M1/M2 (16GB+) | Q6_K or Q8_0 | | NVIDIA RTX 3060 (12GB) | Q6_K or Q8_0 | | NVIDIA RTX 4060 (8GB) | Q4_K_M (this release) | | NVIDIA RTX 3090 (24GB) | Q8_0 or FP16 | | CPU-only (16GB RAM) | Q4_K_M (this release) | | CPU-only (32GB+ RAM) | Q6_K or Q8_0 |