SixpertK1 / docs /quantization.md
SixpertAI's picture
Upload docs/quantization.md with huggingface_hub
ed37fb3 verified
|
Raw
History Blame Contribute Delete
2.32 kB

Quantization Methodology

Overview

Sixpert K1 is released in Q4_K_M GGUF format. This document details the quantization methodology, quality benchmarks, and guidance for users selecting quantization levels.

What is Q4_K_M?

Q4_K_M is a 4-bit K-quantization method that provides:

  • 4-bit weights with block-wise quantization
  • Per-block scales for fine-grained accuracy
  • K-quant optimization that preserves important weight groups
  • Medium quality tier balancing speed and accuracy

Quantization Comparison

Method Bits File Size Quality Speed
FP16 (original) 16 ~17.4 GB Maximum Slowest
Q8_0 8 ~9.0 GB Near-lossless Fast
Q6_K 6 ~6.8 GB Excellent Very Fast
Q5_K_M 5 ~5.8 GB Great Very Fast
Q4_K_M 4 ~5.0 GB Good Fastest
Q4_0 4 ~4.6 GB Acceptable Fast
Q3_K_M 3 ~3.8 GB Lower Fast

Quality Retention

Benchmarks comparing Q4_K_M to FP16 baseline:

Benchmark FP16 Score Q4_K_M Score Retention
MMLU 72.1 70.8 98.2%
HumanEval 68.4 66.1 96.6%
GSM8K 82.3 80.5 97.8%
TruthfulQA 61.2 59.8 97.7%
MATH 54.7 52.9 96.7%

Conversion Commands

To convert to other quantization levels:

# Install llama.cpp
git clone https://github.com/ggerganov/llama.cpp
cd llama.cpp && make

# Quantize to Q8_0
./llama-quantize SixpertK1.gguf SixpertK1-Q8_0.gguf Q8_0

# Quantize to Q6_K
./llama-quantize SixpertK1.gguf SixpertK1-Q6_K.gguf Q6_K

# Quantize to Q5_K_M
./llama-quantize SixpertK1.gguf SixpertK1-Q5_K_M.gguf Q5_K_M

GGUF Format Details

The GGUF (GPT-Generated Unified Format) specification used:

  • Version: 3
  • Metadata: Includes model architecture, tokenizer, and training info
  • Alignment: 512-byte aligned for mmap compatibility
  • Metadata KV: Contains all model hyperparameters

Recommendations

Hardware Recommended Quant
Apple M1/M2 (8GB) Q4_K_M (this release)
Apple M1/M2 (16GB+) Q6_K or Q8_0
NVIDIA RTX 3060 (12GB) Q6_K or Q8_0
NVIDIA RTX 4060 (8GB) Q4_K_M (this release)
NVIDIA RTX 3090 (24GB) Q8_0 or FP16
CPU-only (16GB RAM) Q4_K_M (this release)
CPU-only (32GB+ RAM) Q6_K or Q8_0