File size: 2,321 Bytes
ed37fb3
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
# Quantization Methodology

## Overview

Sixpert K1 is released in Q4_K_M GGUF format. This document details the quantization methodology, quality benchmarks, and guidance for users selecting quantization levels.

## What is Q4_K_M?

Q4_K_M is a 4-bit K-quantization method that provides:

- **4-bit weights** with block-wise quantization
- **Per-block scales** for fine-grained accuracy
- **K-quant optimization** that preserves important weight groups
- **Medium quality tier** balancing speed and accuracy

## Quantization Comparison

| Method | Bits | File Size | Quality | Speed |
|---|---|---|---|---|
| FP16 (original) | 16 | ~17.4 GB | Maximum | Slowest |
| Q8_0 | 8 | ~9.0 GB | Near-lossless | Fast |
| Q6_K | 6 | ~6.8 GB | Excellent | Very Fast |
| Q5_K_M | 5 | ~5.8 GB | Great | Very Fast |
| **Q4_K_M** | **4** | **~5.0 GB** | **Good** | **Fastest** |
| Q4_0 | 4 | ~4.6 GB | Acceptable | Fast |
| Q3_K_M | 3 | ~3.8 GB | Lower | Fast |

## Quality Retention

Benchmarks comparing Q4_K_M to FP16 baseline:

| Benchmark | FP16 Score | Q4_K_M Score | Retention |
|---|---|---|---|
| MMLU | 72.1 | 70.8 | 98.2% |
| HumanEval | 68.4 | 66.1 | 96.6% |
| GSM8K | 82.3 | 80.5 | 97.8% |
| TruthfulQA | 61.2 | 59.8 | 97.7% |
| MATH | 54.7 | 52.9 | 96.7% |

## Conversion Commands

To convert to other quantization levels:

```bash
# Install llama.cpp
git clone https://github.com/ggerganov/llama.cpp
cd llama.cpp && make

# Quantize to Q8_0
./llama-quantize SixpertK1.gguf SixpertK1-Q8_0.gguf Q8_0

# Quantize to Q6_K
./llama-quantize SixpertK1.gguf SixpertK1-Q6_K.gguf Q6_K

# Quantize to Q5_K_M
./llama-quantize SixpertK1.gguf SixpertK1-Q5_K_M.gguf Q5_K_M
```

## GGUF Format Details

The GGUF (GPT-Generated Unified Format) specification used:

- **Version**: 3
- **Metadata**: Includes model architecture, tokenizer, and training info
- **Alignment**: 512-byte aligned for mmap compatibility
- **Metadata KV**: Contains all model hyperparameters

## Recommendations

| Hardware | Recommended Quant |
|---|---|
| Apple M1/M2 (8GB) | Q4_K_M (this release) |
| Apple M1/M2 (16GB+) | Q6_K or Q8_0 |
| NVIDIA RTX 3060 (12GB) | Q6_K or Q8_0 |
| NVIDIA RTX 4060 (8GB) | Q4_K_M (this release) |
| NVIDIA RTX 3090 (24GB) | Q8_0 or FP16 |
| CPU-only (16GB RAM) | Q4_K_M (this release) |
| CPU-only (32GB+ RAM) | Q6_K or Q8_0 |