gemma / README.md
zenpeach's picture
README: report the QAT E4B results from both test machines
3ac7045 verified
|
Raw History Blame Contribute Delete
1.71 kB
---
license: gemma
---
Gemma 4 model weights in GGUF format, used by Understand's AI features and
served locally by ullama (llama.cpp).
- `gemma-4-E4B-it-qat-UD-Q4_K_XL.gguf` — what Understand's Gemma4-E4B choice
downloads. Google's quantization-aware-trained (QAT) E4B as quantized by
Unsloth, copied unmodified from
[unsloth/gemma-4-E4B-it-qat-GGUF](https://huggingface.co/unsloth/gemma-4-E4B-it-qat-GGUF)
(SHA-256 `df0fd4ee07072c607c29a0a1cb4f98918426cca12f45a2776bdd6ee6d09a4de3`).
Tested against the file below on an Apple M5 MacBook Pro and an RTX 4090
Linux machine, it qualified on both. It followed instructions better (0.944 vs
0.897 on the M5, 0.906 vs 0.844 on Linux), answered sooner (51 s vs 62 s median
on the M5) and wrote more accurate code summaries (0.663 vs 0.639 on the M5),
in a 4.2 GB download instead of 5.0 GB. Its chat answers matched the code about
as closely (0.694 vs 0.721 on the M5, 0.749 vs 0.727 on Linux).
- `gemma-4-E4B-it-Q4_K_M.gguf` — the earlier E4B choice, qualified in SciTools'
two-workload evaluation (chat quality 148.6 on the frozen GitAhead scale,
3/3 qualifying runs). Kept for installs that already use it.
Ship and use exactly these files: verdicts do not carry across quantizations.
**Gemma is provided under and subject to the Gemma Terms of Use found at
[ai.google.dev/gemma/terms](https://ai.google.dev/gemma/terms).** By
downloading these weights you agree to those terms, including the Gemma
Prohibited Use Policy ([ai.google.dev/gemma/prohibited_use_policy](https://ai.google.dev/gemma/prohibited_use_policy)).
This repository redistributes the weights unmodified apart from quantization;
it is not endorsed by Google.