--- title: How much of this pack is actually 4-bit? emoji: 🔍 colorFrom: blue colorTo: gray sdk: static app_file: index.html pinned: false license: apache-2.0 short_description: What is really 4-bit in a quantized pack --- # How much of this pack is actually 4-bit? A name like `W4A16` or `-AWQ-4bit` reads as though everything in the file is four bits. It never is. Embeddings, the output head, the norms and any preserved auxiliary head stay at the source's precision — commonly a fifth to two fifths of the bytes — and no model page says which, or how much. The set matters more than the total. An `lm_head` left whole is a deliberate, defensible cost. An `lm_head` quantized is a decision the pack made on the user's behalf without mentioning it. This reads **safetensors headers only**, over ranged requests from the visitor's browser: a few hundred kilobytes against packs of many gigabytes. It never loads the model and never runs it. ## Why it reads dtypes and not tensor names The first version classified by name, learned from compressed-tensors packs where the payload is called `weight_packed`. Probing nine toolchains showed that convention is the minority: **bitsandbytes, ModelOpt FP8, ModelOpt NVFP4 and compressed-tensors' own `int-quantized` and `naive-quantized` formats all store the quantized payload under the plain name `weight`.** A name-driven reader calls every one of those packs 100% full precision — confidently, and about most of what is published. So dtype is the primary evidence. Names separate quantization metadata from model weights, which is the one distinction with no dtype signature. ## Versions No naming convention here is a stable public API. Two versions are reported, answering different questions: * **The producer's**, read out of the pack itself wherever that tool records it — `compressed-tensors` writes `version`, AutoRound writes `autoround_version`, GPTQModel writes `meta.quantizer`, ModelOpt writes `producer` into a separate `hf_quant_config.json`. That is the version that actually wrote those bytes. * **This reader's** — the verification date shown in the footer, with the pack each convention was checked against. If a release moves a name, that date says how stale the table is instead of leaving a wrong answer looking authoritative. AutoAWQ's `version` field is deliberately not reported as a tool version: it holds `gemm`, the kernel variant. ## Development The rules live twice — `precision_map.py` (where they were developed, and where they were caught being wrong) and `precision_map.js` (what ships to the browser, because the free Space tier is static). A parity gate holds them together over fixtures harvested from real packs of nine quantizers: ```bash python harvest_cases.py # re-sample fixtures from published packs python gen_golden.py # record the Python answers node parity.mjs # fail if the JavaScript disagrees ./deploy.sh aleada/pack-precision-map ``` `deploy.sh` runs the gate before it uploads, so a drifted port cannot ship. ## Companion tools * [pack-integrity-check](https://huggingface.co/spaces/aleada/pack-integrity-check) — does this pack's exclusion list match what is actually in it? * [reasoning-parser-advisor](https://huggingface.co/spaces/aleada/reasoning-parser-advisor) — does this model need `--reasoning-parser`, and which one? Built from the tooling behind the [aleada](https://huggingface.co/aleada) W4A16 packs.