pack-precision-map / README.md
aleada's picture
pack precision map
1a5245f verified
|
Raw History Blame Contribute Delete
3.49 kB
metadata
title: How much of this pack is actually 4-bit?
emoji: 🔍
colorFrom: blue
colorTo: gray
sdk: static
app_file: index.html
pinned: false
license: apache-2.0
short_description: What is really 4-bit in a quantized pack

How much of this pack is actually 4-bit?

A name like W4A16 or -AWQ-4bit reads as though everything in the file is four bits. It never is. Embeddings, the output head, the norms and any preserved auxiliary head stay at the source's precision — commonly a fifth to two fifths of the bytes — and no model page says which, or how much.

The set matters more than the total. An lm_head left whole is a deliberate, defensible cost. An lm_head quantized is a decision the pack made on the user's behalf without mentioning it.

This reads safetensors headers only, over ranged requests from the visitor's browser: a few hundred kilobytes against packs of many gigabytes. It never loads the model and never runs it.

Why it reads dtypes and not tensor names

The first version classified by name, learned from compressed-tensors packs where the payload is called weight_packed. Probing nine toolchains showed that convention is the minority: bitsandbytes, ModelOpt FP8, ModelOpt NVFP4 and compressed-tensors' own int-quantized and naive-quantized formats all store the quantized payload under the plain name weight. A name-driven reader calls every one of those packs 100% full precision — confidently, and about most of what is published.

So dtype is the primary evidence. Names separate quantization metadata from model weights, which is the one distinction with no dtype signature.

Versions

No naming convention here is a stable public API. Two versions are reported, answering different questions:

  • The producer's, read out of the pack itself wherever that tool records it — compressed-tensors writes version, AutoRound writes autoround_version, GPTQModel writes meta.quantizer, ModelOpt writes producer into a separate hf_quant_config.json. That is the version that actually wrote those bytes.
  • This reader's — the verification date shown in the footer, with the pack each convention was checked against. If a release moves a name, that date says how stale the table is instead of leaving a wrong answer looking authoritative.

AutoAWQ's version field is deliberately not reported as a tool version: it holds gemm, the kernel variant.

Development

The rules live twice — precision_map.py (where they were developed, and where they were caught being wrong) and precision_map.js (what ships to the browser, because the free Space tier is static). A parity gate holds them together over fixtures harvested from real packs of nine quantizers:

python harvest_cases.py    # re-sample fixtures from published packs
python gen_golden.py       # record the Python answers
node parity.mjs            # fail if the JavaScript disagrees
./deploy.sh aleada/pack-precision-map

deploy.sh runs the gate before it uploads, so a drifted port cannot ship.

Companion tools

Built from the tooling behind the aleada W4A16 packs.