Spaces:
Running
Download README.md from aleada/pack-precision-map: direct link, hf CLI and curl.
- Browser
- Download file 3.49 kB
-
https://huggingface.co/spaces/aleada/pack-precision-map/resolve/main/README.md
- Command line
-
hf download hf://spaces/aleada/pack-precision-map/README.md
-
curl -L -o README.md https://huggingface.co/spaces/aleada/pack-precision-map/resolve/main/README.md
title: How much of this pack is actually 4-bit?
emoji: 🔍
colorFrom: blue
colorTo: gray
sdk: static
app_file: index.html
pinned: false
license: apache-2.0
short_description: What is really 4-bit in a quantized pack
How much of this pack is actually 4-bit?
A name like W4A16 or -AWQ-4bit reads as though everything in the file
is four bits. It never is. Embeddings, the output head, the norms and any
preserved auxiliary head stay at the source's precision — commonly a
fifth to two fifths of the bytes — and no model page says which, or how
much.
The set matters more than the total. An lm_head left whole is a
deliberate, defensible cost. An lm_head quantized is a decision the
pack made on the user's behalf without mentioning it.
This reads safetensors headers only, over ranged requests from the visitor's browser: a few hundred kilobytes against packs of many gigabytes. It never loads the model and never runs it.
Why it reads dtypes and not tensor names
The first version classified by name, learned from compressed-tensors
packs where the payload is called weight_packed. Probing nine
toolchains showed that convention is the minority: bitsandbytes,
ModelOpt FP8, ModelOpt NVFP4 and compressed-tensors' own int-quantized
and naive-quantized formats all store the quantized payload under the
plain name weight. A name-driven reader calls every one of those
packs 100% full precision — confidently, and about most of what is
published.
So dtype is the primary evidence. Names separate quantization metadata from model weights, which is the one distinction with no dtype signature.
Versions
No naming convention here is a stable public API. Two versions are reported, answering different questions:
- The producer's, read out of the pack itself wherever that tool
records it —
compressed-tensorswritesversion, AutoRound writesautoround_version, GPTQModel writesmeta.quantizer, ModelOpt writesproducerinto a separatehf_quant_config.json. That is the version that actually wrote those bytes. - This reader's — the verification date shown in the footer, with the pack each convention was checked against. If a release moves a name, that date says how stale the table is instead of leaving a wrong answer looking authoritative.
AutoAWQ's version field is deliberately not reported as a tool version:
it holds gemm, the kernel variant.
Development
The rules live twice — precision_map.py (where they were developed, and
where they were caught being wrong) and precision_map.js (what ships to
the browser, because the free Space tier is static). A parity gate holds
them together over fixtures harvested from real packs of nine
quantizers:
python harvest_cases.py # re-sample fixtures from published packs
python gen_golden.py # record the Python answers
node parity.mjs # fail if the JavaScript disagrees
./deploy.sh aleada/pack-precision-map
deploy.sh runs the gate before it uploads, so a drifted port cannot
ship.
Companion tools
- pack-integrity-check — does this pack's exclusion list match what is actually in it?
- reasoning-parser-advisor
— does this model need
--reasoning-parser, and which one?
Built from the tooling behind the aleada W4A16 packs.