pack-precision-map / index.html
aleada's picture
pack precision map
e73d8ef verified
Raw History Blame Contribute Delete
12.2 kB
<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>How much of this pack is actually 4-bit?</title>
<!-- The Space renders inside an iframe and huggingface.co sends
x-frame-options: DENY, so any link without a target tries to load a
refused page in the frame and reads as broken. -->
<base target="_blank">
<style>
:root {
--bg: #fbfbfa; --panel: #ffffff; --ink: #1a1a19; --muted: #6b6b66;
--line: #e4e3df; --accent: #3d5a80;
--problem: #a03030; --problem-bg: #fdf3f2;
--check: #8a6212; --check-bg: #fdf9ee;
--clear: #2f6b45; --clear-bg: #f1f8f3;
--q: #3d5a80; --m: #9a9790; --f: #b5793a; --bfr: #d9d6d0;
--mono: ui-monospace, "SF Mono", "Cascadia Mono", Menlo, monospace;
}
@media (prefers-color-scheme: dark) {
:root {
--bg: #16171a; --panel: #1d1f23; --ink: #e8e6e3; --muted: #9a978f;
--line: #2e3137; --accent: #8ab0d9;
--problem: #e08a86; --problem-bg: #2a1e1e;
--check: #d9b76a; --check-bg: #29241a;
--clear: #8fc9a6; --clear-bg: #1b2620;
--q: #7fa3cc; --m: #6f6c66; --f: #d19a5e; --bfr: #3a3d43;
}
}
* { box-sizing: border-box; }
body {
margin: 0; background: var(--bg); color: var(--ink);
font: 16px/1.65 ui-sans-serif, system-ui, -apple-system, "Segoe UI", sans-serif;
-webkit-font-smoothing: antialiased;
}
.wrap { max-width: 50rem; margin: 0 auto; padding: 3.5rem 1.25rem 5rem; }
.brand { display: inline-block; margin-bottom: 1.75rem; }
.brand img { height: 1.5rem; width: auto; display: block; filter: invert(1); opacity: .78; }
@media (prefers-color-scheme: dark) { .brand img { filter: none; opacity: .9; } }
header h1 {
font-size: clamp(1.6rem, 4.5vw, 2.2rem); line-height: 1.2;
letter-spacing: -0.02em; margin: 0 0 .75rem; font-weight: 620;
}
header p { color: var(--muted); margin: 0 0 .5rem; max-width: 42rem; }
form { display: flex; gap: .5rem; flex-wrap: wrap; margin: 2rem 0 .75rem; }
input[type=text] {
flex: 1 1 20rem; min-width: 0; padding: .7rem .85rem;
border: 1px solid var(--line); border-radius: .5rem;
background: var(--panel); color: var(--ink);
font: inherit; font-family: var(--mono); font-size: .92rem;
}
input[type=text]:focus { outline: 2px solid var(--accent); outline-offset: -1px; border-color: transparent; }
button {
padding: .7rem 1.4rem; border: 0; border-radius: .5rem;
background: var(--accent); color: #fff; font: inherit; font-weight: 560; cursor: pointer;
}
button:hover { filter: brightness(1.08); }
button:disabled { opacity: .55; cursor: progress; }
.examples { font-size: .86rem; color: var(--muted); margin-bottom: 2.5rem; line-height: 2; }
.examples .group { display: inline-block; min-width: 8.5rem; }
.examples button {
background: none; color: var(--accent); padding: 0 .15rem; font-size: .86rem;
font-family: var(--mono); text-decoration: underline; text-underline-offset: 2px; font-weight: 400;
}
.card {
background: var(--panel); border: 1px solid var(--line);
border-radius: .7rem; padding: 1.1rem 1.3rem; margin: .85rem 0;
}
.card h3 { margin: 0 0 .5rem; font-size: 1.02rem; font-weight: 600; }
.card h4 { margin: 1.4rem 0 .4rem; font-size: .85rem; font-weight: 600;
text-transform: uppercase; letter-spacing: .05em; color: var(--muted); }
.card p { margin: .5rem 0; }
.problem { border-left: 3px solid var(--problem); background: var(--problem-bg); }
.problem h3 { color: var(--problem); }
.check { border-left: 3px solid var(--check); background: var(--check-bg); }
.check h3 { color: var(--check); }
.clear { border-left: 3px solid var(--clear); background: var(--clear-bg); }
.clear h3 { color: var(--clear); }
.warn { color: var(--problem); }
.note { color: var(--muted); font-size: .92rem; }
.bar {
display: flex; height: 1.15rem; border-radius: .3rem; overflow: hidden;
margin: 1rem 0 .3rem; background: var(--line);
}
.seg { display: block; height: 100%; }
.seg + .seg { box-shadow: inset 2px 0 0 var(--panel); }
.seg.q { background: var(--q); }
.seg.m { background: var(--m); }
.seg.f { background: var(--f); }
.seg.b { background: var(--bfr); }
.key { display: flex; flex-wrap: wrap; gap: .35rem 1.1rem; font-size: .82rem;
color: var(--muted); margin: .6rem 0 0; }
.key i { display: inline-block; width: .62rem; height: .62rem; border-radius: .15rem;
margin-right: .35rem; vertical-align: baseline; }
.meta {
display: flex; flex-wrap: wrap; gap: .35rem 1.5rem; font-size: .86rem;
color: var(--muted); padding-bottom: .9rem; margin-bottom: .3rem;
border-bottom: 1px solid var(--line);
}
.meta code { color: var(--ink); }
code { font-family: var(--mono); font-size: .86em; }
.muted { color: var(--muted); }
.small { font-size: .84rem; color: var(--muted); }
h2 { font-size: 1.05rem; font-weight: 600; letter-spacing: -0.01em; margin: 3rem 0 .75rem; }
table { border-collapse: collapse; width: 100%; font-size: .87rem; }
th, td { text-align: left; padding: .45rem .7rem .45rem 0; border-bottom: 1px solid var(--line); vertical-align: top; }
th { font-weight: 600; color: var(--muted); font-size: .8rem; text-transform: uppercase; letter-spacing: .04em; }
td.num { text-align: right; font-family: var(--mono); font-size: .85rem; white-space: nowrap; }
table.rows { margin: .5rem 0; }
table.rows td:first-child { width: 45%; }
.scroll { overflow-x: auto; }
footer { margin-top: 3.5rem; padding-top: 1.5rem; border-top: 1px solid var(--line); font-size: .87rem; color: var(--muted); }
footer a { color: var(--accent); }
.spin { color: var(--muted); font-size: .9rem; }
</style>
</head>
<body>
<div class="wrap">
<header>
<a class="brand" href="https://assert.gr" rel="noopener">
<img src="./assert-logo.png" alt="ASSERT" width="420" height="87">
</a>
<h1>How much of this pack is actually 4-bit?</h1>
<p>
A name like <code>W4A16</code> or <code>-AWQ-4bit</code> reads as though
everything in the file is four bits. It never is. Embeddings, the output
head, the norms and any preserved auxiliary head stay at the source's
precision — commonly a fifth to two fifths of the bytes — and no model page
says which, or how much.
</p>
<p>
<strong>The set matters more than the total.</strong> An <code>lm_head</code>
left whole is a deliberate, defensible cost. An <code>lm_head</code>
quantized is a decision the pack made on your behalf without mentioning it:
measured on our own weights, quantizing it flips the model's chosen token on
about a fifth of positions.
</p>
<p>
This reads <strong>safetensors headers only</strong>, over ranged requests
from your browser — a few hundred kilobytes against packs of many gigabytes.
It never loads the model and never runs it.
</p>
</header>
<form id="form">
<input type="text" id="repo" value="aleada/Nemotron-3.5-Lightning-30B-A3B-W4A16"
placeholder="owner/model" autocomplete="off" spellcheck="false"
aria-label="Model repo id">
<button type="submit" id="go">Inspect</button>
</form>
<div class="examples">
<span class="group">Different tools:</span>
<button type="button" data-ex="casperhansen/llama-3-8b-instruct-awq">AutoAWQ</button> ·
<button type="button" data-ex="TheBloke/Llama-2-7B-Chat-GPTQ">GPTQ</button> ·
<button type="button" data-ex="unsloth/Meta-Llama-3.1-8B-Instruct-bnb-4bit">bitsandbytes</button> ·
<button type="button" data-ex="nvidia/Llama-3.3-70B-Instruct-FP4">ModelOpt&nbsp;NVFP4</button> ·
<button type="button" data-ex="openai/gpt-oss-20b">MXFP4</button>
<br>
<span class="group">Worth comparing:</span>
<button type="button" data-ex="cyankiwi/Qwen3-VL-8B-Instruct-AWQ-4bit">a pack whose name says AWQ</button> ·
<button type="button" data-ex="Qwen/Qwen2.5-0.5B-Instruct">an unquantized model</button>
<br>
<span class="group">Our packs:</span>
<button type="button" data-ex="aleada/Nemotron-3.5-Lightning-30B-A3B-W4A16">Nemotron-3.5</button> ·
<button type="button" data-ex="aleada/Qwen3.8-27B-W4A16">Qwen3.8-27B</button>
</div>
<div id="out"></div>
<div class="key">
<span><i style="background:var(--q)"></i>quantized payload</span>
<span><i style="background:var(--m)"></i>quantization data</span>
<span><i style="background:var(--f)"></i>full precision</span>
<span><i style="background:var(--bfr)"></i>index buffers</span>
</div>
<h2>Why this reads dtypes, not tensor names</h2>
<p>
The first version of this tool classified tensors by <em>name</em>, learned
from compressed-tensors packs where the payload is called
<code>weight_packed</code>. Probing nine toolchains showed that convention is
the minority. <strong>bitsandbytes, ModelOpt FP8, ModelOpt NVFP4 and
compressed-tensors' own <code>int-quantized</code> and
<code>naive-quantized</code> formats all store the quantized payload under the
plain name <code>weight</code></strong> — so a name-driven reader calls every
one of those packs 100% full precision, confidently, and about most of what is
published.
</p>
<p>
So the dtype decides: an <code>I32</code>, <code>I8</code>, <code>U8</code> or
<code>F8</code> tensor is not full precision whatever it is called. Names are
used only to separate quantization metadata — scales, zero points, group
indices — from model weights, a distinction that genuinely has no dtype
signature.
</p>
<p>
The declared method is read from <code>config.json</code>, or from
<code>hf_quant_config.json</code> where the tool writes it there — never from
the repo name. <code>cyankiwi/Qwen3-VL-8B-Instruct-<strong>AWQ</strong>-4bit</code>
declares <code>compressed-tensors</code>.
</p>
<h2>Conventions, and where each was verified</h2>
<p>
None of these naming conventions is a stable public API — they are internal
choices of each tool, and they have moved before. Two versions therefore
appear, answering different questions: the <strong>producer's</strong>, read
out of the pack itself wherever that tool records it, which is the version
that actually wrote those bytes; and <strong>this reader's</strong>, the date
below, which says how stale the table is rather than leaving a wrong answer
looking authoritative.
</p>
<div class="scroll">
<table>
<thead><tr><th>Declared as</th><th>Tool</th><th>Records its version in</th><th>Verified against</th></tr></thead>
<tbody id="formats"></tbody>
</table>
</div>
<footer>
<p>
<strong>A high full-precision share is not a fault.</strong> It is usually
the pack being careful — an output head or a vision tower left whole costs
size and buys accuracy. This says what was kept and what it cost in bytes;
it does not say whether the trade was a good one.
</p>
<p>
Conventions verified <strong id="measured">—</strong> against one published
pack per tool, listed above. AutoAWQ's <code>version</code> field is
deliberately not reported as a tool version: it holds <code>gemm</code>, the
kernel variant. Packs that ship a <code>.pt</code> or <code>.bin</code>
instead of safetensors — torchao and HQQ commonly do — cannot be read from
metadata and are reported as unreadable rather than as unquantized. Gated
and private repos cannot be read.
</p>
<p>
Built from the tooling behind the
<a href="https://huggingface.co/aleada" rel="noopener">aleada</a> W4A16 packs.
See also the <a href="https://huggingface.co/spaces/aleada/pack-integrity-check" rel="noopener">pack integrity checker</a>, the <a href="https://huggingface.co/spaces/aleada/reasoning-parser-advisor" rel="noopener">reasoning-parser advisor</a> and the <a href="https://huggingface.co/spaces/aleada/model-search-that-answers" rel="noopener">model search</a>.
</p>
</footer>
</div>
<script type="module" src="./app.js"></script>
</body>
</html>