Download loopq_quantization/README.md from JunYoungLee/ut-depth-probe-artifacts: direct link, hf CLI and curl.
- Browser
- Download file 10.1 kB
-
https://huggingface.co/JunYoungLee/ut-depth-probe-artifacts/resolve/main/loopq_quantization/README.md
- Command line
-
hf download hf://JunYoungLee/ut-depth-probe-artifacts/loopq_quantization/README.md
-
curl -L -o README.md https://huggingface.co/JunYoungLee/ut-depth-probe-artifacts/resolve/main/loopq_quantization/README.md
LoopQ 4-bit quantization of Ouro-1.4B β raw artifacts
A reproduction of LoopQ (arXiv:2605.16343v1),
post-training quantization designed for depth-recurrent transformers, on
Ouro-1.4B (ByteDance/Ouro-1.4B, revision
e3b1e0993b1231a51d6069a870476dda4162c00a) β 24 layers run 4 times per token.
Everything here is raw output. No number below is copied from a write-up; each
comes from results/**/harness/**/results_*.json produced by lm-eval at the
pinned commit 64f3d0924fc695efd6d776a5ac91f97138085516.
Layout
| path | what |
|---|---|
results/<arm>/<task>/harness/ |
lm-eval output per task: results_*.json, samples_*.jsonl, installed_task_snapshot |
results/<arm>/<task>/run_metadata.json, timing.json |
protocol note, pinned versions, wall clock |
results/slt_selection_screen/ |
the four selection arms of the SLT stability screen |
results/loopq_w4a4_step_curve/ |
the completed calibration evaluated at five step counts, with its own README |
artifacts/*.components.pt |
the calibrated LoopQ parameters (LAS scales, shared and selected transforms, CTA) |
artifacts/cal_a4_s0_journal_step1050.resume.pt |
the calibration journal that survived the crashed run β carries per-step loss and gradient traces |
artifacts/curve_journals/ |
nine journals of the completed run, one per step target, with its own README |
parity/*.json |
HF-vs-vLLM prefill agreement for each deployed artifact |
jobs/*.job.json |
the scheduler's job records: exact command lines, digests, status |
scripts/ |
the implementation snapshot that produced all of this |
The arms
| arm | what it is |
|---|---|
bf16_a16 |
unquantized reference |
rtn_w4a4_group32, rtn_w4a8_group32 |
plain symmetric round-to-nearest, activations in groups of 32 β the granularity LoopQ itself uses (paper Table 6) |
rtn_w4a4_per_token |
the same RTN code with one activation scale per token β the field-standard W4A4 baseline setting |
loopq_w4a4_untrained |
LoopQ's structure at initialization: random orthogonal Kronecker transforms, LAS clip 1.0, CTA at exact identity, zero optimizer updates |
loopq_w4a4_step0975 |
LoopQ after 975 final optimization steps, recovered from the journal of a run that crashed at step 1065 |
loopq_w4a4_step1500 |
the same calibration run to its configured 1500 steps |
loopq_w4a4_completed |
a calibration run to completion with singular values held in log space, so they cannot reach zero β this is the arm that reproduces the paper |
Protocol
Zero-shot everywhere except MMLU (5-shot). Plain likelihood scoring, no chat
template, prompts exactly as the pinned harness produces them. HellaSwag and
ARC-Challenge report acc_norm; WinoGrande, LAMBADA and MMLU report acc;
WikiText reports word_perplexity. Calibration follows the paper: 1,024 Pile
validation samples at 256 tokens, Ξ» = 0.1, KL temperature 1, teacher top-k 1000,
gradient clip 1.0, SLT budget 4, CTA rank 8, activation group size 32.
Serving is a vLLM 0.10.2 out-of-tree backend; quantization is simulated (quantize-dequantize) so that a 4-bit numerical result is measured without a 4-bit kernel. Every evaluated artifact carries a prefill parity report comparing it against the Transformers reference.
What the runs show
The unquantized reference matches the paper on all seven measures.
| HellaSwag | WinoGrande | LAMBADA | ARC-C | MMLU | WikiText ppl | LAMBADA ppl | |
|---|---|---|---|---|---|---|---|
| paper BF16 | 0.7159 | 0.6717 | 0.6505 | 0.5256 | 0.6823 | 12.03 | 5.30 |
| this repo | 0.7166 | 0.6654 | 0.6577 | 0.5239 | 0.6760 | 12.02 | 5.29 |
Activation granularity, not the method, is what separates the paper's baselines from its LoopQ row. The paper reports every W4A4 baseline collapsing to chance. Changing one setting in the same RTN code reproduces that collapse β and leaving it at LoopQ's own group size does not.
| W4A4 | HellaSwag | WinoGrande | LAMBADA | ARC-C | WikiText ppl |
|---|---|---|---|---|---|
| paper Symmetric / QuaRot / SpinQuant / FlatQuant | 0.27β0.29 | 0.45β0.50 | 0.00β0.03 | 0.23β0.24 | 451β10,500 |
rtn_w4a4_per_token |
0.2645 | 0.5043 | 0.0000 | 0.2457 | 39,089 |
rtn_w4a4_group32 |
0.6705 | 0.6259 | 0.5719 | 0.4923 | 16.13 |
LoopQ reproduces the paper, and on four of six measures it is ahead.
loopq_w4a4_completed is the calibration that ran to completion without
diverging; every column here is within the acceptance band registered before
these numbers were measured (accuracy Β±2 pp, perplexity Β±10%).
| paper LoopQ W4A4 | loopq_w4a4_completed |
difference | |
|---|---|---|---|
| MMLU | 0.5465 | 0.5906 | +4.41 pp |
| WinoGrande | 0.5951 | 0.6290 | +3.39 pp |
| HellaSwag | 0.6582 | 0.6809 | +2.27 pp |
| ARC-Challenge | 0.4881 | 0.4906 | +0.25 pp |
| WikiText ppl | 15.81 | 14.3211 | β9.4% (lower is better) |
| LAMBADA acc | 0.5659 | 0.5605 | β0.54 pp |
| LAMBADA ppl | 8.18 | 8.3831 | +2.48% |
Three accuracy columns land outside the Β±2 pp band that was registered before these numbers existed β all three above the paper. The registered failure condition (5 pp or more below the paper) is not met on any column, so by the letter of that rule this is a hold rather than a pass, and the rule is not being rewritten now that the numbers are in. What the columns say substantively is that this implementation loses less to quantization than the paper reports: from BF16 to W4A4, MMLU falls 8.54 pp here against the paper's 13.58 pp.
The step count is what stood between the two. loopq_w4a4_step0975, taken
from the diverging run, misses LAMBADA badly; the completed run does not. Same
paper contract, same data, same hyperparameters β only the transform
parameterization differs.
step0975 |
completed |
change | |
|---|---|---|---|
| LAMBADA acc | 0.5057 | 0.5605 | +5.48 pp |
| LAMBADA ppl | 11.5369 | 8.3831 | β27% |
| HellaSwag | 0.6452 | 0.6809 | +3.57 pp |
| WikiText ppl | 15.9149 | 14.3211 | β10% |
step0975 does reproduce MMLU to 0.01 points (0.5464 against the paper's
0.5465) and WikiText to 0.66%, which is why the failure read as one axis rather
than a broken implementation.
Running the calibration to its configured 1500 steps destroys the model. Not
degrades β destroys: loopq_w4a4_step1500 scores at or below chance on four
tasks and its LAMBADA perplexity is 5.3 million. The journal in artifacts/
records why. Across the 975 final steps the loss rises β median 5.42e4 over the
first quarter against 1.50e5 over the third β and the transforms' smallest
singular value falls from 8.3e-3 at step 1050 to 6.8e-6 at step 1575. Weight
folding inverts those matrices, so an inverse that amplifies by roughly 1.5e5
lands directly in the folded weights.
The paper specifies Ξ», KL temperature, teacher top-k, gradient clipping, group size, CTA rank, SLT budget, calibration sample count and sequence length. It does not specify the optimizer, the learning rate, the number of optimization steps, the batch size, the weight decay, or how the Kronecker transforms are kept invertible β it defers the last to FlatQuant. The collapse lives entirely in that second list.
The transform costs accuracy before calibration earns it back, and it does not
earn back the same share on every measure. loopq_w4a4_untrained isolates
this: applying a random orthogonal Kronecker transform costs 4.7 points of
LAMBADA perplexity relative to plain RTN before a single optimizer step.
| RTN group-32 | untrained LoopQ | step 975 | completed | paper LoopQ | |
|---|---|---|---|---|---|
| HellaSwag | 0.6705 | 0.6300 | 0.6452 | 0.6809 | 0.6582 |
| WinoGrande | 0.6259 | 0.6062 | 0.6117 | 0.6290 | 0.5951 |
| LAMBADA acc | 0.5719 | 0.5015 | 0.5057 | 0.5605 | 0.5659 |
| ARC-Challenge | 0.4923 | 0.4343 | 0.4744 | 0.4906 | 0.4881 |
| MMLU | 0.5340 | 0.4970 | 0.5464 | 0.5906 | 0.5465 |
| WikiText ppl | 16.13 | 18.57 | 15.91 | 14.3211 | 15.81 |
| LAMBADA ppl | 8.23 | 12.94 | 11.54 | 8.3831 | 8.18 |
Share of the untrained-to-paper gap that calibration closes, by arm:
| step 975 | completed | |
|---|---|---|
| WikiText ppl | 96% | 154% |
| HellaSwag | 54% | 181% |
| LAMBADA ppl | 30% | 96% |
| LAMBADA acc | 6.5% | 92% |
Calibration improves every measure in both arms β so the objective was never actively harming LAMBADA. The diverging run simply gave back most of what it earned; the completed run keeps it.
The step count is a free parameter, and it moves LAMBADA. The paper does not
fix the number of optimization steps; 1,500 was our choice.
results/loopq_w4a4_step_curve/ measures the same calibration at five step
counts. WikiText perplexity improves monotonically and settles (14.53 β 14.32).
LAMBADA accuracy does not: it peaks at step 200 and then moves inside a 3.8 pp
band with no trend β 0.5911, 0.5535, 0.5727, 0.5670, 0.5605 β which is 5.4
standard errors wide on 5,153 examples, so it is not sampling noise. The paper's
0.5659 sits inside that band. A single step count therefore cannot be read on its
own as reproducing or failing to reproduce this column.
The SLT selection is stable. results/slt_selection_screen/ varies the
Fisher estimator, the sample count and the seed. The first three of four
selections are identical in every arm, and 128 calibration samples reproduce the
selection made with 1,024. Only the fourth slot moves, and it moves between two
seeds of the same estimator β a near-tie, not an unstable procedure.
Reproducing
Implementation: scripts/ here, or the quantization_LoopQ branch of the
working repository. Pinned stack: torch 2.8.0, vllm 0.10.2, transformers 4.56.2,
lm-eval at 64f3d0924fc695efd6d776a5ac91f97138085516. Hardware: 2 Γ NVIDIA A40.
One calibration is 11β14 hours; the SLT scan is about 80% of that. A full
seven-task evaluation is 4.6 hours, of which MMLU is 95%.