|
Download loopq_quantization/README.md from JunYoungLee/ut-depth-probe-artifacts: direct link, hf CLI and curl.
- Browser
- Download file 10.1 kB
-
https://huggingface.co/JunYoungLee/ut-depth-probe-artifacts/resolve/main/loopq_quantization/README.md
- Command line
-
hf download hf://JunYoungLee/ut-depth-probe-artifacts/loopq_quantization/README.md
-
curl -L -o README.md https://huggingface.co/JunYoungLee/ut-depth-probe-artifacts/resolve/main/loopq_quantization/README.md
10.1 kB
| # LoopQ 4-bit quantization of Ouro-1.4B β raw artifacts | |
| A reproduction of **LoopQ** ([arXiv:2605.16343v1](https://arxiv.org/abs/2605.16343)), | |
| post-training quantization designed for *depth-recurrent* transformers, on | |
| **Ouro-1.4B** (`ByteDance/Ouro-1.4B`, revision | |
| `e3b1e0993b1231a51d6069a870476dda4162c00a`) β 24 layers run 4 times per token. | |
| Everything here is raw output. No number below is copied from a write-up; each | |
| comes from `results/**/harness/**/results_*.json` produced by lm-eval at the | |
| pinned commit `64f3d0924fc695efd6d776a5ac91f97138085516`. | |
| ## Layout | |
| | path | what | | |
| |---|---| | |
| | `results/<arm>/<task>/harness/` | lm-eval output per task: `results_*.json`, `samples_*.jsonl`, `installed_task_snapshot` | | |
| | `results/<arm>/<task>/run_metadata.json`, `timing.json` | protocol note, pinned versions, wall clock | | |
| | `results/slt_selection_screen/` | the four selection arms of the SLT stability screen | | |
| | `results/loopq_w4a4_step_curve/` | the completed calibration evaluated at five step counts, with its own README | | |
| | `artifacts/*.components.pt` | the calibrated LoopQ parameters (LAS scales, shared and selected transforms, CTA) | | |
| | `artifacts/cal_a4_s0_journal_step1050.resume.pt` | the calibration journal that survived the crashed run β carries per-step loss and gradient traces | | |
| | `artifacts/curve_journals/` | nine journals of the completed run, one per step target, with its own README | | |
| | `parity/*.json` | HF-vs-vLLM prefill agreement for each deployed artifact | | |
| | `jobs/*.job.json` | the scheduler's job records: exact command lines, digests, status | | |
| | `scripts/` | the implementation snapshot that produced all of this | | |
| ## The arms | |
| | arm | what it is | | |
| |---|---| | |
| | `bf16_a16` | unquantized reference | | |
| | `rtn_w4a4_group32`, `rtn_w4a8_group32` | plain symmetric round-to-nearest, activations in groups of 32 β the granularity LoopQ itself uses (paper Table 6) | | |
| | `rtn_w4a4_per_token` | the same RTN code with **one activation scale per token** β the field-standard W4A4 baseline setting | | |
| | `loopq_w4a4_untrained` | LoopQ's structure at initialization: random orthogonal Kronecker transforms, LAS clip 1.0, CTA at exact identity, **zero optimizer updates** | | |
| | `loopq_w4a4_step0975` | LoopQ after 975 final optimization steps, recovered from the journal of a run that crashed at step 1065 | | |
| | `loopq_w4a4_step1500` | the same calibration run to its configured 1500 steps | | |
| | `loopq_w4a4_completed` | a calibration run to completion with singular values held in log space, so they cannot reach zero β **this is the arm that reproduces the paper** | | |
| ## Protocol | |
| Zero-shot everywhere except MMLU (5-shot). Plain likelihood scoring, no chat | |
| template, prompts exactly as the pinned harness produces them. HellaSwag and | |
| ARC-Challenge report `acc_norm`; WinoGrande, LAMBADA and MMLU report `acc`; | |
| WikiText reports `word_perplexity`. Calibration follows the paper: 1,024 Pile | |
| validation samples at 256 tokens, Ξ» = 0.1, KL temperature 1, teacher top-k 1000, | |
| gradient clip 1.0, SLT budget 4, CTA rank 8, activation group size 32. | |
| Serving is a vLLM 0.10.2 out-of-tree backend; quantization is simulated | |
| (quantize-dequantize) so that a 4-bit *numerical* result is measured without a | |
| 4-bit kernel. Every evaluated artifact carries a prefill parity report comparing | |
| it against the Transformers reference. | |
| ## What the runs show | |
| **The unquantized reference matches the paper on all seven measures.** | |
| | | HellaSwag | WinoGrande | LAMBADA | ARC-C | MMLU | WikiText ppl | LAMBADA ppl | | |
| |---|---|---|---|---|---|---|---| | |
| | paper BF16 | 0.7159 | 0.6717 | 0.6505 | 0.5256 | 0.6823 | 12.03 | 5.30 | | |
| | **this repo** | **0.7166** | **0.6654** | **0.6577** | **0.5239** | **0.6760** | **12.02** | **5.29** | | |
| **Activation granularity, not the method, is what separates the paper's baselines | |
| from its LoopQ row.** The paper reports every W4A4 baseline collapsing to chance. | |
| Changing one setting in the same RTN code reproduces that collapse β and leaving | |
| it at LoopQ's own group size does not. | |
| | W4A4 | HellaSwag | WinoGrande | LAMBADA | ARC-C | WikiText ppl | | |
| |---|---|---|---|---|---| | |
| | paper Symmetric / QuaRot / SpinQuant / FlatQuant | 0.27β0.29 | 0.45β0.50 | 0.00β0.03 | 0.23β0.24 | 451β10,500 | | |
| | `rtn_w4a4_per_token` | 0.2645 | 0.5043 | 0.0000 | 0.2457 | 39,089 | | |
| | `rtn_w4a4_group32` | 0.6705 | 0.6259 | 0.5719 | 0.4923 | 16.13 | | |
| **LoopQ reproduces the paper, and on four of six measures it is ahead.** | |
| `loopq_w4a4_completed` is the calibration that ran to completion without | |
| diverging; every column here is within the acceptance band registered before | |
| these numbers were measured (accuracy Β±2 pp, perplexity Β±10%). | |
| | | paper LoopQ W4A4 | `loopq_w4a4_completed` | difference | | |
| |---|---|---|---| | |
| | MMLU | 0.5465 | **0.5906** | +4.41 pp | | |
| | WinoGrande | 0.5951 | **0.6290** | +3.39 pp | | |
| | HellaSwag | 0.6582 | **0.6809** | +2.27 pp | | |
| | ARC-Challenge | 0.4881 | **0.4906** | +0.25 pp | | |
| | WikiText ppl | 15.81 | **14.3211** | β9.4% (lower is better) | | |
| | LAMBADA acc | 0.5659 | 0.5605 | β0.54 pp | | |
| | LAMBADA ppl | 8.18 | 8.3831 | +2.48% | | |
| Three accuracy columns land outside the Β±2 pp band that was registered before | |
| these numbers existed β all three above the paper. The registered failure | |
| condition (5 pp or more *below* the paper) is not met on any column, so by the | |
| letter of that rule this is a hold rather than a pass, and the rule is not being | |
| rewritten now that the numbers are in. What the columns say substantively is that | |
| this implementation loses less to quantization than the paper reports: from BF16 | |
| to W4A4, MMLU falls 8.54 pp here against the paper's 13.58 pp. | |
| **The step count is what stood between the two.** `loopq_w4a4_step0975`, taken | |
| from the diverging run, misses LAMBADA badly; the completed run does not. Same | |
| paper contract, same data, same hyperparameters β only the transform | |
| parameterization differs. | |
| | | `step0975` | `completed` | change | | |
| |---|---|---|---| | |
| | LAMBADA acc | 0.5057 | **0.5605** | +5.48 pp | | |
| | LAMBADA ppl | 11.5369 | **8.3831** | β27% | | |
| | HellaSwag | 0.6452 | **0.6809** | +3.57 pp | | |
| | WikiText ppl | 15.9149 | **14.3211** | β10% | | |
| `step0975` does reproduce MMLU to 0.01 points (0.5464 against the paper's | |
| 0.5465) and WikiText to 0.66%, which is why the failure read as one axis rather | |
| than a broken implementation. | |
| **Running the calibration to its configured 1500 steps destroys the model.** Not | |
| degrades β destroys: `loopq_w4a4_step1500` scores at or below chance on four | |
| tasks and its LAMBADA perplexity is 5.3 million. The journal in `artifacts/` | |
| records why. Across the 975 final steps the loss *rises* β median 5.42e4 over the | |
| first quarter against 1.50e5 over the third β and the transforms' smallest | |
| singular value falls from 8.3e-3 at step 1050 to 6.8e-6 at step 1575. Weight | |
| folding inverts those matrices, so an inverse that amplifies by roughly 1.5e5 | |
| lands directly in the folded weights. | |
| The paper specifies Ξ», KL temperature, teacher top-k, gradient clipping, group | |
| size, CTA rank, SLT budget, calibration sample count and sequence length. It does | |
| not specify the optimizer, the learning rate, the number of optimization steps, | |
| the batch size, the weight decay, or how the Kronecker transforms are kept | |
| invertible β it defers the last to FlatQuant. The collapse lives entirely in that | |
| second list. | |
| **The transform costs accuracy before calibration earns it back, and it does not | |
| earn back the same share on every measure.** `loopq_w4a4_untrained` isolates | |
| this: applying a random orthogonal Kronecker transform costs 4.7 points of | |
| LAMBADA perplexity relative to plain RTN before a single optimizer step. | |
| | | RTN group-32 | untrained LoopQ | step 975 | completed | paper LoopQ | | |
| |---|---|---|---|---|---| | |
| | HellaSwag | 0.6705 | 0.6300 | 0.6452 | **0.6809** | 0.6582 | | |
| | WinoGrande | 0.6259 | 0.6062 | 0.6117 | **0.6290** | 0.5951 | | |
| | LAMBADA acc | 0.5719 | 0.5015 | 0.5057 | **0.5605** | 0.5659 | | |
| | ARC-Challenge | 0.4923 | 0.4343 | 0.4744 | **0.4906** | 0.4881 | | |
| | MMLU | 0.5340 | 0.4970 | 0.5464 | **0.5906** | 0.5465 | | |
| | WikiText ppl | 16.13 | 18.57 | 15.91 | **14.3211** | 15.81 | | |
| | LAMBADA ppl | 8.23 | 12.94 | 11.54 | **8.3831** | 8.18 | | |
| Share of the untrained-to-paper gap that calibration closes, by arm: | |
| | | step 975 | completed | | |
| |---|---|---| | |
| | WikiText ppl | 96% | 154% | | |
| | HellaSwag | 54% | 181% | | |
| | LAMBADA ppl | 30% | 96% | | |
| | **LAMBADA acc** | **6.5%** | **92%** | | |
| Calibration improves every measure in both arms β so the objective was never | |
| actively harming LAMBADA. The diverging run simply gave back most of what it | |
| earned; the completed run keeps it. | |
| **The step count is a free parameter, and it moves LAMBADA.** The paper does not | |
| fix the number of optimization steps; 1,500 was our choice. | |
| `results/loopq_w4a4_step_curve/` measures the same calibration at five step | |
| counts. WikiText perplexity improves monotonically and settles (14.53 β 14.32). | |
| LAMBADA accuracy does not: it peaks at step 200 and then moves inside a 3.8 pp | |
| band with no trend β 0.5911, 0.5535, 0.5727, 0.5670, 0.5605 β which is 5.4 | |
| standard errors wide on 5,153 examples, so it is not sampling noise. The paper's | |
| 0.5659 sits inside that band. A single step count therefore cannot be read on its | |
| own as reproducing or failing to reproduce this column. | |
| **The SLT selection is stable.** `results/slt_selection_screen/` varies the | |
| Fisher estimator, the sample count and the seed. The first three of four | |
| selections are identical in every arm, and 128 calibration samples reproduce the | |
| selection made with 1,024. Only the fourth slot moves, and it moves between two | |
| seeds of the *same* estimator β a near-tie, not an unstable procedure. | |
| ## Reproducing | |
| Implementation: `scripts/` here, or the `quantization_LoopQ` branch of the | |
| working repository. Pinned stack: torch 2.8.0, vllm 0.10.2, transformers 4.56.2, | |
| lm-eval at `64f3d0924fc695efd6d776a5ac91f97138085516`. Hardware: 2 Γ NVIDIA A40. | |
| One calibration is 11β14 hours; the SLT scan is about 80% of that. A full | |
| seven-task evaluation is 4.6 hours, of which MMLU is 95%. | |