# LoopQ 4-bit quantization of Ouro-1.4B — raw artifacts A reproduction of **LoopQ** ([arXiv:2605.16343v1](https://arxiv.org/abs/2605.16343)), post-training quantization designed for *depth-recurrent* transformers, on **Ouro-1.4B** (`ByteDance/Ouro-1.4B`, revision `e3b1e0993b1231a51d6069a870476dda4162c00a`) — 24 layers run 4 times per token. Everything here is raw output. No number below is copied from a write-up; each comes from `results/**/harness/**/results_*.json` produced by lm-eval at the pinned commit `64f3d0924fc695efd6d776a5ac91f97138085516`. ## Layout | path | what | |---|---| | `results///harness/` | lm-eval output per task: `results_*.json`, `samples_*.jsonl`, `installed_task_snapshot` | | `results///run_metadata.json`, `timing.json` | protocol note, pinned versions, wall clock | | `results/slt_selection_screen/` | the four selection arms of the SLT stability screen | | `results/loopq_w4a4_step_curve/` | the completed calibration evaluated at five step counts, with its own README | | `artifacts/*.components.pt` | the calibrated LoopQ parameters (LAS scales, shared and selected transforms, CTA) | | `artifacts/cal_a4_s0_journal_step1050.resume.pt` | the calibration journal that survived the crashed run — carries per-step loss and gradient traces | | `artifacts/curve_journals/` | nine journals of the completed run, one per step target, with its own README | | `parity/*.json` | HF-vs-vLLM prefill agreement for each deployed artifact | | `jobs/*.job.json` | the scheduler's job records: exact command lines, digests, status | | `scripts/` | the implementation snapshot that produced all of this | ## The arms | arm | what it is | |---|---| | `bf16_a16` | unquantized reference | | `rtn_w4a4_group32`, `rtn_w4a8_group32` | plain symmetric round-to-nearest, activations in groups of 32 — the granularity LoopQ itself uses (paper Table 6) | | `rtn_w4a4_per_token` | the same RTN code with **one activation scale per token** — the field-standard W4A4 baseline setting | | `loopq_w4a4_untrained` | LoopQ's structure at initialization: random orthogonal Kronecker transforms, LAS clip 1.0, CTA at exact identity, **zero optimizer updates** | | `loopq_w4a4_step0975` | LoopQ after 975 final optimization steps, recovered from the journal of a run that crashed at step 1065 | | `loopq_w4a4_step1500` | the same calibration run to its configured 1500 steps | | `loopq_w4a4_completed` | a calibration run to completion with singular values held in log space, so they cannot reach zero — **this is the arm that reproduces the paper** | ## Protocol Zero-shot everywhere except MMLU (5-shot). Plain likelihood scoring, no chat template, prompts exactly as the pinned harness produces them. HellaSwag and ARC-Challenge report `acc_norm`; WinoGrande, LAMBADA and MMLU report `acc`; WikiText reports `word_perplexity`. Calibration follows the paper: 1,024 Pile validation samples at 256 tokens, λ = 0.1, KL temperature 1, teacher top-k 1000, gradient clip 1.0, SLT budget 4, CTA rank 8, activation group size 32. Serving is a vLLM 0.10.2 out-of-tree backend; quantization is simulated (quantize-dequantize) so that a 4-bit *numerical* result is measured without a 4-bit kernel. Every evaluated artifact carries a prefill parity report comparing it against the Transformers reference. ## What the runs show **The unquantized reference matches the paper on all seven measures.** | | HellaSwag | WinoGrande | LAMBADA | ARC-C | MMLU | WikiText ppl | LAMBADA ppl | |---|---|---|---|---|---|---|---| | paper BF16 | 0.7159 | 0.6717 | 0.6505 | 0.5256 | 0.6823 | 12.03 | 5.30 | | **this repo** | **0.7166** | **0.6654** | **0.6577** | **0.5239** | **0.6760** | **12.02** | **5.29** | **Activation granularity, not the method, is what separates the paper's baselines from its LoopQ row.** The paper reports every W4A4 baseline collapsing to chance. Changing one setting in the same RTN code reproduces that collapse — and leaving it at LoopQ's own group size does not. | W4A4 | HellaSwag | WinoGrande | LAMBADA | ARC-C | WikiText ppl | |---|---|---|---|---|---| | paper Symmetric / QuaRot / SpinQuant / FlatQuant | 0.27–0.29 | 0.45–0.50 | 0.00–0.03 | 0.23–0.24 | 451–10,500 | | `rtn_w4a4_per_token` | 0.2645 | 0.5043 | 0.0000 | 0.2457 | 39,089 | | `rtn_w4a4_group32` | 0.6705 | 0.6259 | 0.5719 | 0.4923 | 16.13 | **LoopQ reproduces the paper, and on four of six measures it is ahead.** `loopq_w4a4_completed` is the calibration that ran to completion without diverging; every column here is within the acceptance band registered before these numbers were measured (accuracy ±2 pp, perplexity ±10%). | | paper LoopQ W4A4 | `loopq_w4a4_completed` | difference | |---|---|---|---| | MMLU | 0.5465 | **0.5906** | +4.41 pp | | WinoGrande | 0.5951 | **0.6290** | +3.39 pp | | HellaSwag | 0.6582 | **0.6809** | +2.27 pp | | ARC-Challenge | 0.4881 | **0.4906** | +0.25 pp | | WikiText ppl | 15.81 | **14.3211** | −9.4% (lower is better) | | LAMBADA acc | 0.5659 | 0.5605 | −0.54 pp | | LAMBADA ppl | 8.18 | 8.3831 | +2.48% | Three accuracy columns land outside the ±2 pp band that was registered before these numbers existed — all three above the paper. The registered failure condition (5 pp or more *below* the paper) is not met on any column, so by the letter of that rule this is a hold rather than a pass, and the rule is not being rewritten now that the numbers are in. What the columns say substantively is that this implementation loses less to quantization than the paper reports: from BF16 to W4A4, MMLU falls 8.54 pp here against the paper's 13.58 pp. **The step count is what stood between the two.** `loopq_w4a4_step0975`, taken from the diverging run, misses LAMBADA badly; the completed run does not. Same paper contract, same data, same hyperparameters — only the transform parameterization differs. | | `step0975` | `completed` | change | |---|---|---|---| | LAMBADA acc | 0.5057 | **0.5605** | +5.48 pp | | LAMBADA ppl | 11.5369 | **8.3831** | −27% | | HellaSwag | 0.6452 | **0.6809** | +3.57 pp | | WikiText ppl | 15.9149 | **14.3211** | −10% | `step0975` does reproduce MMLU to 0.01 points (0.5464 against the paper's 0.5465) and WikiText to 0.66%, which is why the failure read as one axis rather than a broken implementation. **Running the calibration to its configured 1500 steps destroys the model.** Not degrades — destroys: `loopq_w4a4_step1500` scores at or below chance on four tasks and its LAMBADA perplexity is 5.3 million. The journal in `artifacts/` records why. Across the 975 final steps the loss *rises* — median 5.42e4 over the first quarter against 1.50e5 over the third — and the transforms' smallest singular value falls from 8.3e-3 at step 1050 to 6.8e-6 at step 1575. Weight folding inverts those matrices, so an inverse that amplifies by roughly 1.5e5 lands directly in the folded weights. The paper specifies λ, KL temperature, teacher top-k, gradient clipping, group size, CTA rank, SLT budget, calibration sample count and sequence length. It does not specify the optimizer, the learning rate, the number of optimization steps, the batch size, the weight decay, or how the Kronecker transforms are kept invertible — it defers the last to FlatQuant. The collapse lives entirely in that second list. **The transform costs accuracy before calibration earns it back, and it does not earn back the same share on every measure.** `loopq_w4a4_untrained` isolates this: applying a random orthogonal Kronecker transform costs 4.7 points of LAMBADA perplexity relative to plain RTN before a single optimizer step. | | RTN group-32 | untrained LoopQ | step 975 | completed | paper LoopQ | |---|---|---|---|---|---| | HellaSwag | 0.6705 | 0.6300 | 0.6452 | **0.6809** | 0.6582 | | WinoGrande | 0.6259 | 0.6062 | 0.6117 | **0.6290** | 0.5951 | | LAMBADA acc | 0.5719 | 0.5015 | 0.5057 | **0.5605** | 0.5659 | | ARC-Challenge | 0.4923 | 0.4343 | 0.4744 | **0.4906** | 0.4881 | | MMLU | 0.5340 | 0.4970 | 0.5464 | **0.5906** | 0.5465 | | WikiText ppl | 16.13 | 18.57 | 15.91 | **14.3211** | 15.81 | | LAMBADA ppl | 8.23 | 12.94 | 11.54 | **8.3831** | 8.18 | Share of the untrained-to-paper gap that calibration closes, by arm: | | step 975 | completed | |---|---|---| | WikiText ppl | 96% | 154% | | HellaSwag | 54% | 181% | | LAMBADA ppl | 30% | 96% | | **LAMBADA acc** | **6.5%** | **92%** | Calibration improves every measure in both arms — so the objective was never actively harming LAMBADA. The diverging run simply gave back most of what it earned; the completed run keeps it. **The step count is a free parameter, and it moves LAMBADA.** The paper does not fix the number of optimization steps; 1,500 was our choice. `results/loopq_w4a4_step_curve/` measures the same calibration at five step counts. WikiText perplexity improves monotonically and settles (14.53 → 14.32). LAMBADA accuracy does not: it peaks at step 200 and then moves inside a 3.8 pp band with no trend — 0.5911, 0.5535, 0.5727, 0.5670, 0.5605 — which is 5.4 standard errors wide on 5,153 examples, so it is not sampling noise. The paper's 0.5659 sits inside that band. A single step count therefore cannot be read on its own as reproducing or failing to reproduce this column. **The SLT selection is stable.** `results/slt_selection_screen/` varies the Fisher estimator, the sample count and the seed. The first three of four selections are identical in every arm, and 128 calibration samples reproduce the selection made with 1,024. Only the fourth slot moves, and it moves between two seeds of the *same* estimator — a near-tie, not an unstable procedure. ## Reproducing Implementation: `scripts/` here, or the `quantization_LoopQ` branch of the working repository. Pinned stack: torch 2.8.0, vllm 0.10.2, transformers 4.56.2, lm-eval at `64f3d0924fc695efd6d776a5ac91f97138085516`. Hardware: 2 × NVIDIA A40. One calibration is 11–14 hours; the SLT scan is about 80% of that. A full seven-task evaluation is 4.6 hours, of which MMLU is 95%.