JunYoungLee's picture
Archive compute node 1 step-curve results (stage 3)
85c74af verified
|
Raw History Blame Contribute Delete
10.1 kB

LoopQ 4-bit quantization of Ouro-1.4B β€” raw artifacts

A reproduction of LoopQ (arXiv:2605.16343v1), post-training quantization designed for depth-recurrent transformers, on Ouro-1.4B (ByteDance/Ouro-1.4B, revision e3b1e0993b1231a51d6069a870476dda4162c00a) β€” 24 layers run 4 times per token.

Everything here is raw output. No number below is copied from a write-up; each comes from results/**/harness/**/results_*.json produced by lm-eval at the pinned commit 64f3d0924fc695efd6d776a5ac91f97138085516.

Layout

path what
results/<arm>/<task>/harness/ lm-eval output per task: results_*.json, samples_*.jsonl, installed_task_snapshot
results/<arm>/<task>/run_metadata.json, timing.json protocol note, pinned versions, wall clock
results/slt_selection_screen/ the four selection arms of the SLT stability screen
results/loopq_w4a4_step_curve/ the completed calibration evaluated at five step counts, with its own README
artifacts/*.components.pt the calibrated LoopQ parameters (LAS scales, shared and selected transforms, CTA)
artifacts/cal_a4_s0_journal_step1050.resume.pt the calibration journal that survived the crashed run β€” carries per-step loss and gradient traces
artifacts/curve_journals/ nine journals of the completed run, one per step target, with its own README
parity/*.json HF-vs-vLLM prefill agreement for each deployed artifact
jobs/*.job.json the scheduler's job records: exact command lines, digests, status
scripts/ the implementation snapshot that produced all of this

The arms

arm what it is
bf16_a16 unquantized reference
rtn_w4a4_group32, rtn_w4a8_group32 plain symmetric round-to-nearest, activations in groups of 32 β€” the granularity LoopQ itself uses (paper Table 6)
rtn_w4a4_per_token the same RTN code with one activation scale per token β€” the field-standard W4A4 baseline setting
loopq_w4a4_untrained LoopQ's structure at initialization: random orthogonal Kronecker transforms, LAS clip 1.0, CTA at exact identity, zero optimizer updates
loopq_w4a4_step0975 LoopQ after 975 final optimization steps, recovered from the journal of a run that crashed at step 1065
loopq_w4a4_step1500 the same calibration run to its configured 1500 steps
loopq_w4a4_completed a calibration run to completion with singular values held in log space, so they cannot reach zero β€” this is the arm that reproduces the paper

Protocol

Zero-shot everywhere except MMLU (5-shot). Plain likelihood scoring, no chat template, prompts exactly as the pinned harness produces them. HellaSwag and ARC-Challenge report acc_norm; WinoGrande, LAMBADA and MMLU report acc; WikiText reports word_perplexity. Calibration follows the paper: 1,024 Pile validation samples at 256 tokens, Ξ» = 0.1, KL temperature 1, teacher top-k 1000, gradient clip 1.0, SLT budget 4, CTA rank 8, activation group size 32.

Serving is a vLLM 0.10.2 out-of-tree backend; quantization is simulated (quantize-dequantize) so that a 4-bit numerical result is measured without a 4-bit kernel. Every evaluated artifact carries a prefill parity report comparing it against the Transformers reference.

What the runs show

The unquantized reference matches the paper on all seven measures.

HellaSwag WinoGrande LAMBADA ARC-C MMLU WikiText ppl LAMBADA ppl
paper BF16 0.7159 0.6717 0.6505 0.5256 0.6823 12.03 5.30
this repo 0.7166 0.6654 0.6577 0.5239 0.6760 12.02 5.29

Activation granularity, not the method, is what separates the paper's baselines from its LoopQ row. The paper reports every W4A4 baseline collapsing to chance. Changing one setting in the same RTN code reproduces that collapse β€” and leaving it at LoopQ's own group size does not.

W4A4 HellaSwag WinoGrande LAMBADA ARC-C WikiText ppl
paper Symmetric / QuaRot / SpinQuant / FlatQuant 0.27–0.29 0.45–0.50 0.00–0.03 0.23–0.24 451–10,500
rtn_w4a4_per_token 0.2645 0.5043 0.0000 0.2457 39,089
rtn_w4a4_group32 0.6705 0.6259 0.5719 0.4923 16.13

LoopQ reproduces the paper, and on four of six measures it is ahead. loopq_w4a4_completed is the calibration that ran to completion without diverging; every column here is within the acceptance band registered before these numbers were measured (accuracy Β±2 pp, perplexity Β±10%).

paper LoopQ W4A4 loopq_w4a4_completed difference
MMLU 0.5465 0.5906 +4.41 pp
WinoGrande 0.5951 0.6290 +3.39 pp
HellaSwag 0.6582 0.6809 +2.27 pp
ARC-Challenge 0.4881 0.4906 +0.25 pp
WikiText ppl 15.81 14.3211 βˆ’9.4% (lower is better)
LAMBADA acc 0.5659 0.5605 βˆ’0.54 pp
LAMBADA ppl 8.18 8.3831 +2.48%

Three accuracy columns land outside the Β±2 pp band that was registered before these numbers existed β€” all three above the paper. The registered failure condition (5 pp or more below the paper) is not met on any column, so by the letter of that rule this is a hold rather than a pass, and the rule is not being rewritten now that the numbers are in. What the columns say substantively is that this implementation loses less to quantization than the paper reports: from BF16 to W4A4, MMLU falls 8.54 pp here against the paper's 13.58 pp.

The step count is what stood between the two. loopq_w4a4_step0975, taken from the diverging run, misses LAMBADA badly; the completed run does not. Same paper contract, same data, same hyperparameters β€” only the transform parameterization differs.

step0975 completed change
LAMBADA acc 0.5057 0.5605 +5.48 pp
LAMBADA ppl 11.5369 8.3831 βˆ’27%
HellaSwag 0.6452 0.6809 +3.57 pp
WikiText ppl 15.9149 14.3211 βˆ’10%

step0975 does reproduce MMLU to 0.01 points (0.5464 against the paper's 0.5465) and WikiText to 0.66%, which is why the failure read as one axis rather than a broken implementation.

Running the calibration to its configured 1500 steps destroys the model. Not degrades β€” destroys: loopq_w4a4_step1500 scores at or below chance on four tasks and its LAMBADA perplexity is 5.3 million. The journal in artifacts/ records why. Across the 975 final steps the loss rises β€” median 5.42e4 over the first quarter against 1.50e5 over the third β€” and the transforms' smallest singular value falls from 8.3e-3 at step 1050 to 6.8e-6 at step 1575. Weight folding inverts those matrices, so an inverse that amplifies by roughly 1.5e5 lands directly in the folded weights.

The paper specifies Ξ», KL temperature, teacher top-k, gradient clipping, group size, CTA rank, SLT budget, calibration sample count and sequence length. It does not specify the optimizer, the learning rate, the number of optimization steps, the batch size, the weight decay, or how the Kronecker transforms are kept invertible β€” it defers the last to FlatQuant. The collapse lives entirely in that second list.

The transform costs accuracy before calibration earns it back, and it does not earn back the same share on every measure. loopq_w4a4_untrained isolates this: applying a random orthogonal Kronecker transform costs 4.7 points of LAMBADA perplexity relative to plain RTN before a single optimizer step.

RTN group-32 untrained LoopQ step 975 completed paper LoopQ
HellaSwag 0.6705 0.6300 0.6452 0.6809 0.6582
WinoGrande 0.6259 0.6062 0.6117 0.6290 0.5951
LAMBADA acc 0.5719 0.5015 0.5057 0.5605 0.5659
ARC-Challenge 0.4923 0.4343 0.4744 0.4906 0.4881
MMLU 0.5340 0.4970 0.5464 0.5906 0.5465
WikiText ppl 16.13 18.57 15.91 14.3211 15.81
LAMBADA ppl 8.23 12.94 11.54 8.3831 8.18

Share of the untrained-to-paper gap that calibration closes, by arm:

step 975 completed
WikiText ppl 96% 154%
HellaSwag 54% 181%
LAMBADA ppl 30% 96%
LAMBADA acc 6.5% 92%

Calibration improves every measure in both arms β€” so the objective was never actively harming LAMBADA. The diverging run simply gave back most of what it earned; the completed run keeps it.

The step count is a free parameter, and it moves LAMBADA. The paper does not fix the number of optimization steps; 1,500 was our choice. results/loopq_w4a4_step_curve/ measures the same calibration at five step counts. WikiText perplexity improves monotonically and settles (14.53 β†’ 14.32). LAMBADA accuracy does not: it peaks at step 200 and then moves inside a 3.8 pp band with no trend β€” 0.5911, 0.5535, 0.5727, 0.5670, 0.5605 β€” which is 5.4 standard errors wide on 5,153 examples, so it is not sampling noise. The paper's 0.5659 sits inside that band. A single step count therefore cannot be read on its own as reproducing or failing to reproduce this column.

The SLT selection is stable. results/slt_selection_screen/ varies the Fisher estimator, the sample count and the seed. The first three of four selections are identical in every arm, and 128 calibration samples reproduce the selection made with 1,024. Only the fourth slot moves, and it moves between two seeds of the same estimator β€” a near-tie, not an unstable procedure.

Reproducing

Implementation: scripts/ here, or the quantization_LoopQ branch of the working repository. Pinned stack: torch 2.8.0, vllm 0.10.2, transformers 4.56.2, lm-eval at 64f3d0924fc695efd6d776a5ac91f97138085516. Hardware: 2 Γ— NVIDIA A40. One calibration is 11–14 hours; the SLT scan is about 80% of that. A full seven-task evaluation is 4.6 hours, of which MMLU is 95%.