JunYoungLee's picture
Add LoopQ 4-bit quantization of Ouro-1.4B
9118991 verified
Raw History Blame Contribute Delete
5.28 kB
paper:
title: "LoopQ: Quantization for Recursive Transformers"
arxiv_id: "2605.16343v1"
local_pdf: "artifacts/papers/loopq/loopq_2605.16343v1.pdf"
lq1_quantization:
equation:
source: "Section 2, Equation (4)"
definition: "Q_a(x; c_l) = c_l * clip(round(x / c_l), q_min, q_max)"
method:
source: "Section 5.1, Implementation"
value: "symmetric uniform post-training quantization with round-to-nearest (RTN)"
evaluated_precisions:
source: "Section 5.1, Baselines"
weight_bits: 4
activation_bits: [4, 8]
granularity:
source: "Section 5.1, Implementation; Appendix B.3, Table 6"
weights: "group-wise"
activations: "group-wise"
group_size: 32
loopq_components:
shared_transformed_backbone:
source: "Section 2, Equations (2)-(3); Section 4 introduction"
contract: >-
Preserve a shared quantized backbone and fold an invertible transform
P_l into activation and weight sides.
las:
source: "Section 4.1"
contract: "Replace shared activation scale c_l with c_{t,l}, one per module and loop."
slt:
source: "Section 4.2, Equations (5)-(6); Appendix B.1, Table 4"
contract: >-
Progressively select transform-weight groups by curvature-normalized
cross-loop gradient variance; selected groups receive loop-dependent
transforms and weights.
cta:
source: "Section 4.3, Equation (7)"
contract: >-
Apply an affine plus gated low-rank adapter only at loop transitions;
U and V are shared, while a_t, b_t, and eta_t are loop-dependent.
trajectory_calibration:
source: "Section 4.4, Equation (8); Appendix B.4"
contract: >-
Joint KL, same-loop teacher state, adaptive final-state guidance, and
transition-alignment loss over the true recurrent trajectory.
experiment_hyperparameters:
source: "Appendix B.3, Table 6"
base_model_dtype: bf16
calibration_dataset: "mit-han-lab/pile-val-backup"
calibration_split: validation
calibration_field: text
calibration_samples: 1024
maximum_sequence_length: 256
las_activation_scales: "per module, per loop"
slt_budget:
ouro_1_4b: 4
ouro_2_6b: 8
loopformer: 1
parcae: 1
cta_rank: 8
trajectory_loss_lambda: 0.1
kl_temperature: 1
teacher_top_k_logits: 1000
gradient_clipping: 1.0
unquantized_modules:
- lm_head
- token_embeddings
- position_embeddings
- architecture_specific_time_or_modulation_modules
paper_underspecified_items:
source: "Absence verified in arXiv:2605.16343v1 Sections 2, 4, 5.1 and Appendices B.1-B.4"
items:
- integer_bounds
- rtn_tie_breaking
- initial_scale_estimator
- group_axis_and_tail_layout
- optimizer
- learning_rate
- total_calibration_steps
- fisher_estimator_details
- epsilon_values
- progressive_slt_update_schedule
- flatquant_kronecker_factorization_details
- adaptive_mu_update_frequency
implementation_status:
lq0: >-
Paper contracts and known omissions are recorded here. Each later module
must cite the matching source and add any new assumption to the decision log.
lq1: "Implemented in loopq/quantization.py with executable hand-computed tests."
lq2: >-
Implemented in loopq/las.py: positive per-module/per-loop clipping scalars
applied to dynamic token/group absmax scales, explicit routing and export.
Feature-group static LAS is not the active Ouro paper profile.
lq3: >-
Minimal shared Kronecker transform and exact no-quant inverse-transpose
weight folding implemented in loopq/transforms.py. Calibration
uses an explicit local SVD parameterization for Ouro. Exact author factors
and the required sensitivity evidence remain unresolved.
lq4: >-
Equation (6) population-variance scorer and callback-driven progressive
atomic-group selection implemented in loopq/sharing_gap.py, including Ouro
count budget and explicit Huginn parameter-fraction rule.
lq5: >-
Equation (7) CTA implemented in loopq/cta.py with shared U,V, exact
transition-dependent affine/gate parameters, identity initialization, and
exact Ouro/Huginn transition counts.
lq6: >-
Both real drivers implement the recurrent objective, frozen-backbone
allowlist, progressive scans and atomic checkpoint/resume. A real Ouro
interrupted/resumed run exactly matches continuous component tensors,
selection and loss traces. Full calibration and sensitivity are pending.
lq7: >-
Ouro calibrated smoke exports load and generate in vLLM. The bridge uses
BF16 QDQ and eager selected-weight dispatch. Trained full-trajectory parity,
full calibration and packed-INT4 deployment equivalence remain pending.
lq8: >-
Huginn 32-recurrence calibration/export/runtime smoke is validated. Full
calibration, real resume parity and trained full-trajectory parity remain
pending. Huginn is an architecture extension absent from the paper.
lq9_through_lq10: >-
Both architectures' GSM8K-128 BF16/direct controls are complete. Retraining
ablations and export modes are implemented but require GPU validation.
Huginn full controls are running; calibrated full LoopQ evaluation remains
pending. Smoke/study/ablation artifacts cannot masquerade as full LoopQ rows.