ut-depth-probe-artifacts / loopq_quantization /scripts /configs /loopq_compatibility_decisions.yaml
JunYoungLee's picture
Add LoopQ 4-bit quantization of Ouro-1.4B
9118991 verified
Raw History Blame Contribute Delete
27.5 kB
schema_version: 1
paper: "arXiv:2605.16343v1"
decisions:
- id: LQ1_INTEGER_RANGE
status: assumption
paper_status: >-
Equation (4) names q_min and q_max but does not give their numerical
values or state whether the most-negative two's-complement code is used.
decision: >-
Use the cited FlatQuant symmetric RTN convention: q_min = -2^(b-1) and
q_max = 2^(b-1)-1 ([-8, 7] for 4-bit and [-128, 127] for 8-bit).
sensitivity_required: true
- id: LQ1_SCALE_INITIALIZATION
status: assumption
paper_status: >-
The paper learns activation ranges during calibration but does not state
the initial static RTN scale estimator for weights or activations.
decision: >-
Initialize each group scale as max(abs(group)) / q_max. An all-zero
group receives scale 1 and therefore remains exactly zero.
sensitivity_required: true
- id: LQ1_ROUNDING_TIES
status: assumption
paper_status: "The paper says round-to-nearest but does not specify tie handling."
decision: "Round exact half-way values away from zero deterministically."
sensitivity_required: false
- id: LQ1_GROUP_AXIS
status: assumption
paper_status: "The paper specifies group-wise quantization but not tensor-axis layout."
decision: >-
Form consecutive groups along the last logical dimension for both
weights and activations. Architecture adapters must document any layout
conversion before invoking this quantizer.
sensitivity_required: true
- id: LQ1_TAIL_GROUP
status: compatibility_decision
paper_status: "Tail handling is not specified."
decision: >-
Quantize a final group shorter than 32 independently, without persistent
padding. Its scale is computed only from real elements.
sensitivity_required: false
- id: LQ1_FAKE_QUANT_SCOPE
status: compatibility_decision
paper_status: >-
The paper reports packed weights and quantized runtime kernels, but LQ1
is the numerical quantizer acceptance phase.
decision: >-
LQ1 implements deterministic QDQ/fake quantization and emits integer
codes plus scales. Packing and runtime kernels are deferred to LQ7/LQ8.
sensitivity_required: false
- id: LQ2_GROUP_SCALE_LAYOUT
status: compatibility_decision
paper_status: >-
Section 4.1 specifies c_{t,l} per module and loop and explicitly counts
O(TL) additional scalars. Appendix B.3 says "Per module, per loop".
decision: >-
Store one positive clipping multiplier for each module and loop. Apply it
to dynamic per-token, per-group absmax scales (group size 32), matching
the FlatQuant-style learned activation-clipping construction cited by the
paper. This gives exactly O(TL) LAS parameters.
sensitivity_required: false
- id: LQ2_CALIBRATION_REDUCTION
status: assumption
paper_status: >-
The initialization procedure for learned LAS ranges is not specified.
decision: >-
Initialize each module-loop clipping multiplier to 1.0. Dynamic base
scales use each input token/group absmax. Subsequent trajectory
calibration optimizes the log multiplier.
sensitivity_required: true
- id: LQ2_POSITIVE_PARAMETERIZATION
status: compatibility_decision
paper_status: "The paper requires scales but does not specify their parameterization."
decision: >-
Store trainable log-scales and exponentiate at use time, ensuring scales
remain strictly positive during later gradient calibration.
sensitivity_required: false
- id: LQ2_ARCHITECTURE_LOOP_ROUTING
status: compatibility_decision
paper_status: >-
The paper evaluates Ouro with four loops and does not evaluate Huginn.
decision: >-
Provide exact zero-based routing for Ouro loops [0,4) and Huginn
recurrences [0,32), with no modulo, anchor, or fallback routing.
sensitivity_required: false
- id: LQ3_FLATQUANT_REFERENCE_PIN
status: compatibility_decision
paper_status: >-
LoopQ cites FlatQuant and states that transforms are Kronecker-decomposed,
but does not identify a code revision.
decision: >-
Pin the official ruikangliu/FlatQuant repository at commit
9d88ffcb7d2c6bda59fb5c44dad36adc101aadb1. File hashes and the inspected
folding contract are recorded in flatquant_reference.yaml.
sensitivity_required: false
- id: LQ3_FLATQUANT_SVD_PARAMETERIZATION
status: compatibility_decision
paper_status: >-
LoopQ specifies FlatQuant-style Kronecker-decomposed transforms but does
not state whether it uses FlatQuant's SVD or direct-inverse class, factor
dimensions, add-diagonal flag, or architecture-specific initialization.
decision: >-
Use the default path of the paper-pinned official FlatQuant commit:
SVDDecomposeTransMatrix with independently random-orthogonal U/V factors,
trainable singular diagonals, and Cayley orthogonal parameterizations.
Use explicit factor dimensions and omit FlatQuant's optional add_diag
because LoopQ does not state that flag or its SmoothQuant initialization.
Export effective Kronecker factors so runtime is independent of the
training parameterization; retain legacy identity transform artifacts.
sensitivity_required: true
- id: LQ3_CALIBRATION_BLOCKER
status: assumption
paper_status: >-
LoopQ omits optimizer, learning rate, epochs/steps, Kronecker factor
dimensions, inverse method, and diagonal initialization. FlatQuant's
official code uses AdamW and cosine scheduling, but repository defaults
and published scripts differ and its layer-wise MSE objective is not
LoopQ's trajectory objective.
decision: >-
Do not import FlatQuant's optimizer settings into LoopQ. Defer transform
calibration parameterization and optimizer selection to LQ6, record the
chosen sensitivity study there, and keep LQ3 limited to fold parity.
sensitivity_required: true
- id: LQ4_VARIANCE_ESTIMATOR
status: compatibility_decision
paper_status: >-
Equation (6) explicitly expands variance with a 1/T factor but does not
discuss library estimator settings.
decision: >-
Use population variance across the true loop dimension (unbiased=false),
accumulate the score in deterministic CPU FP64, and sum every coordinate
in the transform-weight group.
sensitivity_required: false
- id: LQ4_FISHER_AND_EPSILON
status: assumption
paper_status: >-
The paper identifies psi as a diagonal Fisher estimate and epsilon as a
positive stabilizer, but does not specify the Fisher sampling estimator,
reduction, or epsilon value.
decision: >-
LQ4 consumes non-negative saved Fisher diagonals without inventing their
collection procedure and uses epsilon=1e-8 by default. Fisher collection
must be resolved and tested with trajectory calibration in LQ6.
sensitivity_required: true
- id: LQ4_PROGRESSIVE_RECOMPUTATION
status: compatibility_decision
paper_status: >-
Section 4.2 requires alternation between adapting selected transforms and
recomputing scores, but does not specify iteration counts or tie handling.
decision: >-
Select exactly one atomic transform-weight group per round, require fresh
callback statistics after each selection, and break equal-score ties by
lexical candidate name for deterministic artifacts.
sensitivity_required: false
- id: LQ4_HUGINN_BUDGET_ROUNDING
status: assumption
paper_status: >-
Huginn is not evaluated. LoopQ reports that about 5% of transform
parameters are selected, without a rounding rule for unequal atomic groups.
decision: >-
Set the Huginn parameter target to ceil(0.05 * all candidate transform
parameters), select whole groups progressively, and stop after the first
selection whose cumulative count meets or exceeds the target. Record the
resulting atomic-group overshoot.
sensitivity_required: true
- id: LQ5_RMSNORM_DEFINITION
status: assumption
paper_status: >-
Equation (7) says RMSNorm along the feature dimension but does not give
epsilon or specify reuse of a backbone norm's learned weight.
decision: >-
Use unweighted RMS normalization x / sqrt(mean(x^2) + 1e-6), computed in
FP32 and returned in the hidden-state dtype. Evaluate epsilon sensitivity
before full calibration.
sensitivity_required: true
- id: LQ5_PARAMETER_SHAPES
status: compatibility_decision
paper_status: >-
Equation (7) defines shared U,V and loop-dependent a_t,b_t,eta_t but only
specifies adapter rank 8 in Table 6.
decision: >-
For hidden width d and rank r, use U,V in R^(d x r), a_t,b_t in R^d,
and eta_t in R^r. Store parameters for actual transitions only: T-1 rows.
sensitivity_required: false
- id: LQ5_IDENTITY_INITIALIZATION
status: assumption
paper_status: "The paper does not state CTA initialization."
decision: >-
Initialize a_t=1, b_t=0, eta_t=0. Initialize shared U and V to the first
r columns of the identity for deterministic nonzero gate gradients. This
makes every CTA exactly identity before calibration.
sensitivity_required: true
- id: LQ5_TRANSITION_ROUTING
status: compatibility_decision
paper_status: >-
Equation (7) applies A_t between loop t and t+1. Ouro has four loops;
Huginn is not evaluated by the paper.
decision: >-
Route exactly three Ouro transitions and thirty-one Huginn transitions
with zero-based indices. Reject out-of-range indices; never use modulo,
Residual/Momentum anchors, or mixed-precision schedule boundaries.
sensitivity_required: false
- id: LQ6_LOGIT_TOPK_KL
status: compatibility_decision
paper_status: >-
Table 6 specifies teacher top-k logits 1000 and temperature 1, but does
not state whether KL is renormalized over the retained support or how
reductions are applied.
decision: >-
Select the teacher's top-1000 indices, renormalize teacher and student
distributions on that same support, and compute KL(teacher || student)
with PyTorch batchmean reduction. For synthetic vocabularies smaller than
1000 only, use the whole vocabulary.
sensitivity_required: true
- id: LQ6_TRAJECTORY_REDUCTION
status: assumption
paper_status: >-
Equation (8) shows squared L2 norms and sums over loops but does not give
batch/token reduction or normalization details.
decision: >-
Use squared L2 sums across every non-loop coordinate and sum across loops
and transitions exactly as displayed. Apply lambda=0.1 to all trajectory
and transition terms. Record each weighted term separately.
sensitivity_required: true
- id: LQ6_ADAPTIVE_MU_REFRESH
status: compatibility_decision
paper_status: >-
Appendix B.4 states that mu_t is updated periodically, for example every
100 calibration steps; this is illustrative rather than an immutable value.
decision: >-
Default to the paper's example interval 100, cache detached mu values
between refreshes, serialize the cache, and expose the interval in the
required calibration config for sensitivity analysis.
sensitivity_required: true
- id: LQ6_OPTIMIZER_SCHEDULE_STEPS
status: assumption
paper_status: >-
LoopQ does not specify optimizer, learning rate, total calibration steps,
weight decay, warmup, or scheduler. FlatQuant's settings use a different
layer-wise reconstruction objective and cannot be silently inherited.
decision: >-
Require optimizer name, base learning rate, and final-phase steps explicitly
with no defaults. Permit explicit LAS, transform, and CTA learning-rate
overrides derived from the exact parameter allowlist so omitted component
rates can be studied without changing the objective. Initially permit Adam
or AdamW as caller-selected alternatives. Record both final-phase and global
optimizer-step counts in every checkpoint, run sensitivity before claiming
reproduction fidelity, and do not add an implicit scheduler.
sensitivity_required: true
- id: LQ6_PARAMETER_ALLOWLIST
status: compatibility_decision
paper_status: >-
Section 4.4 allows shared transforms, LAS, selected loop transforms, and
CTA parameters to update while the shared backbone remains fixed.
decision: >-
Construct an exact optimizer allowlist from only those four categories,
reject duplicated ownership or any identity overlap with backbone
parameters, verify the optimizer on every step, and clip their joint norm
to the paper value 1.0.
sensitivity_required: false
- id: LQ7_OURO_CHECKPOINT_PIN
status: compatibility_decision
paper_status: >-
The paper names Ouro 1.4B but does not provide a checkpoint revision or
hash in the implementation appendix.
decision: >-
Validate against local ByteDance/Ouro-1.4B revision
e3b1e0993b1231a51d6069a870476dda4162c00a. Pin config and modeling-code
SHA256 values in ouro_mapping.yaml and inspect safetensors keys without
loading parameter tensors.
sensitivity_required: false
- id: LQ7_VLLM_PACKED_MAPPING
status: compatibility_decision
paper_status: >-
Appendix B.1 lists separate HF projection names, whereas the local vLLM
backend packs Q/K/V and Gate/Up projections.
decision: >-
Preserve the four paper atomic groups while mapping Q/K/V to qkv_proj and
Gate/Up to gate_up_proj consumers. Artifact provenance retains both the
original HF weight names and packed runtime names.
sensitivity_required: false
- id: LQ7_RUNTIME_INTEGRATION_BOUNDARY
status: compatibility_decision
paper_status: "The paper does not specify vLLM hook APIs."
decision: >-
Keep the adapter backend-neutral: expose transform/LAS/QDQ preparation and
exact CTA boundaries, but do not import the existing NVFP4/Residual plugin.
Runtime model.py must later call these hooks at four listed consumers and
three true loop boundaries before a smoke result can be claimed.
sensitivity_required: false
- id: LQ7_SELECTED_SLT_PACKED_WEIGHT_BLOCKER
status: compatibility_decision
paper_status: >-
Appendix B.1 requires selected groups to use loop-dependent transforms
and quantized weights. The local vLLM projections own one packed weight
parameter shared by all recurrent calls.
decision: >-
Keep one shared base parameter for every projection and allocate four
loop-specific packed RTN variants only for the paper-budget four selected
groups. Merge Q/K/V and Gate/Up source shards with vLLM's existing weight
loaders, then dispatch by exact loop index and restore the shared pointer.
This correctness path requires eager execution.
sensitivity_required: false
- id: LQ7_SELECTED_WEIGHT_EAGER_DISPATCH
status: compatibility_decision
paper_status: "The paper does not specify vLLM compiled packed-weight dispatch."
decision: >-
Use scoped parameter-data substitution only in the explicit eager smoke
path. Do not claim compiled/runtime speed results. A future native kernel
or quant-method weight argument is required before torch.compile use.
sensitivity_required: false
- id: LQ7_CALIBRATION_STE
status: compatibility_decision
paper_status: >-
The paper optimizes quantization parameters through RTN but does not
specify the rounding surrogate used by its calibration implementation.
decision: >-
Use quantized values in the forward pass. For activation QDQ, retain the
differentiable learned-LAS scale path and add an identity input path
around rounding. For weight QDQ, RTN scales are inferred rather than
learned parameters, so detach the quantizer internals and use the
standard identity STE to the folded weight. This preserves transform
gradients without retaining full-weight rounding graphs. Record this
surrogate in every calibration artifact.
sensitivity_required: true
- id: LQ7_DIAGONAL_FISHER_COLLECTION
status: assumption
paper_status: >-
Equation (8) defines the calibration loss and Equation (6) names a
diagonal Fisher term, but the per-example/recurrent reduction is omitted.
decision: >-
For every one of the 24x4 projection groups, compute per-loop transform
gradients of the full Equation (8) loss. During each statistics forward,
use value-identical loop-specific shadow transforms so autograd assigns
both activation-path and folded-weight-path contributions to the correct
loop without changing the forward result. Average gradients over calibration
examples and estimate diagonal Fisher as E[g^2] over examples and loops.
Q/K/V and Gate/Up member VJPs are summed within their paper atomic group.
sensitivity_required: true
- id: LQ7_PROGRESSIVE_ADAPTATION_STEPS
status: assumption
paper_status: >-
Appendix B.1 requires progressive selection with adaptation and score
recomputation, but gives no optimizer-step count between rounds.
decision: >-
Require slt_round_steps explicitly. After each newly selected group,
rebuild the exact allowlisted optimizer, adapt for that many examples,
then recompute all remaining scores. Optimizer state is intentionally not
carried across changed parameter ownership; study this choice.
sensitivity_required: true
- id: LQ8_HUGINN_FULL_GRADIENT_RECURRENCES
status: compatibility_decision
paper_status: >-
LoopQ calibrates recurrent trajectories but does not evaluate Huginn or
specify how its no-grad/with-grad recurrence sampler should be configured.
decision: >-
During LoopQ calibration, pass Huginn num_steps=[0,32] so every recurrence
participates in the differentiable student trajectory. Teacher execution
uses 32 no-grad recurrences. Do not accept Huginn's training default that
would silently place early recurrence steps outside the calibration graph.
sensitivity_required: true
- id: LQ8_SAVED_TENSOR_CPU_OFFLOAD
status: compatibility_decision
paper_status: >-
The paper does not specify where calibration autograd intermediates are
stored or a peak-memory implementation for recurrent full-weight QDQ.
decision: >-
Store tensors saved for backward in CPU memory and restore the exact
tensors to their original GPU
on demand in backward. This changes placement and transfer cost only: it
does not detach, compress, checkpoint, approximate, or change any LoopQ
objective, gradient, quantizer, selection rule, or trainable parameter.
Ouro uses pinned storage after inferred weight-quantizer internals are
detached; Huginn conservatively uses pageable storage because its
32-recurrence saved set may exceed the CUDA pinned-host allocator budget.
Record the offload mode in every calibration bundle and make no
calibration-speed claim from this path.
sensitivity_required: false
- id: LQ9_DIRECT_RTN_CONTROL
status: compatibility_decision
paper_status: >-
LoopQ compares against direct W4A4 and W4A8 RTN but does not publish a
vLLM integration contract for those controls.
decision: >-
Quantize every mapped projection weight with symmetric group-32 W4 RTN
once at load time and quantize each mapped linear input dynamically with
symmetric group-32 A4 or A8 RTN. Use identity transforms, no LAS, no SLT,
and no CTA. Keep the same architecture-specific vLLM GSM8K protocol as
the corresponding BF16 and LoopQ rows. Label this direct RTN rather than
LoopQ and make the mode mutually exclusive with calibrated artifacts and
all ResidualQuant, MomentumQuant, and NVFP4 policies.
sensitivity_required: false
- id: LQ_TRAJECTORY_TRANSITION_TARGET
status: compatibility_decision
paper_status: "Equation (8) and Appendix B.4 target H_{t+1,0}, the next loop input."
decision: >-
The Ouro normalized loop output is directly fed into the next loop.
Align CTA output with teacher_hidden[:-1], not teacher_hidden[1:].
Huginn uses the same recurrent-state boundary before the input adapter;
input injection belongs to its recurrent function, not to CTA.
sensitivity_required: false
- id: LQ_HUGINN_MATCHED_RANDOM_STATE
status: compatibility_decision
paper_status: "Huginn is an architecture extension, not a paper backbone."
decision: >-
Fork and restore CPU/CUDA RNG around the teacher forward so the student
sees the same random initial recurrent state and configured noise.
Advance the outer generator only through the student forward. Record
the seed; do not supervise independently sampled initial trajectories.
sensitivity_required: false
- id: LQ_GROUP_BATCHING_AND_WEIGHT_STE_STORAGE
status: compatibility_decision
paper_status: "The paper does not specify PyTorch allocation scheduling."
decision: >-
Vectorize group-wise RTN, preserving inferred-scale tie handling and
supplied-scale derivatives. Detach folded weights only on the inferred
quantizer input, retaining the existing identity STE to folded weights.
This prevents saving an autograd quantizer graph that is discarded by
the STE. The inverse-transform and full recurrent gradient paths remain.
sensitivity_required: false
- id: LQ_LINEAR_RECOMPUTATION
status: compatibility_decision
paper_status: "The paper does not specify autograd intermediate storage."
decision: >-
Offer --checkpoint-linears as an explicit non-reentrant recomputation
option. Recompute only the deterministic transform/LAS/QDQ/linear function,
capture the exact transform object and recurrence index, and never replay
routing hooks. Retain every recurrent edge and all calibration parameters.
Default saved-tensor offload remains available as the parity reference.
Validate loss, SLT statistics and final component tensors before use.
sensitivity_required: false
- id: LQ_ACTIVATION_STE_EXACT_FORWARD
status: compatibility_decision
paper_status: "RTN forward values must match the exported QDQ runtime."
decision: >-
Compute activation STE as quantized + (original - original.detach()).
Subtract identical values before adding to QDQ so the forward value is
exact even in BF16. Preserve both the identity input gradient and the
learned-LAS quantized branch gradient. Left-associated addition and
subtraction introduce avoidable BF16 rounding and are not equivalent.
sensitivity_required: false
- id: LQ6_FP32_OBJECTIVE
status: compatibility_decision
paper_status: "Loss accumulation dtype is not specified."
decision: >-
Keep the backbone BF16 but compute selected-logit softmax/KL and hidden
squared-error reductions in FP32. Preserve the sum and batchmean objective
reductions; do not silently normalize by token count or hidden width.
sensitivity_required: false
- id: LQ6_SCAN_MU_CLOCK
status: assumption
paper_status: "Appendix B.4 gives a periodic calibration-step update, but not scan semantics."
decision: >-
The training mu cache refreshes once per eligible optimizer step. Each
SLT sample uses independently computed detached mu at the current model;
score collection never advances or replaces the training mu cache.
sensitivity_required: true
- id: LQ6_STREAMING_FISHER
status: compatibility_decision
paper_status: "Dataset reduction/storage precision is not specified."
decision: >-
Accumulate per-loop gradient sums and empirical squared-gradient sums in
CPU FP64. Divide by the actual sample count; Fisher also averages loops.
Store sufficient statistics and scan cursors in atomic resume checkpoints.
sensitivity_required: false
- id: LQ6_RESUME_PHASES
status: assumption
paper_status: "Optimizer-state behavior between progressive phases is unspecified."
decision: >-
Preserve optimizer moments within each SLT adaptation or final training
phase; reset at phase boundaries as in the original real drivers. Resume
restores component parameters, optimizer moments, mu, CPU/CUDA RNG, traces,
completed scan summaries and partial scan sufficient statistics. Refuse
changed source, configuration or calibration-text identities.
sensitivity_required: true
- id: LQ9_ABLATION_TRAINING
status: assumption
paper_status: "Table 2 names component removals but does not give exact ablation code."
decision: >-
Retrain each ablation. No-LAS learns a single shared scale vector per
module initialized by the maximum across loop scales. No-SLT retains
shared transforms and selects no groups. No-CTA bypasses the adapter,
freezes its parameters and omits the transition loss. Keep the other
recurrent loss terms. Persist component mode in exports and reject an
ablation artifact supplied as a full LoopQ row.
sensitivity_required: true
- id: LQ6_ACTIVATION_STE_SENSITIVITY
status: assumption
paper_status: >-
LoopQ does not specify the quantizer's backward estimator. The pinned
FlatQuant quant_utils.py uses an STE on normalized values before clipping.
decision: >-
Expose identity and rounding activation STEs explicitly. Identity retains
the prior direct activation gradient plus code-times-scale derivative.
Rounding uses unit derivative through normalized rounding followed by
clipping, yielding code-minus-normalized-value scale derivatives inside
the range. Both preserve identical RTN forward values, group layout and
bounds. Compare calibration convergence before choosing full-run settings.
sensitivity_required: true
- id: LQ_EXPORT_EXACT_LAS_PARAMETERS
status: compatibility_decision
paper_status: "Serialization precision is not specified."
decision: >-
Store raw FP32 trained log-scales alongside expanded positive scales and
restore raw parameters exactly. Validate consistency of both payloads.
Legacy exports without raw parameters remain readable. Export validation
also checks component-ablation modes and calibration activation precision.
sensitivity_required: false
- id: LQ6_FP64_STATISTICS_AND_CLIP_NORM
status: compatibility_decision
paper_status: >-
Fisher uses squared transform gradients and gradient clipping has norm 1;
the paper does not specify the reduction dtype.
decision: >-
Square per-loop transform gradients in FP64 before Fisher reduction.
Compute the global gradient clipping norm in FP64, retaining the norm-1
rule and rejection of genuine NaN/Inf. This prevents FP32 overflow from
finite gradients observed in Ouro full calibration sample index 6.
This changes floating-point training traces; do not bypass source-contract
checks to resume an old run with the new implementation.
sensitivity_required: false