schema_version: 1 paper: "arXiv:2605.16343v1" decisions: - id: LQ1_INTEGER_RANGE status: assumption paper_status: >- Equation (4) names q_min and q_max but does not give their numerical values or state whether the most-negative two's-complement code is used. decision: >- Use the cited FlatQuant symmetric RTN convention: q_min = -2^(b-1) and q_max = 2^(b-1)-1 ([-8, 7] for 4-bit and [-128, 127] for 8-bit). sensitivity_required: true - id: LQ1_SCALE_INITIALIZATION status: assumption paper_status: >- The paper learns activation ranges during calibration but does not state the initial static RTN scale estimator for weights or activations. decision: >- Initialize each group scale as max(abs(group)) / q_max. An all-zero group receives scale 1 and therefore remains exactly zero. sensitivity_required: true - id: LQ1_ROUNDING_TIES status: assumption paper_status: "The paper says round-to-nearest but does not specify tie handling." decision: "Round exact half-way values away from zero deterministically." sensitivity_required: false - id: LQ1_GROUP_AXIS status: assumption paper_status: "The paper specifies group-wise quantization but not tensor-axis layout." decision: >- Form consecutive groups along the last logical dimension for both weights and activations. Architecture adapters must document any layout conversion before invoking this quantizer. sensitivity_required: true - id: LQ1_TAIL_GROUP status: compatibility_decision paper_status: "Tail handling is not specified." decision: >- Quantize a final group shorter than 32 independently, without persistent padding. Its scale is computed only from real elements. sensitivity_required: false - id: LQ1_FAKE_QUANT_SCOPE status: compatibility_decision paper_status: >- The paper reports packed weights and quantized runtime kernels, but LQ1 is the numerical quantizer acceptance phase. decision: >- LQ1 implements deterministic QDQ/fake quantization and emits integer codes plus scales. Packing and runtime kernels are deferred to LQ7/LQ8. sensitivity_required: false - id: LQ2_GROUP_SCALE_LAYOUT status: compatibility_decision paper_status: >- Section 4.1 specifies c_{t,l} per module and loop and explicitly counts O(TL) additional scalars. Appendix B.3 says "Per module, per loop". decision: >- Store one positive clipping multiplier for each module and loop. Apply it to dynamic per-token, per-group absmax scales (group size 32), matching the FlatQuant-style learned activation-clipping construction cited by the paper. This gives exactly O(TL) LAS parameters. sensitivity_required: false - id: LQ2_CALIBRATION_REDUCTION status: assumption paper_status: >- The initialization procedure for learned LAS ranges is not specified. decision: >- Initialize each module-loop clipping multiplier to 1.0. Dynamic base scales use each input token/group absmax. Subsequent trajectory calibration optimizes the log multiplier. sensitivity_required: true - id: LQ2_POSITIVE_PARAMETERIZATION status: compatibility_decision paper_status: "The paper requires scales but does not specify their parameterization." decision: >- Store trainable log-scales and exponentiate at use time, ensuring scales remain strictly positive during later gradient calibration. sensitivity_required: false - id: LQ2_ARCHITECTURE_LOOP_ROUTING status: compatibility_decision paper_status: >- The paper evaluates Ouro with four loops and does not evaluate Huginn. decision: >- Provide exact zero-based routing for Ouro loops [0,4) and Huginn recurrences [0,32), with no modulo, anchor, or fallback routing. sensitivity_required: false - id: LQ3_FLATQUANT_REFERENCE_PIN status: compatibility_decision paper_status: >- LoopQ cites FlatQuant and states that transforms are Kronecker-decomposed, but does not identify a code revision. decision: >- Pin the official ruikangliu/FlatQuant repository at commit 9d88ffcb7d2c6bda59fb5c44dad36adc101aadb1. File hashes and the inspected folding contract are recorded in flatquant_reference.yaml. sensitivity_required: false - id: LQ3_FLATQUANT_SVD_PARAMETERIZATION status: compatibility_decision paper_status: >- LoopQ specifies FlatQuant-style Kronecker-decomposed transforms but does not state whether it uses FlatQuant's SVD or direct-inverse class, factor dimensions, add-diagonal flag, or architecture-specific initialization. decision: >- Use the default path of the paper-pinned official FlatQuant commit: SVDDecomposeTransMatrix with independently random-orthogonal U/V factors, trainable singular diagonals, and Cayley orthogonal parameterizations. Use explicit factor dimensions and omit FlatQuant's optional add_diag because LoopQ does not state that flag or its SmoothQuant initialization. Export effective Kronecker factors so runtime is independent of the training parameterization; retain legacy identity transform artifacts. sensitivity_required: true - id: LQ3_CALIBRATION_BLOCKER status: assumption paper_status: >- LoopQ omits optimizer, learning rate, epochs/steps, Kronecker factor dimensions, inverse method, and diagonal initialization. FlatQuant's official code uses AdamW and cosine scheduling, but repository defaults and published scripts differ and its layer-wise MSE objective is not LoopQ's trajectory objective. decision: >- Do not import FlatQuant's optimizer settings into LoopQ. Defer transform calibration parameterization and optimizer selection to LQ6, record the chosen sensitivity study there, and keep LQ3 limited to fold parity. sensitivity_required: true - id: LQ4_VARIANCE_ESTIMATOR status: compatibility_decision paper_status: >- Equation (6) explicitly expands variance with a 1/T factor but does not discuss library estimator settings. decision: >- Use population variance across the true loop dimension (unbiased=false), accumulate the score in deterministic CPU FP64, and sum every coordinate in the transform-weight group. sensitivity_required: false - id: LQ4_FISHER_AND_EPSILON status: assumption paper_status: >- The paper identifies psi as a diagonal Fisher estimate and epsilon as a positive stabilizer, but does not specify the Fisher sampling estimator, reduction, or epsilon value. decision: >- LQ4 consumes non-negative saved Fisher diagonals without inventing their collection procedure and uses epsilon=1e-8 by default. Fisher collection must be resolved and tested with trajectory calibration in LQ6. sensitivity_required: true - id: LQ4_PROGRESSIVE_RECOMPUTATION status: compatibility_decision paper_status: >- Section 4.2 requires alternation between adapting selected transforms and recomputing scores, but does not specify iteration counts or tie handling. decision: >- Select exactly one atomic transform-weight group per round, require fresh callback statistics after each selection, and break equal-score ties by lexical candidate name for deterministic artifacts. sensitivity_required: false - id: LQ4_HUGINN_BUDGET_ROUNDING status: assumption paper_status: >- Huginn is not evaluated. LoopQ reports that about 5% of transform parameters are selected, without a rounding rule for unequal atomic groups. decision: >- Set the Huginn parameter target to ceil(0.05 * all candidate transform parameters), select whole groups progressively, and stop after the first selection whose cumulative count meets or exceeds the target. Record the resulting atomic-group overshoot. sensitivity_required: true - id: LQ5_RMSNORM_DEFINITION status: assumption paper_status: >- Equation (7) says RMSNorm along the feature dimension but does not give epsilon or specify reuse of a backbone norm's learned weight. decision: >- Use unweighted RMS normalization x / sqrt(mean(x^2) + 1e-6), computed in FP32 and returned in the hidden-state dtype. Evaluate epsilon sensitivity before full calibration. sensitivity_required: true - id: LQ5_PARAMETER_SHAPES status: compatibility_decision paper_status: >- Equation (7) defines shared U,V and loop-dependent a_t,b_t,eta_t but only specifies adapter rank 8 in Table 6. decision: >- For hidden width d and rank r, use U,V in R^(d x r), a_t,b_t in R^d, and eta_t in R^r. Store parameters for actual transitions only: T-1 rows. sensitivity_required: false - id: LQ5_IDENTITY_INITIALIZATION status: assumption paper_status: "The paper does not state CTA initialization." decision: >- Initialize a_t=1, b_t=0, eta_t=0. Initialize shared U and V to the first r columns of the identity for deterministic nonzero gate gradients. This makes every CTA exactly identity before calibration. sensitivity_required: true - id: LQ5_TRANSITION_ROUTING status: compatibility_decision paper_status: >- Equation (7) applies A_t between loop t and t+1. Ouro has four loops; Huginn is not evaluated by the paper. decision: >- Route exactly three Ouro transitions and thirty-one Huginn transitions with zero-based indices. Reject out-of-range indices; never use modulo, Residual/Momentum anchors, or mixed-precision schedule boundaries. sensitivity_required: false - id: LQ6_LOGIT_TOPK_KL status: compatibility_decision paper_status: >- Table 6 specifies teacher top-k logits 1000 and temperature 1, but does not state whether KL is renormalized over the retained support or how reductions are applied. decision: >- Select the teacher's top-1000 indices, renormalize teacher and student distributions on that same support, and compute KL(teacher || student) with PyTorch batchmean reduction. For synthetic vocabularies smaller than 1000 only, use the whole vocabulary. sensitivity_required: true - id: LQ6_TRAJECTORY_REDUCTION status: assumption paper_status: >- Equation (8) shows squared L2 norms and sums over loops but does not give batch/token reduction or normalization details. decision: >- Use squared L2 sums across every non-loop coordinate and sum across loops and transitions exactly as displayed. Apply lambda=0.1 to all trajectory and transition terms. Record each weighted term separately. sensitivity_required: true - id: LQ6_ADAPTIVE_MU_REFRESH status: compatibility_decision paper_status: >- Appendix B.4 states that mu_t is updated periodically, for example every 100 calibration steps; this is illustrative rather than an immutable value. decision: >- Default to the paper's example interval 100, cache detached mu values between refreshes, serialize the cache, and expose the interval in the required calibration config for sensitivity analysis. sensitivity_required: true - id: LQ6_OPTIMIZER_SCHEDULE_STEPS status: assumption paper_status: >- LoopQ does not specify optimizer, learning rate, total calibration steps, weight decay, warmup, or scheduler. FlatQuant's settings use a different layer-wise reconstruction objective and cannot be silently inherited. decision: >- Require optimizer name, base learning rate, and final-phase steps explicitly with no defaults. Permit explicit LAS, transform, and CTA learning-rate overrides derived from the exact parameter allowlist so omitted component rates can be studied without changing the objective. Initially permit Adam or AdamW as caller-selected alternatives. Record both final-phase and global optimizer-step counts in every checkpoint, run sensitivity before claiming reproduction fidelity, and do not add an implicit scheduler. sensitivity_required: true - id: LQ6_PARAMETER_ALLOWLIST status: compatibility_decision paper_status: >- Section 4.4 allows shared transforms, LAS, selected loop transforms, and CTA parameters to update while the shared backbone remains fixed. decision: >- Construct an exact optimizer allowlist from only those four categories, reject duplicated ownership or any identity overlap with backbone parameters, verify the optimizer on every step, and clip their joint norm to the paper value 1.0. sensitivity_required: false - id: LQ7_OURO_CHECKPOINT_PIN status: compatibility_decision paper_status: >- The paper names Ouro 1.4B but does not provide a checkpoint revision or hash in the implementation appendix. decision: >- Validate against local ByteDance/Ouro-1.4B revision e3b1e0993b1231a51d6069a870476dda4162c00a. Pin config and modeling-code SHA256 values in ouro_mapping.yaml and inspect safetensors keys without loading parameter tensors. sensitivity_required: false - id: LQ7_VLLM_PACKED_MAPPING status: compatibility_decision paper_status: >- Appendix B.1 lists separate HF projection names, whereas the local vLLM backend packs Q/K/V and Gate/Up projections. decision: >- Preserve the four paper atomic groups while mapping Q/K/V to qkv_proj and Gate/Up to gate_up_proj consumers. Artifact provenance retains both the original HF weight names and packed runtime names. sensitivity_required: false - id: LQ7_RUNTIME_INTEGRATION_BOUNDARY status: compatibility_decision paper_status: "The paper does not specify vLLM hook APIs." decision: >- Keep the adapter backend-neutral: expose transform/LAS/QDQ preparation and exact CTA boundaries, but do not import the existing NVFP4/Residual plugin. Runtime model.py must later call these hooks at four listed consumers and three true loop boundaries before a smoke result can be claimed. sensitivity_required: false - id: LQ7_SELECTED_SLT_PACKED_WEIGHT_BLOCKER status: compatibility_decision paper_status: >- Appendix B.1 requires selected groups to use loop-dependent transforms and quantized weights. The local vLLM projections own one packed weight parameter shared by all recurrent calls. decision: >- Keep one shared base parameter for every projection and allocate four loop-specific packed RTN variants only for the paper-budget four selected groups. Merge Q/K/V and Gate/Up source shards with vLLM's existing weight loaders, then dispatch by exact loop index and restore the shared pointer. This correctness path requires eager execution. sensitivity_required: false - id: LQ7_SELECTED_WEIGHT_EAGER_DISPATCH status: compatibility_decision paper_status: "The paper does not specify vLLM compiled packed-weight dispatch." decision: >- Use scoped parameter-data substitution only in the explicit eager smoke path. Do not claim compiled/runtime speed results. A future native kernel or quant-method weight argument is required before torch.compile use. sensitivity_required: false - id: LQ7_CALIBRATION_STE status: compatibility_decision paper_status: >- The paper optimizes quantization parameters through RTN but does not specify the rounding surrogate used by its calibration implementation. decision: >- Use quantized values in the forward pass. For activation QDQ, retain the differentiable learned-LAS scale path and add an identity input path around rounding. For weight QDQ, RTN scales are inferred rather than learned parameters, so detach the quantizer internals and use the standard identity STE to the folded weight. This preserves transform gradients without retaining full-weight rounding graphs. Record this surrogate in every calibration artifact. sensitivity_required: true - id: LQ7_DIAGONAL_FISHER_COLLECTION status: assumption paper_status: >- Equation (8) defines the calibration loss and Equation (6) names a diagonal Fisher term, but the per-example/recurrent reduction is omitted. decision: >- For every one of the 24x4 projection groups, compute per-loop transform gradients of the full Equation (8) loss. During each statistics forward, use value-identical loop-specific shadow transforms so autograd assigns both activation-path and folded-weight-path contributions to the correct loop without changing the forward result. Average gradients over calibration examples and estimate diagonal Fisher as E[g^2] over examples and loops. Q/K/V and Gate/Up member VJPs are summed within their paper atomic group. sensitivity_required: true - id: LQ7_PROGRESSIVE_ADAPTATION_STEPS status: assumption paper_status: >- Appendix B.1 requires progressive selection with adaptation and score recomputation, but gives no optimizer-step count between rounds. decision: >- Require slt_round_steps explicitly. After each newly selected group, rebuild the exact allowlisted optimizer, adapt for that many examples, then recompute all remaining scores. Optimizer state is intentionally not carried across changed parameter ownership; study this choice. sensitivity_required: true - id: LQ8_HUGINN_FULL_GRADIENT_RECURRENCES status: compatibility_decision paper_status: >- LoopQ calibrates recurrent trajectories but does not evaluate Huginn or specify how its no-grad/with-grad recurrence sampler should be configured. decision: >- During LoopQ calibration, pass Huginn num_steps=[0,32] so every recurrence participates in the differentiable student trajectory. Teacher execution uses 32 no-grad recurrences. Do not accept Huginn's training default that would silently place early recurrence steps outside the calibration graph. sensitivity_required: true - id: LQ8_SAVED_TENSOR_CPU_OFFLOAD status: compatibility_decision paper_status: >- The paper does not specify where calibration autograd intermediates are stored or a peak-memory implementation for recurrent full-weight QDQ. decision: >- Store tensors saved for backward in CPU memory and restore the exact tensors to their original GPU on demand in backward. This changes placement and transfer cost only: it does not detach, compress, checkpoint, approximate, or change any LoopQ objective, gradient, quantizer, selection rule, or trainable parameter. Ouro uses pinned storage after inferred weight-quantizer internals are detached; Huginn conservatively uses pageable storage because its 32-recurrence saved set may exceed the CUDA pinned-host allocator budget. Record the offload mode in every calibration bundle and make no calibration-speed claim from this path. sensitivity_required: false - id: LQ9_DIRECT_RTN_CONTROL status: compatibility_decision paper_status: >- LoopQ compares against direct W4A4 and W4A8 RTN but does not publish a vLLM integration contract for those controls. decision: >- Quantize every mapped projection weight with symmetric group-32 W4 RTN once at load time and quantize each mapped linear input dynamically with symmetric group-32 A4 or A8 RTN. Use identity transforms, no LAS, no SLT, and no CTA. Keep the same architecture-specific vLLM GSM8K protocol as the corresponding BF16 and LoopQ rows. Label this direct RTN rather than LoopQ and make the mode mutually exclusive with calibrated artifacts and all ResidualQuant, MomentumQuant, and NVFP4 policies. sensitivity_required: false - id: LQ_TRAJECTORY_TRANSITION_TARGET status: compatibility_decision paper_status: "Equation (8) and Appendix B.4 target H_{t+1,0}, the next loop input." decision: >- The Ouro normalized loop output is directly fed into the next loop. Align CTA output with teacher_hidden[:-1], not teacher_hidden[1:]. Huginn uses the same recurrent-state boundary before the input adapter; input injection belongs to its recurrent function, not to CTA. sensitivity_required: false - id: LQ_HUGINN_MATCHED_RANDOM_STATE status: compatibility_decision paper_status: "Huginn is an architecture extension, not a paper backbone." decision: >- Fork and restore CPU/CUDA RNG around the teacher forward so the student sees the same random initial recurrent state and configured noise. Advance the outer generator only through the student forward. Record the seed; do not supervise independently sampled initial trajectories. sensitivity_required: false - id: LQ_GROUP_BATCHING_AND_WEIGHT_STE_STORAGE status: compatibility_decision paper_status: "The paper does not specify PyTorch allocation scheduling." decision: >- Vectorize group-wise RTN, preserving inferred-scale tie handling and supplied-scale derivatives. Detach folded weights only on the inferred quantizer input, retaining the existing identity STE to folded weights. This prevents saving an autograd quantizer graph that is discarded by the STE. The inverse-transform and full recurrent gradient paths remain. sensitivity_required: false - id: LQ_LINEAR_RECOMPUTATION status: compatibility_decision paper_status: "The paper does not specify autograd intermediate storage." decision: >- Offer --checkpoint-linears as an explicit non-reentrant recomputation option. Recompute only the deterministic transform/LAS/QDQ/linear function, capture the exact transform object and recurrence index, and never replay routing hooks. Retain every recurrent edge and all calibration parameters. Default saved-tensor offload remains available as the parity reference. Validate loss, SLT statistics and final component tensors before use. sensitivity_required: false - id: LQ_ACTIVATION_STE_EXACT_FORWARD status: compatibility_decision paper_status: "RTN forward values must match the exported QDQ runtime." decision: >- Compute activation STE as quantized + (original - original.detach()). Subtract identical values before adding to QDQ so the forward value is exact even in BF16. Preserve both the identity input gradient and the learned-LAS quantized branch gradient. Left-associated addition and subtraction introduce avoidable BF16 rounding and are not equivalent. sensitivity_required: false - id: LQ6_FP32_OBJECTIVE status: compatibility_decision paper_status: "Loss accumulation dtype is not specified." decision: >- Keep the backbone BF16 but compute selected-logit softmax/KL and hidden squared-error reductions in FP32. Preserve the sum and batchmean objective reductions; do not silently normalize by token count or hidden width. sensitivity_required: false - id: LQ6_SCAN_MU_CLOCK status: assumption paper_status: "Appendix B.4 gives a periodic calibration-step update, but not scan semantics." decision: >- The training mu cache refreshes once per eligible optimizer step. Each SLT sample uses independently computed detached mu at the current model; score collection never advances or replaces the training mu cache. sensitivity_required: true - id: LQ6_STREAMING_FISHER status: compatibility_decision paper_status: "Dataset reduction/storage precision is not specified." decision: >- Accumulate per-loop gradient sums and empirical squared-gradient sums in CPU FP64. Divide by the actual sample count; Fisher also averages loops. Store sufficient statistics and scan cursors in atomic resume checkpoints. sensitivity_required: false - id: LQ6_RESUME_PHASES status: assumption paper_status: "Optimizer-state behavior between progressive phases is unspecified." decision: >- Preserve optimizer moments within each SLT adaptation or final training phase; reset at phase boundaries as in the original real drivers. Resume restores component parameters, optimizer moments, mu, CPU/CUDA RNG, traces, completed scan summaries and partial scan sufficient statistics. Refuse changed source, configuration or calibration-text identities. sensitivity_required: true - id: LQ9_ABLATION_TRAINING status: assumption paper_status: "Table 2 names component removals but does not give exact ablation code." decision: >- Retrain each ablation. No-LAS learns a single shared scale vector per module initialized by the maximum across loop scales. No-SLT retains shared transforms and selects no groups. No-CTA bypasses the adapter, freezes its parameters and omits the transition loss. Keep the other recurrent loss terms. Persist component mode in exports and reject an ablation artifact supplied as a full LoopQ row. sensitivity_required: true - id: LQ6_ACTIVATION_STE_SENSITIVITY status: assumption paper_status: >- LoopQ does not specify the quantizer's backward estimator. The pinned FlatQuant quant_utils.py uses an STE on normalized values before clipping. decision: >- Expose identity and rounding activation STEs explicitly. Identity retains the prior direct activation gradient plus code-times-scale derivative. Rounding uses unit derivative through normalized rounding followed by clipping, yielding code-minus-normalized-value scale derivatives inside the range. Both preserve identical RTN forward values, group layout and bounds. Compare calibration convergence before choosing full-run settings. sensitivity_required: true - id: LQ_EXPORT_EXACT_LAS_PARAMETERS status: compatibility_decision paper_status: "Serialization precision is not specified." decision: >- Store raw FP32 trained log-scales alongside expanded positive scales and restore raw parameters exactly. Validate consistency of both payloads. Legacy exports without raw parameters remain readable. Export validation also checks component-ablation modes and calibration activation precision. sensitivity_required: false - id: LQ6_FP64_STATISTICS_AND_CLIP_NORM status: compatibility_decision paper_status: >- Fisher uses squared transform gradients and gradient clipping has norm 1; the paper does not specify the reduction dtype. decision: >- Square per-loop transform gradients in FP64 before Fisher reduction. Compute the global gradient clipping norm in FP64, retaining the norm-1 rule and rejection of genuine NaN/Inf. This prevents FP32 overflow from finite gradients observed in Ouro full calibration sample index 6. This changes floating-point training traces; do not bypass source-contract checks to resume an old run with the new implementation. sensitivity_required: false