Download loopq_quantization/scripts/configs/loopq_compatibility_decisions.yaml from JunYoungLee/ut-depth-probe-artifacts: direct link, hf CLI and curl.
- Browser
- Download file 27.5 kB
-
https://huggingface.co/JunYoungLee/ut-depth-probe-artifacts/resolve/main/loopq_quantization/scripts/configs/loopq_compatibility_decisions.yaml
- Command line
-
hf download hf://JunYoungLee/ut-depth-probe-artifacts/loopq_quantization/scripts/configs/loopq_compatibility_decisions.yaml
-
curl -L -o loopq_compatibility_decisions.yaml https://huggingface.co/JunYoungLee/ut-depth-probe-artifacts/resolve/main/loopq_quantization/scripts/configs/loopq_compatibility_decisions.yaml
27.5 kB
| schema_version: 1 | |
| paper: "arXiv:2605.16343v1" | |
| decisions: | |
| - id: LQ1_INTEGER_RANGE | |
| status: assumption | |
| paper_status: >- | |
| Equation (4) names q_min and q_max but does not give their numerical | |
| values or state whether the most-negative two's-complement code is used. | |
| decision: >- | |
| Use the cited FlatQuant symmetric RTN convention: q_min = -2^(b-1) and | |
| q_max = 2^(b-1)-1 ([-8, 7] for 4-bit and [-128, 127] for 8-bit). | |
| sensitivity_required: true | |
| - id: LQ1_SCALE_INITIALIZATION | |
| status: assumption | |
| paper_status: >- | |
| The paper learns activation ranges during calibration but does not state | |
| the initial static RTN scale estimator for weights or activations. | |
| decision: >- | |
| Initialize each group scale as max(abs(group)) / q_max. An all-zero | |
| group receives scale 1 and therefore remains exactly zero. | |
| sensitivity_required: true | |
| - id: LQ1_ROUNDING_TIES | |
| status: assumption | |
| paper_status: "The paper says round-to-nearest but does not specify tie handling." | |
| decision: "Round exact half-way values away from zero deterministically." | |
| sensitivity_required: false | |
| - id: LQ1_GROUP_AXIS | |
| status: assumption | |
| paper_status: "The paper specifies group-wise quantization but not tensor-axis layout." | |
| decision: >- | |
| Form consecutive groups along the last logical dimension for both | |
| weights and activations. Architecture adapters must document any layout | |
| conversion before invoking this quantizer. | |
| sensitivity_required: true | |
| - id: LQ1_TAIL_GROUP | |
| status: compatibility_decision | |
| paper_status: "Tail handling is not specified." | |
| decision: >- | |
| Quantize a final group shorter than 32 independently, without persistent | |
| padding. Its scale is computed only from real elements. | |
| sensitivity_required: false | |
| - id: LQ1_FAKE_QUANT_SCOPE | |
| status: compatibility_decision | |
| paper_status: >- | |
| The paper reports packed weights and quantized runtime kernels, but LQ1 | |
| is the numerical quantizer acceptance phase. | |
| decision: >- | |
| LQ1 implements deterministic QDQ/fake quantization and emits integer | |
| codes plus scales. Packing and runtime kernels are deferred to LQ7/LQ8. | |
| sensitivity_required: false | |
| - id: LQ2_GROUP_SCALE_LAYOUT | |
| status: compatibility_decision | |
| paper_status: >- | |
| Section 4.1 specifies c_{t,l} per module and loop and explicitly counts | |
| O(TL) additional scalars. Appendix B.3 says "Per module, per loop". | |
| decision: >- | |
| Store one positive clipping multiplier for each module and loop. Apply it | |
| to dynamic per-token, per-group absmax scales (group size 32), matching | |
| the FlatQuant-style learned activation-clipping construction cited by the | |
| paper. This gives exactly O(TL) LAS parameters. | |
| sensitivity_required: false | |
| - id: LQ2_CALIBRATION_REDUCTION | |
| status: assumption | |
| paper_status: >- | |
| The initialization procedure for learned LAS ranges is not specified. | |
| decision: >- | |
| Initialize each module-loop clipping multiplier to 1.0. Dynamic base | |
| scales use each input token/group absmax. Subsequent trajectory | |
| calibration optimizes the log multiplier. | |
| sensitivity_required: true | |
| - id: LQ2_POSITIVE_PARAMETERIZATION | |
| status: compatibility_decision | |
| paper_status: "The paper requires scales but does not specify their parameterization." | |
| decision: >- | |
| Store trainable log-scales and exponentiate at use time, ensuring scales | |
| remain strictly positive during later gradient calibration. | |
| sensitivity_required: false | |
| - id: LQ2_ARCHITECTURE_LOOP_ROUTING | |
| status: compatibility_decision | |
| paper_status: >- | |
| The paper evaluates Ouro with four loops and does not evaluate Huginn. | |
| decision: >- | |
| Provide exact zero-based routing for Ouro loops [0,4) and Huginn | |
| recurrences [0,32), with no modulo, anchor, or fallback routing. | |
| sensitivity_required: false | |
| - id: LQ3_FLATQUANT_REFERENCE_PIN | |
| status: compatibility_decision | |
| paper_status: >- | |
| LoopQ cites FlatQuant and states that transforms are Kronecker-decomposed, | |
| but does not identify a code revision. | |
| decision: >- | |
| Pin the official ruikangliu/FlatQuant repository at commit | |
| 9d88ffcb7d2c6bda59fb5c44dad36adc101aadb1. File hashes and the inspected | |
| folding contract are recorded in flatquant_reference.yaml. | |
| sensitivity_required: false | |
| - id: LQ3_FLATQUANT_SVD_PARAMETERIZATION | |
| status: compatibility_decision | |
| paper_status: >- | |
| LoopQ specifies FlatQuant-style Kronecker-decomposed transforms but does | |
| not state whether it uses FlatQuant's SVD or direct-inverse class, factor | |
| dimensions, add-diagonal flag, or architecture-specific initialization. | |
| decision: >- | |
| Use the default path of the paper-pinned official FlatQuant commit: | |
| SVDDecomposeTransMatrix with independently random-orthogonal U/V factors, | |
| trainable singular diagonals, and Cayley orthogonal parameterizations. | |
| Use explicit factor dimensions and omit FlatQuant's optional add_diag | |
| because LoopQ does not state that flag or its SmoothQuant initialization. | |
| Export effective Kronecker factors so runtime is independent of the | |
| training parameterization; retain legacy identity transform artifacts. | |
| sensitivity_required: true | |
| - id: LQ3_CALIBRATION_BLOCKER | |
| status: assumption | |
| paper_status: >- | |
| LoopQ omits optimizer, learning rate, epochs/steps, Kronecker factor | |
| dimensions, inverse method, and diagonal initialization. FlatQuant's | |
| official code uses AdamW and cosine scheduling, but repository defaults | |
| and published scripts differ and its layer-wise MSE objective is not | |
| LoopQ's trajectory objective. | |
| decision: >- | |
| Do not import FlatQuant's optimizer settings into LoopQ. Defer transform | |
| calibration parameterization and optimizer selection to LQ6, record the | |
| chosen sensitivity study there, and keep LQ3 limited to fold parity. | |
| sensitivity_required: true | |
| - id: LQ4_VARIANCE_ESTIMATOR | |
| status: compatibility_decision | |
| paper_status: >- | |
| Equation (6) explicitly expands variance with a 1/T factor but does not | |
| discuss library estimator settings. | |
| decision: >- | |
| Use population variance across the true loop dimension (unbiased=false), | |
| accumulate the score in deterministic CPU FP64, and sum every coordinate | |
| in the transform-weight group. | |
| sensitivity_required: false | |
| - id: LQ4_FISHER_AND_EPSILON | |
| status: assumption | |
| paper_status: >- | |
| The paper identifies psi as a diagonal Fisher estimate and epsilon as a | |
| positive stabilizer, but does not specify the Fisher sampling estimator, | |
| reduction, or epsilon value. | |
| decision: >- | |
| LQ4 consumes non-negative saved Fisher diagonals without inventing their | |
| collection procedure and uses epsilon=1e-8 by default. Fisher collection | |
| must be resolved and tested with trajectory calibration in LQ6. | |
| sensitivity_required: true | |
| - id: LQ4_PROGRESSIVE_RECOMPUTATION | |
| status: compatibility_decision | |
| paper_status: >- | |
| Section 4.2 requires alternation between adapting selected transforms and | |
| recomputing scores, but does not specify iteration counts or tie handling. | |
| decision: >- | |
| Select exactly one atomic transform-weight group per round, require fresh | |
| callback statistics after each selection, and break equal-score ties by | |
| lexical candidate name for deterministic artifacts. | |
| sensitivity_required: false | |
| - id: LQ4_HUGINN_BUDGET_ROUNDING | |
| status: assumption | |
| paper_status: >- | |
| Huginn is not evaluated. LoopQ reports that about 5% of transform | |
| parameters are selected, without a rounding rule for unequal atomic groups. | |
| decision: >- | |
| Set the Huginn parameter target to ceil(0.05 * all candidate transform | |
| parameters), select whole groups progressively, and stop after the first | |
| selection whose cumulative count meets or exceeds the target. Record the | |
| resulting atomic-group overshoot. | |
| sensitivity_required: true | |
| - id: LQ5_RMSNORM_DEFINITION | |
| status: assumption | |
| paper_status: >- | |
| Equation (7) says RMSNorm along the feature dimension but does not give | |
| epsilon or specify reuse of a backbone norm's learned weight. | |
| decision: >- | |
| Use unweighted RMS normalization x / sqrt(mean(x^2) + 1e-6), computed in | |
| FP32 and returned in the hidden-state dtype. Evaluate epsilon sensitivity | |
| before full calibration. | |
| sensitivity_required: true | |
| - id: LQ5_PARAMETER_SHAPES | |
| status: compatibility_decision | |
| paper_status: >- | |
| Equation (7) defines shared U,V and loop-dependent a_t,b_t,eta_t but only | |
| specifies adapter rank 8 in Table 6. | |
| decision: >- | |
| For hidden width d and rank r, use U,V in R^(d x r), a_t,b_t in R^d, | |
| and eta_t in R^r. Store parameters for actual transitions only: T-1 rows. | |
| sensitivity_required: false | |
| - id: LQ5_IDENTITY_INITIALIZATION | |
| status: assumption | |
| paper_status: "The paper does not state CTA initialization." | |
| decision: >- | |
| Initialize a_t=1, b_t=0, eta_t=0. Initialize shared U and V to the first | |
| r columns of the identity for deterministic nonzero gate gradients. This | |
| makes every CTA exactly identity before calibration. | |
| sensitivity_required: true | |
| - id: LQ5_TRANSITION_ROUTING | |
| status: compatibility_decision | |
| paper_status: >- | |
| Equation (7) applies A_t between loop t and t+1. Ouro has four loops; | |
| Huginn is not evaluated by the paper. | |
| decision: >- | |
| Route exactly three Ouro transitions and thirty-one Huginn transitions | |
| with zero-based indices. Reject out-of-range indices; never use modulo, | |
| Residual/Momentum anchors, or mixed-precision schedule boundaries. | |
| sensitivity_required: false | |
| - id: LQ6_LOGIT_TOPK_KL | |
| status: compatibility_decision | |
| paper_status: >- | |
| Table 6 specifies teacher top-k logits 1000 and temperature 1, but does | |
| not state whether KL is renormalized over the retained support or how | |
| reductions are applied. | |
| decision: >- | |
| Select the teacher's top-1000 indices, renormalize teacher and student | |
| distributions on that same support, and compute KL(teacher || student) | |
| with PyTorch batchmean reduction. For synthetic vocabularies smaller than | |
| 1000 only, use the whole vocabulary. | |
| sensitivity_required: true | |
| - id: LQ6_TRAJECTORY_REDUCTION | |
| status: assumption | |
| paper_status: >- | |
| Equation (8) shows squared L2 norms and sums over loops but does not give | |
| batch/token reduction or normalization details. | |
| decision: >- | |
| Use squared L2 sums across every non-loop coordinate and sum across loops | |
| and transitions exactly as displayed. Apply lambda=0.1 to all trajectory | |
| and transition terms. Record each weighted term separately. | |
| sensitivity_required: true | |
| - id: LQ6_ADAPTIVE_MU_REFRESH | |
| status: compatibility_decision | |
| paper_status: >- | |
| Appendix B.4 states that mu_t is updated periodically, for example every | |
| 100 calibration steps; this is illustrative rather than an immutable value. | |
| decision: >- | |
| Default to the paper's example interval 100, cache detached mu values | |
| between refreshes, serialize the cache, and expose the interval in the | |
| required calibration config for sensitivity analysis. | |
| sensitivity_required: true | |
| - id: LQ6_OPTIMIZER_SCHEDULE_STEPS | |
| status: assumption | |
| paper_status: >- | |
| LoopQ does not specify optimizer, learning rate, total calibration steps, | |
| weight decay, warmup, or scheduler. FlatQuant's settings use a different | |
| layer-wise reconstruction objective and cannot be silently inherited. | |
| decision: >- | |
| Require optimizer name, base learning rate, and final-phase steps explicitly | |
| with no defaults. Permit explicit LAS, transform, and CTA learning-rate | |
| overrides derived from the exact parameter allowlist so omitted component | |
| rates can be studied without changing the objective. Initially permit Adam | |
| or AdamW as caller-selected alternatives. Record both final-phase and global | |
| optimizer-step counts in every checkpoint, run sensitivity before claiming | |
| reproduction fidelity, and do not add an implicit scheduler. | |
| sensitivity_required: true | |
| - id: LQ6_PARAMETER_ALLOWLIST | |
| status: compatibility_decision | |
| paper_status: >- | |
| Section 4.4 allows shared transforms, LAS, selected loop transforms, and | |
| CTA parameters to update while the shared backbone remains fixed. | |
| decision: >- | |
| Construct an exact optimizer allowlist from only those four categories, | |
| reject duplicated ownership or any identity overlap with backbone | |
| parameters, verify the optimizer on every step, and clip their joint norm | |
| to the paper value 1.0. | |
| sensitivity_required: false | |
| - id: LQ7_OURO_CHECKPOINT_PIN | |
| status: compatibility_decision | |
| paper_status: >- | |
| The paper names Ouro 1.4B but does not provide a checkpoint revision or | |
| hash in the implementation appendix. | |
| decision: >- | |
| Validate against local ByteDance/Ouro-1.4B revision | |
| e3b1e0993b1231a51d6069a870476dda4162c00a. Pin config and modeling-code | |
| SHA256 values in ouro_mapping.yaml and inspect safetensors keys without | |
| loading parameter tensors. | |
| sensitivity_required: false | |
| - id: LQ7_VLLM_PACKED_MAPPING | |
| status: compatibility_decision | |
| paper_status: >- | |
| Appendix B.1 lists separate HF projection names, whereas the local vLLM | |
| backend packs Q/K/V and Gate/Up projections. | |
| decision: >- | |
| Preserve the four paper atomic groups while mapping Q/K/V to qkv_proj and | |
| Gate/Up to gate_up_proj consumers. Artifact provenance retains both the | |
| original HF weight names and packed runtime names. | |
| sensitivity_required: false | |
| - id: LQ7_RUNTIME_INTEGRATION_BOUNDARY | |
| status: compatibility_decision | |
| paper_status: "The paper does not specify vLLM hook APIs." | |
| decision: >- | |
| Keep the adapter backend-neutral: expose transform/LAS/QDQ preparation and | |
| exact CTA boundaries, but do not import the existing NVFP4/Residual plugin. | |
| Runtime model.py must later call these hooks at four listed consumers and | |
| three true loop boundaries before a smoke result can be claimed. | |
| sensitivity_required: false | |
| - id: LQ7_SELECTED_SLT_PACKED_WEIGHT_BLOCKER | |
| status: compatibility_decision | |
| paper_status: >- | |
| Appendix B.1 requires selected groups to use loop-dependent transforms | |
| and quantized weights. The local vLLM projections own one packed weight | |
| parameter shared by all recurrent calls. | |
| decision: >- | |
| Keep one shared base parameter for every projection and allocate four | |
| loop-specific packed RTN variants only for the paper-budget four selected | |
| groups. Merge Q/K/V and Gate/Up source shards with vLLM's existing weight | |
| loaders, then dispatch by exact loop index and restore the shared pointer. | |
| This correctness path requires eager execution. | |
| sensitivity_required: false | |
| - id: LQ7_SELECTED_WEIGHT_EAGER_DISPATCH | |
| status: compatibility_decision | |
| paper_status: "The paper does not specify vLLM compiled packed-weight dispatch." | |
| decision: >- | |
| Use scoped parameter-data substitution only in the explicit eager smoke | |
| path. Do not claim compiled/runtime speed results. A future native kernel | |
| or quant-method weight argument is required before torch.compile use. | |
| sensitivity_required: false | |
| - id: LQ7_CALIBRATION_STE | |
| status: compatibility_decision | |
| paper_status: >- | |
| The paper optimizes quantization parameters through RTN but does not | |
| specify the rounding surrogate used by its calibration implementation. | |
| decision: >- | |
| Use quantized values in the forward pass. For activation QDQ, retain the | |
| differentiable learned-LAS scale path and add an identity input path | |
| around rounding. For weight QDQ, RTN scales are inferred rather than | |
| learned parameters, so detach the quantizer internals and use the | |
| standard identity STE to the folded weight. This preserves transform | |
| gradients without retaining full-weight rounding graphs. Record this | |
| surrogate in every calibration artifact. | |
| sensitivity_required: true | |
| - id: LQ7_DIAGONAL_FISHER_COLLECTION | |
| status: assumption | |
| paper_status: >- | |
| Equation (8) defines the calibration loss and Equation (6) names a | |
| diagonal Fisher term, but the per-example/recurrent reduction is omitted. | |
| decision: >- | |
| For every one of the 24x4 projection groups, compute per-loop transform | |
| gradients of the full Equation (8) loss. During each statistics forward, | |
| use value-identical loop-specific shadow transforms so autograd assigns | |
| both activation-path and folded-weight-path contributions to the correct | |
| loop without changing the forward result. Average gradients over calibration | |
| examples and estimate diagonal Fisher as E[g^2] over examples and loops. | |
| Q/K/V and Gate/Up member VJPs are summed within their paper atomic group. | |
| sensitivity_required: true | |
| - id: LQ7_PROGRESSIVE_ADAPTATION_STEPS | |
| status: assumption | |
| paper_status: >- | |
| Appendix B.1 requires progressive selection with adaptation and score | |
| recomputation, but gives no optimizer-step count between rounds. | |
| decision: >- | |
| Require slt_round_steps explicitly. After each newly selected group, | |
| rebuild the exact allowlisted optimizer, adapt for that many examples, | |
| then recompute all remaining scores. Optimizer state is intentionally not | |
| carried across changed parameter ownership; study this choice. | |
| sensitivity_required: true | |
| - id: LQ8_HUGINN_FULL_GRADIENT_RECURRENCES | |
| status: compatibility_decision | |
| paper_status: >- | |
| LoopQ calibrates recurrent trajectories but does not evaluate Huginn or | |
| specify how its no-grad/with-grad recurrence sampler should be configured. | |
| decision: >- | |
| During LoopQ calibration, pass Huginn num_steps=[0,32] so every recurrence | |
| participates in the differentiable student trajectory. Teacher execution | |
| uses 32 no-grad recurrences. Do not accept Huginn's training default that | |
| would silently place early recurrence steps outside the calibration graph. | |
| sensitivity_required: true | |
| - id: LQ8_SAVED_TENSOR_CPU_OFFLOAD | |
| status: compatibility_decision | |
| paper_status: >- | |
| The paper does not specify where calibration autograd intermediates are | |
| stored or a peak-memory implementation for recurrent full-weight QDQ. | |
| decision: >- | |
| Store tensors saved for backward in CPU memory and restore the exact | |
| tensors to their original GPU | |
| on demand in backward. This changes placement and transfer cost only: it | |
| does not detach, compress, checkpoint, approximate, or change any LoopQ | |
| objective, gradient, quantizer, selection rule, or trainable parameter. | |
| Ouro uses pinned storage after inferred weight-quantizer internals are | |
| detached; Huginn conservatively uses pageable storage because its | |
| 32-recurrence saved set may exceed the CUDA pinned-host allocator budget. | |
| Record the offload mode in every calibration bundle and make no | |
| calibration-speed claim from this path. | |
| sensitivity_required: false | |
| - id: LQ9_DIRECT_RTN_CONTROL | |
| status: compatibility_decision | |
| paper_status: >- | |
| LoopQ compares against direct W4A4 and W4A8 RTN but does not publish a | |
| vLLM integration contract for those controls. | |
| decision: >- | |
| Quantize every mapped projection weight with symmetric group-32 W4 RTN | |
| once at load time and quantize each mapped linear input dynamically with | |
| symmetric group-32 A4 or A8 RTN. Use identity transforms, no LAS, no SLT, | |
| and no CTA. Keep the same architecture-specific vLLM GSM8K protocol as | |
| the corresponding BF16 and LoopQ rows. Label this direct RTN rather than | |
| LoopQ and make the mode mutually exclusive with calibrated artifacts and | |
| all ResidualQuant, MomentumQuant, and NVFP4 policies. | |
| sensitivity_required: false | |
| - id: LQ_TRAJECTORY_TRANSITION_TARGET | |
| status: compatibility_decision | |
| paper_status: "Equation (8) and Appendix B.4 target H_{t+1,0}, the next loop input." | |
| decision: >- | |
| The Ouro normalized loop output is directly fed into the next loop. | |
| Align CTA output with teacher_hidden[:-1], not teacher_hidden[1:]. | |
| Huginn uses the same recurrent-state boundary before the input adapter; | |
| input injection belongs to its recurrent function, not to CTA. | |
| sensitivity_required: false | |
| - id: LQ_HUGINN_MATCHED_RANDOM_STATE | |
| status: compatibility_decision | |
| paper_status: "Huginn is an architecture extension, not a paper backbone." | |
| decision: >- | |
| Fork and restore CPU/CUDA RNG around the teacher forward so the student | |
| sees the same random initial recurrent state and configured noise. | |
| Advance the outer generator only through the student forward. Record | |
| the seed; do not supervise independently sampled initial trajectories. | |
| sensitivity_required: false | |
| - id: LQ_GROUP_BATCHING_AND_WEIGHT_STE_STORAGE | |
| status: compatibility_decision | |
| paper_status: "The paper does not specify PyTorch allocation scheduling." | |
| decision: >- | |
| Vectorize group-wise RTN, preserving inferred-scale tie handling and | |
| supplied-scale derivatives. Detach folded weights only on the inferred | |
| quantizer input, retaining the existing identity STE to folded weights. | |
| This prevents saving an autograd quantizer graph that is discarded by | |
| the STE. The inverse-transform and full recurrent gradient paths remain. | |
| sensitivity_required: false | |
| - id: LQ_LINEAR_RECOMPUTATION | |
| status: compatibility_decision | |
| paper_status: "The paper does not specify autograd intermediate storage." | |
| decision: >- | |
| Offer --checkpoint-linears as an explicit non-reentrant recomputation | |
| option. Recompute only the deterministic transform/LAS/QDQ/linear function, | |
| capture the exact transform object and recurrence index, and never replay | |
| routing hooks. Retain every recurrent edge and all calibration parameters. | |
| Default saved-tensor offload remains available as the parity reference. | |
| Validate loss, SLT statistics and final component tensors before use. | |
| sensitivity_required: false | |
| - id: LQ_ACTIVATION_STE_EXACT_FORWARD | |
| status: compatibility_decision | |
| paper_status: "RTN forward values must match the exported QDQ runtime." | |
| decision: >- | |
| Compute activation STE as quantized + (original - original.detach()). | |
| Subtract identical values before adding to QDQ so the forward value is | |
| exact even in BF16. Preserve both the identity input gradient and the | |
| learned-LAS quantized branch gradient. Left-associated addition and | |
| subtraction introduce avoidable BF16 rounding and are not equivalent. | |
| sensitivity_required: false | |
| - id: LQ6_FP32_OBJECTIVE | |
| status: compatibility_decision | |
| paper_status: "Loss accumulation dtype is not specified." | |
| decision: >- | |
| Keep the backbone BF16 but compute selected-logit softmax/KL and hidden | |
| squared-error reductions in FP32. Preserve the sum and batchmean objective | |
| reductions; do not silently normalize by token count or hidden width. | |
| sensitivity_required: false | |
| - id: LQ6_SCAN_MU_CLOCK | |
| status: assumption | |
| paper_status: "Appendix B.4 gives a periodic calibration-step update, but not scan semantics." | |
| decision: >- | |
| The training mu cache refreshes once per eligible optimizer step. Each | |
| SLT sample uses independently computed detached mu at the current model; | |
| score collection never advances or replaces the training mu cache. | |
| sensitivity_required: true | |
| - id: LQ6_STREAMING_FISHER | |
| status: compatibility_decision | |
| paper_status: "Dataset reduction/storage precision is not specified." | |
| decision: >- | |
| Accumulate per-loop gradient sums and empirical squared-gradient sums in | |
| CPU FP64. Divide by the actual sample count; Fisher also averages loops. | |
| Store sufficient statistics and scan cursors in atomic resume checkpoints. | |
| sensitivity_required: false | |
| - id: LQ6_RESUME_PHASES | |
| status: assumption | |
| paper_status: "Optimizer-state behavior between progressive phases is unspecified." | |
| decision: >- | |
| Preserve optimizer moments within each SLT adaptation or final training | |
| phase; reset at phase boundaries as in the original real drivers. Resume | |
| restores component parameters, optimizer moments, mu, CPU/CUDA RNG, traces, | |
| completed scan summaries and partial scan sufficient statistics. Refuse | |
| changed source, configuration or calibration-text identities. | |
| sensitivity_required: true | |
| - id: LQ9_ABLATION_TRAINING | |
| status: assumption | |
| paper_status: "Table 2 names component removals but does not give exact ablation code." | |
| decision: >- | |
| Retrain each ablation. No-LAS learns a single shared scale vector per | |
| module initialized by the maximum across loop scales. No-SLT retains | |
| shared transforms and selects no groups. No-CTA bypasses the adapter, | |
| freezes its parameters and omits the transition loss. Keep the other | |
| recurrent loss terms. Persist component mode in exports and reject an | |
| ablation artifact supplied as a full LoopQ row. | |
| sensitivity_required: true | |
| - id: LQ6_ACTIVATION_STE_SENSITIVITY | |
| status: assumption | |
| paper_status: >- | |
| LoopQ does not specify the quantizer's backward estimator. The pinned | |
| FlatQuant quant_utils.py uses an STE on normalized values before clipping. | |
| decision: >- | |
| Expose identity and rounding activation STEs explicitly. Identity retains | |
| the prior direct activation gradient plus code-times-scale derivative. | |
| Rounding uses unit derivative through normalized rounding followed by | |
| clipping, yielding code-minus-normalized-value scale derivatives inside | |
| the range. Both preserve identical RTN forward values, group layout and | |
| bounds. Compare calibration convergence before choosing full-run settings. | |
| sensitivity_required: true | |
| - id: LQ_EXPORT_EXACT_LAS_PARAMETERS | |
| status: compatibility_decision | |
| paper_status: "Serialization precision is not specified." | |
| decision: >- | |
| Store raw FP32 trained log-scales alongside expanded positive scales and | |
| restore raw parameters exactly. Validate consistency of both payloads. | |
| Legacy exports without raw parameters remain readable. Export validation | |
| also checks component-ablation modes and calibration activation precision. | |
| sensitivity_required: false | |
| - id: LQ6_FP64_STATISTICS_AND_CLIP_NORM | |
| status: compatibility_decision | |
| paper_status: >- | |
| Fisher uses squared transform gradients and gradient clipping has norm 1; | |
| the paper does not specify the reduction dtype. | |
| decision: >- | |
| Square per-loop transform gradients in FP64 before Fisher reduction. | |
| Compute the global gradient clipping norm in FP64, retaining the norm-1 | |
| rule and rejection of genuine NaN/Inf. This prevents FP32 overflow from | |
| finite gradients observed in Ouro full calibration sample index 6. | |
| This changes floating-point training traces; do not bypass source-contract | |
| checks to resume an old run with the new implementation. | |
| sensitivity_required: false | |