File size: 27,509 Bytes
9118991
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
schema_version: 1
paper: "arXiv:2605.16343v1"
decisions:
  - id: LQ1_INTEGER_RANGE
    status: assumption
    paper_status: >-
      Equation (4) names q_min and q_max but does not give their numerical
      values or state whether the most-negative two's-complement code is used.
    decision: >-
      Use the cited FlatQuant symmetric RTN convention: q_min = -2^(b-1) and
      q_max = 2^(b-1)-1 ([-8, 7] for 4-bit and [-128, 127] for 8-bit).
    sensitivity_required: true

  - id: LQ1_SCALE_INITIALIZATION
    status: assumption
    paper_status: >-
      The paper learns activation ranges during calibration but does not state
      the initial static RTN scale estimator for weights or activations.
    decision: >-
      Initialize each group scale as max(abs(group)) / q_max. An all-zero
      group receives scale 1 and therefore remains exactly zero.
    sensitivity_required: true

  - id: LQ1_ROUNDING_TIES
    status: assumption
    paper_status: "The paper says round-to-nearest but does not specify tie handling."
    decision: "Round exact half-way values away from zero deterministically."
    sensitivity_required: false

  - id: LQ1_GROUP_AXIS
    status: assumption
    paper_status: "The paper specifies group-wise quantization but not tensor-axis layout."
    decision: >-
      Form consecutive groups along the last logical dimension for both
      weights and activations. Architecture adapters must document any layout
      conversion before invoking this quantizer.
    sensitivity_required: true

  - id: LQ1_TAIL_GROUP
    status: compatibility_decision
    paper_status: "Tail handling is not specified."
    decision: >-
      Quantize a final group shorter than 32 independently, without persistent
      padding. Its scale is computed only from real elements.
    sensitivity_required: false

  - id: LQ1_FAKE_QUANT_SCOPE
    status: compatibility_decision
    paper_status: >-
      The paper reports packed weights and quantized runtime kernels, but LQ1
      is the numerical quantizer acceptance phase.
    decision: >-
      LQ1 implements deterministic QDQ/fake quantization and emits integer
      codes plus scales. Packing and runtime kernels are deferred to LQ7/LQ8.
    sensitivity_required: false

  - id: LQ2_GROUP_SCALE_LAYOUT
    status: compatibility_decision
    paper_status: >-
      Section 4.1 specifies c_{t,l} per module and loop and explicitly counts
      O(TL) additional scalars. Appendix B.3 says "Per module, per loop".
    decision: >-
      Store one positive clipping multiplier for each module and loop. Apply it
      to dynamic per-token, per-group absmax scales (group size 32), matching
      the FlatQuant-style learned activation-clipping construction cited by the
      paper. This gives exactly O(TL) LAS parameters.
    sensitivity_required: false

  - id: LQ2_CALIBRATION_REDUCTION
    status: assumption
    paper_status: >-
      The initialization procedure for learned LAS ranges is not specified.
    decision: >-
      Initialize each module-loop clipping multiplier to 1.0. Dynamic base
      scales use each input token/group absmax. Subsequent trajectory
      calibration optimizes the log multiplier.
    sensitivity_required: true

  - id: LQ2_POSITIVE_PARAMETERIZATION
    status: compatibility_decision
    paper_status: "The paper requires scales but does not specify their parameterization."
    decision: >-
      Store trainable log-scales and exponentiate at use time, ensuring scales
      remain strictly positive during later gradient calibration.
    sensitivity_required: false

  - id: LQ2_ARCHITECTURE_LOOP_ROUTING
    status: compatibility_decision
    paper_status: >-
      The paper evaluates Ouro with four loops and does not evaluate Huginn.
    decision: >-
      Provide exact zero-based routing for Ouro loops [0,4) and Huginn
      recurrences [0,32), with no modulo, anchor, or fallback routing.
    sensitivity_required: false

  - id: LQ3_FLATQUANT_REFERENCE_PIN
    status: compatibility_decision
    paper_status: >-
      LoopQ cites FlatQuant and states that transforms are Kronecker-decomposed,
      but does not identify a code revision.
    decision: >-
      Pin the official ruikangliu/FlatQuant repository at commit
      9d88ffcb7d2c6bda59fb5c44dad36adc101aadb1. File hashes and the inspected
      folding contract are recorded in flatquant_reference.yaml.
    sensitivity_required: false

  - id: LQ3_FLATQUANT_SVD_PARAMETERIZATION
    status: compatibility_decision
    paper_status: >-
      LoopQ specifies FlatQuant-style Kronecker-decomposed transforms but does
      not state whether it uses FlatQuant's SVD or direct-inverse class, factor
      dimensions, add-diagonal flag, or architecture-specific initialization.
    decision: >-
      Use the default path of the paper-pinned official FlatQuant commit:
      SVDDecomposeTransMatrix with independently random-orthogonal U/V factors,
      trainable singular diagonals, and Cayley orthogonal parameterizations.
      Use explicit factor dimensions and omit FlatQuant's optional add_diag
      because LoopQ does not state that flag or its SmoothQuant initialization.
      Export effective Kronecker factors so runtime is independent of the
      training parameterization; retain legacy identity transform artifacts.
    sensitivity_required: true

  - id: LQ3_CALIBRATION_BLOCKER
    status: assumption
    paper_status: >-
      LoopQ omits optimizer, learning rate, epochs/steps, Kronecker factor
      dimensions, inverse method, and diagonal initialization. FlatQuant's
      official code uses AdamW and cosine scheduling, but repository defaults
      and published scripts differ and its layer-wise MSE objective is not
      LoopQ's trajectory objective.
    decision: >-
      Do not import FlatQuant's optimizer settings into LoopQ. Defer transform
      calibration parameterization and optimizer selection to LQ6, record the
      chosen sensitivity study there, and keep LQ3 limited to fold parity.
    sensitivity_required: true

  - id: LQ4_VARIANCE_ESTIMATOR
    status: compatibility_decision
    paper_status: >-
      Equation (6) explicitly expands variance with a 1/T factor but does not
      discuss library estimator settings.
    decision: >-
      Use population variance across the true loop dimension (unbiased=false),
      accumulate the score in deterministic CPU FP64, and sum every coordinate
      in the transform-weight group.
    sensitivity_required: false

  - id: LQ4_FISHER_AND_EPSILON
    status: assumption
    paper_status: >-
      The paper identifies psi as a diagonal Fisher estimate and epsilon as a
      positive stabilizer, but does not specify the Fisher sampling estimator,
      reduction, or epsilon value.
    decision: >-
      LQ4 consumes non-negative saved Fisher diagonals without inventing their
      collection procedure and uses epsilon=1e-8 by default. Fisher collection
      must be resolved and tested with trajectory calibration in LQ6.
    sensitivity_required: true

  - id: LQ4_PROGRESSIVE_RECOMPUTATION
    status: compatibility_decision
    paper_status: >-
      Section 4.2 requires alternation between adapting selected transforms and
      recomputing scores, but does not specify iteration counts or tie handling.
    decision: >-
      Select exactly one atomic transform-weight group per round, require fresh
      callback statistics after each selection, and break equal-score ties by
      lexical candidate name for deterministic artifacts.
    sensitivity_required: false

  - id: LQ4_HUGINN_BUDGET_ROUNDING
    status: assumption
    paper_status: >-
      Huginn is not evaluated. LoopQ reports that about 5% of transform
      parameters are selected, without a rounding rule for unequal atomic groups.
    decision: >-
      Set the Huginn parameter target to ceil(0.05 * all candidate transform
      parameters), select whole groups progressively, and stop after the first
      selection whose cumulative count meets or exceeds the target. Record the
      resulting atomic-group overshoot.
    sensitivity_required: true

  - id: LQ5_RMSNORM_DEFINITION
    status: assumption
    paper_status: >-
      Equation (7) says RMSNorm along the feature dimension but does not give
      epsilon or specify reuse of a backbone norm's learned weight.
    decision: >-
      Use unweighted RMS normalization x / sqrt(mean(x^2) + 1e-6), computed in
      FP32 and returned in the hidden-state dtype. Evaluate epsilon sensitivity
      before full calibration.
    sensitivity_required: true

  - id: LQ5_PARAMETER_SHAPES
    status: compatibility_decision
    paper_status: >-
      Equation (7) defines shared U,V and loop-dependent a_t,b_t,eta_t but only
      specifies adapter rank 8 in Table 6.
    decision: >-
      For hidden width d and rank r, use U,V in R^(d x r), a_t,b_t in R^d,
      and eta_t in R^r. Store parameters for actual transitions only: T-1 rows.
    sensitivity_required: false

  - id: LQ5_IDENTITY_INITIALIZATION
    status: assumption
    paper_status: "The paper does not state CTA initialization."
    decision: >-
      Initialize a_t=1, b_t=0, eta_t=0. Initialize shared U and V to the first
      r columns of the identity for deterministic nonzero gate gradients. This
      makes every CTA exactly identity before calibration.
    sensitivity_required: true

  - id: LQ5_TRANSITION_ROUTING
    status: compatibility_decision
    paper_status: >-
      Equation (7) applies A_t between loop t and t+1. Ouro has four loops;
      Huginn is not evaluated by the paper.
    decision: >-
      Route exactly three Ouro transitions and thirty-one Huginn transitions
      with zero-based indices. Reject out-of-range indices; never use modulo,
      Residual/Momentum anchors, or mixed-precision schedule boundaries.
    sensitivity_required: false

  - id: LQ6_LOGIT_TOPK_KL
    status: compatibility_decision
    paper_status: >-
      Table 6 specifies teacher top-k logits 1000 and temperature 1, but does
      not state whether KL is renormalized over the retained support or how
      reductions are applied.
    decision: >-
      Select the teacher's top-1000 indices, renormalize teacher and student
      distributions on that same support, and compute KL(teacher || student)
      with PyTorch batchmean reduction. For synthetic vocabularies smaller than
      1000 only, use the whole vocabulary.
    sensitivity_required: true

  - id: LQ6_TRAJECTORY_REDUCTION
    status: assumption
    paper_status: >-
      Equation (8) shows squared L2 norms and sums over loops but does not give
      batch/token reduction or normalization details.
    decision: >-
      Use squared L2 sums across every non-loop coordinate and sum across loops
      and transitions exactly as displayed. Apply lambda=0.1 to all trajectory
      and transition terms. Record each weighted term separately.
    sensitivity_required: true

  - id: LQ6_ADAPTIVE_MU_REFRESH
    status: compatibility_decision
    paper_status: >-
      Appendix B.4 states that mu_t is updated periodically, for example every
      100 calibration steps; this is illustrative rather than an immutable value.
    decision: >-
      Default to the paper's example interval 100, cache detached mu values
      between refreshes, serialize the cache, and expose the interval in the
      required calibration config for sensitivity analysis.
    sensitivity_required: true

  - id: LQ6_OPTIMIZER_SCHEDULE_STEPS
    status: assumption
    paper_status: >-
      LoopQ does not specify optimizer, learning rate, total calibration steps,
      weight decay, warmup, or scheduler. FlatQuant's settings use a different
      layer-wise reconstruction objective and cannot be silently inherited.
    decision: >-
      Require optimizer name, base learning rate, and final-phase steps explicitly
      with no defaults. Permit explicit LAS, transform, and CTA learning-rate
      overrides derived from the exact parameter allowlist so omitted component
      rates can be studied without changing the objective. Initially permit Adam
      or AdamW as caller-selected alternatives. Record both final-phase and global
      optimizer-step counts in every checkpoint, run sensitivity before claiming
      reproduction fidelity, and do not add an implicit scheduler.
    sensitivity_required: true

  - id: LQ6_PARAMETER_ALLOWLIST
    status: compatibility_decision
    paper_status: >-
      Section 4.4 allows shared transforms, LAS, selected loop transforms, and
      CTA parameters to update while the shared backbone remains fixed.
    decision: >-
      Construct an exact optimizer allowlist from only those four categories,
      reject duplicated ownership or any identity overlap with backbone
      parameters, verify the optimizer on every step, and clip their joint norm
      to the paper value 1.0.
    sensitivity_required: false

  - id: LQ7_OURO_CHECKPOINT_PIN
    status: compatibility_decision
    paper_status: >-
      The paper names Ouro 1.4B but does not provide a checkpoint revision or
      hash in the implementation appendix.
    decision: >-
      Validate against local ByteDance/Ouro-1.4B revision
      e3b1e0993b1231a51d6069a870476dda4162c00a. Pin config and modeling-code
      SHA256 values in ouro_mapping.yaml and inspect safetensors keys without
      loading parameter tensors.
    sensitivity_required: false

  - id: LQ7_VLLM_PACKED_MAPPING
    status: compatibility_decision
    paper_status: >-
      Appendix B.1 lists separate HF projection names, whereas the local vLLM
      backend packs Q/K/V and Gate/Up projections.
    decision: >-
      Preserve the four paper atomic groups while mapping Q/K/V to qkv_proj and
      Gate/Up to gate_up_proj consumers. Artifact provenance retains both the
      original HF weight names and packed runtime names.
    sensitivity_required: false

  - id: LQ7_RUNTIME_INTEGRATION_BOUNDARY
    status: compatibility_decision
    paper_status: "The paper does not specify vLLM hook APIs."
    decision: >-
      Keep the adapter backend-neutral: expose transform/LAS/QDQ preparation and
      exact CTA boundaries, but do not import the existing NVFP4/Residual plugin.
      Runtime model.py must later call these hooks at four listed consumers and
      three true loop boundaries before a smoke result can be claimed.
    sensitivity_required: false

  - id: LQ7_SELECTED_SLT_PACKED_WEIGHT_BLOCKER
    status: compatibility_decision
    paper_status: >-
      Appendix B.1 requires selected groups to use loop-dependent transforms
      and quantized weights. The local vLLM projections own one packed weight
      parameter shared by all recurrent calls.
    decision: >-
      Keep one shared base parameter for every projection and allocate four
      loop-specific packed RTN variants only for the paper-budget four selected
      groups. Merge Q/K/V and Gate/Up source shards with vLLM's existing weight
      loaders, then dispatch by exact loop index and restore the shared pointer.
      This correctness path requires eager execution.
    sensitivity_required: false

  - id: LQ7_SELECTED_WEIGHT_EAGER_DISPATCH
    status: compatibility_decision
    paper_status: "The paper does not specify vLLM compiled packed-weight dispatch."
    decision: >-
      Use scoped parameter-data substitution only in the explicit eager smoke
      path. Do not claim compiled/runtime speed results. A future native kernel
      or quant-method weight argument is required before torch.compile use.
    sensitivity_required: false

  - id: LQ7_CALIBRATION_STE
    status: compatibility_decision
    paper_status: >-
      The paper optimizes quantization parameters through RTN but does not
      specify the rounding surrogate used by its calibration implementation.
    decision: >-
      Use quantized values in the forward pass. For activation QDQ, retain the
      differentiable learned-LAS scale path and add an identity input path
      around rounding. For weight QDQ, RTN scales are inferred rather than
      learned parameters, so detach the quantizer internals and use the
      standard identity STE to the folded weight. This preserves transform
      gradients without retaining full-weight rounding graphs. Record this
      surrogate in every calibration artifact.
    sensitivity_required: true

  - id: LQ7_DIAGONAL_FISHER_COLLECTION
    status: assumption
    paper_status: >-
      Equation (8) defines the calibration loss and Equation (6) names a
      diagonal Fisher term, but the per-example/recurrent reduction is omitted.
    decision: >-
      For every one of the 24x4 projection groups, compute per-loop transform
      gradients of the full Equation (8) loss. During each statistics forward,
      use value-identical loop-specific shadow transforms so autograd assigns
      both activation-path and folded-weight-path contributions to the correct
      loop without changing the forward result. Average gradients over calibration
      examples and estimate diagonal Fisher as E[g^2] over examples and loops.
      Q/K/V and Gate/Up member VJPs are summed within their paper atomic group.
    sensitivity_required: true

  - id: LQ7_PROGRESSIVE_ADAPTATION_STEPS
    status: assumption
    paper_status: >-
      Appendix B.1 requires progressive selection with adaptation and score
      recomputation, but gives no optimizer-step count between rounds.
    decision: >-
      Require slt_round_steps explicitly. After each newly selected group,
      rebuild the exact allowlisted optimizer, adapt for that many examples,
      then recompute all remaining scores. Optimizer state is intentionally not
      carried across changed parameter ownership; study this choice.
    sensitivity_required: true

  - id: LQ8_HUGINN_FULL_GRADIENT_RECURRENCES
    status: compatibility_decision
    paper_status: >-
      LoopQ calibrates recurrent trajectories but does not evaluate Huginn or
      specify how its no-grad/with-grad recurrence sampler should be configured.
    decision: >-
      During LoopQ calibration, pass Huginn num_steps=[0,32] so every recurrence
      participates in the differentiable student trajectory. Teacher execution
      uses 32 no-grad recurrences. Do not accept Huginn's training default that
      would silently place early recurrence steps outside the calibration graph.
    sensitivity_required: true

  - id: LQ8_SAVED_TENSOR_CPU_OFFLOAD
    status: compatibility_decision
    paper_status: >-
      The paper does not specify where calibration autograd intermediates are
      stored or a peak-memory implementation for recurrent full-weight QDQ.
    decision: >-
      Store tensors saved for backward in CPU memory and restore the exact
      tensors to their original GPU
      on demand in backward. This changes placement and transfer cost only: it
      does not detach, compress, checkpoint, approximate, or change any LoopQ
      objective, gradient, quantizer, selection rule, or trainable parameter.
      Ouro uses pinned storage after inferred weight-quantizer internals are
      detached; Huginn conservatively uses pageable storage because its
      32-recurrence saved set may exceed the CUDA pinned-host allocator budget.
      Record the offload mode in every calibration bundle and make no
      calibration-speed claim from this path.
    sensitivity_required: false

  - id: LQ9_DIRECT_RTN_CONTROL
    status: compatibility_decision
    paper_status: >-
      LoopQ compares against direct W4A4 and W4A8 RTN but does not publish a
      vLLM integration contract for those controls.
    decision: >-
      Quantize every mapped projection weight with symmetric group-32 W4 RTN
      once at load time and quantize each mapped linear input dynamically with
      symmetric group-32 A4 or A8 RTN. Use identity transforms, no LAS, no SLT,
      and no CTA. Keep the same architecture-specific vLLM GSM8K protocol as
      the corresponding BF16 and LoopQ rows. Label this direct RTN rather than
      LoopQ and make the mode mutually exclusive with calibrated artifacts and
      all ResidualQuant, MomentumQuant, and NVFP4 policies.
    sensitivity_required: false

  - id: LQ_TRAJECTORY_TRANSITION_TARGET
    status: compatibility_decision
    paper_status: "Equation (8) and Appendix B.4 target H_{t+1,0}, the next loop input."
    decision: >-
      The Ouro normalized loop output is directly fed into the next loop.
      Align CTA output with teacher_hidden[:-1], not teacher_hidden[1:].
      Huginn uses the same recurrent-state boundary before the input adapter;
      input injection belongs to its recurrent function, not to CTA.
    sensitivity_required: false

  - id: LQ_HUGINN_MATCHED_RANDOM_STATE
    status: compatibility_decision
    paper_status: "Huginn is an architecture extension, not a paper backbone."
    decision: >-
      Fork and restore CPU/CUDA RNG around the teacher forward so the student
      sees the same random initial recurrent state and configured noise.
      Advance the outer generator only through the student forward. Record
      the seed; do not supervise independently sampled initial trajectories.
    sensitivity_required: false

  - id: LQ_GROUP_BATCHING_AND_WEIGHT_STE_STORAGE
    status: compatibility_decision
    paper_status: "The paper does not specify PyTorch allocation scheduling."
    decision: >-
      Vectorize group-wise RTN, preserving inferred-scale tie handling and
      supplied-scale derivatives. Detach folded weights only on the inferred
      quantizer input, retaining the existing identity STE to folded weights.
      This prevents saving an autograd quantizer graph that is discarded by
      the STE. The inverse-transform and full recurrent gradient paths remain.
    sensitivity_required: false

  - id: LQ_LINEAR_RECOMPUTATION
    status: compatibility_decision
    paper_status: "The paper does not specify autograd intermediate storage."
    decision: >-
      Offer --checkpoint-linears as an explicit non-reentrant recomputation
      option. Recompute only the deterministic transform/LAS/QDQ/linear function,
      capture the exact transform object and recurrence index, and never replay
      routing hooks. Retain every recurrent edge and all calibration parameters.
      Default saved-tensor offload remains available as the parity reference.
      Validate loss, SLT statistics and final component tensors before use.
    sensitivity_required: false

  - id: LQ_ACTIVATION_STE_EXACT_FORWARD
    status: compatibility_decision
    paper_status: "RTN forward values must match the exported QDQ runtime."
    decision: >-
      Compute activation STE as quantized + (original - original.detach()).
      Subtract identical values before adding to QDQ so the forward value is
      exact even in BF16. Preserve both the identity input gradient and the
      learned-LAS quantized branch gradient. Left-associated addition and
      subtraction introduce avoidable BF16 rounding and are not equivalent.
    sensitivity_required: false

  - id: LQ6_FP32_OBJECTIVE
    status: compatibility_decision
    paper_status: "Loss accumulation dtype is not specified."
    decision: >-
      Keep the backbone BF16 but compute selected-logit softmax/KL and hidden
      squared-error reductions in FP32. Preserve the sum and batchmean objective
      reductions; do not silently normalize by token count or hidden width.
    sensitivity_required: false

  - id: LQ6_SCAN_MU_CLOCK
    status: assumption
    paper_status: "Appendix B.4 gives a periodic calibration-step update, but not scan semantics."
    decision: >-
      The training mu cache refreshes once per eligible optimizer step. Each
      SLT sample uses independently computed detached mu at the current model;
      score collection never advances or replaces the training mu cache.
    sensitivity_required: true

  - id: LQ6_STREAMING_FISHER
    status: compatibility_decision
    paper_status: "Dataset reduction/storage precision is not specified."
    decision: >-
      Accumulate per-loop gradient sums and empirical squared-gradient sums in
      CPU FP64. Divide by the actual sample count; Fisher also averages loops.
      Store sufficient statistics and scan cursors in atomic resume checkpoints.
    sensitivity_required: false

  - id: LQ6_RESUME_PHASES
    status: assumption
    paper_status: "Optimizer-state behavior between progressive phases is unspecified."
    decision: >-
      Preserve optimizer moments within each SLT adaptation or final training
      phase; reset at phase boundaries as in the original real drivers. Resume
      restores component parameters, optimizer moments, mu, CPU/CUDA RNG, traces,
      completed scan summaries and partial scan sufficient statistics. Refuse
      changed source, configuration or calibration-text identities.
    sensitivity_required: true

  - id: LQ9_ABLATION_TRAINING
    status: assumption
    paper_status: "Table 2 names component removals but does not give exact ablation code."
    decision: >-
      Retrain each ablation. No-LAS learns a single shared scale vector per
      module initialized by the maximum across loop scales. No-SLT retains
      shared transforms and selects no groups. No-CTA bypasses the adapter,
      freezes its parameters and omits the transition loss. Keep the other
      recurrent loss terms. Persist component mode in exports and reject an
      ablation artifact supplied as a full LoopQ row.
    sensitivity_required: true

  - id: LQ6_ACTIVATION_STE_SENSITIVITY
    status: assumption
    paper_status: >-
      LoopQ does not specify the quantizer's backward estimator. The pinned
      FlatQuant quant_utils.py uses an STE on normalized values before clipping.
    decision: >-
      Expose identity and rounding activation STEs explicitly. Identity retains
      the prior direct activation gradient plus code-times-scale derivative.
      Rounding uses unit derivative through normalized rounding followed by
      clipping, yielding code-minus-normalized-value scale derivatives inside
      the range. Both preserve identical RTN forward values, group layout and
      bounds. Compare calibration convergence before choosing full-run settings.
    sensitivity_required: true

  - id: LQ_EXPORT_EXACT_LAS_PARAMETERS
    status: compatibility_decision
    paper_status: "Serialization precision is not specified."
    decision: >-
      Store raw FP32 trained log-scales alongside expanded positive scales and
      restore raw parameters exactly. Validate consistency of both payloads.
      Legacy exports without raw parameters remain readable. Export validation
      also checks component-ablation modes and calibration activation precision.
    sensitivity_required: false

  - id: LQ6_FP64_STATISTICS_AND_CLIP_NORM
    status: compatibility_decision
    paper_status: >-
      Fisher uses squared transform gradients and gradient clipping has norm 1;
      the paper does not specify the reduction dtype.
    decision: >-
      Square per-loop transform gradients in FP64 before Fisher reduction.
      Compute the global gradient clipping norm in FP64, retaining the norm-1
      rule and rejection of genuine NaN/Inf. This prevents FP32 overflow from
      finite gradients observed in Ouro full calibration sample index 6.
      This changes floating-point training traces; do not bypass source-contract
      checks to resume an old run with the new implementation.
    sensitivity_required: false