Spectral-World-Models-Reproducibility / STABILITY_REFACTOR.md
kiruluta's picture
Upload folder using huggingface_hub
4fd79a1 verified
|
Raw History Blame Contribute Delete
2.25 kB

SWM stability-attribution refactor

This revision responds directly to the five-seed H=5/10/20/30 finding that full SWM remained much flatter than the Koopman baseline, while the existing Frobenius-energy stability penalty had effectively no measurable effect and removing the spectral transition improved rollout MSE.

New controlled variant

swm_spectral_norm retains the SWM wavelet image state, DCT text state, fusion, and low-rank action-conditioned operator dictionary. For each discrete action it explicitly constructs the effective linear operator

A(a) = W0 + sum_r beta_r(a) U_r V_r

and caps its largest singular value at rho=0.95 in the forward pass. The unconstrained nonlinear residual is omitted in this diagnostic variant, making the contraction mechanism directly testable rather than conflating a linear operator bound with an unbounded learned correction. tanh is retained and is non-expansive.

The auxiliary stability loss now measures violations of the same spectral-norm threshold before forward rescaling. The benchmark records raw and effective operator norms so a claimed constraint can be checked from the results rather than inferred from a training hyperparameter.

Attribution experiment

The targeted benchmark compares koopman, neural_operator, swm, swm_spectral_norm, no_spectral_transition, and no_stability_penalty using identical seeds. Default rollout horizons are 5, 10, 20, 30, 50, and 100. A separate held-out long-sequence dataset is generated automatically, so the original short training data are not changed to make the long-horizon test easier.

Run on CUDA:

PYTHONPATH=. python scripts/run_stability_benchmark.py --seeds 0 1 2 3 4 --horizons 5 10 20 30 50 100 --device cuda

Primary output: results/stability_attribution/metrics_summary.csv.

Interpretation rule

The hard-constrained variant supports the stability-mechanism hypothesis only if its long-horizon behavior improves reproducibly relative to ordinary SWM/current-penalty SWM without an unacceptable collapse in one-step predictive or multimodal metrics. If it does not, the correct conclusion is that this benchmark does not attribute SWM's observed rollout flatness to the proposed operator-norm mechanism.