J-Lens: Jacobians, readouts, and analysis outputs
A red-team characterisation of the Jacobian lens
(J_l = E[∂h_target/∂h_l]) — a lens computed from weights, not trained.
Code: https://github.com/jeeva2812/jlens-transformation
The usable result
Read with eigenvectors; steer with singular vectors.
| task | eigen | SVD | n per family |
|---|---|---|---|
| clears an axis probe | 39.6% | 20.8% | 96 |
| steers as its readout predicts | 44.4% | 75.0% | 36 |
The two tasks want different decompositions, and Weyl's inequality says why:
σ₁ ≥ |λ₁| for any matrix, so at a fixed injection norm a singular direction
must produce the larger downstream perturbation. Steering rewards magnitude;
reading rewards coherence.
The sharp version holds too. The size of SVD's steering advantage tracks the
departure from normality σ₁/|λ₁| at r = +0.68, and vanishes at the target
layer — where Φ(T,T) = I is normal and the two families must coincide.
Why J behaves this way
Treating depth as time makes a residual network dh/dt = F(t,h), whose
linearised sensitivity is the state-transition matrix. That identifies
J_ℓ = Φ(T,ℓ), confirmed by the semigroup property J_8 = J_16·M at
rel err 0.256 / cos 0.968 against an identity control at 0.744.
It follows that Φᵀ Φ is the Cauchy–Green strain tensor, so SVD gives
finite-time Lyapunov exponents and eigen gives asymptotic ones; that J − I ≈ A·Δt
recovers the generator; and that Jᵀ is the gradient propagator, so |λ|>1 is
an exploding-gradient channel. Training collapses those from 2102 → 32 at
layer 8 while rotation doubles.
A retraction, recorded rather than removed
An earlier version of this README claimed J amplifies the directions the model occupies least, and built a "control interface, not interpretation interface" framing on it. That was an artefact of one massive-activation direction carrying 53–99.8% of all variance. Projecting it out reverses the sign at every layer (corr(rank, occupancy) +0.229 → −0.856 at layer 8). The same artefact also inflated an "ablation rewards occupancy" result from −0.01 to +0.64.
What survives is narrower: the network gives its own sink direction (dimension 507) only 0.51–0.67× the gain of content directions, and that sink does not exist at initialisation — training builds it.
Methodological note: both claims passed a random-direction null. The null they needed was "remove the outlier activation dimensions first."
Contents
| path | what |
|---|---|
analysis/*.json |
~9,900 direction readouts with σ, share and axis-probe flags; eigen bundles; every intervention result |
jacobians/qwen05_em_*.pt |
Qwen2.5-0.5B base + 3 published EM organisms + a benign control |
jacobians/llama1b_em_*.pt |
Llama-3.2-1B base + 3 organisms — the EM result does not replicate here |
Jall_main.pt, lenses/ |
Olmo 3 7B Jacobians (original upload) |
INDEX.html |
47 entries: every question, what came back, and its status |
report.html, explorer.html |
the write-up and an interactive browser |
Caveats
- Every J is averaged over a fixed prompt set, and that is not a detail: the same ΔJ analysis gave opposite answers on Pile text and on chat-formatted prompts.
- Steering numbers are the corrected path (inject
v, readu). Anything citing 46%, 42%, or a sharp depth profile came from a version that injectedu, which is a type error. - The EM results are unreplicated cross-architecture and should not be cited.