J-Lens: Jacobians, readouts, and analysis outputs

A red-team characterisation of the Jacobian lens (J_l = E[∂h_target/∂h_l]) — a lens computed from weights, not trained.

Code: https://github.com/jeeva2812/jlens-transformation

The usable result

Read with eigenvectors; steer with singular vectors.

task eigen SVD n per family
clears an axis probe 39.6% 20.8% 96
steers as its readout predicts 44.4% 75.0% 36

The two tasks want different decompositions, and Weyl's inequality says why: σ₁ ≥ |λ₁| for any matrix, so at a fixed injection norm a singular direction must produce the larger downstream perturbation. Steering rewards magnitude; reading rewards coherence.

The sharp version holds too. The size of SVD's steering advantage tracks the departure from normality σ₁/|λ₁| at r = +0.68, and vanishes at the target layer — where Φ(T,T) = I is normal and the two families must coincide.

Why J behaves this way

Treating depth as time makes a residual network dh/dt = F(t,h), whose linearised sensitivity is the state-transition matrix. That identifies J_ℓ = Φ(T,ℓ), confirmed by the semigroup property J_8 = J_16·M at rel err 0.256 / cos 0.968 against an identity control at 0.744.

It follows that Φᵀ Φ is the Cauchy–Green strain tensor, so SVD gives finite-time Lyapunov exponents and eigen gives asymptotic ones; that J − I ≈ A·Δt recovers the generator; and that Jᵀ is the gradient propagator, so |λ|>1 is an exploding-gradient channel. Training collapses those from 2102 → 32 at layer 8 while rotation doubles.

A retraction, recorded rather than removed

An earlier version of this README claimed J amplifies the directions the model occupies least, and built a "control interface, not interpretation interface" framing on it. That was an artefact of one massive-activation direction carrying 53–99.8% of all variance. Projecting it out reverses the sign at every layer (corr(rank, occupancy) +0.229 → −0.856 at layer 8). The same artefact also inflated an "ablation rewards occupancy" result from −0.01 to +0.64.

What survives is narrower: the network gives its own sink direction (dimension 507) only 0.51–0.67× the gain of content directions, and that sink does not exist at initialisation — training builds it.

Methodological note: both claims passed a random-direction null. The null they needed was "remove the outlier activation dimensions first."

Contents

path what
analysis/*.json ~9,900 direction readouts with σ, share and axis-probe flags; eigen bundles; every intervention result
jacobians/qwen05_em_*.pt Qwen2.5-0.5B base + 3 published EM organisms + a benign control
jacobians/llama1b_em_*.pt Llama-3.2-1B base + 3 organisms — the EM result does not replicate here
Jall_main.pt, lenses/ Olmo 3 7B Jacobians (original upload)
INDEX.html 47 entries: every question, what came back, and its status
report.html, explorer.html the write-up and an interactive browser

Caveats

  • Every J is averaged over a fixed prompt set, and that is not a detail: the same ΔJ analysis gave opposite answers on Pile text and on chat-formatted prompts.
  • Steering numbers are the corrected path (inject v, read u). Anything citing 46%, 42%, or a sharp depth profile came from a version that injected u, which is a type error.
  • The EM results are unreplicated cross-architecture and should not be cited.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support