# Architecture and tensor flow Let `S(x) = x / sqrt(mean(x²) + 1e-6)` along the feature dimension. This normalization gives approximately unit RMS, so the radius in 256-dimensional Euclidean coordinates is approximately 16. Learned RMSNorm layers additionally multiply by a learned feature-wise gain. For layer `l`, write its input as `h`, its post-Attention state as `a`, and its final state as `y`. ## Attention content and variance From `RMSNorm(h)`, linear projections produce Q, K, V with 8 heads of width 32. Q and K receive RoPE (base 10,000). Causal, document-isolated Attention weights are conceptually: `P = softmax(Q Kᵀ / sqrt(32) + mask)` The implementation computes two moments with the same weights: `mean = P V` `second = P (V²)` `variance = max(second - mean², 0)` A paired kernel evaluates `[V, V²]`; zero-padding Q/K to width 64 allows the paired call while explicitly preserving the original `1/sqrt(32)` scale. Variance is calculated in each head's own Value coordinates, before the output gate and output projection. It measures variation across attended positions, not a comparison of unrelated coordinates across heads. The first token in each document has exactly one eligible Value. Its variance is explicitly set to zero to avoid amplifying BF16 subtraction noise. The weighted mean follows the content path: `gate = sigmoid(W_gate RMSNorm(h))` `a = S(h + W_o concat_heads(gate * mean))` The flattened variance follows its own path. With `r = log1p(variance / 0.1)` and `m = mean(r²)`: `direction = r / sqrt(m + eps)` `magnitude = log1p(sqrt(m + eps)) - log1p(sqrt(eps))` The code uses an algebraically equivalent rationalized expression for magnitude to reduce cancellation. Concatenation yields 257 features. A bias-free 257 → 1,792 projection maps them into the two SwiGLU branches. This projection starts at zero during pretraining. ## Previous-layer differences For the previous layer, record: `dA = S(a_previous - h_previous)` `dF = S(y_previous - a_previous)` Both remain 256-dimensional. A learned query from the current `a` and shared keys from `dA` and `dF` are 16-dimensional. Their dot products divided by 4 produce two softmax weights, `pA` and `pF`. The contribution is: `delta_input = pA * W_A dA + pF * W_F dF` `W_A` and `W_F` are independent 256 → 1,792 matrices. The first layer omits this reader because it has no previous layer. Later layers consume only the immediately previous pair. Gradients flow through both differences and routing weights during training. ## SwiGLU and output `pre = W_base RMSNorm(a) + delta_input + W_var [direction; magnitude]` `g, u = split(pre)` `y = S(a + W_down dropout(SiLU(g) * u))` After layer 7, a final learned RMSNorm and an untied 256 → 16,384 LM head produce logits. The main residual stream remains full-width throughout. There is no cross-token recurrent cache or separate 16-dimensional persistent state in this model. ## Implementation boundaries Training used BF16 autocast, FP32 parameters and residuals, and FlashAttention. The public default uses SDPA for portability; backend rounding can differ. The published evaluation values came from the original evaluation backend, not a rerun of the entire suite on CPU. The public config adds standard Transformers shape fields, registers AutoModel, sets `use_cache=False`, and selects SDPA. The trained model core and all parameter tensors are retained. See the release test report for CPU loading and adapter checks.