hv-two-factor-lm-1024

A NumPy-only two-factor language model trained without any deep learning framework. Hand-coded gradients for the LM head, closed-form ridge regression for the two diagnostic heads. Demonstrates the loss- informativeness phenomenon in a non-backprop setting.

Model size: 44.8 KB (two_factor_lm.npz) Training time: ~15 seconds on a laptop CPU Dependencies: NumPy only

Headline numbers

quantity without marginal loss with marginal loss
` w_alpha
mean output KL at (α=0.65, f=0.00) 0.108 0.073
mean output KL at (α=0.50, f=0.30) 0.015 0.004
mean output KL at (α=0.35, f=0.30) 0.017 0.005
comp R² (ridge head) 0.822 0.822
struct R² (ridge head) 0.054 0.054

The interesting row is the KL: marginal supervision improves the batch- mean output distribution by 2–4× across the conditioning grid, with no regression at any point.

Model Description

A causal language model over an 8-class alphabet with an explicit conditioning path (alpha, inject) → output. Trained on synthetic sequences generated by a controlled compositional motif process.

No PyTorch. No autograd framework. The encoder is a fixed random hypervector context. The LM head is trained by hand-coded Adam with hand-derived gradients. The two diagnostic heads (compositional, structural) are fit by closed-form ridge regression.

Architecture

Encoder (frozen, no learning):

h_t      = decay · h_{t-1} + token_hv[x_{t-1}]
h_norm_t = h_t / ||h_t||

LM head (learned):

scores_t = W_lm @ h_norm_t + alpha · w_alpha + inject · w_inject + b_lm
P(x_t | x_<t, alpha, inject) = softmax(scores_t)

Diagnostic heads (ridge, closed form):

r_hat    = W_comp   @ [h_norm; 1]    -> 4-simplex compositional target
s_hat    = W_struct @ [h_norm; 1]    -> scalar log1p(CSI_struct)

Training objective

L = L_lm + lambda_marginal * L_marginal
  • L_lm = per-token cross-entropy on the next token.
  • L_marginal = KL(mean_{b,t} P || p_target(alpha, f)).

The phenomenon. Without lambda_marginal, the encoder already carries a weak next-token predictor, so the LM loss gives only mild gradient toward conditioning-dependent output. With lambda_marginal > 0, the marginal distribution is directly constrained and the batch-mean output KL drops by 2-4x across the (alpha, f) grid.

Evaluation

Training (D=1024, seq_len=32, batch=16, 1500 steps, Adam 1e-3):

step L_lm (lambda=0) L_lm (lambda=3) L_marg (lambda=0) L_marg (lambda=3)
0 2.077 2.077 0.000 0.104
500 1.767 1.785 0.000 0.009
1000 1.786 1.795 0.000 0.002
1499 1.895 1.903 0.000 0.002

Marginal check — batch-mean output KL at final weights:

alpha f KL (lambda=0) KL (lambda=3) improvement
0.35 0.00 0.0229 0.0232 ~1x
0.35 0.15 0.0047 0.0034 1.4x
0.35 0.30 0.0171 0.0048 3.6x
0.50 0.00 0.0352 0.0253 1.4x
0.50 0.15 0.0008 0.0003 2.7x
0.50 0.30 0.0146 0.0042 3.5x
0.65 0.00 0.1080 0.0734 1.5x
0.65 0.15 0.0080 0.0051 1.6x
0.65 0.30 0.0150 0.0055 2.7x

Encoder diagnostics (ridge heads, closed-form):

head R^2
comp (r_target) 0.822
struct (log1p CSI_struct) 0.054

Intended use

  • Reference implementation of the two-factor architecture without any framework dependency.
  • Loss-informativeness demonstration.
  • Educational.

Limitations

  • Not competitive with any modern LM.
  • Struct head fails. R^2 = 0.054.
  • Synthetic data only.
  • ||w_alpha|| is a misleading metric.
  • Adam is hand-coded.

Comparison to the PyTorch companion

property hv-two-factor-lm two-factor-lm (PyTorch)
framework NumPy PyTorch
encoder hypervector context causal transformer
LM head hand-coded gradients autograd
diagnostic heads ridge learned heads
model size 44.8 KB ~2 MB
training time ~15 s CPU ~4 min CPU
comp R^2 0.82 0.95
struct R^2 0.05 0.65
marginal gain 2-4x ~50x

How to use

import numpy as np
from hv_two_factor_lm import HvTwoFactorLM

data = np.load('two_factor_lm.npz')
model = HvTwoFactorLM(D=int(data['D']), K=int(data['K']),
                      decay=float(data['decay']))
model.token_hv   = data['token_hv']
model.W_lm       = data['W_lm']
model.b_lm       = data['b_lm']
model.w_alpha    = data['w_alpha']
model.w_inject   = data['w_inject']
model.W_comp     = data['W_comp']
model.W_struct   = data['W_struct']
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support