hv-two-factor-lm-1024
A NumPy-only two-factor language model trained without any deep learning framework. Hand-coded gradients for the LM head, closed-form ridge regression for the two diagnostic heads. Demonstrates the loss- informativeness phenomenon in a non-backprop setting.
Model size: 44.8 KB (two_factor_lm.npz)
Training time: ~15 seconds on a laptop CPU
Dependencies: NumPy only
Headline numbers
| quantity | without marginal loss | with marginal loss |
|---|---|---|
| ` | w_alpha | |
| mean output KL at (α=0.65, f=0.00) | 0.108 | 0.073 |
| mean output KL at (α=0.50, f=0.30) | 0.015 | 0.004 |
| mean output KL at (α=0.35, f=0.30) | 0.017 | 0.005 |
| comp R² (ridge head) | 0.822 | 0.822 |
| struct R² (ridge head) | 0.054 | 0.054 |
The interesting row is the KL: marginal supervision improves the batch- mean output distribution by 2–4× across the conditioning grid, with no regression at any point.
Model Description
A causal language model over an 8-class alphabet with an explicit
conditioning path (alpha, inject) → output. Trained on synthetic
sequences generated by a controlled compositional motif process.
No PyTorch. No autograd framework. The encoder is a fixed random hypervector context. The LM head is trained by hand-coded Adam with hand-derived gradients. The two diagnostic heads (compositional, structural) are fit by closed-form ridge regression.
Architecture
Encoder (frozen, no learning):
h_t = decay · h_{t-1} + token_hv[x_{t-1}]
h_norm_t = h_t / ||h_t||
LM head (learned):
scores_t = W_lm @ h_norm_t + alpha · w_alpha + inject · w_inject + b_lm
P(x_t | x_<t, alpha, inject) = softmax(scores_t)
Diagnostic heads (ridge, closed form):
r_hat = W_comp @ [h_norm; 1] -> 4-simplex compositional target
s_hat = W_struct @ [h_norm; 1] -> scalar log1p(CSI_struct)
Training objective
L = L_lm + lambda_marginal * L_marginal
L_lm= per-token cross-entropy on the next token.L_marginal= KL(mean_{b,t} P || p_target(alpha, f)).
The phenomenon. Without lambda_marginal, the encoder already carries
a weak next-token predictor, so the LM loss gives only mild gradient
toward conditioning-dependent output. With lambda_marginal > 0, the
marginal distribution is directly constrained and the batch-mean output
KL drops by 2-4x across the (alpha, f) grid.
Evaluation
Training (D=1024, seq_len=32, batch=16, 1500 steps, Adam 1e-3):
| step | L_lm (lambda=0) | L_lm (lambda=3) | L_marg (lambda=0) | L_marg (lambda=3) |
|---|---|---|---|---|
| 0 | 2.077 | 2.077 | 0.000 | 0.104 |
| 500 | 1.767 | 1.785 | 0.000 | 0.009 |
| 1000 | 1.786 | 1.795 | 0.000 | 0.002 |
| 1499 | 1.895 | 1.903 | 0.000 | 0.002 |
Marginal check — batch-mean output KL at final weights:
| alpha | f | KL (lambda=0) | KL (lambda=3) | improvement |
|---|---|---|---|---|
| 0.35 | 0.00 | 0.0229 | 0.0232 | ~1x |
| 0.35 | 0.15 | 0.0047 | 0.0034 | 1.4x |
| 0.35 | 0.30 | 0.0171 | 0.0048 | 3.6x |
| 0.50 | 0.00 | 0.0352 | 0.0253 | 1.4x |
| 0.50 | 0.15 | 0.0008 | 0.0003 | 2.7x |
| 0.50 | 0.30 | 0.0146 | 0.0042 | 3.5x |
| 0.65 | 0.00 | 0.1080 | 0.0734 | 1.5x |
| 0.65 | 0.15 | 0.0080 | 0.0051 | 1.6x |
| 0.65 | 0.30 | 0.0150 | 0.0055 | 2.7x |
Encoder diagnostics (ridge heads, closed-form):
| head | R^2 |
|---|---|
| comp (r_target) | 0.822 |
| struct (log1p CSI_struct) | 0.054 |
Intended use
- Reference implementation of the two-factor architecture without any framework dependency.
- Loss-informativeness demonstration.
- Educational.
Limitations
- Not competitive with any modern LM.
- Struct head fails. R^2 = 0.054.
- Synthetic data only.
||w_alpha||is a misleading metric.- Adam is hand-coded.
Comparison to the PyTorch companion
| property | hv-two-factor-lm | two-factor-lm (PyTorch) |
|---|---|---|
| framework | NumPy | PyTorch |
| encoder | hypervector context | causal transformer |
| LM head | hand-coded gradients | autograd |
| diagnostic heads | ridge | learned heads |
| model size | 44.8 KB | ~2 MB |
| training time | ~15 s CPU | ~4 min CPU |
| comp R^2 | 0.82 | 0.95 |
| struct R^2 | 0.05 | 0.65 |
| marginal gain | 2-4x | ~50x |
How to use
import numpy as np
from hv_two_factor_lm import HvTwoFactorLM
data = np.load('two_factor_lm.npz')
model = HvTwoFactorLM(D=int(data['D']), K=int(data['K']),
decay=float(data['decay']))
model.token_hv = data['token_hv']
model.W_lm = data['W_lm']
model.b_lm = data['b_lm']
model.w_alpha = data['w_alpha']
model.w_inject = data['w_inject']
model.W_comp = data['W_comp']
model.W_struct = data['W_struct']
- Downloads last month
- -