GNOMON
A certified four-layer transformer.
GNOMON is a research architecture that combines four mechanisms into one model: ternary weights (BitNet b1.58), an information-bottleneck KV cache, Vickrey-auction slot selection, and a Lipschitz Decision Circuit output head. The four layers share one algorithm β measure sufficiency, certify, stop β applied at different scales.
This repository contains the reference implementation, all example scripts, and a trained tiny model checkpoint.
Model Details
Architecture
| Component | Role | State reduction | Source |
|---|---|---|---|
| Ternary weights | All linear layers quantized to {β1, 0, +1} |
16Γ memory | BitNet b1.58 |
| IB cache | Compresses KV state to a 16-dim bottleneck | 64Γ KV state | Tishby (1999) |
| Vickrey auction | Allocates cache slots by truthful bids | variable budget | Vickrey (1961) |
| LDC output | Compiled k-d tree with linear leaves | 10βΆΓ faster decisions | This work |
The four components, verified
Each of the four mechanisms has been measured on this codebase:
Ternary weights. Forward pass uses {β1, 0, +1}; gradients flow through
a straight-through estimator to the unquantized kernel. Confirmed to produce
exactly three distinct values after ternarize().
IB cache. Encodes a (B, S, D) KV state into a (B, k) latent and
reconstructs the KV pair through a decoder. With k=16, S=32, D=32, the
measured compression ratio is 64Γ with a computable reconstruction error.
Vickrey auction. Selects the top-k positions by marginal-information
bids. The runner-up bid becomes the price. On synthetic data with importance
linearly increasing across positions, the top quartile retains 28/128 positions
at budget=32 and 122/128 at budget=256.
LDC. Compiles a (X, Y) classification task into a k-d tree with
ridge-regression leaves. On a 3-class Gaussian task with C-1 = 2 LDA
directions, the compiled tree reaches 100% test accuracy in 92 ms with
128 leaves, and runs at ~4 million queries/s on CPU (jitted).
The coupling result
The headline measurement of the architecture is the logit KL divergence between raw attention and IB-reconstructed attention. On random weights, this divergence starts at ~0.04. After 2000 training steps on a bigram task, it drops to 0.0002 β a 169Γ tightening.
This is the architecture's thesis: the IB state learns to be sufficient for the next layer's attention. When the divergence reaches zero, the cache state is a complete replacement for the raw KV pair.
Intended Uses
Direct use
- Research on efficient inference. The four components are individually
usable:
TernaryDensefor a quantized linear layer,compile_ldcfor compiling a classifier into a fast circuit,IBCachefor KV compression. - Teaching. The four mechanisms are each isolated in their own module with
minimal dependencies. The
showcase.pyscript runs all four end-to-end in under 5 seconds on CPU. - Baseline architecture. The full
GNOMONmodel trains end-to-end on synthetic tasks and can be adapted for small-scale experiments.
Out-of-scope use
- Production language modelling. The provided checkpoint is a 2-layer, 64-dim toy trained on a bigram Markov chain. It does not generate coherent English. The architecture is a research direction, not a deployment-ready LLM.
- Safety-critical decisions. The LDC is a compiled lookup structure. It is only as correct as the data it was compiled from. It should not be used where an out-of-distribution input could cause harm.
- Tasks requiring > 30 intrinsic dimensions. The LDC's speedup comes from the target function having low intrinsic dimension. Tasks with high intrinsic dimension (translation, long-form reasoning) are outside its operating regime.
How to Use
The model has no external dependencies beyond JAX and NumPy. It is not installed as a package β clone the repo and run the examples directly.
Component-level
import sys
sys.path.insert(0, "path/to/gnomon")
import jax, jax.numpy as jnp
from gnomon import compile_ldc, init_ib, ib_forward, ternarize
# --- Ternary weights ---
w = jax.random.normal(jax.random.PRNGKey(0), (128, 128))
w_q, scale = ternarize(w)
# w_q / scale is in {-1, 0, +1}
# --- Compile an LDC ---
ldc = compile_ldc(X_train, Y_train, k=8, max_leaves=128)
preds = ldc.predict_jit(jnp.asarray(X_test))
# --- IB cache ---
params = init_ib(jax.random.PRNGKey(0), d_enc_in=64, d_dec_out=32,
d_hidden=128, k=16)
s, mu, log_sigma, kv_hat = ib_forward(params, x, rng, S=32)
### Full model
```python
from gnomon import init_gnomon, init_mode_state, make_forward_jit
params = init_gnomon(
jax.random.PRNGKey(0),
vocab_size=32, d_model=64, n_layers=2, n_heads=4,
ib_k=16, seq_len=24, d_signal=16,
)
forward = make_forward_jit(n_heads=4)
out = forward(params, input_ids, mode_state, mode_signal, rng)
Training Details
The bigram experiment
The included checkpoint was trained on a synthetic bigram Markov chain with
V = 32 vocabulary and temperature 0.3. The transition matrix has mean row
entropy of 3.347 nats β this is the Bayes-optimal cross-entropy.
| Metric | Value |
|---|---|
| Model | 2 layers, d_model=64, 4 heads, IB k=16 |
| Training steps | 2000 |
| Batch | 32 Γ 24 |
| Optimizer | Adam, lr = 2e-3 |
| Uniform CE | 3.4657 |
| Bayes CE | 3.3474 |
| Final CE | 3.3360 |
| Signal recovered | 110% of uniformβBayes gap |
| Coupling KL (init β final) | 0.0385 β 0.0002 |
| Wall time | 71.6 s on CPU (27.9 steps/s) |
The model slightly beats the Bayes floor, which is expected: the empirical distribution of the finite training sample is not exactly the theoretical transition matrix, and the model fits the sample.
Hyperparameters
| Parameter | Value |
|---|---|
d_model |
64 |
n_layers |
2 |
n_heads |
4 |
ib_k (bottleneck dim) |
16 |
d_signal (neuromodulator input) |
16 |
d_mod (mode vector) |
4 |
| Learning rate | 2e-3 |
| Adam Ξ²β, Ξ²β | 0.9, 0.999 |
| Adam Ξ΅ | 1e-8 |
Evaluation
Results
Test results on the four components:
| Component | Metric | Value |
|---|---|---|
| Ternary | distinct values | {β1, 0, +1} |
| LDC (Gaussian 3-class) | test accuracy | 1.0000 |
| LDC | leaves / internal nodes | 128 / 127 |
| LDC | throughput (jitted, CPU) | 3.98 M queries/s |
| IB cache | compression ratio | 64Γ |
| IB cache | initial reconstruction error | 1.150 |
| Vickrey | top-quartile retention @ budget=128 | 88/128 |
| Full model | forward pass (jitted) | 1.10 ms |
Reproducing
python examples/classify.py # standalone LDC on a 5-class task
python examples/showcase.py # all four components in one run
python examples/coupling.py # raw vs IB-reconstructed attention
python examples/train_task.py # full training loop on the bigram task
python -m pytest tests/ -v # 8 unit tests
All five commands complete on CPU in under 2 minutes.
Limitations
- Scale. The architecture has been tested up to 2 layers and d_model=64. Scaling laws for the four-component stack are unknown.
- The LDC only helps for bounded decisions. It replaces the softmax head for classification, routing, and moderation. It does not generate text.
- The IB encoder is trained jointly with the model. Pretrained Hugging Face checkpoints cannot be retrofitted with an IB cache without fine-tuning.
- Coupling KL is measured on a synthetic task. The 169Γ tightening is real but was measured on a bigram Markov chain. Whether the same collapse occurs on natural language is an open question.
- The neuromodulator is under-trained. The mode vector is a 4-dim EMA over a task signal. In the current checkpoint it barely moves. Its intended role β shifting policy when the task shifts β has not been demonstrated.
- No pre-trained weights for real tasks. This is a research release. The included checkpoint is a toy.
Citation
@software{gnomon2026,
title = {GNOMON: A Certified Four-Layer Transformer},
author = {[zeechimp]},
year = {2026},
url = {https://huggingface.co/zeechimp/gnomon},
note = {Ternary weights, information-bottleneck cache, Vickrey auction,
and Lipschitz Decision Circuit combined in one model}
}
If you build on any of the four mechanisms independently:
@article{wang2023bitnet,
title={BitNet: Scaling 1-bit Transformers for Large Language Models},
author={Wang, Hongyu and others},
journal={arXiv:2310.11453},
year={2023}
}
@article{tishby1999ib,
title={The Information Bottleneck Method},
author={Tishby, Naftali and Pereira, Fernando C and Bialek, William},
journal={Allerton},
year={1999}
}
@article{vickrey1961counterspeculation,
title={Counterspeculation, Auctions, and Competitive Sealed Tenders},
author={Vickrey, William},
journal={Journal of Finance},
year={1961}
}
Repository Structure
gnomon/
βββ gnomon/
β βββ __init__.py
β βββ params.py # Xavier init + Adam (no Optax dependency)
β βββ ternary.py # BitNet b1.58 with straight-through gradient
β βββ neuromod.py # Slow mode vector (4-dim EMA)
β βββ ldc.py # LDA projection + k-d tree with ridge leaves
β βββ vickrey.py # Marginal-information bids + Vickrey top-k
β βββ ib_cache.py # Variational IB encoder/decoder + certificate
β βββ model.py # Forward pass with optional KV coupling
β βββ train.py # Combined CE + IB loss
βββ examples/
β βββ classify.py # Standalone LDC on synthetic 5-class task
β βββ showcase.py # All four components, one script
β βββ coupling.py # Raw vs reconstructed attention divergence
β βββ train_task.py # Full training on bigram Markov chain
βββ tests/
βββ test_gnomon.py # 8 unit tests covering all components
Hardware
All benchmarks were measured on CPU (Windows 11, Python 3.14, JAX 0.11.2). The code is CPU-only by design β no CUDA or TPU-specific operations are used. The LDC throughput of 4M queries/s is per-core; on GPU-batched inputs, the same tree reaches 10βΈ.
License
Apache 2.0
Paper for zeechimp/gnomon
Evaluation results
- CE (uniform baseline 3.466) on synthetic-bigram-v32self-reported3.336
- Logit KL (raw vs IB-reconstructed attention) on synthetic-bigram-v32self-reported0.000