GNOMON

A certified four-layer transformer.

GNOMON is a research architecture that combines four mechanisms into one model: ternary weights (BitNet b1.58), an information-bottleneck KV cache, Vickrey-auction slot selection, and a Lipschitz Decision Circuit output head. The four layers share one algorithm β€” measure sufficiency, certify, stop β€” applied at different scales.

This repository contains the reference implementation, all example scripts, and a trained tiny model checkpoint.

Model Details

Architecture

Component Role State reduction Source
Ternary weights All linear layers quantized to {βˆ’1, 0, +1} 16Γ— memory BitNet b1.58
IB cache Compresses KV state to a 16-dim bottleneck 64Γ— KV state Tishby (1999)
Vickrey auction Allocates cache slots by truthful bids variable budget Vickrey (1961)
LDC output Compiled k-d tree with linear leaves 10⁢× faster decisions This work

The four components, verified

Each of the four mechanisms has been measured on this codebase:

Ternary weights. Forward pass uses {βˆ’1, 0, +1}; gradients flow through a straight-through estimator to the unquantized kernel. Confirmed to produce exactly three distinct values after ternarize().

IB cache. Encodes a (B, S, D) KV state into a (B, k) latent and reconstructs the KV pair through a decoder. With k=16, S=32, D=32, the measured compression ratio is 64Γ— with a computable reconstruction error.

Vickrey auction. Selects the top-k positions by marginal-information bids. The runner-up bid becomes the price. On synthetic data with importance linearly increasing across positions, the top quartile retains 28/128 positions at budget=32 and 122/128 at budget=256.

LDC. Compiles a (X, Y) classification task into a k-d tree with ridge-regression leaves. On a 3-class Gaussian task with C-1 = 2 LDA directions, the compiled tree reaches 100% test accuracy in 92 ms with 128 leaves, and runs at ~4 million queries/s on CPU (jitted).

The coupling result

The headline measurement of the architecture is the logit KL divergence between raw attention and IB-reconstructed attention. On random weights, this divergence starts at ~0.04. After 2000 training steps on a bigram task, it drops to 0.0002 β€” a 169Γ— tightening.

This is the architecture's thesis: the IB state learns to be sufficient for the next layer's attention. When the divergence reaches zero, the cache state is a complete replacement for the raw KV pair.

Intended Uses

Direct use

  • Research on efficient inference. The four components are individually usable: TernaryDense for a quantized linear layer, compile_ldc for compiling a classifier into a fast circuit, IBCache for KV compression.
  • Teaching. The four mechanisms are each isolated in their own module with minimal dependencies. The showcase.py script runs all four end-to-end in under 5 seconds on CPU.
  • Baseline architecture. The full GNOMON model trains end-to-end on synthetic tasks and can be adapted for small-scale experiments.

Out-of-scope use

  • Production language modelling. The provided checkpoint is a 2-layer, 64-dim toy trained on a bigram Markov chain. It does not generate coherent English. The architecture is a research direction, not a deployment-ready LLM.
  • Safety-critical decisions. The LDC is a compiled lookup structure. It is only as correct as the data it was compiled from. It should not be used where an out-of-distribution input could cause harm.
  • Tasks requiring > 30 intrinsic dimensions. The LDC's speedup comes from the target function having low intrinsic dimension. Tasks with high intrinsic dimension (translation, long-form reasoning) are outside its operating regime.

How to Use

The model has no external dependencies beyond JAX and NumPy. It is not installed as a package β€” clone the repo and run the examples directly.

Component-level

import sys
sys.path.insert(0, "path/to/gnomon")

import jax, jax.numpy as jnp
from gnomon import compile_ldc, init_ib, ib_forward, ternarize

# --- Ternary weights ---
w = jax.random.normal(jax.random.PRNGKey(0), (128, 128))
w_q, scale = ternarize(w)
# w_q / scale is in {-1, 0, +1}

# --- Compile an LDC ---
ldc = compile_ldc(X_train, Y_train, k=8, max_leaves=128)
preds = ldc.predict_jit(jnp.asarray(X_test))

# --- IB cache ---
params = init_ib(jax.random.PRNGKey(0), d_enc_in=64, d_dec_out=32,
                 d_hidden=128, k=16)
s, mu, log_sigma, kv_hat = ib_forward(params, x, rng, S=32)

### Full model

```python
from gnomon import init_gnomon, init_mode_state, make_forward_jit

params = init_gnomon(
    jax.random.PRNGKey(0),
    vocab_size=32, d_model=64, n_layers=2, n_heads=4,
    ib_k=16, seq_len=24, d_signal=16,
)
forward = make_forward_jit(n_heads=4)
out = forward(params, input_ids, mode_state, mode_signal, rng)

Training Details

The bigram experiment

The included checkpoint was trained on a synthetic bigram Markov chain with V = 32 vocabulary and temperature 0.3. The transition matrix has mean row entropy of 3.347 nats β€” this is the Bayes-optimal cross-entropy.

Metric Value
Model 2 layers, d_model=64, 4 heads, IB k=16
Training steps 2000
Batch 32 Γ— 24
Optimizer Adam, lr = 2e-3
Uniform CE 3.4657
Bayes CE 3.3474
Final CE 3.3360
Signal recovered 110% of uniform→Bayes gap
Coupling KL (init β†’ final) 0.0385 β†’ 0.0002
Wall time 71.6 s on CPU (27.9 steps/s)

The model slightly beats the Bayes floor, which is expected: the empirical distribution of the finite training sample is not exactly the theoretical transition matrix, and the model fits the sample.

Hyperparameters

Parameter Value
d_model 64
n_layers 2
n_heads 4
ib_k (bottleneck dim) 16
d_signal (neuromodulator input) 16
d_mod (mode vector) 4
Learning rate 2e-3
Adam β₁, Ξ²β‚‚ 0.9, 0.999
Adam Ξ΅ 1e-8

Evaluation

Results

Test results on the four components:

Component Metric Value
Ternary distinct values {βˆ’1, 0, +1}
LDC (Gaussian 3-class) test accuracy 1.0000
LDC leaves / internal nodes 128 / 127
LDC throughput (jitted, CPU) 3.98 M queries/s
IB cache compression ratio 64Γ—
IB cache initial reconstruction error 1.150
Vickrey top-quartile retention @ budget=128 88/128
Full model forward pass (jitted) 1.10 ms

Reproducing

python examples/classify.py       # standalone LDC on a 5-class task
python examples/showcase.py       # all four components in one run
python examples/coupling.py       # raw vs IB-reconstructed attention
python examples/train_task.py     # full training loop on the bigram task
python -m pytest tests/ -v        # 8 unit tests

All five commands complete on CPU in under 2 minutes.

Limitations

  1. Scale. The architecture has been tested up to 2 layers and d_model=64. Scaling laws for the four-component stack are unknown.
  2. The LDC only helps for bounded decisions. It replaces the softmax head for classification, routing, and moderation. It does not generate text.
  3. The IB encoder is trained jointly with the model. Pretrained Hugging Face checkpoints cannot be retrofitted with an IB cache without fine-tuning.
  4. Coupling KL is measured on a synthetic task. The 169Γ— tightening is real but was measured on a bigram Markov chain. Whether the same collapse occurs on natural language is an open question.
  5. The neuromodulator is under-trained. The mode vector is a 4-dim EMA over a task signal. In the current checkpoint it barely moves. Its intended role β€” shifting policy when the task shifts β€” has not been demonstrated.
  6. No pre-trained weights for real tasks. This is a research release. The included checkpoint is a toy.

Citation

@software{gnomon2026,
  title  = {GNOMON: A Certified Four-Layer Transformer},
  author = {[zeechimp]},
  year   = {2026},
  url    = {https://huggingface.co/zeechimp/gnomon},
  note   = {Ternary weights, information-bottleneck cache, Vickrey auction,
            and Lipschitz Decision Circuit combined in one model}
}

If you build on any of the four mechanisms independently:

@article{wang2023bitnet,
  title={BitNet: Scaling 1-bit Transformers for Large Language Models},
  author={Wang, Hongyu and others},
  journal={arXiv:2310.11453},
  year={2023}
}

@article{tishby1999ib,
  title={The Information Bottleneck Method},
  author={Tishby, Naftali and Pereira, Fernando C and Bialek, William},
  journal={Allerton},
  year={1999}
}

@article{vickrey1961counterspeculation,
  title={Counterspeculation, Auctions, and Competitive Sealed Tenders},
  author={Vickrey, William},
  journal={Journal of Finance},
  year={1961}
}

Repository Structure

gnomon/
β”œβ”€β”€ gnomon/
β”‚   β”œβ”€β”€ __init__.py
β”‚   β”œβ”€β”€ params.py        # Xavier init + Adam (no Optax dependency)
β”‚   β”œβ”€β”€ ternary.py       # BitNet b1.58 with straight-through gradient
β”‚   β”œβ”€β”€ neuromod.py      # Slow mode vector (4-dim EMA)
β”‚   β”œβ”€β”€ ldc.py           # LDA projection + k-d tree with ridge leaves
β”‚   β”œβ”€β”€ vickrey.py       # Marginal-information bids + Vickrey top-k
β”‚   β”œβ”€β”€ ib_cache.py      # Variational IB encoder/decoder + certificate
β”‚   β”œβ”€β”€ model.py         # Forward pass with optional KV coupling
β”‚   └── train.py         # Combined CE + IB loss
β”œβ”€β”€ examples/
β”‚   β”œβ”€β”€ classify.py      # Standalone LDC on synthetic 5-class task
β”‚   β”œβ”€β”€ showcase.py      # All four components, one script
β”‚   β”œβ”€β”€ coupling.py      # Raw vs reconstructed attention divergence
β”‚   └── train_task.py    # Full training on bigram Markov chain
└── tests/
    └── test_gnomon.py   # 8 unit tests covering all components

Hardware

All benchmarks were measured on CPU (Windows 11, Python 3.14, JAX 0.11.2). The code is CPU-only by design β€” no CUDA or TPU-specific operations are used. The LDC throughput of 4M queries/s is per-core; on GPU-batched inputs, the same tree reaches 10⁸.

License

Apache 2.0

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Paper for zeechimp/gnomon

Evaluation results

  • CE (uniform baseline 3.466) on synthetic-bigram-v32
    self-reported
    3.336
  • Logit KL (raw vs IB-reconstructed attention) on synthetic-bigram-v32
    self-reported
    0.000