Tern-0.1
A language model whose weights are born ternary. Every weight matrix in Tern, including the tied embedding and output table, is a code in {−1, 0, +1} from the first step of training to the last. There is no full-precision copy of the weights, no post-training quantization, and no GPU: Tern-0.1 was trained from scratch on a CPU by OkeyMeta Ltd.
The whole model is 959 KB.
Tern-0.1 is the first public checkpoint of the Tern line, released as a proof of the training method, not as a useful assistant. Read the evaluation before you use it.
At a glance
| Parameters | 3.68M (3.67M ternary, 10K full-precision scales, norms and biases) |
| Weight format | Ternary {−1, 0, +1}, packed 2 bits per weight, one scale per matrix |
| Activations | 8-bit per token |
| Architecture | Looped transformer: 2 unique blocks applied 2 times (depth 4), d_model 256, 4 heads |
| Positional encoding | Compass (see below) |
| Tokenizer | Oji, byte-level BPE, 8,192 tokens + end-of-text |
| Context | 128 tokens |
| Training data | ~12M tokens of FineWeb-Edu |
| Training hardware | 2 CPU threads, no GPU |
| Optimizer | OkeyMeta in-house ternary optimizer (no Adam, no FP latent weights) |
| File size | 959 KB |
What is different
Native ternary training. Most 1.58-bit models, BitNet b1.58 included, keep a full-precision "latent" copy of every weight during training and round it to ternary in the forward pass. Tern never stores one. The ternary code itself is the weight, and the optimizer edits codes directly. Training memory for the weights is therefore close to their deployed size.
Looped depth. Two transformer blocks are reused twice, with a learned loop-index vector telling the shared blocks which pass they are on. Depth costs compute, not parameters.
Compass positional encoding. Rotary phases come from three axes instead of one: token order, position inside the current word, and whether the token sits inside a quotation. The model knows where a word starts and where quoted speech begins without having to infer it.
Oji tokenizer. A byte-level BPE trained with a language-parity objective, so text in under-served languages is not split into far more tokens than English.
Evaluation
All models were scored on the same held-out FineWeb-Edu text (92 documents, 302,981 bytes, never seen by Tern) in bits per byte, which is fair across different tokenizers. Lower is better.
| Model | Params | Weight bits | bits/byte ↓ | LAMBADA (500) ↑ |
|---|---|---|---|---|
| Pythia-160M | 162.3M | 16 | 0.978 | 38.8% |
| GPT-2 124M | 124.4M | 16/32 | 0.991 | 33.2% |
| Pythia-70M | 70.4M | 16 | 1.107 | 19.0% |
| Tern-0.1 | 3.68M | 1.58 | 2.257 | 0.0% |
Plainly: Tern-0.1 is far behind these baselines. It is 19 to 44 times smaller, stores weights in about a tenth of the bits, and saw roughly 12M training tokens against their hundreds of billions. Its bits-per-byte sits just under a bigram model's (about 2.32 bpb), so it has learned word and local-phrase statistics, not meaning. Its validation loss was still falling when training stopped.
Sample (prompt in bold, temperature 0.8, top-k 40):
The history of science to write they is a find in advived. They also the world say think of members in honths draw things is a gittle slaully in the time the test and the same
English-shaped, not sensible. That is the honest state of a 3.7M-parameter, 12M-token model.
Why release it
Tern-0.1 shows the full pipeline works end to end on commodity hardware: a model trained from random ternary codes to a stable loss curve with no full-precision weights, no GPU and no Adam. The research question for the Tern line is how far that method goes as data and compute grow while weights stay ternary. Tern-0.2, trained on 1.35B tokens, is in progress and will be published against the same benchmark.
Usage
pip install torch safetensors regex huggingface_hub
from huggingface_hub import snapshot_download
import sys
path = snapshot_download("OkeyMetaLtd/Tern-0.1")
sys.path.insert(0, path)
from oji_tokenizer import Oji
from modeling_tern import TernLM
tok = Oji(f"{path}/tokenizer.json")
model = TernLM.from_pretrained(path, tok)
print(model.generate(tok, "Water is made of", max_new_tokens=60, seed=0))
Or from the downloaded folder: python generate.py "Water is made of".
This repository holds inference code only. The forward pass runs ternary weights as float matrices for portability; dedicated ternary kernels would make it much faster.
Limitations
- Not an assistant. No instruction tuning, no safety tuning, no factual reliability.
- 128-token context.
- Trained on English educational web text only, despite a multilingual tokenizer.
- Will produce ungrammatical and meaningless text.
License
Weights and code are released under CC BY-NC 4.0 for research and non-commercial use. For commercial licensing, contact OkeyMeta Ltd.
Citation
@misc{okeymeta2026tern01,
title = {Tern-0.1: A Natively Ternary Language Model Trained on CPU},
author = {Nwaozor, Okechukwu and {OkeyMeta Ltd}},
year = {2026},
url = {https://huggingface.co/OkeyMetaLtd/Tern-0.1}
}
- Downloads last month
- -