tern_tc

verify

A 9M-parameter ternary code model published with the C engine, the verification harness and bit-exact cross-engine receipts.

This README is generated from model_card.t27 -- the spec owns every fact stated here. The same facts feed the Queen WARS arena; the arena's own source of truth is the public wars.t27.

Quickstart

git clone https://huggingface.co/playra/tern-tc-9m && cd tern-tc-9m

sh scripts/verify.sh
# last line: ALL CHECKS DONE

python3 scripts/decode_ids.py 1 3 204 276 405 659 85 1516
# prints the board text from Receipts below with the shipped tokenizer, no dependencies

One command checks the whole package: every file hash, the C engine build, and the greedy ids against the FPGA board receipt. With python3 and torch present it also runs the PyTorch parity harness; without torch those parts skip and the C checks still stand alone.

What this is

This is a pilot, not a coding assistant. At 9M parameters and 2B training tokens, greedy continuations write plausible code-shaped text and turn repetitive within tens of tokens. What this artifact is for: ternary code-model inference that anyone can check end to end, from the container bytes to the FPGA fabric.

No quality comparison with any other model is claimed here. The Queen WARS arena scores configurations only on runs it witnesses, and this checkpoint has not run an arena issue.

Architecture

Decoder-only transformer: RMSNorm, RoPE, grouped-query attention (5 query heads to 1 KV head), SwiGLU, tied embeddings (t27 #1034 P1.1).

layers 6
d_model 320
attention 5 query heads / 1 KV head, head_dim 64 (GQA)
d_ff (SwiGLU) 864
vocab 8192
context 2048
RoPE theta 10000.0

Parameters

9,076,800 total = 6,451,200 ternary block weights + 2,625,600 fp32 (embeddings 2,621,440 tied output head + 4,160 RMSNorm weights).

The 6,451,200 block weights are ternary {-1, 0, +1}: one byte each in the container, one signed product each in the engine. Every weight matrix carries one fp32 scale (7 matrices per layer). Embeddings (tied output head) and all RMSNorm weights stay fp32.

Container (TC02)

3102abdf35057924e077a86db4fdac726e4fe1f6574df47f39a638ba99a3ce9c -- 16,953,800 bytes. little-endian: 'TC02' magic + 7 i32 (n_layer, n_head, n_kv_head, d_model, d_ff, head_dim, vocab), then fp32 wte, then per layer n1 fp32, q/k/v/o ternary + fp32 scale each, n2 fp32, gate/up/down ternary + fp32 scale each, then the final fp32 norm.

Every byte is accounted for by the spec:

header 32 + wte 10,485,760 + 6 x block 1,077,788 + final norm 1,280 = 16,953,800

Tokenizer

GPT-2 byte-level BPE, vocab 8192, fill-in-the-middle special ids 0-5.

81355717cd63d2d92f2d1553eee8e072c05cf49b41990d03b20e7898df0ed9d8

Training

Trained from scratch (seed 0) by the IGLA coder pilot; training data details and per-language counts are in results/data8k/manifest.json.

  • data: codeparrot/github-code-clean, shards 0-96, files up to 100,000 bytes, 15 languages (Python 25.2%, C++ 11.2%, Java 11.0%, C 10.1%, JS 10.0%, Go 5.9%, TS 5.9%, PHP 4.4%, C# 4.3%, Markdown 3.2%, Ruby 2.2%, Rust 2.1%, Shell 2.1%, Scala 1.2%, Lua 0.8%, Julia 0.4%)
  • licenses: permissive only: apache-2.0, bsd-2-clause, bsd-3-clause, cc0-1.0, isc, mit, unlicense
  • tokens seen: 1,999,896,576 (7629 steps x 262,144); corpus 2,414,120,338 tokens
  • peak LR 0.002688478461886226, warmup 152 steps, decay from step 6103, seed 0
  • hardware: NVIDIA RTX PRO 4500 Blackwell

Receipts

Greedy 8 tokens from BOS: 1 3 204 276 405 659 85 1516 -- gen_tokens_8_reopen (trinity-fpga b2bd5a30) on a Xilinx XC7A200T: the ternary weights ran as ternary matvec in FPGA fabric and produced these ids from BOS.

Decoded with the published tokenizer the ids read '', a newline, eight spaces, then 'self._parameters'.

The same TC02 bytes produce the greedy ids above in the C engine, in PyTorch on CPU and in PyTorch on the MPS GPU. Teacher-forced C vs PyTorch CPU: argmax agrees at all 8 positions, max |dNLL| 1.9e-2; an exact f32 port of the C engine reproduces C to 1.5e-6, so the 1.9e-2 is the ternary matvec summation order (sequential adds in C against a BLAS dot in PyTorch) moving 8-bit activation re-quantisation rounding decisions, not a weight mismatch.

Measured speed

Single-stream greedy decode with KV cache, best of 3-5 runs, one Apple M1 Pro (8 CPU cores, 14 GPU cores): C engine compiled with cc -O2 on one thread 36 tokens/s, PyTorch CPU 287, PyTorch MPS 164. The GPU passes the CPU only from batch 32 on decode (1.29x) and from about 4K prefill tokens (up to 1.41x): a 9M fp32 model does not saturate it.

engine mode tokens/s
C (cc -O2, 1 thread) greedy decode 36
PyTorch CPU greedy decode 287
PyTorch MPS (GPU) greedy decode 164

Benchmarks

HumanEval-family pass rate

CPU decode: parameters vs tokens/s

Every number on the charts is either OBSERVED by this repository, with the measurement protocol recorded per row, or SOURCE-CLAIM: the number as printed in the cited source, with its read date. Suites differ row by row, nothing cross-suite is compared, and no claim that tern-tc-9m beats any model is made. The machine-readable projection is charts/benchmarks.json; the source of truth for every chart number is charts/benchmarks.t27, the same spec the charts are rendered from.

Verify it yourself

The C engine needs only a C99 compiler and libm. The Python harness needs torch; the MPS check runs when a Metal GPU is present and is skipped otherwise. scripts/bench_infer.py reproduces the speed table.

Checked continuously, not just at publish time: a public GitHub runner downloads this repository from the hub and runs the same scripts/verify.sh on the downloaded bytes -- on every change to the verify repo and weekly. The badge above and the run history are public.

make -C c_infer && ./c_infer/tc_infer c_infer/model.bin --greedy --max 8
# ids printed: 1 3 204 276 405 659 85 1516

python3 scripts/infer_check.py
# last line:  VERDICT: PASS

Files

file what it is
c_infer/model.bin the checkpoint, TC02 container
c_infer/tc_infer.c, c_infer/Makefile the C99 reference engine
model.py the PyTorch definition (loads the TC02 bytes back into tensors)
scripts/infer_check.py the verification harness (C vs CPU vs GPU parity)
scripts/bench_infer.py the speed benchmark
scripts/export_tc_bin.py the TC02 writer (documents the container)
scripts/verify.sh the end-to-end check the quickstart runs
scripts/decode_ids.py decode token ids with the shipped tokenizer
results/data8k/tokenizer.json the 8192-token tokenizer
charts/bench_pass.svg, charts/bench_speed.svg the benchmark charts
charts/benchmarks.json the machine-readable benchmark projection
charts/benchmarks.t27 the spec every chart number is rendered from
results/data8k/manifest.json training-data provenance
model_card.t27 the source of truth this README is generated from
sha256sums.txt every file above, hashed

License

APACHE-2.0 (see LICENSE).

Spec tests: 7 blocks, 37 asserts, all hold at render time.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support