Instructions to use Ternarycore/ternarycore-sst2-student with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Ternarycore/ternarycore-sst2-student with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Ternarycore/ternarycore-sst2-student", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("Ternarycore/ternarycore-sst2-student", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Ternarycore/ternarycore-sst2-student with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Ternarycore/ternarycore-sst2-student" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Ternarycore/ternarycore-sst2-student", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Ternarycore/ternarycore-sst2-student
- SGLang
How to use Ternarycore/ternarycore-sst2-student with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Ternarycore/ternarycore-sst2-student" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Ternarycore/ternarycore-sst2-student", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Ternarycore/ternarycore-sst2-student" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Ternarycore/ternarycore-sst2-student", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Ternarycore/ternarycore-sst2-student with Docker Model Runner:
docker model run hf.co/Ternarycore/ternarycore-sst2-student
TernaryCore SST-2 Student (W1.58A8)
A ternary-weight transformer distilled from a full-precision teacher for binary
sentiment classification on GLUE SST-2. Every linear-projection weight is constrained
to balanced ternary {-1, 0, +1} (~1.58 bits/weight, absmean-scaled) and paired with
8-bit activations (W1.58A8), so each weight-activation product is +x, -x, or 0
and the matrix multiplies reduce to conditional accumulation - no hardware multiplier
required.
It is trained with Microsoft’s BitNet Distillation recipe
(arXiv:2510.13998): replace the linear layers with
BitLinear, insert SubLN normalization, warm up with quantization-aware training, then
distill from the teacher with logit (and attention-relation) losses. The task is framed
generatively ("Review: ...\nSentiment:" -> " positive" / " negative") so the
deployment path is the same decoder-LM datapath the hardware runs.
This is the reproducible demonstrator workload for TernaryCore, an open-source multiplier-free ternary MAC/dot/GEMM accelerator that runs on a low-cost Artix-7 FPGA, and it is the exact checkpoint behind the accompanying manuscript’s on-hardware energy measurements.
Version note. This card documents a frozen release. Cite this specific revision / DOI (not
main) - the accompanying measured numbers are pinned to this snapshot.
Usage
The model adds SubLN layers that stock Qwen3ForCausalLM does not have, so it ships with
its own modeling code and needs trust_remote_code=True:
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
mid = "Ternarycore/ternarycore-sst2-student"
tok = AutoTokenizer.from_pretrained(mid, trust_remote_code=True)
m = AutoModelForCausalLM.from_pretrained(
mid, trust_remote_code=True, dtype=torch.bfloat16).eval()
ids = tok("Review: a gorgeous, witty, seductive movie.\nSentiment:",
return_tensors="pt")
with torch.no_grad():
logits = m(**ids).logits[0, -1]
print(tok.decode(logits.argmax())) # " positive"
The safetensors hold the ternary values already multiplied by their per-layer absmean
scale, stored as bf16. Inference is therefore ordinary bf16 matmul, numerically
identical to the ternary datapath, and every projection tensor contains exactly three
distinct values:
import numpy as np
w = m.model.layers[0].self_attn.q_proj.weight.detach().float().numpy()
np.unique(w) # [-0.02478, 0.0, 0.02478]
(w == 0).mean() # 0.3151
The genuinely 1.58-bit artifact - 2 bits per weight as the accelerator consumes it - is
under hardware/ (see below).
Results
| Task | Teacher | Ternary student | % of teacher |
|---|---|---|---|
| SST-2 (sentiment) | 94.4% | 91.4% | 96.8% |
(A sibling MNLI student reached 82.8% vs an 88.2% teacher. This checkpoint is the SST-2 student.)
Model details
- Teacher / skeleton:
Qwen/Qwen3-0.6B(Apache-2.0, 596M params). The student inherits the teacher’s dimensions, so it fits the accelerator by construction. It is not architecturally identical: BitNet-Distillation inserts a SubLN (RMSNorm) beforeo_projanddown_projin every block - 56 extra tensors - which is why the custom modeling code is required. - Architecture: decoder transformer, grouped-query attention (
q_proj2048x1024,k_proj/v_proj1024x1024,o_proj1024x2048) with a SwiGLU MLP (gate_proj/up_proj3072x1024,down_proj1024x3072). - Depth: 28 blocks -> 196 weight-layer GEMMs.
- Hidden size: 1024 - MLP intermediate: 3072 - head dim: 128 - 16 attention heads, 8 KV heads - vocab: 151,936 - tied embeddings.
- Weight precision: balanced ternary {-1, 0, +1} via BitLinear (absmean scale applied once after accumulation); activations INT8 -> W1.58A8.
- Total parameters: 596.2M. The ternary GEMMs hold 440,545,280 weights and pack to
110.1 MB in the accelerator’s 2-bit format; 31.4% are exact zeros (each a
skipon the MAC array - the model learned roughly the sparsity the BitNet papers report). Embeddings, the tied LM head, SubLN and the normalization layers are not ternary and run in the activation domain. - Task: SST-2 binary sentiment (positive / negative), framed generatively.
hardware/ - the packed format the FPGA consumes
layers/- 196 per-projection blobs,addr = k * GROUPS + g, 4 codes per byte, LSB-first, code map00 -> 0,01 -> +1,10 -> -1. One 1024x1024 layer packs to exactly 262,144 bytes.manifest.json- per-layer shape, absmean scale, zero fraction and file path.reference.npz- the same 196 projections as int8 ternary matrices (-1/0/+1), for checking a packer or an FPGA readback element by element.weights.bin+pages.json- the whole model pre-sliced into 420 pages of 256 KB in the order the block executor walks them, so the board can DMA by page index.
The bf16 safetensors were reconstructed from layers/ and verified against
reference.npz: 196/196 projections match exactly, and every zero fraction matches
manifest.json.
Intended use
- Reproducing the TernaryCore silicon and energy results.
- A compact, honestly-sized ternary checkpoint for research on multiplier-free / sub-watt LLM inference on reconfigurable hardware.
- An educational reference for BitNet b1.58 W1.58A8 arithmetic and the BitNet-distillation workflow at fine-tuning scale.
Out of scope: a task-specific SST-2 classifier, not a general chat/instruction model. Do not use for open-ended generation or as a safety-sensitive classifier without independent validation.
How the arithmetic maps to hardware
| Ternary property | Hardware image |
|---|---|
| weight = 0 | skip (no accumulation) |
| weight = +1 / -1 | add / subtract |
| no sign bit | no multiplier |
On the TernaryCore array a projection runs at one element/clock (64 MACs/cycle, 12.7 GOPS at 100 MHz) using zero DSP slices. On the deployed Artix-7 SoC (DDR3 + Ethernet active) one forward pass of this model’s 196 weight-layer GEMMs was measured at ~310 mJ / ~88 ms by inline current sensing - the weight-layer portion only (attention scores, softmax, normalization, embeddings, and the head are activation-domain and excluded).
Provenance, license and attribution
- Model weights: released under Apache-2.0, following the teacher.
- Teacher:
Qwen/Qwen3-0.6B(Apache-2.0) - attribution retained per that license. The student was distilled to ternary from this teacher; it is not a Qwen release and carries no Qwen endorsement. - Method: Microsoft BitNet Distillation, arXiv:2510.13998.
- Data: GLUE SST-2 (Socher et al., 2013), research-permissive terms.
- Accelerator RTL/tooling (separate repo) is licensed CERN-OHL-S v2; this license note covers the model weights, which are Apache-2.0.
Citation
The accelerator this model was built for:
@software{ternarycore,
title = {TernaryCore: Multiplier-Free Balanced Ternary GEMM Accelerator for BitNet
b1.58 Inference},
author = {Oladapo, Ifedayo and Vasilev, Dmitrii and Mohammed, Abubakar},
year = {2026},
doi = {10.5281/zenodo.22837568},
url = {https://github.com/Ternarycore/ternarycore},
license = {CERN-OHL-S-2.0}
}
A manuscript describing the design, silicon validation and same-fabric comparison is in preparation; this card will be updated with its citation once it is available.
Limitations
Ternary quantization and distillation shift decision boundaries versus the full-precision teacher (SST-2 lands at 96.8% of teacher accuracy; the sibling MNLI student missed its bar by about 1 point). Validate on your own split before relying on it. The model inherits biases of its teacher and of the SST-2 training data, and is English- and sentiment-only.
- Downloads last month
- 145
Model tree for Ternarycore/ternarycore-sst2-student
Dataset used to train Ternarycore/ternarycore-sst2-student
Paper for Ternarycore/ternarycore-sst2-student
Evaluation results
- Validation accuracy on GLUE SST-2self-reported0.914