Instructions to use q-project/Q-164M with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use q-project/Q-164M with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="q-project/Q-164M", trust_remote_code=True)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("q-project/Q-164M", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use q-project/Q-164M with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "q-project/Q-164M" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "q-project/Q-164M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/q-project/Q-164M
- SGLang
How to use q-project/Q-164M with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "q-project/Q-164M" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "q-project/Q-164M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "q-project/Q-164M" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "q-project/Q-164M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use q-project/Q-164M with Docker Model Runner:
docker model run hf.co/q-project/Q-164M
This is a research release, not a product. It is small and often factually wrong outside its trained tool-calling domain. See Results below and the blog post before using it for anything beyond experimentation.
Model summary
Q-164M is a 164,648,648-parameter decoder-only model. Its weights are ternary ({-1, 0, +1},
quantization-aware from step one of training), and instead of a trained embedding table it uses a frozen
512-bit fingerprint per token plus a small learned correction, covering the whole 32,832-token
vocabulary. A fixed circuit intercepts <CALC>expr<EQ> spans during generation and force-decodes the
exact result from a deterministic evaluator — the model's job is only ever to decide when and which
tool to call (arithmetic, dates, unit conversion, sort, compare, gcd/lcm, primality, and 12 other kinds),
never to compute the answer itself. Pretrained from scratch on a single 16GB V100, no cluster: 200,000
steps, ~92.4 hours, 13.1B tokens, zero NaNs over the whole run.
An instruction-tuned version, Q-U-164M, is available as its own repo — chat-SFT then RL, fine-tuned specifically for dialogue. Code for both is at github.com/Q-Project-LM/Q-164M.
Results (final checkpoint, step 200,000)
| capability | result | verdict |
|---|---|---|
| validation perplexity (held-out mix) | 16.98 | steadily improving to the end |
| FineWeb holdout perplexity (plain text, no tool rows) | 25.61 | generalizes past its own tool-call rows |
| fixed 8-task tool-calling probe | 7/8 | — |
— the one failure: sort: |
never converged | the model calls a wrong tool for this one task family |
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("q-project/Q-164M")
model = AutoModelForCausalLM.from_pretrained("q-project/Q-164M", trust_remote_code=True)
ids = tok("Once upon a time,", return_tensors="pt").input_ids
print(tok.decode(model.generate(ids, max_new_tokens=60, do_sample=True, top_k=40,
suppress_tokens=list(range(32768, model.config.vocab_size)))[0], skip_special_tokens=True))
For exact tool calls, build [<bos>, <USER>, *question_ids, <EOT>, <MODEL>] and generate with
circuits.CircuitLogitsProcessor(tokenizer):
from circuits import USER, MODEL, EOT, CircuitLogitsProcessor
ids = [1, USER] + tok.encode("What is 91 divided by 7?").ids + [EOT, MODEL]
out = model.generate(torch.tensor([ids]), max_new_tokens=60,
logits_processor=[CircuitLogitsProcessor(tok)])
Architecture
- 10 layers, hidden 1280, 20 query / 4 KV heads (head dim 64), plain 2-matrix SiLU MLP 1280→3520→1280, per-head QK-norm, per-head sigmoid output gate on attention, RoPE (θ=1e5), context 4096 (trained at 2048).
- Ternary weights from step 1: every attention/MLP projection is re-quantized to
{-1,0,+1}·α(per-row absmean α) at each forward pass with a straight-through estimator. - Frozen fingerprint vocabulary: each token is a fixed ±1 code of 512 bits instead of a trained embedding. Input =
C[id]·W_in; readout =(h·W_out)·Cᵀ/√512 + frozen unigram log-prior. Every id carries its own untieddelta_in/delta_outcorrection on top of the frozen code. - Exact circuits (
circuits.py): the model writes<CALC>expr<EQ>;CircuitLogitsProcessorforces the exact result and<ECALC>— arithmetic, dates, weekday, duration, unit conversion, comparison, sorting, filtering, counting, mean/max/min/sum, gcd/lcm, primality, string reversal, factorial, square root, percent change, and word counts.
Token layout
| ids | content |
|---|---|
| 0–32767 | text BPE (32k); 0 <pad>, 1 <bos>, 2 <eos> |
| 32768–32831 | control: <USER> <MODEL> <EOT> <THINK> <ETHINK> <CALC> <EQ> <ECALC> <NEED> <QUOTE> and reserved ids |
How it compares to other small models
Measured with QBench — see the leaderboard for the live table and methodology. See the Q-U-164M model card for the instruct variant's numbers, which is what QBench actually targets (this base model is a plain-text continuation model, not instruction-tuned for chat).
Limitations
- Facts are frequently wrong; thin general knowledge; repetitive under greedy decoding; weak creative writing. No safety tuning.
- Sorting lists (
sort:) never converged during pretraining — the model reliably calls the wrong tool for this one task family. - Ternary weights at this size are a research setting; do not expect competitive benchmark scores from a single-V100, 164M-parameter model.
License and data terms
Code and weights: Apache-2.0 (see LICENSE). Training data keeps its own terms — check before commercial use: FineWeb / FineWeb-Edu and the SmolLM corpus (ODC-By 1.0 / Apache-2.0).
- Downloads last month
- 67