Tokle-3M

Model Summary

Tokle-3M is a decoder-only language model with 2.91M trainable parameters, trained on 12B tokens. Its main architectural addition is SPAB (Static Pairwise Attention Bias), a frozen table of token-pair association scores built from Pointwise Mutual Information (PMI) over the training corpus and added to the attention logits of the first layer.

For every query-key pair, SPAB hashes the two token IDs into the table, pulls out their PMI value, multiplies it by a learned per-head scale, and adds it to the attention logits before softmax. The bias ignores position and depends only on which tokens are involved, so the model starts training already knowing which tokens tend to co-occur. It only has to learn how much to trust that prior.

Model Architecture

Parameter Value
Architecture Custom decoder-only transformer + SPAB (TokleForCausalLM)
Layers 9
Hidden size (d_model) 144
Attention heads 3
KV heads (GQA) 1 (multi-query attention)
Head dim 48
FFN intermediate size 432
Max sequence length 512
Trainable parameters 2,908,947
Frozen SPAB table 8,388,608 (float32 buffer)

How to use

This model uses a custom architecture, so it needs trust_remote_code=True.

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "techdotus/Tokle-3M"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, trust_remote_code=True).eval()

ids = tok("The capital of France is", return_tensors="pt")
with torch.no_grad():
    out = model.generate(**ids, max_new_tokens=32, do_sample=False)  # greedy
print(tok.decode(out[0], skip_special_tokens=True))

Benchmark Results

All scores are 0-shot acc_norm, using the Open SLM Leaderboard methodology:

Hellaswag ARC-Easy ARC-Challenge PIQA Arithmark-3
27.22% 34.68% 24.49% 54.95% 41.70%

Comparison Results

All scores are 0-shot acc_norm, using the Open SLM Leaderboard methodology. Scores for the other models are from the Open SLM Leaderboard. Bold marks the best result in each column.

Model Params Int Index HellaSwag ARC-Easy ARC-Chal PIQA ArithMark-3
Tokle-3M (Tech.us) 2.9M 9.16 27.22% 34.68% 24.49% 54.95% 41.70%
Ember-2 (SurjoLabs) 2.96Mx2 7.21 27.28% 33.42% 22.01% 55.11% 35.90%
BananaMind-2-Micro (BananaMind) 2.9M 6.01 28.27% 33.12% 21.93% 53.21% 34.00%
GPT-S-1.4M (Axiomic Labs) 1.4M 5.40 26.89% 31.57% 21.93% 55.17% 30.20%

Training Data Details

We trained on a curated mixture with a strict cleaning pipeline that also removed topics not useful for a model of this size.

Source Percentage
FineWeb-Edu 43.1%
Cosmopedia 24.3%
OpenMathInstruct-2 13.5%
Tiny Strange Textbooks 9.0%
MegaScience (medicine & biology, custom curated) 5.0%
High-Quality English Sentences 3.0%
ScienceQA 1.2%
Orca-Math Word Problems 200k 0.9%
Total 100%
  • Tokenizer: all data was tokenized with the model's 5,048-token BPE tokenizer, and 1% was held out for validation.
  • Blending: sources were blended per dataset using the weights above.

Limitations

  • Tiny model: with ~2.9M trainable parameters and 144-dim hidden states, generations are often repetitive, incoherent or factually wrong. The model is a research artifact for studying small-scale LMs, not an assistant.
  • Short context: 512 tokens maximum. RoPE tables are not built beyond that length.
  • English only: trained on English web, educational, synthetic and math text.
  • Not instruction-tuned or safety-aligned: it may reproduce biases present in web data.

Licenses

Model weights and code: MIT.

Citation

@misc{tokle2026,
  title        = {{Tokle-3M}: Pointwise Mutual Information as an Inductive Bias for Self-Attention},
  author       = {{Tech.us Team}},
  year         = {2026},
  publisher    = {Hugging Face},
  howpublished = {\url{https://huggingface.co/techdotus/Tokle-3M}}
}
Downloads last month
-
Safetensors
Model size
2.91M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train techdotus/Tokle-3M

Collection including techdotus/Tokle-3M