Text Generation
Transformers
Safetensors
English
tokle
causal-lm
custom-architecture
custom_code
slm
small-language-model
Instructions to use techdotus/Tokle-3M with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use techdotus/Tokle-3M with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="techdotus/Tokle-3M", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("techdotus/Tokle-3M", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use techdotus/Tokle-3M with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "techdotus/Tokle-3M" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "techdotus/Tokle-3M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/techdotus/Tokle-3M
- SGLang
How to use techdotus/Tokle-3M with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "techdotus/Tokle-3M" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "techdotus/Tokle-3M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "techdotus/Tokle-3M" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "techdotus/Tokle-3M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use techdotus/Tokle-3M with Docker Model Runner:
docker model run hf.co/techdotus/Tokle-3M
File size: 5,919 Bytes
933698f ff59975 713054f ff59975 933698f ff59975 ad27408 ff59975 3b1e47e 912fed0 3b1e47e ff59975 912fed0 73a614c ff59975 713054f ff59975 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 | ---
language:
- en
license: cc-by-nc-sa-4.0
library_name: transformers
pipeline_tag: text-generation
tags:
- text-generation
- causal-lm
- custom-architecture
- custom_code
- slm
- small-language-model
datasets:
- HuggingFaceFW/fineweb-edu
- HuggingFaceTB/cosmopedia
- agentlans/high-quality-english-sentences
- nampdn-ai/tiny-strange-textbooks
- armanc/ScienceQA
- nvidia/OpenMathInstruct-2
- microsoft/orca-math-word-problems-200k
---
# Tokle-3M
## Model Summary
Tokle-3M is a decoder-only language model with 2.91M parameters. It was first trained on 12B tokens with SPAB (Static Pairwise Attention Bias), a frozen table of 8.39M token-pair association scores built from Pointwise Mutual Information (PMI) over the training corpus, giving 11.3M parameters in total during this stage. During training, for every query-key pair, SPAB hashed the two token IDs into the table, retrieved their PMI value, scaled it by a learned per-head factor, and added it to the attention logits before softmax.
After this stage, the SPAB table was removed and the model was trained for an additional 0.5B tokens to distill the knowledge in the SPAB matrix into its own layers. As a result, Tokle-3M runs entirely on its 2.91M parameters at inference, with no SPAB table required.
## Model Architecture
| Parameter | Value |
|---|---|
| Architecture | Decoder-only transformer (RMSNorm, RoPE, GQA, SwiGLU) |
| Layers | 9 |
| Hidden size (d_model) | 144 |
| Attention heads | 3 |
| KV heads (GQA) | 1 (multi-query attention) |
| Head dim | 48 |
| FFN intermediate size | 432 |
| Max sequence length | 512 |
| Tie word embeddings | Yes |
| Precision | FP32 weights |
| Parameters | 2.91M |
## How to use
```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "techdotus/Tokle-3M"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, trust_remote_code=True).eval()
ids = tok("The climate change", return_tensors="pt")
with torch.no_grad():
out = model.generate(**ids, max_new_tokens=32, do_sample=False,
repetition_penalty=1.3) # greedy
print(tok.decode(out[0], skip_special_tokens=True))
```
## Benchmark Results
All scores are 0-shot acc_norm, using the Open SLM Leaderboard methodology.
| HellaSwag | ARC-Easy | ARC-Challenge | PIQA | ArithMark-3 |
|---|---|---|---|---|
| 27.20% | 34.85% | 23.98% | 55.01% | 40.80% |
### Ablation: SPAB vs. Distilled
Stage 1 model (SPAB active) vs. the released Tokle-3M (SPAB removed and distilled).
| Model | Params | Int Index | HellaSwag | ARC-Easy | ARC-Chal | PIQA | ArithMark-3 |
|---|---|---|---|---|---|---|---|
| [Tokle-SPAB-3M](https://huggingface.co/techdotus/Tokle-SPAB-3M) | 11.3M (2.91M trainable + 8.39M frozen) | 9.16 | 27.22% | 34.68% | 24.49% | 54.95% | 41.70% |
| Tokle-3M | 2.91M | 8.91 | 27.20% | 34.85% | 23.98% | 55.01% | 40.80% |
## Comparison Results
All scores are 0-shot acc_norm, using the Open SLM Leaderboard methodology. Scores for the other models are from the Open SLM Leaderboard. Bold marks the best result in each column.
| Model | Params | Int Index | HellaSwag | ARC-Easy | ARC-Chal | PIQA | ArithMark-3 |
|---|---|---|---|---|---|---|---|
| Tokle-3M (Tech.us) | 2.91M | **8.91** | 27.20% | **34.85%** | **23.98%** | 55.01% | **40.80%** |
| pulvis-v2 | 2.96M×3* | 8.62 | 27.93% | 31.86% | 22.78% | **56.86%** | 37.40% |
| Purrence-3M | 2.99M×6* | 8.49 | 27.81% | 34.55% | 23.63% | 56.64% | 34.80% |
| Ember-2 | 2.96M×2* | 7.21 | 27.28% | 33.42% | 22.01% | 55.11% | 35.90% |
## Training Details
Tokle-3M was trained in two stages on the same data mixture.
| Stage | Tokens | SPAB | Parameters |
|---|---|---|---|
| 1. Pretraining | 12B | Active (frozen PMI table) | 11.3M (2.91M trainable + 8.39M frozen) |
| 2. Distillation | 0.5B | Removed | 2.91M |
Stage 2 lets the trained weights absorb the prior the SPAB table had been providing, so the released model is self-contained rather than losing that knowledge when the table is removed.
## Training Data
We trained on a curated mixture with a strict cleaning pipeline that also removed topics not useful for a model of this size.
| Source | Percentage |
|---|---|
| FineWeb-Edu | 43.1% |
| Cosmopedia | 24.3% |
| OpenMathInstruct-2 | 13.5% |
| Tiny Strange Textbooks | 9.0% |
| MegaScience (medicine & biology, custom curated) | 5.0% |
| High-Quality English Sentences | 3.0% |
| ScienceQA | 1.2% |
| Orca-Math Word Problems 200k | 0.9% |
| **Total** | **100%** |
- **Tokenizer:** all data was tokenized with the model's 5,048-token BPE tokenizer, and 1% was held out for validation.
- **Blending:** sources were blended per dataset using the weights above.
## Limitations
- **Tiny model:** with 2.91M parameters and 144-dim hidden states, generations are often repetitive, incoherent or factually wrong. The model is a research artifact for studying small-scale LMs, not an assistant.
- **Short context:** 512 tokens maximum. RoPE tables are not built beyond that length.
- **English only:** trained on English web, educational, synthetic and math text.
- **Not instruction-tuned or safety-aligned:** it may reproduce biases present in web data.
## License
**Code**: MIT. The modeling code, tokenizer, and training scripts are released under the MIT license.
**Weights**: CC BY-NC-SA 4.0. The training data includes MegaScience (CC BY-NC-SA 4.0), so the weights are released under the same terms: attribution required, non-commercial use only, and derivatives (including fine-tunes) must be shared under the same license.
## Citation
```bibtex
@misc{tokle2026,
title = {{Tokle-3M}: Pointwise Mutual Information as a Removable
Inductive Bias for Self-Attention},
author = {{Tech.us Team}},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/techdotus/Tokle-3M}}
}
``` |