TrynMini v2 — a tiny static sentence embedder (7.7M params · 7.8 MB · no GPU)
Semantic similarity, search, clustering and RAG retrieval at ~8,000 sentences/sec on a single CPU core — from a 7.8 MB file, with only numpy + tokenizers. No PyTorch. No GPU. No transformer at inference.
TrynMini v2 is a static embedding model: instead of running a transformer at inference, each token is a row in a learned table. Encoding a sentence is just:
tokenize → look up rows → SIF/Zipf-weighted mean pool → small residual matmul → L2-normalize
That makes it ~10× smaller and dramatically faster than encoder models like all-MiniLM-L6-v2, while keeping most of their quality. It's trained — not hashed — by distilling a strong teacher (BAAI/bge-small-en-v1.5) and fine-tuning with a ranking objective, so the vectors are genuinely useful, not just fast.
Why use this?
| TrynMini v2 (this model) | all-MiniLM-L6-v2 | potion-base-8M | |
|---|---|---|---|
| Size on disk | 7.8 MB (int8) | ~90 MB | ~30 MB |
| Runtime deps | numpy, tokenizers |
PyTorch / ONNX | numpy |
| GPU needed | No (0 VRAM) | Optional | No |
| Speed (1 CPU core) | ~8,300 sent/sec | ~hundreds/sec | fast |
| STSB-dev Spearman | 0.715 | ~0.82 | ~0.75 |
Use it when you want good-enough semantic vectors that are trivial to ship — serverless functions, edge devices, browsers (via Pyodide), CI, or anywhere a 90 MB PyTorch dependency is too heavy. You get ~87% of all-MiniLM's STS quality at under 1/10th the size and no GPU.
Reach for a full transformer encoder instead when you need the last few points of accuracy and can afford the size/latency.
Performance (measured)
Single CPU core, numpy only, batched encode of short sentences:
| Metric | Value |
|---|---|
| Throughput (batched) | ~8,300 sentences/sec |
| Latency (one at a time) | 0.23 ms / sentence |
| Model load time | ~0.4 s |
| RAM resident | ~75 MB |
| VRAM | 0 |
Benchmarks — STSBenchmark dev (Spearman)
Matryoshka training means you can truncate the vector to trade a little accuracy for a lot of memory/speed:
| Dim | Spearman | Notes |
|---|---|---|
| 64 | 0.700 | smallest / fastest; already beats v1 at 256 |
| 128 | 0.710 | great default |
| 256 | 0.715 | full quality |
Reference points (reported by their authors, same STSB task): all-MiniLM-L6-v2 ≈ 0.82, Model2Vec potion-base-8M ≈ 0.75.
Install
pip install -U huggingface_hub numpy tokenizers safetensors
Quickstart (5 lines)
import sys; from huggingface_hub import snapshot_download
d = snapshot_download("LNTTushar/trynmini-v2-static-7m-v2"); sys.path.insert(0, d)
from modeling_trynmini import TrynMiniV2
m = TrynMiniV2.from_pretrained(d)
emb = m.encode(["a man plays guitar", "someone plays a guitar"], dim=256) # [2, 256], L2-normalized
print("cosine similarity:", float(emb[0] @ emb[1]))
m.encode(texts, dim=64|128|256) returns L2-normalized vectors, so cosine similarity is just a dot product. The repo also ships example.py (a runnable demo + a small built-in STS sanity check).
How it was trained
- Distillation — encode a large sentence corpus with the teacher (
BAAI/bge-small-en-v1.5, 384-d), reduce to 256-d, and fit a static token table weighted by Zipf/SIF frequency. - Fine-tuning — a post-pool residual is trained with Multiple-Negatives Ranking Loss (MNRL) + Matryoshka on ~156k NLI / STS / paraphrase pairs (SNLI/MNLI, Quora, STS-B).
- Quantization — the table is stored as per-row int8 (≈4× smaller, lossless on STSB here).
Specs
- Params: ~7.7M (30,000 × 256 table + 256² residual)
- Dims: 256, Matryoshka-truncatable to 128 / 64
- Tokenizer: 30k WordPiece (
tokenizer.json) - Files:
model.safetensors(int8 table),residual.npy,sif.npy,vocab.json,tokenizer.json,modeling_trynmini.py
Limitations
A static model has no attention at inference, so it can't model word order or long-range context the way a transformer can — it trades those last accuracy points for size and speed. It's English, trained for sentence-level similarity (not a drop-in for long-document or multilingual retrieval). Further gains are possible with a stronger teacher (bge-base / e5) and more training.
License
Apache-2.0.
- Downloads last month
- 21
Evaluation results
- Spearman cosine (dim 256) on STSBenchmark (dev)self-reported0.715
- Spearman cosine (dim 128) on STSBenchmark (dev)self-reported0.710
- Spearman cosine (dim 64) on STSBenchmark (dev)self-reported0.700