TrynMini v2 — a tiny static sentence embedder (7.7M params · 7.8 MB · no GPU)

Semantic similarity, search, clustering and RAG retrieval at ~8,000 sentences/sec on a single CPU core — from a 7.8 MB file, with only numpy + tokenizers. No PyTorch. No GPU. No transformer at inference.

TrynMini v2 is a static embedding model: instead of running a transformer at inference, each token is a row in a learned table. Encoding a sentence is just:

tokenize → look up rows → SIF/Zipf-weighted mean pool → small residual matmul → L2-normalize

That makes it ~10× smaller and dramatically faster than encoder models like all-MiniLM-L6-v2, while keeping most of their quality. It's trained — not hashed — by distilling a strong teacher (BAAI/bge-small-en-v1.5) and fine-tuning with a ranking objective, so the vectors are genuinely useful, not just fast.


Why use this?

TrynMini v2 (this model) all-MiniLM-L6-v2 potion-base-8M
Size on disk 7.8 MB (int8) ~90 MB ~30 MB
Runtime deps numpy, tokenizers PyTorch / ONNX numpy
GPU needed No (0 VRAM) Optional No
Speed (1 CPU core) ~8,300 sent/sec ~hundreds/sec fast
STSB-dev Spearman 0.715 ~0.82 ~0.75

Use it when you want good-enough semantic vectors that are trivial to ship — serverless functions, edge devices, browsers (via Pyodide), CI, or anywhere a 90 MB PyTorch dependency is too heavy. You get ~87% of all-MiniLM's STS quality at under 1/10th the size and no GPU.

Reach for a full transformer encoder instead when you need the last few points of accuracy and can afford the size/latency.


Performance (measured)

Single CPU core, numpy only, batched encode of short sentences:

Metric Value
Throughput (batched) ~8,300 sentences/sec
Latency (one at a time) 0.23 ms / sentence
Model load time ~0.4 s
RAM resident ~75 MB
VRAM 0

Benchmarks — STSBenchmark dev (Spearman)

Matryoshka training means you can truncate the vector to trade a little accuracy for a lot of memory/speed:

Dim Spearman Notes
64 0.700 smallest / fastest; already beats v1 at 256
128 0.710 great default
256 0.715 full quality

Reference points (reported by their authors, same STSB task): all-MiniLM-L6-v2 ≈ 0.82, Model2Vec potion-base-8M ≈ 0.75.


Install

pip install -U huggingface_hub numpy tokenizers safetensors

Quickstart (5 lines)

import sys; from huggingface_hub import snapshot_download
d = snapshot_download("LNTTushar/trynmini-v2-static-7m-v2"); sys.path.insert(0, d)
from modeling_trynmini import TrynMiniV2
m = TrynMiniV2.from_pretrained(d)
emb = m.encode(["a man plays guitar", "someone plays a guitar"], dim=256)   # [2, 256], L2-normalized
print("cosine similarity:", float(emb[0] @ emb[1]))

m.encode(texts, dim=64|128|256) returns L2-normalized vectors, so cosine similarity is just a dot product. The repo also ships example.py (a runnable demo + a small built-in STS sanity check).


How it was trained

  1. Distillation — encode a large sentence corpus with the teacher (BAAI/bge-small-en-v1.5, 384-d), reduce to 256-d, and fit a static token table weighted by Zipf/SIF frequency.
  2. Fine-tuning — a post-pool residual is trained with Multiple-Negatives Ranking Loss (MNRL) + Matryoshka on ~156k NLI / STS / paraphrase pairs (SNLI/MNLI, Quora, STS-B).
  3. Quantization — the table is stored as per-row int8 (≈4× smaller, lossless on STSB here).

Specs

  • Params: ~7.7M (30,000 × 256 table + 256² residual)
  • Dims: 256, Matryoshka-truncatable to 128 / 64
  • Tokenizer: 30k WordPiece (tokenizer.json)
  • Files: model.safetensors (int8 table), residual.npy, sif.npy, vocab.json, tokenizer.json, modeling_trynmini.py

Limitations

A static model has no attention at inference, so it can't model word order or long-range context the way a transformer can — it trades those last accuracy points for size and speed. It's English, trained for sentence-level similarity (not a drop-in for long-document or multilingual retrieval). Further gains are possible with a stronger teacher (bge-base / e5) and more training.

License

Apache-2.0.

Downloads last month
21
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Evaluation results

  • Spearman cosine (dim 256) on STSBenchmark (dev)
    self-reported
    0.715
  • Spearman cosine (dim 128) on STSBenchmark (dev)
    self-reported
    0.710
  • Spearman cosine (dim 64) on STSBenchmark (dev)
    self-reported
    0.700