RUNE-4M

Russian Ultra-efficient Encoder · 4.25M parameters

Компактный русский bidirectional encoder с local attention (окно 64) и контекстом до 8192 токенов.
Дистиллирован из deepvk/RuModernBERT-small (~34.6M).

Ultra-light · fast inference · teacher-compatible tokenizer (vocab 50368)


Why RUNE-4M

Size 4.25M — ~7× smaller than rubert-tiny2, ~8× smaller than RuModernBERT-small
Speed On Apple MPS: 2.1 ms @ seq 128, 62 ms @ seq 2048 (batch=1)
Quality (wiki MLM) Top-1 48% — ahead of rubert-tiny2 (44%), behind the teacher (67%)
Context RoPE, up to 8192 positions
Attention Layer 1: local window 64 · Layer 2: global

Сильная сторона — скорость и размер. По MLM на wiki holdout модель обходит tiny2, но не догоняет учителя. На OOD-доменах (новости, отзывы) tiny2 пока сильнее — это честный статус.


Wiki MLM quality

Holdout: 128 Russian Wikipedia texts, seq 256, BERT 15% mask, seed 42.
Each model uses its own tokenizer (same raw documents).

Wiki MLM Top-1

Model Params CE ↓ Top-1 ↑ Top-5 ↑
RuModernBERT-small 34.6M 1.71 67% 81%
RUNE-4M 4.25M 3.17 48% 62%
rubert-tiny2 29.3M 3.57 44% 59%

Latency

Apple MPS · batch=1 · warmup=8 · reps=30 · encoder + MLM head.

Latency across sequence lengths

Model Params seq 128 seq 256 seq 512 seq 1024 seq 2048
RUNE-4M 4.25M 2.1 ms 8.2 ms 13 ms 33 ms 62 ms
rubert-tiny2 29.3M 19 ms 25 ms 49 ms 111 ms 277 ms
RuModernBERT-small 34.6M 20 ms 28 ms 60 ms 127 ms 331 ms

Short (128): RUNE ≈ ×9 faster than tiny2 / ModernBERT.
Long (2048): RUNE ≈ ×4.5 vs tiny2, ≈ ×5.3 vs ModernBERT.


Architecture

BERT-style encoder (Masked LM, bidirectional).

tokens (L ≤ 8192)
        │
        ▼
 Embedding + LayerNorm          [H=80, vocab=50368]
        │
        ▼
 Local attention · window 64    [O(L · 64)]
   + GLU MLP
        │
        ▼
 Global softmax attention       [O(L²)]
   + GLU MLP
        │
        ▼
 Final LN → tied MLM head
value
Parameters 4.25M
Hidden size 80
Heads 4
Layers 2 (local/64 → global)
MLP GLU, intermediate 240
Position encoding RoPE
Max positions 8192
Tokenizer compatible with RuModernBERT-small

Intended use

  • On-device / CPU / edge Russian NLP
  • Fast reranking, lightweight classifiers, embeddings
  • Distillation / student baselines where latency and checkpoint size matter

Not a drop-in replacement for full-size RuModernBERT on heavy NLU suites.


Limitations

  • MLM quality trails the teacher; on non-wiki domains may trail rubert-tiny2.
  • Global layer is still (O(L^2)) — long contexts are dominated by it.
  • Custom architecture: needs registered model class / wrapper (not plain BertModel weights).

Training (short)

  • Teacher: deepvk/RuModernBERT-small
  • Objective: MLM CE + KL + hidden MSE (distillation)
  • Data: Russian Wikipedia + news (multi-shard)
  • Best checkpoint: train-val MLM CE ≈ 2.40 @ step 193500

Usage

from transformers import AutoTokenizer

repo = "vvvvtrt/RUNE-4M"
tok = AutoTokenizer.from_pretrained(repo)

# Weights use HybridMultiscaleForMaskedLM.
# Load via the project class or export ONNX for production.

Special tokens (RuModernBERT-compatible): BOS/CLS 50281, EOS/SEP 50282, pad 50283.


Citation

@misc{rune4m2026,
  title        = {RUNE-4M: Russian Ultra-efficient Encoder (4.25M)},
  author       = {Mikhail Chashin},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/vvvvtrt/RUNE-4M}}
}

Also cite the teacher:

@misc{deepvk2025rumodernbert,
  title  = {RuModernBERT: Modernized BERT for Russian},
  author = {Spirin, Egor and Malashenko, Boris and Sokolov, Andrey},
  year   = {2025},
  url    = {https://huggingface.co/deepvk/RuModernBERT-small}
}

License

Apache License 2.0

Aligned with the teacher model deepvk/RuModernBERT-small (apache-2.0).
You redistribute weights of a distilled student, not the pretraining corpora. If you redistribute training data separately, follow each dataset’s license (Wikipedia CC BY-SA, etc.).

Downloads last month
26
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for vvvvtrt/RUNE-4M

Finetuned
(10)
this model