Instructions to use vvvvtrt/RUNE-4M with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use vvvvtrt/RUNE-4M with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("fill-mask", model="vvvvtrt/RUNE-4M")# Load model directly from transformers import AutoModelForMaskedLM model = AutoModelForMaskedLM.from_pretrained("vvvvtrt/RUNE-4M", device_map="auto") - Notebooks
- Google Colab
- Kaggle
RUNE-4M
Russian Ultra-efficient Encoder · 4.25M parameters
Компактный русский bidirectional encoder с local attention (окно 64) и контекстом до 8192 токенов.
Дистиллирован из deepvk/RuModernBERT-small (~34.6M).
Ultra-light · fast inference · teacher-compatible tokenizer (vocab 50368)
Why RUNE-4M
| Size | 4.25M — ~7× smaller than rubert-tiny2, ~8× smaller than RuModernBERT-small |
| Speed | On Apple MPS: 2.1 ms @ seq 128, 62 ms @ seq 2048 (batch=1) |
| Quality (wiki MLM) | Top-1 48% — ahead of rubert-tiny2 (44%), behind the teacher (67%) |
| Context | RoPE, up to 8192 positions |
| Attention | Layer 1: local window 64 · Layer 2: global |
Сильная сторона — скорость и размер. По MLM на wiki holdout модель обходит tiny2, но не догоняет учителя. На OOD-доменах (новости, отзывы) tiny2 пока сильнее — это честный статус.
Wiki MLM quality
Holdout: 128 Russian Wikipedia texts, seq 256, BERT 15% mask, seed 42.
Each model uses its own tokenizer (same raw documents).
| Model | Params | CE ↓ | Top-1 ↑ | Top-5 ↑ |
|---|---|---|---|---|
| RuModernBERT-small | 34.6M | 1.71 | 67% | 81% |
| RUNE-4M | 4.25M | 3.17 | 48% | 62% |
| rubert-tiny2 | 29.3M | 3.57 | 44% | 59% |
Latency
Apple MPS · batch=1 · warmup=8 · reps=30 · encoder + MLM head.
| Model | Params | seq 128 | seq 256 | seq 512 | seq 1024 | seq 2048 |
|---|---|---|---|---|---|---|
| RUNE-4M | 4.25M | 2.1 ms | 8.2 ms | 13 ms | 33 ms | 62 ms |
| rubert-tiny2 | 29.3M | 19 ms | 25 ms | 49 ms | 111 ms | 277 ms |
| RuModernBERT-small | 34.6M | 20 ms | 28 ms | 60 ms | 127 ms | 331 ms |
Short (128): RUNE ≈ ×9 faster than tiny2 / ModernBERT.
Long (2048): RUNE ≈ ×4.5 vs tiny2, ≈ ×5.3 vs ModernBERT.
Architecture
BERT-style encoder (Masked LM, bidirectional).
tokens (L ≤ 8192)
│
▼
Embedding + LayerNorm [H=80, vocab=50368]
│
▼
Local attention · window 64 [O(L · 64)]
+ GLU MLP
│
▼
Global softmax attention [O(L²)]
+ GLU MLP
│
▼
Final LN → tied MLM head
| value | |
|---|---|
| Parameters | 4.25M |
| Hidden size | 80 |
| Heads | 4 |
| Layers | 2 (local/64 → global) |
| MLP | GLU, intermediate 240 |
| Position encoding | RoPE |
| Max positions | 8192 |
| Tokenizer | compatible with RuModernBERT-small |
Intended use
- On-device / CPU / edge Russian NLP
- Fast reranking, lightweight classifiers, embeddings
- Distillation / student baselines where latency and checkpoint size matter
Not a drop-in replacement for full-size RuModernBERT on heavy NLU suites.
Limitations
- MLM quality trails the teacher; on non-wiki domains may trail rubert-tiny2.
- Global layer is still (O(L^2)) — long contexts are dominated by it.
- Custom architecture: needs registered model class / wrapper (not plain
BertModelweights).
Training (short)
- Teacher:
deepvk/RuModernBERT-small - Objective: MLM CE + KL + hidden MSE (distillation)
- Data: Russian Wikipedia + news (multi-shard)
- Best checkpoint: train-val MLM CE ≈ 2.40 @ step 193500
Usage
from transformers import AutoTokenizer
repo = "vvvvtrt/RUNE-4M"
tok = AutoTokenizer.from_pretrained(repo)
# Weights use HybridMultiscaleForMaskedLM.
# Load via the project class or export ONNX for production.
Special tokens (RuModernBERT-compatible): BOS/CLS 50281, EOS/SEP 50282, pad 50283.
Citation
@misc{rune4m2026,
title = {RUNE-4M: Russian Ultra-efficient Encoder (4.25M)},
author = {Mikhail Chashin},
year = {2026},
howpublished = {\url{https://huggingface.co/vvvvtrt/RUNE-4M}}
}
Also cite the teacher:
@misc{deepvk2025rumodernbert,
title = {RuModernBERT: Modernized BERT for Russian},
author = {Spirin, Egor and Malashenko, Boris and Sokolov, Andrey},
year = {2025},
url = {https://huggingface.co/deepvk/RuModernBERT-small}
}
License
Apache License 2.0
Aligned with the teacher model deepvk/RuModernBERT-small (apache-2.0).
You redistribute weights of a distilled student, not the pretraining corpora. If you redistribute training data separately, follow each dataset’s license (Wikipedia CC BY-SA, etc.).
- Downloads last month
- 26
Model tree for vvvvtrt/RUNE-4M
Base model
deepvk/RuModernBERT-small