Instructions to use dmis-lab/UltraBERT-380M with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use dmis-lab/UltraBERT-380M with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("fill-mask", model="dmis-lab/UltraBERT-380M", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForMaskedLM model = AutoModelForMaskedLM.from_pretrained("dmis-lab/UltraBERT-380M", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
UltraBERT-380M
UltraBERT is a family of English bidirectional encoders with a native 32K context that reach state of the art on both classification and retrieval, two task families on which previous encoders had to choose a side.
| Model | Parameters | Hidden | Context | Description |
|---|---|---|---|---|
| UltraBERT-150M | 150M | 768 | 32K | Pre-trained encoder |
| UltraBERT-380M (this model) | 380M | 1,024 | 32K | Pre-trained encoder |
| UltraBERT-Embed-150M | 150M | 768 | 32K | Text embedding model built on UltraBERT-150M |
For more information about UltraBERT, please check out our paper.
✨ Highlights
- Classification and retrieval at once. Earlier encoders each excel in one area: Ettin on GLUE, RoBERTa on NER, LFM2.5-Encoder on BEIR. UltraBERT is the first encoder to take the top score on both classification and retrieval at the same size.
- Native 32K context. The model is trained at 2K tokens and extended to 32K context length. On the LongEmbed passkey test it keeps 99% accuracy at every length up to 32K, where 8K-context encoders drop sharply.
- Decoupled LM heads. Each pre-training batch mixes sequences masked at 5% and at 50%, and each masking ratio has its own LM head. Low ratios favor classification and high ratios favor retrieval; separate heads let one encoder learn both without the interference a shared head causes.
- Pruning and distillation from a 1.3B teacher. The 380M model is width-pruned from a 1.3B encoder and the 150M model from the 380M model. Both are distilled from the 1.3B teacher with a KL loss, applied only to the 5% masked sequences, while the 50% sequences use the standard MLM loss.
🚀 Usage
The model needs transformers >= 5 (tested with 5.16.1) and trust_remote_code=True. FlashAttention 2 is recommended for long inputs; SDPA and eager attention are also supported.
import torch
from transformers import AutoModelForMaskedLM, AutoTokenizer
model_id = "dmis-lab/UltraBERT-380M"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForMaskedLM.from_pretrained(
model_id,
trust_remote_code=True,
dtype=torch.bfloat16,
attn_implementation="flash_attention_2", # or "sdpa"
device_map="cuda",
)
text = f"The capital of France is {tokenizer.mask_token}."
inputs = tokenizer(text, return_tensors="pt").to(model.device)
with torch.no_grad():
logits = model(**inputs).logits
mask_index = (inputs.input_ids == tokenizer.mask_token_id).nonzero()[0, 1]
print(tokenizer.decode(logits[0, mask_index].argmax()))
Fine-Tuning
The checkpoint loads into the usual task classes:
from transformers import AutoModelForSequenceClassification
model = AutoModelForSequenceClassification.from_pretrained(
"dmis-lab/UltraBERT-380M", num_labels=2, trust_remote_code=True
)
We recommend starting from the settings used in the paper. For classification: AdamW (β1 0.9, β2 0.98, ε 1e-6), weight decay 0.1, a linear schedule with 5% warm-up, and a learning rate swept over {5e-6,8e-6,1e-5,3e-5,5e-5,8e-5}. For dense retrieval: mean pooling with a contrastive loss and a wider learning-rate sweep of {1e-5,3e-5,5e-5,8e-5,1e-4,2e-4,3e-4,5e-4,8e-4,1e-3}.
💡 Note: UltraBERT consistently prefers lower learning rates than other encoders, by up to 10x. For example, the best learning rates for UltraBERT-150M and UltraBERT-380M are 5e-5 and 3e-5, respectively, whereas other encoders typically prefer 3e-4 to 5e-4.
📊 Evaluation
Every model is fine-tuned per task with the same grid search. For the retrieval tasks, each encoder is fine-tuned on a subset of MS MARCO or MLDR under the same sweep, so these scores reflect the potential information retrieval capability of each encoder. Dedicated embedding models are usually trained further on hundreds of millions of pairs with much longer training schedules.
Base Size (at most 200M parameters)
| Model | Params | GLUE | SQuAD 2.0 | MultiCoNER v2 | Few-NERD | Reward-Bench | BEIR | RARb Code | RARb Math | MLDR | LongEmbed |
|---|---|---|---|---|---|---|---|---|---|---|---|
| BERT | 110M | 84.7 | 76.3 | 53.8 | 66.6 | 59.4 | 36.7 | 6.6 | 27.5 | 29.0 | 29.3 |
| RoBERTa | 125M | 86.4 | 83.7 | 56.5 | 67.1 | 63.2 | 39.6 | 11.7 | 33.4 | 28.7 | 33.8 |
| MosaicBERT | 137M | 85.4 | 79.4 | 57.8 | 65.7 | 62.5 | 41.4 | 12.7 | 30.9 | 29.6 | 40.5 |
| NomicBERT | 137M | 84.0 | 82.7 | 60.1 | 66.9 | 64.1 | 42.0 | 13.8 | 30.2 | 29.9 | 45.0 |
| GTE-en-MLM | 137M | 85.6 | 82.7 | 59.2 | 66.4 | 63.4 | 42.4 | 12.0 | 33.6 | 37.1 | 58.5 |
| ModernBERT | 149M | 88.4 | 86.7 | 57.4 | 66.7 | 68.0 | 42.6 | 29.6 | 46.1 | 37.7 | 63.2 |
| Ettin Encoder | 149M | 88.9 | 87.4 | 57.4 | 66.5 | 70.3 | 42.4 | 31.8 | 47.1 | 35.9 | 60.1 |
| UltraBERT | 152M | 89.5 | 89.2 | 57.9 | 67.4 | 73.0 | 45.2 | 39.1 | 55.4 | 43.5 | 72.0 |
Medium Size (at most 500M parameters)
| Model | Params | GLUE | SQuAD 2.0 | MultiCoNER v2 | Few-NERD | Reward-Bench | BEIR | RARb Code | RARb Math | MLDR | LongEmbed |
|---|---|---|---|---|---|---|---|---|---|---|---|
| BERT | 334M | 85.2 | 81.8 | 57.7 | 67.8 | 62.6 | 38.6 | 12.2 | 33.0 | 29.8 | 30.0 |
| RoBERTa | 355M | 88.9 | 89.4 | 62.3 | 68.1 | 68.5 | 42.8 | 20.3 | 37.8 | 29.5 | 32.2 |
| GTE-en-MLM | 434M | 87.6 | 87.2 | 63.0 | 67.1 | 62.2 | 39.4 | 13.5 | 32.0 | 40.8 | 57.0 |
| NeoBERT | 222M | 89.0 | 86.6 | 57.4 | 66.4 | 57.6 | 42.9 | 17.3 | 33.4 | 21.7 | 32.4 |
| LFM2.5-Encoder | 230M | 88.2 | 86.1 | 59.6 | 67.6 | 67.5 | 46.7 | 30.9 | 50.2 | 43.7 | 70.0 |
| LFM2.5-Encoder | 355M | 89.2 | 87.2 | 62.7 | 67.9 | 68.5 | 47.3 | 38.6 | 50.2 | 44.5 | 71.7 |
| ModernBERT | 395M | 90.4 | 89.8 | 60.9 | 67.6 | 72.8 | 45.2 | 34.0 | 46.1 | 39.9 | 63.5 |
| Ettin Encoder | 395M | 90.8 | 90.4 | 62.3 | 67.8 | 76.7 | 45.9 | 38.5 | 54.0 | 40.9 | 64.3 |
| UltraBERT | 380M | 91.0 | 90.9 | 63.5 | 68.3 | 78.8 | 47.7 | 47.0 | 61.7 | 44.1 | 74.9 |
Long-Context Passkey Retrieval
Passkey retrieval accuracy on LongEmbed as a function of document length, for (a) base and (b) medium encoders. Every baseline degrades once the document exceeds its context length: the 512-context models fail beyond 1K tokens and the 8K-context models drop sharply at 16K and 32K. UltraBERT is the only model that retains 99% accuracy at every length up to 32K, at both sizes.
🏛️ Model Description
| UltraBERT-150M | UltraBERT-380M | |
|---|---|---|
| Layers | 24 | 24 |
| Hidden size | 768 | 1,024 |
| SwiGLU size | 1,024 | 3,072 |
| Attention heads | 6 | 8 |
| Head dimension | 128 | 128 |
| Attention pattern | local x2, global | local x2, global |
| Local attention window | 128 tokens each side | 128 tokens each side |
| RoPE base (local / global) | 10K / 2.56M | 10K / 2.56M |
| Context length | 32,768 | 32,768 |
| Vocabulary | 49,152 | 49,152 |
| Normalization | RMSNorm (eps 1e-6) | RMSNorm (eps 1e-6) |
| Parameters | 152M | 380M |
UltraBERT is a pre-norm Transformer encoder with RoPE, SwiGLU feed-forward layers, RMSNorm and no bias terms. Two of every three layers use local attention restricted to 128 tokens on each side, and every third layer attends over the whole sequence. The depth is fixed at 24 layers for every size and only the width is scaled. The LM head is untied from the input embeddings.
What Is in This Checkpoint
model.safetensorsin float32 (the pre-training master weights), withconfig.dtype: float32.- The encoder and the MLM head.
- The MLM head is the one trained on the 5% masked sequences (the distilled head); the 50% head is not included.
- Remote code for
transformers5.x withUltrabertModel,UltrabertForMaskedLM,UltrabertForSequenceClassification,UltrabertForQuestionAnsweringandUltrabertForTokenClassification.
⚙️ Pre-Training
Objective. Masked language modeling with two masking ratios in every batch, 5% and 50% at a 4:1 sequence ratio, each with its own LM head. The 5% head is trained with a KL divergence loss against the logits of a 1.3B teacher; the 50% head with the standard MLM loss.
Pruning and distillation. The 1.3B teacher is width-pruned to 380M and the distilled 380M model is pruned again to 150M. Importance of hidden channels, attention heads and FFN neurons is estimated from activations on 2,000 calibration documents, and the top-scoring units are kept. Both sizes are distilled from the 1.3B teacher directly.
Optimization. AdamW (β1 0.9, β2 0.95, ε 1e-8), weight decay 0.1, gradient clipping 1.0, peak learning rate 1e-4 with a warmup-stable-decay schedule (10B warm-up tokens, linear decay over the last 10%). Batch size 4M tokens, and 2T tokens in total; the context is extended to 32K, raising the global-attention RoPE base from 640K to 2.56M. Trained on a single TPU v4-512, v5e-256 or v6e-256 slice with data parallelism and FSDP.
Data. One fixed mixture for the whole run (Table 4 of the paper).
| Category | Dataset | Unique tokens (B) | Epochs | Trained tokens (B) | Share |
|---|---|---|---|---|---|
| Web | DCLM | 756.4 | 1.17 | 881.7 | 44.1% |
| Web | FineWeb-Edu | 195.4 | 2.33 | 455.5 | 22.8% |
| Code | Stack-Edu | 120.3 | 2.33 | 280.5 | 14.0% |
| Academic | peS2o | 60.7 | 2.33 | 141.6 | 7.1% |
| Math | MegaMath | 20.9 | 2.80 | 58.4 | 2.9% |
| Math | FineMath | 10.9 | 2.80 | 30.5 | 1.5% |
| Math | InfiWebMath | 9.6 | 2.80 | 26.9 | 1.3% |
| Reference | Cosmopedia v2 | 28.1 | 2.33 | 65.6 | 3.3% |
| Reference | Wikipedia (Dolma) | 3.9 | 2.33 | 9.2 | 0.5% |
| Instruction | FLAN | 19.0 | 2.33 | 44.4 | 2.2% |
| Instruction | StackExchange (RedPajama) | 1.5 | 2.33 | 3.4 | 0.2% |
| Instruction | Natural Reasoning | 1.0 | 2.33 | 2.4 | 0.1% |
| Total | 1,227.8 | 1.63 | 2,000.0 | 100% |
Limitations
- English only. The pre-training mixture is English web, code, academic, math and instruction text.
- This is a pre-trained encoder, not a ready-to-use classifier or embedding model: it needs fine-tuning for a task. For a ready-to-use text embedding model built on it, see UltraBERT-Embed-150M.
- The model can reproduce biases of its web-scale training data; evaluate it for your use case.
Citation
@article{park2026ultrabert,
title = {UltraBERT: Beyond the Limit of BERT Pre-Training},
author = {Park, Jungwoo and Lee, Taewhoo and Kang, Jaewoo},
journal = {arXiv preprint arXiv:{ARXIV_ID}},
year = {2026}
}
Acknowledgments
The pre-training was primarily supported by Cloud TPUs from Google's TPU Research Cloud (TRC) program.
- Downloads last month
- 78

