UltraBERT-380M

UltraBERT

UltraBERT is a family of English bidirectional encoders with a native 32K context that reach state of the art on both classification and retrieval, two task families on which previous encoders had to choose a side.

Model Parameters Hidden Context Description
UltraBERT-150M 150M 768 32K Pre-trained encoder
UltraBERT-380M (this model) 380M 1,024 32K Pre-trained encoder
UltraBERT-Embed-150M 150M 768 32K Text embedding model built on UltraBERT-150M

For more information about UltraBERT, please check out our paper.

✨ Highlights

  • Classification and retrieval at once. Earlier encoders each excel in one area: Ettin on GLUE, RoBERTa on NER, LFM2.5-Encoder on BEIR. UltraBERT is the first encoder to take the top score on both classification and retrieval at the same size.
  • Native 32K context. The model is trained at 2K tokens and extended to 32K context length. On the LongEmbed passkey test it keeps 99% accuracy at every length up to 32K, where 8K-context encoders drop sharply.
  • Decoupled LM heads. Each pre-training batch mixes sequences masked at 5% and at 50%, and each masking ratio has its own LM head. Low ratios favor classification and high ratios favor retrieval; separate heads let one encoder learn both without the interference a shared head causes.
  • Pruning and distillation from a 1.3B teacher. The 380M model is width-pruned from a 1.3B encoder and the 150M model from the 380M model. Both are distilled from the 1.3B teacher with a KL loss, applied only to the 5% masked sequences, while the 50% sequences use the standard MLM loss.

🚀 Usage

The model needs transformers >= 5 (tested with 5.16.1) and trust_remote_code=True. FlashAttention 2 is recommended for long inputs; SDPA and eager attention are also supported.

import torch
from transformers import AutoModelForMaskedLM, AutoTokenizer

model_id = "dmis-lab/UltraBERT-380M"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForMaskedLM.from_pretrained(
    model_id,
    trust_remote_code=True,
    dtype=torch.bfloat16,
    attn_implementation="flash_attention_2",  # or "sdpa"
    device_map="cuda",
)

text = f"The capital of France is {tokenizer.mask_token}."
inputs = tokenizer(text, return_tensors="pt").to(model.device)
with torch.no_grad():
    logits = model(**inputs).logits

mask_index = (inputs.input_ids == tokenizer.mask_token_id).nonzero()[0, 1]
print(tokenizer.decode(logits[0, mask_index].argmax()))

Fine-Tuning

The checkpoint loads into the usual task classes:

from transformers import AutoModelForSequenceClassification

model = AutoModelForSequenceClassification.from_pretrained(
    "dmis-lab/UltraBERT-380M", num_labels=2, trust_remote_code=True
)

We recommend starting from the settings used in the paper. For classification: AdamW (β1 0.9, β2 0.98, ε 1e-6), weight decay 0.1, a linear schedule with 5% warm-up, and a learning rate swept over {5e-6,8e-6,1e-5,3e-5,5e-5,8e-5}. For dense retrieval: mean pooling with a contrastive loss and a wider learning-rate sweep of {1e-5,3e-5,5e-5,8e-5,1e-4,2e-4,3e-4,5e-4,8e-4,1e-3}.

💡 Note: UltraBERT consistently prefers lower learning rates than other encoders, by up to 10x. For example, the best learning rates for UltraBERT-150M and UltraBERT-380M are 5e-5 and 3e-5, respectively, whereas other encoders typically prefer 3e-4 to 5e-4.

📊 Evaluation

Classification vs retrieval performance across encoder models

Every model is fine-tuned per task with the same grid search. For the retrieval tasks, each encoder is fine-tuned on a subset of MS MARCO or MLDR under the same sweep, so these scores reflect the potential information retrieval capability of each encoder. Dedicated embedding models are usually trained further on hundreds of millions of pairs with much longer training schedules.

Base Size (at most 200M parameters)

Model Params GLUE SQuAD 2.0 MultiCoNER v2 Few-NERD Reward-Bench BEIR RARb Code RARb Math MLDR LongEmbed
BERT 110M 84.7 76.3 53.8 66.6 59.4 36.7 6.6 27.5 29.0 29.3
RoBERTa 125M 86.4 83.7 56.5 67.1 63.2 39.6 11.7 33.4 28.7 33.8
MosaicBERT 137M 85.4 79.4 57.8 65.7 62.5 41.4 12.7 30.9 29.6 40.5
NomicBERT 137M 84.0 82.7 60.1 66.9 64.1 42.0 13.8 30.2 29.9 45.0
GTE-en-MLM 137M 85.6 82.7 59.2 66.4 63.4 42.4 12.0 33.6 37.1 58.5
ModernBERT 149M 88.4 86.7 57.4 66.7 68.0 42.6 29.6 46.1 37.7 63.2
Ettin Encoder 149M 88.9 87.4 57.4 66.5 70.3 42.4 31.8 47.1 35.9 60.1
UltraBERT 152M 89.5 89.2 57.9 67.4 73.0 45.2 39.1 55.4 43.5 72.0

Medium Size (at most 500M parameters)

Model Params GLUE SQuAD 2.0 MultiCoNER v2 Few-NERD Reward-Bench BEIR RARb Code RARb Math MLDR LongEmbed
BERT 334M 85.2 81.8 57.7 67.8 62.6 38.6 12.2 33.0 29.8 30.0
RoBERTa 355M 88.9 89.4 62.3 68.1 68.5 42.8 20.3 37.8 29.5 32.2
GTE-en-MLM 434M 87.6 87.2 63.0 67.1 62.2 39.4 13.5 32.0 40.8 57.0
NeoBERT 222M 89.0 86.6 57.4 66.4 57.6 42.9 17.3 33.4 21.7 32.4
LFM2.5-Encoder 230M 88.2 86.1 59.6 67.6 67.5 46.7 30.9 50.2 43.7 70.0
LFM2.5-Encoder 355M 89.2 87.2 62.7 67.9 68.5 47.3 38.6 50.2 44.5 71.7
ModernBERT 395M 90.4 89.8 60.9 67.6 72.8 45.2 34.0 46.1 39.9 63.5
Ettin Encoder 395M 90.8 90.4 62.3 67.8 76.7 45.9 38.5 54.0 40.9 64.3
UltraBERT 380M 91.0 90.9 63.5 68.3 78.8 47.7 47.0 61.7 44.1 74.9

Long-Context Passkey Retrieval

Passkey retrieval accuracy on LongEmbed by document length

Passkey retrieval accuracy on LongEmbed as a function of document length, for (a) base and (b) medium encoders. Every baseline degrades once the document exceeds its context length: the 512-context models fail beyond 1K tokens and the 8K-context models drop sharply at 16K and 32K. UltraBERT is the only model that retains 99% accuracy at every length up to 32K, at both sizes.

🏛️ Model Description

UltraBERT-150M UltraBERT-380M
Layers 24 24
Hidden size 768 1,024
SwiGLU size 1,024 3,072
Attention heads 6 8
Head dimension 128 128
Attention pattern local x2, global local x2, global
Local attention window 128 tokens each side 128 tokens each side
RoPE base (local / global) 10K / 2.56M 10K / 2.56M
Context length 32,768 32,768
Vocabulary 49,152 49,152
Normalization RMSNorm (eps 1e-6) RMSNorm (eps 1e-6)
Parameters 152M 380M

UltraBERT is a pre-norm Transformer encoder with RoPE, SwiGLU feed-forward layers, RMSNorm and no bias terms. Two of every three layers use local attention restricted to 128 tokens on each side, and every third layer attends over the whole sequence. The depth is fixed at 24 layers for every size and only the width is scaled. The LM head is untied from the input embeddings.

What Is in This Checkpoint

  • model.safetensors in float32 (the pre-training master weights), with config.dtype: float32.
  • The encoder and the MLM head.
  • The MLM head is the one trained on the 5% masked sequences (the distilled head); the 50% head is not included.
  • Remote code for transformers 5.x with UltrabertModel, UltrabertForMaskedLM, UltrabertForSequenceClassification, UltrabertForQuestionAnswering and UltrabertForTokenClassification.

⚙️ Pre-Training

Objective. Masked language modeling with two masking ratios in every batch, 5% and 50% at a 4:1 sequence ratio, each with its own LM head. The 5% head is trained with a KL divergence loss against the logits of a 1.3B teacher; the 50% head with the standard MLM loss.

Pruning and distillation. The 1.3B teacher is width-pruned to 380M and the distilled 380M model is pruned again to 150M. Importance of hidden channels, attention heads and FFN neurons is estimated from activations on 2,000 calibration documents, and the top-scoring units are kept. Both sizes are distilled from the 1.3B teacher directly.

Optimization. AdamW (β1 0.9, β2 0.95, ε 1e-8), weight decay 0.1, gradient clipping 1.0, peak learning rate 1e-4 with a warmup-stable-decay schedule (10B warm-up tokens, linear decay over the last 10%). Batch size 4M tokens, and 2T tokens in total; the context is extended to 32K, raising the global-attention RoPE base from 640K to 2.56M. Trained on a single TPU v4-512, v5e-256 or v6e-256 slice with data parallelism and FSDP.

Data. One fixed mixture for the whole run (Table 4 of the paper).

Category Dataset Unique tokens (B) Epochs Trained tokens (B) Share
Web DCLM 756.4 1.17 881.7 44.1%
Web FineWeb-Edu 195.4 2.33 455.5 22.8%
Code Stack-Edu 120.3 2.33 280.5 14.0%
Academic peS2o 60.7 2.33 141.6 7.1%
Math MegaMath 20.9 2.80 58.4 2.9%
Math FineMath 10.9 2.80 30.5 1.5%
Math InfiWebMath 9.6 2.80 26.9 1.3%
Reference Cosmopedia v2 28.1 2.33 65.6 3.3%
Reference Wikipedia (Dolma) 3.9 2.33 9.2 0.5%
Instruction FLAN 19.0 2.33 44.4 2.2%
Instruction StackExchange (RedPajama) 1.5 2.33 3.4 0.2%
Instruction Natural Reasoning 1.0 2.33 2.4 0.1%
Total 1,227.8 1.63 2,000.0 100%

Limitations

  • English only. The pre-training mixture is English web, code, academic, math and instruction text.
  • This is a pre-trained encoder, not a ready-to-use classifier or embedding model: it needs fine-tuning for a task. For a ready-to-use text embedding model built on it, see UltraBERT-Embed-150M.
  • The model can reproduce biases of its web-scale training data; evaluate it for your use case.

Citation

@article{park2026ultrabert,
  title   = {UltraBERT: Beyond the Limit of BERT Pre-Training},
  author  = {Park, Jungwoo and Lee, Taewhoo and Kang, Jaewoo},
  journal = {arXiv preprint arXiv:{ARXIV_ID}},
  year    = {2026}
}

Acknowledgments

The pre-training was primarily supported by Cloud TPUs from Google's TPU Research Cloud (TRC) program.

Downloads last month
78
Safetensors
Model size
0.4B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train dmis-lab/UltraBERT-380M

Collection including dmis-lab/UltraBERT-380M