mmBERT-small fine-tuned for multilingual NER

Token classification (PER / ORG / LOC, IOB2) fine-tuned from jhu-clsp/mmBERT-small on a 32-language news NER dataset.

Training

  • lr 8e-5, up to 10 epochs with early stopping (patience 3) on validation F1, warmup 0.1, weight decay 0.01, max length 256
  • Training labels denoised by masking contradictory repeat mentions from the loss
  • Uniform weight average of 3 seeds

Validation results

Leak-free, stratified validation split, seqeval micro scores at 256-token truncation.

metric value
F1 0.7347
precision 0.6662
recall 0.8189

Per class F1: LOC 0.7804, ORG 0.6953, PER 0.7876.

Usage

from transformers import pipeline

ner = pipeline("token-classification", model="clincolnoz/mmbert-small-ner", aggregation_strategy="simple")
ner("Angela Merkel met executives from Siemens in Munich.")
Downloads last month
27
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for clincolnoz/mmbert-small-ner

Finetuned
(56)
this model