BERT NER on CoNLL-2003

Intended use

English newswire named entity recognition (PER, ORG, LOC, MISC) for a course experiment. Not a general-domain or safety-critical extractor.

Model and training

  • Base: bert-base-cased
  • Dataset: lhoestq/conll2003; splits train/validation/test; original shared task: CoNLL-2003.
  • Adaptation: full; initial seed 42; 3 epochs; batch 16; maximum sequence length 256.
  • Two compared alternatives: frozen BERT with trained token head and full fine-tuning. Head learning rate 1e-3; encoder learning rate 2e-5 when trained.
  • Labels: O, B-PER, I-PER, B-ORG, I-ORG, B-LOC, I-LOC, B-MISC, I-MISC. Supervise the first subtoken only; other positions -100.

Results

Entity-level strict IOB2 span F1 on held-out test: 0.9120; precision 0.9113; recall 0.9127; token accuracy 0.9824. Validation F1: 0.9455. Single seed; close differences may be noise.

Limitations

English Reuters news from 1996; domain and time shift can hurt performance. One label per word; truncated sentences beyond 256 subtokens lose supervised words. Named entities and noisy labels may differ across datasets. Review the upstream dataset terms before reuse or redistribution.

References

Downloads last month
9
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for Perry-DLC/upy-tds-bert-ner-conll2003