DistilBERT fine-tuned on WNUT 17 for Named Entity Recognition

Model Description

This model is a fine-tuned version of distilbert/distilbert-base-uncased for Named Entity Recognition (NER).

It was fine-tuned on the WNUT 17 Emerging and Rare Entity Recognition dataset to identify named entities in English text.

The model was created as a practical exercise while following the Hugging Face LLM Course.

  • Developed by: Driw0x
  • Model type: DistilBERT
  • Language: English
  • Base model: distilbert/distilbert-base-uncased
  • Task: Token classification / Named Entity Recognition
  • Number of labels: 13

Entity Labels

The model uses BIO-formatted labels for six entity categories:

  • corporation
  • creative work
  • group
  • location
  • person
  • product

The complete label mapping is:

ID Label
0 O
1 B-corporation
2 I-corporation
3 B-creative-work
4 I-creative-work
5 B-group
6 I-group
7 B-location
8 I-location
9 B-person
10 I-person
11 B-product
12 I-product

Training and Evaluation Data

The model was fine-tuned on:

flaitenberger/wnut_17

The training notebook uses:

  • the train split for training;
  • the test split for evaluation.

The validation split was not used in this experiment.

Because the test split is evaluated during training, it also participates in model development and should not be considered a completely untouched final test set.

Preprocessing

WNUT 17 provides labels at the word level, while DistilBERT uses subword tokenization.

The labels are therefore aligned with the generated subword tokens.

def tokenize_and_align_labels(examples):
    tokenized_inputs = tokenizer(
        examples["tokens"],
        truncation=True,
        is_split_into_words=True
    )
    labels = []
    for i, label in enumerate(examples["ner_tags"]):
        word_ids = tokenized_inputs.word_ids(batch_index=i)
        previous_word_idx = None
        label_ids = []
        for word_idx in word_ids:
            if word_idx is None:
                label_ids.append(-100)
            elif word_idx != previous_word_idx:
                label_ids.append(label[word_idx])
            else:
                label_ids.append(-100)
            previous_word_idx = word_idx
        labels.append(label_ids)
    tokenized_inputs["labels"] = labels
    return tokenized_inputs

Special tokens are assigned the ignored label -100.

When a word is split into multiple subtokens, only the first subtoken keeps the original entity label. The remaining subtokens are ignored during loss computation and evaluation.

Dynamic padding is handled using DataCollatorForTokenClassification.

Training Procedure

The model was fine-tuned using the Hugging Face Trainer API.

Training Hyperparameters

Hyperparameter Value
Learning rate 2e-5
Train batch size 16
Evaluation batch size 16
Number of epochs 2
Weight decay 0.01
Seed 42
Evaluation strategy epoch
Save strategy epoch
LR scheduler linear
Load best model at end True

The optimizer used during training was AdamW (ADAMW_TORCH_FUSED) with:

  • betas=(0.9, 0.999)
  • epsilon=1e-8

Evaluation was performed with seqeval.

The following metrics were computed:

  • precision;
  • recall;
  • F1;
  • token-level accuracy.

Training Results

Epoch Validation Loss Precision Recall F1 Accuracy
1 0.280841 0.490000 0.227062 0.310323 0.937839
2 0.271229 0.545000 0.303058 0.389518 0.942114

The complete training run reported:

  • Training loss: 0.208966
  • Training steps: 426
  • Epochs: 2

The final evaluation results are:

  • Validation loss: 0.271229
  • Precision: 0.545000
  • Recall: 0.303058
  • F1: 0.389518
  • Accuracy: 0.942114

The relatively high token-level accuracy should be interpreted together with the lower entity-level F1 score.

Most tokens do not correspond to named entities and receive the O label, which means token accuracy alone is not sufficient to assess NER performance.

The evaluation also produced a seqeval warning indicating that some labels had no predicted samples.

Usage

Pipeline

from transformers import pipeline
classifier = pipeline(
    "ner",
    model="Driw0x/my_awesome_wnut_model"
)
text = (
    "The Golden State Warriors are an American professional "
    "basketball team based in San Francisco."
)
print(classifier(text))

The example used in the training notebook identifies location tokens corresponding to parts of:

  • Golden State
  • San Francisco

Direct Inference

import torch
from transformers import AutoTokenizer, AutoModelForTokenClassification
tokenizer = AutoTokenizer.from_pretrained(
    "Driw0x/my_awesome_wnut_model"
)
model = AutoModelForTokenClassification.from_pretrained(
    "Driw0x/my_awesome_wnut_model"
)
text = (
    "The Golden State Warriors are an American professional "
    "basketball team based in San Francisco."
)
inputs = tokenizer(
    text,
    return_tensors="pt"
)
with torch.no_grad():
    logits = model(**inputs).logits
predictions = torch.argmax(
    logits,
    dim=2
)
predicted_labels = [
    model.config.id2label[token.item()]
    for token in predictions[0]
]
print(predicted_labels)

Intended Uses

This model is primarily intended for:

  • learning token classification with Transformers;
  • experimenting with Named Entity Recognition;
  • learning word-to-subword label alignment;
  • experimenting with seqeval;
  • reproducing a Hugging Face token-classification workflow.

Limitations

Important limitations include:

  • the model was trained on the relatively small WNUT 17 dataset;
  • WNUT 17 focuses on emerging and rare entities and may not generalize to all NER domains;
  • recall and F1 remain relatively limited;
  • some entity labels may receive very few or no predictions;
  • token-level accuracy is influenced by the dominant non-entity O class;
  • the WNUT test split was used during training for evaluation;
  • biases inherited from DistilBERT or the training dataset may remain.

This model was created as a course exercise and has not been validated for production or high-stakes information extraction.

Framework Versions

  • Transformers 5.17.0
  • PyTorch 2.11.0+cu130
  • Datasets 4.8.5
  • Tokenizers 0.23.2

Training Source

The complete training procedure is available in the following repository:

Driw0x/hf-ai-courses

Notebook:

llm-course/1-transformer-models/notebooks/token_classification.ipynb

Downloads last month
64
Safetensors
Model size
66.4M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Driw0x/my_awesome_wnut_model

Finetuned
(12680)
this model

Dataset used to train Driw0x/my_awesome_wnut_model