LondonLB: personal-name recognition for the London Letter-Books

LondonLB is a fine-tuned version of microsoft/deberta-v3-base that identifies personal names in Reginald R. Sharpe's Calendar of Letter-Books of the City of London, volumes A to I (c. 1275–1422). It was built to support research on medieval London naming practices, in particular the transition from by-names to hereditary family names, by making it possible to extract every named individual from roughly 14,800 calendar entries.

On a held-out gold-standard set of 200 entries containing 1,136 names, the model reaches an exact-match F1 of 0.977, and finds 1,135 of the 1,136 names at least partially.

Model details

Developed by John McEwan
Model type Token classification (BIO tagging), single entity type PER
Labels O, B-PER, I-PER
Base model microsoft/deberta-v3-base
Language English (Sharpe's calendar is an English-language summary of records originally written in Latin and Anglo-Norman French; names appear in the forms given by the editor)
Training date September 2026 (run 2026sep9)
License MIT

Intended use

The model is intended for researchers working with Sharpe's calendar who want to locate personal names automatically, for example to build name indexes, carry out prosopographical work, or study naming patterns across the period. Its output is a set of character spans in the input text, each with a confidence score.

It may also be useful as a starting point for similar English-language calendars of medieval records, though it has not been tested on them.

Out of scope. The model does not recognise places, institutions or other entity types; it does not link names to individuals (two mentions of "John de Prestone" are not identified as the same person); and it does not separate a name into forename, by-name or surname. It has not been evaluated on original Latin or French texts, on other editors' calendars, or on modern prose, and should not be expected to perform well on them without further fine-tuning.

How to use

from transformers import pipeline

ner = pipeline(
    "token-classification",
    model="jmcewan3/LondonLB",
    aggregation_strategy="simple",
    stride=128,   # process long entries as overlapping 512-token windows
)

text = "Nicholas de Suffok, William le Blond, Robert Heyrun."
for ent in ner(text):
    # Slice the original text with start/end rather than using ent["word"],
    # which can differ slightly from the source because of tokenization.
    print(f"{text[ent['start']:ent['end']]!r}  {ent['score']:.3f}")

Output:

'Nicholas de Suffok'  1.000
'William le Blond'  1.000
'Robert Heyrun'  1.000

Most calendar entries fit within the model's 512-token window. A small number, typically long lists of names, do not; the stride argument splits these into overlapping windows so that names near the end are not lost.

Training data

Source

The text comes from the British History Online edition of Sharpe's Calendar of Letter-Books, volumes A–I, which together comprise about 2,770 pages and 14,800 entries. Because the online entries are not individually numbered or dated, entry numbers and dates were added manually before sampling.

Sampling

An initial 1,000 entries were drawn at random from across the whole chronological range, stratified by period to ensure balanced coverage from the late thirteenth to the early fifteenth century. After a first round of fine-tuning, further entries were added to strengthen record types the model found difficult, especially entries containing long lists of names. The final training set contains 1,603 entries and 8,274 annotated names.

Annotation

All entries were annotated manually by the author in Label Studio. A name span covers the baptismal name and the next two significant name elements, as it appears in the text (e.g. Thomas the scrivener in Frydaystret).

Key conventions:

  • Nested names are treated as separate name: "Nicholas de Hockele, fishmonger, son of John de Hockele, late fishmonger", is annotated as two separate names.
  • All titles, such as 'Master', 'Lady', or 'Sir' are omitted.
  • Offices, such as 'mayor' or 'sheriff', are not tagged, but occupations are tagged.
  • Plural occupational descriptors shared by several people are excluded: in "Thomas de Chestrehunte and Vincent de Totenham, cordwainers", only the two names are tagged.
  • Royal names include the regnal number (Henry III), but regnal-year dating formulas (15 Richard II.) are not tagged.
  • Dating formulas that reference saints are not tagged.

Preprocessing

Annotations were exported from Label Studio as character offsets. Leading and trailing whitespace and trailing punctuation (, . ; :) were trimmed from each span, and duplicate spans removed. Each entry was tokenized once into overlapping 512-token windows (stride 128) and the character spans projected onto tokens as BIO labels. A name cut by a window edge is masked from the loss in that window; the overlap guarantees it appears whole in a neighbouring window. All 8,274 names fall entirely within at least one window.

Training procedure

Entries were split 90/10 into training (1,442 entries, 1,454 windows) and validation (161 entries) sets with seed 42. The split was made on whole entries before windowing, so no text is shared between the two sets.

Hyperparameter Value
Learning rate 1e-5, linear schedule, 6% warmup
Batch size 4 per device × 4 gradient-accumulation steps (effective 16)
Optimizer AdamW (non-fused), weight decay 0.01
Gradient clipping 1.0
Max epochs 20, with early stopping (patience 4) on validation F1
Best epoch 9 (training stopped after epoch 13)
Precision float32
Max sequence length / stride 512 / 128
Seed 42

Training used the Hugging Face Trainer with gradient checkpointing on a single Intel GPU via PyTorch's XPU backend (PyTorch 2.7.0, Transformers 4.56.2). The checkpoint with the highest validation F1 is the one published here.

Evaluation

Evaluation set Precision Recall F1
Validation, 161 entries (entity-level, token-based) 0.969 0.982 0.975
Gold standard, 200 entries, exact span match 0.975 0.980 0.977
Gold standard, relaxed (any overlap counts) 0.994 0.999 0.997

The validation score was used to choose the best checkpoint, so it is a slightly optimistic estimate. The gold-standard set is a separate collection of 200 entries containing 1,136 names, annotated by the author and never used for training or model selection. It was scored at the character-span level, comparing the exact start and end of each predicted name with the gold annotation. One very short gold entry (a conventual single-sentence cancellation note) also appears verbatim in the training data; its effect on the scores is negligible.

Error analysis

Of the 1,136 gold names, the model found 1,113 exactly, placed the boundaries incorrectly on 22, and missed 1 entirely. It also returned 3 spurious names. Boundary errors therefore account for almost all of its mistakes, and fall into a few patterns:

  • Stopping short before an extended locational or descriptive element, e.g. Richard Clerk of Waltone for Richard Clerk of Waltone on Thames, or Rohesia de Coventre for Rohesia de Coventre in Westchepe.
  • Running on to include words that are not part of the name: editorial sic (Simon de Swanlond sic), plural occupational descriptors (Vincent de Totenham, cordwainers), and occasionally offices (Ralph de Heyham, Serjeant).
  • Fragmenting a name at internal punctuation, which happened three times, e.g. at the apostrophe in Thomas Guillim of St. Jean d'Angély and at the editorial query in William (John ?) Michel.

The three spurious predictions were two place names (Gascony, Stebenhethe) and a by-name standing alone (le Balauncer); the one missed name was a bare surname (Wychingham).

The strict-match F1 counts each wrong-boundary prediction as both a miss and a false positive, which is why the relaxed score is about two points higher. For most search and indexing purposes the relaxed figure better reflects how often the model finds a name; for analysis that depends on the exact form of a name, the strict figure applies.

Limitations

  • Single annotator. All training and gold data were annotated by one person, so the model learns that annotator's conventions for where a name begins and ends. Inter-annotator agreement has not been measured.
  • One source. The model has been trained and evaluated only on Sharpe's calendar of Letter-Books A–I. Performance on later volumes (J–L), other calendars, or differently edited texts is unknown.
  • Name boundaries. Users who need exact name forms should review predictions in which a name is followed by a comma and a descriptor, contains internal punctuation, or ends in a long place-name qualifier, since these account for most errors.
  • Pipeline decoding. The gold-standard scores were computed with custom decoding that merges overlapping windows by keeping each token's prediction from the window where it has the most surrounding context. The standard pipeline merges windows differently, so results on very long entries may differ slightly.

Training and evaluation code

The notebook used to convert the Label Studio exports and train the model is included in this repository.

Citation

@misc{londonlb2026,
  author       = {AUTHOR NAME},
  title        = {LondonLB: Personal-Name Recognition for the Calendar of Letter-Books of the City of London},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/jmcewan3/LondonLB}}
}

References

Downloads last month
2
Safetensors
Model size
0.2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for jmcewan3/LondonLB

Finetuned
(778)
this model

Paper for jmcewan3/LondonLB