dateline-loc-resolver-v1

A multilingual cross-encoder that picks the right GeoNames place for a news dateline location string, given the article it appeared in.

It is a reranker, not a geocoder: it scores candidates you supply and does not search a gazetteer itself, so you need your own candidate generation (for example a GeoNames name and alternate-name lookup) in front of it.

Usage

from sentence_transformers.cross_encoder import CrossEncoder

model = CrossEncoder("emilys/dateline-loc-resolver-v1", max_length=256)

context = "NAIROBI ; published by nation_ke in KE ; NAIROBI, March 4 (Reuters) - ..."
candidates = [
    "Nairobi ; national capital ; Nairobi County, Kenya ; population 2750547",
    "Nairobi ; town or village ; Karnataka, India ; population 0",
]
scores = model.predict([(context, c) for c in candidates])
best = candidates[int(scores.argmax())]

Each candidate is scored independently, so take the argmax over one span's candidates to get its answer.

Input format

⚠️ The two strings are a contract, not a convention. This is the format the model was trained on; departing from it degrades accuracy substantially, silently, and with no error.

Context

{span} ; published by {outlet} in {country_code} ; {article text}

The published by clause is omitted when the outlet is unknown. The article text is a window of roughly 900 characters centred on the location span.

Candidate

{name} ; {feature word} ; {admin1}, {country} ; population {n}

{feature word} is a short gloss of the GeoNames feature code — national capital, regional capital, district capital, town or village, neighbourhood, province or state, district, country, airport, and so on. The feature word and the population must both be present: they carry a large share of the signal, and omitting them at inference on a model trained with them is the single most damaging way to get the format wrong.

Details

Base model xlm-roberta-base, fine-tuned as a single-logit cross-encoder with binary cross-entropy loss. Maximum sequence length 256 tokens.

Scores are sigmoid-comparable across spans, but no abstention threshold is applied by default — the model answers whenever it is given a non-empty candidate list.

Downloads last month
22
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for emilys/dateline-loc-resolver-v1

Finetuned
(4234)
this model