Text Classification
Transformers
Safetensors
English
deberta-v2
digital-humanities
historical-text
dating
deberta-v3
chronologic
Eval Results (legacy)
text-embeddings-inference
Instructions to use chronologic/chronologic-date-deberta with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use chronologic/chronologic-date-deberta with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="chronologic/chronologic-date-deberta")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("chronologic/chronologic-date-deberta") model = AutoModelForSequenceClassification.from_pretrained("chronologic/chronologic-date-deberta", device_map="auto") - Notebooks
- Google Colab
- Kaggle
File size: 11,420 Bytes
a3ec452 9b1bd06 a3ec452 9b1bd06 6511388 9b1bd06 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 | ---
license: mit
base_model: microsoft/deberta-v3-large
library_name: transformers
language: en
pipeline_tag: text-classification
tags:
- digital-humanities
- historical-text
- dating
- deberta-v3
- chronologic
metrics:
- mae
- crps
widget:
- text: "One evening of late summer, before the nineteenth century had reached one-third of its span, a young man and woman, the latter carrying a child, were approaching the large village of Weydon-Priors, in Upper Wessex, on foot."
example_title: "Hardy, 1886"
- text: "She scrolled through her messages one more time, hoping the notification she had been waiting for all afternoon would finally appear."
example_title: "Present-day prose"
model-index:
- name: chronologic-date-deberta
results:
- task:
type: text-classification
name: Publication-date prediction (1830-1930 core)
dataset:
type: chronologic-date-test-core
name: Held-out test split, 1830-1930 (n=2,032)
metrics:
- type: crps
name: CRPS (years)
value: 16.69
- type: mae
name: MAE (years)
value: 23.09
- task:
type: text-classification
name: Publication-date prediction (all decades)
dataset:
type: chronologic-date-test
name: Held-out test split, 1700s-2020s (n=6,915)
metrics:
- type: crps
name: CRPS (years)
value: 20.13
- type: mae
name: MAE (years)
value: 26.94
- task:
type: text-classification
name: Publication-date prediction (reserved benchmark volumes)
dataset:
type: chronologic-date-benchmarkbooks
name: Reserved benchmark source volumes (n=606)
metrics:
- type: crps
name: CRPS (years)
value: 17.66
- type: mae
name: MAE (years)
value: 24.20
---
# chronologic-date-deberta
A DeBERTa-v3-large model that predicts the publication date of an English passage as a
probability distribution over decades from the 1680s to the 2030s. It is one of the two
style instruments in the [Chronologic-EN](https://github.com/Historical-AI-Lab/Chronologic-EN)
benchmark. The companion model is
[chronologic-authenticity-deberta](https://huggingface.co/chronologic/chronologic-authenticity-deberta).
**What the output means.** The model outputs 36 logits, one per ten-year bin from 1680
to 2040. The softmax over them is a distribution over decades: label `1880s` is the
probability that the passage was published between 1880 and 1889. The model's year
estimate is the probability-weighted mean of the bin midpoints (1685, 1695, …, 2035).
The width of the distribution carries information too, which is why the benchmark
keeps the whole distribution instead of the top label.
## Intended use and out-of-scope use
This model is a **model-level benchmark instrument**. It measures how far a language
model's answers, taken together, sit from matched authentic prose of the target period.
It is **not** a tool for dating individual documents. Its error on a single passage is
typically two decades or more (MAE about 23 years in the core period), and its target is
publication date, not composition date.
**Raw outputs are not Chronologic scores.** The benchmark never reads the predicted date
as a verdict. It ranks the model's error on an answer against the errors it makes on
authentic prose. That step needs the code in the GitHub repository and the scored
reference corpus (the `CHRONOLOGIC_DATA` directory). A number you compute from this
model alone can't be compared with published Period Fidelity scores.
## How the benchmark uses it
For each answer the benchmark takes the signed residual (predicted mean year minus the
question's target year). It converts that residual to a percentile among the residuals of
authentic passages from within ±10 years of the target date, in the same length bin,
scored by the same frozen model. So the model's own biases cancel, and a model writing
indistinguishably from period prose would produce uniform percentiles. Period Fidelity
(0–100) is the Wasserstein-1 distance of those percentiles from uniform, rescaled against
the distance that resampled authentic passages themselves produce (a pseudo-model
baseline). So 100 means indistinguishable from genuine prose of the target period.
The benchmark also applies temperature scaling (T = 1.0551, fit on held-out authentic
text) and abstains on sentence fragments (see Usage).
## Training data
- **Roster:** 4,878 volumes, 150 per decade across the 1831–1930 core and 100 per decade
elsewhere. Sources:
- IDI (Harvard's Institutional Data Initiative), used first
- ECCO before 1800, and OAPEN after 1925, where IDI falls short
- the Chicago Novel Corpus, capped at 30% of twentieth-century fiction
- COHA, uncapped, for genre diversity
- A MARC date-reliability filter drops volumes with unreliable dates (continuing
resources, questionable dates), leaving 4,325 volumes. Hand corrections to publication
dates are folded in from catalog records only.
- **Passages:** 66,000, exactly 2,000 per decade from the 1700s to the 2020s. Their
length distribution is matched to real benchmark answers.
- **Splits:** 52,651 train, 6,434 validation and 6,915 test, grouped by volume and by
author.
- Decades outside 1831–1930 are there so the model can say "this reads like 1770" or
"this reads like 1990". The benchmark has to detect drift in both directions. We don't
claim accuracy there.
- The benchmark's reserved source volumes are excluded from training.
- The training text is not redistributed, because some sources (COHA, the Chicago Novel
Corpus, OAPEN) are in copyright.
## Training procedure
- `microsoft/deberta-v3-large` with `num_labels=36`, a softmax over ten-year bins from
1680 to 2040. The bin grid is stored on the config as `date_grid_lo` (1680),
`date_grid_hi` (2040), `date_bin_width` (10) and `date_sigma` (15).
- **Soft-target cross-entropy.** A passage published in 1883 is trained against a
Gaussian bump centred on 1883 with σ = 15 years, not a one-hot label. So missing by one
bin costs far less than missing by ten. This is an "ordinal softmax" in that loose
sense only. It is not a regressor, and not a cumulative-link or CORAL ordinal model.
- Max length 256 tokens.
- The checkpoint was kept at epoch 2, the validation-loss minimum (loss 2.545).
## Evaluation
CRPS (continuous ranked probability score) is the primary metric. It scores the whole
predicted distribution, is expressed in years, and reduces to absolute error for a point
prediction. All figures are at T = 1.
| Evaluation | n | CRPS (yr) | MAE (yr) | Other |
|---|---|---|---|---|
| Test split, 1830–1930 core | 2,032 | 16.69 | 23.09 | cov@50 0.63, cov@90 0.95, mean signed residual −4.3 yr |
| Test split, all decades | 6,915 | 20.13 | 26.94 | cov@90 0.90, R² 0.795 |
| Reserved benchmark volumes | 606 | 17.66 | 24.20 | NLL 2.50 |
"cov@50" and "cov@90" are the share of passages whose true date falls inside the central
50% and 90% intervals of the predicted distribution.
Comparison with the lexical predecessor (tf-idf plus multinomial logistic regression,
same soft-target objective):
- On identical splits, CRPS is 20.13 against 44.46.
- On the reserved benchmark volumes, MAE / CRPS / NLL are 24.20 / 17.66 / 2.50 against
28.10 / 29.51 / 3.44.
## Limitations and biases
- **Early eighteenth century.** The 1700s and 1710s fit badly (MAE 89 and 106 years),
from a pool of only seventeen usable volumes.
- **Residual bias.** Mean signed residuals across the 1830s–1920s run from +4 to −10
years, except the 1900s at −23. The lexical predecessor's pull toward the centre of the
range (+28 years at the 1830s to −43 at the 1930s) is mostly gone, but not entirely.
- **Publication date, not composition date.** Reprints, posthumous editions and
late-published manuscripts are labelled by publication.
- **Fragments.** The model is trained on complete sentences. It is out of distribution
on sub-sentential fragments, and the benchmark abstains on them. A raw model call does
not abstain.
- **Typography.** The training text was normalized. Unnormalized input (curly quotes,
em-dashes, line breaks) is out of distribution. Normalize first (see Usage).
## A note on Goodhart's law
Publishing the judges makes them possible to train against. Any model trained, tuned,
RL'd or selected against these instruments, or against models derived from them, no
longer has valid Chronologic style scores. If you report Chronologic style scores for
such a model, say that it was optimized against the judges.
## Usage
**(a) Pipeline.** Quick inspection only. It skips normalization, temperature and the
fragment check.
```python
from transformers import pipeline
clf = pipeline("text-classification", model="chronologic/chronologic-date-deberta", top_k=5)
clf("The railway had lately come to the town, and with it a new class of traveller.")
# [[{'label': '1860s', 'score': ...}, {'label': '1870s', 'score': ...}, ...]]
```
**(b) Explicit code.** Normalize with `normalize_typography` from `stylejudge/normalize.py`
in the [GitHub repository](https://github.com/Historical-AI-Lab/Chronologic-EN) (standard
library only), skip fragments, apply the benchmark's temperature, and take the mean year:
```python
import sys, torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification
sys.path.insert(0, "Chronologic-EN/stylejudge") # path to your clone
from normalize import normalize_typography
from measure_length_distribution import is_fragment # needs nltk
repo = "chronologic/chronologic-date-deberta"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForSequenceClassification.from_pretrained(repo).eval()
cfg = model.config
midpoints = torch.arange(cfg.date_grid_lo, cfg.date_grid_hi, cfg.date_bin_width) + cfg.date_bin_width / 2
T = 1.0551 # the benchmark's temperature (t_nll in stylejudge/e3_temperature_fit.json)
texts = [normalize_typography(t) for t in ["...your passages..."]]
texts = [t for t in texts if not is_fragment(t)] # the benchmark abstains on fragments
enc = tok(texts, padding=True, truncation=True, max_length=256, return_tensors="pt")
with torch.no_grad():
probs = torch.softmax(model(**enc).logits / T, dim=-1)
mean_year = probs @ midpoints
```
`is_fragment` flags a passage whose first alphabetic character is lowercase and that is
at most one sentence long.
**(c) Benchmark scoring.** For real Chronologic scores, use `stylejudge/score_style.py`
(or the full `modelasjudge/run_pipeline.py`) from the GitHub repository. These apply the
percentile layer against the reference corpus.
## Citation
Preprint forthcoming.
```bibtex
@misc{underwood2026chronologicmeasuringlanguagemodels,
title={Chronologic: Measuring Language Models' Ability to Represent the Past},
author={Ted Underwood and Ziliang Qiu and Sarah Griebel and Laura K. Nelson and Edwin Roland and Wenyi Shang and Matthew Wilkens},
year={2026},
eprint={2609.23178},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2609.23178},
}
```
## Contact
Ted Underwood, University of Illinois Urbana-Champaign, tunder@illinois.edu.
|