File size: 2,733 Bytes
63271a9
a058480
 
 
 
63271a9
 
 
 
a058480
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
84270c2
a058480
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
84270c2
 
 
 
 
 
 
 
a058480
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
---
title: Chronologic
emoji: 🕰️
colorFrom: yellow
colorTo: gray
sdk: static
pinned: false
---

# Chronologic

**Chronologic-EN** is a benchmark that measures whether language models can respond like
English-language writers placed in specific social and historical contexts between 1831
and 1930. It has 866 questions drawn from 187 period sources. Each question carries a
**metadata frame** naming a context (a date, a nationality, often a genre and an
author's background). Ground-truth answers come from period texts, not from us.

## Two ways to score a model

- **Likelihood scoring** asks whether a model *prefers* an apt answer. Each candidate
  answer is scored by the model's log-probability. This needs an open-weight model.
- **Free generation** asks whether a model can *produce* one. The model answers in its
  own words, and the answer is scored on two axes:
  - **Substance**, judged by an LLM comparing it with ground truth.
  - **Style**, judged by two fine-tuned DeBERTa instruments, one for date and one for
    authenticity. Style scores are model-level by design: they compare a model's answers,
    taken together, with matched authentic prose.

## Resources

- Preprint: [https://arxiv.org/abs/2609.23178](https://arxiv.org/abs/2609.23178)
- Code: [github.com/Historical-AI-Lab/Chronologic-EN](https://github.com/Historical-AI-Lab/Chronologic-EN)
- Date instrument: [chronologic-date-deberta](https://huggingface.co/chronologic/chronologic-date-deberta)
- Authenticity instrument: [chronologic-authenticity-deberta](https://huggingface.co/chronologic/chronologic-authenticity-deberta)
- Full benchmark: [Chronologic-EN-1.0](https://huggingface.co/datasets/chronologic/Chronologic-EN-1.0).
  Access is controlled and requires an agreement not to redistribute the data or use it
  for training.

**Contamination.** To keep the benchmark usable, the questions are not public. The GitHub
repository has a representative 100-question sample; the full item bank is gated. Please
don't post benchmark items where crawlers can reach them.

## People

Ted Underwood, Ziliang Qiu, Sarah Griebel, Laura K. Nelson, Teddy Roland, Wenyi Shang and
Matthew Wilkens. Contact: Ted Underwood, University of Illinois Urbana-Champaign,
tunder@illinois.edu.

Preprint forthcoming.

```bibtex
@misc{underwood2026chronologicmeasuringlanguagemodels,
      title={Chronologic: Measuring Language Models' Ability to Represent the Past}, 
      author={Ted Underwood and Ziliang Qiu and Sarah Griebel and Laura K. Nelson and Edwin Roland and Wenyi Shang and Matthew Wilkens},
      year={2026},
      eprint={2609.23178},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2609.23178}, 
}
```