Spaces:
Running
Running
File size: 2,733 Bytes
63271a9 a058480 63271a9 a058480 84270c2 a058480 84270c2 a058480 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 | ---
title: Chronologic
emoji: 🕰️
colorFrom: yellow
colorTo: gray
sdk: static
pinned: false
---
# Chronologic
**Chronologic-EN** is a benchmark that measures whether language models can respond like
English-language writers placed in specific social and historical contexts between 1831
and 1930. It has 866 questions drawn from 187 period sources. Each question carries a
**metadata frame** naming a context (a date, a nationality, often a genre and an
author's background). Ground-truth answers come from period texts, not from us.
## Two ways to score a model
- **Likelihood scoring** asks whether a model *prefers* an apt answer. Each candidate
answer is scored by the model's log-probability. This needs an open-weight model.
- **Free generation** asks whether a model can *produce* one. The model answers in its
own words, and the answer is scored on two axes:
- **Substance**, judged by an LLM comparing it with ground truth.
- **Style**, judged by two fine-tuned DeBERTa instruments, one for date and one for
authenticity. Style scores are model-level by design: they compare a model's answers,
taken together, with matched authentic prose.
## Resources
- Preprint: [https://arxiv.org/abs/2609.23178](https://arxiv.org/abs/2609.23178)
- Code: [github.com/Historical-AI-Lab/Chronologic-EN](https://github.com/Historical-AI-Lab/Chronologic-EN)
- Date instrument: [chronologic-date-deberta](https://huggingface.co/chronologic/chronologic-date-deberta)
- Authenticity instrument: [chronologic-authenticity-deberta](https://huggingface.co/chronologic/chronologic-authenticity-deberta)
- Full benchmark: [Chronologic-EN-1.0](https://huggingface.co/datasets/chronologic/Chronologic-EN-1.0).
Access is controlled and requires an agreement not to redistribute the data or use it
for training.
**Contamination.** To keep the benchmark usable, the questions are not public. The GitHub
repository has a representative 100-question sample; the full item bank is gated. Please
don't post benchmark items where crawlers can reach them.
## People
Ted Underwood, Ziliang Qiu, Sarah Griebel, Laura K. Nelson, Teddy Roland, Wenyi Shang and
Matthew Wilkens. Contact: Ted Underwood, University of Illinois Urbana-Champaign,
tunder@illinois.edu.
Preprint forthcoming.
```bibtex
@misc{underwood2026chronologicmeasuringlanguagemodels,
title={Chronologic: Measuring Language Models' Ability to Represent the Past},
author={Ted Underwood and Ziliang Qiu and Sarah Griebel and Laura K. Nelson and Edwin Roland and Wenyi Shang and Matthew Wilkens},
year={2026},
eprint={2609.23178},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2609.23178},
}
```
|