README / README.md
tedunderwood's picture
updated with preprint citations
84270c2 verified
|
Raw History Blame Contribute Delete
2.73 kB
---
title: Chronologic
emoji: 🕰️
colorFrom: yellow
colorTo: gray
sdk: static
pinned: false
---
# Chronologic
**Chronologic-EN** is a benchmark that measures whether language models can respond like
English-language writers placed in specific social and historical contexts between 1831
and 1930. It has 866 questions drawn from 187 period sources. Each question carries a
**metadata frame** naming a context (a date, a nationality, often a genre and an
author's background). Ground-truth answers come from period texts, not from us.
## Two ways to score a model
- **Likelihood scoring** asks whether a model *prefers* an apt answer. Each candidate
answer is scored by the model's log-probability. This needs an open-weight model.
- **Free generation** asks whether a model can *produce* one. The model answers in its
own words, and the answer is scored on two axes:
- **Substance**, judged by an LLM comparing it with ground truth.
- **Style**, judged by two fine-tuned DeBERTa instruments, one for date and one for
authenticity. Style scores are model-level by design: they compare a model's answers,
taken together, with matched authentic prose.
## Resources
- Preprint: [https://arxiv.org/abs/2609.23178](https://arxiv.org/abs/2609.23178)
- Code: [github.com/Historical-AI-Lab/Chronologic-EN](https://github.com/Historical-AI-Lab/Chronologic-EN)
- Date instrument: [chronologic-date-deberta](https://huggingface.co/chronologic/chronologic-date-deberta)
- Authenticity instrument: [chronologic-authenticity-deberta](https://huggingface.co/chronologic/chronologic-authenticity-deberta)
- Full benchmark: [Chronologic-EN-1.0](https://huggingface.co/datasets/chronologic/Chronologic-EN-1.0).
Access is controlled and requires an agreement not to redistribute the data or use it
for training.
**Contamination.** To keep the benchmark usable, the questions are not public. The GitHub
repository has a representative 100-question sample; the full item bank is gated. Please
don't post benchmark items where crawlers can reach them.
## People
Ted Underwood, Ziliang Qiu, Sarah Griebel, Laura K. Nelson, Teddy Roland, Wenyi Shang and
Matthew Wilkens. Contact: Ted Underwood, University of Illinois Urbana-Champaign,
tunder@illinois.edu.
Preprint forthcoming.
```bibtex
@misc{underwood2026chronologicmeasuringlanguagemodels,
title={Chronologic: Measuring Language Models' Ability to Represent the Past},
author={Ted Underwood and Ziliang Qiu and Sarah Griebel and Laura K. Nelson and Edwin Roland and Wenyi Shang and Matthew Wilkens},
year={2026},
eprint={2609.23178},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2609.23178},
}
```