Spaces:
Running
Download README.md from chronologic/README: direct link, hf CLI and curl.
- Browser
- Download file 2.73 kB
-
https://huggingface.co/spaces/chronologic/README/resolve/main/README.md
- Command line
-
hf download hf://spaces/chronologic/README/README.md
-
curl -L -o README.md https://huggingface.co/spaces/chronologic/README/resolve/main/README.md
title: Chronologic
emoji: 🕰️
colorFrom: yellow
colorTo: gray
sdk: static
pinned: false
Chronologic
Chronologic-EN is a benchmark that measures whether language models can respond like English-language writers placed in specific social and historical contexts between 1831 and 1930. It has 866 questions drawn from 187 period sources. Each question carries a metadata frame naming a context (a date, a nationality, often a genre and an author's background). Ground-truth answers come from period texts, not from us.
Two ways to score a model
- Likelihood scoring asks whether a model prefers an apt answer. Each candidate answer is scored by the model's log-probability. This needs an open-weight model.
- Free generation asks whether a model can produce one. The model answers in its
own words, and the answer is scored on two axes:
- Substance, judged by an LLM comparing it with ground truth.
- Style, judged by two fine-tuned DeBERTa instruments, one for date and one for authenticity. Style scores are model-level by design: they compare a model's answers, taken together, with matched authentic prose.
Resources
- Preprint: https://arxiv.org/abs/2609.23178
- Code: github.com/Historical-AI-Lab/Chronologic-EN
- Date instrument: chronologic-date-deberta
- Authenticity instrument: chronologic-authenticity-deberta
- Full benchmark: Chronologic-EN-1.0. Access is controlled and requires an agreement not to redistribute the data or use it for training.
Contamination. To keep the benchmark usable, the questions are not public. The GitHub repository has a representative 100-question sample; the full item bank is gated. Please don't post benchmark items where crawlers can reach them.
People
Ted Underwood, Ziliang Qiu, Sarah Griebel, Laura K. Nelson, Teddy Roland, Wenyi Shang and Matthew Wilkens. Contact: Ted Underwood, University of Illinois Urbana-Champaign, tunder@illinois.edu.
Preprint forthcoming.
@misc{underwood2026chronologicmeasuringlanguagemodels,
title={Chronologic: Measuring Language Models' Ability to Represent the Past},
author={Ted Underwood and Ziliang Qiu and Sarah Griebel and Laura K. Nelson and Edwin Roland and Wenyi Shang and Matthew Wilkens},
year={2026},
eprint={2609.23178},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2609.23178},
}