--- title: Chronologic emoji: 🕰️ colorFrom: yellow colorTo: gray sdk: static pinned: false --- # Chronologic **Chronologic-EN** is a benchmark that measures whether language models can respond like English-language writers placed in specific social and historical contexts between 1831 and 1930. It has 866 questions drawn from 187 period sources. Each question carries a **metadata frame** naming a context (a date, a nationality, often a genre and an author's background). Ground-truth answers come from period texts, not from us. ## Two ways to score a model - **Likelihood scoring** asks whether a model *prefers* an apt answer. Each candidate answer is scored by the model's log-probability. This needs an open-weight model. - **Free generation** asks whether a model can *produce* one. The model answers in its own words, and the answer is scored on two axes: - **Substance**, judged by an LLM comparing it with ground truth. - **Style**, judged by two fine-tuned DeBERTa instruments, one for date and one for authenticity. Style scores are model-level by design: they compare a model's answers, taken together, with matched authentic prose. ## Resources - Preprint: [https://arxiv.org/abs/2609.23178](https://arxiv.org/abs/2609.23178) - Code: [github.com/Historical-AI-Lab/Chronologic-EN](https://github.com/Historical-AI-Lab/Chronologic-EN) - Date instrument: [chronologic-date-deberta](https://huggingface.co/chronologic/chronologic-date-deberta) - Authenticity instrument: [chronologic-authenticity-deberta](https://huggingface.co/chronologic/chronologic-authenticity-deberta) - Full benchmark: [Chronologic-EN-1.0](https://huggingface.co/datasets/chronologic/Chronologic-EN-1.0). Access is controlled and requires an agreement not to redistribute the data or use it for training. **Contamination.** To keep the benchmark usable, the questions are not public. The GitHub repository has a representative 100-question sample; the full item bank is gated. Please don't post benchmark items where crawlers can reach them. ## People Ted Underwood, Ziliang Qiu, Sarah Griebel, Laura K. Nelson, Teddy Roland, Wenyi Shang and Matthew Wilkens. Contact: Ted Underwood, University of Illinois Urbana-Champaign, tunder@illinois.edu. Preprint forthcoming. ```bibtex @misc{underwood2026chronologicmeasuringlanguagemodels, title={Chronologic: Measuring Language Models' Ability to Represent the Past}, author={Ted Underwood and Ziliang Qiu and Sarah Griebel and Laura K. Nelson and Edwin Roland and Wenyi Shang and Matthew Wilkens}, year={2026}, eprint={2609.23178}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2609.23178}, } ```