Spaces:
Running
Running
|
Download README.md from chronologic/README: direct link, hf CLI and curl.
- Browser
- Download file 2.73 kB
-
https://huggingface.co/spaces/chronologic/README/resolve/main/README.md
- Command line
-
hf download hf://spaces/chronologic/README/README.md
-
curl -L -o README.md https://huggingface.co/spaces/chronologic/README/resolve/main/README.md
2.73 kB
| title: Chronologic | |
| emoji: 🕰️ | |
| colorFrom: yellow | |
| colorTo: gray | |
| sdk: static | |
| pinned: false | |
| # Chronologic | |
| **Chronologic-EN** is a benchmark that measures whether language models can respond like | |
| English-language writers placed in specific social and historical contexts between 1831 | |
| and 1930. It has 866 questions drawn from 187 period sources. Each question carries a | |
| **metadata frame** naming a context (a date, a nationality, often a genre and an | |
| author's background). Ground-truth answers come from period texts, not from us. | |
| ## Two ways to score a model | |
| - **Likelihood scoring** asks whether a model *prefers* an apt answer. Each candidate | |
| answer is scored by the model's log-probability. This needs an open-weight model. | |
| - **Free generation** asks whether a model can *produce* one. The model answers in its | |
| own words, and the answer is scored on two axes: | |
| - **Substance**, judged by an LLM comparing it with ground truth. | |
| - **Style**, judged by two fine-tuned DeBERTa instruments, one for date and one for | |
| authenticity. Style scores are model-level by design: they compare a model's answers, | |
| taken together, with matched authentic prose. | |
| ## Resources | |
| - Preprint: [https://arxiv.org/abs/2609.23178](https://arxiv.org/abs/2609.23178) | |
| - Code: [github.com/Historical-AI-Lab/Chronologic-EN](https://github.com/Historical-AI-Lab/Chronologic-EN) | |
| - Date instrument: [chronologic-date-deberta](https://huggingface.co/chronologic/chronologic-date-deberta) | |
| - Authenticity instrument: [chronologic-authenticity-deberta](https://huggingface.co/chronologic/chronologic-authenticity-deberta) | |
| - Full benchmark: [Chronologic-EN-1.0](https://huggingface.co/datasets/chronologic/Chronologic-EN-1.0). | |
| Access is controlled and requires an agreement not to redistribute the data or use it | |
| for training. | |
| **Contamination.** To keep the benchmark usable, the questions are not public. The GitHub | |
| repository has a representative 100-question sample; the full item bank is gated. Please | |
| don't post benchmark items where crawlers can reach them. | |
| ## People | |
| Ted Underwood, Ziliang Qiu, Sarah Griebel, Laura K. Nelson, Teddy Roland, Wenyi Shang and | |
| Matthew Wilkens. Contact: Ted Underwood, University of Illinois Urbana-Champaign, | |
| tunder@illinois.edu. | |
| Preprint forthcoming. | |
| ```bibtex | |
| @misc{underwood2026chronologicmeasuringlanguagemodels, | |
| title={Chronologic: Measuring Language Models' Ability to Represent the Past}, | |
| author={Ted Underwood and Ziliang Qiu and Sarah Griebel and Laura K. Nelson and Edwin Roland and Wenyi Shang and Matthew Wilkens}, | |
| year={2026}, | |
| eprint={2609.23178}, | |
| archivePrefix={arXiv}, | |
| primaryClass={cs.CL}, | |
| url={https://arxiv.org/abs/2609.23178}, | |
| } | |
| ``` | |