Edgar-base

Edgar-base is a 632M-parameter language model trained from scratch on English text from 1786โ€“1849, augmented with synthetic data derived from that same material. This is a non instruction-tuned checkpoint.

Model details

  • Developed by: Craig Messner, Johns Hopkins University Center for Digital Humanities
  • Model type: decoder-only transformer (LlamaForCausalLM), trained from scratch
  • Language: English (late 18th to mid 19th century)
  • License: MIT
  • Code: built on historical-perspectival-lm

Architecture

Parameters 631,504,512
Layers 42
Hidden size 1152
Feed-forward size 3072
Attention heads 18 (head dim 64)
Key/value heads 6 (grouped-query attention, 3:1)
Max positions 1024
Vocabulary 32,000 (BPE)
Embeddings tied input/output
Weights float32, safetensors

The design is deep and thin with grouped-query attention and tied embeddings, scaled up from the BabyLlama-2 recipe.

Intended uses

  • Text continuation in the register of 1786โ€“1849 English prose and verse.
  • A lightweight base for further fine-tuning (instruction tuning, and especially author-specific adaptation).
  • Research on leakage and fidelity in domain-limited models.

How to use

from transformers import AutoModelForCausalLM, AutoTokenizer

tok = AutoTokenizer.from_pretrained("hplm/edgar-base")
model = AutoModelForCausalLM.from_pretrained("hplm/edgar-base")

prompt = "It was upon a dreary evening in the autumn of the year"
inputs = tok(prompt, return_tensors="pt")
out = model.generate(**inputs, max_new_tokens=120, do_sample=True, top_p=0.95, temperature=0.8)
print(tok.decode(out[0], skip_special_tokens=True))

Training data

The corpus mixes two streams:

  • Real: English-language texts dated 1786โ€“1849. Primarily Gutenberg derived, with some news articles.
  • Synthetic: gated synthetic rewrites and augmentations.

The assembled training set is 363.7M tokens; the dev set is 5.8M tokens.

Training procedure

Training follows the BabyLlama-style ensemble distillation recipe:

  1. A single 32k BPE tokenizer was trained on the corpus.
  2. Two teacher models of the same architecture were trained independently on the corpus.
  3. The student (this model) was trained on the same corpus with an equal-weight blend of cross-entropy on the data and KL divergence to the averaged teacher logits.

Overparamaterization relative to data sizing is intentional; the lever is high weight decay

Hyperparameter Value
Epochs 4 (about 1.45B tokens seen)
Optimizer steps 22,196
Sequence length 512
Effective batch size 128
Learning rate 7e-4
Weight decay 5.0
Warmup steps 600
Distillation alpha / temperature 0.5 / 1.0
Precision bf16

Evaluation

Results on the held-out dev set at the end of training:

Model Cross-entropy loss Perplexity
Teacher 1 3.297 27.0
Teacher 2 3.295 27.0
Student (this model) 3.202 24.6

Filtered BLIMP: 0.72

Limitations and biases

  • The model reflects the language, knowledge and attitudes of its source period, including views that are offensive or false by present standards. It has no knowledge of anything after 1849 by design.

Citation

Currently, if you use this model, please cite the paper the training framework accompanies: Pretraining Language Models for Diachronic Linguistic Change Discovery. An additional preprint will come shortly.

Downloads last month
-
Safetensors
Model size
0.6B params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support