word

A GPT-2 language model trained from random initialisation on WikiText-2, rather than adapted from OpenAI's released weights. The architecture is GPT-2 small โ€” twelve layers, 768 embedding dimensions, 50257-token vocabulary โ€” which works out at about 124M parameters; the previous version of this card put the figure at 117M, which does not match the 496 MB float32 checkpoint. config.json carries no _name_or_path, confirming nothing was fine-tuned.

The tokenizer is GPT-2's own byte-pair vocabulary, taken from openai-community/gpt2, because the checkpoint was trained from scratch and therefore ships no vocabulary of its own. It matches: both are 50257 entries, so the token ids in the weights line up with this tokenizer exactly. An earlier version of this repository shipped no tokenizer at all, and AutoTokenizer then returned an empty vocabulary of length one rather than raising, which meant every input collapsed to the unknown token and generation returned noise without any error. That is fixed. GPT-2 uses byte-pair encoding without a distinct unknown token, so rare inputs are segmented rather than discarded โ€” one reason a model of this size produces fluent-looking nonsense.

Usage

from transformers import pipeline

generator = pipeline("text-generation", model="harpertoken/word")
print(generator("The quick brown fox", max_new_tokens=60)[0]["generated_text"])

Limitations

WikiText-2 is about two million tokens drawn from Wikipedia verification articles. Training on it for a few epochs produces a model that reproduces that register โ€” encyclopaedic, declarative English โ€” and little else. The previous card reported three epochs at batch size one completed in roughly ten minutes on CPU or MPS; for a 124M-parameter model over that corpus at that batch size the arithmetic does not work, so I have left the training time unstated rather than repeat it.

There are no evaluation figures. Perplexity on WikiText-2 would be the obvious measurement and has not been recorded here. Treat this as an architectural and training demonstration: it is a working GPT-2 implementation at small scale, not a language model to build on. It has no instruction tuning, no dialogue behaviour, and no pretraining beyond a narrow slice of one corpus.

Attribution

The architecture follows Radford et al., Improving Language Understanding by Generative Pre-Training (2018). WikiText-2 is described in Merity et al., Pointer Sentinel Mixture Models (2016).

Downloads last month
11
Safetensors
Model size
0.1B params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support