Whetstone-124M-Base

A 124M-parameter base language model trained from scratch on 2.5 billion tokens (2.49B at this checkpoint) with Whetstone, a from-scratch training platform (data → tokenizer → training). It is a plain next-token predictor: not instruction-tuned, not safety-tuned, and it will continue text, not answer questions. It is a research and baseline model.

How it compares, in one line: on four zero-shot benchmarks it averages 39.4, against 40.3 for GPT-2 124M, using roughly a quarter of GPT-2's training tokens (and 1/70 to 1/240 of the tokens of OPT, GPT-Neo, Pythia and SmolLM). It matches that class on ARC and trails it by 1 to 3 points on HellaSwag and PIQA. See Evaluation.

Quick start

from transformers import AutoModelForCausalLM, AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("harsimran2004/whetstone-124m-base")
model = AutoModelForCausalLM.from_pretrained("harsimran2004/whetstone-124m-base")

inputs = tokenizer("The history of the printing press began", return_tensors="pt")
output = model.generate(**inputs, max_new_tokens=50, do_sample=True, temperature=0.7, top_k=40)
print(tokenizer.decode(output[0], skip_special_tokens=True))

Tested with transformers 5.19. The weights are a standard LlamaForCausalLM, so no trust_remote_code is needed. Notes:

  • Documents were separated by <eos> (id 32001) during training and no <bos> was ever used; the tokenizer does not add one. The benchmark numbers below were scored with <eos> prepended as the start-of-text token.
  • Special-token strings that appear inside training documents (for example <pad> in a web page) were tokenized as ordinary text. To get that behaviour: AutoTokenizer.from_pretrained(..., split_special_tokens=True).
  • Sampling from a model this small drifts and repeats; expect to tune temperature, top_k and repetition_penalty. Example of what it does well and badly: it writes fluent, on-topic prose but gets facts wrong ("The history of the printing press began in the late 18th century. It emerged as a significant invention in the 21st century.").

Model

Architecture decoder-only Transformer, Llama-style (pre-norm RMSNorm, RoPE, SwiGLU, no biases)
Parameters 123,694,080 (embeddings tied)
Layers / width / heads 14 / 768 / 12 (head dim 64, multi-head attention)
MLP hidden size 2048 (SwiGLU)
Context length 2048 tokens
Positions RoPE, theta 10,000
Vocabulary 32,007 = 32,000 byte-level BPE tokens + 7 special tokens (` <
Tokenizer custom byte-level BPE trained on a 400 MB sample of the training mix (45% code, 25% math/technical, 30% prose); regex pre-tokenizer, NFC + LF normalisation
Weights float32 (model.safetensors, 495 MB)

The 14-layer shape (a GPT-2 shape has 12) is what brings the parameter count to 124M with a 32K vocabulary.

Training data

The released checkpoint (step 4,750) had seen 2,490,368,000 tokens; the run ended at step 4,769 (2,500,329,472 tokens). Every token was seen once (the prepared set held 3.22B training tokens). Documents are mixed by token share:

Source (pinned revision) Share What was used
FineWeb-Edu 87f09149 40% sample/10BT, 2 shards
GitHub code (clean) c48d40f9 20% 13 shards, filtered to files whose license field is MIT, Apache-2.0, BSD-2/3, ISC, CC0 or Unlicense, and to Python, C, C++, Rust, Go, Shell, Julia, Makefile, CMake, Dockerfile, TeX
Cosmopedia v2 3ba9d605 15% 2 shards
FineMath e92b25a6 (4+) 10% 2 shards
OpenWebMath fde8ef8d 5% 2 shards
peS2o 636a503e (v2) 5% 1 shard
ML-systems code (KernelBook) 5% planned not used: its license was not resolved, so the realised mix has no such data

Preparation: documents shorter than 100 or longer than 1,000,000 characters dropped; exact deduplication (SHA-256 of normalised text); near-deduplication with MinHash (128 hashes, 5-token shingles where a token is a word, a whitespace run or a punctuation run, 16 bands of 8 rows) where a pair is a duplicate when its estimated Jaccard similarity reaches 0.8 (no exact re-check). Validation documents come from shards disjoint from training. No benchmark decontamination was performed.

Training

Optimizer AdamW, betas (0.9, 0.95), eps 1e-8, weight decay 0.1
Learning rate peak 6e-4, 200 warmup steps, cosine decay to 10% (6e-5)
Batch 524,288 tokens per step (micro-batch 4 × 64 accumulation × 2048)
Steps 4,769
Gradient clipping 1.0
Precision bf16 mixed precision
Seed 1337
Hardware one NVIDIA L4 (24 GB), Google Cloud
Speed 21,081 tokens/s, 24.9 s/step, 17.5% model-FLOPs utilisation, 33.5 hours
Selection lowest loss on the 256-sequence validation subset: checkpoint step-00004750 (this release); the last checkpoints differ by less than 0.01

Exact configuration: training_config.json. Per-step history: training_metrics.json.

Training and validation loss

Held-out loss. On the entire held-out split (11,278 sequences, 23,086,066 tokens) this checkpoint's cross-entropy is 1.6629 nats per token (perplexity 5.27, standard error ±0.0037). The curve above was logged during training on a fixed 256-sequence subset of that split (524,288 tokens): it reads 1.6164 at the end (perplexity 5.03, standard error ±0.024), about 0.05 lower than the whole-split value, because those rows happen to be somewhat easier; compare checkpoints with it, not absolute levels. The subset loss fell at all 19 evaluations from 3.42 at step 250. Train loss and held-out loss stay within about 0.05 nats of each other, so there is no sign of overfitting.

Evaluation

Zero-shot, scoring the log-likelihood of each answer ending (the lm-evaluation-harness prompts and normalisation), on the full HellaSwag validation, ARC test and PIQA validation sets. Every model in the table was scored by the same code on the same machine; the scorer reproduces GPT-2's published numbers. Results: benchmark_scores.json.

Model tokens HellaSwag (acc_norm) ARC-Easy (acc) ARC-Challenge (acc_norm) PIQA (acc) average
Whetstone-124M-Base 2.5B 29.3 44.0 24.3 59.8 39.4
GPT-2 124M ~10B 31.4 43.9 22.8 63.1 40.3
OPT-125M ~180B 31.8 43.1 22.5 62.9 40.1
GPT-Neo-125M ~300B 30.4 43.8 23.2 63.2 40.1
Pythia-160M ~300B 30.7 44.6 24.8 61.6 40.4
SmolLM-135M ~600B 43.9 61.1 28.8 68.2 50.5

Benchmarks

Standard errors are about 0.5 (HellaSwag), 1.0 to 1.3 (ARC) and 1.1 (PIQA) points; chance is 25 (HellaSwag, ARC) and 50 (PIQA). The token counts of the other models are their published training sizes, approximate and not verified here.

How to read it: the model is within about a point of the 125M class on average and equal to it on ARC, on far fewer tokens; the gaps are on HellaSwag and PIQA, the tests that reward having read more everyday text. SmolLM-135M, trained on roughly 240× more tokens, is about 11 points ahead of every model here.

Intended use and limitations

Intended for: studying small-model pretraining, a baseline for training or evaluation experiments, and as a starting point for fine-tuning.

Not intended for: answering questions, chat, or any use where wrong output matters. It has had no instruction tuning and no safety training.

Known limitations:

  • It states falsehoods fluently. Facts, dates, arithmetic and code logic are frequently wrong.
  • Scale. At 124M parameters and 2.5B tokens it is far below modern small models; see SmolLM-135M above.
  • English and web-centric. Training data is English-language web text, math, scientific papers and code in a dozen languages; other languages are weak.
  • No decontamination. Benchmark text may appear in the training data (FineWeb-Edu and others are web crawls); the scores could be slightly inflated, and this has not been measured.
  • One run, one seed. Differences of one or two benchmark points are within noise.
  • Licensing. The weights are released under Apache-2.0. The training sources are mostly ODC-By (attribution required: see the dataset links above); the code was filtered by each file's declared license field, which this release has not independently verified.
  • Duplicates and memorisation. Near-duplicates were removed approximately (estimate-only MinHash at 0.8), so some duplicated text remains and may be reproduced verbatim.

Reproducibility and links

  • Training logbook: LOGBOOK.md in this repository: what was run, what broke, what it cost, and the corrections made along the way.
  • Source code: the Whetstone platform is not public yet; the exact recipe is in training_config.json and the data section above, and the pinned dataset revisions are linked.
  • Re-running the benchmarks: lm_eval --model hf --model_args pretrained=harsimran2004/whetstone-124m-base --tasks hellaswag,arc_easy,arc_challenge,piqa --num_fewshot 0. The numbers above come from a small custom scorer that follows the same prompts; expect differences of a few tenths of a point from tokenisation and start-token details.
  • Export checks: the weights were converted from the training framework's format to LlamaForCausalLM with a test that the logits match, and the tokenizer was checked to give identical ids on 1,500 real documents.

Files

model.safetensors, config.json, generation_config.json, tokenizer.json, tokenizer_config.json, special_tokens_map.json, training_config.json, training_metrics.json, heldout_loss.json, benchmark_scores.json, figures/, LOGBOOK.md, LICENSE.

Citation

@misc{whetstone124m,
  title  = {Whetstone-124M-Base: a 124M-parameter language model trained from scratch on 2.5B tokens},
  year   = {2026},
  note   = {Trained with the Whetstone platform},
  url    = {https://huggingface.co/harsimran2004/whetstone-124m-base}
}
Downloads last month
246
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train harsimran2004/whetstone-124m-base

Evaluation results