Drummer-540M

Drummer-540M is a 542,310,720-parameter decoder-only language model I trained from scratch. This repository contains the continued Foundation 0.12.0 checkpoint at 15.85B training tokens. It is a research project, not a product.

Architecture

  • Llama-style decoder: 24 layers, hidden size 1,344, feed-forward size 3,584, 21 attention heads and 21 key/value heads, head size 64, RMSNorm, SiLU, rotary positions (theta 10,000), tied embeddings.
  • Vocabulary 16,384, byte-level BPE; context length 2,048 tokens.
  • 542,310,720 parameters. The weights in this repository are stored in float32 (2.17 GB); the public demo serves a bfloat16 copy of the same checkpoint.
  • Weights SHA-256: 84d5b2284b983533fc126c4d3799fa1659f73a0b741c3a3796d5f6414585bc7d.

Training

The base run used 10,846,208,000 tokens in 82,750 updates of 131,072 tokens, seed 20260913. Its mix was FineWeb-Edu 50%, Cosmopedia v2 15%, Wikipedia 15%, Project Gutenberg 10%, and dialogue 10%. The base run used warmup-stable-decay with peak learning rate 5e-4 and a final 10%-of-peak floor.

The continuation carried the model to 15.85B tokens using new FineWeb-Edu, DCLM-Edu and FineMath shards plus replay of the base sources. Relative to Foundation 0.11.0, held-out loss improved 3.46% on the old distribution and 3.43% on the new distribution; replay loss improved 2.08%. BLiMP rose from 0.7435 to 0.8133. The eight-task zero-shot lm-eval mean rose from 0.4698 to 0.4821; the worst single-task change was -0.0016. The preregistered publication gate passed every clause.

About 8% of the base tokens were smol-smoltalk, text generated by Llama-3.1-405B; that model's license includes a naming clause. The continuation used no Llama-derived text.

Evaluation

Benchmark comparison

All rows below were measured by me with lm-evaluation-harness 0.4.13, zero-shot, float32, context 2,048 and seed 20260915.

Model Params Train tokens Eight-task mean ARC-e ARC-c BoolQ PIQA SIQA HellaSwag OBQA WinoGrande
SmolLM2-360M 362M 4T .543 .680 .384 .617 .719 .408 .564 .382 .590
Qwen2.5-0.5B 494M 18T .514 .587 .324 .624 .700 .442 .521 .352 .564
Drummer-540M 0.12.0 542M 15.85B .482 .554 .290 .565 .685 .411 .467 .352 .533
Drummer-540M 0.11.0 542M 10.85B .470 .528 .281 .540 .687 .412 .440 .348 .522
Pythia-410M 405M 300B .450 .457 .243 .606 .672 .390 .406 .294 .533
GPT-2 medium 355M ~10B .444 .436 .250 .586 .664 .391 .394 .302 .531

MMLU, 57 subjects, zero-shot, example-weighted aggregate: Foundation 0.11.0 at 10.85B tokens scored 0.257, chance level. The released Foundation 0.12.0 checkpoint at 15.85B tokens has no MMLU run on record.

BLiMP, 67 paradigms x 1,000 minimal pairs: Foundation 0.12.0 0.8133; Foundation 0.11.0 0.7435.

IFEval on Chat 0.18.0, a separate fine-tune of the 8.13B-token checkpoint, not this one: 30 of 541 prompts satisfied every instruction under prompt-level strict accuracy. This is included to show how far this model family remains from reliable novel-instruction following.

Intended use

Research on small-model pretraining, controlled continuation and how limited training budgets affect language-model behavior. The model can be used for completion experiments and as a base for narrow fine-tunes.

Limitations and safety

  • No safety tuning. Do not rely on this model for safety-critical, medical, legal, financial or factual decisions.
  • It invents facts, produces biased or offensive text inherited from its training data, and contradicts itself across turns.
  • The 10.85B-token Foundation 0.11.0 parent scored 0.257 on MMLU, at chance level; the released 15.85B-token checkpoint has not been run on MMLU. General knowledge and reasoning remain weak.
  • Context is limited to 2,048 tokens.
  • It does not use tools that are only declared in a prompt: a tools fine-tune of this checkpoint made 0 of 24 valid calls to unseen tools, as the one from 10.85B did. Instruction following was not measured on it; the Chat fine-tune of the 8.13B checkpoint scored 30 of 541 on IFEval.
  • English is the primary language represented and evaluated.

Licenses and attribution

Weights and code: Apache-2.0. This card and project-authored documentation: CC BY 4.0. Training sources retain their own terms. Base run: FineWeb-Edu and Cosmopedia v2 (ODC-By 1.0), Wikipedia (CC BY-SA 4.0; the Hub dataset wikimedia/wikipedia is tagged CC BY-SA 3.0/GFDL), Project Gutenberg via common-pile/project_gutenberg_filtered (public domain in the US), SODA (CC BY 4.0; GPT-3.5 output), and smol-smoltalk (Apache-2.0 data; Llama 3.1 naming clause on the generating model). Continuation: DCLM-Edu (CC BY 4.0), FineMath (ODC-By 1.0), and a small human-transcript share from OANC Switchboard (unrestricted), QuAC (MIT) and Taskmaster (CC BY 4.0).

Evaluation used lm-evaluation-harness (MIT), BLiMP (CC BY 4.0), IFEval (Apache-2.0) and MMLU (MIT); scores are reported, no benchmark data is redistributed here.

Author and citation

Luke Steuber - lukesteuber.com - GitHub

Also by me: Chime-SMS-2M, a 2M-parameter next-word model for phone keyboards, with a demo that runs in your browser.

Steuber, L. (2026). Drummer-540M: a small language model trained from scratch, with matched studies of context. https://drummer-demo.dr.eamer.dev

Downloads last month
1,012
Safetensors
Model size
0.5B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support