Instructions to use lukeslp/drummer2-540m with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use lukeslp/drummer2-540m with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="lukeslp/drummer2-540m")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("lukeslp/drummer2-540m") model = AutoModelForCausalLM.from_pretrained("lukeslp/drummer2-540m", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use lukeslp/drummer2-540m with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "lukeslp/drummer2-540m" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "lukeslp/drummer2-540m", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/lukeslp/drummer2-540m
- SGLang
How to use lukeslp/drummer2-540m with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "lukeslp/drummer2-540m" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "lukeslp/drummer2-540m", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "lukeslp/drummer2-540m" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "lukeslp/drummer2-540m", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use lukeslp/drummer2-540m with Docker Model Runner:
docker model run hf.co/lukeslp/drummer2-540m
Drummer-540M
Drummer-540M is a 542,310,720-parameter decoder-only language model I trained from scratch. This repository contains the continued Foundation 0.12.0 checkpoint at 15.85B training tokens. It is a research project, not a product.
Architecture
- Llama-style decoder: 24 layers, hidden size 1,344, feed-forward size 3,584, 21 attention heads and 21 key/value heads, head size 64, RMSNorm, SiLU, rotary positions (theta 10,000), tied embeddings.
- Vocabulary 16,384, byte-level BPE; context length 2,048 tokens.
- 542,310,720 parameters. The weights in this repository are stored in float32 (2.17 GB); the public demo serves a bfloat16 copy of the same checkpoint.
- Weights SHA-256:
84d5b2284b983533fc126c4d3799fa1659f73a0b741c3a3796d5f6414585bc7d.
Training
The base run used 10,846,208,000 tokens in 82,750 updates of 131,072 tokens, seed 20260913. Its mix was FineWeb-Edu 50%, Cosmopedia v2 15%, Wikipedia 15%, Project Gutenberg 10%, and dialogue 10%. The base run used warmup-stable-decay with peak learning rate 5e-4 and a final 10%-of-peak floor.
The continuation carried the model to 15.85B tokens using new FineWeb-Edu, DCLM-Edu and FineMath shards plus replay of the base sources. Relative to Foundation 0.11.0, held-out loss improved 3.46% on the old distribution and 3.43% on the new distribution; replay loss improved 2.08%. BLiMP rose from 0.7435 to 0.8133. The eight-task zero-shot lm-eval mean rose from 0.4698 to 0.4821; the worst single-task change was -0.0016. The preregistered publication gate passed every clause.
About 8% of the base tokens were smol-smoltalk, text generated by Llama-3.1-405B; that model's license includes a naming clause. The continuation used no Llama-derived text.
Evaluation
All rows below were measured by me with lm-evaluation-harness 0.4.13, zero-shot, float32, context 2,048 and seed 20260915.
| Model | Params | Train tokens | Eight-task mean | ARC-e | ARC-c | BoolQ | PIQA | SIQA | HellaSwag | OBQA | WinoGrande |
|---|---|---|---|---|---|---|---|---|---|---|---|
| SmolLM2-360M | 362M | 4T | .543 | .680 | .384 | .617 | .719 | .408 | .564 | .382 | .590 |
| Qwen2.5-0.5B | 494M | 18T | .514 | .587 | .324 | .624 | .700 | .442 | .521 | .352 | .564 |
| Drummer-540M 0.12.0 | 542M | 15.85B | .482 | .554 | .290 | .565 | .685 | .411 | .467 | .352 | .533 |
| Drummer-540M 0.11.0 | 542M | 10.85B | .470 | .528 | .281 | .540 | .687 | .412 | .440 | .348 | .522 |
| Pythia-410M | 405M | 300B | .450 | .457 | .243 | .606 | .672 | .390 | .406 | .294 | .533 |
| GPT-2 medium | 355M | ~10B | .444 | .436 | .250 | .586 | .664 | .391 | .394 | .302 | .531 |
MMLU, 57 subjects, zero-shot, example-weighted aggregate: Foundation 0.11.0 at 10.85B tokens scored 0.257, chance level. The released Foundation 0.12.0 checkpoint at 15.85B tokens has no MMLU run on record.
BLiMP, 67 paradigms x 1,000 minimal pairs: Foundation 0.12.0 0.8133; Foundation 0.11.0 0.7435.
IFEval on Chat 0.18.0, a separate fine-tune of the 8.13B-token checkpoint, not this one: 30 of 541 prompts satisfied every instruction under prompt-level strict accuracy. This is included to show how far this model family remains from reliable novel-instruction following.
Intended use
Research on small-model pretraining, controlled continuation and how limited training budgets affect language-model behavior. The model can be used for completion experiments and as a base for narrow fine-tunes.
Limitations and safety
- No safety tuning. Do not rely on this model for safety-critical, medical, legal, financial or factual decisions.
- It invents facts, produces biased or offensive text inherited from its training data, and contradicts itself across turns.
- The 10.85B-token Foundation 0.11.0 parent scored 0.257 on MMLU, at chance level; the released 15.85B-token checkpoint has not been run on MMLU. General knowledge and reasoning remain weak.
- Context is limited to 2,048 tokens.
- It does not use tools that are only declared in a prompt: a tools fine-tune of this checkpoint made 0 of 24 valid calls to unseen tools, as the one from 10.85B did. Instruction following was not measured on it; the Chat fine-tune of the 8.13B checkpoint scored 30 of 541 on IFEval.
- English is the primary language represented and evaluated.
Licenses and attribution
Weights and code: Apache-2.0. This card and project-authored documentation: CC BY 4.0. Training sources retain their own terms. Base run: FineWeb-Edu and Cosmopedia v2 (ODC-By 1.0), Wikipedia (CC BY-SA 4.0; the Hub dataset wikimedia/wikipedia is tagged CC BY-SA 3.0/GFDL), Project Gutenberg via common-pile/project_gutenberg_filtered (public domain in the US), SODA (CC BY 4.0; GPT-3.5 output), and smol-smoltalk (Apache-2.0 data; Llama 3.1 naming clause on the generating model). Continuation: DCLM-Edu (CC BY 4.0), FineMath (ODC-By 1.0), and a small human-transcript share from OANC Switchboard (unrestricted), QuAC (MIT) and Taskmaster (CC BY 4.0).
Evaluation used lm-evaluation-harness (MIT), BLiMP (CC BY 4.0), IFEval (Apache-2.0) and MMLU (MIT); scores are reported, no benchmark data is redistributed here.
Author and citation
Luke Steuber - lukesteuber.com - GitHub
Also by me: Chime-SMS-2M, a 2M-parameter next-word model for phone keyboards, with a demo that runs in your browser.
Steuber, L. (2026). Drummer-540M: a small language model trained from scratch, with matched studies of context. https://drummer-demo.dr.eamer.dev
- Downloads last month
- 1,012
