Instructions to use harsimran2004/whetstone-124m-base with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use harsimran2004/whetstone-124m-base with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="harsimran2004/whetstone-124m-base")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("harsimran2004/whetstone-124m-base") model = AutoModelForCausalLM.from_pretrained("harsimran2004/whetstone-124m-base", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use harsimran2004/whetstone-124m-base with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "harsimran2004/whetstone-124m-base" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "harsimran2004/whetstone-124m-base", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/harsimran2004/whetstone-124m-base
- SGLang
How to use harsimran2004/whetstone-124m-base with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "harsimran2004/whetstone-124m-base" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "harsimran2004/whetstone-124m-base", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "harsimran2004/whetstone-124m-base" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "harsimran2004/whetstone-124m-base", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use harsimran2004/whetstone-124m-base with Docker Model Runner:
docker model run hf.co/harsimran2004/whetstone-124m-base
Whetstone-124M-Base
A 124M-parameter base language model trained from scratch on 2.5 billion tokens (2.49B at this checkpoint) with Whetstone, a from-scratch training platform (data → tokenizer → training). It is a plain next-token predictor: not instruction-tuned, not safety-tuned, and it will continue text, not answer questions. It is a research and baseline model.
How it compares, in one line: on four zero-shot benchmarks it averages 39.4, against 40.3 for GPT-2 124M, using roughly a quarter of GPT-2's training tokens (and 1/70 to 1/240 of the tokens of OPT, GPT-Neo, Pythia and SmolLM). It matches that class on ARC and trails it by 1 to 3 points on HellaSwag and PIQA. See Evaluation.
Quick start
from transformers import AutoModelForCausalLM, AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("harsimran2004/whetstone-124m-base")
model = AutoModelForCausalLM.from_pretrained("harsimran2004/whetstone-124m-base")
inputs = tokenizer("The history of the printing press began", return_tensors="pt")
output = model.generate(**inputs, max_new_tokens=50, do_sample=True, temperature=0.7, top_k=40)
print(tokenizer.decode(output[0], skip_special_tokens=True))
Tested with transformers 5.19. The weights are a standard LlamaForCausalLM, so no trust_remote_code is needed. Notes:
- Documents were separated by
<eos>(id 32001) during training and no<bos>was ever used; the tokenizer does not add one. The benchmark numbers below were scored with<eos>prepended as the start-of-text token. - Special-token strings that appear inside training documents (for example
<pad>in a web page) were tokenized as ordinary text. To get that behaviour:AutoTokenizer.from_pretrained(..., split_special_tokens=True). - Sampling from a model this small drifts and repeats; expect to tune temperature,
top_kandrepetition_penalty. Example of what it does well and badly: it writes fluent, on-topic prose but gets facts wrong ("The history of the printing press began in the late 18th century. It emerged as a significant invention in the 21st century.").
Model
| Architecture | decoder-only Transformer, Llama-style (pre-norm RMSNorm, RoPE, SwiGLU, no biases) |
| Parameters | 123,694,080 (embeddings tied) |
| Layers / width / heads | 14 / 768 / 12 (head dim 64, multi-head attention) |
| MLP hidden size | 2048 (SwiGLU) |
| Context length | 2048 tokens |
| Positions | RoPE, theta 10,000 |
| Vocabulary | 32,007 = 32,000 byte-level BPE tokens + 7 special tokens (` < |
| Tokenizer | custom byte-level BPE trained on a 400 MB sample of the training mix (45% code, 25% math/technical, 30% prose); regex pre-tokenizer, NFC + LF normalisation |
| Weights | float32 (model.safetensors, 495 MB) |
The 14-layer shape (a GPT-2 shape has 12) is what brings the parameter count to 124M with a 32K vocabulary.
Training data
The released checkpoint (step 4,750) had seen 2,490,368,000 tokens; the run ended at step 4,769 (2,500,329,472 tokens). Every token was seen once (the prepared set held 3.22B training tokens). Documents are mixed by token share:
| Source (pinned revision) | Share | What was used |
|---|---|---|
FineWeb-Edu 87f09149 |
40% | sample/10BT, 2 shards |
GitHub code (clean) c48d40f9 |
20% | 13 shards, filtered to files whose license field is MIT, Apache-2.0, BSD-2/3, ISC, CC0 or Unlicense, and to Python, C, C++, Rust, Go, Shell, Julia, Makefile, CMake, Dockerfile, TeX |
Cosmopedia v2 3ba9d605 |
15% | 2 shards |
FineMath e92b25a6 (4+) |
10% | 2 shards |
OpenWebMath fde8ef8d |
5% | 2 shards |
peS2o 636a503e (v2) |
5% | 1 shard |
| ML-systems code (KernelBook) | 5% planned | not used: its license was not resolved, so the realised mix has no such data |
Preparation: documents shorter than 100 or longer than 1,000,000 characters dropped; exact deduplication (SHA-256 of normalised text); near-deduplication with MinHash (128 hashes, 5-token shingles where a token is a word, a whitespace run or a punctuation run, 16 bands of 8 rows) where a pair is a duplicate when its estimated Jaccard similarity reaches 0.8 (no exact re-check). Validation documents come from shards disjoint from training. No benchmark decontamination was performed.
Training
| Optimizer | AdamW, betas (0.9, 0.95), eps 1e-8, weight decay 0.1 |
| Learning rate | peak 6e-4, 200 warmup steps, cosine decay to 10% (6e-5) |
| Batch | 524,288 tokens per step (micro-batch 4 × 64 accumulation × 2048) |
| Steps | 4,769 |
| Gradient clipping | 1.0 |
| Precision | bf16 mixed precision |
| Seed | 1337 |
| Hardware | one NVIDIA L4 (24 GB), Google Cloud |
| Speed | 21,081 tokens/s, 24.9 s/step, 17.5% model-FLOPs utilisation, 33.5 hours |
| Selection | lowest loss on the 256-sequence validation subset: checkpoint step-00004750 (this release); the last checkpoints differ by less than 0.01 |
Exact configuration: training_config.json. Per-step history: training_metrics.json.
Held-out loss. On the entire held-out split (11,278 sequences, 23,086,066 tokens) this checkpoint's cross-entropy is 1.6629 nats per token (perplexity 5.27, standard error ±0.0037). The curve above was logged during training on a fixed 256-sequence subset of that split (524,288 tokens): it reads 1.6164 at the end (perplexity 5.03, standard error ±0.024), about 0.05 lower than the whole-split value, because those rows happen to be somewhat easier; compare checkpoints with it, not absolute levels. The subset loss fell at all 19 evaluations from 3.42 at step 250. Train loss and held-out loss stay within about 0.05 nats of each other, so there is no sign of overfitting.
Evaluation
Zero-shot, scoring the log-likelihood of each answer ending (the lm-evaluation-harness prompts and normalisation), on the full HellaSwag validation, ARC test and PIQA validation sets. Every model in the table was scored by the same code on the same machine; the scorer reproduces GPT-2's published numbers. Results: benchmark_scores.json.
| Model | tokens | HellaSwag (acc_norm) | ARC-Easy (acc) | ARC-Challenge (acc_norm) | PIQA (acc) | average |
|---|---|---|---|---|---|---|
| Whetstone-124M-Base | 2.5B | 29.3 | 44.0 | 24.3 | 59.8 | 39.4 |
| GPT-2 124M | ~10B | 31.4 | 43.9 | 22.8 | 63.1 | 40.3 |
| OPT-125M | ~180B | 31.8 | 43.1 | 22.5 | 62.9 | 40.1 |
| GPT-Neo-125M | ~300B | 30.4 | 43.8 | 23.2 | 63.2 | 40.1 |
| Pythia-160M | ~300B | 30.7 | 44.6 | 24.8 | 61.6 | 40.4 |
| SmolLM-135M | ~600B | 43.9 | 61.1 | 28.8 | 68.2 | 50.5 |
Standard errors are about 0.5 (HellaSwag), 1.0 to 1.3 (ARC) and 1.1 (PIQA) points; chance is 25 (HellaSwag, ARC) and 50 (PIQA). The token counts of the other models are their published training sizes, approximate and not verified here.
How to read it: the model is within about a point of the 125M class on average and equal to it on ARC, on far fewer tokens; the gaps are on HellaSwag and PIQA, the tests that reward having read more everyday text. SmolLM-135M, trained on roughly 240× more tokens, is about 11 points ahead of every model here.
Intended use and limitations
Intended for: studying small-model pretraining, a baseline for training or evaluation experiments, and as a starting point for fine-tuning.
Not intended for: answering questions, chat, or any use where wrong output matters. It has had no instruction tuning and no safety training.
Known limitations:
- It states falsehoods fluently. Facts, dates, arithmetic and code logic are frequently wrong.
- Scale. At 124M parameters and 2.5B tokens it is far below modern small models; see SmolLM-135M above.
- English and web-centric. Training data is English-language web text, math, scientific papers and code in a dozen languages; other languages are weak.
- No decontamination. Benchmark text may appear in the training data (FineWeb-Edu and others are web crawls); the scores could be slightly inflated, and this has not been measured.
- One run, one seed. Differences of one or two benchmark points are within noise.
- Licensing. The weights are released under Apache-2.0. The training sources are mostly ODC-By (attribution required: see the dataset links above); the code was filtered by each file's declared license field, which this release has not independently verified.
- Duplicates and memorisation. Near-duplicates were removed approximately (estimate-only MinHash at 0.8), so some duplicated text remains and may be reproduced verbatim.
Reproducibility and links
- Training logbook:
LOGBOOK.mdin this repository: what was run, what broke, what it cost, and the corrections made along the way. - Source code: the Whetstone platform is not public yet; the exact recipe is in
training_config.jsonand the data section above, and the pinned dataset revisions are linked. - Re-running the benchmarks:
lm_eval --model hf --model_args pretrained=harsimran2004/whetstone-124m-base --tasks hellaswag,arc_easy,arc_challenge,piqa --num_fewshot 0. The numbers above come from a small custom scorer that follows the same prompts; expect differences of a few tenths of a point from tokenisation and start-token details. - Export checks: the weights were converted from the training framework's format to
LlamaForCausalLMwith a test that the logits match, and the tokenizer was checked to give identical ids on 1,500 real documents.
Files
model.safetensors, config.json, generation_config.json, tokenizer.json, tokenizer_config.json, special_tokens_map.json, training_config.json, training_metrics.json, heldout_loss.json, benchmark_scores.json, figures/, LOGBOOK.md, LICENSE.
Citation
@misc{whetstone124m,
title = {Whetstone-124M-Base: a 124M-parameter language model trained from scratch on 2.5B tokens},
year = {2026},
note = {Trained with the Whetstone platform},
url = {https://huggingface.co/harsimran2004/whetstone-124m-base}
}
- Downloads last month
- 246
Datasets used to train harsimran2004/whetstone-124m-base
HuggingFaceTB/smollm-corpus
HuggingFaceTB/finemath
Evaluation results
- acc_norm (zero-shot) on HellaSwagvalidation set self-reported29.347
- acc (zero-shot) on ARC-Easytest set self-reported44.024
- acc_norm (zero-shot) on ARC-Challengetest set self-reported24.317
- acc (zero-shot) on PIQAvalidation set self-reported59.848

