pulvis-v1

A 2.96M parameter decoder-only language model pretrained from scratch on 20B tokens. On the Open SLM Leaderboard Index it scores 7.85, which is the highest score among models under 3M parameters. It is not yet listed on the board.

The name is Latin for dust, after the alchemists' pulvis projectionis, the pinch of powder that was supposed to transmute a whole crucible.

Results

Zero-shot, measured on the public weights in this repository with exactly the commands under Reproducing the evaluation: the lm-evaluation-harness hf backend and the leaderboard's official ArithMark-3 script, both in float32.

Benchmark Metric Score
HellaSwag acc_norm 27.91
ARC-Easy acc_norm 31.61
ARC-Challenge acc_norm 22.27
PIQA acc_norm 55.93
ArithMark-3 acc_norm 36.90
Open SLM Index 7.85
ArithMark-2 (not part of the Index) acc 48.04

The Index is the leaderboard's own formula, (N(HellaSwag,25) + N(CombinedARC,25) + N(PIQA,50) + 0.65*N(ArithMark,25)) / 3.65 where N(v,c) = 100(v-c)/(100-c) and CombinedARC is the mean of ARC-Easy and ARC-Challenge.

The top of the sub-3M field, using the leaderboard's published figures on 30 September 2026:

Model Params Index
pulvis-v1 2.96M 7.85
Ember-2 2.96M 7.21
ZeroS-Micro-v1.0 2.87M 6.49
Sol Nano 2.90M 6.07
BananaMind-2-Micro 2.93M 6.01
Ember 2.87M 5.94
CMA-1M-Mini 0.96M 5.78

Index against the sub-3M field

Per-benchmark comparison with Ember-2

How big the lead really is

At this size every benchmark sits only a few points above chance, and a single evaluation carries about 0.8 Index of sampling noise, most of it from PIQA's 1,838 questions. The 0.64 gap to Ember-2 is inside that noise, and it is worth saying where it comes from.

To get a steadier comparison, both models were scored on the training splits of the same benchmarks, which lm-evaluation-harness does not report and which neither model trained on: 39,905 HellaSwag items, 16,113 PIQA items, 4,239 ARC items, plus ArithMark-2's 2,500. Same code, same prompts, same scoring for both.

Model HellaSwag ARC-Easy ARC-Challenge PIQA ArithMark-2 Index over HS, ARC, PIQA
Ember-2 27.59 31.87 22.28 54.76 27.92 5.25
pulvis-v1 27.85 30.66 21.86 54.51 48.04 4.84

On general knowledge the two are level within noise, with Ember-2 slightly ahead. On arithmetic pulvis-v1 is far ahead: 48 against 28 on ArithMark-2, where chance is 25. Part of pulvis-v1's 7.85 is a favorable PIQA draw. It scores 55.93 on the PIQA test set and 54.51 on the much larger training split, and at 54.5 the Index would be about 7.1. Ember-2's ARC-Easy shows the same effect in the other direction, 33.42 on test against 31.87 on the training split.

What made the difference

The architecture work and the first full run got to 6.60, below Ember-2. The step that moved it past was the data in the final 25% of training.

During training

The run uses a warmup-stable-decay schedule. Through the stable phase the benchmark scores barely move, and HellaSwag in particular sits at about 27.6 from the first checkpoint to the last. Almost everything happens when the learning rate decays. The first full run spent its decay on a mathematics-heavy mixture drawn from FineMath and OpenMathInstruct-2 and reached 6.60 on the public evaluation.

The decay phase was then run again from the same pre-decay weights, with the mathematics share moved to grade-school word problems: the GSM8K-derived half of OpenMathInstruct-2 and Microsoft's Orca-Math. ArithMark-2 went from 32.6 to 47.9 on the training-time evaluation while the general benchmarks stayed level. On the public evaluation the released model scores 7.85.

Training loss

The loss in the second decay phase is lower mostly because worked arithmetic is easier to predict than web text. It is not a like-for-like comparison.

How the model was chosen

Everything that picked a direction was measured on held-out splits, not on the reported test sets, with two disclosed exceptions.

  1. Pilot ablations. About 30 runs of 0.5B to 3B tokens, most of them on one RTX 5080, compared tokenizers, widths, layer reuse, optimizers, attention variants and data mixtures, scored by validation bits per byte and the training-split benchmarks above. Run-to-run noise was measured with repeat seeds and turned out to be about 0.01 bits per byte, so several apparent wins were treated as ties.
  2. Mathematics mixture (first exception). ArithMark-2's bare expressions proved a poor stand-in for ArithMark-3's word problems, so the official ArithMark-3 script was run on ten pilot models to choose how much mathematics to use and when. This is use of a test set for a design decision.
  3. Decay-phase branches. Three decay phases were run from the same pre-decay weights: an exact repeat of the first run's decay as a control, a grade-school mathematics mixture, and a heavier version of it. The selection rule was written down before any result was seen: highest score on the training-split benchmarks plus ArithMark-2. The heavier mixture won, and only it was evaluated on the test sets.
  4. Test evaluations of release candidates (second exception). Two models were scored on the full public evaluation: the first full run (6.60) and this one (7.85). Both numbers are reported here.
Decay-phase branch ArithMark-2 Index over HS, ARC, PIQA
Control, the first run's decay repeated 31.36 5.15
Grade-school mathematics 43.80 5.07
Heavier grade-school mathematics (released) 47.88 4.90

These are training-time evaluations on the training splits. The control landed 0.5 below the first run on the general benchmarks with an identical recipe, which gives a sense of how much of this is noise.

No data was generated for this model, and no benchmark items or benchmark generators were used in training. OpenMathInstruct-2 and Orca-Math are themselves synthetic datasets, released by NVIDIA and Microsoft and built from GSM8K-style seed problems. They were checked against ArithMark-3: none of its 1,000 items shares a single 10-word sequence with any of the 200,035 Orca-Math problems, or with a sample of 160,904 of the GSM8K-derived OpenMathInstruct-2 rows (2 of its 32 files).

Findings from the ablations

At 3M parameters the budget is dominated by what the embedding table costs, and compute is cheap, so the choices differ from larger models.

Choice Result
Tokenizer A 4,096-token BPE with single-digit splitting tied with 2,048 and beat 1,024 and raw bytes on bits per byte. Bytes gave the best HellaSwag but lost elsewhere.
Layer reuse Running the middle blocks several times was the clearest architectural gain, worth about 0.02 bits per byte over no reuse at the same parameter count. Three passes were as good as four.
Width 192 beat 160 even though the embedding then takes 26.5% of the parameters. 224 was too wide.
XSA Removing each head's component along the token's own value vector helped by about 0.014 bits per byte.
Optimizer Muon matched tuned AdamW at best and trained 13% slower at this size. AdamW was kept.
Learning rate Anywhere from 4e-3 to 1.6e-2 was within noise. 8e-3 was used.
Data mixture No stable-phase mixture was distinguishable from the others. The decay-phase mixture was the only data choice that mattered.

Model details

Field Value
Parameters 2,962,240
Non-embedding parameters 73.5%
Hidden size 192
Blocks 1 prelude, 3 core blocks run 3 times, 3 coda
Blocks stored / applied 7 / 13
Intermediate size 368
Attention heads 6
Key/value heads 2
Head dimension 32
Attention Grouped query attention with QK-norm and XSA
Activation SwiGLU
Normalization RMSNorm, eps 1e-6
Positional encoding RoPE, theta 10,000
Context length 1,024
Vocabulary 4,096 byte-level BPE, digits split individually
Embeddings Tied input and output
Logit cap 15.0
Weights float32 safetensors

Each pass through the core blocks adds a small learned vector first, so the model can tell the passes apart. trust_remote_code=True is required because PulvisForCausalLM is not part of transformers.

Training data

Source Stable phase, 15B tokens Decay phase, 5B tokens
FinePhrase faq (rephrased FineWeb-Edu) 15% 10%
FinePhrase tutorial 15% 10%
FinePhrase math 8% 10%
FinePhrase table 4% 0%
FineWeb-Edu 18% 12%
DCLM-Baseline 22% 10%
FineMath 4+ 10% 10%
OpenMathInstruct-2, all sources 8% 0%
OpenMathInstruct-2, GSM8K-derived problems only 0% 32%
Orca-Math word problems 0% 6%

FinePhrase is used through its rephrased text, not the original web page it was written from. OpenMathInstruct-2 and Orca-Math are used as problem followed by worked solution.

Training setup

Field Value
Optimizer AdamW, betas 0.9 and 0.95, weight decay 0.1 on matrices
Peak learning rate 8e-3
Schedule Warmup-stable-decay
Stable phase 500 warmup steps, then constant to step 57,222 (15.0B tokens)
Decay phase Fresh optimizer state, 200 warmup steps, then linear decay to zero over 19,073 steps (5.0B tokens)
Tokens per step 262,144 (256 sequences of 1,024)
Total tokens 20.0B
Precision bfloat16 with float32 master weights
Hardware eight RTX 5090, data parallel, about 12M tokens per second
Wall clock about 25 minutes for the stable phase and 8 for the decay phase

Usage

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "bench-labs/pulvis-v1"
tok = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(model_id, trust_remote_code=True, dtype=torch.float32)

ids = tok("The capital of France is", return_tensors="pt")
out = model.generate(**ids, max_new_tokens=32, do_sample=False)
print(tok.decode(out[0], skip_special_tokens=True))

This is a base model with no instruction tuning and no chat template. It completes text.

Samples

Twelve fixed prompts, one sample each, unedited. Sampling at temperature 0.7, top-p 0.9, repetition penalty 1.1, 48 new tokens, seed 0 before every prompt. These are the same prompts used for tinctura-v1 and cagliostro-v3. At 3M parameters it strings together plausible phrasing, often with a mathematical flavor from the decay phase, and very little of it is true.

All 12 samples

1. The water cycle begins when

The water cycle begins when the water is moving at a constant rate.
Speed = 10 km/hr
Average speed of the water = Speed * Time = 10 * 30^(-6)

2. Photosynthesis is the process by which

Photosynthesis is the process by which your product will be produced in a single product. If you want to know how to choose from the manufacturer, use the formula for the product of two products, and use the formula for the product of the product of two products.

3. To solve 47 + 68, first

To solve 47 + 68, first finding the value of x which is 337. This means that 337 = 47 and 48 = 27 which implies that x = 90.
So the number of

4. The Roman Empire fell because

The Roman Empire fell because of the struggle between the Trojanes and the Apollo of Hague.

2. **Growth**
   - The Catholic Church of Tori took control of the Roman

5. A prime number is a number that

A prime number is a number that can be either 1 or 2.
Since 100 is an integer and 1 is 50, we can choose 50 that is an integer.
This means that the number 5

6. The three states of matter are

The three states of matter are:
1. Silver: 60% of the original material is being added.
2. Temperature: 20% of the original material is being added.
3. Magnetic Field

7. Gravity is the force that

Gravity is the force that can be applied to a person's body.

### 4. How does an intelligentness work?

The intelligentness of a person is a process where one's body is able to use a number of

8. In 1969, astronauts

In 1969, astronauts had discovered the oldest atom.
Seven years later, the gravitational field of star formed by the sun and a star formed by a new stars forming the same dichotomy

9. The heart pumps blood through

The heart pumps blood through the blood. This blood supply is a part of the body. This process can be transmitted to other body parts and can occur on other things.

**Step 6: Promote Your Career**

10. To find the area of a rectangle you

To find the area of a rectangle you can find by multiplying the length of each side of the rectangle by their width and then multiplying by their respective lengths.

The area of one rectangle is given by the product of

11. Volcanoes form when

Volcanoes form when there is a shaded or southeast it is likely to have an event that occurs in a large area of 1000 feet.

What is the significance of this documentary on the historical significance of

12. The difference between weather and climate is

The difference between weather and climate is 53.5%.
Since there are 400000 people in the city, and 40% of these people have a dry weather, the number of people who have a dry weather is

Reproducing the evaluation

pip install lm-eval
python -m lm_eval --model hf \
  --model_args pretrained=bench-labs/pulvis-v1,dtype=float32,trust_remote_code=True \
  --tasks hellaswag,arc_easy,arc_challenge,piqa \
  --num_fewshot 0 --batch_size 64 --device cuda:0

ArithMark-3 uses the leaderboard's official script from AxiomicLabs/ArithMark-3.0. The script defaults to bfloat16, so pass --dtype float32:

python bencharithmark-3.py --model bench-labs/pulvis-v1 --data-path arithmark-3.jsonl --device cuda --dtype float32

ArithMark-2 uses the official script from AxiomicLabs/ArithMark-2.0. It runs in bfloat16 whenever a GPU is visible and has no dtype option, so the number above comes from a CPU run, which is float32:

CUDA_VISIBLE_DEVICES="" python benchmark_arithmark-2.0.py --model bench-labs/pulvis-v1 --data-path arithmark_2.0.jsonl

Evaluate in float32 throughout. The model uses a logit cap of 15, and bfloat16 shifts the ArithMark scores of logit-capped models. The exported weights agree with the training checkpoint exactly: the largest logit difference between the two is zero.

Provenance

checkpoints/ holds the stable-phase weights every 2.5B tokens, including the weights the released decay phase started from (step 57,222, three steps after the first run began its own decay), and the end of the first full run that scored 6.60. Each subfolder loads with subfolder= in from_pretrained.

Limitations

English only. 1,024 token context. No instruction tuning, no safety tuning, no RLHF. At 3M parameters it confabulates freely, as the samples show, and should not be relied on for anything factual. Its arithmetic ability is two-digit arithmetic in short word problems, not mathematical reasoning.

License

Apache-2.0. The training data is drawn from FinePhrase, FineWeb-Edu and FineMath (ODC-By), DCLM-Baseline and OpenMathInstruct-2 (CC-BY-4.0), and Orca-Math (MIT).

Downloads last month
-
Safetensors
Model size
2.96M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train bench-labs/pulvis-v1