pulvis-v2

A 2.96M parameter decoder-only language model pretrained from scratch on 20B tokens. On the Open SLM Leaderboard Index it scores 8.62, up from 7.85 for pulvis-v1. Among models under 3M parameters that is second, behind Tokle-3M at 8.91. It is not yet listed on the board.

The main change from v1 is the shape of the network: ten stored blocks run as sixteen, with three of them looped. It uses no benchmark training data of any kind.

Results

Zero-shot, measured on the public weights in this repository with exactly the commands under Reproducing the evaluation: the lm-evaluation-harness hf backend and the leaderboard's official ArithMark-3 script, both in float32.

Benchmark Metric pulvis-v2 pulvis-v1
HellaSwag acc_norm 27.93 27.91
ARC-Easy acc_norm 31.86 31.61
ARC-Challenge acc_norm 22.78 22.27
PIQA acc_norm 56.86 55.93
ArithMark-3 acc_norm 37.40 36.90
Open SLM Index 8.62 7.85

The Index is the leaderboard's own formula, (N(HellaSwag,25) + N(CombinedARC,25) + N(PIQA,50) + 0.65*N(ArithMark,25)) / 3.65 where N(v,c) = 100(v-c)/(100-c) and CombinedARC is the mean of ARC-Easy and ARC-Challenge.

The top of the sub-3M field, using the leaderboard's published figures on 3 October 2026:

Model Params Index
Tokle-3M 2.91M 8.91
pulvis-v2 2.96M 8.62
pulvis-v1 2.96M 7.85
Ember-2 2.96M 7.21
ZeroS-Micro-v1.0 2.87M 6.49
Sol Nano 2.90M 6.07
BananaMind-2-Micro 2.93M 6.01

Index against the sub-3M field

Per-benchmark comparison with pulvis-v1

How much of the gain is real

A single evaluation at this size carries about 0.8 Index of sampling noise, most of it from PIQA's 1,838 questions, so a 0.77 gain on the test sets alone would not settle much. Both models were also scored on held-out questions that neither trained on, with the same code and scoring for both: 4,037 HellaSwag and 1,587 PIQA items held out from the training splits, the 869 ARC validation questions, and the 2,097 integer-answer word problems in the ASDiv validation set.

Model HellaSwag PIQA ARC-Easy ARC-Challenge Index over HS, ARC, PIQA ASDiv word problems
pulvis-v1 27.62 54.19 28.42 21.74 3.99 35.00
pulvis-v2 28.24 59.61 30.70 23.75 8.83 42.30

On held-out questions v2 is ahead on every benchmark, and by more than on the test sets: 8.83 against 3.99 on the knowledge index and 42.3 against 35.0 on word problems. PIQA is the one result worth flagging. v2 gains 5.4 points on the held-out PIQA training questions but only 0.9 on the PIQA test set. The decay-phase tutorial data was checked for overlap with both PIQA splits and overlaps them less than FineWeb-Edu does (0.47% of PIQA training items against 0.63%), so this does not look like contamination, but the gap between the two splits is unexplained.

What changed from v1

The shape of the network

pulvis-v1 stores seven blocks and applies thirteen: one block in, three core blocks run three times, three blocks out, at a width of 192. Running the core several times buys depth without spending parameters, and at 3M parameters that depth is most of what makes the model know things.

It also turned out to be what held the model back on arithmetic word problems. Every model in this family knows its times tables: on bare a x b = questions every shape below scores between 92 and 98 percent. The looped model loses ground on ASDiv's word problems, where the numbers sit inside a sentence and have to be found and combined, which is also the ArithMark-3 format.

To pin this down, six models of looped, plain and hybrid shapes were trained on the same data mixture with the same schedule and parameter budget, then scored on the held-out questions above.

Knowledge against word problems

Shape Width Blocks stored / applied Index over HS, ARC, PIQA ASDiv word problems Bare a x b =
Looped, the v1 shape 192 7 / 13 8.74 38.87 92.6
Plain 192 7 / 7 7.22 39.77 93.4
Plain 144 9 / 9 6.90 41.92 97.5
Plain 160 10 / 10 6.82 42.49 93.4
Hybrid 160 10 / 14 6.30 41.06 97.5
Hybrid, pulvis-v2 160 10 / 16 8.83 42.30 96.7

One training run per shape. The ARC part of the knowledge index rests on 869 validation questions, so differences of a few tenths are within noise.

Word problems track the number of distinct blocks: nine or ten read numbers well, seven do not, looped or not. Knowledge favors models that apply many blocks at a reasonable width, and the looped v1 shape and pulvis-v2 score highest on it. At 3M parameters a plain stack can only reach ten distinct blocks by getting narrower, and it pays for that in knowledge. pulvis-v2 keeps ten distinct blocks at width 160 and loops three of them three times, so it applies sixteen. A version that applied fourteen did not show the knowledge gain. With one run per shape, that comparison is the least certain part of the table.

Tokenizer

The 4,096-entry BPE vocabulary of v1 splits every digit. v2 drops its 100 rarest merges and adds the 100 two-digit strings 00 to 99, matched left to right, so most numbers in grade-school arithmetic become one or two tokens. The vocabulary size is unchanged.

Data

The stable phase adds Cosmopedia, MegaScience and Cosmopedia's web samples written as WikiHow-style tutorials to v1's mixture, and brings GSM8K-derived word problems in from the start instead of only at the end. The first 172 steps train on balanced bracket sequences instead of text, a warm-up borrowed from the formal-language pretraining literature. The decay phase gives 40% of its tokens to Cosmopedia's WikiHow-style tutorials, which lifted PIQA by about two points on the held-out set in every run that used it.

Training loss

How the model was chosen

Everything that picked a direction was measured on held-out data, not on the reported test sets, with two disclosed exceptions.

  1. Held-out selection. Candidates were compared on held-out slices of the HellaSwag, PIQA and ARC training splits plus ArithMark-2 at first, and later ASDiv in place of ArithMark-2, once it was clear that ArithMark-2's bare expressions could not see the word-problem gap. Selection rules and the score a candidate had to beat were written down before each round's results.
  2. Tokenizer pilots (first exception). The official ArithMark-3 script was run on fourteen 1.5B-token pilots to choose between digit, two-digit and three-digit number tokens. This is use of a test set for a design decision.
  3. Test evaluations of release candidates (second exception). Five candidates on the way to v2 were scored on the full public evaluation. In order: 6.59, 7.20, 7.90, 8.07 and 8.62. Each was evaluated only after winning its round on held-out data. The best of five noisy measurements is biased upward, which is one more reason to read the held-out table above alongside the headline number.

No benchmark items, benchmark training splits or benchmark generators were used in training. The two WikiHow-style sources were checked against the evaluation sets with a 10-word overlap test. Cosmopedia's WikiHow subset overlapped no more than general web text did. A second WikiHow-format source drawn from Cosmopedia-v2 overlapped the HellaSwag validation set at 7.25% of items against 5.38% for FineWeb-Edu, and was excluded. ASDiv was used only to choose between candidates and never trained on. The bracket sequences in the warm-up are the only generated data and contain no language.

Model details

Field Value
Parameters 2,963,840
Hidden size 160
Blocks 2 in, 3 core blocks run 3 times, 5 out
Blocks stored / applied 10 / 16
Intermediate size 352
Attention heads 5
Key/value heads 1
Head dimension 32
Attention Multi-query attention with QK-norm and XSA
Activation SwiGLU
Normalization RMSNorm, eps 1e-6
Positional encoding RoPE, theta 10,000
Context length 1,024
Vocabulary 4,096 byte-level BPE with two-digit number tokens
Embeddings Tied input and output
Logit cap 15.0
Weights float32 safetensors

Each pass through the core blocks adds a small learned vector first, so the model can tell the passes apart. trust_remote_code=True is required because PulvisForCausalLM is not part of transformers. The modeling code is the v1 code with one change: models without looping skip the loop vector.

Training data

Source Stable phase, 15B tokens Decay phase, 5B tokens
FinePhrase faq (rephrased FineWeb-Edu) 7% 4%
FinePhrase tutorial 7% 2%
FinePhrase math 5% 4%
FineWeb-Edu 20% 12%
DCLM-Baseline 3% 0%
FineMath 4+ 7% 4%
Cosmopedia-v2 (SmolLM corpus) 10% 3%
Cosmopedia science (OpenStax, Khan Academy, Stanford) 6% 4%
MegaScience 3% 2%
OpenMathInstruct-2, all sources 10% 0%
OpenMathInstruct-2, GSM8K-derived problems only 7% 20%
Orca-Math word problems 0% 5%
Cosmopedia web samples in WikiHow format 15% 0%
Cosmopedia WikiHow tutorials 0% 40%

The first 172 steps of the stable phase train only on balanced bracket sequences over four bracket types, nested up to 16 deep.

FinePhrase is used through its rephrased text. OpenMathInstruct-2 and Orca-Math are used as problem followed by worked solution. MegaScience is used as question followed by answer.

Training setup

Field Value
Optimizer AdamW, betas 0.9 and 0.95, weight decay 0.1 on matrices
Peak learning rate 8e-3
Schedule Warmup-stable-decay
Warm-up data Balanced bracket sequences for the first 172 steps (about 45M tokens)
Stable phase 500 warmup steps, then constant to step 57,220 (15.0B tokens)
Decay phase Fresh optimizer state, 200 warmup steps, then linear decay to zero over 19,073 steps (5.0B tokens)
Tokens per step 262,144 (256 sequences of 1,024)
Total tokens 20.0B
Precision bfloat16 with float32 master weights
Hardware two RTX 5090, data parallel, about 3.2M tokens per second
Wall clock about 77 minutes for the stable phase and 26 for the decay phase

Usage

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "bench-labs/pulvis-v2"
tok = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(model_id, trust_remote_code=True, dtype=torch.float32)

ids = tok("The capital of France is", return_tensors="pt")
out = model.generate(**ids, max_new_tokens=32, do_sample=False)
print(tok.decode(out[0], skip_special_tokens=True))

This is a base model with no instruction tuning and no chat template. It completes text.

Samples

The same twelve prompts as pulvis-v1, one sample each, unedited. Sampling at temperature 0.7, top-p 0.9, repetition penalty 1.1, 48 new tokens, seed 0 before every prompt. At 3M parameters it strings together plausible phrasing and very little of it is true. The arithmetic in the third sample is a good illustration.

All 12 samples

1. The water cycle begins when

The water cycle begins when the water will be dripping up. If there are 20000 gallons of water in a jar, how many jars of water would it take for the water cycle to be dripped?

2. Photosynthesis is the process by which

Photosynthesis is the process by which photographs are made in a single image to represent each other.
The first step to creating the image is to use the first step in the photography process of photographs and their form.
To create the image, we

3. To solve 47 + 68, first

To solve 47 + 68, first find the value of x:

x = 47 + 68
x = 110

Therefore, 110 + 68 = 186.

Answer: \boxed{186

4. The Roman Empire fell because

The Roman Empire fell because of its murder in the southern and southern parts of the world, the Roman Empire fell on the west. The Civil War was the last of the years that the Turkey

5. A prime number is a number that

A prime number is a number that can be expressed as a sum of numbers 3, 4, 6, or 9.

2. **Expand and Divide:**
   - For the largest prime, the sum of a number is

6. The three states of matter are

The three states of matter are 100% and 40%.
Since 20000 is 30000, we can find the amount of air that fires the remaining particles by subtracting the amount of air that f

7. Gravity is the force that

Gravity is the force that can be defined as the sum of the masses of a mass and the mass of its object. The mass of a mass is calculated as follows:

P_squared_A_s^2 =

8. In 1969, astronauts

In 1969, astronauts came to the Northwest at a higher speile of 2000 astronauts.
The first time they were to visit the Northwest, the astronauts were traveling on the southern side of

9. The heart pumps blood through

The heart pumps blood through the blood, which helps in maintaining blood flow and supports its function.
3. *Avoid Turning Out*: Breathing deeply into the blood can help regulate blood flow, which makes it

10. To find the area of a rectangle you

To find the area of a rectangle you can complete in 5 hours, we can divide the area of the rectangle by the number of hours it takes to complete one rectangle.

Area = (Sea area / Time) *

11. Volcanoes form when

Volcanoes form when they are not moving. If the southern and north-center southern regions are in a south-center region, the southern region is the largest. The southern southern regions of

12. The difference between weather and climate is

The difference between weather and climate is 100%. If the temperature of the water is 80 degrees Fahrenheit and the temperature of the atmosphere is 30 degrees Fahrenheit, calculate the total temperature of the rainwater that is fl

Reproducing the evaluation

pip install lm-eval
python -m lm_eval --model hf \
  --model_args pretrained=bench-labs/pulvis-v2,dtype=float32,trust_remote_code=True \
  --tasks hellaswag,arc_easy,arc_challenge,piqa \
  --num_fewshot 0 --batch_size 64 --device cuda:0

ArithMark-3 uses the leaderboard's official script from AxiomicLabs/ArithMark-3.0. The script defaults to bfloat16, so pass --dtype float32:

python bencharithmark-3.py --model bench-labs/pulvis-v2 --data-path arithmark-3.jsonl --device cuda --dtype float32

Evaluate in float32 throughout. The model uses a logit cap of 15, and bfloat16 shifts the ArithMark scores of logit-capped models. The exported weights agree with the training checkpoint exactly: the largest logit difference between the two is zero.

Provenance

checkpoints/step_057220 holds the weights at the end of the stable phase, the point the decay phase started from. It loads with subfolder="checkpoints/step_057220" in from_pretrained.

Limitations

English only. 1,024 token context. No instruction tuning, no safety tuning, no RLHF. At 3M parameters it confabulates freely, as the samples show, and should not be relied on for anything factual. Its arithmetic ability is one- and two-step grade-school word problems with small numbers, not mathematical reasoning.

License

CC-BY-NC-SA-4.0. The training data includes MegaScience, which is released under CC-BY-NC-SA-4.0, so the weights carry the same terms. The rest of the data is drawn from FinePhrase, FineWeb-Edu, FineMath and the SmolLM corpus (ODC-By), Cosmopedia (Apache-2.0), DCLM-Baseline and OpenMathInstruct-2 (CC-BY-4.0), and Orca-Math (MIT).

Downloads last month
-
Safetensors
Model size
2.96M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train bench-labs/pulvis-v2