Instructions to use bench-labs/pulvis-v2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use bench-labs/pulvis-v2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="bench-labs/pulvis-v2", trust_remote_code=True)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("bench-labs/pulvis-v2", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use bench-labs/pulvis-v2 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "bench-labs/pulvis-v2" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bench-labs/pulvis-v2", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/bench-labs/pulvis-v2
- SGLang
How to use bench-labs/pulvis-v2 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "bench-labs/pulvis-v2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bench-labs/pulvis-v2", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "bench-labs/pulvis-v2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bench-labs/pulvis-v2", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use bench-labs/pulvis-v2 with Docker Model Runner:
docker model run hf.co/bench-labs/pulvis-v2
pulvis-v2
A 2.96M parameter decoder-only language model pretrained from scratch on 20B tokens. On the Open SLM Leaderboard Index it scores 8.62, up from 7.85 for pulvis-v1. Among models under 3M parameters that is second, behind Tokle-3M at 8.91. It is not yet listed on the board.
The main change from v1 is the shape of the network: ten stored blocks run as sixteen, with three of them looped. It uses no benchmark training data of any kind.
Results
Zero-shot, measured on the public weights in this repository with exactly the commands under Reproducing the evaluation: the lm-evaluation-harness hf backend and the leaderboard's official ArithMark-3 script, both in float32.
| Benchmark | Metric | pulvis-v2 | pulvis-v1 |
|---|---|---|---|
| HellaSwag | acc_norm | 27.93 | 27.91 |
| ARC-Easy | acc_norm | 31.86 | 31.61 |
| ARC-Challenge | acc_norm | 22.78 | 22.27 |
| PIQA | acc_norm | 56.86 | 55.93 |
| ArithMark-3 | acc_norm | 37.40 | 36.90 |
| Open SLM Index | 8.62 | 7.85 |
The Index is the leaderboard's own formula, (N(HellaSwag,25) + N(CombinedARC,25) + N(PIQA,50) + 0.65*N(ArithMark,25)) / 3.65 where N(v,c) = 100(v-c)/(100-c) and CombinedARC is the mean of ARC-Easy and ARC-Challenge.
The top of the sub-3M field, using the leaderboard's published figures on 3 October 2026:
| Model | Params | Index |
|---|---|---|
| Tokle-3M | 2.91M | 8.91 |
| pulvis-v2 | 2.96M | 8.62 |
| pulvis-v1 | 2.96M | 7.85 |
| Ember-2 | 2.96M | 7.21 |
| ZeroS-Micro-v1.0 | 2.87M | 6.49 |
| Sol Nano | 2.90M | 6.07 |
| BananaMind-2-Micro | 2.93M | 6.01 |
How much of the gain is real
A single evaluation at this size carries about 0.8 Index of sampling noise, most of it from PIQA's 1,838 questions, so a 0.77 gain on the test sets alone would not settle much. Both models were also scored on held-out questions that neither trained on, with the same code and scoring for both: 4,037 HellaSwag and 1,587 PIQA items held out from the training splits, the 869 ARC validation questions, and the 2,097 integer-answer word problems in the ASDiv validation set.
| Model | HellaSwag | PIQA | ARC-Easy | ARC-Challenge | Index over HS, ARC, PIQA | ASDiv word problems |
|---|---|---|---|---|---|---|
| pulvis-v1 | 27.62 | 54.19 | 28.42 | 21.74 | 3.99 | 35.00 |
| pulvis-v2 | 28.24 | 59.61 | 30.70 | 23.75 | 8.83 | 42.30 |
On held-out questions v2 is ahead on every benchmark, and by more than on the test sets: 8.83 against 3.99 on the knowledge index and 42.3 against 35.0 on word problems. PIQA is the one result worth flagging. v2 gains 5.4 points on the held-out PIQA training questions but only 0.9 on the PIQA test set. The decay-phase tutorial data was checked for overlap with both PIQA splits and overlaps them less than FineWeb-Edu does (0.47% of PIQA training items against 0.63%), so this does not look like contamination, but the gap between the two splits is unexplained.
What changed from v1
The shape of the network
pulvis-v1 stores seven blocks and applies thirteen: one block in, three core blocks run three times, three blocks out, at a width of 192. Running the core several times buys depth without spending parameters, and at 3M parameters that depth is most of what makes the model know things.
It also turned out to be what held the model back on arithmetic word problems. Every model in this family knows its times tables: on bare a x b = questions every shape below scores between 92 and 98 percent. The looped model loses ground on ASDiv's word problems, where the numbers sit inside a sentence and have to be found and combined, which is also the ArithMark-3 format.
To pin this down, six models of looped, plain and hybrid shapes were trained on the same data mixture with the same schedule and parameter budget, then scored on the held-out questions above.
| Shape | Width | Blocks stored / applied | Index over HS, ARC, PIQA | ASDiv word problems | Bare a x b = |
|---|---|---|---|---|---|
| Looped, the v1 shape | 192 | 7 / 13 | 8.74 | 38.87 | 92.6 |
| Plain | 192 | 7 / 7 | 7.22 | 39.77 | 93.4 |
| Plain | 144 | 9 / 9 | 6.90 | 41.92 | 97.5 |
| Plain | 160 | 10 / 10 | 6.82 | 42.49 | 93.4 |
| Hybrid | 160 | 10 / 14 | 6.30 | 41.06 | 97.5 |
| Hybrid, pulvis-v2 | 160 | 10 / 16 | 8.83 | 42.30 | 96.7 |
One training run per shape. The ARC part of the knowledge index rests on 869 validation questions, so differences of a few tenths are within noise.
Word problems track the number of distinct blocks: nine or ten read numbers well, seven do not, looped or not. Knowledge favors models that apply many blocks at a reasonable width, and the looped v1 shape and pulvis-v2 score highest on it. At 3M parameters a plain stack can only reach ten distinct blocks by getting narrower, and it pays for that in knowledge. pulvis-v2 keeps ten distinct blocks at width 160 and loops three of them three times, so it applies sixteen. A version that applied fourteen did not show the knowledge gain. With one run per shape, that comparison is the least certain part of the table.
Tokenizer
The 4,096-entry BPE vocabulary of v1 splits every digit. v2 drops its 100 rarest merges and adds the 100 two-digit strings 00 to 99, matched left to right, so most numbers in grade-school arithmetic become one or two tokens. The vocabulary size is unchanged.
Data
The stable phase adds Cosmopedia, MegaScience and Cosmopedia's web samples written as WikiHow-style tutorials to v1's mixture, and brings GSM8K-derived word problems in from the start instead of only at the end. The first 172 steps train on balanced bracket sequences instead of text, a warm-up borrowed from the formal-language pretraining literature. The decay phase gives 40% of its tokens to Cosmopedia's WikiHow-style tutorials, which lifted PIQA by about two points on the held-out set in every run that used it.
How the model was chosen
Everything that picked a direction was measured on held-out data, not on the reported test sets, with two disclosed exceptions.
- Held-out selection. Candidates were compared on held-out slices of the HellaSwag, PIQA and ARC training splits plus ArithMark-2 at first, and later ASDiv in place of ArithMark-2, once it was clear that ArithMark-2's bare expressions could not see the word-problem gap. Selection rules and the score a candidate had to beat were written down before each round's results.
- Tokenizer pilots (first exception). The official ArithMark-3 script was run on fourteen 1.5B-token pilots to choose between digit, two-digit and three-digit number tokens. This is use of a test set for a design decision.
- Test evaluations of release candidates (second exception). Five candidates on the way to v2 were scored on the full public evaluation. In order: 6.59, 7.20, 7.90, 8.07 and 8.62. Each was evaluated only after winning its round on held-out data. The best of five noisy measurements is biased upward, which is one more reason to read the held-out table above alongside the headline number.
No benchmark items, benchmark training splits or benchmark generators were used in training. The two WikiHow-style sources were checked against the evaluation sets with a 10-word overlap test. Cosmopedia's WikiHow subset overlapped no more than general web text did. A second WikiHow-format source drawn from Cosmopedia-v2 overlapped the HellaSwag validation set at 7.25% of items against 5.38% for FineWeb-Edu, and was excluded. ASDiv was used only to choose between candidates and never trained on. The bracket sequences in the warm-up are the only generated data and contain no language.
Model details
| Field | Value |
|---|---|
| Parameters | 2,963,840 |
| Hidden size | 160 |
| Blocks | 2 in, 3 core blocks run 3 times, 5 out |
| Blocks stored / applied | 10 / 16 |
| Intermediate size | 352 |
| Attention heads | 5 |
| Key/value heads | 1 |
| Head dimension | 32 |
| Attention | Multi-query attention with QK-norm and XSA |
| Activation | SwiGLU |
| Normalization | RMSNorm, eps 1e-6 |
| Positional encoding | RoPE, theta 10,000 |
| Context length | 1,024 |
| Vocabulary | 4,096 byte-level BPE with two-digit number tokens |
| Embeddings | Tied input and output |
| Logit cap | 15.0 |
| Weights | float32 safetensors |
Each pass through the core blocks adds a small learned vector first, so the model can tell the passes apart. trust_remote_code=True is required because PulvisForCausalLM is not part of transformers. The modeling code is the v1 code with one change: models without looping skip the loop vector.
Training data
| Source | Stable phase, 15B tokens | Decay phase, 5B tokens |
|---|---|---|
| FinePhrase faq (rephrased FineWeb-Edu) | 7% | 4% |
| FinePhrase tutorial | 7% | 2% |
| FinePhrase math | 5% | 4% |
| FineWeb-Edu | 20% | 12% |
| DCLM-Baseline | 3% | 0% |
| FineMath 4+ | 7% | 4% |
| Cosmopedia-v2 (SmolLM corpus) | 10% | 3% |
| Cosmopedia science (OpenStax, Khan Academy, Stanford) | 6% | 4% |
| MegaScience | 3% | 2% |
| OpenMathInstruct-2, all sources | 10% | 0% |
| OpenMathInstruct-2, GSM8K-derived problems only | 7% | 20% |
| Orca-Math word problems | 0% | 5% |
| Cosmopedia web samples in WikiHow format | 15% | 0% |
| Cosmopedia WikiHow tutorials | 0% | 40% |
The first 172 steps of the stable phase train only on balanced bracket sequences over four bracket types, nested up to 16 deep.
FinePhrase is used through its rephrased text. OpenMathInstruct-2 and Orca-Math are used as problem followed by worked solution. MegaScience is used as question followed by answer.
Training setup
| Field | Value |
|---|---|
| Optimizer | AdamW, betas 0.9 and 0.95, weight decay 0.1 on matrices |
| Peak learning rate | 8e-3 |
| Schedule | Warmup-stable-decay |
| Warm-up data | Balanced bracket sequences for the first 172 steps (about 45M tokens) |
| Stable phase | 500 warmup steps, then constant to step 57,220 (15.0B tokens) |
| Decay phase | Fresh optimizer state, 200 warmup steps, then linear decay to zero over 19,073 steps (5.0B tokens) |
| Tokens per step | 262,144 (256 sequences of 1,024) |
| Total tokens | 20.0B |
| Precision | bfloat16 with float32 master weights |
| Hardware | two RTX 5090, data parallel, about 3.2M tokens per second |
| Wall clock | about 77 minutes for the stable phase and 26 for the decay phase |
Usage
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "bench-labs/pulvis-v2"
tok = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(model_id, trust_remote_code=True, dtype=torch.float32)
ids = tok("The capital of France is", return_tensors="pt")
out = model.generate(**ids, max_new_tokens=32, do_sample=False)
print(tok.decode(out[0], skip_special_tokens=True))
This is a base model with no instruction tuning and no chat template. It completes text.
Samples
The same twelve prompts as pulvis-v1, one sample each, unedited. Sampling at temperature 0.7, top-p 0.9, repetition penalty 1.1, 48 new tokens, seed 0 before every prompt. At 3M parameters it strings together plausible phrasing and very little of it is true. The arithmetic in the third sample is a good illustration.
All 12 samples
1. The water cycle begins when
The water cycle begins when the water will be dripping up. If there are 20000 gallons of water in a jar, how many jars of water would it take for the water cycle to be dripped?
2. Photosynthesis is the process by which
Photosynthesis is the process by which photographs are made in a single image to represent each other.
The first step to creating the image is to use the first step in the photography process of photographs and their form.
To create the image, we
3. To solve 47 + 68, first
To solve 47 + 68, first find the value of x:
x = 47 + 68
x = 110
Therefore, 110 + 68 = 186.
Answer: \boxed{186
4. The Roman Empire fell because
The Roman Empire fell because of its murder in the southern and southern parts of the world, the Roman Empire fell on the west. The Civil War was the last of the years that the Turkey
5. A prime number is a number that
A prime number is a number that can be expressed as a sum of numbers 3, 4, 6, or 9.
2. **Expand and Divide:**
- For the largest prime, the sum of a number is
6. The three states of matter are
The three states of matter are 100% and 40%.
Since 20000 is 30000, we can find the amount of air that fires the remaining particles by subtracting the amount of air that f
7. Gravity is the force that
Gravity is the force that can be defined as the sum of the masses of a mass and the mass of its object. The mass of a mass is calculated as follows:
P_squared_A_s^2 =
8. In 1969, astronauts
In 1969, astronauts came to the Northwest at a higher speile of 2000 astronauts.
The first time they were to visit the Northwest, the astronauts were traveling on the southern side of
9. The heart pumps blood through
The heart pumps blood through the blood, which helps in maintaining blood flow and supports its function.
3. *Avoid Turning Out*: Breathing deeply into the blood can help regulate blood flow, which makes it
10. To find the area of a rectangle you
To find the area of a rectangle you can complete in 5 hours, we can divide the area of the rectangle by the number of hours it takes to complete one rectangle.
Area = (Sea area / Time) *
11. Volcanoes form when
Volcanoes form when they are not moving. If the southern and north-center southern regions are in a south-center region, the southern region is the largest. The southern southern regions of
12. The difference between weather and climate is
The difference between weather and climate is 100%. If the temperature of the water is 80 degrees Fahrenheit and the temperature of the atmosphere is 30 degrees Fahrenheit, calculate the total temperature of the rainwater that is fl
Reproducing the evaluation
pip install lm-eval
python -m lm_eval --model hf \
--model_args pretrained=bench-labs/pulvis-v2,dtype=float32,trust_remote_code=True \
--tasks hellaswag,arc_easy,arc_challenge,piqa \
--num_fewshot 0 --batch_size 64 --device cuda:0
ArithMark-3 uses the leaderboard's official script from AxiomicLabs/ArithMark-3.0. The script defaults to bfloat16, so pass --dtype float32:
python bencharithmark-3.py --model bench-labs/pulvis-v2 --data-path arithmark-3.jsonl --device cuda --dtype float32
Evaluate in float32 throughout. The model uses a logit cap of 15, and bfloat16 shifts the ArithMark scores of logit-capped models. The exported weights agree with the training checkpoint exactly: the largest logit difference between the two is zero.
Provenance
checkpoints/step_057220 holds the weights at the end of the stable phase, the point the decay phase started from. It loads with subfolder="checkpoints/step_057220" in from_pretrained.
Limitations
English only. 1,024 token context. No instruction tuning, no safety tuning, no RLHF. At 3M parameters it confabulates freely, as the samples show, and should not be relied on for anything factual. Its arithmetic ability is one- and two-step grade-school word problems with small numbers, not mathematical reasoning.
License
CC-BY-NC-SA-4.0. The training data includes MegaScience, which is released under CC-BY-NC-SA-4.0, so the weights carry the same terms. The rest of the data is drawn from FinePhrase, FineWeb-Edu, FineMath and the SmolLM corpus (ODC-By), Cosmopedia (Apache-2.0), DCLM-Baseline and OpenMathInstruct-2 (CC-BY-4.0), and Orca-Math (MIT).
- Downloads last month
- -



