min-spark 1.5

min-spark 1.5

min-spark 1.5 is a 9.36M-parameter language model with native effort levels. It reaches an Intelligence Index of 9.79, higher than any sub-10M model currently on the Open SLM leaderboard, after training on just 2.0B tokens.

What's new in 1.5

Much better results for far less training. min-spark 1.5 beats min-spark 1.1 while training on 8× fewer tokens with 5.7× less compute. It also beats Hummingbird-V2, the strongest other sub-10M model on the leaderboard, which trained on 10B tokens, five times as many.

min-spark 1.1 min-spark 1.5
Int Index 9.44 9.79
Training tokens 16.81B 2.0B (8× fewer)
Training compute 2.2 × 10¹⁸ FLOPs 3.9 × 10¹⁷ FLOPs (5.7× less)
Parameters 5.76M 9.36M

min-spark 1.5 Int Index against training tokens

min-spark runs a few shared transformer blocks several times in a row. Each pass has its own small adapter, so it can do something different from the pass before. In earlier versions those adapters were held back. Version 1.5 fixes that:

  • Adapters on every layer. Every pass now adapts all four projections of each shared block instead of only the attention input, at rank 24 instead of 16.
  • Stable adapters. Adapters read normalized inputs and train at a steadier learning rate, so none of them can grow out of control and shut off part of a block.
  • No wasted passes. Earlier versions trained a fourth pass that the model learned to ignore. Version 1.5 trains three, and all three do real work.
  • A bigger core and new data. There are five shared blocks instead of three. Cosmopedia v2 textbooks and a STEM web crawl make up a fifth of the training data.

Generation also gets faster: 1.5 adds a KV cache, which makes generation about 3× faster on CPU.

Effort levels

Effort levels set how many passes the model makes through its shared core. Large reasoning models usually do this by changing a thinking-token budget. min-spark changes the computation inside the network instead.

Effort Passes Use
low 2 Fastest completion
medium 3 (default) General use; best on most tasks
high 4 Experimental: one pass more than the model was trained with. It still gives the best HellaSwag score

min-spark 1.5 accuracy by effort

Effort ARC-Easy ARC-Challenge HellaSwag PIQA ArithMark-3 Int Index
min-spark 1.5 low 37.21% 23.55% 28.11% 56.69% 34.30% 8.98
min-spark 1.5 medium 38.64% 22.70% 28.15% 57.34% 34.90% 9.59
min-spark 1.5 high 36.99% 23.21% 28.27% 56.58% 34.00% 8.80
best per task 38.64% 23.55% 28.27% 57.34% 34.90% 9.79

Evaluation

All scores are zero-shot and length-normalized. ARC, HellaSwag and PIQA use lm-eval 0.4.12. ArithMark-3 uses the official ArithMark-3.0 script. As on earlier min-spark cards, each benchmark is reported at its best effort level.

The Intelligence Index is the Open SLM leaderboard's summary score. Each benchmark is rescaled so that random guessing scores 0 and a perfect score is 100, then averaged, with ARC as the mean of Easy and Challenge:

Int Index = (HellaSwag + ARC + PIQA + 0.65 × ArithMark-3) / 3.65

The table below compares 1.5 with the leading sub-10M models on the leaderboard that have a public model card. Peer scores come from the leaderboard and training tokens from each model's card.

min-spark 1.5 compared with sub-10M peers

Model Params Training tokens ARC-Easy ARC-Challenge HellaSwag PIQA ArithMark-3 Int Index
min-spark 1.5 9.36M 2.0B 38.64% 23.55% 28.27% 57.34% 34.90% 9.79
Hummingbird-V2 9.59M 10B 39.65% 21.16% 27.61% 57.40% 36.00% 9.59
min-spark 1.1 5.76M 16.81B 36.45% 23.04% 28.24% 57.40% 35.40% 9.44
ForgePlex-M2-9M 9.95M 30B 36.53% 23.46% 28.02% 57.18% 34.60% 9.14
Spark-2A 9.15M 34.26% 22.18% 28.13% 57.24% 37.00% 9.14

Usage

min-spark 1.5 loads through Transformers with remote code:

from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained(
    "MinimaLabs/min-spark-1.5",
    trust_remote_code=True,
).to("cuda")
tokenizer = AutoTokenizer.from_pretrained(
    "MinimaLabs/min-spark-1.5",
    trust_remote_code=True,
)

inputs = tokenizer("The meaning of life is", return_tensors="pt").to("cuda")
outputs = model.generate(**inputs, effort="medium", max_new_tokens=64)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))

effort works the same way with the pipeline API. To skip Transformers, use the bundled script:

python generate.py -p "The meaning of life is" -e medium

Generation runs one sequence at a time and uses the KV cache by default. Right-padded batches work for scoring.

Architecture

Field Value
Parameters 9,360,596
Architecture Meiosis (tied-embedding looped decoder)
Vocabulary 4,096-token byte-level BPE
Embedding width 288
Heads 6 query · 2 KV (GQA)
FFN hidden size 832
Blocks 1 prelude · 5 shared body blocks · 1 coda · final RMSNorm
Per-pass adapters LoRA rank 24 on attention QKV, attention output, FFN gate/up and FFN down
Context window 512 tokens
Effort (loop count) low = 2 · medium = 3 · high = 4 (trained up to 3)

Training

Field Value
Training tokens 2.0B, from scratch
Precision bf16 autocast
Global batch size 128 sequences × 512 tokens
Optimizer Muon (matrices) + NAdamW (other parameters), weight decay 0.01
Learning-rate schedule 2,000-step warmup, stable phase, linear decay to 10% over the final 20%
Objective Next-token loss at 3 passes, plus 0.1 × the loss at 2 passes and 0.1 × a term that keeps the 2-pass state close to the 3-pass one
Attention masking Intra-document
Data source Share
FineWeb-Edu (score 4+) ~63%
Cosmopedia v2 (synthetic textbooks), new in 1.5 15%
DCLM-baseline filtered for commonsense, plus WikiHow ~11%
FineMath (finemath-4plus) ~5%
Dolmino STEM crawl, new in 1.5 4%
Synthetic bracket-matching sequences, at the start of training ~1%

Both new sources were decontaminated: any document sharing a 13-word sequence with an item from the five evaluation sets was dropped.

Reproducing the evaluation

python run_lmeval.py --effort low,medium,high
python bencharithmark-3.py --model MinimaLabs/min-spark-1.5 --dtype float32

The first command scores ARC, HellaSwag and PIQA at each effort and prints each task's best (add --limit N for a quick run). The second is the official ArithMark-3 script, which runs at the default medium effort.

Limitations

This is a small base model. It has not been instruction-tuned, so it has no conversational behavior, alignment or safety filtering, and its knowledge and reasoning are limited by its size. The context window is 512 tokens; all results were measured within it.

License

Apache-2.0. See LICENSE.

Downloads last month
-
Safetensors
Model size
9.36M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train MinimaLabs/min-spark-1.5