ZeroS-Micro-v1.0

ZeroS-Micro-v1.0 is a 3M-parameter SLM trained on 2.95B tokens of a high-quality dataset mix, featuring a custom architecture designed for parameter efficiency and representational quality at sub-10M scale.

Originally, this was supposed to be math-focused model called ZeroS-Micro-Math, but it started doing well on general language benchmarks during pretraining, so we've adopted it as our general micro model.

Architecture

The architecture of ZeroS-Micro-v1.0 is built around an interleaved feed-forward design that packs 18 layers into just 3 million parameters. Rather than placing dense feed-forward blocks at every layer, ZeroS-Micro-v1.0 alternates between parameter-free Hadamard FFNs and SwiGLU FFNs. This allows the model to achieve greater depth while bounding the total parameter count.

  • Hidden Size: 144
  • Vocab Size: 2564
  • Number of Layers: 18 (9 SwiGLU, 9 Hadamard)
  • Intermediate Size (SwiGLU FFN Capacity): 384
  • Number of Attention Heads: 3
  • Number of KV Heads: 1 (Grouped Query Attention / GQA)
  • Dimensions Per Head: 48 (144 / 3)
  • XSA Attention: true
  • RoPE Theta: 7500.0
  • Sequence Length: 1536 (can generate past this at inference)
  • Tied Word Embeddings: true
  • Total Parameters: 2,869,813 (2.87M)

Training Dataset

ZeroS-Micro-v1.0 was trained on 2.95 billion tokens consisting of high-density mathematical reasoning, formal logic, code, and academic knowledge.

Dataset Share Domain
FineMath (HuggingFaceTB/finemath) 19.4% Mathematical reasoning, educational math
The Stack v3 (HuggingFaceCode/stack-v3-train) 18.4% code
FinePhrase (HuggingFaceFW/finephrase) 16.5% High-quality synthetic math
Common-Pile: DOAB Filtered (common-pile/doab_filtered) 13.6% Academic textbooks
Algebraic Stack (typeof/algebraic-stack) 8.7% Formal mathematics & symbolic code
OpenWebMath (open-web-math/open-web-math) 7.8% Web-extracted mathematical text & LaTeX
Common-Pile: ArXiv Filtered (common-pile/arxiv_papers_filtered) 7.8% STEM & mathematics
UltraData-Math (openbmb/UltraData-Math) 7.8% Curated math problem solving & proofs

Benchmark Results

Task Score
HellaSwag 28.46%
ARC-Easy 31.27%
ARC-Challenge 22.01%
PIQA 53.86%
ArithMark-3 35.60%
Average 34.24%

License

Apache 2.0.

Citation

@misc{zeros-micro-v1.0,
  title        = {ZeroS-Micro-v1.0},
  organization = {FromZero},
  authors      = {Paul Courneya},
  year         = {2026},
  url          = {https://huggingface.co/fromziro/ZeroS-Micro-v1.0}
}
Downloads last month
63
Safetensors
Model size
3.24M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train fromziro/ZeroS-Micro-v1.0