Sol Nano
Sol Nano has 2,895,188 parameters, and its training run used 5 billion tokens. It uses the Sol Lite SolForCausalLM implementation, with TN-Gram local memory and a 1,024-entry tokenizer.
This is a base model for text completion. It hasn't been instruction-tuned, so the benchmark results below shouldn't be read as a measure of how well it works as a chat assistant.
Architecture
| Setting | Value |
|---|---|
| Parameters | 2,895,188 |
| TN-Gram parameters | 209,748 |
| Hidden width | 128 |
| Context | 512 tokens |
| Vocabulary | 1,024 |
| Stored blocks / effective applications | 10 / 14 |
| Query heads / KV heads | 4 / 2 |
| FFN width | 536 |
| Training tokens | 5,000,000,000 |
| Optimizer updates | 19,074 |
| Weights | FP32 safetensors |
The transformer uses causal grouped-query attention with RoPE and QK normalization. It reuses blocks with loop conditioning, and the output head shares the token embeddings. TN-Gram stores factorized local patterns for orders 2-5.
Run it
The included implementation needs CUDA, Triton, and FlexAttention support. Alongside PyTorch, install huggingface_hub, tokenizers, and safetensors. This example loads the released weights and predicts one token:
import os
import sys
from pathlib import Path
import torch
from huggingface_hub import snapshot_download
from safetensors.torch import load_file
from tokenizers import Tokenizer
os.environ["SOL_NANO_ATTENTION"] = "triton"
model_dir = Path(snapshot_download("solintellegence/sol-nano"))
sys.path.insert(0, str(model_dir))
from modeling_sol_lite import SolForCausalLM, variant_config
model = SolForCausalLM(variant_config("sol_nano_2p9m_tn_gram"))
model.load_state_dict(load_file(str(model_dir / "model.safetensors")), strict=True)
model = model.cuda().eval()
tokenizer = Tokenizer.from_file(str(model_dir / "tokenizer.json"))
prompt = "The sum of 12 and 7 is"
ids = tokenizer.encode(prompt, add_special_tokens=False).ids
inputs = torch.tensor([ids], dtype=torch.long, device="cuda")
with torch.inference_mode():
logits = model(inputs) # [batch, sequence, vocabulary]
next_id = logits[0, -1].argmax().item()
print(tokenizer.decode([next_id]))
The training run
Training ran on one RTX PRO 6000 Blackwell Server Edition with fused AdamW. The model used BF16; optimizer states stayed in FP32. A full optimizer update covered 512 sequences of 512 tokens. CPU workers streamed and tokenized the data while the GPU trained, with the complete scheduled mixture in each update.
The learning rate peaked at 0.001. The WSD schedule warmed up linearly during the first 2% of updates, held that rate through 90%, then decayed linearly to zero over the last 10%.
| Phase | FineWeb-Edu | FineMath | OpenWebMath | Generated math | Procedural | Physical science | Code |
|---|---|---|---|---|---|---|---|
| Opening, about 0-1.333B tokens | 65% | 7.5% | 4.5% | 3% | 12% | 4% | 4% |
| Main, about 1.4-4.5B tokens | 45% | 20% | 12% | 8% | 8% | 3% | 4% |
| Final 10% of optimizer steps | 30% | 30% | 20% | 10% | 4% | 2% | 4% |
The opening phase moves into the main phase through a 66.85M-token ramp. In the final phase, FineWeb-Edu examples need a score of at least 3.5 and FineMath examples at least 4.5. The procedural subset comes from Cosmopedia-v2, the physical-science subset from FineWeb-Edu, and the code from CoRNStack Python. run.json records the exact phase boundaries and source settings.
The training environment used PyTorch 2.11.0+cu130 and Triton 3.6.0.
Measured results
| Benchmark | Examples | Normalized accuracy |
|---|---|---|
| HellaSwag | 10,042 | 28.40% |
| ARC-Easy | 2,376 | 32.07% |
| ARC-Challenge | 1,172 | 21.16% |
| PIQA | 1,838 | 53.92% |
| ArithMark-3 | 1,000 | 33.80% |
The Axiomic Labs Open SLM Intelligence Index is 6.0684. Evaluation used the full zero-shot splits, LM Evaluation Harness 0.4.12, and the official ArithMark-3.0 dataset. Scoring ran in float32 with a 512-token context under PyTorch 2.14.0+cu130. None of the candidate requests needed truncation.
The calculation follows the published methodology. These are local results that Axiomic Labs hasn't independently verified. evaluation/summary.json contains the full-precision scores and checkpoint hashes.
Files and limits
model.safetensors is the 11,591,920-byte weight file. The repository also has the matching tokenizer and modeling_sol_lite.py, along with the configuration and training metadata. It doesn't include optimizer state.
Nano can give incorrect answers. These benchmark scores measure multiple-choice likelihood accuracy; they don't establish reliable free-form problem solving.
- Downloads last month
- 11
