Sol Lite 2

Sol Lite 2

Sol Lite 2 is a 14,942,592-parameter base language model trained from scratch on exactly 20 billion tokens. We screened 50 architecture runs before choosing its configuration. The released checkpoint keeps recurrence and loop conditioning.

The final checkpoint scored 9.6120 on the Axiomic Open SLM Intelligence Index. It was the highest Index score measured during this production run. Sol Lite 2 completes text and has no instruction tuning.

Architecture

Setting Value
Trainable parameters 14,942,592
Hidden width 256
Context 2,048 tokens
Vocabulary 4,096, with the existing tokenizer kept fixed
Stored blocks / effective applications 10 / 14
Recurrent layout 1 prelude, 4 middle blocks used twice, 5 coda blocks
Attention Causal GQA, 8 query heads, 4 KV heads, 32 dimensions per head
Position encoding RoPE, theta 20,000
Normalization Pre-RMSNorm and learned Q/K RMSNorm
FFN SwiGLU, width 1,552
Recurrence Learned pass embeddings and loop gates
Embeddings Tied input and output weights
Model class Sol2ForCausalLM
Released weights FP32 safetensors, about 59.8 MB

Four middle blocks run a second time, so the model performs fourteen block applications while storing ten blocks. Pass embeddings and learned gates condition the repeated pass. The Transformers wrapper uses PyTorch SDPA for inference; training used compiled FlexAttention. KV caching isn't implemented.

Load and generate

Install the dependencies and sign in to Hugging Face with an account that can access this private repository:

pip install torch transformers safetensors tokenizers huggingface_hub
hf auth login

The repository includes the custom model code. Load it through Transformers:

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "solintellegence/Sol-Lite-2"
device = "cuda" if torch.cuda.is_available() else "cpu"

tokenizer = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    repo,
    trust_remote_code=True,
).to(device).eval()

inputs = tokenizer("A small language model can", return_tensors="pt")
inputs = {name: value.to(device) for name, value in inputs.items()}

with torch.inference_mode():
    output = model.generate(
        **inputs,
        max_new_tokens=48,
        do_sample=False,
        use_cache=False,
    )

print(tokenizer.decode(output[0], skip_special_tokens=True))

Keep the prompt and generated continuation within the 2,048-token context. This loading example was checked with PyTorch 2.13.0 and Transformers 5.16.1.

Final evaluation

Benchmark Examples Normalized accuracy
HellaSwag 10,042 28.6895%
ARC-Easy 2,376 36.2795%
ARC-Challenge 1,172 24.9147%
PIQA 1,838 56.8009%
ArithMark-3 1,000 35.5000%

The full-precision Index is 9.612010364814575. Evaluation used the complete zero-shot splits, float32 scoring, lm-eval 0.4.12, and a 2,048-token context. No requests were truncated. The calculation follows the Axiomic methodology. These are local results, without independent leaderboard verification.

evaluation/index.json records the exact scores and the evaluated training-checkpoint, tokenizer, and ArithMark hashes. The final checkpoint followed the fixed 20B-token schedule. Architecture selection used a separate development proxy and held-out loss guards rather than these full test scores.

Training and data

The run finished at 305,176 optimizer steps on an RTX PRO 6000 Blackwell. We used BF16, fused AdamW with betas 0.9 and 0.95, weight decay 0.1, gradient clipping 1.0, batch size 32, and 2,048-token sequences. The peak learning rate was 0.001. WSD allocated 2% of updates to warmup, 88% to the peak-rate hold, and 10% to linear decay to zero.

Six matched mixture pilots selected more_web. Its starting quotas were:

Source Starting token share
FineWeb-Edu 65%
Cosmopedia v2 20%
OpenMathInstruct-2 8%
MegaScience biology / medicine 3%
High-Quality English Sentences 2%
Text-only ScienceQA 1.5%
Orca-Math 0.5%

The mixture changed during production. Tiny Strange Textbooks was inaccessible and excluded before selection. MegaScience later produced no accepted documents within the rejection limit, so its quota moved to FineWeb. FineMath4+ was added at 14% of the remaining forward mixture by taking 14 percentage points from FineWeb. Smaller finite sources had a three-pass cap, with exhausted quotas redirected to FineWeb. The table describes the starting recipe, not the realized shares across all 20B tokens.

We streamed pinned sources with English, length, repetition, and exact-duplicate filters. FineWeb required an education score of at least 3; ScienceQA used image-free training examples whose text didn't refer to an image. Content hashes split documents 95/5, and question identity kept alternative QA solutions in the same split. Deduplication used bounded caches. No frozen training corpus was prepared, and these filters don't establish complete benchmark decontamination.

Selection and limits

The architecture screen tested 25 treatments at two seeds, each for 20M tokens. Three fresh-seed confirmations and correctness checks followed. Optional treatments didn't meet the promotion rules, so we retained the baseline. These short runs support the choice made for this campaign; they don't establish that memory or attention-bias methods fail at other sizes or budgets.

Sol Lite 2 is intended for text-completion and small-model research. It can repeat itself, lose coherence, or give incorrect answers. Multiple-choice likelihood scores don't establish reliable free-form reasoning or instruction following.

Files and license

model.safetensors contains the inference weights. configuration_sol2.py, modeling_sol2.py, and sol2_core.py define the custom model. The repository also includes its tokenizer, generation configuration, banner, and evaluation record. The inference download doesn't include optimizer or data-stream recovery state.

Sol Lite 2 is released under the Apache License 2.0. See LICENSE for the full terms. The training datasets retain their own terms.

Downloads last month
-
Safetensors
Model size
14.9M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train solintellegence/Sol-Lite-2

Collection including solintellegence/Sol-Lite-2