Instructions to use solintellegence/Sol-Lite-2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use solintellegence/Sol-Lite-2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="solintellegence/Sol-Lite-2", trust_remote_code=True)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("solintellegence/Sol-Lite-2", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use solintellegence/Sol-Lite-2 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "solintellegence/Sol-Lite-2" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "solintellegence/Sol-Lite-2", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/solintellegence/Sol-Lite-2
- SGLang
How to use solintellegence/Sol-Lite-2 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "solintellegence/Sol-Lite-2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "solintellegence/Sol-Lite-2", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "solintellegence/Sol-Lite-2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "solintellegence/Sol-Lite-2", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use solintellegence/Sol-Lite-2 with Docker Model Runner:
docker model run hf.co/solintellegence/Sol-Lite-2
Sol Lite 2
Sol Lite 2 is a 14,942,592-parameter base language model trained from scratch on exactly 20 billion tokens. We screened 50 architecture runs before choosing its configuration. The released checkpoint keeps recurrence and loop conditioning.
The final checkpoint scored 9.6120 on the Axiomic Open SLM Intelligence Index. It was the highest Index score measured during this production run. Sol Lite 2 completes text and has no instruction tuning.
Architecture
| Setting | Value |
|---|---|
| Trainable parameters | 14,942,592 |
| Hidden width | 256 |
| Context | 2,048 tokens |
| Vocabulary | 4,096, with the existing tokenizer kept fixed |
| Stored blocks / effective applications | 10 / 14 |
| Recurrent layout | 1 prelude, 4 middle blocks used twice, 5 coda blocks |
| Attention | Causal GQA, 8 query heads, 4 KV heads, 32 dimensions per head |
| Position encoding | RoPE, theta 20,000 |
| Normalization | Pre-RMSNorm and learned Q/K RMSNorm |
| FFN | SwiGLU, width 1,552 |
| Recurrence | Learned pass embeddings and loop gates |
| Embeddings | Tied input and output weights |
| Model class | Sol2ForCausalLM |
| Released weights | FP32 safetensors, about 59.8 MB |
Four middle blocks run a second time, so the model performs fourteen block applications while storing ten blocks. Pass embeddings and learned gates condition the repeated pass. The Transformers wrapper uses PyTorch SDPA for inference; training used compiled FlexAttention. KV caching isn't implemented.
Load and generate
Install the dependencies and sign in to Hugging Face with an account that can access this private repository:
pip install torch transformers safetensors tokenizers huggingface_hub
hf auth login
The repository includes the custom model code. Load it through Transformers:
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "solintellegence/Sol-Lite-2"
device = "cuda" if torch.cuda.is_available() else "cpu"
tokenizer = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
repo,
trust_remote_code=True,
).to(device).eval()
inputs = tokenizer("A small language model can", return_tensors="pt")
inputs = {name: value.to(device) for name, value in inputs.items()}
with torch.inference_mode():
output = model.generate(
**inputs,
max_new_tokens=48,
do_sample=False,
use_cache=False,
)
print(tokenizer.decode(output[0], skip_special_tokens=True))
Keep the prompt and generated continuation within the 2,048-token context. This loading example was checked with PyTorch 2.13.0 and Transformers 5.16.1.
Final evaluation
| Benchmark | Examples | Normalized accuracy |
|---|---|---|
| HellaSwag | 10,042 | 28.6895% |
| ARC-Easy | 2,376 | 36.2795% |
| ARC-Challenge | 1,172 | 24.9147% |
| PIQA | 1,838 | 56.8009% |
| ArithMark-3 | 1,000 | 35.5000% |
The full-precision Index is 9.612010364814575. Evaluation used the complete zero-shot splits, float32 scoring, lm-eval 0.4.12, and a 2,048-token context. No requests were truncated. The calculation follows the Axiomic methodology. These are local results, without independent leaderboard verification.
evaluation/index.json records the exact scores and the evaluated training-checkpoint, tokenizer, and ArithMark hashes. The final checkpoint followed the fixed 20B-token schedule. Architecture selection used a separate development proxy and held-out loss guards rather than these full test scores.
Training and data
The run finished at 305,176 optimizer steps on an RTX PRO 6000 Blackwell. We used BF16, fused AdamW with betas 0.9 and 0.95, weight decay 0.1, gradient clipping 1.0, batch size 32, and 2,048-token sequences. The peak learning rate was 0.001. WSD allocated 2% of updates to warmup, 88% to the peak-rate hold, and 10% to linear decay to zero.
Six matched mixture pilots selected more_web. Its starting quotas were:
| Source | Starting token share |
|---|---|
| FineWeb-Edu | 65% |
| Cosmopedia v2 | 20% |
| OpenMathInstruct-2 | 8% |
| MegaScience biology / medicine | 3% |
| High-Quality English Sentences | 2% |
| Text-only ScienceQA | 1.5% |
| Orca-Math | 0.5% |
The mixture changed during production. Tiny Strange Textbooks was inaccessible and excluded before selection. MegaScience later produced no accepted documents within the rejection limit, so its quota moved to FineWeb. FineMath4+ was added at 14% of the remaining forward mixture by taking 14 percentage points from FineWeb. Smaller finite sources had a three-pass cap, with exhausted quotas redirected to FineWeb. The table describes the starting recipe, not the realized shares across all 20B tokens.
We streamed pinned sources with English, length, repetition, and exact-duplicate filters. FineWeb required an education score of at least 3; ScienceQA used image-free training examples whose text didn't refer to an image. Content hashes split documents 95/5, and question identity kept alternative QA solutions in the same split. Deduplication used bounded caches. No frozen training corpus was prepared, and these filters don't establish complete benchmark decontamination.
Selection and limits
The architecture screen tested 25 treatments at two seeds, each for 20M tokens. Three fresh-seed confirmations and correctness checks followed. Optional treatments didn't meet the promotion rules, so we retained the baseline. These short runs support the choice made for this campaign; they don't establish that memory or attention-bias methods fail at other sizes or budgets.
Sol Lite 2 is intended for text-completion and small-model research. It can repeat itself, lose coherence, or give incorrect answers. Multiple-choice likelihood scores don't establish reliable free-form reasoning or instruction following.
Files and license
model.safetensors contains the inference weights. configuration_sol2.py, modeling_sol2.py, and sol2_core.py define the custom model. The repository also includes its tokenizer, generation configuration, banner, and evaluation record. The inference download doesn't include optimizer or data-stream recovery state.
Sol Lite 2 is released under the Apache License 2.0. See LICENSE for the full terms. The training datasets retain their own terms.
- Downloads last month
- -
