|
Download README.md from solintellegence/Sol-Lite-Base: direct link, hf CLI and curl.
- Browser
- Download file 3.57 kB
-
https://huggingface.co/solintellegence/Sol-Lite-Base/resolve/main/README.md
- Command line
-
hf download hf://solintellegence/Sol-Lite-Base/README.md
-
curl -L -o README.md https://huggingface.co/solintellegence/Sol-Lite-Base/resolve/main/README.md
3.57 kB
| license: cc-by-4.0 | |
| language: | |
| - en | |
| pipeline_tag: text-generation | |
| tags: | |
| - causal-lm | |
| - decoder-only | |
| - small-language-model | |
| - recurrent-depth | |
| - grouped-query-attention | |
| - research | |
|  | |
| # Sol Lite Base | |
| Lite Base reuses four transformer blocks on a second pass. That gives it fourteen block applications while storing ten blocks. We trained the 14,995,843-parameter model from scratch on 8,153,333,760 tokens. | |
| The released weights are for text completion. There was no instruction tuning or assistant safety tuning. The repository includes a PyTorch loader and generation helper. | |
| ## Architecture | |
| | Setting | Value | | |
| |---|---| | |
| | Parameters | 14,995,843 | | |
| | Training tokens | 8,153,333,760 | | |
| | Context | 2,048 tokens | | |
| | Tokenizer | 4,096-entry digit-aware byte-level BPE | | |
| | Hidden width | 256 | | |
| | Stored blocks / effective applications | 10 / 14 | | |
| | Recurrent layout | 1 prelude, 4 middle blocks used twice, 5 coda blocks | | |
| | Attention | 8 query heads, 2 KV heads, head dimension 32 | | |
| | Q/K normalization | Per-head RMSNorm with RoPE | | |
| | Recurrence | Learned pass embeddings and channel-wise refresh gates | | |
| | FFN | Gated SiLU, width 1,465 | | |
| | Local memory | 2,048-entry bigram/trigram EngramLite | | |
| | Token embeddings | Tied to the output head | | |
| | Weights | FP32 safetensors | | |
| Grouped-query attention and loop conditioning support the repeated pass. XSA subtracts attention values, and EngramLite supplies local memory. Learned pass embeddings and channel-wise refresh gates distinguish the two uses of the middle blocks. | |
| ## Generate text | |
| Install the dependencies, then download the repository and import its standalone implementation: | |
| ```bash | |
| pip install "torch>=2.5" "transformers>=5" safetensors huggingface_hub | |
| ``` | |
| ```python | |
| from huggingface_hub import snapshot_download | |
| import sys | |
| model_dir = snapshot_download("solintellegence/Sol-Lite-Base") | |
| sys.path.insert(0, model_dir) | |
| from modeling_sol_lite import load_model, generate | |
| model, tokenizer = load_model(model_dir, device="cpu") | |
| text = generate( | |
| model, | |
| tokenizer, | |
| "The future of efficient language models is", | |
| max_new_tokens=64, | |
| ) | |
| print(text) | |
| ``` | |
| ## Benchmark results | |
| | Benchmark | Examples | Normalized accuracy | | |
| |---|---:|---:| | |
| | HellaSwag | 10,042 | 27.72% | | |
| | ARC-Easy | 2,376 | 34.72% | | |
| | ARC-Challenge | 1,172 | 22.87% | | |
| | PIQA | 1,838 | 57.73% | | |
| | ArithMark-3 | 1,000 | 34.20% | | |
| The local Intelligence Index score is 8.799. We ran the complete zero-shot language-model splits with `lm-eval` 0.4.12. ArithMark-3 used independent tokenization and normalized accuracy. Raw evaluation outputs are saved in `evals/`. | |
| ## Training sources | |
| The English curriculum combined FineWeb-Edu, DCLM, UltraFineWeb levels 1-3, FineWeb-HQ, FineMath, and FinePhrase. We prepared the stream beforehand and kept the tokenizer fixed. Public benchmark examples weren't used to choose the checkpoint. | |
| ## Files | |
| The checkpoint is `model.safetensors`. Use `modeling_sol_lite.py` for `SolForCausalLM`, loading, and generation; `config.json` supplies the architecture settings. `tokenizer.json` and `tokenizer_config.json` are the matching tokenizer files. Run provenance is in `training_state.json`. | |
| ## Intended use and license | |
| Lite Base is for experiments with small language models, recurrence, memory, and distillation. Generated text can repeat itself, lose coherence, or give wrong answers. Its training doesn't cover instruction following. | |
| The model is licensed under [CC BY 4.0](LICENSE). Upstream datasets retain their own terms. | |