--- license: cc-by-4.0 language: - en pipeline_tag: text-generation tags: - causal-lm - decoder-only - small-language-model - mlx - experimental - sol-intelligence --- ![Sol Lassi](./sol-lassi-banner.png) # Sol Lassi 600K Lassi is an M-SimOW optimizer experiment: 600,000 parameters, trained from scratch with MLX on Apple Silicon. The run reached 1 billion token exposures using a balanced FinePhrase stream. There are no Open SLM, ArithMark, or Intelligence Index results for this checkpoint. The available loss logs come from individual training minibatches; they don't measure held-out performance. ## Model and training settings | Setting | Value | |---|---| | Parameters | 600,000 | | Blocks | 6 independent transformer blocks | | Hidden width | 96 | | Attention | 3 heads, head dimension 32 | | FFN | Gated SiLU, width 104 | | Tokenizer | 2,048-entry byte-level BPE | | Token embeddings | Tied to the output head | | Training context | 128 tokens | | Training exposures | 1,000,000,000 | | Weights | MLX NPZ | The data uses FinePhrase's `faq`, `math`, `table`, and `tutorial` configurations. `FinePhrase-balanced-500m-2k-v2` records 500M stream tokens, while the run accumulated 1B exposures. We trained with 128-token sequences, batch size 32, and seed 7. M-SimOW used momentum beta 0.8 and decoupled weight decay 0.1. The learning rate was 0.006 for the first 500M exposures and 0.003 for the remaining 500M. ## Load and generate Download the repository, install `requirements.txt`, and add the downloaded directory to Python's import path: ```bash pip install -r requirements.txt ``` ```python import sys from huggingface_hub import snapshot_download model_dir = snapshot_download("solintellegence/sol-lassi") sys.path.insert(0, model_dir) from modeling_sol_lassi import load_model, generate model, tokenizer = load_model(model_dir) print(generate(model, tokenizer, "A tiny language model can", max_new_tokens=48)) ``` Lassi is a base model for optimizer tests, with no instruction tuning. Generated text may repeat, stop making sense, or state incorrect facts. ## Files and license `model.npz` contains the weights. The loader and generator in `modeling_sol_lassi.py` depend on `keystone_mlx/`. The tokenizer and architecture configuration are included; `training_state.json` records the run. [CC BY 4.0](LICENSE) covers the model. FinePhrase and its source datasets retain their own terms.