Sol-Lassi / README.md
j0no12's picture
Rewrite model card in plain language
01657f9 verified
|
Raw History Blame Contribute Delete
2.43 kB
---
license: cc-by-4.0
language:
- en
pipeline_tag: text-generation
tags:
- causal-lm
- decoder-only
- small-language-model
- mlx
- experimental
- sol-intelligence
---
![Sol Lassi](./sol-lassi-banner.png)
# Sol Lassi 600K
Lassi is an M-SimOW optimizer experiment: 600,000 parameters, trained from scratch with MLX on Apple Silicon. The run reached 1 billion token exposures using a balanced FinePhrase stream.
There are no Open SLM, ArithMark, or Intelligence Index results for this checkpoint. The available loss logs come from individual training minibatches; they don't measure held-out performance.
## Model and training settings
| Setting | Value |
|---|---|
| Parameters | 600,000 |
| Blocks | 6 independent transformer blocks |
| Hidden width | 96 |
| Attention | 3 heads, head dimension 32 |
| FFN | Gated SiLU, width 104 |
| Tokenizer | 2,048-entry byte-level BPE |
| Token embeddings | Tied to the output head |
| Training context | 128 tokens |
| Training exposures | 1,000,000,000 |
| Weights | MLX NPZ |
The data uses FinePhrase's `faq`, `math`, `table`, and `tutorial` configurations. `FinePhrase-balanced-500m-2k-v2` records 500M stream tokens, while the run accumulated 1B exposures.
We trained with 128-token sequences, batch size 32, and seed 7. M-SimOW used momentum beta 0.8 and decoupled weight decay 0.1. The learning rate was 0.006 for the first 500M exposures and 0.003 for the remaining 500M.
## Load and generate
Download the repository, install `requirements.txt`, and add the downloaded directory to Python's import path:
```bash
pip install -r requirements.txt
```
```python
import sys
from huggingface_hub import snapshot_download
model_dir = snapshot_download("solintellegence/sol-lassi")
sys.path.insert(0, model_dir)
from modeling_sol_lassi import load_model, generate
model, tokenizer = load_model(model_dir)
print(generate(model, tokenizer, "A tiny language model can", max_new_tokens=48))
```
Lassi is a base model for optimizer tests, with no instruction tuning. Generated text may repeat, stop making sense, or state incorrect facts.
## Files and license
`model.npz` contains the weights. The loader and generator in `modeling_sol_lassi.py` depend on `keystone_mlx/`. The tokenizer and architecture configuration are included; `training_state.json` records the run.
[CC BY 4.0](LICENSE) covers the model. FinePhrase and its source datasets retain their own terms.