File size: 2,434 Bytes
063093a
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
60d7740
063093a
01657f9
063093a
01657f9
bfa218e
01657f9
063093a
60d7740
063093a
60d7740
 
063093a
60d7740
 
 
 
063093a
60d7740
 
063093a
01657f9
063093a
01657f9
 
 
 
 
063093a
 
 
 
 
 
60d7740
063093a
 
 
60d7740
 
 
063093a
 
 
 
01657f9
60d7740
 
 
01657f9
60d7740
01657f9
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
---
license: cc-by-4.0
language:
  - en
pipeline_tag: text-generation
tags:
  - causal-lm
  - decoder-only
  - small-language-model
  - mlx
  - experimental
  - sol-intelligence
---

![Sol Lassi](./sol-lassi-banner.png)

# Sol Lassi 600K

Lassi is an M-SimOW optimizer experiment: 600,000 parameters, trained from scratch with MLX on Apple Silicon. The run reached 1 billion token exposures using a balanced FinePhrase stream.

There are no Open SLM, ArithMark, or Intelligence Index results for this checkpoint. The available loss logs come from individual training minibatches; they don't measure held-out performance.

## Model and training settings

| Setting | Value |
|---|---|
| Parameters | 600,000 |
| Blocks | 6 independent transformer blocks |
| Hidden width | 96 |
| Attention | 3 heads, head dimension 32 |
| FFN | Gated SiLU, width 104 |
| Tokenizer | 2,048-entry byte-level BPE |
| Token embeddings | Tied to the output head |
| Training context | 128 tokens |
| Training exposures | 1,000,000,000 |
| Weights | MLX NPZ |

The data uses FinePhrase's `faq`, `math`, `table`, and `tutorial` configurations. `FinePhrase-balanced-500m-2k-v2` records 500M stream tokens, while the run accumulated 1B exposures.

We trained with 128-token sequences, batch size 32, and seed 7. M-SimOW used momentum beta 0.8 and decoupled weight decay 0.1. The learning rate was 0.006 for the first 500M exposures and 0.003 for the remaining 500M.

## Load and generate

Download the repository, install `requirements.txt`, and add the downloaded directory to Python's import path:

```bash
pip install -r requirements.txt
```

```python
import sys
from huggingface_hub import snapshot_download

model_dir = snapshot_download("solintellegence/sol-lassi")
sys.path.insert(0, model_dir)
from modeling_sol_lassi import load_model, generate

model, tokenizer = load_model(model_dir)
print(generate(model, tokenizer, "A tiny language model can", max_new_tokens=48))
```

Lassi is a base model for optimizer tests, with no instruction tuning. Generated text may repeat, stop making sense, or state incorrect facts.

## Files and license

`model.npz` contains the weights. The loader and generator in `modeling_sol_lassi.py` depend on `keystone_mlx/`. The tokenizer and architecture configuration are included; `training_state.json` records the run.

[CC BY 4.0](LICENSE) covers the model. FinePhrase and its source datasets retain their own terms.