Text Generation
MLX
English
sol_lassi
causal-lm
decoder-only
small-language-model
experimental
sol-intelligence
Instructions to use solintellegence/Sol-Lassi with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use solintellegence/Sol-Lassi with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # if on a CUDA device, also pip install mlx[cuda] # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("solintellegence/Sol-Lassi") prompt = "Once upon a time in" text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- MLX LM
How to use solintellegence/Sol-Lassi with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Generate some text mlx_lm.generate --model "solintellegence/Sol-Lassi" --prompt "Once upon a time"
- Atomic Chat
Rewrite model card in plain language
Browse files
README.md
CHANGED
|
@@ -16,11 +16,11 @@ tags:
|
|
| 16 |
|
| 17 |
# Sol Lassi 600K
|
| 18 |
|
| 19 |
-
|
| 20 |
|
| 21 |
-
|
| 22 |
|
| 23 |
-
## Model
|
| 24 |
|
| 25 |
| Setting | Value |
|
| 26 |
|---|---|
|
|
@@ -35,9 +35,13 @@ Lassi is a base model for text completion. Its size makes it useful for optimize
|
|
| 35 |
| Training exposures | 1,000,000,000 |
|
| 36 |
| Weights | MLX NPZ |
|
| 37 |
|
| 38 |
-
|
| 39 |
|
| 40 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 41 |
|
| 42 |
```bash
|
| 43 |
pip install -r requirements.txt
|
|
@@ -55,18 +59,10 @@ model, tokenizer = load_model(model_dir)
|
|
| 55 |
print(generate(model, tokenizer, "A tiny language model can", max_new_tokens=48))
|
| 56 |
```
|
| 57 |
|
| 58 |
-
|
| 59 |
-
|
| 60 |
-
Training used a balanced FinePhrase stream from the `faq`, `math`, `table`, and `tutorial` configurations. The FinePhrase-balanced-500m-2k-v2 manifest has 500M stream tokens; the run accumulated 1B exposures.
|
| 61 |
-
|
| 62 |
-
The sequence length was 128 tokens, with a batch size of 32 and seed 7. M-SimOW used momentum beta 0.8 and decoupled weight decay 0.1. The first 500M exposures used a learning rate of 0.006, followed by 0.003 for the remaining 500M.
|
| 63 |
-
|
| 64 |
-
## What has been measured
|
| 65 |
-
|
| 66 |
-
This repository doesn't include an Open SLM, ArithMark, or Intelligence Index evaluation. The loss logs record individual training minibatches, so they don't give a held-out score.
|
| 67 |
|
| 68 |
## Files and license
|
| 69 |
|
| 70 |
-
|
| 71 |
|
| 72 |
-
|
|
|
|
| 16 |
|
| 17 |
# Sol Lassi 600K
|
| 18 |
|
| 19 |
+
Lassi is an M-SimOW optimizer experiment: 600,000 parameters, trained from scratch with MLX on Apple Silicon. The run reached 1 billion token exposures using a balanced FinePhrase stream.
|
| 20 |
|
| 21 |
+
There are no Open SLM, ArithMark, or Intelligence Index results for this checkpoint. The available loss logs come from individual training minibatches; they don't measure held-out performance.
|
| 22 |
|
| 23 |
+
## Model and training settings
|
| 24 |
|
| 25 |
| Setting | Value |
|
| 26 |
|---|---|
|
|
|
|
| 35 |
| Training exposures | 1,000,000,000 |
|
| 36 |
| Weights | MLX NPZ |
|
| 37 |
|
| 38 |
+
The data uses FinePhrase's `faq`, `math`, `table`, and `tutorial` configurations. `FinePhrase-balanced-500m-2k-v2` records 500M stream tokens, while the run accumulated 1B exposures.
|
| 39 |
|
| 40 |
+
We trained with 128-token sequences, batch size 32, and seed 7. M-SimOW used momentum beta 0.8 and decoupled weight decay 0.1. The learning rate was 0.006 for the first 500M exposures and 0.003 for the remaining 500M.
|
| 41 |
+
|
| 42 |
+
## Load and generate
|
| 43 |
+
|
| 44 |
+
Download the repository, install `requirements.txt`, and add the downloaded directory to Python's import path:
|
| 45 |
|
| 46 |
```bash
|
| 47 |
pip install -r requirements.txt
|
|
|
|
| 59 |
print(generate(model, tokenizer, "A tiny language model can", max_new_tokens=48))
|
| 60 |
```
|
| 61 |
|
| 62 |
+
Lassi is a base model for optimizer tests, with no instruction tuning. Generated text may repeat, stop making sense, or state incorrect facts.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 63 |
|
| 64 |
## Files and license
|
| 65 |
|
| 66 |
+
`model.npz` contains the weights. The loader and generator in `modeling_sol_lassi.py` depend on `keystone_mlx/`. The tokenizer and architecture configuration are included; `training_state.json` records the run.
|
| 67 |
|
| 68 |
+
[CC BY 4.0](LICENSE) covers the model. FinePhrase and its source datasets retain their own terms.
|