Text Generation
MLX
English
sol_lassi
causal-lm
decoder-only
small-language-model
experimental
sol-intelligence
Instructions to use solintellegence/Sol-Lassi with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use solintellegence/Sol-Lassi with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # if on a CUDA device, also pip install mlx[cuda] # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("solintellegence/Sol-Lassi") prompt = "Once upon a time in" text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- MLX LM
How to use solintellegence/Sol-Lassi with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Generate some text mlx_lm.generate --model "solintellegence/Sol-Lassi" --prompt "Once upon a time"
- Atomic Chat
Rewrite model card in plain language
Browse files
README.md
CHANGED
|
@@ -16,9 +16,11 @@ tags:
|
|
| 16 |
|
| 17 |
# Sol Lassi 600K
|
| 18 |
|
| 19 |
-
Sol Lassi is a 600,000-parameter
|
| 20 |
|
| 21 |
-
|
|
|
|
|
|
|
| 22 |
|
| 23 |
| Setting | Value |
|
| 24 |
|---|---|
|
|
@@ -33,9 +35,9 @@ Sol Lassi is a 600,000-parameter language model trained from scratch on 1 billio
|
|
| 33 |
| Training exposures | 1,000,000,000 |
|
| 34 |
| Weights | MLX NPZ |
|
| 35 |
|
| 36 |
-
##
|
| 37 |
|
| 38 |
-
Download the repository
|
| 39 |
|
| 40 |
```bash
|
| 41 |
pip install -r requirements.txt
|
|
@@ -53,20 +55,18 @@ model, tokenizer = load_model(model_dir)
|
|
| 53 |
print(generate(model, tokenizer, "A tiny language model can", max_new_tokens=48))
|
| 54 |
```
|
| 55 |
|
| 56 |
-
##
|
| 57 |
-
|
| 58 |
-
The model used a balanced FinePhrase stream drawn from the `faq`, `math`, `table`, and `tutorial` configurations. The FinePhrase-balanced-500m-2k-v2 manifest contains 500M stream tokens, which produced 1B training exposures.
|
| 59 |
|
| 60 |
-
Training used
|
| 61 |
|
| 62 |
-
|
| 63 |
|
| 64 |
-
|
| 65 |
|
| 66 |
-
|
| 67 |
|
| 68 |
## Files and license
|
| 69 |
|
| 70 |
-
`model.npz`
|
| 71 |
|
| 72 |
-
The model
|
|
|
|
| 16 |
|
| 17 |
# Sol Lassi 600K
|
| 18 |
|
| 19 |
+
Sol Lassi is a 600,000-parameter experiment in training a language model with M-SimOW. It runs with MLX on Apple Silicon. Training started from scratch and used 1 billion token exposures.
|
| 20 |
|
| 21 |
+
Lassi is a base model for text completion. Its size makes it useful for optimizer experiments, but completions can repeat, lose coherence, or give incorrect answers. It hasn't been instruction-tuned.
|
| 22 |
+
|
| 23 |
+
## Model details
|
| 24 |
|
| 25 |
| Setting | Value |
|
| 26 |
|---|---|
|
|
|
|
| 35 |
| Training exposures | 1,000,000,000 |
|
| 36 |
| Weights | MLX NPZ |
|
| 37 |
|
| 38 |
+
## Run a completion
|
| 39 |
|
| 40 |
+
Download the repository and install the dependencies in `requirements.txt`. Add the downloaded directory to Python's path before importing the loader:
|
| 41 |
|
| 42 |
```bash
|
| 43 |
pip install -r requirements.txt
|
|
|
|
| 55 |
print(generate(model, tokenizer, "A tiny language model can", max_new_tokens=48))
|
| 56 |
```
|
| 57 |
|
| 58 |
+
## Data and optimizer settings
|
|
|
|
|
|
|
| 59 |
|
| 60 |
+
Training used a balanced FinePhrase stream from the `faq`, `math`, `table`, and `tutorial` configurations. The FinePhrase-balanced-500m-2k-v2 manifest has 500M stream tokens; the run accumulated 1B exposures.
|
| 61 |
|
| 62 |
+
The sequence length was 128 tokens, with a batch size of 32 and seed 7. M-SimOW used momentum beta 0.8 and decoupled weight decay 0.1. The first 500M exposures used a learning rate of 0.006, followed by 0.003 for the remaining 500M.
|
| 63 |
|
| 64 |
+
## What has been measured
|
| 65 |
|
| 66 |
+
This repository doesn't include an Open SLM, ArithMark, or Intelligence Index evaluation. The loss logs record individual training minibatches, so they don't give a held-out score.
|
| 67 |
|
| 68 |
## Files and license
|
| 69 |
|
| 70 |
+
The weights are in `model.npz`. `modeling_sol_lassi.py` supplies the loader and generator and depends on `keystone_mlx/`. The repository also has the tokenizer and architecture configuration, with run metadata in `training_state.json`.
|
| 71 |
|
| 72 |
+
The model uses [CC BY 4.0](LICENSE). FinePhrase and its source datasets keep their own terms.
|