Sol-Lassi / README.md
j0no12's picture
Rewrite model card in plain language
01657f9 verified
|
Raw History Blame Contribute Delete
2.43 kB
metadata
license: cc-by-4.0
language:
  - en
pipeline_tag: text-generation
tags:
  - causal-lm
  - decoder-only
  - small-language-model
  - mlx
  - experimental
  - sol-intelligence

Sol Lassi

Sol Lassi 600K

Lassi is an M-SimOW optimizer experiment: 600,000 parameters, trained from scratch with MLX on Apple Silicon. The run reached 1 billion token exposures using a balanced FinePhrase stream.

There are no Open SLM, ArithMark, or Intelligence Index results for this checkpoint. The available loss logs come from individual training minibatches; they don't measure held-out performance.

Model and training settings

Setting Value
Parameters 600,000
Blocks 6 independent transformer blocks
Hidden width 96
Attention 3 heads, head dimension 32
FFN Gated SiLU, width 104
Tokenizer 2,048-entry byte-level BPE
Token embeddings Tied to the output head
Training context 128 tokens
Training exposures 1,000,000,000
Weights MLX NPZ

The data uses FinePhrase's faq, math, table, and tutorial configurations. FinePhrase-balanced-500m-2k-v2 records 500M stream tokens, while the run accumulated 1B exposures.

We trained with 128-token sequences, batch size 32, and seed 7. M-SimOW used momentum beta 0.8 and decoupled weight decay 0.1. The learning rate was 0.006 for the first 500M exposures and 0.003 for the remaining 500M.

Load and generate

Download the repository, install requirements.txt, and add the downloaded directory to Python's import path:

pip install -r requirements.txt
import sys
from huggingface_hub import snapshot_download

model_dir = snapshot_download("solintellegence/sol-lassi")
sys.path.insert(0, model_dir)
from modeling_sol_lassi import load_model, generate

model, tokenizer = load_model(model_dir)
print(generate(model, tokenizer, "A tiny language model can", max_new_tokens=48))

Lassi is a base model for optimizer tests, with no instruction tuning. Generated text may repeat, stop making sense, or state incorrect facts.

Files and license

model.npz contains the weights. The loader and generator in modeling_sol_lassi.py depend on keystone_mlx/. The tokenizer and architecture configuration are included; training_state.json records the run.

CC BY 4.0 covers the model. FinePhrase and its source datasets retain their own terms.