Text Generation
MLX
English
sol_lassi
causal-lm
decoder-only
small-language-model
experimental
sol-intelligence
Instructions to use solintellegence/Sol-Lassi with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use solintellegence/Sol-Lassi with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # if on a CUDA device, also pip install mlx[cuda] # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("solintellegence/Sol-Lassi") prompt = "Once upon a time in" text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- MLX LM
How to use solintellegence/Sol-Lassi with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Generate some text mlx_lm.generate --model "solintellegence/Sol-Lassi" --prompt "Once upon a time"
- Atomic Chat
Download modeling_sol_lassi.py from solintellegence/Sol-Lassi: direct link, hf CLI and curl.
- Browser
- Download file 1.8 kB
-
https://huggingface.co/solintellegence/Sol-Lassi/resolve/main/modeling_sol_lassi.py
- Command line
-
hf download hf://solintellegence/Sol-Lassi/modeling_sol_lassi.py
-
curl -L -o modeling_sol_lassi.py https://huggingface.co/solintellegence/Sol-Lassi/resolve/main/modeling_sol_lassi.py
1.8 kB
| """Standalone MLX loader and greedy/sampled generation for Sol Lassi 600K.""" | |
| from __future__ import annotations | |
| from pathlib import Path | |
| import mlx.core as mx | |
| from tokenizers import Tokenizer | |
| from keystone_mlx.dense_control import ( | |
| DENSE_CONTROL_600K_DEEP, | |
| DenseControlLM, | |
| parameter_count, | |
| ) | |
| def load_model(model_dir: str | Path): | |
| """Load the frozen 600K MLX weights and tokenizer from a local snapshot.""" | |
| directory = Path(model_dir) | |
| model = DenseControlLM(DENSE_CONTROL_600K_DEEP) | |
| model.load_weights(str(directory / "model.npz")) | |
| mx.eval(model.parameters()) | |
| if parameter_count(model) != 600_000: | |
| raise RuntimeError("Sol Lassi checkpoint does not contain exactly 600,000 parameters") | |
| tokenizer = Tokenizer.from_file(str(directory / "tokenizer.json")) | |
| return model, tokenizer | |
| def generate( | |
| model, | |
| tokenizer: Tokenizer, | |
| prompt: str, | |
| max_new_tokens: int = 64, | |
| temperature: float = 0.0, | |
| seed: int = 0, | |
| ) -> str: | |
| """Generate a continuation; context is truncated to the trained 128 tokens.""" | |
| if max_new_tokens < 0: | |
| raise ValueError("max_new_tokens must be nonnegative") | |
| if temperature < 0: | |
| raise ValueError("temperature must be nonnegative") | |
| mx.random.seed(seed) | |
| ids = tokenizer.encode(prompt).ids or [1] | |
| eos_id = tokenizer.token_to_id("<|eos|>") | |
| for _ in range(max_new_tokens): | |
| context = ids[-128:] | |
| logits = model(mx.array([context], dtype=mx.int32))[0, -1] | |
| if temperature > 0: | |
| next_id = int(mx.random.categorical(logits / temperature).item()) | |
| else: | |
| next_id = int(mx.argmax(logits).item()) | |
| ids.append(next_id) | |
| if eos_id is not None and next_id == eos_id: | |
| break | |
| return tokenizer.decode(ids) | |