Text Generation
MLX
English
sol_milkshake
causal-lm
decoder-only
small-language-model
recurrent-depth
ngpt
research
Instructions to use solintellegence/Sol-Milkshake with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use solintellegence/Sol-Milkshake with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # if on a CUDA device, also pip install mlx[cuda] # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("solintellegence/Sol-Milkshake") prompt = "Once upon a time in" text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- MLX LM
How to use solintellegence/Sol-Milkshake with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Generate some text mlx_lm.generate --model "solintellegence/Sol-Milkshake" --prompt "Once upon a time"
- Atomic Chat
Download sol_config.py from solintellegence/Sol-Milkshake: direct link, hf CLI and curl.
- Browser
- Download file 1.22 kB
-
https://huggingface.co/solintellegence/Sol-Milkshake/resolve/main/sol_config.py
- Command line
-
hf download hf://solintellegence/Sol-Milkshake/sol_config.py
-
curl -L -o sol_config.py https://huggingface.co/solintellegence/Sol-Milkshake/resolve/main/sol_config.py
1.22 kB
| from __future__ import annotations | |
| from dataclasses import asdict, dataclass | |
| MODEL_NAME = "Sol Milkshake" | |
| ORGANIZATION = "Sol Labs" | |
| DEPLOYED_PARAMS = 2_990_000 | |
| VOCAB_SIZE = 2_048 | |
| D_MODEL = 192 | |
| MAX_CONTEXT = 2_048 | |
| N_PHYSICAL_BLOCKS = 5 | |
| N_Q_HEADS = 6 | |
| N_KV_HEADS = 2 | |
| HEAD_DIM = 32 | |
| FFN_HIDDEN = 512 | |
| TN_RANK = 26 | |
| TIED_EMBEDDING_HEAD = True | |
| XSA_PARAMETER_COUNT = 0 | |
| TOKENS_PER_PARAMETER = 544 | |
| GLOBAL_TOKENS = 1_626_560_000 | |
| GLOBAL_BATCH = 32_768 | |
| COMPLETE_UPDATES = 49_638 | |
| FINAL_PARTIAL_TOKENS = 22_016 | |
| assert DEPLOYED_PARAMS * TOKENS_PER_PARAMETER == GLOBAL_TOKENS | |
| assert COMPLETE_UPDATES * GLOBAL_BATCH + FINAL_PARTIAL_TOKENS == GLOBAL_TOKENS | |
| class SolConfig: | |
| model_name: str = MODEL_NAME | |
| vocab_size: int = VOCAB_SIZE | |
| d_model: int = D_MODEL | |
| max_context: int = MAX_CONTEXT | |
| n_blocks: int = N_PHYSICAL_BLOCKS | |
| n_q_heads: int = N_Q_HEADS | |
| n_kv_heads: int = N_KV_HEADS | |
| head_dim: int = HEAD_DIM | |
| ffn_hidden: int = FFN_HIDDEN | |
| tn_rank: int = TN_RANK | |
| rope_theta: float = 20_000.0 | |
| memory_width: int = 64 | |
| memory_slots: int = 32 | |
| chunk_size: int = 32 | |
| passes: int = 3 | |
| dropout: float = 0.0 | |
| def as_dict(self) -> dict: | |
| return asdict(self) | |