Sol-Milkshake / README.md
j0no12's picture
Rewrite model card in plain language
eb48963 verified
|
Raw History Blame Contribute Delete
3.77 kB
---
license: cc-by-4.0
language:
- en
pipeline_tag: text-generation
tags:
- causal-lm
- decoder-only
- small-language-model
- recurrent-depth
- mlx
- ngpt
- research
---
![Sol Milkshake](./sol-milkshake)
# Sol Milkshake
Milkshake tests recurrence and hyperspherical optimization in MLX on Apple Silicon. Five stored transformer blocks produce eleven block applications. Local token-pattern memory and rolling chunk memory accompany the shared blocks.
The model has 2,990,000 parameters. Its released checkpoint received 2,526,565,888 token exposures across pretraining and recovery. It generates text completions and hasn't been instruction-tuned.
## Configuration
| Setting | Value |
|---|---|
| Parameters | 2,990,000 |
| Token exposures | 2,526,565,888 |
| Context | 2,048 tokens |
| Tokenizer | 2,048-entry byte-level BPE |
| Hidden width | 192 |
| Stored blocks / effective applications | 5 / 11 |
| Recurrent layout | 1 prelude, 3 middle blocks used three times, 1 coda |
| Attention | 6 query heads, 2 KV heads, head dimension 32 |
| Q/K normalization | Unit normalization with RoPE |
| Attention features | XSA value subtraction and value residuals |
| Routing | Full first pass, then 75% and 50% token capacity |
| FFN | Gated SiLU, width 512 |
| Local memory | Rank-26 tensorized 2-5-gram memory |
| Rolling memory | 32-token chunks, 32 slots, width 64 |
| Token embeddings | Tied to the output head |
| Weights | MLX NPZ |
## Run with MLX
Install MLX, the tokenizer library, and the Hugging Face Hub client:
```bash
pip install "mlx>=0.29" "tokenizers>=0.22" huggingface_hub
```
The downloaded repository contains the model implementation, loader, and generation function:
```python
from huggingface_hub import snapshot_download
import sys
model_dir = snapshot_download("solintellegence/Sol-Milkshake")
sys.path.insert(0, model_dir)
from modeling_sol_milkshake import load_model, generate
model, tokenizer = load_model(model_dir)
text = generate(
model,
tokenizer,
"The future of efficient language models is",
max_new_tokens=64,
)
print(text)
```
## Evaluation and selection
| Benchmark | Examples | Normalized accuracy |
|---|---:|---:|
| HellaSwag | 10,042 | 25.02% |
| ARC-Easy | 2,376 | 29.76% |
| ARC-Challenge | 1,172 | 22.70% |
| PIQA | 1,838 | 52.50% |
| ArithMark-3 | 1,000 | 30.60% |
Milkshake's local Intelligence Index score is 3.158. The four language-model tasks used their complete zero-shot splits and `lm-eval` 0.4.12. ArithMark-3 used independent tokenization and normalized accuracy. See `evals/` for the raw outputs.
We used this same benchmark suite to select the recovery checkpoint, so selection affects the reported scores. They aren't an untouched held-out estimate, and no independent evaluation has verified them.
## Pretraining and recovery
Pretraining drew on FineWeb-Edu, FinePDFs-Edu, English UltraFineWeb multi-domain and question-answer subsets, Cosmopedia, and FineMath. For recovery, the mixture was 70% original frozen curriculum, 15% Cosmopedia-v2 English, and 15% FinePhrase.
Recovery kept the tokenizer and prepared streams fixed. Filtering, deduplication, and decontamination rules also stayed the same. The run record is `training_state.json`.
## Repository contents
Load the MLX weights from `model.npz` with `modeling_sol_milkshake.py`. That file also provides generation; `sol_config.py` defines the architecture. The download includes `config.json`, tokenizer files, training metadata, and evaluation outputs.
Milkshake is intended for research on recurrence, memory, hyperspherical optimization, and MLX inference. Completions can be repetitive, incoherent, or incorrect.
The model license is [CC BY 4.0](LICENSE). Dataset licenses apply separately.