File size: 3,767 Bytes
7aa9bc4
 
 
 
 
 
 
 
 
 
 
 
 
 
 
f68bd14
7aa9bc4
f71be18
7aa9bc4
eb48963
c5c0500
eb48963
c5c0500
eb48963
c5c0500
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
7aa9bc4
eb48963
7aa9bc4
eb48963
7aa9bc4
 
 
 
 
eb48963
7aa9bc4
 
 
 
 
c5c0500
7aa9bc4
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
eb48963
7aa9bc4
c5c0500
 
7aa9bc4
 
 
 
 
 
eb48963
7aa9bc4
eb48963
7aa9bc4
eb48963
7aa9bc4
eb48963
7aa9bc4
eb48963
7aa9bc4
eb48963
7aa9bc4
eb48963
7aa9bc4
eb48963
7aa9bc4
eb48963
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
---
license: cc-by-4.0
language:
  - en
pipeline_tag: text-generation
tags:
  - causal-lm
  - decoder-only
  - small-language-model
  - recurrent-depth
  - mlx
  - ngpt
  - research
---

![Sol Milkshake](./sol-milkshake)

# Sol Milkshake

Milkshake tests recurrence and hyperspherical optimization in MLX on Apple Silicon. Five stored transformer blocks produce eleven block applications. Local token-pattern memory and rolling chunk memory accompany the shared blocks.

The model has 2,990,000 parameters. Its released checkpoint received 2,526,565,888 token exposures across pretraining and recovery. It generates text completions and hasn't been instruction-tuned.

## Configuration

| Setting | Value |
|---|---|
| Parameters | 2,990,000 |
| Token exposures | 2,526,565,888 |
| Context | 2,048 tokens |
| Tokenizer | 2,048-entry byte-level BPE |
| Hidden width | 192 |
| Stored blocks / effective applications | 5 / 11 |
| Recurrent layout | 1 prelude, 3 middle blocks used three times, 1 coda |
| Attention | 6 query heads, 2 KV heads, head dimension 32 |
| Q/K normalization | Unit normalization with RoPE |
| Attention features | XSA value subtraction and value residuals |
| Routing | Full first pass, then 75% and 50% token capacity |
| FFN | Gated SiLU, width 512 |
| Local memory | Rank-26 tensorized 2-5-gram memory |
| Rolling memory | 32-token chunks, 32 slots, width 64 |
| Token embeddings | Tied to the output head |
| Weights | MLX NPZ |

## Run with MLX

Install MLX, the tokenizer library, and the Hugging Face Hub client:

```bash
pip install "mlx>=0.29" "tokenizers>=0.22" huggingface_hub
```

The downloaded repository contains the model implementation, loader, and generation function:

```python
from huggingface_hub import snapshot_download
import sys

model_dir = snapshot_download("solintellegence/Sol-Milkshake")
sys.path.insert(0, model_dir)

from modeling_sol_milkshake import load_model, generate

model, tokenizer = load_model(model_dir)

text = generate(
    model,
    tokenizer,
    "The future of efficient language models is",
    max_new_tokens=64,
)

print(text)
```

## Evaluation and selection

| Benchmark | Examples | Normalized accuracy |
|---|---:|---:|
| HellaSwag | 10,042 | 25.02% |
| ARC-Easy | 2,376 | 29.76% |
| ARC-Challenge | 1,172 | 22.70% |
| PIQA | 1,838 | 52.50% |
| ArithMark-3 | 1,000 | 30.60% |

Milkshake's local Intelligence Index score is 3.158. The four language-model tasks used their complete zero-shot splits and `lm-eval` 0.4.12. ArithMark-3 used independent tokenization and normalized accuracy. See `evals/` for the raw outputs.

We used this same benchmark suite to select the recovery checkpoint, so selection affects the reported scores. They aren't an untouched held-out estimate, and no independent evaluation has verified them.

## Pretraining and recovery

Pretraining drew on FineWeb-Edu, FinePDFs-Edu, English UltraFineWeb multi-domain and question-answer subsets, Cosmopedia, and FineMath. For recovery, the mixture was 70% original frozen curriculum, 15% Cosmopedia-v2 English, and 15% FinePhrase.

Recovery kept the tokenizer and prepared streams fixed. Filtering, deduplication, and decontamination rules also stayed the same. The run record is `training_state.json`.

## Repository contents

Load the MLX weights from `model.npz` with `modeling_sol_milkshake.py`. That file also provides generation; `sol_config.py` defines the architecture. The download includes `config.json`, tokenizer files, training metadata, and evaluation outputs.

Milkshake is intended for research on recurrence, memory, hyperspherical optimization, and MLX inference. Completions can be repetitive, incoherent, or incorrect.

The model license is [CC BY 4.0](LICENSE). Dataset licenses apply separately.