j0no12 commited on
Commit
01657f9
·
verified ·
1 Parent(s): bfa218e

Rewrite model card in plain language

Browse files
Files changed (1) hide show
  1. README.md +12 -16
README.md CHANGED
@@ -16,11 +16,11 @@ tags:
16
 
17
  # Sol Lassi 600K
18
 
19
- Sol Lassi is a 600,000-parameter experiment in training a language model with M-SimOW. It runs with MLX on Apple Silicon. Training started from scratch and used 1 billion token exposures.
20
 
21
- Lassi is a base model for text completion. Its size makes it useful for optimizer experiments, but completions can repeat, lose coherence, or give incorrect answers. It hasn't been instruction-tuned.
22
 
23
- ## Model details
24
 
25
  | Setting | Value |
26
  |---|---|
@@ -35,9 +35,13 @@ Lassi is a base model for text completion. Its size makes it useful for optimize
35
  | Training exposures | 1,000,000,000 |
36
  | Weights | MLX NPZ |
37
 
38
- ## Run a completion
39
 
40
- Download the repository and install the dependencies in `requirements.txt`. Add the downloaded directory to Python's path before importing the loader:
 
 
 
 
41
 
42
  ```bash
43
  pip install -r requirements.txt
@@ -55,18 +59,10 @@ model, tokenizer = load_model(model_dir)
55
  print(generate(model, tokenizer, "A tiny language model can", max_new_tokens=48))
56
  ```
57
 
58
- ## Data and optimizer settings
59
-
60
- Training used a balanced FinePhrase stream from the `faq`, `math`, `table`, and `tutorial` configurations. The FinePhrase-balanced-500m-2k-v2 manifest has 500M stream tokens; the run accumulated 1B exposures.
61
-
62
- The sequence length was 128 tokens, with a batch size of 32 and seed 7. M-SimOW used momentum beta 0.8 and decoupled weight decay 0.1. The first 500M exposures used a learning rate of 0.006, followed by 0.003 for the remaining 500M.
63
-
64
- ## What has been measured
65
-
66
- This repository doesn't include an Open SLM, ArithMark, or Intelligence Index evaluation. The loss logs record individual training minibatches, so they don't give a held-out score.
67
 
68
  ## Files and license
69
 
70
- The weights are in `model.npz`. `modeling_sol_lassi.py` supplies the loader and generator and depends on `keystone_mlx/`. The repository also has the tokenizer and architecture configuration, with run metadata in `training_state.json`.
71
 
72
- The model uses [CC BY 4.0](LICENSE). FinePhrase and its source datasets keep their own terms.
 
16
 
17
  # Sol Lassi 600K
18
 
19
+ Lassi is an M-SimOW optimizer experiment: 600,000 parameters, trained from scratch with MLX on Apple Silicon. The run reached 1 billion token exposures using a balanced FinePhrase stream.
20
 
21
+ There are no Open SLM, ArithMark, or Intelligence Index results for this checkpoint. The available loss logs come from individual training minibatches; they don't measure held-out performance.
22
 
23
+ ## Model and training settings
24
 
25
  | Setting | Value |
26
  |---|---|
 
35
  | Training exposures | 1,000,000,000 |
36
  | Weights | MLX NPZ |
37
 
38
+ The data uses FinePhrase's `faq`, `math`, `table`, and `tutorial` configurations. `FinePhrase-balanced-500m-2k-v2` records 500M stream tokens, while the run accumulated 1B exposures.
39
 
40
+ We trained with 128-token sequences, batch size 32, and seed 7. M-SimOW used momentum beta 0.8 and decoupled weight decay 0.1. The learning rate was 0.006 for the first 500M exposures and 0.003 for the remaining 500M.
41
+
42
+ ## Load and generate
43
+
44
+ Download the repository, install `requirements.txt`, and add the downloaded directory to Python's import path:
45
 
46
  ```bash
47
  pip install -r requirements.txt
 
59
  print(generate(model, tokenizer, "A tiny language model can", max_new_tokens=48))
60
  ```
61
 
62
+ Lassi is a base model for optimizer tests, with no instruction tuning. Generated text may repeat, stop making sense, or state incorrect facts.
 
 
 
 
 
 
 
 
63
 
64
  ## Files and license
65
 
66
+ `model.npz` contains the weights. The loader and generator in `modeling_sol_lassi.py` depend on `keystone_mlx/`. The tokenizer and architecture configuration are included; `training_state.json` records the run.
67
 
68
+ [CC BY 4.0](LICENSE) covers the model. FinePhrase and its source datasets retain their own terms.