j0no12 commited on
Commit
bfa218e
·
verified ·
1 Parent(s): 60d7740

Rewrite model card in plain language

Browse files
Files changed (1) hide show
  1. README.md +13 -13
README.md CHANGED
@@ -16,9 +16,11 @@ tags:
16
 
17
  # Sol Lassi 600K
18
 
19
- Sol Lassi is a 600,000-parameter language model trained from scratch on 1 billion token exposures. It uses M-SimOW and runs with MLX on Apple Silicon.
20
 
21
- ## Model
 
 
22
 
23
  | Setting | Value |
24
  |---|---|
@@ -33,9 +35,9 @@ Sol Lassi is a 600,000-parameter language model trained from scratch on 1 billio
33
  | Training exposures | 1,000,000,000 |
34
  | Weights | MLX NPZ |
35
 
36
- ## Use
37
 
38
- Download the repository, install its dependencies, and add its directory to the Python path before importing the loader:
39
 
40
  ```bash
41
  pip install -r requirements.txt
@@ -53,20 +55,18 @@ model, tokenizer = load_model(model_dir)
53
  print(generate(model, tokenizer, "A tiny language model can", max_new_tokens=48))
54
  ```
55
 
56
- ## Training
57
-
58
- The model used a balanced FinePhrase stream drawn from the `faq`, `math`, `table`, and `tutorial` configurations. The FinePhrase-balanced-500m-2k-v2 manifest contains 500M stream tokens, which produced 1B training exposures.
59
 
60
- Training used 128-token sequences, a batch size of 32, and seed 7. M-SimOW used momentum beta 0.8 and decoupled weight decay 0.1. The learning rate was 0.006 for the first 500M exposures, then 0.003 for the remaining 500M.
61
 
62
- ## Evaluation and limits
63
 
64
- No Open SLM, ArithMark, or Intelligence Index evaluation is included. The training-loss logs measure individual minibatches and are not held-out scores.
65
 
66
- Lassi is a base-model and optimizer experiment. At this size, completions can be incoherent, repetitive, or incorrect. It has not been instruction-tuned.
67
 
68
  ## Files and license
69
 
70
- `model.npz` contains the weights. `modeling_sol_lassi.py` provides the loader and generator, which depend on `keystone_mlx/`. The repository also includes the tokenizer, architecture configuration, and `training_state.json`.
71
 
72
- The model is licensed under [CC BY 4.0](LICENSE). FinePhrase and its sources retain their own terms.
 
16
 
17
  # Sol Lassi 600K
18
 
19
+ Sol Lassi is a 600,000-parameter experiment in training a language model with M-SimOW. It runs with MLX on Apple Silicon. Training started from scratch and used 1 billion token exposures.
20
 
21
+ Lassi is a base model for text completion. Its size makes it useful for optimizer experiments, but completions can repeat, lose coherence, or give incorrect answers. It hasn't been instruction-tuned.
22
+
23
+ ## Model details
24
 
25
  | Setting | Value |
26
  |---|---|
 
35
  | Training exposures | 1,000,000,000 |
36
  | Weights | MLX NPZ |
37
 
38
+ ## Run a completion
39
 
40
+ Download the repository and install the dependencies in `requirements.txt`. Add the downloaded directory to Python's path before importing the loader:
41
 
42
  ```bash
43
  pip install -r requirements.txt
 
55
  print(generate(model, tokenizer, "A tiny language model can", max_new_tokens=48))
56
  ```
57
 
58
+ ## Data and optimizer settings
 
 
59
 
60
+ Training used a balanced FinePhrase stream from the `faq`, `math`, `table`, and `tutorial` configurations. The FinePhrase-balanced-500m-2k-v2 manifest has 500M stream tokens; the run accumulated 1B exposures.
61
 
62
+ The sequence length was 128 tokens, with a batch size of 32 and seed 7. M-SimOW used momentum beta 0.8 and decoupled weight decay 0.1. The first 500M exposures used a learning rate of 0.006, followed by 0.003 for the remaining 500M.
63
 
64
+ ## What has been measured
65
 
66
+ This repository doesn't include an Open SLM, ArithMark, or Intelligence Index evaluation. The loss logs record individual training minibatches, so they don't give a held-out score.
67
 
68
  ## Files and license
69
 
70
+ The weights are in `model.npz`. `modeling_sol_lassi.py` supplies the loader and generator and depends on `keystone_mlx/`. The repository also has the tokenizer and architecture configuration, with run metadata in `training_state.json`.
71
 
72
+ The model uses [CC BY 4.0](LICENSE). FinePhrase and its source datasets keep their own terms.