File size: 6,201 Bytes
3f431df
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
7857dd2
3f431df
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
7857dd2
 
3f431df
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
---
license: apache-2.0
language:
- en
library_name: pytorch
pipeline_tag: text-generation
base_model: SlayerLab/gollem-v5-ckpts
tags:
- gollem
- tiny-lm
- muon
- value-residual
- continued-pretraining
- glint-tiny-ml-leaderboard
---

# Slayer149 — 149M English language model

A **148,910,738-parameter English causal language model**, grown from the GoLLeM-v5
122.8M final checkpoint and continued for **20,000,014,336 additional tokens**.
This is a base completion model, not an instruction-tuned assistant.

## Evaluation

Full final-checkpoint evaluation, with no best-checkpoint selection:

| Metric | Result | Coverage |
|---|---:|---|
| BLiMP accuracy | **79.47%** | 67,000 pairs / 67 tasks |
| ARC-Easy accuracy | **54.97%** | 2,376 test examples |
| WikiText-2 byte perplexity | **2.212573** | Full test text |
| WikiText-2 token perplexity | 21.452430 | Same evaluation |
| GLINT overall score, fixed board snapshot | **77.11235** | Calculated |
| GLINT efficiency score, fixed board snapshot | **77.13585** | Calculated |

The likelihood protocol follows `Glint-Research/Glint-1.3/benchmark.py` at
`c3f99e246aa2c64382f9668dc864533617482d0a`: first-256-token truncation for BLiMP
and ARC, summed unnormalized log probabilities, zero-shot ARC scoring, and
non-overlapping 256-token WikiText contexts. Inference uses BF16 autocast with
FP32 log probabilities. Byte perplexity is converted from token perplexity using
the measured token/UTF-8-byte ratio. Dataset revisions are pinned in
[evaluation.json](evaluation.json). Loader order and complete task counts are checked.

Against the [GLINT board](https://huggingface.co/spaces/Glint-Research/Tiny-ML-Leaderboard)
at revision `2c5ea9682babfe2c410ebbbf78113cc98f08a2f9`, this would place **4th in
overall score** and **8th in size-adjusted efficiency** among the existing entries
plus this model. This is a calculated, self-reported placement, not an accepted
leaderboard entry. Competitor scores were not independently re-evaluated here;
some competitors use lm-eval-harness and different context/scoring conventions.
The ranking and normalized score can change as the board changes. This release
does not establish a leaderboard win.

## Use

This release uses a small custom PyTorch architecture. Use the included loader;
it is not packaged for `transformers.AutoModelForCausalLM`.

```bash
pip install huggingface-hub
hf download SlayerLab/Slayer149 --local-dir Slayer149
cd Slayer149
pip install -r requirements.txt
python generate.py --prompt "The scientific method is" --max-new-tokens 64
```

A matching
CUDA-enabled PyTorch installation is needed for GPU inference/evaluation.

```python
from load_model import load_model
model, tokenizer = load_model('.', device='cuda')
# model(input_ids) returns (logits, optional_loss)
```

Weights remain FP32 in `model.safetensors`; tied embeddings are restored by
`safetensors.torch.load_model`. Export verification checks every model tensor
and compares logits against the original checkpoint at lengths 16, 256 and 1024.
The generation CLI uses BF16 autocast on CUDA, FP32 on CPU, and greedy decoding
by default. It recomputes context rather than using a KV cache.

## Architecture and training

- 19 transformer layers; hidden size 768; 12 attention heads; FFN size 2160.
- Vocabulary 12,288; tied input/output embeddings; context length 1024.
- RMSNorm, RoPE (theta 100,000), QK normalization, SwiGLU and value residuals.
- Source: `SlayerLab/gollem-v5-ckpts`, revision
  `963440ce6ab4ada7e95da4c1faa28ebb00c082d4`, file
  `final_128m_16x768/ckpt_760k.pt` (24,903,680,000 source training tokens).
- FFNs were widened with zero new down-projection columns; three residual blocks
  were appended with zero output projections. Initial FP32 logits exactly matched
  the donor on the tested contexts. Optimizer state was reset for continuation.
- Continuation: 20,000,014,336 tokens, 39,147 actual optimizer updates, seed 1337.
  Inherited token lineage totals 44,903,694,336; newly added parameters only saw
  the 20B-token continuation.
- Cleaned ARC-MIX pool: 9,391,706,576 tokens, reconstructed from the published
  SlayerLab corpus/removal list. SHA256:
  `ecfd0a4040a0ece6f07f728f092a1b373b6d253b4adfbe259c5107c1fc16681d`.
  Repeated shuffled passes; training corpus is not redistributed here. This work
  did not independently audit all training/evaluation overlap.
- Muon for hidden matrices and AdamW for auxiliary parameters; cosine learning
  rate 2e-4 to 2e-5, 250 reference-unit warmup; Muon LR ratio 33.3333.
- Eight H100 PCIe GPUs. Effective batch started at 256 sequences, then became
  512 after 524,288,000 continuation tokens. Learning rate remained indexed by
  token position. Grouped Muon, BF16 gradient communication and later CUDA graphs
  improved throughput. Numerical ordering changed; this is not a bitwise replay
  of the original trainer.
- Optimized sustained full-segment throughput: 1.045M tokens/sec including
  startup, validation and saves; later ordinary windows reached about 1.14M.
  These are measurements on this host, not expected performance on every H100 node.

`training_manifest.json` records configuration and hashes. Its `reference_step`
uses 262,144 tokens/unit; it is not the actual optimizer-update count after the
batch change. `training_metrics.jsonl` records the training history.

## Reproduce evaluation

```bash
python evaluate_glint.py --model-dir . --out reproduced-results.json
```

This evaluates all three datasets on CUDA using pinned dataset revisions. It can
take tens of minutes. `board_snapshot.json` pins the normalization constants.
See `release_verification.json` for export and official-function parity checks.

## Limitations and license

English base model intended for research. It can generate incorrect, repetitive,
biased or inappropriate text and is not optimized for instruction following.
Small changes in precision, context length, tokenizer handling or benchmark
harness can change scores. Overall score and size-adjusted efficiency are distinct.

Weights and GoLLeM implementation: Apache-2.0, following the source model's
published license. See `LICENSE` and `NOTICE` for source and evaluation attribution.