Document fresh balanced training experiment
Browse files
README.md
ADDED
|
@@ -0,0 +1,41 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
language: en
|
| 3 |
+
license: apache-2.0
|
| 4 |
+
library_name: pytorch
|
| 5 |
+
pipeline_tag: text-generation
|
| 6 |
+
tags:
|
| 7 |
+
- experimental
|
| 8 |
+
- from-scratch
|
| 9 |
+
---
|
| 10 |
+
# Slayer149 Balanced
|
| 11 |
+
|
| 12 |
+
Experimental 149,333,081-parameter English causal model trained **from scratch**.
|
| 13 |
+
26 layers, hidden width 576, FFN 2496, 9 query heads / 3 KV heads, tied 24,576-token
|
| 14 |
+
byte-level BPE vocabulary, QK normalization and value residuals. Architecture
|
| 15 |
+
informed by Qwen3 and AltSlate's Apache-2.0 Jugnu recipe; no Jugnu weights reused.
|
| 16 |
+
This is a separate experiment from SlayerLab/Slayer149, which remains available.
|
| 17 |
+
|
| 18 |
+
Training may still be in progress. See `reports/status.json`, training metrics and
|
| 19 |
+
checkpoint manifests. No leaderboard rank or winning result is claimed.
|
| 20 |
+
The final full GLINT-derived evaluation is `reports/evaluation.json`, when present.
|
| 21 |
+
Competitor scores may use different evaluation conventions.
|
| 22 |
+
|
| 23 |
+
`tokenizer.json` preserves Unicode/whitespace and uses EOS ID 0. Tokenizer selection
|
| 24 |
+
and dataset provenance are recorded in reports. Training sources are pinned
|
| 25 |
+
FineWeb-Edu, DCLM-edu, Cosmopedia-v2 and FineMath, with document-disjoint validation
|
| 26 |
+
and exact-overlap filtering against the benchmark suite. Filtering is not a
|
| 27 |
+
semantic contamination guarantee. No training text or credentials are uploaded.
|
| 28 |
+
|
| 29 |
+
`checkpoints/<step>/training-state.pt` contains model and both optimizer states.
|
| 30 |
+
Restore with the matching tokenizer, config, source files and training code.
|
| 31 |
+
Final safetensors appear under the final checkpoint and at repository root.
|
| 32 |
+
Use the supplied custom PyTorch loader; this is not yet an AutoModel package.
|
| 33 |
+
|
| 34 |
+
```python
|
| 35 |
+
from load_model import load_model
|
| 36 |
+
model, tokenizer = load_model("/path/to/downloaded/repository", device="cuda")
|
| 37 |
+
```
|
| 38 |
+
|
| 39 |
+
Goal: have a durable checkpoint and evaluation report by 15:00 Warsaw time on
|
| 40 |
+
2026-10-01, within the user's confirmed free GPU reservation. This is not a
|
| 41 |
+
promise of model quality or benchmark rank.
|