kacperwikiel commited on
Commit
fd0936c
·
verified ·
1 Parent(s): b19c3d2

Document fresh balanced training experiment

Browse files
Files changed (1) hide show
  1. README.md +41 -0
README.md ADDED
@@ -0,0 +1,41 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ language: en
3
+ license: apache-2.0
4
+ library_name: pytorch
5
+ pipeline_tag: text-generation
6
+ tags:
7
+ - experimental
8
+ - from-scratch
9
+ ---
10
+ # Slayer149 Balanced
11
+
12
+ Experimental 149,333,081-parameter English causal model trained **from scratch**.
13
+ 26 layers, hidden width 576, FFN 2496, 9 query heads / 3 KV heads, tied 24,576-token
14
+ byte-level BPE vocabulary, QK normalization and value residuals. Architecture
15
+ informed by Qwen3 and AltSlate's Apache-2.0 Jugnu recipe; no Jugnu weights reused.
16
+ This is a separate experiment from SlayerLab/Slayer149, which remains available.
17
+
18
+ Training may still be in progress. See `reports/status.json`, training metrics and
19
+ checkpoint manifests. No leaderboard rank or winning result is claimed.
20
+ The final full GLINT-derived evaluation is `reports/evaluation.json`, when present.
21
+ Competitor scores may use different evaluation conventions.
22
+
23
+ `tokenizer.json` preserves Unicode/whitespace and uses EOS ID 0. Tokenizer selection
24
+ and dataset provenance are recorded in reports. Training sources are pinned
25
+ FineWeb-Edu, DCLM-edu, Cosmopedia-v2 and FineMath, with document-disjoint validation
26
+ and exact-overlap filtering against the benchmark suite. Filtering is not a
27
+ semantic contamination guarantee. No training text or credentials are uploaded.
28
+
29
+ `checkpoints/<step>/training-state.pt` contains model and both optimizer states.
30
+ Restore with the matching tokenizer, config, source files and training code.
31
+ Final safetensors appear under the final checkpoint and at repository root.
32
+ Use the supplied custom PyTorch loader; this is not yet an AutoModel package.
33
+
34
+ ```python
35
+ from load_model import load_model
36
+ model, tokenizer = load_model("/path/to/downloaded/repository", device="cuda")
37
+ ```
38
+
39
+ Goal: have a durable checkpoint and evaluation report by 15:00 Warsaw time on
40
+ 2026-10-01, within the user's confirmed free GPU reservation. This is not a
41
+ promise of model quality or benchmark rank.