Release Pollock 1.5 (r007)

#3
CHANGELOG.md CHANGED
@@ -2,6 +2,17 @@
2
 
3
  Revision numbers identify published model states independently of release names.
4
 
 
 
 
 
 
 
 
 
 
 
 
5
  ## r006 — Pollock 1.4
6
 
7
  - Returned from the r005 5B corpus to V2 of `SlayerLab/minimal-en-corpus-2.5b`, including its corrected cleaning and record boundaries.
 
2
 
3
  Revision numbers identify published model states independently of release names.
4
 
5
+ ## r007 — Pollock 1.5
6
+
7
+ - Changed the architecture from 12/14/896 with a 1,024-token context to 16/12/768 with a 3,200-wide MLP and a 2,048-token context while remaining below 128M native trainable parameters.
8
+ - Returned from the r006 2.5B V2 corpus to `SlayerLab/minimal-en-corpus-5b` at revision `38bebbd`, with its separate 12,288-entry byte-level BPE tokenizer.
9
+ - Replaced AdamW-only training with Muon for eligible hidden attention and MLP matrices plus fused AdamW for embeddings, the tied output head, normalization parameters, and biases.
10
+ - Replaced random-with-replacement sampling and fixed-subset validation with deterministic shuffled corpus passes, exact resumable loader state, a fixed corpus-wide train sample, and complete validation-split evaluation.
11
+ - Completed exactly four corpus passes at the unchanged 491,520-token effective batch and selected the best full-validation checkpoint at update 43,750.
12
+ - Updated the seven-task English zero-shot suite, Transformers weights, locked E4-best diagnostics, and fixed cross-revision inference samples.
13
+
14
+ Full record: [`training-history/r007.md`](./training-history/r007.md)
15
+
16
  ## r006 — Pollock 1.4
17
 
18
  - Returned from the r005 5B corpus to V2 of `SlayerLab/minimal-en-corpus-2.5b`, including its corrected cleaning and record boundaries.
LICENSE.md CHANGED
@@ -16,14 +16,14 @@ Code components distributed with the project remain subject to their respective
16
 
17
  ## Korpus i wagi / Corpus and weights
18
 
19
- Model został wytrenowany na `SlayerLab/minimal-en-corpus-2.5b`, agregacie danych z wielu źródeł. Korpus nie nadaje dokumentom jednej wspólnej licencji; każdy dokument zachowuje identyfikator źródła i podlega warunkom, licencjom oraz ograniczeniom właściwego upstreamowego datasetu.
20
 
21
  Ze względu na mieszany charakter tych warunków repozytorium modelu używa metadanej Hugging Face `license: other`. Nie jest to przyznanie dodatkowych praw do materiałów źródłowych. Użytkownik powinien przed użyciem, redystrybucją lub zastosowaniem komercyjnym zapoznać się z kartą korpusu i warunkami wszystkich właściwych źródeł:
22
 
23
- <https://huggingface.co/datasets/SlayerLab/minimal-en-corpus-2.5b>
24
 
25
- The model was trained on `SlayerLab/minimal-en-corpus-2.5b`, an aggregate of multiple data sources. The corpus does not apply a single common license to its documents; each document retains its source identifier and remains subject to the terms, licenses, and restrictions of the applicable upstream dataset.
26
 
27
  Because these terms are mixed, the model repository uses the Hugging Face metadata value `license: other`. This notice does not grant additional rights to upstream materials. Before use, redistribution, or commercial application, users should review the corpus card and the terms of every applicable source:
28
 
29
- <https://huggingface.co/datasets/SlayerLab/minimal-en-corpus-2.5b>
 
16
 
17
  ## Korpus i wagi / Corpus and weights
18
 
19
+ Model został wytrenowany na `SlayerLab/minimal-en-corpus-5b`, agregacie danych z wielu źródeł. Korpus nie nadaje dokumentom jednej wspólnej licencji; każdy dokument zachowuje identyfikator źródła i podlega warunkom, licencjom oraz ograniczeniom właściwego upstreamowego datasetu.
20
 
21
  Ze względu na mieszany charakter tych warunków repozytorium modelu używa metadanej Hugging Face `license: other`. Nie jest to przyznanie dodatkowych praw do materiałów źródłowych. Użytkownik powinien przed użyciem, redystrybucją lub zastosowaniem komercyjnym zapoznać się z kartą korpusu i warunkami wszystkich właściwych źródeł:
22
 
23
+ <https://huggingface.co/datasets/SlayerLab/minimal-en-corpus-5b>
24
 
25
+ The model was trained on `SlayerLab/minimal-en-corpus-5b`, an aggregate of multiple data sources. The corpus does not apply a single common license to its documents; each document retains its source identifier and remains subject to the terms, licenses, and restrictions of the applicable upstream dataset.
26
 
27
  Because these terms are mixed, the model repository uses the Hugging Face metadata value `license: other`. This notice does not grant additional rights to upstream materials. Before use, redistribution, or commercial application, users should review the corpus card and the terms of every applicable source:
28
 
29
+ <https://huggingface.co/datasets/SlayerLab/minimal-en-corpus-5b>
README.md CHANGED
@@ -5,38 +5,38 @@ pipeline_tag: text-generation
5
  license: other
6
  license_name: mixed-upstream-dataset-terms
7
  license_link: https://huggingface.co/SlayerLab/pollock-mini-lm-125m/blob/main/LICENSE.md
8
- datasets: [SlayerLab/minimal-en-corpus-2.5b]
9
  tags: [causal-lm, gpt2, nanogpt, bpe, educational, base-model]
10
  model-index:
11
- - name: Pollock 1.4
12
  results:
13
  - task: {type: text-generation, name: Language modeling}
14
- dataset: {type: SlayerLab/minimal-en-corpus-2.5b, name: Minimal EN 2.5B V2 validation (fixed subset), split: validation}
15
- metrics: [{type: loss, value: 2.5363565445, name: Final fixed-subset validation loss}]
16
- - task: {type: text-generation, name: Zero-shot evaluation}
17
- dataset: {type: blimp, name: BLiMP, split: train}
18
- metrics: [{type: acc, value: 0.7691641791}]
19
- - task: {type: text-generation, name: Zero-shot evaluation}
20
- dataset: {type: EleutherAI/lambada_openai, name: LAMBADA OpenAI, split: test}
21
- metrics: [{type: acc, value: 0.2769260625}, {type: perplexity, value: 49.9072065519}]
22
- - task: {type: text-generation, name: Zero-shot evaluation}
23
- dataset: {type: hellaswag, name: HellaSwag, split: validation}
24
- metrics: [{type: acc_norm, value: 0.3013343955}]
25
- - task: {type: text-generation, name: Zero-shot evaluation}
26
- dataset: {type: piqa, name: PIQA, split: validation}
27
- metrics: [{type: acc_norm, value: 0.5968443961}]
28
- - task: {type: text-generation, name: Zero-shot evaluation}
29
- dataset: {type: sciq, name: SciQ, split: test}
30
- metrics: [{type: acc_norm, value: 0.6640000000}]
31
- - task: {type: text-generation, name: Zero-shot evaluation}
32
- dataset: {type: allenai/ai2_arc, config: ARC-Easy, name: ARC-Easy, split: test}
33
- metrics: [{type: acc_norm, value: 0.4343434343}]
34
- - task: {type: text-generation, name: Zero-shot evaluation}
35
- dataset: {type: allenai/ai2_arc, config: ARC-Challenge, name: ARC-Challenge, split: test}
36
- metrics: [{type: acc_norm, value: 0.2414675768}]
37
  ---
38
 
39
- # Pollock 1.4 — r006
40
 
41
  ![Pollock avatar](./assets/pollock-mini-lm-avatar-320.png)
42
 
@@ -48,60 +48,67 @@ model-index:
48
 
49
  Pollock to niewielki, anglojęzyczny model bazowy typu decoder-only, wytrenowany od zera jako czytelny eksperyment edukacyjny. Implementacja bazuje na [nanoGPT](https://github.com/karpathy/nanoGPT) i własnym tokenizerze byte-level BPE. Jest to model do uzupełniania tekstu, nie asystent konwersacyjny.
50
 
51
- Nazwa luźno nawiązuje do gestu malarskiego Jacksona Pollocka: nanoGPT jest płótnem, na którym dane, konfiguracja i decyzje treningowe tworzą różne wzorce zachowania. Pełne dane techniczne tej wersji znajdują się w [`training-history/r006.md`](./training-history/r006.md), a różnice między wydaniami w [`CHANGELOG.md`](./CHANGELOG.md).
52
 
53
  ### Architektura i tokenizer
54
 
55
  | Właściwość | Wartość |
56
  |---|---:|
57
- | Rewizja / wydanie | r006 / Pollock 1.4 |
58
  | Typ | decoder-only Transformer w stylu GPT-2 |
59
- | Warstwy / głowy / embedding | 12 / 14 / 896 |
60
- | Maksymalny kontekst | 1024 tokeny |
 
61
  | Słownik | 12 288 tokenów |
62
- | Parametry nanoGPT | 126 637 952 |
63
- | Łączne unikalne parametry trenowalne | 127 555 456 |
64
  | Tokenizer | byte-level BPE, pretokenizacja w stylu GPT-2 |
65
  | Tokeny specjalne | <code>&lt;&#124;endoftext&#124;&gt;</code>, <code>&lt;&#124;im_start&#124;&gt;</code>, <code>&lt;&#124;im_end&#124;&gt;</code> |
66
 
67
- Artefakt Transformers ma 127 674 624 parametry, w tym 119 168 zerowych parametrów bias dla zgodności z `GPT2LMHeadModel`. Natywny model był trenowany z `bias=False`.
 
 
 
 
68
 
69
  ### Dane i trening
70
 
71
- Model wytrenowano na wersji V2 datasetu [`SlayerLab/minimal-en-corpus-2.5b`](https://huggingface.co/datasets/SlayerLab/minimal-en-corpus-2.5b) przypiętej do commita `d68d992`. Jest to subiektywnie dobrana mieszanka 15 anglojęzycznych źródeł po filtrowaniu języka, deduplikacji dokładnej i przybliżonej, decontaminacji benchmarków oraz dodatkowym czyszczeniu boilerplate'u i błędnie połączonych rekordów. R006 wraca do tokenizera używanego przez r003-r004; nie jest on zgodny z tokenizerem r005 poza rozmiarem słownika i identyfikatorami tokenów specjalnych.
 
 
72
 
73
  | Parametr | Wartość |
74
  |---|---:|
75
- | Tokeny treningowe / walidacyjne | 2 689 323 439 / 5 236 486 |
76
- | Finalny checkpoint | aktualizacja 22 003 |
77
- | Przetworzone tokeny | 10 814 914 560 (około 4,02 epoki) |
78
- | Sekwencja / micro-batch na GPU | 1024 / 12 |
79
  | Akumulacja globalna / na GPU | 40 / 20 micro-stepów |
80
  | Effective batch | 491 520 tokenów |
81
- | Optymalizator | fused AdamW, betas 0.9/0.95 |
82
- | Learning rate | 4e-4 → 4e-5, cosine decay |
83
- | Warmup / weight decay / grad clip | 440 / 0.1 / 1.0 |
84
  | Precyzja | BF16 |
85
  | Sprzęt | 2× NVIDIA GeForce RTX 4090 |
86
- | Framework | PyTorch 2.8.0+cu128, nanoGPT commit `3adf61e` |
87
 
88
  ### Ewaluacja
89
 
90
- Loss treningowy szacowano na stałych podzbiorach po 1 228 800 tokenów na split. Finalny checkpoint uzyskał validation loss **2.536357**; najlepszy wynik to **2.536231** po 10 813 440 000 przetworzonych tokenów. Nie należy porównywać tych wartości bezpośrednio z r005, ponieważ r006 używa innego korpusu, tokenizera i zbioru walidacyjnego.
91
 
92
- Benchmarki wykonano zero-shot na pełnych splitach przy użyciu `lm-evaluation-harness` 0.4.12, batch size 8 i BF16.
93
 
94
  | Benchmark | Główna metryka | Wynik | Próbki |
95
  |---|---|---:|---:|
96
- | BLiMP | acc | 0.769164 | 67 000 |
97
- | LAMBADA | acc | 0.276926 | 5 153 |
98
- | HellaSwag | acc_norm | 0.301334 | 10 042 |
99
  | PIQA | acc_norm | 0.596844 | 1 838 |
100
- | SciQ | acc_norm | 0.664000 | 1 000 |
101
- | ARC-Easy | acc_norm | 0.434343 | 2 376 |
102
- | ARC-Challenge | acc_norm | 0.241468 | 1 172 |
103
 
104
- LAMBADA osiągnęła perplexity 49.907207. Pełne metryki i protokół zapisano w [`benchmarks/english.json`](./benchmarks/english.json).
105
 
106
  ### Użycie z Transformers
107
 
@@ -123,11 +130,12 @@ Model używa standardowego `GPT2LMHeadModel`; `trust_remote_code=True` nie jest
123
 
124
  ### Stałe próbki inferencji
125
 
126
- Wspólny zestaw [`fixed-sampling-v1`](./inference-samples/README.md) pokazuje te same cztery prompty wygenerowane przez r001-r006. Wszystkie opublikowane rewizje są ładowane z pełnych, niezmiennych SHA commitów, a SHA-256 każdego pliku z wagami jest sprawdzane przed inferencją. Referencyjny protokół używa macOS 26.2 na arm64, CPU, float32, jednego wątku, Transformers 5.15.1, seed 1337 resetowanego dla każdego promptu, temperature 0.7, top-k 50 i limitu 100 nowych tokenów.
127
 
128
  | Rewizja | Wydanie | Wagi | Historia |
129
  |---|---|---|---|
130
- | **r006** | **Pollock 1.4** | [`a2e53d9`](https://huggingface.co/SlayerLab/pollock-mini-lm-125m/commit/a2e53d9f1b690ebfbea375cf8bd3b5bdf69dece9) | [`r006.md`](./training-history/r006.md) |
 
131
  | r005 | Pollock 1.3 | [`e780025`](https://huggingface.co/SlayerLab/pollock-mini-lm-125m/commit/e780025f15e06bc3765a74d973906eb8a11c022c) | [`r005.md`](./training-history/r005.md) |
132
  | r004 | Pollock 1.2 | [`30feb81`](https://huggingface.co/SlayerLab/pollock-mini-lm-125m/commit/30feb81e5097eef13b939e1827d98e0458bf602d) | [`r004.md`](./training-history/r004.md) |
133
  | r003 | Pollock 1.1 | [`698984b`](https://huggingface.co/SlayerLab/pollock-mini-lm-125m/commit/698984b1d1b96c9b6ffaff00c7fc2e78e140e842) | [`r003.md`](./training-history/r003.md) |
@@ -151,60 +159,67 @@ Pełne teksty są przeznaczone do porównywania zachowania, nie są benchmarkiem
151
 
152
  Pollock is a small English decoder-only base language model trained from scratch as a readable educational experiment. It is based on [nanoGPT](https://github.com/karpathy/nanoGPT) and a custom byte-level BPE tokenizer. It is a completion model, not a conversational assistant.
153
 
154
- The name loosely refers to Jackson Pollock's painterly gesture: nanoGPT is the canvas on which data, configuration, and training decisions create different behavioral patterns. See [`training-history/r006.md`](./training-history/r006.md) for the complete technical record and [`CHANGELOG.md`](./CHANGELOG.md) for release-to-release changes.
155
 
156
  ### Architecture and tokenizer
157
 
158
  | Property | Value |
159
  |---|---:|
160
- | Revision / release | r006 / Pollock 1.4 |
161
  | Type | GPT-2-style decoder-only Transformer |
162
- | Layers / heads / width | 12 / 14 / 896 |
163
- | Maximum context | 1,024 tokens |
 
164
  | Vocabulary | 12,288 tokens |
165
- | nanoGPT parameters | 126,637,952 |
166
- | Total unique trainable parameters | 127,555,456 |
167
  | Tokenizer | byte-level BPE, GPT-2-style pretokenization |
168
  | Special tokens | `<|endoftext|>`, `<|im_start|>`, `<|im_end|>` |
169
 
170
- The Transformers artifact has 127,674,624 parameters, including 119,168 zero-valued compatibility bias parameters required by `GPT2LMHeadModel`. The native model was trained with `bias=False`.
 
 
 
 
171
 
172
  ### Data and training
173
 
174
- The model was trained on V2 of [`SlayerLab/minimal-en-corpus-2.5b`](https://huggingface.co/datasets/SlayerLab/minimal-en-corpus-2.5b), pinned to commit `d68d992`. It is a subjectively selected mixture of 15 English-language sources after language filtering, exact and approximate deduplication, benchmark decontamination, and additional cleanup of boilerplate and incorrectly concatenated records. R006 returns to the tokenizer used by r003-r004; it is not compatible with r005's tokenizer beyond the vocabulary size and special-token IDs.
 
 
175
 
176
  | Setting | Value |
177
  |---|---:|
178
- | Training / validation tokens | 2,689,323,439 / 5,236,486 |
179
- | Final checkpoint | update 22,003 |
180
- | Token presentations | 10,814,914,560 (approximately 4.02 epochs) |
181
- | Sequence / micro-batch per GPU | 1,024 / 12 |
182
  | Global / per-GPU accumulation | 40 / 20 micro-steps |
183
  | Effective batch | 491,520 tokens |
184
- | Optimizer | fused AdamW, betas 0.9/0.95 |
185
- | Learning rate | 4e-4 → 4e-5, cosine decay |
186
- | Warmup / weight decay / grad clip | 440 / 0.1 / 1.0 |
187
  | Precision | BF16 |
188
  | Hardware | 2× NVIDIA GeForce RTX 4090 |
189
- | Framework | PyTorch 2.8.0+cu128, nanoGPT commit `3adf61e` |
190
 
191
  ### Evaluation
192
 
193
- Training-time loss was estimated on fixed subsets of 1,228,800 tokens per split. The final checkpoint achieved validation loss **2.536357**; the best result was **2.536231** after 10,813,440,000 token presentations. These values are not directly comparable with r005 because r006 uses a different corpus, tokenizer, and validation set.
194
 
195
- Benchmarks used complete splits with `lm-evaluation-harness` 0.4.12, zero few-shot examples, batch size 8, and BF16.
196
 
197
  | Benchmark | Primary metric | Score | Samples |
198
  |---|---|---:|---:|
199
- | BLiMP | acc | 0.769164 | 67,000 |
200
- | LAMBADA | acc | 0.276926 | 5,153 |
201
- | HellaSwag | acc_norm | 0.301334 | 10,042 |
202
  | PIQA | acc_norm | 0.596844 | 1,838 |
203
- | SciQ | acc_norm | 0.664000 | 1,000 |
204
- | ARC-Easy | acc_norm | 0.434343 | 2,376 |
205
- | ARC-Challenge | acc_norm | 0.241468 | 1,172 |
206
 
207
- LAMBADA perplexity was 49.907207. Full metrics and protocol details are recorded in [`benchmarks/english.json`](./benchmarks/english.json).
208
 
209
  ### Usage
210
 
@@ -212,11 +227,12 @@ Use the Transformers example in the Polish section. The artifact uses standard `
212
 
213
  ### Fixed inference samples
214
 
215
- The shared [`fixed-sampling-v1`](./inference-samples/README.md) suite runs the same four prompts on r001-r006. All published revisions are loaded from full immutable commit SHAs, and every weight-file SHA-256 is verified before inference. The reference protocol uses macOS 26.2 on arm64, CPU float32 with one thread, Transformers 5.15.1, seed 1337 reset for every prompt, temperature 0.7, top-k 50, and a 100-new-token limit.
216
 
217
  | Revision | Release | Weights | History |
218
  |---|---|---|---|
219
- | **r006** | **Pollock 1.4** | [`a2e53d9`](https://huggingface.co/SlayerLab/pollock-mini-lm-125m/commit/a2e53d9f1b690ebfbea375cf8bd3b5bdf69dece9) | [`r006.md`](./training-history/r006.md) |
 
220
  | r005 | Pollock 1.3 | [`e780025`](https://huggingface.co/SlayerLab/pollock-mini-lm-125m/commit/e780025f15e06bc3765a74d973906eb8a11c022c) | [`r005.md`](./training-history/r005.md) |
221
  | r004 | Pollock 1.2 | [`30feb81`](https://huggingface.co/SlayerLab/pollock-mini-lm-125m/commit/30feb81e5097eef13b939e1827d98e0458bf602d) | [`r004.md`](./training-history/r004.md) |
222
  | r003 | Pollock 1.1 | [`698984b`](https://huggingface.co/SlayerLab/pollock-mini-lm-125m/commit/698984b1d1b96c9b6ffaff00c7fc2e78e140e842) | [`r003.md`](./training-history/r003.md) |
 
5
  license: other
6
  license_name: mixed-upstream-dataset-terms
7
  license_link: https://huggingface.co/SlayerLab/pollock-mini-lm-125m/blob/main/LICENSE.md
8
+ datasets: [SlayerLab/minimal-en-corpus-5b]
9
  tags: [causal-lm, gpt2, nanogpt, bpe, educational, base-model]
10
  model-index:
11
+ - name: Pollock 1.5
12
  results:
13
  - task: {type: text-generation, name: Language modeling}
14
+ dataset: {type: SlayerLab/minimal-en-corpus-5b, name: Minimal EN validation (full split), split: validation}
15
+ metrics: [{type: loss, value: 2.4585190566, name: Full-validation loss}]
16
+ - task: {type: text-generation, name: Linguistic minimal pairs}
17
+ dataset: {type: blimp, name: BLiMP}
18
+ metrics: [{type: accuracy, value: 0.7849701493, name: Accuracy}]
19
+ - task: {type: text-generation, name: Language modeling}
20
+ dataset: {type: lambada_openai, name: LAMBADA OpenAI}
21
+ metrics: [{type: accuracy, value: 0.3007956530, name: Accuracy}, {type: perplexity, value: 45.0303617702, name: Perplexity}]
22
+ - task: {type: text-generation, name: Multiple-choice completion}
23
+ dataset: {type: hellaswag, name: HellaSwag}
24
+ metrics: [{type: accuracy, value: 0.3083051185, name: Normalized accuracy}]
25
+ - task: {type: text-generation, name: Physical commonsense reasoning}
26
+ dataset: {type: piqa, name: PIQA}
27
+ metrics: [{type: accuracy, value: 0.5968443961, name: Normalized accuracy}]
28
+ - task: {type: text-generation, name: Science question answering}
29
+ dataset: {type: sciq, name: SciQ}
30
+ metrics: [{type: accuracy, value: 0.675, name: Normalized accuracy}]
31
+ - task: {type: text-generation, name: Grade-school science questions}
32
+ dataset: {type: arc_easy, name: ARC-Easy}
33
+ metrics: [{type: accuracy, value: 0.4183501684, name: Normalized accuracy}]
34
+ - task: {type: text-generation, name: Science challenge questions}
35
+ dataset: {type: arc_challenge, name: ARC-Challenge}
36
+ metrics: [{type: accuracy, value: 0.25, name: Normalized accuracy}]
37
  ---
38
 
39
+ # Pollock 1.5 — r007
40
 
41
  ![Pollock avatar](./assets/pollock-mini-lm-avatar-320.png)
42
 
 
48
 
49
  Pollock to niewielki, anglojęzyczny model bazowy typu decoder-only, wytrenowany od zera jako czytelny eksperyment edukacyjny. Implementacja bazuje na [nanoGPT](https://github.com/karpathy/nanoGPT) i własnym tokenizerze byte-level BPE. Jest to model do uzupełniania tekstu, nie asystent konwersacyjny.
50
 
51
+ Nazwa luźno nawiązuje do gestu malarskiego Jacksona Pollocka: nanoGPT jest płótnem, na którym dane, konfiguracja i decyzje treningowe tworzą różne wzorce zachowania. Pełne dane techniczne tej wersji znajdują się w [`training-history/r007.md`](./training-history/r007.md), a różnice między wydaniami w [`CHANGELOG.md`](./CHANGELOG.md).
52
 
53
  ### Architektura i tokenizer
54
 
55
  | Właściwość | Wartość |
56
  |---|---:|
57
+ | Rewizja / wydanie | r007 / Pollock 1.5 |
58
  | Typ | decoder-only Transformer w stylu GPT-2 |
59
+ | Warstwy / głowy / embedding | 16 / 12 / 768 |
60
+ | Szerokość MLP | 3 200 |
61
+ | Maksymalny kontekst | 2048 tokenów |
62
  | Słownik | 12 288 tokenów |
63
+ | Parametry nanoGPT | 125 854 464 |
64
+ | Łączne unikalne parametry trenowalne | 127 427 328 |
65
  | Tokenizer | byte-level BPE, pretokenizacja w stylu GPT-2 |
66
  | Tokeny specjalne | <code>&lt;&#124;endoftext&#124;&gt;</code>, <code>&lt;&#124;im_start&#124;&gt;</code>, <code>&lt;&#124;im_end&#124;&gt;</code> |
67
 
68
+ Artefakt Transformers ma 127 565 312 parametrów, w tym 137 984 zerowe parametry bias dla zgodności z `GPT2LMHeadModel`. Natywny model był trenowany z `bias=False`.
69
+
70
+ Wybrany checkpoint przekonwertowano przy użyciu Transformers 5.15.1. Deterministyczny test `[2, 64]` wykazał maksymalną bezwzględną różnicę logits **0.0**; SHA-256 pliku `model.safetensors` to `4a88432225dac6255ae1e2a085019f36dcdaf55fa33c77629dba6e7445c6c2fb`.
71
+
72
+ Muon optymalizuje kwalifikujące się ukryte macierze attention i MLP. AdamW obsługuje embeddingi tokenów, powiązaną głowicę wyjściową, parametry normalizacji i biasy. Podział parametrów jest sprawdzany przed treningiem pod kątem nakładania się i kompletności.
73
 
74
  ### Dane i trening
75
 
76
+ Model wytrenowano na [`SlayerLab/minimal-en-corpus-5b`](https://huggingface.co/datasets/SlayerLab/minimal-en-corpus-5b) przypiętym do commita `38bebbd`. Jest to subiektywnie dobrana mieszanka 15 anglojęzycznych źródeł po filtrowaniu języka, deduplikacji dokładnej i przybliżonej oraz decontaminacji benchmarków. R007 używa tokenizera z tego korpusu; zachowuje rozmiar słownika i identyfikatory tokenów specjalnych, ale pozostałe identyfikatory tokenów nie są zgodne z tokenizerem r006.
77
+
78
+ Loader stosuje globalną permutację ze stałym seedem do niezmiennych okien tokenów w każdym przebiegu, przydziela rozłączne pozycje między rangi DDP, maskuje końcowy niepełny batch i zapisuje globalny kursor, topologię oraz ustawienia tasowania w każdym checkpoincie.
79
 
80
  | Parametr | Wartość |
81
  |---|---:|
82
+ | Tokeny treningowe / walidacyjne | 5 396 605 407 / 5 238 223 |
83
+ | Wybrany checkpoint | aktualizacja 43 750 |
84
+ | Przetworzone tokeny | 21 503 919 992 (około 3,98 epoki) |
85
+ | Sekwencja / micro-batch na GPU | 2048 / 6 |
86
  | Akumulacja globalna / na GPU | 40 / 20 micro-stepów |
87
  | Effective batch | 491 520 tokenów |
88
+ | Optymalizator | Muon plus fused AdamW, betas AdamW 0.9/0.95 |
89
+ | Learning rate | Muon 1e-2 → 1e-3; AdamW 3e-4 → 3e-5 |
90
+ | Warmup / weight decay / grad clip | 200 / 0.01 / 1.0 |
91
  | Precyzja | BF16 |
92
  | Sprzęt | 2× NVIDIA GeForce RTX 4090 |
93
+ | Framework | PyTorch 2.11.0+cu128, nanoGPT commit `3adf61e` |
94
 
95
  ### Ewaluacja
96
 
97
+ Loss treningowy szacowano na stałej, tasowanej próbie 2 457 600 tokenów z całego splitu treningowego. Każda ewaluacja walidacyjna obejmowała pełny split, czyli 5 238 222 targety. Wybrany checkpoint zapisał validation loss **2.458616** po 21 503 919 992 przetworzonych tokenach. Niezależna ewaluacja pełnego splitu odtworzyła cross-entropy **2.458519**, perplexity **11.687490** i predictive entropy **2.437529**. Trening zakończył się po dokładnie czterech przebiegach korpusu, na aktualizacji 43 918 i po 21 586 421 624 przetworzonych tokenach.
98
 
99
+ Benchmarki wykonano zero-shot na pełnych splitach przy użyciu `lm-evaluation-harness` 0.4.12, batch size 8, maksymalnego kontekstu 1024 i BF16 na Apple M1 Max MPS.
100
 
101
  | Benchmark | Główna metryka | Wynik | Próbki |
102
  |---|---|---:|---:|
103
+ | BLiMP | acc | 0.784970 | 67 000 |
104
+ | LAMBADA | acc | 0.300796 | 5 153 |
105
+ | HellaSwag | acc_norm | 0.308305 | 10 042 |
106
  | PIQA | acc_norm | 0.596844 | 1 838 |
107
+ | SciQ | acc_norm | 0.675000 | 1 000 |
108
+ | ARC-Easy | acc_norm | 0.418350 | 2 376 |
109
+ | ARC-Challenge | acc_norm | 0.250000 | 1 172 |
110
 
111
+ Perplexity LAMBADA wynosi **45.030362**. W porównaniu z pomiarem E2 checkpoint E4-best poprawia BLiMP i ARC-Challenge, ale uzyskuje niższe wyniki na pozostałych pięciu zadaniach. Spadek cross-entropy walidacyjnej nie przekłada się więc jednolicie na te benchmarki downstream. Pełne metryki i protokół znajdują się w [`benchmarks/english.json`](./benchmarks/english.json).
112
 
113
  ### Użycie z Transformers
114
 
 
130
 
131
  ### Stałe próbki inferencji
132
 
133
+ Wspólny zestaw [`fixed-sampling-v1`](./inference-samples/README.md) pokazuje te same cztery prompty wygenerowane przez r001-r007. Wszystkie opublikowane rewizje są ładowane z pełnych, niezmiennych SHA commitów, a SHA-256 każdego pliku z wagami jest sprawdzane przed inferencją. Referencyjny protokół używa macOS 26.2 na arm64, CPU, float32, jednego wątku, Transformers 5.15.1, seed 1337 resetowanego dla każdego promptu, temperature 0.7, top-k 50 i limitu 100 nowych tokenów.
134
 
135
  | Rewizja | Wydanie | Wagi | Historia |
136
  |---|---|---|---|
137
+ | **r007** | **Pollock 1.5** | [`f2a1ec0`](https://huggingface.co/SlayerLab/pollock-mini-lm-125m/commit/f2a1ec072098d257c36b129b0e554e2e6de9c46e) | [`r007.md`](./training-history/r007.md) |
138
+ | r006 | Pollock 1.4 | [`a2e53d9`](https://huggingface.co/SlayerLab/pollock-mini-lm-125m/commit/a2e53d9f1b690ebfbea375cf8bd3b5bdf69dece9) | [`r006.md`](./training-history/r006.md) |
139
  | r005 | Pollock 1.3 | [`e780025`](https://huggingface.co/SlayerLab/pollock-mini-lm-125m/commit/e780025f15e06bc3765a74d973906eb8a11c022c) | [`r005.md`](./training-history/r005.md) |
140
  | r004 | Pollock 1.2 | [`30feb81`](https://huggingface.co/SlayerLab/pollock-mini-lm-125m/commit/30feb81e5097eef13b939e1827d98e0458bf602d) | [`r004.md`](./training-history/r004.md) |
141
  | r003 | Pollock 1.1 | [`698984b`](https://huggingface.co/SlayerLab/pollock-mini-lm-125m/commit/698984b1d1b96c9b6ffaff00c7fc2e78e140e842) | [`r003.md`](./training-history/r003.md) |
 
159
 
160
  Pollock is a small English decoder-only base language model trained from scratch as a readable educational experiment. It is based on [nanoGPT](https://github.com/karpathy/nanoGPT) and a custom byte-level BPE tokenizer. It is a completion model, not a conversational assistant.
161
 
162
+ The name loosely refers to Jackson Pollock's painterly gesture: nanoGPT is the canvas on which data, configuration, and training decisions create different behavioral patterns. See [`training-history/r007.md`](./training-history/r007.md) for the complete technical record and [`CHANGELOG.md`](./CHANGELOG.md) for release-to-release changes.
163
 
164
  ### Architecture and tokenizer
165
 
166
  | Property | Value |
167
  |---|---:|
168
+ | Revision / release | r007 / Pollock 1.5 |
169
  | Type | GPT-2-style decoder-only Transformer |
170
+ | Layers / heads / width | 16 / 12 / 768 |
171
+ | MLP width | 3,200 |
172
+ | Maximum context | 2,048 tokens |
173
  | Vocabulary | 12,288 tokens |
174
+ | nanoGPT parameters | 125,854,464 |
175
+ | Total unique trainable parameters | 127,427,328 |
176
  | Tokenizer | byte-level BPE, GPT-2-style pretokenization |
177
  | Special tokens | `<|endoftext|>`, `<|im_start|>`, `<|im_end|>` |
178
 
179
+ The Transformers artifact has 127,565,312 parameters, including 137,984 zero-valued compatibility bias parameters required by `GPT2LMHeadModel`. The native model was trained with `bias=False`.
180
+
181
+ The selected checkpoint was converted with Transformers 5.15.1. A deterministic `[2, 64]` probe produced maximum absolute logit error **0.0**; the `model.safetensors` SHA-256 is `4a88432225dac6255ae1e2a085019f36dcdaf55fa33c77629dba6e7445c6c2fb`.
182
+
183
+ Muon optimizes eligible hidden attention and MLP matrices. AdamW handles token embeddings, the tied output head, normalization parameters, and biases. The parameter partition is checked for overlap and completeness before training.
184
 
185
  ### Data and training
186
 
187
+ The model was trained on [`SlayerLab/minimal-en-corpus-5b`](https://huggingface.co/datasets/SlayerLab/minimal-en-corpus-5b), pinned to commit `38bebbd`. It is a subjectively selected mixture of 15 English-language sources after language filtering, exact and approximate deduplication, and benchmark decontamination. R007 uses the tokenizer from this corpus; it retains the vocabulary size and special-token IDs, but other token IDs are not compatible with r006's tokenizer.
188
+
189
+ The loader applies a seeded global permutation to the fixed token windows in each pass, assigns disjoint positions across DDP ranks, masks the final partial batch, and stores its global cursor, topology, and shuffle settings in every checkpoint.
190
 
191
  | Setting | Value |
192
  |---|---:|
193
+ | Training / validation tokens | 5,396,605,407 / 5,238,223 |
194
+ | Selected checkpoint | update 43,750 |
195
+ | Token presentations | 21,503,919,992 (approximately 3.98 epochs) |
196
+ | Sequence / micro-batch per GPU | 2,048 / 6 |
197
  | Global / per-GPU accumulation | 40 / 20 micro-steps |
198
  | Effective batch | 491,520 tokens |
199
+ | Optimizer | Muon plus fused AdamW, AdamW betas 0.9/0.95 |
200
+ | Learning rate | Muon 1e-2 → 1e-3; AdamW 3e-4 → 3e-5 |
201
+ | Warmup / weight decay / grad clip | 200 / 0.01 / 1.0 |
202
  | Precision | BF16 |
203
  | Hardware | 2× NVIDIA GeForce RTX 4090 |
204
+ | Framework | PyTorch 2.11.0+cu128, nanoGPT commit `3adf61e` |
205
 
206
  ### Evaluation
207
 
208
+ Training loss was estimated on a fixed shuffled sample of 2,457,600 tokens from the complete training split. Every validation evaluation covered the full split, or 5,238,222 targets. The selected checkpoint recorded validation loss **2.458616** after 21,503,919,992 token presentations. An independent full-split evaluation reproduced cross-entropy **2.458519**, perplexity **11.687490**, and predictive entropy **2.437529**. Training ended after exactly four corpus passes, at update 43,918 and 21,586,421,624 token presentations.
209
 
210
+ Benchmarks used complete splits with `lm-evaluation-harness` 0.4.12, zero few-shot examples, batch size 8, maximum context 1024, and BF16 on Apple M1 Max MPS.
211
 
212
  | Benchmark | Primary metric | Score | Samples |
213
  |---|---|---:|---:|
214
+ | BLiMP | acc | 0.784970 | 67,000 |
215
+ | LAMBADA | acc | 0.300796 | 5,153 |
216
+ | HellaSwag | acc_norm | 0.308305 | 10,042 |
217
  | PIQA | acc_norm | 0.596844 | 1,838 |
218
+ | SciQ | acc_norm | 0.675000 | 1,000 |
219
+ | ARC-Easy | acc_norm | 0.418350 | 2,376 |
220
+ | ARC-Challenge | acc_norm | 0.250000 | 1,172 |
221
 
222
+ LAMBADA perplexity is **45.030362**. Compared with the E2 measurement, the E4-best checkpoint improves BLiMP and ARC-Challenge but scores lower on the other five tasks. The validation cross-entropy gain therefore does not transfer uniformly to these downstream benchmarks. Full metrics and protocol details are recorded in [`benchmarks/english.json`](./benchmarks/english.json).
223
 
224
  ### Usage
225
 
 
227
 
228
  ### Fixed inference samples
229
 
230
+ The shared [`fixed-sampling-v1`](./inference-samples/README.md) suite runs the same four prompts on r001-r007. All published revisions are loaded from full immutable commit SHAs, and every weight-file SHA-256 is verified before inference. The reference protocol uses macOS 26.2 on arm64, CPU float32 with one thread, Transformers 5.15.1, seed 1337 reset for every prompt, temperature 0.7, top-k 50, and a 100-new-token limit.
231
 
232
  | Revision | Release | Weights | History |
233
  |---|---|---|---|
234
+ | **r007** | **Pollock 1.5** | [`f2a1ec0`](https://huggingface.co/SlayerLab/pollock-mini-lm-125m/commit/f2a1ec072098d257c36b129b0e554e2e6de9c46e) | [`r007.md`](./training-history/r007.md) |
235
+ | r006 | Pollock 1.4 | [`a2e53d9`](https://huggingface.co/SlayerLab/pollock-mini-lm-125m/commit/a2e53d9f1b690ebfbea375cf8bd3b5bdf69dece9) | [`r006.md`](./training-history/r006.md) |
236
  | r005 | Pollock 1.3 | [`e780025`](https://huggingface.co/SlayerLab/pollock-mini-lm-125m/commit/e780025f15e06bc3765a74d973906eb8a11c022c) | [`r005.md`](./training-history/r005.md) |
237
  | r004 | Pollock 1.2 | [`30feb81`](https://huggingface.co/SlayerLab/pollock-mini-lm-125m/commit/30feb81e5097eef13b939e1827d98e0458bf602d) | [`r004.md`](./training-history/r004.md) |
238
  | r003 | Pollock 1.1 | [`698984b`](https://huggingface.co/SlayerLab/pollock-mini-lm-125m/commit/698984b1d1b96c9b6ffaff00c7fc2e78e140e842) | [`r003.md`](./training-history/r003.md) |
benchmarks/english.json CHANGED
@@ -1,72 +1,76 @@
1
  {
2
  "schema_version": 1,
3
- "revision_id": "r006",
4
- "release": "Pollock 1.4",
5
- "source_result": "runs/r006-corrected-corpus-v2-s1337-lr0.0004/benchmarks/results.json",
6
  "checkpoint": {
7
- "path": "runs/r006-corrected-corpus-v2-s1337-lr0.0004/checkpoints/ckpt-final.pt",
8
- "sha256": "580fbb3c039ea1f349285a9bf7a67731fb6d455de931b6b7601d4ed09fce9437",
9
- "source_iteration": 22003
 
10
  },
11
  "benchmark_checkpoint": {
12
- "path": "/workspace/cache/dmpod-tokenizers/r006-final-benchmark.pt",
13
- "sha256": "9339d20ea32d24bc9777518a2f0ec553e4a3871647d342db92d7c5b6c18c7b56",
14
- "trusted_input_required": true
15
  },
16
  "tokenizer": {
17
- "dataset_revision": "d68d992622e9fc11f19e7d7fb8547c4e653439a4",
18
- "sha256": "6cda4e5ec8293f3b821e02253f9c0e88e43ad4f7c6e1324763e1fb91749fef51",
19
- "version": "V2.0@d68d992622e9fc11f19e7d7fb8547c4e653439a4"
20
  },
21
  "execution": {
22
  "harness": "lm-evaluation-harness 0.4.12",
 
 
 
23
  "num_fewshot": 0,
24
  "batch_size": 8,
25
  "precision": "bfloat16",
26
  "max_length": 1024,
27
  "limit_per_benchmark": null,
28
- "truncated_requests": 0
 
29
  },
30
  "results": {
31
  "blimp": {
32
- "acc,none": 0.7691641791044775,
33
- "sample_len": 67000.0,
34
  "samples": 67000
35
  },
36
  "lambada_openai": {
37
- "acc,none": 0.27692606248787116,
38
- "perplexity,none": 49.907206551888116,
39
- "sample_len": 5153.0,
40
  "samples": 5153
41
  },
42
  "hellaswag": {
43
- "acc,none": 0.2830113523202549,
44
- "acc_norm,none": 0.3013343955387373,
45
- "sample_len": 10042.0,
46
  "samples": 10042
47
  },
48
  "piqa": {
49
- "acc,none": 0.6071817192600653,
50
  "acc_norm,none": 0.5968443960826986,
51
- "sample_len": 1838.0,
52
  "samples": 1838
53
  },
54
  "sciq": {
55
- "acc,none": 0.768,
56
- "acc_norm,none": 0.664,
57
- "sample_len": 1000.0,
58
  "samples": 1000
59
  },
60
  "arc_easy": {
61
- "acc,none": 0.4877946127946128,
62
- "acc_norm,none": 0.43434343434343436,
63
- "sample_len": 2376.0,
64
  "samples": 2376
65
  },
66
  "arc_challenge": {
67
- "acc,none": 0.21075085324232082,
68
- "acc_norm,none": 0.24146757679180889,
69
- "sample_len": 1172.0,
70
  "samples": 1172
71
  }
72
  }
 
1
  {
2
  "schema_version": 1,
3
+ "revision_id": "r007",
4
+ "source_result": "out-pollock-r007-e4-final/benchmarks/full/release-draft/results_2026-09-20T13-15-46.016579.json",
5
+ "source_result_sha256": "bfbda3e30e936d975d5ab110be17067b418048c10beb6fe5a4e8e760aa8a258c",
6
  "checkpoint": {
7
+ "path": "out-pollock-r007-e4-final/ckpt-best.pt",
8
+ "sha256": "384d912cc8b7ebf872b66d1f8ad46f223296a069e906167076f1899296f9a788",
9
+ "source_iteration": 43750,
10
+ "tokens_processed": 21503919992
11
  },
12
  "benchmark_checkpoint": {
13
+ "path": "release-draft/model.safetensors",
14
+ "sha256": "4a88432225dac6255ae1e2a085019f36dcdaf55fa33c77629dba6e7445c6c2fb",
15
+ "conversion_max_absolute_logit_delta": 0.0
16
  },
17
  "tokenizer": {
18
+ "dataset_revision": "38bebbd2ae36858c25aa0799fd3296fbfa9432aa",
19
+ "sha256": "3733307577230bb4802d2c774d8e2323f7f64e4139c13712736a57daf91bdda1"
 
20
  },
21
  "execution": {
22
  "harness": "lm-evaluation-harness 0.4.12",
23
+ "transformers": "5.15.1",
24
+ "torch": "2.14.0",
25
+ "device": "Apple M1 Max MPS",
26
  "num_fewshot": 0,
27
  "batch_size": 8,
28
  "precision": "bfloat16",
29
  "max_length": 1024,
30
  "limit_per_benchmark": null,
31
+ "truncated_requests": 0,
32
+ "evaluation_time_seconds": 1552.1888507499825
33
  },
34
  "results": {
35
  "blimp": {
36
+ "acc,none": 0.7849701492537313,
37
+ "sample_len": 67000,
38
  "samples": 67000
39
  },
40
  "lambada_openai": {
41
+ "acc,none": 0.3007956530176596,
42
+ "perplexity,none": 45.030361770214945,
43
+ "sample_len": 5153,
44
  "samples": 5153
45
  },
46
  "hellaswag": {
47
+ "acc,none": 0.2876916948814977,
48
+ "acc_norm,none": 0.30830511850229037,
49
+ "sample_len": 10042,
50
  "samples": 10042
51
  },
52
  "piqa": {
53
+ "acc,none": 0.6077257889009793,
54
  "acc_norm,none": 0.5968443960826986,
55
+ "sample_len": 1838,
56
  "samples": 1838
57
  },
58
  "sciq": {
59
+ "acc,none": 0.766,
60
+ "acc_norm,none": 0.675,
61
+ "sample_len": 1000,
62
  "samples": 1000
63
  },
64
  "arc_easy": {
65
+ "acc,none": 0.4772727272727273,
66
+ "acc_norm,none": 0.41835016835016836,
67
+ "sample_len": 2376,
68
  "samples": 2376
69
  },
70
  "arc_challenge": {
71
+ "acc,none": 0.2158703071672355,
72
+ "acc_norm,none": 0.25,
73
+ "sample_len": 1172,
74
  "samples": 1172
75
  }
76
  }
config.json CHANGED
@@ -12,13 +12,12 @@
12
  "initializer_range": 0.02,
13
  "layer_norm_epsilon": 1e-05,
14
  "model_type": "gpt2",
15
- "model_version": "1.4",
16
- "n_ctx": 1024,
17
- "n_embd": 896,
18
- "n_head": 14,
19
- "n_inner": null,
20
- "n_layer": 12,
21
- "n_positions": 1024,
22
  "pad_token_id": 12285,
23
  "reorder_and_upcast_attn": false,
24
  "resid_pdrop": 0.0,
 
12
  "initializer_range": 0.02,
13
  "layer_norm_epsilon": 1e-05,
14
  "model_type": "gpt2",
15
+ "n_ctx": 2048,
16
+ "n_embd": 768,
17
+ "n_head": 12,
18
+ "n_inner": 3200,
19
+ "n_layer": 16,
20
+ "n_positions": 2048,
 
21
  "pad_token_id": 12285,
22
  "reorder_and_upcast_attn": false,
23
  "resid_pdrop": 0.0,
conversion_manifest.json ADDED
@@ -0,0 +1,28 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "checkpoint": "out-pollock-r007-e4-final/ckpt-best.pt",
3
+ "checkpoint_sha256": "384d912cc8b7ebf872b66d1f8ad46f223296a069e906167076f1899296f9a788",
4
+ "tokenizer": "data/minimal_en_corpus_5b/tokenizer/tokenizer.json",
5
+ "tokenizer_sha256": "3733307577230bb4802d2c774d8e2323f7f64e4139c13712736a57daf91bdda1",
6
+ "model_sha256": "4a88432225dac6255ae1e2a085019f36dcdaf55fa33c77629dba6e7445c6c2fb",
7
+ "source_iteration": 43750,
8
+ "tokens_processed": 21503919992,
9
+ "model_args": {
10
+ "n_layer": 16,
11
+ "n_head": 12,
12
+ "n_embd": 768,
13
+ "n_inner": 3200,
14
+ "block_size": 2048,
15
+ "bias": false,
16
+ "vocab_size": 12288,
17
+ "dropout": 0.0
18
+ },
19
+ "transformers_version": "5.15.1",
20
+ "verification": {
21
+ "seed": 1337,
22
+ "input_shape": [
23
+ 2,
24
+ 64
25
+ ],
26
+ "max_absolute_logit_delta": 0.0
27
+ }
28
+ }
inference-samples/README.md CHANGED
@@ -24,6 +24,7 @@ Exact token replay was verified in the recorded environment. Sampling may diverg
24
  | r004 | Pollock 1.2 | [`30feb81`](https://huggingface.co/SlayerLab/pollock-mini-lm-125m/commit/30feb81e5097eef13b939e1827d98e0458bf602d) | `3c453f22d4bb70228e5cca79e183a425e0f0221e54bd782d12027c42c943c880` |
25
  | r005 | Pollock 1.3 | [`e780025`](https://huggingface.co/SlayerLab/pollock-mini-lm-125m/commit/e780025f15e06bc3765a74d973906eb8a11c022c) | `3db43dfa44622e8156f060fee6f16875e108db4f5f949969cb91c0ca8cdce32c` |
26
  | r006 | Pollock 1.4 | [`a2e53d9`](https://huggingface.co/SlayerLab/pollock-mini-lm-125m/commit/a2e53d9f1b690ebfbea375cf8bd3b5bdf69dece9) | `8e74aa4d34464229e86a1831fb34d442f91e22aa600e0cf2503891957fdc6a43` |
 
27
 
28
  ## Results
29
 
@@ -120,6 +121,22 @@ The municipality of Aisne is situated on the banks of the Neuve-Brieu river, whi
120
 
121
  Output SHA-256: `b47fcb42354b8f735c7b98cfdf93605c6e095ed8290e2ddb3329886df244e99f`
122
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
123
  ### Explanation
124
 
125
  Prompt: `Photosynthesis is the process by which`
@@ -197,6 +214,19 @@ Carbon gas: A mixture of gases in which carbon dioxide is replaced by a gas; in
197
 
198
  Output SHA-256: `c09d560713828851410d37fcd7c648d49c2bedf6e9c8325bb47cb32a8015bfd4`
199
 
 
 
 
 
 
 
 
 
 
 
 
 
 
200
  ### Story
201
 
202
  Prompt: `In a distant future, humanity discovered`
@@ -276,6 +306,20 @@ This is fascinating! Can you tell me more about the specific cultural values tha
276
 
277
  Output SHA-256: `6455cdbe15a71ae11ff476739071f31b3f86b8bf4d97f3dabe99a7d1595c5c7c`
278
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
279
  ### Code
280
 
281
  Prompt: `def fibonacci(n):`
@@ -408,6 +452,27 @@ def fibonacci(n):
408
 
409
  Output SHA-256: `ad3aa2ea2cab9a46ae5f52a9cc109bcf5f7362e1d214ceb791ee2dd6995a345f`
410
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
411
  ## Reproduce or verify
412
 
413
  Run from this directory. The check downloads approximately 2.4 GB of published model artifacts if they are not already cached.
 
24
  | r004 | Pollock 1.2 | [`30feb81`](https://huggingface.co/SlayerLab/pollock-mini-lm-125m/commit/30feb81e5097eef13b939e1827d98e0458bf602d) | `3c453f22d4bb70228e5cca79e183a425e0f0221e54bd782d12027c42c943c880` |
25
  | r005 | Pollock 1.3 | [`e780025`](https://huggingface.co/SlayerLab/pollock-mini-lm-125m/commit/e780025f15e06bc3765a74d973906eb8a11c022c) | `3db43dfa44622e8156f060fee6f16875e108db4f5f949969cb91c0ca8cdce32c` |
26
  | r006 | Pollock 1.4 | [`a2e53d9`](https://huggingface.co/SlayerLab/pollock-mini-lm-125m/commit/a2e53d9f1b690ebfbea375cf8bd3b5bdf69dece9) | `8e74aa4d34464229e86a1831fb34d442f91e22aa600e0cf2503891957fdc6a43` |
27
+ | r007 | Pollock 1.5 | [`f2a1ec0`](https://huggingface.co/SlayerLab/pollock-mini-lm-125m/commit/f2a1ec072098d257c36b129b0e554e2e6de9c46e) | `4a88432225dac6255ae1e2a085019f36dcdaf55fa33c77629dba6e7445c6c2fb` |
28
 
29
  ## Results
30
 
 
121
 
122
  Output SHA-256: `b47fcb42354b8f735c7b98cfdf93605c6e095ed8290e2ddb3329886df244e99f`
123
 
124
+ #### r007 - Pollock 1.5
125
+
126
+ ````text
127
+ The capital of France is the port city of Château-Thierry.
128
+
129
+ The capital of France is the port city of Mont-Royal.
130
+
131
+ History
132
+
133
+ The city of Mont-Royal was first mentioned by the French colonial administration in the 15th century. When the French occupation began in 1793, the city was captured by the French army and the French were forced to surrender to the British.
134
+
135
+ The city is thought to have been the first port city in
136
+ ````
137
+
138
+ Output SHA-256: `4b8907f330be8f045959ddb3ec1c50c88c82a1884c219aee6d3c6537a442ca84`
139
+
140
  ### Explanation
141
 
142
  Prompt: `Photosynthesis is the process by which`
 
214
 
215
  Output SHA-256: `c09d560713828851410d37fcd7c648d49c2bedf6e9c8325bb47cb32a8015bfd4`
216
 
217
+ #### r007 - Pollock 1.5
218
+
219
+ ````text
220
+ Photosynthesis is the process by which plants and animals use energy from sunlight and other sources to synthesize proteins and carbohydrates. The process is vital to survival and growth of plants and animals.
221
+
222
+ 2. Energy Capture in Plants:
223
+
224
+ Plants are the primary source of energy for the Earth's ecosystems. The ability of plants to capture solar radiation is essential for their survival. Plants also convert atmospheric carbon dioxide into sugars, which are used by plants to build their cell walls.
225
+
226
+ ````
227
+
228
+ Output SHA-256: `a8457bf7fa743d8397b5a221f4f11c5e0b978372eca68f068c17052307853591`
229
+
230
  ### Story
231
 
232
  Prompt: `In a distant future, humanity discovered`
 
306
 
307
  Output SHA-256: `6455cdbe15a71ae11ff476739071f31b3f86b8bf4d97f3dabe99a7d1595c5c7c`
308
 
309
+ #### r007 - Pollock 1.5
310
+
311
+ ````text
312
+ In a distant future, humanity discovered the potential of the moon to guide humanity in their quest for knowledge and understanding.
313
+ user
314
+ Can you tell me more about the moon and its impact on humanity?
315
+ assistant
316
+ Sure! The moon is a natural feature of the Earth's surface and is a crucial part of the natural landscape of the Earth. Here are some ways in which the moon impacts human life:
317
+
318
+ 1. Earth's gravitational pull: The moon has a significant impact on human life. It
319
+ ````
320
+
321
+ Output SHA-256: `00fb4213a5eeb36d18cd947a0876ad711f00bee4a8f576d75fd551d66183ba82`
322
+
323
  ### Code
324
 
325
  Prompt: `def fibonacci(n):`
 
452
 
453
  Output SHA-256: `ad3aa2ea2cab9a46ae5f52a9cc109bcf5f7362e1d214ceb791ee2dd6995a345f`
454
 
455
+ #### r007 - Pollock 1.5
456
+
457
+ ````text
458
+ def fibonacci(n):
459
+ if n < 1:
460
+ return 0
461
+
462
+ if fibonacci(n):
463
+ return fibonacci(n-1)
464
+
465
+ return fibonacci(n-1) + fibonacci(n-2)
466
+
467
+ print fibonacci(0) + fibonacci(0)
468
+
469
+ return fibonacci(0) + fibonacci(0) + fibonacci(1)
470
+
471
+ print fib
472
+ ````
473
+
474
+ Output SHA-256: `012a9d1312d787e5e759f1af7a242f20f4c9fdfc895ae673b0889fd08a0e8fe1`
475
+
476
  ## Reproduce or verify
477
 
478
  Run from this directory. The check downloads approximately 2.4 GB of published model artifacts if they are not already cached.
inference-samples/config.json CHANGED
@@ -60,6 +60,12 @@
60
  "release": "Pollock 1.4",
61
  "commit": "a2e53d9f1b690ebfbea375cf8bd3b5bdf69dece9",
62
  "model_sha256": "8e74aa4d34464229e86a1831fb34d442f91e22aa600e0cf2503891957fdc6a43"
 
 
 
 
 
 
63
  }
64
  ],
65
  "prompts": [
 
60
  "release": "Pollock 1.4",
61
  "commit": "a2e53d9f1b690ebfbea375cf8bd3b5bdf69dece9",
62
  "model_sha256": "8e74aa4d34464229e86a1831fb34d442f91e22aa600e0cf2503891957fdc6a43"
63
+ },
64
+ {
65
+ "revision_id": "r007",
66
+ "release": "Pollock 1.5",
67
+ "commit": "f2a1ec072098d257c36b129b0e554e2e6de9c46e",
68
+ "model_sha256": "4a88432225dac6255ae1e2a085019f36dcdaf55fa33c77629dba6e7445c6c2fb"
69
  }
70
  ],
71
  "prompts": [
inference-samples/generate.py CHANGED
@@ -218,6 +218,9 @@ def render_markdown(config: dict[str, Any], results: dict[str, Any]) -> str:
218
  protocol = results["protocol"]
219
  environment = results["environment"]
220
  revisions = {item["revision_id"]: item for item in results["revisions"]}
 
 
 
221
  lines = [
222
  "# Fixed cross-revision inference samples",
223
  "",
@@ -250,6 +253,15 @@ def render_markdown(config: dict[str, Any], results: dict[str, Any]) -> str:
250
  "diverge on another operating system or CPU architecture even when package "
251
  "versions and seeds match.",
252
  "",
 
 
 
 
 
 
 
 
 
253
  "## Revisions",
254
  "",
255
  "| Revision | Release | Immutable weights | `model.safetensors` SHA-256 |",
@@ -257,8 +269,11 @@ def render_markdown(config: dict[str, Any], results: dict[str, Any]) -> str:
257
  ]
258
  for revision in results["revisions"]:
259
  commit = revision["commit"]
260
- commit_url = f"https://huggingface.co/{results['model_id']}/commit/{commit}"
261
- weights = f"[`{commit[:7]}`]({commit_url})"
 
 
 
262
  lines.append(
263
  f"| {revision['revision_id']} | {revision['release']} | "
264
  f"{weights} | `{revision['model_sha256']}` |"
 
218
  protocol = results["protocol"]
219
  environment = results["environment"]
220
  revisions = {item["revision_id"]: item for item in results["revisions"]}
221
+ has_pending_commit = any(
222
+ item["commit"].startswith("__PENDING_") for item in results["revisions"]
223
+ )
224
  lines = [
225
  "# Fixed cross-revision inference samples",
226
  "",
 
253
  "diverge on another operating system or CPU architecture even when package "
254
  "versions and seeds match.",
255
  "",
256
+ *(
257
+ [
258
+ "The r007 outputs currently use the local artifact whose weight hash is "
259
+ "shown below. Replay from its immutable publication commit remains pending.",
260
+ "",
261
+ ]
262
+ if has_pending_commit
263
+ else []
264
+ ),
265
  "## Revisions",
266
  "",
267
  "| Revision | Release | Immutable weights | `model.safetensors` SHA-256 |",
 
269
  ]
270
  for revision in results["revisions"]:
271
  commit = revision["commit"]
272
+ if commit.startswith("__PENDING_"):
273
+ weights = "local artifact; immutable commit pending"
274
+ else:
275
+ commit_url = f"https://huggingface.co/{results['model_id']}/commit/{commit}"
276
+ weights = f"[`{commit[:7]}`]({commit_url})"
277
  lines.append(
278
  f"| {revision['revision_id']} | {revision['release']} | "
279
  f"{weights} | `{revision['model_sha256']}` |"
inference-samples/results.json CHANGED
@@ -2945,6 +2945,498 @@
2945
  "full_text_sha256": "ad3aa2ea2cab9a46ae5f52a9cc109bcf5f7362e1d214ceb791ee2dd6995a345f"
2946
  }
2947
  ]
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2948
  }
2949
  ]
2950
  }
 
2945
  "full_text_sha256": "ad3aa2ea2cab9a46ae5f52a9cc109bcf5f7362e1d214ceb791ee2dd6995a345f"
2946
  }
2947
  ]
2948
+ },
2949
+ {
2950
+ "revision_id": "r007",
2951
+ "release": "Pollock 1.5",
2952
+ "commit": "f2a1ec072098d257c36b129b0e554e2e6de9c46e",
2953
+ "model_sha256": "4a88432225dac6255ae1e2a085019f36dcdaf55fa33c77629dba6e7445c6c2fb",
2954
+ "loaded_model_version": null,
2955
+ "results": [
2956
+ {
2957
+ "prompt_id": "general_knowledge",
2958
+ "prompt": "The capital of France is",
2959
+ "prompt_token_ids": [
2960
+ 489,
2961
+ 4496,
2962
+ 286,
2963
+ 4392,
2964
+ 311
2965
+ ],
2966
+ "completion_token_ids": [
2967
+ 263,
2968
+ 2242,
2969
+ 2405,
2970
+ 286,
2971
+ 689,
2972
+ 11913,
2973
+ 690,
2974
+ 2183,
2975
+ 45,
2976
+ 980,
2977
+ 1291,
2978
+ 595,
2979
+ 46,
2980
+ 10,
2981
+ 10,
2982
+ 489,
2983
+ 4496,
2984
+ 286,
2985
+ 4392,
2986
+ 311,
2987
+ 263,
2988
+ 2242,
2989
+ 2405,
2990
+ 286,
2991
+ 6334,
2992
+ 45,
2993
+ 82,
2994
+ 4792,
2995
+ 46,
2996
+ 10,
2997
+ 10,
2998
+ 8231,
2999
+ 10,
3000
+ 10,
3001
+ 489,
3002
+ 2405,
3003
+ 286,
3004
+ 6334,
3005
+ 45,
3006
+ 82,
3007
+ 4792,
3008
+ 399,
3009
+ 770,
3010
+ 4888,
3011
+ 410,
3012
+ 263,
3013
+ 3369,
3014
+ 4683,
3015
+ 460,
3016
+ 7484,
3017
+ 287,
3018
+ 263,
3019
+ 1487,
3020
+ 373,
3021
+ 2882,
3022
+ 46,
3023
+ 1769,
3024
+ 263,
3025
+ 3369,
3026
+ 4727,
3027
+ 328,
3028
+ 2716,
3029
+ 287,
3030
+ 1430,
3031
+ 5091,
3032
+ 44,
3033
+ 263,
3034
+ 2405,
3035
+ 399,
3036
+ 5533,
3037
+ 410,
3038
+ 263,
3039
+ 3369,
3040
+ 5762,
3041
+ 292,
3042
+ 263,
3043
+ 3369,
3044
+ 586,
3045
+ 6113,
3046
+ 290,
3047
+ 968,
3048
+ 9811,
3049
+ 290,
3050
+ 263,
3051
+ 3036,
3052
+ 46,
3053
+ 10,
3054
+ 10,
3055
+ 489,
3056
+ 2405,
3057
+ 311,
3058
+ 1873,
3059
+ 290,
3060
+ 451,
3061
+ 698,
3062
+ 263,
3063
+ 770,
3064
+ 2242,
3065
+ 2405,
3066
+ 287
3067
+ ],
3068
+ "completion_tokens": 100,
3069
+ "stopped_on_eos": false,
3070
+ "completion": " the port city of Château-Thierry.\n\nThe capital of France is the port city of Mont-Royal.\n\nHistory\n\nThe city of Mont-Royal was first mentioned by the French colonial administration in the 15th century. When the French occupation began in 1793, the city was captured by the French army and the French were forced to surrender to the British.\n\nThe city is thought to have been the first port city in",
3071
+ "full_text": "The capital of France is the port city of Château-Thierry.\n\nThe capital of France is the port city of Mont-Royal.\n\nHistory\n\nThe city of Mont-Royal was first mentioned by the French colonial administration in the 15th century. When the French occupation began in 1793, the city was captured by the French army and the French were forced to surrender to the British.\n\nThe city is thought to have been the first port city in",
3072
+ "full_text_sha256": "4b8907f330be8f045959ddb3ec1c50c88c82a1884c219aee6d3c6537a442ca84"
3073
+ },
3074
+ {
3075
+ "prompt_id": "explanation",
3076
+ "prompt": "Photosynthesis is the process by which",
3077
+ "prompt_token_ids": [
3078
+ 3561,
3079
+ 316,
3080
+ 394,
3081
+ 10904,
3082
+ 7834,
3083
+ 311,
3084
+ 263,
3085
+ 1198,
3086
+ 410,
3087
+ 510
3088
+ ],
3089
+ "completion_token_ids": [
3090
+ 3301,
3091
+ 292,
3092
+ 3581,
3093
+ 709,
3094
+ 1552,
3095
+ 434,
3096
+ 8653,
3097
+ 292,
3098
+ 602,
3099
+ 3760,
3100
+ 290,
3101
+ 5190,
3102
+ 915,
3103
+ 869,
3104
+ 7965,
3105
+ 292,
3106
+ 11698,
3107
+ 1240,
3108
+ 3465,
3109
+ 7969,
3110
+ 46,
3111
+ 383,
3112
+ 1198,
3113
+ 311,
3114
+ 8440,
3115
+ 290,
3116
+ 9161,
3117
+ 292,
3118
+ 3093,
3119
+ 286,
3120
+ 3301,
3121
+ 292,
3122
+ 3581,
3123
+ 46,
3124
+ 10,
3125
+ 10,
3126
+ 50,
3127
+ 46,
3128
+ 8645,
3129
+ 7569,
3130
+ 443,
3131
+ 287,
3132
+ 1788,
3133
+ 1222,
3134
+ 58,
3135
+ 10,
3136
+ 10,
3137
+ 4315,
3138
+ 1222,
3139
+ 376,
3140
+ 263,
3141
+ 3919,
3142
+ 2350,
3143
+ 286,
3144
+ 1552,
3145
+ 332,
3146
+ 263,
3147
+ 2921,
3148
+ 456,
3149
+ 8124,
3150
+ 115,
3151
+ 46,
3152
+ 383,
3153
+ 3315,
3154
+ 286,
3155
+ 3301,
3156
+ 290,
3157
+ 7396,
3158
+ 4487,
3159
+ 5888,
3160
+ 311,
3161
+ 3480,
3162
+ 332,
3163
+ 548,
3164
+ 9161,
3165
+ 46,
3166
+ 1788,
3167
+ 1222,
3168
+ 641,
3169
+ 5759,
3170
+ 11784,
3171
+ 3699,
3172
+ 9079,
3173
+ 686,
3174
+ 422,
3175
+ 103,
3176
+ 887,
3177
+ 44,
3178
+ 510,
3179
+ 376,
3180
+ 851,
3181
+ 410,
3182
+ 3301,
3183
+ 290,
3184
+ 1472,
3185
+ 548,
3186
+ 1442,
3187
+ 6800,
3188
+ 46,
3189
+ 10
3190
+ ],
3191
+ "completion_tokens": 100,
3192
+ "stopped_on_eos": false,
3193
+ "completion": " plants and animals use energy from sunlight and other sources to synthesize proteins and carbohydrates. The process is vital to survival and growth of plants and animals.\n\n2. Energy Capture in Plants:\n\nPlants are the primary source of energy for the Earth's ecosystems. The ability of plants to capture solar radiation is essential for their survival. Plants also convert atmospheric carbon dioxide into sugars, which are used by plants to build their cell walls.\n",
3194
+ "full_text": "Photosynthesis is the process by which plants and animals use energy from sunlight and other sources to synthesize proteins and carbohydrates. The process is vital to survival and growth of plants and animals.\n\n2. Energy Capture in Plants:\n\nPlants are the primary source of energy for the Earth's ecosystems. The ability of plants to capture solar radiation is essential for their survival. Plants also convert atmospheric carbon dioxide into sugars, which are used by plants to build their cell walls.\n",
3195
+ "full_text_sha256": "a8457bf7fa743d8397b5a221f4f11c5e0b978372eca68f068c17052307853591"
3196
+ },
3197
+ {
3198
+ "prompt_id": "story",
3199
+ "prompt": "In a distant future, humanity discovered",
3200
+ "prompt_token_ids": [
3201
+ 750,
3202
+ 257,
3203
+ 10043,
3204
+ 2639,
3205
+ 44,
3206
+ 1871,
3207
+ 421,
3208
+ 4919
3209
+ ],
3210
+ "completion_token_ids": [
3211
+ 263,
3212
+ 2291,
3213
+ 286,
3214
+ 263,
3215
+ 6601,
3216
+ 290,
3217
+ 5396,
3218
+ 1871,
3219
+ 421,
3220
+ 287,
3221
+ 548,
3222
+ 1264,
3223
+ 332,
3224
+ 2981,
3225
+ 292,
3226
+ 3167,
3227
+ 46,
3228
+ 12287,
3229
+ 10,
3230
+ 12286,
3231
+ 1518,
3232
+ 10,
3233
+ 3648,
3234
+ 359,
3235
+ 2032,
3236
+ 525,
3237
+ 584,
3238
+ 652,
3239
+ 263,
3240
+ 6601,
3241
+ 292,
3242
+ 671,
3243
+ 2539,
3244
+ 337,
3245
+ 1871,
3246
+ 421,
3247
+ 63,
3248
+ 12287,
3249
+ 10,
3250
+ 12286,
3251
+ 2146,
3252
+ 10,
3253
+ 8661,
3254
+ 33,
3255
+ 383,
3256
+ 6601,
3257
+ 311,
3258
+ 257,
3259
+ 2387,
3260
+ 3629,
3261
+ 286,
3262
+ 263,
3263
+ 2921,
3264
+ 456,
3265
+ 2395,
3266
+ 292,
3267
+ 311,
3268
+ 257,
3269
+ 7166,
3270
+ 647,
3271
+ 286,
3272
+ 263,
3273
+ 2387,
3274
+ 7462,
3275
+ 286,
3276
+ 263,
3277
+ 2921,
3278
+ 46,
3279
+ 2843,
3280
+ 376,
3281
+ 649,
3282
+ 2658,
3283
+ 287,
3284
+ 510,
3285
+ 263,
3286
+ 6601,
3287
+ 9220,
3288
+ 1871,
3289
+ 1227,
3290
+ 58,
3291
+ 10,
3292
+ 10,
3293
+ 49,
3294
+ 46,
3295
+ 2921,
3296
+ 456,
3297
+ 11026,
3298
+ 4326,
3299
+ 58,
3300
+ 383,
3301
+ 6601,
3302
+ 564,
3303
+ 257,
3304
+ 2239,
3305
+ 2539,
3306
+ 337,
3307
+ 1871,
3308
+ 1227,
3309
+ 46,
3310
+ 672
3311
+ ],
3312
+ "completion_tokens": 100,
3313
+ "stopped_on_eos": false,
3314
+ "completion": " the potential of the moon to guide humanity in their quest for knowledge and understanding.\nuser\nCan you tell me more about the moon and its impact on humanity?\nassistant\nSure! The moon is a natural feature of the Earth's surface and is a crucial part of the natural landscape of the Earth. Here are some ways in which the moon impacts human life:\n\n1. Earth's gravitational pull: The moon has a significant impact on human life. It",
3315
+ "full_text": "In a distant future, humanity discovered the potential of the moon to guide humanity in their quest for knowledge and understanding.\nuser\nCan you tell me more about the moon and its impact on humanity?\nassistant\nSure! The moon is a natural feature of the Earth's surface and is a crucial part of the natural landscape of the Earth. Here are some ways in which the moon impacts human life:\n\n1. Earth's gravitational pull: The moon has a significant impact on human life. It",
3316
+ "full_text_sha256": "00fb4213a5eeb36d18cd947a0876ad711f00bee4a8f576d75fd551d66183ba82"
3317
+ },
3318
+ {
3319
+ "prompt_id": "code",
3320
+ "prompt": "def fibonacci(n):",
3321
+ "prompt_token_ids": [
3322
+ 1806,
3323
+ 8247,
3324
+ 262,
3325
+ 3119,
3326
+ 105,
3327
+ 40,
3328
+ 110,
3329
+ 992
3330
+ ],
3331
+ "completion_token_ids": [
3332
+ 10,
3333
+ 351,
3334
+ 303,
3335
+ 684,
3336
+ 345,
3337
+ 58,
3338
+ 10,
3339
+ 4941,
3340
+ 620,
3341
+ 10,
3342
+ 10,
3343
+ 351,
3344
+ 8247,
3345
+ 262,
3346
+ 3119,
3347
+ 105,
3348
+ 40,
3349
+ 110,
3350
+ 992,
3351
+ 10,
3352
+ 4941,
3353
+ 8247,
3354
+ 262,
3355
+ 3119,
3356
+ 105,
3357
+ 40,
3358
+ 110,
3359
+ 45,
3360
+ 49,
3361
+ 41,
3362
+ 10,
3363
+ 10,
3364
+ 4941,
3365
+ 8247,
3366
+ 262,
3367
+ 3119,
3368
+ 105,
3369
+ 40,
3370
+ 110,
3371
+ 45,
3372
+ 49,
3373
+ 41,
3374
+ 757,
3375
+ 8247,
3376
+ 262,
3377
+ 3119,
3378
+ 105,
3379
+ 40,
3380
+ 110,
3381
+ 45,
3382
+ 50,
3383
+ 41,
3384
+ 10,
3385
+ 10,
3386
+ 2647,
3387
+ 8247,
3388
+ 262,
3389
+ 3119,
3390
+ 105,
3391
+ 40,
3392
+ 48,
3393
+ 41,
3394
+ 757,
3395
+ 8247,
3396
+ 262,
3397
+ 3119,
3398
+ 105,
3399
+ 40,
3400
+ 48,
3401
+ 41,
3402
+ 10,
3403
+ 10,
3404
+ 4941,
3405
+ 8247,
3406
+ 262,
3407
+ 3119,
3408
+ 105,
3409
+ 40,
3410
+ 48,
3411
+ 41,
3412
+ 757,
3413
+ 8247,
3414
+ 262,
3415
+ 3119,
3416
+ 105,
3417
+ 40,
3418
+ 48,
3419
+ 41,
3420
+ 757,
3421
+ 8247,
3422
+ 262,
3423
+ 3119,
3424
+ 105,
3425
+ 40,
3426
+ 49,
3427
+ 41,
3428
+ 10,
3429
+ 10,
3430
+ 2647,
3431
+ 8247
3432
+ ],
3433
+ "completion_tokens": 100,
3434
+ "stopped_on_eos": false,
3435
+ "completion": "\nif n < 1:\nreturn 0\n\nif fibonacci(n):\nreturn fibonacci(n-1)\n\nreturn fibonacci(n-1) + fibonacci(n-2)\n\nprint fibonacci(0) + fibonacci(0)\n\nreturn fibonacci(0) + fibonacci(0) + fibonacci(1)\n\nprint fib",
3436
+ "full_text": "def fibonacci(n):\nif n < 1:\nreturn 0\n\nif fibonacci(n):\nreturn fibonacci(n-1)\n\nreturn fibonacci(n-1) + fibonacci(n-2)\n\nprint fibonacci(0) + fibonacci(0)\n\nreturn fibonacci(0) + fibonacci(0) + fibonacci(1)\n\nprint fib",
3437
+ "full_text_sha256": "012a9d1312d787e5e759f1af7a242f20f4c9fdfc895ae673b0889fd08a0e8fe1"
3438
+ }
3439
+ ]
3440
  }
3441
  ]
3442
  }
model.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:8e74aa4d34464229e86a1831fb34d442f91e22aa600e0cf2503891957fdc6a43
3
- size 510713512
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:4a88432225dac6255ae1e2a085019f36dcdaf55fa33c77629dba6e7445c6c2fb
3
+ size 510281160
release_manifest.json CHANGED
@@ -1,81 +1,118 @@
1
  {
2
  "schema_version": 3,
3
- "revision": 6,
4
- "revision_id": "r006",
5
- "release": "Pollock 1.4",
6
  "model_id": "SlayerLab/pollock-mini-lm-125m",
 
7
  "publication": {
8
- "weights_commit": "a2e53d9f1b690ebfbea375cf8bd3b5bdf69dece9",
9
- "version_tag": "v1.4",
10
- "publication_commit_pending": true
11
  },
12
  "source_checkpoint": {
13
- "path_in_training_workspace": "runs/r006-corrected-corpus-v2-s1337-lr0.0004/checkpoints/ckpt-final.pt",
14
- "sha256": "580fbb3c039ea1f349285a9bf7a67731fb6d455de931b6b7601d4ed09fce9437",
15
- "iteration": 22003,
16
- "tokens_seen": 10814914560,
17
- "native_nanogpt_parameters": 126637952,
18
- "native_unique_trainable_parameters": 127555456
 
 
 
 
 
 
 
 
 
19
  },
20
  "architecture": {
21
- "n_layer": 12,
22
- "n_head": 14,
23
- "n_embd": 896,
24
- "block_size": 1024,
 
25
  "vocab_size": 12288,
26
  "dropout": 0.0,
27
  "bias": false,
28
  "tied_word_embeddings": true
29
  },
30
  "training": {
31
- "dataset": "SlayerLab/minimal-en-corpus-2.5b",
32
- "dataset_revision": "d68d992622e9fc11f19e7d7fb8547c4e653439a4",
33
- "training_tokens": 2689323439,
34
- "validation_tokens": 5236486,
35
- "tokenizer_sha256": "6cda4e5ec8293f3b821e02253f9c0e88e43ad4f7c6e1324763e1fb91749fef51",
36
- "native_tokenizer_sha256": "d0d126f0a59e51e8cb1b26d77b5527870be0e8ea1e84adefd730fff906c48234",
37
  "init_from": "scratch",
38
- "micro_batch_per_gpu": 12,
 
 
39
  "gradient_accumulation_global": 40,
40
  "gradient_accumulation_per_gpu": 20,
41
  "ddp_world_size": 2,
42
  "effective_batch_tokens": 491520,
43
- "optimizer": "fused AdamW",
44
- "learning_rate": 0.0004,
45
- "min_learning_rate": 0.00004,
46
- "schedule": "cosine",
47
- "warmup_iters": 440,
 
 
 
 
48
  "beta1": 0.9,
49
  "beta2": 0.95,
50
- "weight_decay": 0.1,
 
51
  "grad_clip": 1.0,
52
  "precision": "bfloat16",
53
  "hardware": "2x NVIDIA GeForce RTX 4090",
54
- "framework": "PyTorch 2.8.0+cu128",
55
  "nanogpt_commit": "3adf61e154c3fe3fca428ad6bc3818b27a3b8291",
56
- "data_pass_equivalent": 4.021425762020438,
57
- "runtime_hours": 15.226956854563,
58
- "mean_tokens_per_second": 199894.6379482684
59
  },
60
  "evaluation": {
61
- "protocol": "fixed sampled subset",
62
- "subset_tokens_per_split": 1228800,
63
- "final_train_loss": 2.48974818944931,
64
- "final_validation_loss": 2.536356544494629,
65
- "best_validation_loss": 2.5362311387062073,
66
- "best_validation_update": 22000,
67
- "best_validation_tokens_seen": 10813440000,
68
- "final_validation_perplexity": 12.633557211556399,
 
 
 
 
 
 
 
 
 
 
 
69
  "benchmark_harness": "lm-evaluation-harness 0.4.12",
70
  "benchmark_num_fewshot": 0,
71
  "benchmark_batch_size": 8,
72
  "truncated_benchmark_requests": 0,
73
- "benchmark_checkpoint_sha256": "9339d20ea32d24bc9777518a2f0ec553e4a3871647d342db92d7c5b6c18c7b56",
74
- "benchmark_tokenizer_sha256": "6cda4e5ec8293f3b821e02253f9c0e88e43ad4f7c6e1324763e1fb91749fef51",
75
- "results_file": "benchmarks/english.json"
 
 
 
 
 
 
 
 
76
  },
77
  "fixed_inference": {
78
  "suite_id": "fixed-sampling-v1",
 
79
  "comparison_file": "inference-samples/README.md",
80
  "config_file": "inference-samples/config.json",
81
  "generator_file": "inference-samples/generate.py",
@@ -86,44 +123,47 @@
86
  "dtype": "float32",
87
  "seed": 1337,
88
  "prompt_count": 4,
89
- "revision_count": 6,
90
- "local_review_artifact": false,
91
  "exact_replay_verified": true,
92
- "publication_commit_pending": true
 
93
  },
94
  "conversion": {
 
95
  "target_class": "GPT2LMHeadModel",
96
  "transformers_version": "5.15.1",
97
- "checkpoint_sha256": "580fbb3c039ea1f349285a9bf7a67731fb6d455de931b6b7601d4ed09fce9437",
98
- "model_sha256": "8e74aa4d34464229e86a1831fb34d442f91e22aa600e0cf2503891957fdc6a43",
99
- "unique_serialized_parameters": 127674624,
100
- "compatibility_zero_bias_parameters": 119168,
101
  "validation_probe_shape": [2, 64],
102
  "max_absolute_logit_error": 0.0
103
  },
104
  "artifacts": {
105
  ".gitattributes": {"sha256": "c759491a998899dbefaec4d51cc791e68c714a37dd9f9829020a368788fe3063"},
106
- "CHANGELOG.md": {"sha256": "8422ad337b22be7adb561d1b2388cf53e5a6d1c6d49d628468498c72f078769a"},
107
- "LICENSE.md": {"sha256": "46cbe928ed0aa24875f02f774313b27ec9d8abdf41f0c9ef67e0adb4e0614de4"},
108
- "README.md": {"sha256": "875fa9cd748691e376df905e036a0c471db77aaa1272d71c5ea10d0f76fa888d"},
109
  "assets/pollock-mini-lm-avatar-320.png": {"sha256": "7be10cc9d0f4aedeb298d9a5d722b2b2ac5b2e219b3884916dbca16198f70750"},
110
- "benchmarks/english.json": {"sha256": "646a3aff7c97859e6974baa2c191702829a982af87675870dd1a339f8073a146"},
111
- "config.json": {"sha256": "1529f8a8fc31b4fa68fc18fffbad7e155dc8f9e8663a40b4f3306b4012ea38c0"},
 
112
  "generation_config.json": {"sha256": "435beb27be51f0ed054f4a011e5109d125cdadc118b8799b18b155cc798d94d2"},
113
- "inference-samples/README.md": {"sha256": "6559c85e1ed70f47f2d2978f28151fdd9e3355e397c5591cbe0b8ab767fa5ac5"},
114
- "inference-samples/config.json": {"sha256": "a1b306a1e1f3a70942706d00107a83dd98289a534d17af78adbb95d48c9b69c1"},
115
- "inference-samples/generate.py": {"sha256": "807a9946472bd0087dd2b9260ee90d1dea09785d24d4f923f028667c53a23933"},
116
  "inference-samples/requirements.txt": {"sha256": "584583335ffb3061aa62058aee725b0de25fed7523904fbd13fb4622f601943d"},
117
- "inference-samples/results.json": {"sha256": "a279806ea225130b0e0aa36d9bd13653401e236395fe64b5f0aa7c825d445620"},
118
- "model.safetensors": {"sha256": "8e74aa4d34464229e86a1831fb34d442f91e22aa600e0cf2503891957fdc6a43"},
119
  "special_tokens_map.json": {"sha256": "8b2257a17ea997bb038f43b133aefec82344ad2b8abc2b8a02a6c0a994ed624e"},
120
- "tokenizer.json": {"sha256": "6cda4e5ec8293f3b821e02253f9c0e88e43ad4f7c6e1324763e1fb91749fef51"},
121
- "tokenizer_config.json": {"sha256": "4cdabe37dbdc1adfcc017ee9a1f86ab89bdf184827d2d9f05181cae0f8af19bf"},
122
  "training-history/r001.md": {"sha256": "5b0ac39265921ed3604e7dfe8181b00d07902640d8024f591462eeb508bcfece"},
123
  "training-history/r002.md": {"sha256": "80031a84b2b67dc4cb4f81a178fa354feea15cb184d17368ee7e64ec3d0a2629"},
124
  "training-history/r003.md": {"sha256": "7ea9e6b15a0af635ff238417b9fff2a98c12a9090800a0923d622253b7eab02d"},
125
  "training-history/r004.md": {"sha256": "e260bfba69f154caf16f405d6944c2f963f467da9ed1e6c3c94e08932feae9b9"},
126
  "training-history/r005.md": {"sha256": "a25c2932701695dd7402582510adbe71e50a6fcc2da0fb3c6224a846dfdb22df"},
127
- "training-history/r006.md": {"sha256": "47d53f9aa8a233eb7a319f559209936df988e83d96f6f6517fa3670b5dd2cdd3"}
 
128
  }
129
  }
 
1
  {
2
  "schema_version": 3,
3
+ "revision": 7,
4
+ "revision_id": "r007",
5
+ "release": "Pollock 1.5",
6
  "model_id": "SlayerLab/pollock-mini-lm-125m",
7
+ "draft": false,
8
  "publication": {
9
+ "weights_commit": "f2a1ec072098d257c36b129b0e554e2e6de9c46e",
10
+ "version_tag": "v1.5",
11
+ "publication_commit_pending": false
12
  },
13
  "source_checkpoint": {
14
+ "role": "best-full-validation",
15
+ "path_in_training_workspace": "out-pollock-r007-e4-final/ckpt-best.pt",
16
+ "sha256": "384d912cc8b7ebf872b66d1f8ad46f223296a069e906167076f1899296f9a788",
17
+ "iteration": 43750,
18
+ "tokens_seen": 21503919992,
19
+ "data_pass_equivalent": 3.9847123089835352,
20
+ "native_nanogpt_parameters": 125854464,
21
+ "native_unique_trainable_parameters": 127427328
22
+ },
23
+ "terminal_checkpoint": {
24
+ "path_in_training_workspace": "out-pollock-r007-e4-final/ckpt-e4-terminal-step-00043918.pt",
25
+ "sha256": "ef6952d47b717266ff57272b0890a7c54ed36111bb65a2c22bb2630df87f3016",
26
+ "iteration": 43918,
27
+ "tokens_seen": 21586421624,
28
+ "data_pass_equivalent": 4.0
29
  },
30
  "architecture": {
31
+ "n_layer": 16,
32
+ "n_head": 12,
33
+ "n_embd": 768,
34
+ "n_inner": 3200,
35
+ "block_size": 2048,
36
  "vocab_size": 12288,
37
  "dropout": 0.0,
38
  "bias": false,
39
  "tied_word_embeddings": true
40
  },
41
  "training": {
42
+ "dataset": "SlayerLab/minimal-en-corpus-5b",
43
+ "dataset_revision": "38bebbd2ae36858c25aa0799fd3296fbfa9432aa",
44
+ "training_tokens": 5396605407,
45
+ "validation_tokens": 5238223,
46
+ "tokenizer_sha256": "3733307577230bb4802d2c774d8e2323f7f64e4139c13712736a57daf91bdda1",
 
47
  "init_from": "scratch",
48
+ "data_sampling": "seeded global permutation of fixed token windows per pass",
49
+ "data_shuffle_seed": 1337,
50
+ "micro_batch_per_gpu": 6,
51
  "gradient_accumulation_global": 40,
52
  "gradient_accumulation_per_gpu": 20,
53
  "ddp_world_size": 2,
54
  "effective_batch_tokens": 491520,
55
+ "optimizer": "Muon plus fused AdamW",
56
+ "muon_parameters": 116391936,
57
+ "adamw_parameters": 11035392,
58
+ "muon_learning_rate": 0.01,
59
+ "muon_min_learning_rate": 0.001,
60
+ "adamw_learning_rate": 0.0003,
61
+ "adamw_min_learning_rate": 0.00003,
62
+ "schedule": "200-update warmup and cosine decay through update 21959; learning-rate floors thereafter",
63
+ "warmup_iters": 200,
64
  "beta1": 0.9,
65
  "beta2": 0.95,
66
+ "muon_momentum": 0.95,
67
+ "weight_decay": 0.01,
68
  "grad_clip": 1.0,
69
  "precision": "bfloat16",
70
  "hardware": "2x NVIDIA GeForce RTX 4090",
71
+ "framework": "PyTorch 2.11.0+cu128",
72
  "nanogpt_commit": "3adf61e154c3fe3fca428ad6bc3818b27a3b8291",
73
+ "completed_data_passes": 4.0,
74
+ "runtime_hours": null,
75
+ "mean_tokens_per_second": null
76
  },
77
  "evaluation": {
78
+ "protocol": "complete validation split plus fixed corpus-wide shuffled train sample",
79
+ "train_subset_tokens": 2457600,
80
+ "validation_targets": 5238222,
81
+ "selected_train_loss_rounded": 2.4227,
82
+ "selected_validation_loss": 2.458615977563123,
83
+ "selected_validation_perplexity": 11.688623023409297,
84
+ "selected_validation_update": 43750,
85
+ "selected_validation_tokens_seen": 21503919992,
86
+ "terminal_train_loss_rounded": 2.4226,
87
+ "terminal_validation_loss_rounded": 2.4587,
88
+ "e4_best_diagnostic_status": "completed",
89
+ "e4_best_training_workspace_results_file": "evaluations/r007_epochs/results/E4-best/metrics.json",
90
+ "e4_best_cross_entropy": 2.458519056568919,
91
+ "e4_best_perplexity": 11.687490205342804,
92
+ "e4_best_predictive_entropy": 2.4375294293734733,
93
+ "e4_best_repetition_4gram": 0.33997844253584,
94
+ "e4_best_repetition_8gram": 0.2570576867781918,
95
+ "e4_best_repetition_16gram": 0.20893125597814885,
96
+ "benchmark_status": "completed",
97
  "benchmark_harness": "lm-evaluation-harness 0.4.12",
98
  "benchmark_num_fewshot": 0,
99
  "benchmark_batch_size": 8,
100
  "truncated_benchmark_requests": 0,
101
+ "benchmark_checkpoint_sha256": "4a88432225dac6255ae1e2a085019f36dcdaf55fa33c77629dba6e7445c6c2fb",
102
+ "benchmark_tokenizer_sha256": "3733307577230bb4802d2c774d8e2323f7f64e4139c13712736a57daf91bdda1",
103
+ "results_file": "benchmarks/english.json",
104
+ "blimp_accuracy": 0.7849701492537313,
105
+ "lambada_openai_accuracy": 0.3007956530176596,
106
+ "lambada_openai_perplexity": 45.030361770214945,
107
+ "hellaswag_accuracy_normalized": 0.30830511850229037,
108
+ "piqa_accuracy_normalized": 0.5968443960826986,
109
+ "sciq_accuracy_normalized": 0.675,
110
+ "arc_easy_accuracy_normalized": 0.41835016835016836,
111
+ "arc_challenge_accuracy_normalized": 0.25
112
  },
113
  "fixed_inference": {
114
  "suite_id": "fixed-sampling-v1",
115
+ "status": "completed",
116
  "comparison_file": "inference-samples/README.md",
117
  "config_file": "inference-samples/config.json",
118
  "generator_file": "inference-samples/generate.py",
 
123
  "dtype": "float32",
124
  "seed": 1337,
125
  "prompt_count": 4,
126
+ "revision_count": 7,
 
127
  "exact_replay_verified": true,
128
+ "publication_commit_pending": false,
129
+ "r007_weights_commit": "f2a1ec072098d257c36b129b0e554e2e6de9c46e"
130
  },
131
  "conversion": {
132
+ "status": "completed",
133
  "target_class": "GPT2LMHeadModel",
134
  "transformers_version": "5.15.1",
135
+ "checkpoint_sha256": "384d912cc8b7ebf872b66d1f8ad46f223296a069e906167076f1899296f9a788",
136
+ "model_sha256": "4a88432225dac6255ae1e2a085019f36dcdaf55fa33c77629dba6e7445c6c2fb",
137
+ "unique_serialized_parameters": 127565312,
138
+ "compatibility_zero_bias_parameters": 137984,
139
  "validation_probe_shape": [2, 64],
140
  "max_absolute_logit_error": 0.0
141
  },
142
  "artifacts": {
143
  ".gitattributes": {"sha256": "c759491a998899dbefaec4d51cc791e68c714a37dd9f9829020a368788fe3063"},
144
+ "CHANGELOG.md": {"sha256": "c2f62c0227cb8ccf18417bf91ff4dbfcc1c08069815fe215785e3207cbabde11"},
145
+ "LICENSE.md": {"sha256": "162b97898a956d43d0a5428d1999bac3e42e2a97eb01a79f73ad719a4a15c412"},
146
+ "README.md": {"sha256": "b9e153c60d9dd14caf19e02515522c7ef17d14fa1fc7037a0acbaf48a6a7d49a"},
147
  "assets/pollock-mini-lm-avatar-320.png": {"sha256": "7be10cc9d0f4aedeb298d9a5d722b2b2ac5b2e219b3884916dbca16198f70750"},
148
+ "benchmarks/english.json": {"sha256": "430feae7aefd6d8e226f61403a9e1080274daac538361d9ab6dcd5e1129218e2"},
149
+ "config.json": {"sha256": "aedc3eea385b4295b3ba257eb8533b06046eea0cca1ff009515ad3f1adbb4854"},
150
+ "conversion_manifest.json": {"sha256": "eb207d676eca255f7dd9415049b825f55ea81f8eed11cc9b0fc85a3f8b2bea5d"},
151
  "generation_config.json": {"sha256": "435beb27be51f0ed054f4a011e5109d125cdadc118b8799b18b155cc798d94d2"},
152
+ "inference-samples/README.md": {"sha256": "66a15dba47cc44864f9accaad51955ed11c61d2103472ecf7a0c3451aa2fb291"},
153
+ "inference-samples/config.json": {"sha256": "3e908dfa8ce6085092dc665f7bdf7b69490fe49779c74a2c69adbd7265403ed9"},
154
+ "inference-samples/generate.py": {"sha256": "b4bb2edc64924ebdcbbd4ed7d0e72c44e6d9d05e58c96541c00cdeeb075b917d"},
155
  "inference-samples/requirements.txt": {"sha256": "584583335ffb3061aa62058aee725b0de25fed7523904fbd13fb4622f601943d"},
156
+ "inference-samples/results.json": {"sha256": "ab07a2bf47297ab17f12fbd00aa5ca10569eef31a07984ba1af6e0321d689841"},
157
+ "model.safetensors": {"sha256": "4a88432225dac6255ae1e2a085019f36dcdaf55fa33c77629dba6e7445c6c2fb"},
158
  "special_tokens_map.json": {"sha256": "8b2257a17ea997bb038f43b133aefec82344ad2b8abc2b8a02a6c0a994ed624e"},
159
+ "tokenizer.json": {"sha256": "3733307577230bb4802d2c774d8e2323f7f64e4139c13712736a57daf91bdda1"},
160
+ "tokenizer_config.json": {"sha256": "39a075b24f05384c6d759521563d879cab4bb80b0fe94381e6e792966ab4c2ba"},
161
  "training-history/r001.md": {"sha256": "5b0ac39265921ed3604e7dfe8181b00d07902640d8024f591462eeb508bcfece"},
162
  "training-history/r002.md": {"sha256": "80031a84b2b67dc4cb4f81a178fa354feea15cb184d17368ee7e64ec3d0a2629"},
163
  "training-history/r003.md": {"sha256": "7ea9e6b15a0af635ff238417b9fff2a98c12a9090800a0923d622253b7eab02d"},
164
  "training-history/r004.md": {"sha256": "e260bfba69f154caf16f405d6944c2f963f467da9ed1e6c3c94e08932feae9b9"},
165
  "training-history/r005.md": {"sha256": "a25c2932701695dd7402582510adbe71e50a6fcc2da0fb3c6224a846dfdb22df"},
166
+ "training-history/r006.md": {"sha256": "47d53f9aa8a233eb7a319f559209936df988e83d96f6f6517fa3670b5dd2cdd3"},
167
+ "training-history/r007.md": {"sha256": "fc9a4153eb8c073287f76b82b89ffaad62f3a606ba000e53bf450e393a5807d6"}
168
  }
169
  }
tokenizer.json CHANGED
The diff for this file is too large to render. See raw diff
 
tokenizer_config.json CHANGED
@@ -6,7 +6,7 @@
6
  "<|im_start|>",
7
  "<|im_end|>"
8
  ],
9
- "model_max_length": 1024,
10
  "pad_token": "<|endoftext|>",
11
  "tokenizer_class": "TokenizersBackend"
12
  }
 
6
  "<|im_start|>",
7
  "<|im_end|>"
8
  ],
9
+ "model_max_length": 2048,
10
  "pad_token": "<|endoftext|>",
11
  "tokenizer_class": "TokenizersBackend"
12
  }
training-history/r007.md ADDED
@@ -0,0 +1,71 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # r007 — Pollock 1.5 training record
2
+
3
+ ## Identity and provenance
4
+
5
+ | Field | Value |
6
+ |---|---|
7
+ | Revision / release | `r007` / Pollock 1.5 |
8
+ | Source run | `r007-muon-128m` |
9
+ | Selected checkpoint | update 43,750 |
10
+ | Selected checkpoint SHA-256 | `384d912cc8b7ebf872b66d1f8ad46f223296a069e906167076f1899296f9a788` |
11
+ | Terminal E4 checkpoint | update 43,918 |
12
+ | Terminal E4 checkpoint SHA-256 | `ef6952d47b717266ff57272b0890a7c54ed36111bb65a2c22bb2630df87f3016` |
13
+ | `model.safetensors` SHA-256 | `4a88432225dac6255ae1e2a085019f36dcdaf55fa33c77629dba6e7445c6c2fb` |
14
+
15
+ ## Architecture and training
16
+
17
+ | Setting | Value |
18
+ |---|---:|
19
+ | Layers / heads / width | 16 / 12 / 768 |
20
+ | MLP width | 3,200 |
21
+ | Context / vocabulary | 2048 / 12288 |
22
+ | Native trainable parameters | 127,427,328 |
23
+ | nanoGPT non-embedding parameters | 125,854,464 |
24
+ | Dataset | SlayerLab/minimal-en-corpus-5b (`38bebbd`) |
25
+ | Tokenizer SHA-256 | `3733307577230bb4802d2c774d8e2323f7f64e4139c13712736a57daf91bdda1` |
26
+ | Selected token presentations / passes | 21,503,919,992 / 3.984712 |
27
+ | Terminal token presentations / passes | 21,586,421,624 / 4.000000 |
28
+ | Effective batch | 491,520 tokens |
29
+ | Optimizer | Muon plus fused AdamW |
30
+ | Muon LR range | 1e-02 → 1e-03 |
31
+ | AdamW LR range | 3e-04 → 3e-05 |
32
+ | Precision / hardware | BF16 / 2× NVIDIA GeForce RTX 4090 |
33
+ | Framework | PyTorch 2.11.0+cu128 |
34
+
35
+ Muon optimized 116,391,936 parameters in eligible hidden attention and MLP matrices. Fused AdamW optimized the remaining 11,035,392 parameters in token and position embeddings, the tied output head, normalization parameters, and biases. The parameter partition was checked for overlap and completeness before training.
36
+
37
+ The loader applied a seeded global permutation to fixed 2,048-token windows in every pass, assigned disjoint positions across the two DDP ranks, masked each final partial batch, and stored its global cursor, topology, and shuffle settings in every checkpoint. Training completed exactly four passes over the 5,396,605,407-token training split.
38
+
39
+ The first two passes used a 200-update linear warmup followed by cosine decay through update 21,959. The E3 and E4 continuation kept both learning rates at their floors without a restart or second cosine cycle.
40
+
41
+ ## Training-time evaluation
42
+
43
+ Every validation evaluation covered the complete 5,238,222-target validation split. Train evaluation used a separate fixed, corpus-wide shuffled sample of 2,457,600 tokens without advancing the training loader.
44
+
45
+ The selected update-43,750 checkpoint recorded train-eval loss **2.4227** and full-validation loss **2.458616** (perplexity **11.688623**) after 21,503,919,992 token presentations. The terminal update-43,918 evaluation recorded train-eval loss **2.4226** and full-validation loss **2.4587**; these terminal values are printed to four decimal places in the training log.
46
+
47
+ The independent locked `E4-best` diagnostic evaluation reproduced full-validation cross-entropy **2.458519**, perplexity **11.687490**, and predictive entropy **2.437529** on Apple MPS with BF16. It improved CE over E3 in every reported domain, every context-position bucket, and every document-length bucket. The fixed 45-prompt generation suite measured repeated 4-, 8-, and 16-gram ratios of **0.339978**, **0.257058**, and **0.208931**. These are lower than E3's ratios but remain above E2's.
48
+
49
+ The complete machine-readable diagnostic outputs are preserved in the training workspace under `evaluations/r007_epochs/results/E4-best/`.
50
+
51
+ ## English zero-shot benchmarks
52
+
53
+ The selected update-43,750 checkpoint was evaluated on complete splits with zero few-shot examples, `lm-evaluation-harness` 0.4.12, BF16, batch size 8, and maximum context 1024 on Apple M1 Max MPS. Exact benchmark-checkpoint and tokenizer hashes are preserved in [`../benchmarks/english.json`](../benchmarks/english.json).
54
+
55
+ | Benchmark | Primary metric | E2 | E4-best | Delta |
56
+ |---|---:|---:|---:|---:|
57
+ | BLiMP | accuracy | 0.784299 | 0.784970 | +0.000672 |
58
+ | LAMBADA OpenAI | accuracy | 0.314186 | 0.300796 | -0.013390 |
59
+ | HellaSwag | normalized accuracy | 0.310795 | 0.308305 | -0.002490 |
60
+ | PIQA | normalized accuracy | 0.603373 | 0.596844 | -0.006529 |
61
+ | SciQ | normalized accuracy | 0.681000 | 0.675000 | -0.006000 |
62
+ | ARC-Easy | normalized accuracy | 0.429714 | 0.418350 | -0.011364 |
63
+ | ARC-Challenge | normalized accuracy | 0.247440 | 0.250000 | +0.002560 |
64
+
65
+ LAMBADA perplexity changed from **44.145462** at E2 to **45.030362** at E4-best. E4-best improves validation cross-entropy and the locked generation repetition suite relative to E3, but its downstream benchmark movement relative to E2 is mixed.
66
+
67
+ ## Conversion and release status
68
+
69
+ The selected native checkpoint was converted to `GPT2LMHeadModel` with Transformers 5.15.1. A deterministic `[2, 64]` parity probe produced maximum absolute logit error **0.0**. The resulting `model.safetensors` SHA-256 is `4a88432225dac6255ae1e2a085019f36dcdaf55fa33c77629dba6e7445c6c2fb`.
70
+
71
+ The `fixed-sampling-v1` suite completed in its pinned environment and passed exact regeneration checks for all r001-r007 token sequences. R007 was replayed directly from immutable Hugging Face commit [`f2a1ec0`](https://huggingface.co/SlayerLab/pollock-mini-lm-125m/commit/f2a1ec072098d257c36b129b0e554e2e6de9c46e), whose `model.safetensors` hash matches the local converted artifact. Merge and version tag `v1.5` remain pending.