|
Download README.md from SlayerLab/GoLLeM-v6-250M: direct link, hf CLI and curl.
- Browser
- Download file 17.8 kB
-
https://huggingface.co/SlayerLab/GoLLeM-v6-250M/resolve/main/README.md
- Command line
-
hf download hf://SlayerLab/GoLLeM-v6-250M/README.md
-
curl -L -o README.md https://huggingface.co/SlayerLab/GoLLeM-v6-250M/resolve/main/README.md
17.8 kB
| license: cc-by-sa-4.0 | |
| language: | |
| - pl | |
| - en | |
| library_name: pytorch | |
| tags: | |
| - polish | |
| - english | |
| - language-model | |
| - research | |
| - gollem | |
| pipeline_tag: text-generation | |
| # GoLLeM v6 — 250M (Polish–English research base model) | |
| > **STATUS: TRAINING COMPLETE — RESEARCH BASE MODEL.** Final checkpoint of a single 760,000-step run. This is a **base | |
| > (pretrained) model, not instruction-tuned**: it continues text and does not follow commands. An instruction-tuned | |
| > version is a separate model with its own card. Results below use **our protocol, not official** benchmark submissions, | |
| > and are compared only with our own earlier models (GoLLeM v4 and v5) measured the same way. | |
| GoLLeM v6 is a 252.8M-parameter decoder-only language model trained from scratch on a Polish–English mix | |
| (about two thirds Polish, one third English). It is a Fabryka AI project (formerly SlayerLab). | |
| **Author:** Arkadiusz Słota (Fabryka AI). | |
| ## What it is for (and what it is not) | |
| **Intended use:** research on small bilingual language models, comparison with GoLLeM v4 (Polish only) and v5 | |
| (English only), and text continuation in Polish and English. | |
| **Not intended for:** production use, question answering without context, or following instructions. Like any base | |
| model of this size it has limited factual knowledge and may produce fluent but wrong text. | |
| ## Results (final checkpoint, step 760,000) | |
| All numbers: final checkpoint `221b1286…`, single run, seed 1337. **Our protocol, not official.** | |
| **English** — our implementation of the Glint leaderboard metrics (the same harness we used for GoLLeM v5): | |
| | | GoLLeM v6 250M | GoLLeM v5 128M (reference) | | |
| |---|---|---| | |
| | ARC-Easy (test, 2,376 questions) | 50.21 | 53.24 | | |
| | BLiMP (67,000 pairs) | 77.61 | 79.09 | | |
| | WikiText-2 test, byte perplexity | 2.2884 | 2.2538 | | |
| | **eff** (with size multiplier) | **74.71** (× 1.0) | **76.94** (76.30 × 1.0084) | | |
| eff = mean(ARC-Easy, BLiMP, WikiScore) × size multiplier, WikiScore = 100·(ln 500 − ln byte_ppl)/(ln 500 − ln 1.86), | |
| clipped to 0–100; the multiplier follows the leaderboard formula (1.0 at 150M parameters and above, 1.0084 for v5's | |
| declared 122.8M). Without the multiplier the English gap between v6 and v5 128M is −1.59 points, with it −2.23. v6 is | |
| below v5 128M in English: v6 saw about 8.9B English tokens (35.8 % of 24.9B), v5 128M was trained on English only. | |
| **Polish** — our eff-PL protocol (definition fixed on 2026-09-29, before training started): | |
| | | GoLLeM v6 250M | GoLLeM v4 250M (reference) | | |
| |---|---|---| | |
| | G — MultiBLiMP, Polish (3,272 pairs) | 98.14 | 98.47 | | |
| | K — ARC-Easy-PL (EU20 translation, test, 2,376 questions) | 42.09 | 41.96 | | |
| | W — byte perplexity on 463 Polish Wikipedia articles | 1.9578 | 1.9507 | | |
| | **eff-PL** = mean(G, K, WikiScore on W) | **79.77** | **79.86** | | |
| | **PL clean** = mean(G on 944 pairs, K on 2,362 questions) | **69.42 ± 0.58** | **69.62** | | |
| - **PL clean** removes MultiBLiMP pairs and ARC-Easy-PL questions found in **GoLLeM v4's** training corpus, so that v4 | |
| and v6 are compared on the same items: 2,207 of the 3,272 MultiBLiMP pairs are flagged by our scan of the v4 corpus, and pairs | |
| with a sentence shorter than 20 characters are also excluded (944 pairs remain); 14 ARC-Easy-PL test questions are in | |
| the v4 corpus (2,362 remain). "Clean" is therefore defined relative to the v4 corpus, not the v6 pool. The v6 pool was | |
| filtered before training against a protected list that includes these MultiBLiMP sentences (738 documents removed); | |
| a separate scan of the finished v6 pool against all 3,272 pairs has not been run. ± is one standard error (binomial, | |
| per axis, combined). | |
| - **Condition fixed in advance:** PL clean ≥ v4 − 1 point. The condition and the two subsets were fixed on 2026-09-29, | |
| before any v6 checkpoint was evaluated in Polish (the subsets about one hour after training started); the threshold | |
| value, **68.62**, depends only on the v4 measurement. The final checkpoint reaches 69.42, so the condition is met. | |
| - **Reading:** in Polish, v6 is **within one standard error of v4** (eff-PL −0.09, PL clean −0.20), while also being | |
| trained on about one third English. We do not claim it is better than v4 in Polish. | |
| - W articles were selected by a fixed hash rule and removed if more than 5 % of their 13-grams appear in the v6 | |
| training pool or if they were on our privacy drop list (589 of 1,052 removed; 463 kept). | |
| **Training progress** (evaluated every 6 data slots during training; intermediate checkpoints, not results): | |
| | step | ARC-Easy | BLiMP | WikiText-2 byte_ppl | G | K | W byte_ppl | | |
| |---|---|---|---|---|---|---| | |
| | 60,000 | 43.27 | 73.80 | 2.5193 | 96.49 | 35.82 | 2.1560 | | |
| | 180,000 | 45.29 | 75.54 | 2.4281 | 96.61 | 39.98 | 2.0787 | | |
| | 300,000 | 44.91 | 76.60 | 2.3996 | 97.31 | 39.44 | 2.0514 | | |
| | 420,000 | 47.14 | 76.55 | 2.3922 | 97.04 | 39.23 | 2.0368 | | |
| | 540,000 | 47.56 | 74.99 | 2.3709 | 97.46 | 40.49 | 2.0273 | | |
| | 600,000 (kept trunk checkpoint) | 48.86 | 76.59 | 2.3610 | 97.19 | 40.91 | 2.0239 | | |
| | 720,000 | 49.20 | 77.54 | 2.3069 | 97.83 | 41.25 | 1.9715 | | |
| | **760,000 (final)** | **50.21** | **77.61** | **2.2884** | **98.14** | **42.09** | **1.9578** | | |
| The final checkpoint is the best of all 13 evaluated points on every axis; most of the late gain came during the | |
| learning-rate decay (last 20 % of steps). With a single run we cannot separate the effect of the decay from the effect | |
| of more training. | |
| ## Training | |
| | | | | |
| |---|---| | |
| | Parameters | 252,760,019 | | |
| | Shape | 20 layers, d_model 960, 15 heads, context 1,024 tokens | | |
| | Block | RMSNorm, RoPE (θ = 100,000), SwiGLU (×2.667), QK-norm, value residuals | | |
| | Tokenizer | 32,768-token BPE, Polish + English | | |
| | Optimizer | Muon for the 221,184,000 parameters in 2-D hidden matrices (attention QKV and output projection, SwiGLU gate/up/down): momentum 0.95 (Nesterov), 5 Newton–Schulz iterations, learning rate 0.02 on the same schedule, no weight decay. AdamW (β 0.9 / 0.95, weight decay 0.1, learning rate 6e-4 → 6e-5) for the other 31,576,019: the token embedding tied with the output head, normalisation and QK-norm gains, attention biases and value-residual gates. 2,000 warmup steps. | | |
| | Schedule | warmup–stable–decay: constant learning rate, decay over the last 20 % of steps (from step 608,000) | | |
| | Batch / steps / tokens | 32 × 1,024 tokens per step × 760,000 steps = 24.90B tokens (about one pass over the pool) | | |
| | Validation loss (per token, our held-out set) | 2.5745 at step 760,000 | | |
| | Precision | bf16 autocast with fp32 weights; `torch.compile`. Released weights are fp32. | | |
| | Hardware / time | 1 × NVIDIA RTX PRO 6000 Blackwell Server Edition (cloud), single GPU; about 99,000 tokens/s during training. 2026-09-29 18:23 → 2026-10-03 00:19 UTC (about 78 h wall-clock, including a pause after each data slot for the control check: median about 6 min, range 5–35 min, about 8 h in total). | | |
| **Data queue with a rule.** Training data arrived in 76 slots of 10,000 steps. After each slot the checkpoint was scored | |
| on five fixed control sets (English educational, encyclopedic and Q&A text; Polish encyclopedic and general text) on a | |
| separate machine. The first 6 slots ran in shadow mode: the rule was not applied, and its noise level was calibrated on | |
| them (per-set residual standard deviation σ, at least 0.002), so the rule's thresholds were set during the run, after | |
| the shadow slots. From slot 6 on, a slot would be rejected if the loss on any of the **three English** control sets rose | |
| by more than 3σ; the two Polish sets were scored and logged but did not decide. **All 76 slots were accepted, none | |
| rejected.** | |
| ## Training data | |
| GoLLeM v6 250M was pretrained on a fixed pool of **25,000,424,215 tokens** (26,124,390 documents; tokenizer sha256 | |
| `0640d3bd…`); each token was seen at most once (24.90B of the 25.00B tokens were used). The pool is an aggregate of public | |
| sources; **each source keeps its upstream terms.** | |
| | part | share | tokens | origin | upstream license | | |
| |---|---:|---:|---|---| | |
| | English mix (ARC-MIX v5) | 35.79 % | 8.95 B | FineWeb-Edu, DCLM-baseline, FineMath, OpenWebMath, Wikipedia (EN), StackExchange, StarCoderData, LoC PD Books, Project Gutenberg, scientific papers, CC-News, UltraChat, WildChat, tiny-textbooks, OpenSubtitles (small), OpenStax textbooks | mixed: ODC-By, CC BY 4.0, CC BY-SA, CC0/PD, The Stack terms, **unknown** (CC-News, scientific papers), **model-generated** (UltraChat, WildChat, tiny-textbooks) | | |
| | Polish web (HPLT 3.0, filtered) | 60.20 % | 15.05 B | [HPLT/HPLT3.0](https://huggingface.co/datasets/HPLT/HPLT3.0), Polish, our quality filtering | compilation CC0-1.0; texts collected under TDM exception (EU DSM art. 4) with crawl-level opt-out | | |
| | Polish Wikipedia + Wolne Lektury (SA/FAL) | 1.94 % | 0.49 B | [SlayerLab/polish-dynaword](https://huggingface.co/datasets/SlayerLab/polish-dynaword) | CC BY-SA 3.0, CC BY-SA 3.0/4.0, Free Art License 1.3 | | |
| | Polish open: Europarl v7, Wolne Lektury (PD) | 0.33 % | 0.08 B | [statmt.org Europarl v7](https://www.statmt.org/europarl/), polish-dynaword | Europarl: no known copyright restrictions; public domain | | |
| | Biblioteka Nauki (CC BY-SA), e-mails masked | 0.60 % | 0.15 B | polish-dynaword `biblioteka_nauki` | CC BY-SA 4.0 | | |
| | Biblioteka Nauki (CC BY / PD), e-mails masked | 1.14 % | 0.28 B | polish-dynaword `biblioteka_nauki` | CC BY 4.0, public domain | | |
| - **No non-commercial (NC) sources.** Documents marked CC BY-NC-SA were removed from the English mix. | |
| - **Share-alike** sources are about 7 % of tokens. The model weights are released under **CC BY-SA 4.0**. | |
| - **Commercial use:** review upstream terms, in particular CC-News, scientific papers, StarCoderData (The Stack terms) | |
| and model-generated chat data (UltraChat, WildChat, tiny-textbooks: outputs of OpenAI models). | |
| - **Attribution:** the English mix includes text from 55 OpenStax textbooks (CC BY 4.0, © Rice University, | |
| https://openstax.org), obtained from `crumb/openstax-text` (revision `8f502ca4…`), modified (extracted, chunked, | |
| filtered). The full title list is in the appendix *OpenStax attribution* below, because the English mix itself is not | |
| published. | |
| - **Personal data:** in Biblioteka Nauki, e-mail addresses were masked before training (`[email]`). | |
| - **Benchmark decontamination:** documents matching WikiText-2 test/validation or ARC validation/test were removed from | |
| the English mix (4,362 documents). In the Polish part, documents containing an exact copy (after normalization) of any | |
| of 12,755 protected Polish evaluation sentences of at least 20 characters, including MultiBLiMP-pl, were dropped (738 | |
| documents). | |
| ### Data availability | |
| | part | status | | |
| |---|---| | |
| | upstream sources | public (links above) | | |
| | Polish Wikipedia, Wolne Lektury, Biblioteka Nauki | the exact input files are public in SlayerLab/polish-dynaword (sha256 match); our e-mail masking of Biblioteka Nauki is not published | | |
| | Polish web (HPLT, filtered) | **not published** in the exact form used (18 files); upstream HPLT 3.0 is public | | |
| | English mix (ARC-MIX v5) | **not published** as data; recipe described in the SlayerLab/gollem-v5-arcmix-9b card | | |
| | tokenized pool | not published | | |
| ## Usage | |
| The model is a custom PyTorch architecture, not a `transformers` class. The repository contains: | |
| | file | content | sha256 | | |
| |---|---|---| | |
| | `model.safetensors` | fp32 weights (252,760,019 parameters; output head stored as a copy of the tied embedding) | `d30b60d5f617a1d118e4ba363996c5e56ea57037b82dbb707382a0be29e7bbd0` | | |
| | `config.json` | architecture | `61d8594034918d7e49253f3801da79bae9f14f57ba871fccde6d307a01c73a8f` | | |
| | `tokenizer.json` | 32,768-token BPE (`tokenizers`) | `0640d3bd3674d7a6e59c540d945a2ad275a12e1bcaeb6007c37a90751f9815dd` | | |
| | `modeling_gollem_v6.py` | model code (inference) and a small `generate` helper | `ee2ea8c4e957413f3a7c1e8fe93a0a0b4431a2bcabaa6cc5ceeaea103fcd58ee` | | |
| ```python | |
| # pip install torch safetensors tokenizers huggingface_hub | |
| from huggingface_hub import snapshot_download | |
| import sys | |
| path = snapshot_download("SlayerLab/GoLLeM-v6-250M") # pin revision="<commit>" for reproducibility | |
| sys.path.insert(0, path) | |
| from modeling_gollem_v6 import load_gollem_v6, generate | |
| model, tok = load_gollem_v6(path) # strict load, eval mode, CPU by default | |
| print(generate(model, tok, "Ala ma kota", max_new_tokens=40, temperature=0.8, top_k=50, seed=1)) | |
| print(generate(model, tok, "Photosynthesis is the process by which", max_new_tokens=40)) | |
| ``` | |
| - This is a **base model**: it continues text; it does not answer questions or follow instructions. | |
| - Context: 1,024 tokens. RoPE uses the interleaved channel convention; if you port the model to another framework, | |
| keep that convention. | |
| - The module definitions in `modeling_gollem_v6.py` are those used in training (initialisation and training loss | |
| removed). Loading `model.safetensors` into them reproduces the logits of the training checkpoint exactly | |
| (checked with `torch.equal`, also on the files downloaded from this repository). | |
| ## Limitations | |
| - Small **base** model: not instruction-tuned, limited factual knowledge, may reproduce biases and errors present in web | |
| data. Not for production use. | |
| - **Single run, single seed.** Differences below about one standard error (or below the run-to-run noise we measured | |
| on smaller models, about 0.3 eff) should not be read as real. | |
| - English and Polish scores use **our** protocol; they are not official leaderboard results, and self-reported numbers | |
| of other models measured differently are not comparable. | |
| - **Self-description.** Prompted in chat format (ChatML), the model may describe itself as "an AI language model | |
| created by OpenAI". The English pretraining mix contains model-generated chat data (UltraChat, WildChat) with such | |
| self-descriptions. GoLLeM is **not** affiliated with OpenAI; it was built by Fabryka AI (author: Arkadiusz Słota). | |
| ## License | |
| **Weights: CC BY-SA 4.0.** Attribution: GoLLeM v6, Arkadiusz Słota / Fabryka AI, link to this repository; derivative | |
| weights under the same licence. | |
| Training data keep their upstream licences; see *Training data* above. | |
| ## Po polsku (skrót) | |
| GoLLeM v6 250M to bazowy (niedostrojony do poleceń) model językowy polsko-angielski, 252,8 mln parametrów, wytrenowany | |
| od zera na 24,9 mld tokenów (ok. 2/3 polski, 1/3 angielski). Wyniki według **naszego protokołu, nie oficjalne**: | |
| eff (angielski) **74,71**, eff-PL **79,77**, PL czyste **69,42 ± 0,58**. Warunek ustalony przed pierwszym pomiarem | |
| po polsku (PL czyste ≥ 68,62, czyli v4 − 1 pkt) jest spełniony. W polskim v6 jest **w granicach jednego błędu standardowego od | |
| GoLLeM v4** (który trenowano tylko po polsku), w angielskim poniżej GoLLeM v5 128M (tylko angielski). Model do badań, | |
| nie do zastosowań produkcyjnych. W formacie czatu (ChatML) model może przedstawiać się jako „model językowy stworzony | |
| przez OpenAI”, bo angielska część danych zawiera rozmowy wygenerowane modelami OpenAI (UltraChat, WildChat). GoLLeM nie | |
| jest powiązany z OpenAI; stworzyła go Fabryka AI (autor: Arkadiusz Słota). | |
| ## Appendix: OpenStax attribution | |
| The training data of this model includes text extracted from the following OpenStax textbooks, each licensed under the Creative Commons Attribution 4.0 International License (CC BY 4.0), © Rice University. Download for free at https://openstax.org. The texts were obtained from the Hugging Face dataset crumb/openstax-text (revision 8f502ca45f9f05cb5673eae445b7a97a4e8c4349). Modified: text extracted, chunked and filtered (chunks matching benchmark test/validation sets were removed). | |
| | Title (from source file name) | © year | | |
| |---|---| | |
| | APBiology | 2018 | | |
| | APCollege Physics | 2017 | | |
| | APMacroeconomics 2e | 2017 | | |
| | APMicroeconomics 2e | 2017 | | |
| | Algebra and Trigonometry 2e | 2021 | | |
| | American Government 3e | 2021 | | |
| | Anatomy and Physiology 2e | 2022 | | |
| | Anatomyand Physiology | 2017 | | |
| | Astronomy 2e | 2022 | | |
| | Astronomy | 2018 | | |
| | Biology 2e | 2020 | | |
| | Business Ethics | 2018 | | |
| | Chemistry 2e | 2019 | | |
| | Chemistry Atoms First 2e | 2019 | | |
| | College Algebra 2e | 2021 | | |
| | College Algebra Corequisite Support 2e | 2021 | | |
| | College Physics | 2020 | | |
| | College Physics 2e | 2022 | | |
| | College Physics for AP Courses 2e | 2022 | | |
| | College Success | 2020 | | |
| | College Success | 2023 | | |
| | Concepts Biology | 2017 | | |
| | Contemporary Mathematics | 2023 | | |
| | Economics 2e | 2018 | | |
| | Economics 3e | 2022 | | |
| | Elementary Algebra 2e | 2020 | | |
| | Entrepreneurship | 2020 | | |
| | Intermediate Algebra 2e | 2020 | | |
| | Introduction to Intellectual Property | n/a | | |
| | Introduction to Philosophy | 2022 | | |
| | Introduction to Political Science | 2022 | | |
| | Introductionto Anthropology | 2022 | | |
| | Introductionto Sociology 3e | 2021 | | |
| | Introductory Business Statistics | 2018 | | |
| | Introductory Statistics | 2018 | | |
| | Macroeconomics 2e | 2018 | | |
| | Macroeconomics 3e | 2022 | | |
| | Microbiology | 2021 | | |
| | Microeconomics 2e | 2018 | | |
| | Microeconomics 3e | 2022 | | |
| | Physics | n/a | | |
| | Prealgebra 2e | 2020 | | |
| | Precalculus 2e | 2021 | | |
| | Preparing for College Success | 2023 | | |
| | Principles Marketing | 2023 | | |
| | Principlesof Finance | 2022 | | |
| | Psychology 2e | 2020 | | |
| | Statistics | n/a | | |
| | USHistory | 2021 | | |
| | University Physics Vol 1 | 2021 | | |
| | University Physics Volume 2 | 2021 | | |
| | University Physics Volume 3 | 2021 | | |
| | World History Volume 1 | 2023 | | |
| | World History Volume 2 | 2022 | | |
| | Writing Guide | 2021 | | |
| Titles licensed CC BY-NC-SA 4.0, non-English titles, and three CC BY 4.0 titles whose text contains elements marked ‘CC BY-NC-SA’ (Introduction to Business, Organizational Behavior, Principles of Management) were not used. | |
| --- | |
| *GoLLeM v6 — Fabryka AI. **Author: Arkadiusz Słota.** Research base model (final checkpoint).* | |