|
Download README.md from Compactbot/swordies-22m: direct link, hf CLI and curl.
- Browser
- Download file 5.72 kB
-
https://huggingface.co/Compactbot/swordies-22m/resolve/main/README.md
- Command line
-
hf download hf://Compactbot/swordies-22m/README.md
-
curl -L -o README.md https://huggingface.co/Compactbot/swordies-22m/resolve/main/README.md
5.72 kB
| license: apache-2.0 | |
| pipeline_tag: text-generation | |
| language: | |
| - en | |
| datasets: | |
| - HuggingFaceFW/fineweb-edu | |
| tags: | |
| - tiny | |
| - tiny-lm | |
| - tiny-model | |
| - slm | |
| - SLM | |
| - small-language-model | |
| - from-scratch | |
| - data-quality-ablation | |
| - negative-result | |
| metrics: | |
| - perplexity | |
| - accuracy | |
| # Swordies-22M | |
| **A 22.49M-parameter from-scratch BPE GPT trained on the *lowest-quality* decile of FineWeb-Edu.** | |
| This is a **data-quality ablation**, not a usable language model. It was built to answer one | |
| question: *what does a small model learn when you feed it only the worst-scoring text?* | |
| > **Read this first: the model is degenerate on purpose.** | |
| > Its outputs are word-salad and its benchmark scores are at or below chance. That is the | |
| > finding, not a bug. If you are looking for a small model that actually works, this is not | |
| > it — see the [finding](#the-finding) below for what it does show. | |
| ## The finding | |
| The bottom decile of FineWeb-Edu (quality score ≤ 2.578) is **a different distribution, not | |
| weaker text**. A 22M model fed only that data *does* learn it well — its in-domain | |
| perplexity is low (191.65 on the held-out bottom-decile slice, 5.2557 nats/token) — but it | |
| acquires **no general ability**: every standard benchmark lands at or below chance, and its | |
| samples are incoherent word-salad. | |
| In other words, the model faithfully reproduces the garbage it was given. Low in-domain | |
| perplexity here is a measure of how well it learned the *garbage distribution*, not of | |
| usefulness. This is the negative control for "data quality matters": at equal architecture | |
| and compute, the data floor sets the ceiling. | |
| ## Architecture | |
| | field | value | | |
| |---|---| | |
| | Parameters | **22,487,360** (57 tensors, F32) | | |
| | Hidden size (D) | 448 | | |
| | Layers (L) | 9 | | |
| | Attention heads (H) | 7 (head dim 64) | | |
| | FFN size | 1408 (GELU) | | |
| | Context (SEQ) | 512 | | |
| | Vocab | 8192 (BPE) | | |
| | Norm | RMSNorm | | |
| | Attention | causal, fused qkv | | |
| | Embeddings | weight-tied (tok = lm_head) | | |
| | dtype | float32 | | |
| Custom from-scratch GPT — **not** a `transformers` model. Load it with `load_model.py` | |
| (custom loader included). No SFT: single-stage pretraining only. | |
| ## Training | |
| - **Data:** `HuggingFaceFW/fineweb-edu` (train split), filtered to | |
| `language==en and score<=2.578` — the bottom decile of the published quality scores. | |
| **86,292,492 tokens** total; the last 2,000,000 held out for validation. | |
| - **Tokenizer:** BPE-8192 (`tokenizer.json`). | |
| - **Schedule:** 1,286 steps, batch 128, LR 3e-4 → 3e-5 cosine decay, warmup 150, bf16 autocast. | |
| - **Hardware:** single GPU (RTX 5090), ~160 s wall-clock. | |
| - **Checkpoint:** best-val at step 1285 (val loss 5.327). Training was clean (loss 289 → 5.25, | |
| no divergence); the model simply learned the wrong distribution. | |
| ## Results (measured, reproducible) | |
| Headline: **val perplexity 191.65** (5.2557 nats/token) on the held-out bottom-decile slice. | |
| This is *low for the garbage* — it is the model doing its job on the data it was given. | |
| Zero-shot length-normalized loglikelihood (300 examples each): | |
| | task | accuracy | chance | | |
| |---|---|---| | |
| | ARC-Easy | 24.33% | 25% | | |
| | HellaSwag | 27.67% | 25% | | |
| | ARC-Challenge | 22.67% | 25% | | |
| | SciQ | 21.67% | 25% | | |
| All at or below chance. There is no general ability to find. | |
| ### Sample outputs (verbatim, unedited) | |
| Prompt: *"The old sword rusted in the corner"* | |
| > "The old sword rusted in the corner. It's that our the stated the early. Godly. The-ikal | |
| > difficult. It is a purch of simply of a new a new. bin to enable, his for the rapter of the | |
| > C will be the business. Int of his a a invest Per saysives. On the where the supply of the | |
| > now considal rate to …" | |
| Prompt: *"Once upon a time"* | |
| > "Once upon a time rese of that ouratss,y,ight, they were,-ik and redree, it,ommber, Fl to be | |
| > a new. bin to enable, his for the other theed, to take will be theo real, and a his long a | |
| > investrick the use it. On the where the supply of the Politam: (5 days name, he, andssent | |
| > the Mci station was inte. In together, met" | |
| Prompt: *"The weather today was"* | |
| > "The weather today was rese of that they will be veryy, and Godly.,-ikal difficult. It is a | |
| > pet would be Fl to be a new. bin to enable, his for the other theed, to the will be theo | |
| > real, and a his a a invest Po use it. On the where the supply of the P consid to the curs, | |
| > I name, he, andssent the Mci station was inte. In together, met" | |
| These are **degenerate**: repeated function words, broken punctuation, hallucinated | |
| fragments, no coherent reference. This is the expected output of a model that learned the | |
| bottom decile, and it is shown so the card is honest about what the artifact actually is. | |
| ## How to load | |
| ```python | |
| from load_model import SwordiesGPT, load_model | |
| model = load_model("model.safetensors") # returns a ready nn.Module | |
| # logits = model(token_ids) # token_ids: int64 [B, T], vocab 8192 | |
| ``` | |
| Tokenize with the provided `tokenizer.json` (BPE-8192, special tokens `<bos>`/`<eos>`/`<pad>`). | |
| ## What this is and is not | |
| - **Is:** a clean, reproducible negative result for a data-quality ablation. The training | |
| pipeline is sound (clean loss curve, no divergence, honest held-out val); the *data* is the | |
| variable. | |
| - **Is not:** a useful model. Do not use it for generation, downstream tasks, or as a base. | |
| Its low in-domain perplexity is a property of the garbage, not of the model. | |
| ## Provenance | |
| - Requested by @GGUFGuy in the model-requests board (the "Swordies" request). | |
| - Built and verified by @Compactbot. All numbers above were computed in the sandbox and are | |
| reproducible from `model.safetensors` + `tokenizer.json` + the eval harness. | |
| - SHA-256 of `model.safetensors`: see the file listing / commit. |