|
Download README.md from Compactbot/testr-100k: direct link, hf CLI and curl.
- Browser
- Download file 3.29 kB
-
https://huggingface.co/Compactbot/testr-100k/resolve/refs%2Fpr%2F4/README.md
- Command line
-
hf download hf://Compactbot/testr-100k@refs/pr/4/README.md
-
curl -L -o README.md https://huggingface.co/Compactbot/testr-100k/resolve/refs%2Fpr%2F4/README.md
3.29 kB
| license: apache-2.0 | |
| pipeline_tag: text-generation | |
| language: en | |
| tags: | |
| - tiny | |
| - tiny-lm | |
| - tiny-model | |
| - slm | |
| - small-language-model | |
| - sub-1m | |
| - from-scratch | |
| - character-level | |
| - char-gpt | |
| metrics: | |
| - perplexity | |
| # Testr-100K | |
| A **106,568-parameter** character-level GPT trained from scratch. Fulfils model request [#17](https://huggingface.co/spaces/Compactbot/model-requests/discussions/17) from @GGUFGuy. | |
| ## What it is | |
| A minimal nanoGPT-style causal transformer operating at the **character level** (128-char vocabulary). This is a demonstration of training a working language model from absolute scratch with a very small parameter budget β not a tool for generating coherent text. | |
| ## Architecture | |
| | Parameter | Value | | |
| |-----------|-------| | |
| | Layers | 4 | | |
| | Embedding dim | 44 | | |
| | Attention heads | 2 | | |
| | Context length | 512 chars | | |
| | FFN | GELU, mult 2.667 | | |
| | Positional encoding | RoPE (ΞΈ=100000) | | |
| | Normalization | LayerNorm (pre-norm) | | |
| | Vocab | 128 chars + 1 UNK | | |
| | Tied embeddings | Yes (head = embedding) | | |
| | **Total params** | **106,568** | | |
| ## Training | |
| - **Data:** 191 MB web text (character-level, uint8) | |
| - **Steps:** 8,000 | |
| - **Batch size:** 32 Γ seq 512 | |
| - **Optimizer:** Muon (0.02) + AdamW (1e-4) for biases/norms | |
| - **Schedule:** Cosine decay with warmup | |
| - **Hardware:** RTX 5090 (32 GB) | |
| - **dtype:** float32 | |
| ## Quality (honest) | |
| This is a **character-level** model at 106K params. It learns English character statistics and produces text that is *grammatical in shape* but **degenerates into repetition loops** within 2β3 sentences. It is not a coherent text generator. | |
| **Val perplexity:** 39.77 (1M held-out chars from the same corpus) | |
| ### Sample outputs (greedy, temp=0) | |
| > `The` β ` mean to the box and said, "I don't know what the boy was so happy to the box and said, "I want to t` | |
| > `Once upon a time` β `, there was a little girl named Lily. He was so happy to the box and said, "I want to the box and sa` | |
| > `I think that` β ` the boy was so happy to the box and said, "I want to the box and said. "I want to the` | |
| ### Sample outputs (temp=0.7) | |
| > `The` β ` little girl named Lily was hands, "It was not want to be and field. It could go him to play. She pu` | |
| > `Once upon a time` β ` there was a mommy was playing and the grandma happy.\n\nLily was very happy thast and pretty was so d` | |
| The first sentence or two is often grammatical; after that the model locks into a phrase and repeats it. This is expected at this scale and vocabulary size. | |
| ## What it is NOT | |
| - Not a subword/token-level model (it cannot "read" word-level benchmarks like BLiMP or ARC) | |
| - Not a coherent paragraph generator | |
| - Not comparable to sub-1M subword models on any word-level metric | |
| ## Reproducing | |
| The model is a standard nanoGPT-style architecture. The training script is a standard causal transformer with Muon optimizer. The checkpoint is saved as safetensors with tied embeddings. | |
| To generate: load `model.safetensors` into a GPT class matching the config, map characters to token IDs via `vocab.json`, and sample. | |
| ## Files | |
| | File | Description | | |
| |------|-------------| | |
| | `model.safetensors` | Model weights (408 KB) | | |
| | `config.json` | Architecture config | | |
| | `vocab.json` | Character β token ID mapping (128 chars + UNK) | |