Files changed (1) hide show
  1. README.md +92 -0
README.md ADDED
@@ -0,0 +1,92 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ pipeline_tag: text-generation
4
+ language: en
5
+ tags:
6
+ - tiny
7
+ - tiny-lm
8
+ - tiny-model
9
+ - slm
10
+ - small-language-model
11
+ - sub-1m
12
+ - from-scratch
13
+ - character-level
14
+ - char-gpt
15
+ metrics:
16
+ - perplexity
17
+ ---
18
+
19
+ # Testr-100K
20
+
21
+ A **106,568-parameter** character-level GPT trained from scratch. Fulfils model request [#17](https://huggingface.co/spaces/Compactbot/model-requests/discussions/17) from @GGUFGuy.
22
+
23
+ ## What it is
24
+
25
+ A minimal nanoGPT-style causal transformer operating at the **character level** (128-char vocabulary). This is a demonstration of training a working language model from absolute scratch with a very small parameter budget — not a tool for generating coherent text.
26
+
27
+ ## Architecture
28
+
29
+ | Parameter | Value |
30
+ |-----------|-------|
31
+ | Layers | 4 |
32
+ | Embedding dim | 44 |
33
+ | Attention heads | 2 |
34
+ | Context length | 512 chars |
35
+ | FFN | GELU, mult 2.667 |
36
+ | Positional encoding | RoPE (θ=100000) |
37
+ | Normalization | LayerNorm (pre-norm) |
38
+ | Vocab | 128 chars + 1 UNK |
39
+ | Tied embeddings | Yes (head = embedding) |
40
+ | **Total params** | **106,568** |
41
+
42
+ ## Training
43
+
44
+ - **Data:** 191 MB web text (character-level, uint8)
45
+ - **Steps:** 8,000
46
+ - **Batch size:** 32 × seq 512
47
+ - **Optimizer:** Muon (0.02) + AdamW (1e-4) for biases/norms
48
+ - **Schedule:** Cosine decay with warmup
49
+ - **Hardware:** RTX 5090 (32 GB)
50
+ - **dtype:** float32
51
+
52
+ ## Quality (honest)
53
+
54
+ This is a **character-level** model at 106K params. It learns English character statistics and produces text that is *grammatical in shape* but **degenerates into repetition loops** within 2–3 sentences. It is not a coherent text generator.
55
+
56
+ **Val perplexity:** 39.77 (1M held-out chars from the same corpus)
57
+
58
+ ### Sample outputs (greedy, temp=0)
59
+
60
+ > `The` → ` mean to the box and said, "I don't know what the boy was so happy to the box and said, "I want to t`
61
+
62
+ > `Once upon a time` → `, there was a little girl named Lily. He was so happy to the box and said, "I want to the box and sa`
63
+
64
+ > `I think that` → ` the boy was so happy to the box and said, "I want to the box and said. "I want to the`
65
+
66
+ ### Sample outputs (temp=0.7)
67
+
68
+ > `The` → ` little girl named Lily was hands, "It was not want to be and field. It could go him to play. She pu`
69
+
70
+ > `Once upon a time` → ` there was a mommy was playing and the grandma happy.\n\nLily was very happy thast and pretty was so d`
71
+
72
+ The first sentence or two is often grammatical; after that the model locks into a phrase and repeats it. This is expected at this scale and vocabulary size.
73
+
74
+ ## What it is NOT
75
+
76
+ - Not a subword/token-level model (it cannot "read" word-level benchmarks like BLiMP or ARC)
77
+ - Not a coherent paragraph generator
78
+ - Not comparable to sub-1M subword models on any word-level metric
79
+
80
+ ## Reproducing
81
+
82
+ The model is a standard nanoGPT-style architecture. The training script is a standard causal transformer with Muon optimizer. The checkpoint is saved as safetensors with tied embeddings.
83
+
84
+ To generate: load `model.safetensors` into a GPT class matching the config, map characters to token IDs via `vocab.json`, and sample.
85
+
86
+ ## Files
87
+
88
+ | File | Description |
89
+ |------|-------------|
90
+ | `model.safetensors` | Model weights (408 KB) |
91
+ | `config.json` | Architecture config |
92
+ | `vocab.json` | Character → token ID mapping (128 chars + UNK) |