File size: 7,465 Bytes
3c62626
 
 
116cf2b
3c62626
116cf2b
 
 
 
 
 
3c62626
 
 
 
116cf2b
 
 
 
3c62626
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
116cf2b
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
---
license: mit
language:
- en
tags:
- text-generation
- causal-lm
- gpt
- tinygpt
- transformer
- llm
library_name: transformers
pipeline_tag: text-generation
base_model: gpt2
widget:
- text: Once upon a time, a little boy named Sam
  example_title: Story start
- text: The cat and the dog decided to
  example_title: Animals
---

# TinyGPT

A small-scale, **trained-from-scratch** GPT language model. Based on the GPT-2 architecture and fully compatible with `transformers`; downloadable via the Hub, with model source code included in this repo.

- **Architecture:** Decoder-only Transformer (custom, Chapter-1 style)
- **Parameters:** ~29M (`28,974,161`)
- **Context length (`max_seq_len`):** 128 tokens
- **Tokenizer:** GPT-2 (vocab size = 50,257, BPE)
- **Training data:** First 50,000 stories from [TinyStories](https://huggingface.co/datasets/roneneldan/TinyStories)
- **Inference memory (fp32 weights):** ~116MB -- runs comfortably on a single GPU or even CPU

---

## Installation

```bash
pip install transformers torch
```

---

## Usage (with transformers)

TinyGPT uses a custom architecture, so the model code is loaded from this repo at runtime. This requires the **`trust_remote_code=True`** flag.

### 1) Simplest approach: Auto classes

```python
from transformers import AutoTokenizer, AutoModelForCausalLM

repo = "coderian/TinyGPT"

tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(
    repo,
    trust_remote_code=True,   # loads custom architecture code from repo
)

model.eval()

# Generation (sampling)
prompt = "Once upon a time, a little girl named Lily found"
inputs = tokenizer(prompt, return_tensors="pt")

output = model.generate(
    **inputs,
    max_new_tokens=50,
    do_sample=True,
    temperature=0.8,
    top_k=50,
    top_p=0.95,
)

text = tokenizer.decode(output[0], skip_special_tokens=True)
print(text)
```

### 2) Direct class (code-based usage)

If you copied the source code into your own project, `trust_remote_code` is not needed:

```python
from models.tinygpt import TinyGPT

model = TinyGPT.from_pretrained("coderian/TinyGPT", device="cpu")
```

### 3) GPU placement and half-precision

```python
model = AutoModelForCausalLM.from_pretrained(
    repo,
    trust_remote_code=True,
    torch_dtype="auto",          # uses config dtype (fp32)
    device_map="auto",           # places model on GPU if available
)

# Lower memory usage with bf16/fp16:
model = AutoModelForCausalLM.from_pretrained(
    repo,
    trust_remote_code=True,
    torch_dtype="bf16",          # or "float16"
    device_map="auto",
)
```

### 4) Simple greedy chat loop

```python
def reply(prompt, max_new_tokens=40):
    inputs = tokenizer(prompt, return_tensors="pt")
    out = model.generate(**inputs, max_new_tokens=max_new_tokens)
    return tokenizer.decode(out[0], skip_special_tokens=True)

print(reply("One rainy day, Ben and his dog Rex decided to"))
```

> **Note:** The model was trained exclusively on **English** TinyStories data. It works best with simple English story openings.

---

## Repo contents

| File | Description |
|---|---|
| `model.safetensors` | Trained weights (~116MB, fp32) |
| `config.json` | Model configuration (`TinyGPTConfig`) |
| `generation_config.json` | Default `generate()` parameters |
| `tokenizer.json`, `tokenizer_config.json` | GPT-2 tokenizer |
| `models/` | Open-source model code (config, tinygpt, transformer_block, attention) |

---

## Architecture

TinyGPT is a **decoder-only Transformer**. Input tokens pass through `token + position embedding`, are processed by `num_layers` Transformer blocks, and finally a LayerNorm + linear head produces logits over the vocabulary.

```
Input tokens (B, T)
      |
      v
+-----------------------------------+
|  token_embedding  (50,257 -> d)    |  nn.Embedding
|  position_embedding (T -> d)       |  nn.Embedding
|  x = token_emb + pos_emb           |
+------------------+-----------------+
                   v
   +-----------------------------+  x num_layers (4)
   |  LN1 -> CausalSelfAttn      |  Pre-LN, residual
   |       +                     |
   |  LN2 -> FFN (x4)            |  GELU activation
   |       +                     |
   +-----------------------------+
                   v
            ln_f (LayerNorm)
                   v
        lm_head (d -> 50,257) --> logits (B, T, V)
```

### 1) Embedding layers

- **`token_embedding`**: maps each token ID to a `d=256` dimensional vector. `W_e` in R^(50257 x 256)
- **`position_embedding`**: adds positional information. Consecutive positions use the same learned embedding; `W_p` in R^(128 x 256)

Input vector: `x = W_e[input_ids] + W_p[positions]`

### 2) Causal Self-Attention (masked scaled dot-product)

Each position can only attend to **previous** tokens (causal mask).

```
Q = x W_q ,  K = x W_k ,  V = x W_v          W in R^(256x256)

scores     = (Q K^T) / sqrt(d)                 d = 256
scores     = masked_fill(upper_triangle, -inf)  <- blocks future tokens
attn       = softmax(scores, dim=-1)
output     = attn @ V ,  then  output W_o
```

- Scaling by `1/sqrt(d)`: keeps dot-product variance stable at high dimensions,
  preventing softmax saturation.
- `nn.Embedding` and `nn.Linear` weights are initialized from `N(0, 0.02)`.

### 3) Residual connections and Pre-LayerNorm

LayerNorm is applied before each sublayer, followed by a residual add:

```
x = x + attention(ln1(x))
x = x + ffn(ln2(x))
```

Pre-LN keeps training stable in deeper models.

### 4) FFN (Feed-Forward Network)

```
FFN(x) = W2 * GELU(W1 * x + b1) + b2

W1: 256 -> 1024   (expansion x4)
GELU: activation
W2: 1024 -> 256   (contraction)
```

### 5) Output

```
logits = lm_head(ln_f(x))          ln_f: LayerNorm(256)
                                    lm_head: Linear(256 -> 50,257)
```

During training, `CrossEntropyLoss(logits, shift_right(y))` is used for the language modeling objective.
During inference, the `lm_head` logits directly represent the next-token distribution.

### Parameter breakdown

| Component | Parameters | Notes |
|---|---|---|
| `token_embedding` | 12,865,792 | 50,257 x 256 |
| `position_embedding` | 32,768 | 128 x 256 |
| 1 Transformer block | 789,760 | LN + attn + FFN |
| 4 Transformer blocks | 3,159,040 | total body |
| `ln_f` + `lm_head` | 12,916,561 | language head |

**Total: ~28.97M.** Most parameters reside in the embedding/head due to the large vocabulary; the actual reasoning capacity (transformer body) is ~3.1M.

---

## Training details

| Metric | Value |
|---|---|
| Data | TinyStories (first 50,000 stories, ~11.1M tokens) |
| Training samples | 86,636 blocks x 128 tokens |
| `embed_dim` | 256 |
| `num_layers` | 4 |
| `max_seq_len` | 128 |
| Optimizer | AdamW, lr `3e-4` |
| Batch size | 24 |
| Epochs | 2 |
| Duration | Short training run on CPU |

---

## Limitations & Notes

- English only; produces nonsensical output for non-English input.
- 128-token context window; insufficient for tasks requiring long-range context.
- Due to its small size, the model does not perform logic, recall factual knowledge, or produce coherent long-form text.
- Requires `trust_remote_code=True` to run; the code is open for inspection in this repo's `models/` directory.
- No weight decay or other regularization was applied during training.

---

## License

MIT -- model weights and code are freely usable.

**HF Hub:** [huggingface.co/coderian/TinyGPT](https://huggingface.co/coderian/TinyGPT)