File size: 2,069 Bytes
41f5760
f879c07
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
41f5760
 
f879c07
41f5760
f879c07
41f5760
 
 
f879c07
 
 
 
 
41f5760
f879c07
41f5760
f879c07
 
 
 
 
41f5760
f879c07
41f5760
f879c07
41f5760
f879c07
 
41f5760
f879c07
41f5760
f879c07
 
41f5760
f879c07
 
 
 
 
41f5760
f879c07
41f5760
f879c07
 
 
 
41f5760
f879c07
 
41f5760
f879c07
 
41f5760
f879c07
41f5760
f879c07
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
---
license: mit
base_model: gpt2
tags:
  - text-generation
  - causal-lm
  - gpt
  - transformer
  - decoder-only
  - tiny-stories
  - story
  - children
  - stories
  - narrative
  - lm
  - language-model
  - pytorch
  - 29m
  - warmth
  - wisdom
  - inept-vocab
  - story-generation
  - causal-lm-pretraining
language:
  - en
pipeline_tag: text-generation
widget:
  - text: "One day, a little girl named Lily"
---

# StoryGPT

A small [GPT](https://arxiv.org/abs/1710.11815)-style causal language model trained from scratch on a 50k excerpt of the [TinyStories](https://huggingface.co/datasets/roneneldan/TinyStories) dataset.

## Model Details

- **Architecture:** decoder-only transformer with causal self-attention and pre-LN blocks
- **Parameters:** ~29M
- **Vocabulary:** GPT-2 tokenizer (50,257 tokens)
- **Context length:** 512
- **Layers:** 4, **Embedding dim:** 256

## Training

- **Dataset:** TinyStories, 50,000 stories (44 MB of text)
- **Block size:** 128 tokens
- **Batch size:** 16
- **Optimizer:** AdamW, lr = 3e-4
- **Loss:** Cross-entropy (causal LM)

## Usage

Direct loading with `transformers` — no local files needed:

```python
from transformers import GPT2Tokenizer, AutoConfig, AutoModel

tokenizer = GPT2Tokenizer.from_pretrained("coderian/StoryGPT")

model = AutoModel.from_pretrained("coderian/StoryGPT")
model.eval()

text = "One day, a little girl named Lily"
inputs = tokenizer(text, return_tensors="pt")
outputs = model.generate(input_ids=inputs["input_ids"], max_new_tokens=50)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
```

Because `StoryGPT` is a custom model, register its config class once if autoloading fails:

```python
from transformers import AutoConfig, AutoModel
from config_gpt import GPTConfig
from models import StoryGPT

AutoConfig.register("my_gpt", GPTConfig)
AutoModel.register(GPTConfig, StoryGPT)

model = AutoModel.from_pretrained("coderian/StoryGPT")
```

## Limitations

Tiny model trained for only 2 epochs. Stories may be short, repetitive, or contain small errors in grammar or logic.