File size: 2,069 Bytes
41f5760 f879c07 41f5760 f879c07 41f5760 f879c07 41f5760 f879c07 41f5760 f879c07 41f5760 f879c07 41f5760 f879c07 41f5760 f879c07 41f5760 f879c07 41f5760 f879c07 41f5760 f879c07 41f5760 f879c07 41f5760 f879c07 41f5760 f879c07 41f5760 f879c07 41f5760 f879c07 41f5760 f879c07 41f5760 f879c07 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 | ---
license: mit
base_model: gpt2
tags:
- text-generation
- causal-lm
- gpt
- transformer
- decoder-only
- tiny-stories
- story
- children
- stories
- narrative
- lm
- language-model
- pytorch
- 29m
- warmth
- wisdom
- inept-vocab
- story-generation
- causal-lm-pretraining
language:
- en
pipeline_tag: text-generation
widget:
- text: "One day, a little girl named Lily"
---
# StoryGPT
A small [GPT](https://arxiv.org/abs/1710.11815)-style causal language model trained from scratch on a 50k excerpt of the [TinyStories](https://huggingface.co/datasets/roneneldan/TinyStories) dataset.
## Model Details
- **Architecture:** decoder-only transformer with causal self-attention and pre-LN blocks
- **Parameters:** ~29M
- **Vocabulary:** GPT-2 tokenizer (50,257 tokens)
- **Context length:** 512
- **Layers:** 4, **Embedding dim:** 256
## Training
- **Dataset:** TinyStories, 50,000 stories (44 MB of text)
- **Block size:** 128 tokens
- **Batch size:** 16
- **Optimizer:** AdamW, lr = 3e-4
- **Loss:** Cross-entropy (causal LM)
## Usage
Direct loading with `transformers` — no local files needed:
```python
from transformers import GPT2Tokenizer, AutoConfig, AutoModel
tokenizer = GPT2Tokenizer.from_pretrained("coderian/StoryGPT")
model = AutoModel.from_pretrained("coderian/StoryGPT")
model.eval()
text = "One day, a little girl named Lily"
inputs = tokenizer(text, return_tensors="pt")
outputs = model.generate(input_ids=inputs["input_ids"], max_new_tokens=50)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
```
Because `StoryGPT` is a custom model, register its config class once if autoloading fails:
```python
from transformers import AutoConfig, AutoModel
from config_gpt import GPTConfig
from models import StoryGPT
AutoConfig.register("my_gpt", GPTConfig)
AutoModel.register(GPTConfig, StoryGPT)
model = AutoModel.from_pretrained("coderian/StoryGPT")
```
## Limitations
Tiny model trained for only 2 epochs. Stories may be short, repetitive, or contain small errors in grammar or logic. |