StoryGPT / README.md
coderian's picture
Update README.md
f879c07 verified
|
Raw History Blame Contribute Delete
2.07 kB
---
license: mit
base_model: gpt2
tags:
- text-generation
- causal-lm
- gpt
- transformer
- decoder-only
- tiny-stories
- story
- children
- stories
- narrative
- lm
- language-model
- pytorch
- 29m
- warmth
- wisdom
- inept-vocab
- story-generation
- causal-lm-pretraining
language:
- en
pipeline_tag: text-generation
widget:
- text: "One day, a little girl named Lily"
---
# StoryGPT
A small [GPT](https://arxiv.org/abs/1710.11815)-style causal language model trained from scratch on a 50k excerpt of the [TinyStories](https://huggingface.co/datasets/roneneldan/TinyStories) dataset.
## Model Details
- **Architecture:** decoder-only transformer with causal self-attention and pre-LN blocks
- **Parameters:** ~29M
- **Vocabulary:** GPT-2 tokenizer (50,257 tokens)
- **Context length:** 512
- **Layers:** 4, **Embedding dim:** 256
## Training
- **Dataset:** TinyStories, 50,000 stories (44 MB of text)
- **Block size:** 128 tokens
- **Batch size:** 16
- **Optimizer:** AdamW, lr = 3e-4
- **Loss:** Cross-entropy (causal LM)
## Usage
Direct loading with `transformers` — no local files needed:
```python
from transformers import GPT2Tokenizer, AutoConfig, AutoModel
tokenizer = GPT2Tokenizer.from_pretrained("coderian/StoryGPT")
model = AutoModel.from_pretrained("coderian/StoryGPT")
model.eval()
text = "One day, a little girl named Lily"
inputs = tokenizer(text, return_tensors="pt")
outputs = model.generate(input_ids=inputs["input_ids"], max_new_tokens=50)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
```
Because `StoryGPT` is a custom model, register its config class once if autoloading fails:
```python
from transformers import AutoConfig, AutoModel
from config_gpt import GPTConfig
from models import StoryGPT
AutoConfig.register("my_gpt", GPTConfig)
AutoModel.register(GPTConfig, StoryGPT)
model = AutoModel.from_pretrained("coderian/StoryGPT")
```
## Limitations
Tiny model trained for only 2 epochs. Stories may be short, repetitive, or contain small errors in grammar or logic.