|
Download README.md from coderian/StoryGPT: direct link, hf CLI and curl.
- Browser
- Download file 2.07 kB
-
https://huggingface.co/coderian/StoryGPT/resolve/main/README.md
- Command line
-
hf download hf://coderian/StoryGPT/README.md
-
curl -L -o README.md https://huggingface.co/coderian/StoryGPT/resolve/main/README.md
2.07 kB
metadata
license: mit
base_model: gpt2
tags:
- text-generation
- causal-lm
- gpt
- transformer
- decoder-only
- tiny-stories
- story
- children
- stories
- narrative
- lm
- language-model
- pytorch
- 29m
- warmth
- wisdom
- inept-vocab
- story-generation
- causal-lm-pretraining
language:
- en
pipeline_tag: text-generation
widget:
- text: One day, a little girl named Lily
StoryGPT
A small GPT-style causal language model trained from scratch on a 50k excerpt of the TinyStories dataset.
Model Details
- Architecture: decoder-only transformer with causal self-attention and pre-LN blocks
- Parameters: ~29M
- Vocabulary: GPT-2 tokenizer (50,257 tokens)
- Context length: 512
- Layers: 4, Embedding dim: 256
Training
- Dataset: TinyStories, 50,000 stories (44 MB of text)
- Block size: 128 tokens
- Batch size: 16
- Optimizer: AdamW, lr = 3e-4
- Loss: Cross-entropy (causal LM)
Usage
Direct loading with transformers — no local files needed:
from transformers import GPT2Tokenizer, AutoConfig, AutoModel
tokenizer = GPT2Tokenizer.from_pretrained("coderian/StoryGPT")
model = AutoModel.from_pretrained("coderian/StoryGPT")
model.eval()
text = "One day, a little girl named Lily"
inputs = tokenizer(text, return_tensors="pt")
outputs = model.generate(input_ids=inputs["input_ids"], max_new_tokens=50)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Because StoryGPT is a custom model, register its config class once if autoloading fails:
from transformers import AutoConfig, AutoModel
from config_gpt import GPTConfig
from models import StoryGPT
AutoConfig.register("my_gpt", GPTConfig)
AutoModel.register(GPTConfig, StoryGPT)
model = AutoModel.from_pretrained("coderian/StoryGPT")
Limitations
Tiny model trained for only 2 epochs. Stories may be short, repetitive, or contain small errors in grammar or logic.