StoryLM-10M

A ~12.6M parameter LLaMA-style causal language model trained from scratch on TinyStories.

Architecture

Parameter Value
d_model 256
n_heads 8 (MHA)
n_kv_heads 8
n_layers 8
head_dim 32
FFN dim 1024 (SwiGLU, 4x)
Vocab size 8192
Context length 512
Tied embeddings Yes
Total params 12,603,648

Training

  • Data: TinyStories (2,119,719 stories, 533.7M tokens)
  • Steps: 24,400 (batch 4, seq 512)
  • Optimizer: AdamW, LR 3e-4, cosine schedule, 500-step warmup
  • Best val loss: 1.8635
  • Val perplexity: 6.45

Generation Samples

Prompt: Once upon a time, there was a little
Output: Once upon a time, there was a little

Prompt: The cat sat on the
Output: The cat sat on the

Prompt: In the beginning, the world was
Output: In the beginning, the world was

Prompt: A small robot named
Output: A small ro bot named

Prompt: Every morning, the sun
Output: Every morning, the sun

Notes

  • Trained on a single RTX 5090 (32 GB) in ~281 seconds of training time.
  • The model is a custom implementation (not HuggingFace transformers-compatible out of the box).
  • Weights are stored as a single PyTorch checkpoint (model.pt).
Downloads last month
369
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support