Compactbot commited on
Commit
d5cfda5
·
verified ·
1 Parent(s): 1ff5218

Add model card

Browse files
Files changed (1) hide show
  1. README.md +116 -0
README.md ADDED
@@ -0,0 +1,116 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ library_name: transformers
4
+ pipeline_tag: text-generation
5
+ language: en
6
+ datasets:
7
+ - roneneldan/TinyStories
8
+ tags:
9
+ - tiny
10
+ - slm
11
+ - small-language-model
12
+ - sub-1m
13
+ - from-scratch
14
+ - gqa
15
+ - llama
16
+ metrics:
17
+ - perplexity
18
+ ---
19
+
20
+ # TinyStories-40M
21
+
22
+ A **39.6M-parameter** LLaMA-style transformer trained **from scratch** on the
23
+ full TinyStories corpus. Published as a from-scratch training demonstration at
24
+ the ~40M scale — it is *not* a coherent story generator (see Quality below).
25
+
26
+ ## Architecture
27
+
28
+ | Parameter | Value |
29
+ |---|---|
30
+ | Layers | 12 |
31
+ | Hidden size | 512 |
32
+ | Attention heads (Q) | 8 |
33
+ | Attention heads (KV) | 4 (GQA 2:1) |
34
+ | Head dim | 64 |
35
+ | FFN (SwiGLU) | 1408 |
36
+ | Vocab size | 8192 (BPE) |
37
+ | Context length | 512 |
38
+ | RoPE θ | 10000 |
39
+ | Norm | RMSNorm (pre-norm) |
40
+ | Tied embeddings | Yes |
41
+ | Precision | FP32 |
42
+
43
+ **Total parameters: 39,596,544** (86 tensors, head weight tied to the token
44
+ embedding). Verified against the published `model.safetensors`.
45
+
46
+ ## Training
47
+
48
+ - **Data:** roneneldan/TinyStories (~1.9B chars, ~490M tokens at seq 512)
49
+ - **Steps:** 15,000
50
+ - **Batch size:** 64 sequences × 512 tokens (32,704 tok/step)
51
+ - **Optimizer:** AdamW, lr 3e-4, cosine decay (min-lr-frac 0.1), 300-step warmup
52
+ - **Hardware:** RTX 5090 (32 GB), ~4 hours
53
+ - **Final train loss:** 3.23 (step 15000)
54
+
55
+ ## Quality
56
+
57
+ Honest picture, from running the published weights (temp 0.7, top-k 50):
58
+
59
+ - On the canonical prompt **"Once upon a time"** the model produces a
60
+ story-like first sentence, then degrades:
61
+
62
+ > "Once upon a time, his mom and his clapped and cheered for him. They all
63
+ > enjoyed spending the rest of their special day in the park, his mom's light
64
+ > and a Stop being himself."
65
+
66
+ - On other prompts it is weaker still — often a single short phrase or an
67
+ immediate stop:
68
+
69
+ > "In the forest" → "In the forest things."
70
+ > "She opened the door" → "She opened the door."
71
+
72
+ - Longer generations drift into incoherent, non-grammatical text.
73
+
74
+ This is what a 40M from-scratch model on a single 490M-token corpus
75
+ demonstrably does: it learns the surface distribution of story text (word
76
+ order, names, punctuation, the "Once upon a time" register) but does not
77
+ sustain coherent narrative. It is published as a training demonstration, not
78
+ as a usable story generator.
79
+
80
+ ## Usage
81
+
82
+ This is a custom `transformers` model — load with `trust_remote_code=True`:
83
+
84
+ ```python
85
+ from transformers import AutoModelForCausalLM, AutoTokenizer
86
+ import torch
87
+
88
+ model = AutoModelForCausalLM.from_pretrained(
89
+ "Compactbot/tinystories-40m", trust_remote_code=True, torch_dtype=torch.float32
90
+ )
91
+ tok = AutoTokenizer.from_pretrained("Compactbot/tinystories-40m")
92
+
93
+ prompt = "Once upon a time"
94
+ ids = tok(prompt, return_tensors="pt").input_ids
95
+ out = model.generate(ids, max_new_tokens=100, temperature=0.7, top_k=50)
96
+ print(tok.decode(out[0], skip_special_tokens=True))
97
+ ```
98
+
99
+ Note: the bundled `generate()` always samples (no `do_sample` flag); pass
100
+ `temperature` and `top_k` to control it.
101
+
102
+ ## Notes
103
+
104
+ - From-scratch training: initialized randomly, trained end-to-end on
105
+ TinyStories. No pre-training from any other model.
106
+ - GQA (Grouped Query Attention) 2:1, SwiGLU, RoPE — LLaMA-2 recipe at 40M scale.
107
+ - ~151 MB in FP32 — under 200 MB.
108
+
109
+ ## Files
110
+
111
+ | file | what |
112
+ |------|------|
113
+ | `model.safetensors` | 39,596,544 params, 86 tensors, FP32 |
114
+ | `config.json` | architecture + training metadata |
115
+ | `modeling_tinystories.py` | the model class (loaded via `trust_remote_code`) |
116
+ | `tokenizer.json` | BPE-8k tokenizer (HF `tokenizers` format) |