TinyChat-5M
A 4.3M parameter transformer trained from scratch on the TinyChat dataset.
Architecture
| Parameter | Value |
|---|---|
| Parameters | 4,327,872 |
| Layers | 6 |
| Hidden dim | 192 |
| Attention heads | 3 |
| Head dim | 64 |
| FFN mult | 4x (SwiGLU) |
| Context | 512 |
| Vocab | 4,096 (BPE) |
| Norm | RMSNorm |
| Positional | RoPE |
| Embeddings | Tied (input = output) |
Training
- Dataset: starhopp3r/TinyChat (1M rows, ~740M tokens)
- Tokens seen: 1.64B (90% train / 10% val split, multi-epoch)
- Optimizer: AdamW, LR 1e-3, cosine decay + warmup
- Batch: 32 × 512 tokens
- Steps: 100,000 (completed cleanly)
- Hardware: RTX 5090 (32 GB)
- Time: ~46 minutes
Results
| Metric | Value |
|---|---|
| Final val loss (step 100k) | 1.7603 |
| Final val perplexity | 5.81 |
| Best val loss (step 96k) | 1.7531 |
| Best val perplexity | 5.77 |
Generation Samples
Prompt: Hello, how are you today? Output: I feel quite disgusted by the mess in our office today. Yes, it is necessary to have a better work space for everyone. The clutter makes everyone a bit unwell and can affect their mood. It really does, especially when we just want things to be clean again. A tidy environment helps us enjoy our work more clearly each day.
Prompt: Once upon a time Output: when the thread interrupts our plans and makes everything feel more tense. It is so annoying how simple things can turn into chaos, isn't it? Yes, it really is, like trying to find noise that won't stop bothering us. I wish we could just enjoy the sewing process without all these problems.
Prompt: The cat sat on the Output: bench, waiting for lunch to arrive. That is surprising and cute! I hope it unfolds before they try to play outside. Yes, it will be a delightful surprise, but we must stay calm as her. Do you remember that day when the weather was so nice?
Prompt: I think that Output: is a good idea because sometimes sadness can feel so heavy and hard to handle. Yes, it feels like emotions are holding us down like a thin needle poking at our hearts. Do you believe sharing feelings helps lighten the load we carry in some way?
The model produces coherent conversational English appropriate for its size. It does not answer factual questions correctly (expected at 4.3M params) but maintains consistent tone and topic.
Notes
- v4 fixes v3's NaN divergence (LR reduced from 2e-3 to 1e-3, NaN detection added).
- Training completed all 100k steps without divergence.
- Model is for research/educational purposes. Not suitable for production use.