File size: 2,141 Bytes
f725858 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 | ---
language: en
library_name: pytorch
license: mit
tags:
- tiny
- slm
- from-scratch
- text-generation
- tinychat
---
# TinyChat-5M
A 5.1M parameter transformer trained from scratch on the [TinyChat](https://huggingface.co/datasets/starhopp3r/TinyChat) dataset.
## Architecture
| Parameter | Value |
|-----------|-------|
| Parameters | 5,114,304 |
| Layers | 6 |
| Hidden dim | 192 |
| Attention heads | 3 |
| Head dim | 64 |
| FFN mult | 4x (SwiGLU) |
| Context | 512 |
| Vocab | 4,096 (BPE) |
| Norm | RMSNorm |
| Positional | RoPE |
| Embeddings | Tied (input = output) |
## Training
- **Dataset**: starhopp3r/TinyChat (1M rows, ~170M tokens)
- **Tokens seen**: 143.9M (90% train / 10% val split)
- **Optimizer**: AdamW, LR 2e-3, cosine decay + warmup
- **Batch**: 32 × 512 tokens
- **Steps**: 45,000 (diverged to NaN at step 47,490; last stable checkpoint used)
- **Hardware**: RTX 5090 (32 GB)
- **Time**: ~25 minutes
## Results
| Metric | Value |
|--------|-------|
| Val loss (step 45k) | 1.9027 |
| Val perplexity | 6.70 |
| Best val loss (step 47k) | 1.8883 |
## Generation Samples
> **Prompt**: [INST] What is the capital of France? [/INST]
> **Output**: t time in a place me ant for ner ve and I feel quite nervous. But what if something goes wrong during those moments at night...
> **Prompt**: [INST] Tell me a joke. [/INST]
> **Output**: very frustrating, but you are doing your best in the end. Thank you, it is nice to relax and share ideas with someone today...
> **Prompt**: [INST] Explain gravity in simple terms. [/INST]
> **Output**: remind us of how we all have these smaller and more expressive feelings. I miss the days when everything felt bright and happy...
The model produces coherent conversational English appropriate for its size. It does not answer factual questions correctly (expected at 5M params) but maintains consistent tone and topic.
## Notes
- Training diverged to NaN at step 47,490 (likely LR too high for late training).
- Checkpoint at step 45,000 is the last stable one before divergence.
- Model is for research/educational purposes. Not suitable for production use. |