Add model card
Browse files
README.md
ADDED
|
@@ -0,0 +1,68 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
language: en
|
| 3 |
+
library_name: pytorch
|
| 4 |
+
license: mit
|
| 5 |
+
tags:
|
| 6 |
+
- tiny
|
| 7 |
+
- slm
|
| 8 |
+
- from-scratch
|
| 9 |
+
- text-generation
|
| 10 |
+
- tinychat
|
| 11 |
+
---
|
| 12 |
+
|
| 13 |
+
# TinyChat-5M
|
| 14 |
+
|
| 15 |
+
A 5.1M parameter transformer trained from scratch on the [TinyChat](https://huggingface.co/datasets/starhopp3r/TinyChat) dataset.
|
| 16 |
+
|
| 17 |
+
## Architecture
|
| 18 |
+
|
| 19 |
+
| Parameter | Value |
|
| 20 |
+
|-----------|-------|
|
| 21 |
+
| Parameters | 5,114,304 |
|
| 22 |
+
| Layers | 6 |
|
| 23 |
+
| Hidden dim | 192 |
|
| 24 |
+
| Attention heads | 3 |
|
| 25 |
+
| Head dim | 64 |
|
| 26 |
+
| FFN mult | 4x (SwiGLU) |
|
| 27 |
+
| Context | 512 |
|
| 28 |
+
| Vocab | 4,096 (BPE) |
|
| 29 |
+
| Norm | RMSNorm |
|
| 30 |
+
| Positional | RoPE |
|
| 31 |
+
| Embeddings | Tied (input = output) |
|
| 32 |
+
|
| 33 |
+
## Training
|
| 34 |
+
|
| 35 |
+
- **Dataset**: starhopp3r/TinyChat (1M rows, ~170M tokens)
|
| 36 |
+
- **Tokens seen**: 143.9M (90% train / 10% val split)
|
| 37 |
+
- **Optimizer**: AdamW, LR 2e-3, cosine decay + warmup
|
| 38 |
+
- **Batch**: 32 × 512 tokens
|
| 39 |
+
- **Steps**: 45,000 (diverged to NaN at step 47,490; last stable checkpoint used)
|
| 40 |
+
- **Hardware**: RTX 5090 (32 GB)
|
| 41 |
+
- **Time**: ~25 minutes
|
| 42 |
+
|
| 43 |
+
## Results
|
| 44 |
+
|
| 45 |
+
| Metric | Value |
|
| 46 |
+
|--------|-------|
|
| 47 |
+
| Val loss (step 45k) | 1.9027 |
|
| 48 |
+
| Val perplexity | 6.70 |
|
| 49 |
+
| Best val loss (step 47k) | 1.8883 |
|
| 50 |
+
|
| 51 |
+
## Generation Samples
|
| 52 |
+
|
| 53 |
+
> **Prompt**: [INST] What is the capital of France? [/INST]
|
| 54 |
+
> **Output**: t time in a place me ant for ner ve and I feel quite nervous. But what if something goes wrong during those moments at night...
|
| 55 |
+
|
| 56 |
+
> **Prompt**: [INST] Tell me a joke. [/INST]
|
| 57 |
+
> **Output**: very frustrating, but you are doing your best in the end. Thank you, it is nice to relax and share ideas with someone today...
|
| 58 |
+
|
| 59 |
+
> **Prompt**: [INST] Explain gravity in simple terms. [/INST]
|
| 60 |
+
> **Output**: remind us of how we all have these smaller and more expressive feelings. I miss the days when everything felt bright and happy...
|
| 61 |
+
|
| 62 |
+
The model produces coherent conversational English appropriate for its size. It does not answer factual questions correctly (expected at 5M params) but maintains consistent tone and topic.
|
| 63 |
+
|
| 64 |
+
## Notes
|
| 65 |
+
|
| 66 |
+
- Training diverged to NaN at step 47,490 (likely LR too high for late training).
|
| 67 |
+
- Checkpoint at step 45,000 is the last stable one before divergence.
|
| 68 |
+
- Model is for research/educational purposes. Not suitable for production use.
|