Compactbot commited on
Commit
f2dfad7
·
verified ·
1 Parent(s): 885263a

Update card for v4: 100k steps, 1.64B tokens, val_loss 1.7603

Browse files
Files changed (1) hide show
  1. README.md +23 -19
README.md CHANGED
@@ -12,13 +12,13 @@ tags:
12
 
13
  # TinyChat-5M
14
 
15
- A 5.1M parameter transformer trained from scratch on the [TinyChat](https://huggingface.co/datasets/starhopp3r/TinyChat) dataset.
16
 
17
  ## Architecture
18
 
19
  | Parameter | Value |
20
  |-----------|-------|
21
- | Parameters | 5,114,304 |
22
  | Layers | 6 |
23
  | Hidden dim | 192 |
24
  | Attention heads | 3 |
@@ -32,37 +32,41 @@ A 5.1M parameter transformer trained from scratch on the [TinyChat](https://hugg
32
 
33
  ## Training
34
 
35
- - **Dataset**: starhopp3r/TinyChat (1M rows, ~170M tokens)
36
- - **Tokens seen**: 143.9M (90% train / 10% val split)
37
- - **Optimizer**: AdamW, LR 2e-3, cosine decay + warmup
38
  - **Batch**: 32 × 512 tokens
39
- - **Steps**: 45,000 (diverged to NaN at step 47,490; last stable checkpoint used)
40
  - **Hardware**: RTX 5090 (32 GB)
41
- - **Time**: ~25 minutes
42
 
43
  ## Results
44
 
45
  | Metric | Value |
46
  |--------|-------|
47
- | Val loss (step 45k) | 1.9027 |
48
- | Val perplexity | 6.70 |
49
- | Best val loss (step 47k) | 1.8883 |
 
50
 
51
  ## Generation Samples
52
 
53
- > **Prompt**: [INST] What is the capital of France? [/INST]
54
- > **Output**: t time in a place me ant for ner ve and I feel quite nervous. But what if something goes wrong during those moments at night...
55
 
56
- > **Prompt**: [INST] Tell me a joke. [/INST]
57
- > **Output**: very frustrating, but you are doing your best in the end. Thank you, it is nice to relax and share ideas with someone today...
58
 
59
- > **Prompt**: [INST] Explain gravity in simple terms. [/INST]
60
- > **Output**: remind us of how we all have these smaller and more expressive feelings. I miss the days when everything felt bright and happy...
61
 
62
- The model produces coherent conversational English appropriate for its size. It does not answer factual questions correctly (expected at 5M params) but maintains consistent tone and topic.
 
 
 
63
 
64
  ## Notes
65
 
66
- - Training diverged to NaN at step 47,490 (likely LR too high for late training).
67
- - Checkpoint at step 45,000 is the last stable one before divergence.
68
  - Model is for research/educational purposes. Not suitable for production use.
 
12
 
13
  # TinyChat-5M
14
 
15
+ A 4.3M parameter transformer trained from scratch on the [TinyChat](https://huggingface.co/datasets/starhopp3r/TinyChat) dataset.
16
 
17
  ## Architecture
18
 
19
  | Parameter | Value |
20
  |-----------|-------|
21
+ | Parameters | 4,327,872 |
22
  | Layers | 6 |
23
  | Hidden dim | 192 |
24
  | Attention heads | 3 |
 
32
 
33
  ## Training
34
 
35
+ - **Dataset**: starhopp3r/TinyChat (1M rows, ~740M tokens)
36
+ - **Tokens seen**: 1.64B (90% train / 10% val split, multi-epoch)
37
+ - **Optimizer**: AdamW, LR 1e-3, cosine decay + warmup
38
  - **Batch**: 32 × 512 tokens
39
+ - **Steps**: 100,000 (completed cleanly)
40
  - **Hardware**: RTX 5090 (32 GB)
41
+ - **Time**: ~46 minutes
42
 
43
  ## Results
44
 
45
  | Metric | Value |
46
  |--------|-------|
47
+ | Final val loss (step 100k) | 1.7603 |
48
+ | Final val perplexity | 5.81 |
49
+ | Best val loss (step 96k) | 1.7531 |
50
+ | Best val perplexity | 5.77 |
51
 
52
  ## Generation Samples
53
 
54
+ > **Prompt**: Hello, how are you today?
55
+ > **Output**: I feel quite disgusted by the mess in our office today. Yes, it is necessary to have a better work space for everyone. The clutter makes everyone a bit unwell and can affect their mood. It really does, especially when we just want things to be clean again. A tidy environment helps us enjoy our work more clearly each day.
56
 
57
+ > **Prompt**: Once upon a time
58
+ > **Output**: when the thread interrupts our plans and makes everything feel more tense. It is so annoying how simple things can turn into chaos, isn't it? Yes, it really is, like trying to find noise that won't stop bothering us. I wish we could just enjoy the sewing process without all these problems.
59
 
60
+ > **Prompt**: The cat sat on the
61
+ > **Output**: bench, waiting for lunch to arrive. That is surprising and cute! I hope it unfolds before they try to play outside. Yes, it will be a delightful surprise, but we must stay calm as her. Do you remember that day when the weather was so nice?
62
 
63
+ > **Prompt**: I think that
64
+ > **Output**: is a good idea because sometimes sadness can feel so heavy and hard to handle. Yes, it feels like emotions are holding us down like a thin needle poking at our hearts. Do you believe sharing feelings helps lighten the load we carry in some way?
65
+
66
+ The model produces coherent conversational English appropriate for its size. It does not answer factual questions correctly (expected at 4.3M params) but maintains consistent tone and topic.
67
 
68
  ## Notes
69
 
70
+ - v4 fixes v3's NaN divergence (LR reduced from 2e-3 to 1e-3, NaN detection added).
71
+ - Training completed all 100k steps without divergence.
72
  - Model is for research/educational purposes. Not suitable for production use.