File size: 2,541 Bytes
d5b9248
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
---
library_name: pytorch
pipeline_tag: text-generation
language:
- en
datasets:
- roneneldan/TinyStories
tags:
- safetensors
- custom-code
---

# Tiny CED

Tiny CED is a small custom PyTorch encoder-decoder language model trained from
scratch on the English [TinyStories](https://huggingface.co/datasets/roneneldan/TinyStories)
dataset. This package contains inference-only weights in Safetensors format.

This is not a Transformers `PreTrainedModel`. Load it with the included
`model.py` and `safetensors.torch.load_model` as shown below.

## Model details

| Item | Value |
| --- | ---: |
| Independent parameters | 19,667,712 |
| Vocabulary size | 8,192 |
| Hidden size | 384 |
| Encoder layers | 4 |
| Decoder layers | 4 |
| Attention heads | 6 |
| Feed-forward size | 1,024 |
| Local attention window | 64 |
| Trained context length | 512 |
| Weight dtype | FP32 |

The token embedding and output head are tied. `model.safetensors` was exported
with `safetensors.torch.save_model` so the shared tensor is stored only once.

## Usage

Install the runtime dependencies:

```bash
pip install -r requirements.txt
```

Run generation on CPU:

```bash
python generate.py \
  --device cpu \
  --text "Once upon a time, a little rabbit" \
  --max-new-tokens 100
```

Use `--device cuda` when CUDA is available. Set `--temperature 0` for greedy
decoding.

## Training and evaluation

The exported weights come from the existing `best.pt` checkpoint; no retraining
was performed during export.

- Optimizer: AdamW
- Peak learning rate: 3e-4
- Best-checkpoint training tokens: 40,004,782
- Best-checkpoint step: 5,856
- Validation loss: 1.896527
- Validation perplexity: 6.662715

Validation used the TinyStories validation split and the same 8,192-token BPE
tokenizer included in this repository.

## Files

- `model.safetensors`: inference weights
- `model.py`: exact custom PyTorch architecture
- `config.json`: architecture and token IDs
- `tokenizer.json`: BPE tokenizer
- `generate.py`: minimal generation example

## Limitations

This is a small research model trained only on synthetic English children's
stories. It may generate incorrect, repetitive, biased, unsafe, or otherwise
inappropriate text. It is not suitable for factual, safety-critical, or
production use without further evaluation.

The source project did not specify a license for its code or weights. Choose an
appropriate license and confirm the training-data terms before publishing this
package publicly. TinyStories is listed as CDLA-Sharing-1.0 on its dataset page.