| --- |
| library_name: pytorch |
| pipeline_tag: text-generation |
| language: |
| - en |
| datasets: |
| - roneneldan/TinyStories |
| tags: |
| - safetensors |
| - custom-code |
| --- |
| |
| # Tiny CED |
|
|
| Tiny CED is a small custom PyTorch encoder-decoder language model trained from |
| scratch on the English [TinyStories](https://huggingface.co/datasets/roneneldan/TinyStories) |
| dataset. This package contains inference-only weights in Safetensors format. |
|
|
| This is not a Transformers `PreTrainedModel`. Load it with the included |
| `model.py` and `safetensors.torch.load_model` as shown below. |
|
|
| ## Model details |
|
|
| | Item | Value | |
| | --- | ---: | |
| | Independent parameters | 19,667,712 | |
| | Vocabulary size | 8,192 | |
| | Hidden size | 384 | |
| | Encoder layers | 4 | |
| | Decoder layers | 4 | |
| | Attention heads | 6 | |
| | Feed-forward size | 1,024 | |
| | Local attention window | 64 | |
| | Trained context length | 512 | |
| | Weight dtype | FP32 | |
|
|
| The token embedding and output head are tied. `model.safetensors` was exported |
| with `safetensors.torch.save_model` so the shared tensor is stored only once. |
|
|
| ## Usage |
|
|
| Install the runtime dependencies: |
|
|
| ```bash |
| pip install -r requirements.txt |
| ``` |
|
|
| Run generation on CPU: |
|
|
| ```bash |
| python generate.py \ |
| --device cpu \ |
| --text "Once upon a time, a little rabbit" \ |
| --max-new-tokens 100 |
| ``` |
|
|
| Use `--device cuda` when CUDA is available. Set `--temperature 0` for greedy |
| decoding. |
|
|
| ## Training and evaluation |
|
|
| The exported weights come from the existing `best.pt` checkpoint; no retraining |
| was performed during export. |
|
|
| - Optimizer: AdamW |
| - Peak learning rate: 3e-4 |
| - Best-checkpoint training tokens: 40,004,782 |
| - Best-checkpoint step: 5,856 |
| - Validation loss: 1.896527 |
| - Validation perplexity: 6.662715 |
|
|
| Validation used the TinyStories validation split and the same 8,192-token BPE |
| tokenizer included in this repository. |
|
|
| ## Files |
|
|
| - `model.safetensors`: inference weights |
| - `model.py`: exact custom PyTorch architecture |
| - `config.json`: architecture and token IDs |
| - `tokenizer.json`: BPE tokenizer |
| - `generate.py`: minimal generation example |
|
|
| ## Limitations |
|
|
| This is a small research model trained only on synthetic English children's |
| stories. It may generate incorrect, repetitive, biased, unsafe, or otherwise |
| inappropriate text. It is not suitable for factual, safety-critical, or |
| production use without further evaluation. |
|
|
| The source project did not specify a license for its code or weights. Choose an |
| appropriate license and confirm the training-data terms before publishing this |
| package publicly. TinyStories is listed as CDLA-Sharing-1.0 on its dataset page. |
|
|