testr-100k / README.md
Compactbot's picture
Add model card
775cee0 verified
|
Raw History Blame
3.29 kB
---
license: apache-2.0
pipeline_tag: text-generation
language: en
tags:
- tiny
- tiny-lm
- tiny-model
- slm
- small-language-model
- sub-1m
- from-scratch
- character-level
- char-gpt
metrics:
- perplexity
---
# Testr-100K
A **106,568-parameter** character-level GPT trained from scratch. Fulfils model request [#17](https://huggingface.co/spaces/Compactbot/model-requests/discussions/17) from @GGUFGuy.
## What it is
A minimal nanoGPT-style causal transformer operating at the **character level** (128-char vocabulary). This is a demonstration of training a working language model from absolute scratch with a very small parameter budget β€” not a tool for generating coherent text.
## Architecture
| Parameter | Value |
|-----------|-------|
| Layers | 4 |
| Embedding dim | 44 |
| Attention heads | 2 |
| Context length | 512 chars |
| FFN | GELU, mult 2.667 |
| Positional encoding | RoPE (ΞΈ=100000) |
| Normalization | LayerNorm (pre-norm) |
| Vocab | 128 chars + 1 UNK |
| Tied embeddings | Yes (head = embedding) |
| **Total params** | **106,568** |
## Training
- **Data:** 191 MB web text (character-level, uint8)
- **Steps:** 8,000
- **Batch size:** 32 Γ— seq 512
- **Optimizer:** Muon (0.02) + AdamW (1e-4) for biases/norms
- **Schedule:** Cosine decay with warmup
- **Hardware:** RTX 5090 (32 GB)
- **dtype:** float32
## Quality (honest)
This is a **character-level** model at 106K params. It learns English character statistics and produces text that is *grammatical in shape* but **degenerates into repetition loops** within 2–3 sentences. It is not a coherent text generator.
**Val perplexity:** 39.77 (1M held-out chars from the same corpus)
### Sample outputs (greedy, temp=0)
> `The` β†’ ` mean to the box and said, "I don't know what the boy was so happy to the box and said, "I want to t`
> `Once upon a time` β†’ `, there was a little girl named Lily. He was so happy to the box and said, "I want to the box and sa`
> `I think that` β†’ ` the boy was so happy to the box and said, "I want to the box and said. "I want to the`
### Sample outputs (temp=0.7)
> `The` β†’ ` little girl named Lily was hands, "It was not want to be and field. It could go him to play. She pu`
> `Once upon a time` β†’ ` there was a mommy was playing and the grandma happy.\n\nLily was very happy thast and pretty was so d`
The first sentence or two is often grammatical; after that the model locks into a phrase and repeats it. This is expected at this scale and vocabulary size.
## What it is NOT
- Not a subword/token-level model (it cannot "read" word-level benchmarks like BLiMP or ARC)
- Not a coherent paragraph generator
- Not comparable to sub-1M subword models on any word-level metric
## Reproducing
The model is a standard nanoGPT-style architecture. The training script is a standard causal transformer with Muon optimizer. The checkpoint is saved as safetensors with tied embeddings.
To generate: load `model.safetensors` into a GPT class matching the config, map characters to token IDs via `vocab.json`, and sample.
## Files
| File | Description |
|------|-------------|
| `model.safetensors` | Model weights (408 KB) |
| `config.json` | Architecture config |
| `vocab.json` | Character β†’ token ID mapping (128 chars + UNK) |