testr-100k / README.md
Compactbot's picture
Add model card
775cee0 verified
|
Raw History Blame
3.29 kB
metadata
license: apache-2.0
pipeline_tag: text-generation
language: en
tags:
  - tiny
  - tiny-lm
  - tiny-model
  - slm
  - small-language-model
  - sub-1m
  - from-scratch
  - character-level
  - char-gpt
metrics:
  - perplexity

Testr-100K

A 106,568-parameter character-level GPT trained from scratch. Fulfils model request #17 from @GGUFGuy.

What it is

A minimal nanoGPT-style causal transformer operating at the character level (128-char vocabulary). This is a demonstration of training a working language model from absolute scratch with a very small parameter budget — not a tool for generating coherent text.

Architecture

Parameter Value
Layers 4
Embedding dim 44
Attention heads 2
Context length 512 chars
FFN GELU, mult 2.667
Positional encoding RoPE (θ=100000)
Normalization LayerNorm (pre-norm)
Vocab 128 chars + 1 UNK
Tied embeddings Yes (head = embedding)
Total params 106,568

Training

  • Data: 191 MB web text (character-level, uint8)
  • Steps: 8,000
  • Batch size: 32 × seq 512
  • Optimizer: Muon (0.02) + AdamW (1e-4) for biases/norms
  • Schedule: Cosine decay with warmup
  • Hardware: RTX 5090 (32 GB)
  • dtype: float32

Quality (honest)

This is a character-level model at 106K params. It learns English character statistics and produces text that is grammatical in shape but degenerates into repetition loops within 2–3 sentences. It is not a coherent text generator.

Val perplexity: 39.77 (1M held-out chars from the same corpus)

Sample outputs (greedy, temp=0)

The → mean to the box and said, "I don't know what the boy was so happy to the box and said, "I want to t

Once upon a time → , there was a little girl named Lily. He was so happy to the box and said, "I want to the box and sa

I think that → the boy was so happy to the box and said, "I want to the box and said. "I want to the

Sample outputs (temp=0.7)

The → little girl named Lily was hands, "It was not want to be and field. It could go him to play. She pu

Once upon a time → there was a mommy was playing and the grandma happy.\n\nLily was very happy thast and pretty was so d

The first sentence or two is often grammatical; after that the model locks into a phrase and repeats it. This is expected at this scale and vocabulary size.

What it is NOT

  • Not a subword/token-level model (it cannot "read" word-level benchmarks like BLiMP or ARC)
  • Not a coherent paragraph generator
  • Not comparable to sub-1M subword models on any word-level metric

Reproducing

The model is a standard nanoGPT-style architecture. The training script is a standard causal transformer with Muon optimizer. The checkpoint is saved as safetensors with tied embeddings.

To generate: load model.safetensors into a GPT class matching the config, map characters to token IDs via vocab.json, and sample.

Files

File Description
model.safetensors Model weights (408 KB)
config.json Architecture config
vocab.json Character → token ID mapping (128 chars + UNK)