meow-lite-v5

meow-lite logo: the Go gopher in a cat costume

The meow-lite family

  • meow-lite โ€” v4 classic: the 104K meow toy behind the OpenAI/Anthropic shim (497 organic downloads and counting)
  • meow-lite-v5 โ€” the chaotic cat: 6.8M from-scratch BPE GPT, 67.7% held-out comprehension, 3% leak
  • meow-lite-v6 โ€” the calmer cat: same architecture, 70.4% comprehension, 3% leak
  • meow-lite-v6-dataset โ€” the 16,633-pair teacher web + full methodology
  • github.com/tampajohn/meow-lite โ€” server, specs, eval batteries (MIT) The chaotic cat. A 6,836,224-parameter from-scratch BPE GPT that reads English and can only speak cat. Maximum comprehension, minimal manners. It responds to everything a cat would: correctly, and then some.

Two temperaments

v5 (this) v6
Temperament chaotic calmer
Held-out synonym accuracy 67.7% 70.4%
Fresh-neutral leak 3% 3%
Vibe comprehends more, interrupts more comprehends slightly less, minds its manners

Served with action damping d=2 (inference-time prior suppression): the numbers above are the served reality, measured on the full held-out battery. Both were trained on the identical corpus with the identical recipe; the difference is which seed the dice favored. We ship both, because cats are inconsistent.

Architecture

  • GPT-2 (from-scratch, random init): n_layer=6, n_head=8, n_embd=256, n_positions=192
  • Tokenizer: byte-level BPE (~8k) + 10 cat action tokens as specials
  • Trained in minutes on a laptop (the 16,633-pair teacher web took ~13h on 2x GB10)
  • Output hard-masked to cat tokens (CatMask LogitsProcessor): non-cat output is impossible by construction
  • ActionOnce: each action token fires once per response (repeat biting 100% -> 0%)
  • ActionDamping d=2: prior suppression of action logits (the served numbers above)
  • Deterministic: sha256(prompt) seeds sampling; same prompt, same cat

Samples

  • "can I rub her tummy?" -> " Purrr! "
  • "did you see that dog" -> " Mraow Mew!"
  • "How was your day" -> "Mewmew. Mrow!"
  • "it is 3am" -> " Meow ..."

Usage

Server, tokenizer, eval batteries, and the companion calmer model: https://github.com/tampajohn/meow-lite (MEOW_LITE_ENGINE=v5)

Dataset + full methodology: https://huggingface.co/datasets/tampajohn/meow-lite-v6-dataset

Downloads last month
238
Safetensors
Model size
6.84M params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support