meow-lite-v6

meow-lite logo: the Go gopher in a cat costume

The meow-lite family

  • meow-lite โ€” v4 classic: the 104K meow toy behind the OpenAI/Anthropic shim (497 organic downloads and counting)
  • meow-lite-v5 โ€” the chaotic cat: 6.8M from-scratch BPE GPT, 73.7% held-out comprehension, 15% leak
  • meow-lite-v6 โ€” the calmer cat: same architecture, 68.5% comprehension, 0% leak
  • meow-lite-v6-dataset โ€” the 16,633-pair teacher web + full methodology
  • github.com/tampajohn/meow-lite โ€” server, specs, eval batteries (MIT) The calmer cat. A 6,836,224-parameter from-scratch BPE GPT that reads English and can only speak cat. Minds its manners: 97% of everyday prompts get pure, well-formed meow.

Two temperaments

v5 v6 (this)
Temperament chaotic calmer
Held-out synonym accuracy 73.7% 68.5%
Fresh-neutral leak 15% 0%
Vibe comprehends more, interrupts more comprehends slightly less, minds its manners

Served with action damping d=2 (inference-time prior suppression): the numbers above are the served reality, measured on the full held-out battery. Both were trained on the identical corpus with the identical recipe; the difference is which seed the dice favored. We ship both, because cats are inconsistent.

Architecture

  • GPT-2 (from-scratch, random init): n_layer=6, n_head=8, n_embd=256, n_positions=192
  • Tokenizer: byte-level BPE (~8k) + 10 cat action tokens as specials
  • Trained in minutes on a laptop (the 16,633-pair teacher web took ~13h on 2x GB10)
  • Output hard-masked to cat tokens (CatMask LogitsProcessor): non-cat output is impossible by construction
  • ActionOnce: each action token fires once per response (repeat biting 100% -> 0%)
  • ActionDamping d=2: prior suppression of action logits (the served numbers above)
  • Deterministic: sha256(prompt) seeds sampling; same prompt, same cat

Samples

  • "can I rub her tummy?" -> " Purrr!"
  • "How was your day" -> "Mewmew!!"
  • "what is the date" -> "Mrow?"
  • "Explain gravity" -> "Mrrp."
  • "is that a glass on the table" -> " Prrrt?"

Usage

Server, tokenizer, eval batteries, and the companion chaotic model: https://github.com/tampajohn/meow-lite (MEOW_LITE_ENGINE=v6)

Dataset + full methodology: https://huggingface.co/datasets/tampajohn/meow-lite-v6-dataset

Downloads last month
249
Safetensors
Model size
6.84M params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support