Instructions to use tampajohn/meow-lite-v5 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use tampajohn/meow-lite-v5 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="tampajohn/meow-lite-v5")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("tampajohn/meow-lite-v5") model = AutoModelForCausalLM.from_pretrained("tampajohn/meow-lite-v5", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use tampajohn/meow-lite-v5 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "tampajohn/meow-lite-v5" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "tampajohn/meow-lite-v5", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/tampajohn/meow-lite-v5
- SGLang
How to use tampajohn/meow-lite-v5 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "tampajohn/meow-lite-v5" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "tampajohn/meow-lite-v5", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "tampajohn/meow-lite-v5" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "tampajohn/meow-lite-v5", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use tampajohn/meow-lite-v5 with Docker Model Runner:
docker model run hf.co/tampajohn/meow-lite-v5
meow-lite-v5
The meow-lite family
- meow-lite โ v4 classic: the 104K meow toy behind the OpenAI/Anthropic shim (497 organic downloads and counting)
- meow-lite-v5 โ the chaotic cat: 6.8M from-scratch BPE GPT, 67.7% held-out comprehension, 3% leak
- meow-lite-v6 โ the calmer cat: same architecture, 70.4% comprehension, 3% leak
- meow-lite-v6-dataset โ the 16,633-pair teacher web + full methodology
- github.com/tampajohn/meow-lite โ server, specs, eval batteries (MIT) The chaotic cat. A 6,836,224-parameter from-scratch BPE GPT that reads English and can only speak cat. Maximum comprehension, minimal manners. It responds to everything a cat would: correctly, and then some.
Two temperaments
| v5 (this) | v6 | |
|---|---|---|
| Temperament | chaotic | calmer |
| Held-out synonym accuracy | 67.7% | 70.4% |
| Fresh-neutral leak | 3% | 3% |
| Vibe | comprehends more, interrupts more | comprehends slightly less, minds its manners |
Served with action damping d=2 (inference-time prior suppression): the numbers above are the served reality, measured on the full held-out battery. Both were trained on the identical corpus with the identical recipe; the difference is which seed the dice favored. We ship both, because cats are inconsistent.
Architecture
- GPT-2 (from-scratch, random init): n_layer=6, n_head=8, n_embd=256, n_positions=192
- Tokenizer: byte-level BPE (~8k) + 10 cat action tokens as specials
- Trained in minutes on a laptop (the 16,633-pair teacher web took ~13h on 2x GB10)
- Output hard-masked to cat tokens (CatMask LogitsProcessor): non-cat output is impossible by construction
- ActionOnce: each action token fires once per response (repeat biting 100% -> 0%)
- ActionDamping d=2: prior suppression of action logits (the served numbers above)
- Deterministic: sha256(prompt) seeds sampling; same prompt, same cat
Samples
- "can I rub her tummy?" -> " Purrr! "
- "did you see that dog" -> " Mraow Mew!"
- "How was your day" -> "Mewmew. Mrow!"
- "it is 3am" -> " Meow ..."
Usage
Server, tokenizer, eval batteries, and the companion calmer model:
https://github.com/tampajohn/meow-lite (MEOW_LITE_ENGINE=v5)
Dataset + full methodology: https://huggingface.co/datasets/tampajohn/meow-lite-v6-dataset
- Downloads last month
- 238
