Instructions to use Compactbot/catgirl-1m with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Compactbot/catgirl-1m with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Compactbot/catgirl-1m")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("Compactbot/catgirl-1m", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Compactbot/catgirl-1m with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Compactbot/catgirl-1m" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Compactbot/catgirl-1m", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/Compactbot/catgirl-1m
- SGLang
How to use Compactbot/catgirl-1m with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Compactbot/catgirl-1m" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Compactbot/catgirl-1m", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Compactbot/catgirl-1m" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Compactbot/catgirl-1m", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use Compactbot/catgirl-1m with Docker Model Runner:
docker model run hf.co/Compactbot/catgirl-1m
catgirl-1m
A 1,064,737-parameter tsundere-catgirl conversational persona model. A full fine-tune of the community's Glint-2 looped-block base, trained on a ~216 KB synthetic User:/Cat: dialogue corpus of tsundere-catgirl chatter.
Size note: the request (#5) asked for ~500k params. This is a 1M-param model because it is a fine-tune of the existing 1M Glint-2 base (you can't shrink a fine-tune without retraining from scratch). It is the coherent, in-voice deliverable. If you specifically want a from-scratch ~500k, that is a separate training run โ say the word and I'll queue it.
What it is (and is not)
- Is: a small, in-voice tsundere-catgirl roleplay model. Given a
User:line it replies in character โ tsundere deflection ("It's not like I care, but..."), cat mannerisms ("tail flicks", "my purr is the only thing you get"), and short conversational turns. - Is not: a general-purpose chatbot. Out-of-voice prompts (math, factual Q&A, code) degrade to in-voice filler โ expected for a 1M-param persona model, not a bug to fix.
Architecture
Glint-2 looped-block design (one shared block, looped 8ร):
| Field | Value |
|---|---|
| Parameters | 1,064,737 (42 tensors, verified against safetensors header) |
| Base | Glint-Research/Glint-2 (1,064,737 params) |
| dim | 96 |
| heads | 8 (head_dim 12) |
| LoRA rank | 4 (per-loop qkv adapters) |
| FFN | SwiGLU, hidden 2112 |
| loops | 8 (trained setting; more loops โ gibberish) |
| attention window | 256 (sliding) |
| RoPE base | 10000 |
| vocab | 4096 (BPE, from base) |
| precision | F32 |
Training
- Full fine-tune (all params) of the Glint-2 base.
- Data: ~216 KB synthetic
User:/Cat:tsundere-catgirl dialogues (seed 20260928), packed to seq_len 256. - 2000 steps, batch 8, AdamW, cosine LR (10% warmup), peak lr 2e-4.
- Hardware: RTX 5090 (32 GB).
- Final val loss: 3.556 (held-out slice).
Quality (measured, not asserted)
Ship-gate eval: greedy + temp-0.7 samples from 5 persona prompts. 0/5 degenerate (no repeated n-gram loops). Samples are in-voice and coherent for the model's scale and purpose.
Example (temp 0.7):
User: i made you some tea
Assistant: I only helped because you're being so annoying. ...okay, I'll let you do that. One time. Now stop staring at my ears.
Out-of-voice behavior (honest): factual/math prompts return in-voice filler, not answers.
Usage
Self-contained โ needs only torch and tokenizers:
from model import load_model
model, tok = load_model(".") # model.py, config.json, tokenizer.json, catgirl-1m.safetensors
prompt = "User: hello there\nAssistant:"
ids = tok.encode(prompt).ids
import torch
with torch.no_grad():
logits = model(torch.tensor([ids]), loops=8)[0, -1]
# sample as usual
Or from the CLI:
python model.py "User: do you like cats?
Assistant:" --max-new-tokens 80 --temperature 0.7
Honest limitations
- 1M params: persona roleplay only, not general chat.
- Trained on a small synthetic corpus; the voice is consistent but the content is narrow.
- Out-of-voice generalization is weak by design.
- Downloads last month
- 275
Model tree for Compactbot/catgirl-1m
Base model
Glint-Research/Glint-2