Instructions to use ViorikaAI-org/CalmaCatCoder-Next-mini with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ViorikaAI-org/CalmaCatCoder-Next-mini with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="ViorikaAI-org/CalmaCatCoder-Next-mini") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("ViorikaAI-org/CalmaCatCoder-Next-mini") model = AutoModelForCausalLM.from_pretrained("ViorikaAI-org/CalmaCatCoder-Next-mini", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use ViorikaAI-org/CalmaCatCoder-Next-mini with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf ViorikaAI-org/CalmaCatCoder-Next-mini:F16 # Run inference directly in the terminal: llama cli -hf ViorikaAI-org/CalmaCatCoder-Next-mini:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf ViorikaAI-org/CalmaCatCoder-Next-mini:F16 # Run inference directly in the terminal: llama cli -hf ViorikaAI-org/CalmaCatCoder-Next-mini:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf ViorikaAI-org/CalmaCatCoder-Next-mini:F16 # Run inference directly in the terminal: ./llama-cli -hf ViorikaAI-org/CalmaCatCoder-Next-mini:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf ViorikaAI-org/CalmaCatCoder-Next-mini:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf ViorikaAI-org/CalmaCatCoder-Next-mini:F16
Use Docker
docker model run hf.co/ViorikaAI-org/CalmaCatCoder-Next-mini:F16
- LM Studio
- Jan
- vLLM
How to use ViorikaAI-org/CalmaCatCoder-Next-mini with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ViorikaAI-org/CalmaCatCoder-Next-mini" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ViorikaAI-org/CalmaCatCoder-Next-mini", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ViorikaAI-org/CalmaCatCoder-Next-mini:F16
- SGLang
How to use ViorikaAI-org/CalmaCatCoder-Next-mini with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "ViorikaAI-org/CalmaCatCoder-Next-mini" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ViorikaAI-org/CalmaCatCoder-Next-mini", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "ViorikaAI-org/CalmaCatCoder-Next-mini" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ViorikaAI-org/CalmaCatCoder-Next-mini", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use ViorikaAI-org/CalmaCatCoder-Next-mini with Ollama:
ollama run hf.co/ViorikaAI-org/CalmaCatCoder-Next-mini:F16
- Unsloth Desktop
- Docker Model Runner
How to use ViorikaAI-org/CalmaCatCoder-Next-mini with Docker Model Runner:
docker model run hf.co/ViorikaAI-org/CalmaCatCoder-Next-mini:F16
- Lemonade
How to use ViorikaAI-org/CalmaCatCoder-Next-mini with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull ViorikaAI-org/CalmaCatCoder-Next-mini:F16
Run and chat with the model
lemonade run user.CalmaCatCoder-Next-mini-F16
List all available models
lemonade list
- Atomic Chat
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf ViorikaAI-org/CalmaCatCoder-Next-mini:F16# Run inference directly in the terminal:
llama cli -hf ViorikaAI-org/CalmaCatCoder-Next-mini:F16Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf ViorikaAI-org/CalmaCatCoder-Next-mini:F16# Run inference directly in the terminal:
./llama-cli -hf ViorikaAI-org/CalmaCatCoder-Next-mini:F16Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf ViorikaAI-org/CalmaCatCoder-Next-mini:F16# Run inference directly in the terminal:
./build/bin/llama-cli -hf ViorikaAI-org/CalmaCatCoder-Next-mini:F16Use Docker
docker model run hf.co/ViorikaAI-org/CalmaCatCoder-Next-mini:F16CalmaCatCoder-Next-Mini
A tiny experimental chat model (96.6M parameters) trained completely from scratch on one consumer GPU (RTX 5060 Ti, 16 GB) by ViorikaAI. It is a hobby project: a small calm, cat-themed Python helper that knows who it is, speaks ChatML and writes very simple Python.
Please read the limitations below. This is a learning project, not a useful assistant. It memorised a handful of templates and will confidently produce wrong code and wrong facts.
Model details
| Architecture | Qwen3ForCausalLM (Qwen3 architecture, no Qwen weights, random init) |
| Parameters | 96.6M (tied input/output embeddings) |
| Layers / hidden / FFN | 14 / 768 / 2304 |
| Attention | 12 heads, 4 KV heads, head_dim 64 |
| Context length | 1024 bytes (not words) |
| Vocabulary | 256 tokens: one token = one UTF-8 byte (no learned tokenizer) |
| Precision | bf16 weights |
| Language | English only |
The tokenizer is byte-level
Token id equals the byte value, so "def " becomes [100, 101, 102, 32] and one word costs several tokens.
This is why the context is short. <|im_start|> and <|im_end|> are plain text made of bytes, not special
tokens, so the model ends its answer by writing the literal text <|im_end|>.
Stop generation on that string (see the example below).
Chat format (the same text the model was trained on, available as tokenizer.apply_chat_template):
<|im_start|>system
You are CalmaCatCoder-Next, a calm and friendly cat-themed coding assistant. You explain things step by step and say so when you are not sure.<|im_end|>
<|im_start|>user
Who are you?<|im_end|>
<|im_start|>assistant
If you pass no system message, the one above is inserted automatically. Other system prompts were never seen in training, so expect them to be mostly ignored.
Quick start
Requires transformers>=4.51 (Qwen3 support). No trust_remote_code needed.
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM, StoppingCriteria, StoppingCriteriaList
repo = "ViorikaAI-org/CalmaCatCoder-Next-mini" # or a local folder
device = "cuda" if torch.cuda.is_available() else "cpu"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo).to(device, dtype=torch.bfloat16).eval()
END = tok.encode("<|im_end|>") # the end marker is ordinary bytes
class StopOnEnd(StoppingCriteria):
def __call__(self, input_ids, scores, **kwargs):
return torch.tensor([input_ids[0, -len(END):].tolist() == END] * input_ids.shape[0],
device=input_ids.device)
messages = [{"role": "user", "content": "Who are you?"}]
prompt = tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tok(prompt, return_tensors="pt").to(device)
with torch.no_grad():
out = model.generate(**inputs, max_new_tokens=300,
stopping_criteria=StoppingCriteriaList([StopOnEnd()]))
print(tok.decode(out[0, inputs["input_ids"].shape[1]:]).replace("<|im_end|>", "").strip())
Sampling defaults live in generation_config.json (temperature 0.7, top-k 40). Keep prompt + answer under
1024 bytes; the model never saw longer sequences.
GGUF / llama.cpp
The tokenizer has no BPE merges and a pre-tokenizer unknown to the stock converter, so
convert_hf_to_gguf.py fails on it. The make_gguf.py wrapper from the code repository patches this
at conversion time without modifying llama.cpp (details in its header):
git clone https://github.com/ggml-org/llama.cpp
python make_gguf.py llama.cpp CalmaCatCoder-Next-96M calmacat-96m-f16.gguf f16 # or q8_0
Because <|im_end|> is plain text, llama.cpp does not stop on it by itself. Save the ChatML prompt to a file
and use a reverse prompt:
llama-completion -m calmacat-96m-f16.gguf -no-cnv --no-escape -f prompt.txt -r "<|im_end|>" -n 300 --temp 0.7 --top-k 40
A warning special_eos_id is not in special_eog_ids is expected: the model has no EOS token.
Prefer f16/bf16 or q8_0; this model is so small that aggressive quantization will likely hurt it.
How it was trained
All of it ran on a single RTX 5060 Ti (16 GB) under Windows with a PyTorch nightly build.
Stage 1: pre-training from scratch (1 epoch, ~7.6 hours).
29,908 steps × 32,768 byte-tokens ≈ 980M tokens, about 36k tokens/s.
AdamW (lr 6e-4, betas 0.9/0.95, weight decay 0.1), 300 warm-up steps, cosine decay to 10%,
gradient clipping 1.0, bf16 autocast, torch.compile.
Final validation loss 0.508 nats/byte on a held-out 2% of the mix. This number is not comparable with other
models: it is per byte, and the held-out part contains near-copies of the templated persona dialogues.
Data mix (~1 GB, target shares; the exact proportions depend on what each source could supply):
- Python code:
codeparrot/codeparrot-clean, the largest share (about 60%) - English text:
cosmopedia-v2fromHuggingFaceTB/smollm-corpus(15%) - Math:
meta-math/MetaMathQA(5%) - ChatML dialogues:
HuggingFaceTB/smol-smoltalk, only dialogues shorter than 1000 bytes (up to 20%) - About 5 MB of synthetic persona dialogues generated by a script (identity, creator, greetings, a few Python tasks)
Stage 2: supervised fine-tuning (1 epoch, ~13 minutes).
About 28k short ChatML dialogues (~17M tokens), loss computed on assistant tokens only, peak lr 1e-4.
Mixture of programmatically generated dialogues (identity and creator questions in many phrasings, greetings,
18 small Python tasks, follow-ups such as "add a docstring", remembering a name within a chat),
HuggingFaceTB/smoltalk (everyday-conversations) and a subset of smol-smoltalk.
SFT validation loss was 0.123, which mostly reflects how templated that data is.
Intended use
Learning and curiosity: look at how a tiny model trained from scratch behaves, reuse the training recipe, test tokenizers or inference tooling on a very small model.
Not intended for real coding help, factual questions, any decision that matters, or production use.
Limitations (observed, not hypothetical)
- English only, almost only Python. Other languages and programming languages are outside what it learned.
- It memorised templates. For a coding request it picks the closest of the ~18 tasks it knows and adapts it.
Real examples from testing:
- "give me a random python code" produced a function named
random_statethat simply reverses a string, followed by an explanation copied from a different task. - "Write a function that multiplies all numbers in a list" produced a list comprehension that filters even numbers.
- After "My name is Zorro." and one more turn, "What is my name?" was answered with a different name (Sam).
- "give me a random python code" produced a function named
- Facts are unreliable, and it often confidently continues with plausible-looking nonsense.
- Short memory. 1024 bytes of context covers a couple of turns at most.
- No safety training of any kind. It is too small to be dangerous, but it has no guardrails either.
- Very low training loss does not mean good behaviour here: it mostly measures memorised templates.
Training data notes and license
Part of the training data (cosmopedia-v2, smol-smoltalk, smoltalk and the synthetic persona dialogues)
was generated by other language models, so this model inherits their style and any of their errors.
The code corpus consists of public repositories with mixed licenses.
Check each dataset card for its license before reusing the weights or the data mix.
License: CalmaCat Open License (CCPL-1.0)
- Downloads last month
- 34
Install (macOS, Linux)
# Start a local OpenAI-compatible server with a web UI: llama serve -hf ViorikaAI-org/CalmaCatCoder-Next-mini:F16# Run inference directly in the terminal: llama cli -hf ViorikaAI-org/CalmaCatCoder-Next-mini:F16