Text Generation
Transformers
Safetensors
GGUF
English
qwen3
code
100m
conversational
text-generation-inference

CalmaCatCoder-Next-Mini

A tiny experimental chat model (96.6M parameters) trained completely from scratch on one consumer GPU (RTX 5060 Ti, 16 GB) by ViorikaAI. It is a hobby project: a small calm, cat-themed Python helper that knows who it is, speaks ChatML and writes very simple Python.

Please read the limitations below. This is a learning project, not a useful assistant. It memorised a handful of templates and will confidently produce wrong code and wrong facts.

Model details

Architecture Qwen3ForCausalLM (Qwen3 architecture, no Qwen weights, random init)
Parameters 96.6M (tied input/output embeddings)
Layers / hidden / FFN 14 / 768 / 2304
Attention 12 heads, 4 KV heads, head_dim 64
Context length 1024 bytes (not words)
Vocabulary 256 tokens: one token = one UTF-8 byte (no learned tokenizer)
Precision bf16 weights
Language English only

The tokenizer is byte-level

Token id equals the byte value, so "def " becomes [100, 101, 102, 32] and one word costs several tokens. This is why the context is short. <|im_start|> and <|im_end|> are plain text made of bytes, not special tokens, so the model ends its answer by writing the literal text <|im_end|>. Stop generation on that string (see the example below).

Chat format (the same text the model was trained on, available as tokenizer.apply_chat_template):

<|im_start|>system
You are CalmaCatCoder-Next, a calm and friendly cat-themed coding assistant. You explain things step by step and say so when you are not sure.<|im_end|>
<|im_start|>user
Who are you?<|im_end|>
<|im_start|>assistant

If you pass no system message, the one above is inserted automatically. Other system prompts were never seen in training, so expect them to be mostly ignored.

Quick start

Requires transformers>=4.51 (Qwen3 support). No trust_remote_code needed.

import torch
from transformers import AutoTokenizer, AutoModelForCausalLM, StoppingCriteria, StoppingCriteriaList

repo = "ViorikaAI-org/CalmaCatCoder-Next-mini"      # or a local folder
device = "cuda" if torch.cuda.is_available() else "cpu"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo).to(device, dtype=torch.bfloat16).eval()

END = tok.encode("<|im_end|>")                      # the end marker is ordinary bytes
class StopOnEnd(StoppingCriteria):
    def __call__(self, input_ids, scores, **kwargs):
        return torch.tensor([input_ids[0, -len(END):].tolist() == END] * input_ids.shape[0],
                            device=input_ids.device)

messages = [{"role": "user", "content": "Who are you?"}]
prompt = tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tok(prompt, return_tensors="pt").to(device)
with torch.no_grad():
    out = model.generate(**inputs, max_new_tokens=300,
                         stopping_criteria=StoppingCriteriaList([StopOnEnd()]))
print(tok.decode(out[0, inputs["input_ids"].shape[1]:]).replace("<|im_end|>", "").strip())

Sampling defaults live in generation_config.json (temperature 0.7, top-k 40). Keep prompt + answer under 1024 bytes; the model never saw longer sequences.

GGUF / llama.cpp

The tokenizer has no BPE merges and a pre-tokenizer unknown to the stock converter, so convert_hf_to_gguf.py fails on it. The make_gguf.py wrapper from the code repository patches this at conversion time without modifying llama.cpp (details in its header):

git clone https://github.com/ggml-org/llama.cpp
python make_gguf.py llama.cpp CalmaCatCoder-Next-96M calmacat-96m-f16.gguf f16     # or q8_0

Because <|im_end|> is plain text, llama.cpp does not stop on it by itself. Save the ChatML prompt to a file and use a reverse prompt:

llama-completion -m calmacat-96m-f16.gguf -no-cnv --no-escape -f prompt.txt -r "<|im_end|>" -n 300 --temp 0.7 --top-k 40

A warning special_eos_id is not in special_eog_ids is expected: the model has no EOS token. Prefer f16/bf16 or q8_0; this model is so small that aggressive quantization will likely hurt it.

How it was trained

All of it ran on a single RTX 5060 Ti (16 GB) under Windows with a PyTorch nightly build.

Stage 1: pre-training from scratch (1 epoch, ~7.6 hours). 29,908 steps × 32,768 byte-tokens ≈ 980M tokens, about 36k tokens/s. AdamW (lr 6e-4, betas 0.9/0.95, weight decay 0.1), 300 warm-up steps, cosine decay to 10%, gradient clipping 1.0, bf16 autocast, torch.compile. Final validation loss 0.508 nats/byte on a held-out 2% of the mix. This number is not comparable with other models: it is per byte, and the held-out part contains near-copies of the templated persona dialogues.

Data mix (~1 GB, target shares; the exact proportions depend on what each source could supply):

  • Python code: codeparrot/codeparrot-clean, the largest share (about 60%)
  • English text: cosmopedia-v2 from HuggingFaceTB/smollm-corpus (15%)
  • Math: meta-math/MetaMathQA (5%)
  • ChatML dialogues: HuggingFaceTB/smol-smoltalk, only dialogues shorter than 1000 bytes (up to 20%)
  • About 5 MB of synthetic persona dialogues generated by a script (identity, creator, greetings, a few Python tasks)

Stage 2: supervised fine-tuning (1 epoch, ~13 minutes). About 28k short ChatML dialogues (~17M tokens), loss computed on assistant tokens only, peak lr 1e-4. Mixture of programmatically generated dialogues (identity and creator questions in many phrasings, greetings, 18 small Python tasks, follow-ups such as "add a docstring", remembering a name within a chat), HuggingFaceTB/smoltalk (everyday-conversations) and a subset of smol-smoltalk. SFT validation loss was 0.123, which mostly reflects how templated that data is.

Intended use

Learning and curiosity: look at how a tiny model trained from scratch behaves, reuse the training recipe, test tokenizers or inference tooling on a very small model.

Not intended for real coding help, factual questions, any decision that matters, or production use.

Limitations (observed, not hypothetical)

  • English only, almost only Python. Other languages and programming languages are outside what it learned.
  • It memorised templates. For a coding request it picks the closest of the ~18 tasks it knows and adapts it. Real examples from testing:
    • "give me a random python code" produced a function named random_state that simply reverses a string, followed by an explanation copied from a different task.
    • "Write a function that multiplies all numbers in a list" produced a list comprehension that filters even numbers.
    • After "My name is Zorro." and one more turn, "What is my name?" was answered with a different name (Sam).
  • Facts are unreliable, and it often confidently continues with plausible-looking nonsense.
  • Short memory. 1024 bytes of context covers a couple of turns at most.
  • No safety training of any kind. It is too small to be dangerous, but it has no guardrails either.
  • Very low training loss does not mean good behaviour here: it mostly measures memorised templates.

Training data notes and license

Part of the training data (cosmopedia-v2, smol-smoltalk, smoltalk and the synthetic persona dialogues) was generated by other language models, so this model inherits their style and any of their errors. The code corpus consists of public repositories with mixed licenses. Check each dataset card for its license before reusing the weights or the data mix.

License: CalmaCat Open License (CCPL-1.0)

Downloads last month
34
Safetensors
Model size
96.6M params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train ViorikaAI-org/CalmaCatCoder-Next-mini

Collection including ViorikaAI-org/CalmaCatCoder-Next-mini