--- license: other license_name: ccpl-1.0 license_link: LICENSE language: - en pipeline_tag: text-generation tags: - code - qwen3 - 100m library_name: transformers datasets: - HuggingFaceTB/smollm-corpus - HuggingFaceTB/smol-smoltalk - HuggingFaceTB/smoltalk --- # CalmaCatCoder-Next-Mini A tiny **experimental** chat model (96.6M parameters) trained completely **from scratch on one consumer GPU** (RTX 5060 Ti, 16 GB) by ViorikaAI. It is a hobby project: a small calm, cat-themed Python helper that knows who it is, speaks ChatML and writes very simple Python. > **Please read the limitations below.** This is a learning project, not a useful assistant. > It memorised a handful of templates and will confidently produce wrong code and wrong facts. ## Model details | | | |---|---| | Architecture | `Qwen3ForCausalLM` (Qwen3 architecture, **no Qwen weights**, random init) | | Parameters | 96.6M (tied input/output embeddings) | | Layers / hidden / FFN | 14 / 768 / 2304 | | Attention | 12 heads, 4 KV heads, head_dim 64 | | Context length | **1024 bytes** (not words) | | Vocabulary | **256 tokens: one token = one UTF-8 byte** (no learned tokenizer) | | Precision | bf16 weights | | Language | English only | ### The tokenizer is byte-level Token id equals the byte value, so `"def "` becomes `[100, 101, 102, 32]` and one word costs several tokens. This is why the context is short. `<|im_start|>` and `<|im_end|>` are **plain text made of bytes**, not special tokens, so the model ends its answer by writing the literal text `<|im_end|>`. Stop generation on that string (see the example below). Chat format (the same text the model was trained on, available as `tokenizer.apply_chat_template`): ``` <|im_start|>system You are CalmaCatCoder-Next, a calm and friendly cat-themed coding assistant. You explain things step by step and say so when you are not sure.<|im_end|> <|im_start|>user Who are you?<|im_end|> <|im_start|>assistant ``` If you pass no system message, the one above is inserted automatically. Other system prompts were never seen in training, so expect them to be mostly ignored. ## Quick start Requires `transformers>=4.51` (Qwen3 support). No `trust_remote_code` needed. ```python import torch from transformers import AutoTokenizer, AutoModelForCausalLM, StoppingCriteria, StoppingCriteriaList repo = "ViorikaAI-org/CalmaCatCoder-Next-mini" # or a local folder device = "cuda" if torch.cuda.is_available() else "cpu" tok = AutoTokenizer.from_pretrained(repo) model = AutoModelForCausalLM.from_pretrained(repo).to(device, dtype=torch.bfloat16).eval() END = tok.encode("<|im_end|>") # the end marker is ordinary bytes class StopOnEnd(StoppingCriteria): def __call__(self, input_ids, scores, **kwargs): return torch.tensor([input_ids[0, -len(END):].tolist() == END] * input_ids.shape[0], device=input_ids.device) messages = [{"role": "user", "content": "Who are you?"}] prompt = tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True) inputs = tok(prompt, return_tensors="pt").to(device) with torch.no_grad(): out = model.generate(**inputs, max_new_tokens=300, stopping_criteria=StoppingCriteriaList([StopOnEnd()])) print(tok.decode(out[0, inputs["input_ids"].shape[1]:]).replace("<|im_end|>", "").strip()) ``` Sampling defaults live in `generation_config.json` (temperature 0.7, top-k 40). Keep prompt + answer under 1024 bytes; the model never saw longer sequences. ## How it was trained All of it ran on a single RTX 5060 Ti (16 GB) under Windows with a PyTorch nightly build. **Stage 1: pre-training from scratch (1 epoch, ~7.6 hours).** 29,908 steps × 32,768 byte-tokens ≈ 980M tokens, about 36k tokens/s. AdamW (lr 6e-4, betas 0.9/0.95, weight decay 0.1), 300 warm-up steps, cosine decay to 10%, gradient clipping 1.0, bf16 autocast, `torch.compile`. Final validation loss 0.508 nats/byte on a held-out 2% of the mix. This number is **not comparable** with other models: it is per byte, and the held-out part contains near-copies of the templated persona dialogues. Data mix (~1 GB, target shares; the exact proportions depend on what each source could supply): - Python code: `codeparrot/codeparrot-clean`, the largest share (about 60%) - English text: `cosmopedia-v2` from `HuggingFaceTB/smollm-corpus` (15%) - Math: `meta-math/MetaMathQA` (5%) - ChatML dialogues: `HuggingFaceTB/smol-smoltalk`, only dialogues shorter than 1000 bytes (up to 20%) - About 5 MB of synthetic persona dialogues generated by a script (identity, creator, greetings, a few Python tasks) **Stage 2: supervised fine-tuning (1 epoch, ~13 minutes).** About 28k short ChatML dialogues (~17M tokens), loss computed on assistant tokens only, peak lr 1e-4. Mixture of programmatically generated dialogues (identity and creator questions in many phrasings, greetings, 18 small Python tasks, follow-ups such as "add a docstring", remembering a name within a chat), `HuggingFaceTB/smoltalk` (`everyday-conversations`) and a subset of `smol-smoltalk`. SFT validation loss was 0.123, which mostly reflects how templated that data is. ## Intended use Learning and curiosity: look at how a tiny model trained from scratch behaves, reuse the training recipe, test tokenizers or inference tooling on a very small model. **Not intended for** real coding help, factual questions, any decision that matters, or production use. ## Limitations (observed, not hypothetical) - **English only, almost only Python.** Other languages and programming languages are outside what it learned. - **It memorised templates.** For a coding request it picks the closest of the ~18 tasks it knows and adapts it. Real examples from testing: - "give me a random python code" produced a function named `random_state` that simply reverses a string, followed by an explanation copied from a different task. - "Write a function that multiplies all numbers in a list" produced a list comprehension that filters even numbers. - After "My name is Zorro." and one more turn, "What is my name?" was answered with a different name (Sam). - **Facts are unreliable**, and it often confidently continues with plausible-looking nonsense. - **Short memory.** 1024 bytes of context covers a couple of turns at most. - **No safety training** of any kind. It is too small to be dangerous, but it has no guardrails either. - Very low training loss does not mean good behaviour here: it mostly measures memorised templates. ## Training data notes and license Part of the training data (`cosmopedia-v2`, `smol-smoltalk`, `smoltalk` and the synthetic persona dialogues) was generated by other language models, so this model inherits their style and any of their errors. The code corpus consists of public repositories with mixed licenses. Check each dataset card for its license before reusing the weights or the data mix. License: CalmaCat Open License (CCPL-1.0)