Text Generation
Transformers
Safetensors
GGUF
English
qwen3
code
100m
conversational
text-generation-inference
Instructions to use ViorikaAI-org/CalmaCatCoder-Next-mini with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ViorikaAI-org/CalmaCatCoder-Next-mini with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="ViorikaAI-org/CalmaCatCoder-Next-mini") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("ViorikaAI-org/CalmaCatCoder-Next-mini") model = AutoModelForCausalLM.from_pretrained("ViorikaAI-org/CalmaCatCoder-Next-mini", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use ViorikaAI-org/CalmaCatCoder-Next-mini with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf ViorikaAI-org/CalmaCatCoder-Next-mini:F16 # Run inference directly in the terminal: llama cli -hf ViorikaAI-org/CalmaCatCoder-Next-mini:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf ViorikaAI-org/CalmaCatCoder-Next-mini:F16 # Run inference directly in the terminal: llama cli -hf ViorikaAI-org/CalmaCatCoder-Next-mini:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf ViorikaAI-org/CalmaCatCoder-Next-mini:F16 # Run inference directly in the terminal: ./llama-cli -hf ViorikaAI-org/CalmaCatCoder-Next-mini:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf ViorikaAI-org/CalmaCatCoder-Next-mini:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf ViorikaAI-org/CalmaCatCoder-Next-mini:F16
Use Docker
docker model run hf.co/ViorikaAI-org/CalmaCatCoder-Next-mini:F16
- LM Studio
- Jan
- vLLM
How to use ViorikaAI-org/CalmaCatCoder-Next-mini with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ViorikaAI-org/CalmaCatCoder-Next-mini" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ViorikaAI-org/CalmaCatCoder-Next-mini", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ViorikaAI-org/CalmaCatCoder-Next-mini:F16
- SGLang
How to use ViorikaAI-org/CalmaCatCoder-Next-mini with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "ViorikaAI-org/CalmaCatCoder-Next-mini" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ViorikaAI-org/CalmaCatCoder-Next-mini", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "ViorikaAI-org/CalmaCatCoder-Next-mini" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ViorikaAI-org/CalmaCatCoder-Next-mini", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use ViorikaAI-org/CalmaCatCoder-Next-mini with Ollama:
ollama run hf.co/ViorikaAI-org/CalmaCatCoder-Next-mini:F16
- Unsloth Desktop
- Docker Model Runner
How to use ViorikaAI-org/CalmaCatCoder-Next-mini with Docker Model Runner:
docker model run hf.co/ViorikaAI-org/CalmaCatCoder-Next-mini:F16
- Lemonade
How to use ViorikaAI-org/CalmaCatCoder-Next-mini with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull ViorikaAI-org/CalmaCatCoder-Next-mini:F16
Run and chat with the model
lemonade run user.CalmaCatCoder-Next-mini-F16
List all available models
lemonade list
- Atomic Chat
|
Download README.md from ViorikaAI-org/CalmaCatCoder-Next-mini: direct link, hf CLI and curl.
- Browser
- Download file 6.98 kB
-
https://huggingface.co/ViorikaAI-org/CalmaCatCoder-Next-mini/resolve/main/README.md
- Command line
-
hf download hf://ViorikaAI-org/CalmaCatCoder-Next-mini/README.md
-
curl -L -o README.md https://huggingface.co/ViorikaAI-org/CalmaCatCoder-Next-mini/resolve/main/README.md
6.98 kB
| license: other | |
| license_name: ccpl-1.0 | |
| license_link: LICENSE | |
| language: | |
| - en | |
| pipeline_tag: text-generation | |
| tags: | |
| - code | |
| - qwen3 | |
| - 100m | |
| library_name: transformers | |
| datasets: | |
| - HuggingFaceTB/smollm-corpus | |
| - HuggingFaceTB/smol-smoltalk | |
| - HuggingFaceTB/smoltalk | |
| # CalmaCatCoder-Next-Mini | |
| A tiny **experimental** chat model (96.6M parameters) trained completely **from scratch on one consumer GPU** | |
| (RTX 5060 Ti, 16 GB) by ViorikaAI. It is a hobby project: a small calm, cat-themed | |
| Python helper that knows who it is, speaks ChatML and writes very simple Python. | |
| > **Please read the limitations below.** This is a learning project, not a useful assistant. | |
| > It memorised a handful of templates and will confidently produce wrong code and wrong facts. | |
| ## Model details | |
| | | | | |
| |---|---| | |
| | Architecture | `Qwen3ForCausalLM` (Qwen3 architecture, **no Qwen weights**, random init) | | |
| | Parameters | 96.6M (tied input/output embeddings) | | |
| | Layers / hidden / FFN | 14 / 768 / 2304 | | |
| | Attention | 12 heads, 4 KV heads, head_dim 64 | | |
| | Context length | **1024 bytes** (not words) | | |
| | Vocabulary | **256 tokens: one token = one UTF-8 byte** (no learned tokenizer) | | |
| | Precision | bf16 weights | | |
| | Language | English only | | |
| ### The tokenizer is byte-level | |
| Token id equals the byte value, so `"def "` becomes `[100, 101, 102, 32]` and one word costs several tokens. | |
| This is why the context is short. `<|im_start|>` and `<|im_end|>` are **plain text made of bytes**, not special | |
| tokens, so the model ends its answer by writing the literal text `<|im_end|>`. | |
| Stop generation on that string (see the example below). | |
| Chat format (the same text the model was trained on, available as `tokenizer.apply_chat_template`): | |
| ``` | |
| <|im_start|>system | |
| You are CalmaCatCoder-Next, a calm and friendly cat-themed coding assistant. You explain things step by step and say so when you are not sure.<|im_end|> | |
| <|im_start|>user | |
| Who are you?<|im_end|> | |
| <|im_start|>assistant | |
| ``` | |
| If you pass no system message, the one above is inserted automatically. Other system prompts were never seen | |
| in training, so expect them to be mostly ignored. | |
| ## Quick start | |
| Requires `transformers>=4.51` (Qwen3 support). No `trust_remote_code` needed. | |
| ```python | |
| import torch | |
| from transformers import AutoTokenizer, AutoModelForCausalLM, StoppingCriteria, StoppingCriteriaList | |
| repo = "ViorikaAI-org/CalmaCatCoder-Next-mini" # or a local folder | |
| device = "cuda" if torch.cuda.is_available() else "cpu" | |
| tok = AutoTokenizer.from_pretrained(repo) | |
| model = AutoModelForCausalLM.from_pretrained(repo).to(device, dtype=torch.bfloat16).eval() | |
| END = tok.encode("<|im_end|>") # the end marker is ordinary bytes | |
| class StopOnEnd(StoppingCriteria): | |
| def __call__(self, input_ids, scores, **kwargs): | |
| return torch.tensor([input_ids[0, -len(END):].tolist() == END] * input_ids.shape[0], | |
| device=input_ids.device) | |
| messages = [{"role": "user", "content": "Who are you?"}] | |
| prompt = tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True) | |
| inputs = tok(prompt, return_tensors="pt").to(device) | |
| with torch.no_grad(): | |
| out = model.generate(**inputs, max_new_tokens=300, | |
| stopping_criteria=StoppingCriteriaList([StopOnEnd()])) | |
| print(tok.decode(out[0, inputs["input_ids"].shape[1]:]).replace("<|im_end|>", "").strip()) | |
| ``` | |
| Sampling defaults live in `generation_config.json` (temperature 0.7, top-k 40). Keep prompt + answer under | |
| 1024 bytes; the model never saw longer sequences. | |
| ## How it was trained | |
| All of it ran on a single RTX 5060 Ti (16 GB) under Windows with a PyTorch nightly build. | |
| **Stage 1: pre-training from scratch (1 epoch, ~7.6 hours).** | |
| 29,908 steps × 32,768 byte-tokens ≈ 980M tokens, about 36k tokens/s. | |
| AdamW (lr 6e-4, betas 0.9/0.95, weight decay 0.1), 300 warm-up steps, cosine decay to 10%, | |
| gradient clipping 1.0, bf16 autocast, `torch.compile`. | |
| Final validation loss 0.508 nats/byte on a held-out 2% of the mix. This number is **not comparable** with other | |
| models: it is per byte, and the held-out part contains near-copies of the templated persona dialogues. | |
| Data mix (~1 GB, target shares; the exact proportions depend on what each source could supply): | |
| - Python code: `codeparrot/codeparrot-clean`, the largest share (about 60%) | |
| - English text: `cosmopedia-v2` from `HuggingFaceTB/smollm-corpus` (15%) | |
| - Math: `meta-math/MetaMathQA` (5%) | |
| - ChatML dialogues: `HuggingFaceTB/smol-smoltalk`, only dialogues shorter than 1000 bytes (up to 20%) | |
| - About 5 MB of synthetic persona dialogues generated by a script (identity, creator, greetings, | |
| a few Python tasks) | |
| **Stage 2: supervised fine-tuning (1 epoch, ~13 minutes).** | |
| About 28k short ChatML dialogues (~17M tokens), loss computed on assistant tokens only, peak lr 1e-4. | |
| Mixture of programmatically generated dialogues (identity and creator questions in many phrasings, greetings, | |
| 18 small Python tasks, follow-ups such as "add a docstring", remembering a name within a chat), | |
| `HuggingFaceTB/smoltalk` (`everyday-conversations`) and a subset of `smol-smoltalk`. | |
| SFT validation loss was 0.123, which mostly reflects how templated that data is. | |
| ## Intended use | |
| Learning and curiosity: look at how a tiny model trained from scratch behaves, reuse the training recipe, | |
| test tokenizers or inference tooling on a very small model. | |
| **Not intended for** real coding help, factual questions, any decision that matters, or production use. | |
| ## Limitations (observed, not hypothetical) | |
| - **English only, almost only Python.** Other languages and programming languages are outside what it learned. | |
| - **It memorised templates.** For a coding request it picks the closest of the ~18 tasks it knows and adapts it. | |
| Real examples from testing: | |
| - "give me a random python code" produced a function named `random_state` that simply reverses a string, | |
| followed by an explanation copied from a different task. | |
| - "Write a function that multiplies all numbers in a list" produced a list comprehension that filters even numbers. | |
| - After "My name is Zorro." and one more turn, "What is my name?" was answered with a different name (Sam). | |
| - **Facts are unreliable**, and it often confidently continues with plausible-looking nonsense. | |
| - **Short memory.** 1024 bytes of context covers a couple of turns at most. | |
| - **No safety training** of any kind. It is too small to be dangerous, but it has no guardrails either. | |
| - Very low training loss does not mean good behaviour here: it mostly measures memorised templates. | |
| ## Training data notes and license | |
| Part of the training data (`cosmopedia-v2`, `smol-smoltalk`, `smoltalk` and the synthetic persona dialogues) | |
| was generated by other language models, so this model inherits their style and any of their errors. | |
| The code corpus consists of public repositories with mixed licenses. | |
| Check each dataset card for its license before reusing the weights or the data mix. | |
| License: CalmaCat Open License (CCPL-1.0) | |