Main repo:

GGUF / llama.cpp

The tokenizer has no BPE merges and a pre-tokenizer unknown to the stock converter, so convert_hf_to_gguf.py fails on it. The make_gguf.py wrapper from the code repository patches this at conversion time without modifying llama.cpp (details in its header):

git clone https://github.com/ggml-org/llama.cpp
python make_gguf.py llama.cpp CalmaCatCoder-Next-mini calmacatcoder-f16.gguf f16     # or q8_0

Because <|im_end|> is plain text, llama.cpp does not stop on it by itself. Save the ChatML prompt to a file and use a reverse prompt:

llama-completion -m calmacatcoder-f16.gguf -no-cnv --no-escape -f prompt.txt -r "<|im_end|>" -n 300 --temp 0.7 --top-k 40

A warning special_eos_id is not in special_eog_ids is expected: the model has no EOS token. Prefer f16/bf16 or q8_0; this model is so small that aggressive quantization will likely hurt it.

Downloads last month
-
GGUF
Model size
96.6M params
Architecture
qwen3
Hardware compatibility
Log In to add your hardware

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ViorikaAI-org/CalmaCatCoder-next-mini-gguf

Quantized
(1)
this model

Collection including ViorikaAI-org/CalmaCatCoder-next-mini-gguf