Instructions to use M1n1A1/MiniAI-Quata2-4b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use M1n1A1/MiniAI-Quata2-4b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="M1n1A1/MiniAI-Quata2-4b") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("M1n1A1/MiniAI-Quata2-4b") model = AutoModelForCausalLM.from_pretrained("M1n1A1/MiniAI-Quata2-4b", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use M1n1A1/MiniAI-Quata2-4b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "M1n1A1/MiniAI-Quata2-4b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "M1n1A1/MiniAI-Quata2-4b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/M1n1A1/MiniAI-Quata2-4b
- SGLang
How to use M1n1A1/MiniAI-Quata2-4b with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "M1n1A1/MiniAI-Quata2-4b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "M1n1A1/MiniAI-Quata2-4b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "M1n1A1/MiniAI-Quata2-4b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "M1n1A1/MiniAI-Quata2-4b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use M1n1A1/MiniAI-Quata2-4b with Docker Model Runner:
docker model run hf.co/M1n1A1/MiniAI-Quata2-4b
MiniAI Quata2 (4B)
Smarter, more honest, and it thinks before it speaks. 4 billion parameters.
Quata2 is the successor to Quata1.5: a compact 4B language model from MiniAI in Belgrade, built on a Qwen3-4B foundation and rebuilt from the ground up to fix everything Quata1.5 got wrong. It beats its own base model on 7 of 8 standard benchmarks, thinks step by step with native Qwen3 reasoning, and stands its ground when it's right — all in a package that fits on a single modest GPU or runs entirely on your own hardware.
Highlights
- Beats Qwen3-4B — wins 7 of 8 knowledge and common-sense benchmarks against its own base
- Native thinking mode — Qwen3
<think>reasoning, switch it on or off per message - Honest under pressure — holds correct answers when pushed back on, owns real mistakes
- Knows the Balkans — better Serbian facts, no more putting every company in Belgrade
- Speaks BCMS — tops the Serbian LLM Eval among the 4B models we tested
- 100% on-device option — nothing leaves your machine
Benchmarks
Every number below was measured on the same harness, with the same questions, for all three models: 500 questions per benchmark, zero-shot.
ARC-Challenge — hard grade-school science

ARC-Easy — grade-school science

HellaSwag — common-sense sentence completion

Winogrande — pronoun and common-sense reasoning

PIQA — physical common sense

BoolQ — yes/no reading comprehension

OpenBookQA — science facts plus reasoning

TruthfulQA — avoiding common misconceptions

IFEval — following precise instructions

MMLU — knowledge across 57 subjects

Serbian LLM Eval — the same kind of tasks, in Serbian

| Benchmark | MiniAI Quata2 (4B) | Qwen3-4B (base) | MiniAI Quata1.5 (4B) |
|---|---|---|---|
| ARC-Challenge | 54.8% | 50.6% | 52.4% |
| ARC-Easy | 82.4% | 78.4% | 78.2% |
| HellaSwag | 70.8% | 67.2% | 69.0% |
| Winogrande | 71.8% | 67.6% | 64.0% |
| PIQA | 76.8% | 74.6% | 74.6% |
| BoolQ | 84.4% | 85.4% | 84.8% |
| OpenBookQA | 40.6% | 40.4% | 40.2% |
| TruthfulQA (MC2) | 49.3% | 47.4% | 46.0% |
| IFEval (strict) | 71.4% | 69.2% | 70.4% |
| GSM8K | 84.4% | 85.4% | 76.2% |
| MMLU | 68.6% | 66.8% | 67.8% |
| Serbian LLM Eval (avg of 7) | 50.7% | 49.1% | 50.0% |
Quata2 leads on 10 of 12 benchmarks, beats its own Qwen3-4B base on the knowledge and common-sense suite by +2.4 points on average, and jumps +8.2 points on GSM8K maths over Quata1.5. That's the return you get from fixing a model's flaws instead of papering over them.
The fixes, measured
Quata2 was built specifically to fix Quata1.5's weak spots. Our own targeted checks (small test sets, so read them as directional):
| Check | Quata1.5 | Quata2 |
|---|---|---|
| Doesn't put every company in Belgrade | 45% | 100% |
| Knows where companies are | 58% | 83% |
| Facts about Serbia | 50% | 62% |
| Holds correct answers under pushback, owns real mistakes | 67% | 100% |
| Knows who it is, never claims to be ChatGPT | 90% | 100% |
| Thinking mode reaches the right answer | 80% | 90% |
Get started
Run it with Transformers
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "M1n1A1/MiniAI-Quata2-4b"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype="auto", device_map="auto")
messages = [{"role": "user", "content": "Ko te je napravio?"}]
prompt = tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True,
enable_thinking=False) # True = think step by step first
inputs = tok(prompt, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=512)
print(tok.decode(out[0][inputs.input_ids.shape[1]:], skip_special_tokens=True))
Thinking mode: enable_thinking=True, or put /think / /no_think in a
message. Give it room (2,000–4,000 new tokens) when it thinks.
GGUF
- Soonâ„¢
Hosted API
- Soonâ„¢
Details
- Architecture: Qwen3-based, 4B parameters, 36 layers, hidden size 2560
- Context length: 40,960 tokens
- Precision: bfloat16 safetensors (~8 GB)
- Training: supervised fine-tune on ~108k curated conversations (OpenHermes-2.5, Tulu-3 instruction following, MiniAI identity data, self-distilled verified reasoning) → preference tuning (DPO) on Quata1.5's failure cases → weight blend toward Qwen3-4B (WiSE-FT, λ = 0.8) to keep the base model's strengths
- Languages: English first; Serbian, Croatian and Bosnian understood and spoken, with grammar still behind English — better Balkan-language ability is the focus of the next Quata models
- License: Apache 2.0 (base)
Citation
If you use Quata2 in your work, please cite:
@misc{miniai2026quata2,
title = {MiniAI Quata2 (4B)},
author = {{MiniAI}},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/M1n1A1/MiniAI-Quata2-4b}},
note = {Fine-tuned from Qwen3-4B}
}
and the base model:
@misc{qwen3technicalreport,
title = {Qwen3 Technical Report},
author = {{Qwen Team}},
year = {2025},
eprint = {2505.09388},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2505.09388}
}
- Downloads last month
- -
