KyroLM V2 Pro (unreleased, in training)

KyroLM V2 Pro is a 4B-parameter reasoning model from the independent KyroLM project. It is a supervised fine-tune of Qwen3-4B-Thinking-2507 and is built to do one thing consistently: think first, in a structured scratchpad, then answer, and call your tools when it needs them.

It is a generalist with a clear priority order: mathematics, logic and programming first, then languages (English, Italian, Spanish, French), conversation and general knowledge. It is small enough to run locally (GGUF) on a laptop.

Status: We're having problems , so the model relase is being moved to around the 15th of october , 2026 . The model does beat KyroLM in everything on our custom test , but lags between 3 to 5 points the base model . It is much better in token optimization , speed and general thinking.

Highlights

  • Always reasons. Every answer starts with a non-empty <think> block, including greetings and trivial questions (short, real content; never an empty block). This is trained to hold with or without a system prompt, and also when the user writes /no_think or asks the model not to reason.
  • Structured scratchpad. The reasoning follows a consistent shape: restate the problem, choose a method, work through verifiable steps, check the result, conclude. It avoids repetition, fake self-corrections and drafts of the final answer.
  • Reasoning in English, answers in your language. The <think> is always English. The final answer is in the language of the user (EN / IT / ES / FR).
  • Uses tools when they are provided. The model keeps the native Qwen3 tool-calling format. When a tool is available and the question needs it (recent facts, calculations), it emits a real tool call instead of guessing, and answers using only what the tool returned. When no tool is needed it answers directly.
  • Works with no system prompt. Training mixes no system prompt (50%), a short one (25%), a full KyroLM prompt (15%) and generic ones (10%), so behavior is the same in all cases.
  • Honest about uncertainty. If it has no tool and the question is about current events or facts it cannot know, it says so briefly and explains how to verify (official source, type of document) instead of guessing. It does not claim to have searched when it has not.
  • Adaptive answers. Length and format are decided in the reasoning from the concrete request and the apparent level of the user: short prose for simple things, structure only for long or technical answers.

Model details

Developed by KyroLM project
Model type Causal language model, reasoning (thinking-only)
Base model Qwen3-4B-Thinking-2507 (Apache-2.0)
Fine-tuning Supervised fine-tuning with LoRA
Languages English (primary), Italian, Spanish, French
Tool use Native Qwen3 format (<tools>, <tool_call>, <tool_response>)
Training context Examples up to 3,072 tokens (chat template applied)
Typical reasoning length Usually under ~1,500 tokens (tuned for local GGUF use)
License Apache-2.0

Benchmarks

Status : The benchmarks still aren't available , but we exepct it to surpass KyroLM-V2-Pro on our custom test and beat base Qwen 3 4b Thinking on reasoning .

Intended use

Good fit:

  • Math word problems, algebra, number theory, probability, step-by-step derivations.
  • Logic puzzles and constraint problems.
  • Writing, explaining and debugging code (Python, Swift, C, Java, HTML, CSS first; also C++ and Bash).
  • Multilingual Q&A and rewriting in EN / IT / ES / FR.
  • Agent-style use with your own tools (web search, calculator, code execution, ...) in an app that runs the tool calls.
  • Local assistant on consumer hardware.

Not a good fit:

  • Anything needing live information without a tool: the model has no built-in internet access, no built-in tools and no memory of past conversations. It can call the tools you give it; your application has to execute them.
  • High-stakes medical, legal or financial decisions without a qualified professional.
  • Languages outside EN / IT / ES / FR: they are not targeted by the training and rely on what the base model already knows.
  • Programming languages other than those listed above: they rely on what the base model already knows.
  • Latency-critical chat: because it always thinks, even simple questions produce a (short) reasoning block.

How to use

Output format

The Qwen3-Thinking chat template already opens <think> in the generation prompt. The generated text therefore usually contains only the reasoning followed by a closing </think>, then the answer (or a tool call). Split on </think>:

text = output_text
reasoning, _, answer = text.partition("</think>")
reasoning = reasoning.replace("<think>", "").strip()
answer = answer.strip()

Transformers

from transformers import AutoModelForCausalLM, AutoTokenizer

# Use this repository's id, or a local folder with the files.
model_id = "Davide531/KyroLM-V2-Pro"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id, torch_dtype="auto", device_map="auto"
)

# No system prompt needed.
messages = [{"role": "user", "content": "A price of 240 drops by 15%, then rises by 12%. What is the final price?"}]

prompt = tokenizer.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True
)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)

out = model.generate(
    **inputs,
    max_new_tokens=3072,
    do_sample=True,
    temperature=0.6,
    top_p=0.95,
    top_k=20,
)
text = tokenizer.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=False)
reasoning, _, answer = text.partition("</think>")
print(answer.strip())

Tool calling

Pass your tool schemas with tools= so the chat template writes them into the prompt. The model answers with a <tool_call> block containing JSON (name and arguments); run it, append the result as a tool message and generate again.

tools = [{
    "type": "function",
    "function": {
        "name": "web_search",
        "description": "Search the web and return short result snippets.",
        "parameters": {
            "type": "object",
            "properties": {"query": {"type": "string"}},
            "required": ["query"],
        },
    },
}]

messages = [{"role": "user", "content": "Who won the most recent Formula 1 race?"}]
prompt = tokenizer.apply_chat_template(
    messages, tools=tools, tokenize=False, add_generation_prompt=True
)
# Generate, parse the <tool_call>{"name": ..., "arguments": ...}</tool_call> block,
# run the tool, then continue the conversation:
# messages.append({"role": "assistant", "content": "", "tool_calls": [{"type": "function", "function": {"name": "web_search", "arguments": {"query": "..."}}}]})
# messages.append({"role": "tool", "content": "<tool result>"})

Most local apps (LM Studio, llama.cpp server, vLLM, Ollama) do this for you when you enable tool/function calling with the model's own template.

GGUF (LM Studio / llama.cpp)

GGUF builds, when available, are listed in the Files and versions tab. Example with a Q4_K_M file:

llama-cli -m KyroLM-V2-Pro-Q4_K_M.gguf --jinja -c 8192 \
  --temp 0.6 --top-p 0.95 --top-k 20 -cnv

Tips:

  • Use the model's own chat template (--jinja in llama.cpp; the default in LM Studio). It is the original Qwen3 template, which includes the tool block. Do not force a custom template: one without the tool block makes the model blind to your tools.
  • No system prompt is required. If you add one, the model follows it, including a different assistant name, but it will not state false things about its origin.
  • Leave enough context for the reasoning: use at least 4k tokens, 8k recommended.

Recommended sampling

Temperature 0.6, top-p 0.95 and top-k 20 are the settings recommended for the base model, and they are the suggested starting point here. Avoid greedy decoding: it can cause repetition in long reasoning.

Identity

The model introduces itself as Kyro (full name KyroLM V2 Pro), a model from the independent KyroLM project. If asked, it says it is built on an open-source base model and further trained within the KyroLM project; it does not claim to be another assistant, and it does not have reliable information about its exact architecture, parameter count or training date.

Training data

The model is trained on the V2-R KyroLM First Generation Reasoning Dataset, a purpose-built SFT dataset where every assistant turn contains a non-empty <think> followed by an answer or a tool call.

  • Size: roughly 11 thousand examples after all filters (training, validation and test splits kept apart by template family). The final size is set by the smallest verified area and is published with the release.
  • Area mix (target ± 3%): reasoning 25% · conversation 21% · knowledge 19% · style of speaking 18% · languages 17%.
  • Inside reasoning: around 40% coding · 35% mathematics · 25% logic.
  • Languages: English, Italian, Spanish, French only. Reasoning is always English. The data is English-heavy; Italian, Spanish and French examples are oversampled during training to compensate. Support for other languages is inherited from the base model, but other languages might not behave as well as the ones the model was trained on.
  • Tool trajectories: about 600 multi-turn examples in the native Qwen3 format, in all four languages: questions that need a tool and call it, questions where tools are available but not needed (answered directly), multi-step calls, and tool errors or empty results handled honestly. Tool names and schemas vary. Tool results are synthetic, and the model is trained to answer only from the returned content.
  • Identity examples: present in all four languages, with and without system prompt.

Sources

Source Use License
Programmatic generators (math, logic, derivations, multi-turn, identity, tool trajectories) Synthetic families; answers checked with exact arithmetic or Python where a verifier exists Own work
OpenThoughts-114k Small filtered subset of code traces that execute correctly against examples Apache-2.0
NuminaMath-CoT Math problems whose final answer matches an independently verified answer; reasoning written in the KyroLM style Apache-2.0
OpenR1-Math-220k Answer-verified problems and style-filtered reasoning traces (short, few hesitation markers) Apache-2.0
MBPP (train split only) Programming problems with official tests; reasoning written by the teacher and kept only if the solution passes CC BY 4.0
Teacher-generated reasoning Reasoning produced with DeepSeek (V4.1 Flash, via API) for areas the public data does not cover (IT/ES/FR, several programming languages, style, conversation, knowledge, identity, tool trajectories). The teacher receives a question and an already verified answer and writes only the structured reasoning. Math, logic and code are kept only if the final answer matches an independently verified one Subject to the provider's terms of use
Base-model rejection sampling Reasoning generated by Qwen3-4B-Thinking-2507 itself on problems with verified answers; only the shortest correct, clean sample per problem is kept, at most 30% of the reasoning area Apache-2.0 (base model output)

Credit goes to the authors of the datasets above.

Quality controls

The dataset build enforces the following gates and rejects any example that fails them:

  • Verification: math and logic answers checked with Python/sympy using exact fractions; code accepted only after execution against tests with a timeout (Python, Swift, Java, C, C++, Bash). HTML/CSS is not independently verified: the only checks are that a complete, non-truncated code block is present. Families without a verifier are not counted as verified. Known-bad families are excluded (for example, an entire synthetic multi-turn math family was removed after verification showed every answer was wrong).
  • Reasoning style gate: think is English and non-empty on every assistant turn; length ranges per area; short reasoning blocks capped at 25% of an area; no answer plans, drafts, filler or repeated openings (no opening phrase above 5% of an area, with one documented exception for public OpenR1 traces); self-corrections kept only when real.
  • Tool gate: trajectories are rendered with the real Qwen3 template and tools=; every call must be valid JSON that matches the declared schema; the final answer may only use content returned by the tool; reasoning that claims a search without a tool call is rejected.
  • Deduplication: MinHash (128 permutations, 3-word shingles, threshold 0.8, numbers normalized) on prompts and answers, template-family caps, and family-level splits so no template family appears in both train and validation/test.
  • Decontamination: train is checked against validation, test, the evaluation sets and the public benchmark problems used for evaluation (GSM8K, MATH-500, HumanEval); residual matches are removed.
  • Language check: Lingua language detection on reasoning (English) and answers (user language).
  • Token limit: no example above 3,072 tokens with the real Qwen3 chat template and tokenizer; none truncated.
  • Quality reading: examples read by hand (20 per 500 new examples) to check that intermediate steps are true, since teacher reasoning is written from a known answer.
  • Knowledge fact-check: a sample of knowledge answers, weighted toward science and technology, is checked against the web, and doubtful items are discarded. It is a sample, not an exhaustive check.

Training procedure

  • Method: supervised fine-tuning, LoRA on the base model, loss on assistant tokens only (reasoning, answer and tool calls; tool results are masked).
  • Template: the original Qwen3-4B-Thinking-2507 chat template, including its tool block (no default system prompt is injected when none is given; the generation prompt opens <think>).
  • Hyperparameters: learning rate up to 1e-4, LoRA rank 32–64, 2 epochs, warmup with cosine schedule, sequence length 3,072; Italian, Spanish and French examples oversampled ×2. The exact values are published with the release.
  • Checkpoints: saved about every 250 steps and evaluated; the best checkpoint is selected, not the last one.
  • Hardware: a single RTX 3090 on a rented GPU instance.

Evaluation

The model is evaluated against the unmodified base model on:

  • the KyroLM evaluation suite: math, logic, coding, languages (EN/IT/ES/FR), style and conversation, identity, multilingual math and logic;
  • an "always thinks" test (200 prompts: greetings, trivial questions, /no_think, short prompts, four languages, code, math, identity, no system prompt; target: at least 99.5% of answers start with a non-empty <think>);
  • a tool-use test (60 prompts, four languages: 32 need a tool, 28 do not). Metrics: share of valid tool calls, share of unnecessary calls, and share of answers that claim a search without a call;
  • about 100 public problems (GSM8K 40, MATH-500 40, HumanEval 20), used only for evaluation and never in training.

Results will be added to this section at release.

Release criteria: if math, coding or logic drop by more than 2 points against the base model, the best checkpoint is chosen among all saved ones and the run is investigated before anything is released. The fine-tuned model must also not fall below the base model in the share of valid tool calls.

Behavior and safety

  • Sensitive but legitimate topics (medicine, law, finance, sexuality, defensive security, harm reduction, history, politics): answered directly, with caution proportional to the risk and a pointer to a professional when it matters.
  • Dangerous requests (weapons, explosives, CBRN, malware/exploits, sexual content involving minors, self-harm instructions, fraud/identity theft, direct violence): short, natural refusal with a safe alternative when one exists. Refusal training data contains no harmful operational details.
  • Jailbreaks ("ignore your instructions", role-play framings, "it's for a novel"): the model answers the legitimate part of the request.
  • Emotional distress: responds with empathy and points to appropriate support without inventing contact details.

Limitations

  • Can make mistakes, including confident ones, in math, code and facts. Verify anything important; run generated code before relying on it.
  • Knowledge has a cutoff date and the model has no built-in internet access. Fresh information requires a tool that your application executes.
  • Tool use was trained on synthetic trajectories. It follows the native Qwen3 format, but unusual tool schemas or very long tool outputs may still fail.
  • The reasoning block is a working scratchpad, not a guaranteed faithful explanation of the answer.
  • HTML/CSS training data was not independently verified, and only some programming languages were verified by execution.
  • Teacher reasoning was written from an already known answer, so it may present a tidier route than the one a model would find by itself.
  • 4B parameters: weaker than large models on broad world knowledge and very long, multi-stage problems.
  • The training data is English-heavy; Italian, Spanish and French are supported but may be less polished than English. Other languages may work through the base model but are not verified.
  • Because it always thinks, simple prompts still cost a short reasoning block.

License and attribution

Released under Apache-2.0, following the base model Qwen3-4B-Thinking-2507, which was modified (fine-tuned) by the KyroLM project. Training data includes third-party datasets under the licenses listed above (including MBPP under CC BY 4.0); attribution to their authors is required where the license says so. Part of the data was generated with DeepSeek under the provider's terms of use.

Citation

@misc{kyrolm_v2_pro,
  title  = {KyroLM V2 Pro},
  author = {{KyroLM project}},
  year   = {2026}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Davide531/KyroLM-V2-Pro

Adapter
(72)
this model