Toolcall-2B — GGUF

Quantized builds of ajvikram/toolcall-2b, a 2B function-calling model fine-tuned from Qwen3.5-2B for local agent tool routing. Full results, training details and limitations are on the parent model's card.

File Size Use
toolcall-2b-Q4_K_M.gguf 1.22 GB Default. Smallest sensible quality loss, runs on a laptop CPU.
toolcall-2b-Q5_K_M.gguf 1.35 GB A little closer to full precision for modest extra memory.
toolcall-2b-Q8_0.gguf 1.93 GB Near-lossless; use when you have the memory.
toolcall-2b-f16.gguf 3.63 GB Unquantized source for making your own quants.

Measured on the benchmark harness (safetensors, bf16): 36.35 overall on BFCL v4 against 33.85 for the Qwen3.5-2B base, with every group ahead of the base. The quantized builds are not separately scored.

Run it

llama-server -m toolcall-2b-Q4_K_M.gguf --jinja -c 8192
ollama run hf.co/ajvikram/toolcall-2b-gguf:Q4_K_M

The model uses Qwen3.5's native XML tool-call format, so any client that already parses Qwen3.5 tool calls works unchanged:

<tool_call>
<function=get_weather>
<parameter=city>
Berlin
</parameter>
</function>
</tool_call>

Verified with llama-cli on CPU: the Q4_K_M build loads, generates at roughly 33 tokens per second on an ARM CPU, and returns the call above for a get_weather tool given "What is the weather in Berlin?".

Thinking is off by default, matching how the model was trained and evaluated.

Notes

  • Built with llama.cpp (September 2026), which added Qwen3.5 conversion support; older builds cannot convert this architecture.
  • These are text-only builds. The base architecture is vision-capable, but this model was trained and evaluated purely on text tool calling.
Downloads last month
627
GGUF
Model size
2B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ajvikram/toolcall-2b-gguf

Finetuned
Qwen/Qwen3.5-2B
Quantized
(2)
this model