🦙 Llama-3.1-8B-Function-Calling-Agent

Function calling that generalizes to tool schemas it has never seen

Base Method Quant
Unseen schemas VRAM License

Built with Llama


A function-calling fine-tune of Llama-3.1-8B-Instruct, trained to read JSON tool schemas at inference time and emit well-formed <tool_call> blocks — then answer in natural English once the tool result comes back. It runs on one 8 GB consumer GPU at ~40 tok/s.

🎯 Unseen schemas 🔑 Seen schemas ✅ JSON validity 💾 Total VRAM
24 / 26 48 / 56 57 / 57 5055 MiB
92% 85% 100% of 7774 MiB

Those first two columns are the whole result. Near-identical performance on schemas the model has never seen means it learned "read the schema, build the call" rather than the tool names it was trained on.

👀 A session

you > Do you remember what my job is?
agent> Your job is that of a software engineer.
  · prompt 713 tok 0.8s (860 t/s) · gen 10 tok 0.2s (38 t/s)

you > What is the current population of Japan?
  · google_search({'query': 'current population of Japan', 'num_results': 1})
      -> 3 sources (wikipedia): en.wikipedia.org
agent> According to Wikipedia, Japan's population peaked at 128.5 million people in 2010
       and had declined to approximately 123.4 million by April 2025. By 2050, an estimated
       one-third of the population in Japan is expected to be 65 and older due to
       sub-replacement fertility rates and high life expectancy. […]
  · prompt 1390 tok 1.3s (1044 t/s) · gen 137 tok 3.8s (36 t/s)

you > Do you remember which city I live in? Could you check what the weather will be
      like there tomorrow and whether I should take an umbrella?
  · get_weather({'location': 'New York', 'unit': 'celsius', 'days_ahead': 1}) -> ok
agent> In New York, there will be rain tomorrow with a minimum temperature of 19.6 degrees
       Celsius and a maximum of 25.6 degrees Celsius. It would be wise to bring an umbrella.
  · prompt 420 tok 0.3s (1232 t/s) · gen 41 tok 1.1s (37 t/s)

Two things in there are the card's claims, live:

  • 🎓 days_ahead was never in the training data. It is an extension to the get_weather schema added after the fine-tune, and the model reads "tomorrow" → days_ahead: 1 straight out of the parameter description — the same schema-reading ability the 92% measures.
  • 1️⃣ The search turn asked for num_results: 1 and got three sources back: the reference agent overriding a learned habit, see It asks for one search result under Limitations.

📦 What is in this repository

Two separate artifacts. They are not interchangeable.

file what it is needs
🟢 llama31-8b-Q4_K_M.gguf
4.92 GB
Merged model — base + LoRA, imatrix-calibrated llama.cpp / llama-server. Self-contained
🔵 adapter_model.safetensors
adapter_config.json
336 MB
The LoRA adapter only unsloth/Meta-Llama-3.1-8B-Instruct + PEFT

⚡ Quick start

llama.cpp

./llama-server -m llama31-8b-Q4_K_M.gguf \
  -c 8192 -ngl 99 -fa on \
  --cache-type-k q8_0 --cache-type-v q8_0 \
  --host 127.0.0.1 --port 8080

The q8_0 KV cache halves cache cost to ~68 KB/token, which is what makes 8192 context fit alongside the weights on an 8 GB card.

transformers + PEFT

from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

base = "unsloth/Meta-Llama-3.1-8B-Instruct"
model = AutoModelForCausalLM.from_pretrained(base, torch_dtype="bfloat16", device_map="auto")
model = PeftModel.from_pretrained(model, "Berhak/Llama-3.1-8B-Function-Calling-Agent")
tokenizer = AutoTokenizer.from_pretrained(base)

🔌 Prompt contract

⚠️ This matters more than the weights.

The model was trained on one specific shape and behaves poorly outside it. All three pieces below are load-bearing.

1️⃣ System turn

Tool schemas go inside <tools> as a JSON array, in Llama-3.1's native header format:

<|begin_of_text|><|start_header_id|>system<|end_header_id|>

Cutting Knowledge Date: December 2023
Today Date: 12 Sep 2026

You are a function calling AI model. You are provided with function signatures within <tools></tools> XML tags. You may call one or more functions to assist with the user query. Don't make assumptions about what values to plug into functions.
<tools>
[{"type": "function", "function": {"name": "get_weather", "description": "Get current or forecast weather conditions for a location.", "parameters": {"type": "object", "properties": {"location": {"type": "string", "description": "The city name, e.g. 'Ankara'."}, "unit": {"type": "string", "enum": ["celsius", "fahrenheit"], "description": "Temperature unit."}}, "required": ["location"]}}}]
</tools>
For each function call return a json object with function name and arguments within <tool_call></tool_call> XML tags.<|eot_id|>

📅 Pass the real date. strftime("%b") follows the OS locale and produces a non-English month abbreviation under a non-English one, which breaks the template.

2️⃣ What the model emits

<tool_call>
{"name": "get_weather", "arguments": {"location": "Berlin", "unit": "celsius"}}
</tool_call>

🔢 It emits several blocks in one generation when the request needs several tools. Parse every block, not just the first.

3️⃣ Observations — the unusual part

Tool results go back on the ipython role, double-encoded: the <tool_response> block is wrapped in a JSON string, escaped quotes and literal \n included. An artifact of the original Hermes conversion, but the model learned that exact shape:

<|start_header_id|>ipython<|end_header_id|>

"<tool_response>\n{\"name\": \"get_weather\", \"content\": {\"location\": \"Berlin\", \"temperature\": 18.4}}\n</tool_response>"<|eot_id|>
observation = json.dumps(
    "<tool_response>\n" + json.dumps({"name": name, "content": result}) + "\n</tool_response>"
)

🚨 A plain, single-encoded block puts the model off its training distribution.


🔒 Recommended: constrain decoding with a grammar

Build a GBNF grammar from the tool schemas you pass in this request — not from a fixed file — with a root that permits either a run of calls or plain prose:

root      ::= tool-call (ws-nl tool-call)* | prose
tool-call ::= "<tool_call>" ws-nl call ws-nl "</tool_call>"
prose     ::= [^<] ( [^<] | "<" )*
ws-nl     ::= [ \t\n]*
grammar off grammar on
JSON validity 56/59 · 94% 57/57 · 100%
Overall 71/82 · 86% 72/82 · 87%
Unparseable blocks 3 0

No category regressed, and the gain concentrates where the model is weakest unaided — generations needing more than one call. Three details are worth copying verbatim:

  • 🕳️ The root must allow prose, or every question is forced into a tool call.
  • ⚖️ prose bans < only in the first position. Banning it everywhere ([^<]+) made the character unrepresentable — asked for an inequality the model wrote 3 ≤ 5 instead of 3 < 5, a different claim, not a formatting quirk.
  • 🧩 Build it per request. A fixed grammar forbids the caller's own schemas and scores 0/11 on unseen ones.

📊 Evaluation

86 hand-written held-out records, screened against the training data for verbatim matches and at a 0.70 similarity threshold. Every tool name claimed as "unseen" was checked against all 2,985 tool names present in training. Scored with the grammar on:

metric score
🎯 Unseen tool schemas 24 / 26 █████████░ 92%
🔑 Seen tool schemas 48 / 56 ████████░░ 85%
✅ JSON validity 57 / 57 ██████████ 100%
🎓 Calls built correctly from an unseen schema 11 / 11 ██████████ 100%
🚫 Not calling a tool when none is needed 8 / 8 ██████████ 100%
🌐 Live-information questions → search 6 / 6 ██████████ 100%
🧮 Arithmetic routed to a calculator tool 5 / 5 ██████████ 100%
🔗 Two calls in one generation 2 / 2 ██████████ 100%
📋 Several tasks in one message 5 / 7 ███████░░░ 71%
🪤 Keyword collision without picking the wrong tool 3 / 5 ██████░░░░ 60%
👻 Not fabricating a missing required argument 1 / 4 ██░░░░░░░░ 25%
🧩 Overall 72 / 82 █████████░ 87%

Four ambiguous records are reported but never scored — both behaviours are defensible there, and inventing a reference answer to complete a metric would only make the metric worse.

A handful of records also score the language of the final answer against Turkish (see below), so those are not an English measurement. Every row above except Overall measures tool selection and argument construction, which is language-independent.


⚠️ Limitations

🚧 It does not call a tool after seeing an observation.

This is structural, not a weak tendency.

Of the 8,144 assistant turns that follow a tool result across 9,146 training examples, zero are a tool call. The data only ever contains CALL → OBSERVE → ANSWER; a second call appears solely after a new user turn. A directive telling it to continue scored 0/3.

✅ Plan around it: a request needing several tools must be satisfied by the first generation. The model does this reliably, so parse every block it emits. Do not build a loop that expects it to iterate.


👻 It fabricates arguments it does not have. "What's the weather?" with no city produced location='Ankara'; "is there a train from Ankara?" produced destination="user's destination" — a literal placeholder, meaning it knows it does not know and fills the slot anyway. No prompt variant fixed this. Validate arguments against the conversation before dispatching.

1️⃣ It asks for one search result. The training data calls search with num_results=1 in 34 of 55 cases. Treat that parameter as a floor in your tool, not a ceiling.

🎭 It invents explanations for things that do not exist. Asked about a fabricated syndrome, it answers confidently and fluently. Empty or failed tool results should say so explicitly in the observation, and say what the model must not do.

🎯 Search queries drift off subject. Asked to search two topics, the second query sometimes lands on a neighbouring one — worse with accumulated context.

🧮 Arithmetic needs a tool. Unaided it produced "1 FP16 = 2 INT4, so 16 GB becomes 32 GB" — the ratio is 4 and the operation is division — then reused the wrong figure as a premise on the next turn. Route calculations to a calculator tool.

📏 Factual accuracy of direct answers was never measured, and 📉 no bfloat16 baseline exists — the merged model was deleted by the quantization pipeline before one was taken, so every number here is absolute rather than a measurement of quantization loss.

🌍 On Turkish. The training mix includes a small Turkish subset (~5%), and the evaluation set was originally written against Turkish final turns. That path is not validated and is not recommended — the agent around this model was brought to a working standard in English only. Treat this as an English function-calling model.


🏋️ Training

QLoRA on 4× H100, a single run with no retry budget. Source: NousResearch/hermes-function-calling-v1, converted from ChatML into Llama-3.1's native chat template before loss masking — Llama-3.1's tokenizer does not recognize <|im_start|> as a special token, so raw ChatML would shatter into meaningless sub-words.

🎛️ LoRA r=32, α=64, dropout 0.05, 4-bit NF4 + double quant, bf16 compute
🎯 Target modules q_proj k_proj v_proj o_proj gate_proj up_proj down_proj
📏 / 🔁 4096 tokens · 2 epochs, early stopping (patience 3)
📉 LR 2e-4 cosine, 3% warmup, weight decay 0.01
📦 Batch 2 per device × 8 gradient accumulation
🗃️ Data 9,146 examples → 8,781 train / 365 eval

🕳️ One warning worth repeating. English final-turn prose was masked out so it would not compete with the intended final-turn behaviour. But Hermes examples that answer without a tool consist of nothing but a final turn — so masking it masked the whole example, and 803 of 851 vanished silently. The resulting skew toward calling a tool is why the model over-triggers, and why the reference agent adds a restraint clause to the system prompt. If you reuse this recipe: count your examples after masking, not before.

Quantization: bf16 merge → GGUF q8_0 → imatrix → Q4_K_M. The intermediate is q8_0 rather than f16 (half the disk, no measurable K-quant penalty; needs --allow-requantize), and the importance matrix was calibrated on 40 samples drawn from the real training distribution.


💻 Measured performance

RTX 4060 Laptop (8 GB), Vulkan, n_ctx=8192, Q4_K_M, q8_0 KV cache, flash attention:

🧠 Weights 4403 MiB
🗂️ KV cache 544 MiB · 68 KB/token
⚙️ Compute buffer 108 MiB
💾 Total 5055 / 7774 MiB
✍️ Generation ~40 tok/s
📥 Prompt processing ~175 tok/s

🚦 Memory is not the binding constraint — prompt processing is.

At 175 tok/s, every 1000 tokens of context costs about six seconds on every subsequent turn. Trim tool output aggressively and keep the system prefix stable so prefix caching holds (measured: 14 tokens reprocessed instead of 440 on the second request).

Sampling: temperature=0.5 top_k=40 top_p=0.9 repeat_penalty=1.1 repeat_last_n=768. At temperature=0 the model locked into repeating one sentence across turns; repeat_penalty is what broke that, not the temperature. Use temperature=0 for reproducible evaluation.


📜 License

Llama 3.1 Community License. Use is subject to the Llama 3.1 License and the Acceptable Use Policy.

Built with Llama.

Training data: NousResearch/hermes-function-calling-v1, subject to its own terms.

@misc{llama31-8b-function-calling-agent,
  title  = {Llama-3.1-8B-Function-Calling-Agent},
  author = {Tanyıldızı, Mahmut Berhak},
  year   = {2026},
  url    = {https://huggingface.co/Berhak/Llama-3.1-8B-Function-Calling-Agent}
}
Downloads last month
150
GGUF
Model size
8B params
Architecture
llama
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Berhak/Llama-3.1-8B-Function-Calling-Agent

Dataset used to train Berhak/Llama-3.1-8B-Function-Calling-Agent