Instructions to use Berhak/Llama-3.1-8B-Function-Calling-Agent with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Berhak/Llama-3.1-8B-Function-Calling-Agent with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("unsloth/Meta-Llama-3.1-8B-Instruct") model = PeftModel.from_pretrained(base_model, "Berhak/Llama-3.1-8B-Function-Calling-Agent") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Berhak/Llama-3.1-8B-Function-Calling-Agent with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Berhak/Llama-3.1-8B-Function-Calling-Agent:Q4_K_M # Run inference directly in the terminal: llama cli -hf Berhak/Llama-3.1-8B-Function-Calling-Agent:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Berhak/Llama-3.1-8B-Function-Calling-Agent:Q4_K_M # Run inference directly in the terminal: llama cli -hf Berhak/Llama-3.1-8B-Function-Calling-Agent:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Berhak/Llama-3.1-8B-Function-Calling-Agent:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf Berhak/Llama-3.1-8B-Function-Calling-Agent:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Berhak/Llama-3.1-8B-Function-Calling-Agent:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf Berhak/Llama-3.1-8B-Function-Calling-Agent:Q4_K_M
Use Docker
docker model run hf.co/Berhak/Llama-3.1-8B-Function-Calling-Agent:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use Berhak/Llama-3.1-8B-Function-Calling-Agent with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Berhak/Llama-3.1-8B-Function-Calling-Agent" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Berhak/Llama-3.1-8B-Function-Calling-Agent", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Berhak/Llama-3.1-8B-Function-Calling-Agent:Q4_K_M
- Ollama
How to use Berhak/Llama-3.1-8B-Function-Calling-Agent with Ollama:
ollama run hf.co/Berhak/Llama-3.1-8B-Function-Calling-Agent:Q4_K_M
- Unsloth Desktop
- Pi
How to use Berhak/Llama-3.1-8B-Function-Calling-Agent with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Berhak/Llama-3.1-8B-Function-Calling-Agent:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Berhak/Llama-3.1-8B-Function-Calling-Agent:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Berhak/Llama-3.1-8B-Function-Calling-Agent with Docker Model Runner:
docker model run hf.co/Berhak/Llama-3.1-8B-Function-Calling-Agent:Q4_K_M
- Lemonade
How to use Berhak/Llama-3.1-8B-Function-Calling-Agent with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Berhak/Llama-3.1-8B-Function-Calling-Agent:Q4_K_M
Run and chat with the model
lemonade run user.Llama-3.1-8B-Function-Calling-Agent-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use Berhak/Llama-3.1-8B-Function-Calling-Agent with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Berhak/Llama-3.1-8B-Function-Calling-Agent:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Berhak/Llama-3.1-8B-Function-Calling-Agent:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Berhak/Llama-3.1-8B-Function-Calling-Agent with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Berhak/Llama-3.1-8B-Function-Calling-Agent:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Berhak/Llama-3.1-8B-Function-Calling-Agent:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
🦙 Llama-3.1-8B-Function-Calling-Agent
Function calling that generalizes to tool schemas it has never seen
Built with Llama
A function-calling fine-tune of Llama-3.1-8B-Instruct, trained to read JSON tool schemas at
inference time and emit well-formed <tool_call> blocks — then answer in natural English once the
tool result comes back. It runs on one 8 GB consumer GPU at ~40 tok/s.
| 🎯 Unseen schemas | 🔑 Seen schemas | ✅ JSON validity | 💾 Total VRAM |
|---|---|---|---|
| 24 / 26 | 48 / 56 | 57 / 57 | 5055 MiB |
| 92% | 85% | 100% | of 7774 MiB |
Those first two columns are the whole result. Near-identical performance on schemas the model has never seen means it learned "read the schema, build the call" rather than the tool names it was trained on.
👀 A session
you > Do you remember what my job is?
agent> Your job is that of a software engineer.
· prompt 713 tok 0.8s (860 t/s) · gen 10 tok 0.2s (38 t/s)
you > What is the current population of Japan?
· google_search({'query': 'current population of Japan', 'num_results': 1})
-> 3 sources (wikipedia): en.wikipedia.org
agent> According to Wikipedia, Japan's population peaked at 128.5 million people in 2010
and had declined to approximately 123.4 million by April 2025. By 2050, an estimated
one-third of the population in Japan is expected to be 65 and older due to
sub-replacement fertility rates and high life expectancy. […]
· prompt 1390 tok 1.3s (1044 t/s) · gen 137 tok 3.8s (36 t/s)
you > Do you remember which city I live in? Could you check what the weather will be
like there tomorrow and whether I should take an umbrella?
· get_weather({'location': 'New York', 'unit': 'celsius', 'days_ahead': 1}) -> ok
agent> In New York, there will be rain tomorrow with a minimum temperature of 19.6 degrees
Celsius and a maximum of 25.6 degrees Celsius. It would be wise to bring an umbrella.
· prompt 420 tok 0.3s (1232 t/s) · gen 41 tok 1.1s (37 t/s)
Two things in there are the card's claims, live:
- 🎓
days_aheadwas never in the training data. It is an extension to theget_weatherschema added after the fine-tune, and the model reads "tomorrow" →days_ahead: 1straight out of the parameter description — the same schema-reading ability the 92% measures. - 1️⃣ The search turn asked for
num_results: 1and got three sources back: the reference agent overriding a learned habit, see It asks for one search result under Limitations.
📦 What is in this repository
Two separate artifacts. They are not interchangeable.
| file | what it is | needs |
|---|---|---|
🟢 llama31-8b-Q4_K_M.gguf4.92 GB |
Merged model — base + LoRA, imatrix-calibrated | llama.cpp / llama-server. Self-contained |
🔵 adapter_model.safetensorsadapter_config.json336 MB |
The LoRA adapter only | unsloth/Meta-Llama-3.1-8B-Instruct + PEFT |
⚡ Quick start
llama.cpp
./llama-server -m llama31-8b-Q4_K_M.gguf \
-c 8192 -ngl 99 -fa on \
--cache-type-k q8_0 --cache-type-v q8_0 \
--host 127.0.0.1 --port 8080
The q8_0 KV cache halves cache cost to ~68 KB/token, which is what makes 8192 context fit alongside the weights on an 8 GB card.
transformers + PEFT
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base = "unsloth/Meta-Llama-3.1-8B-Instruct"
model = AutoModelForCausalLM.from_pretrained(base, torch_dtype="bfloat16", device_map="auto")
model = PeftModel.from_pretrained(model, "Berhak/Llama-3.1-8B-Function-Calling-Agent")
tokenizer = AutoTokenizer.from_pretrained(base)
🔌 Prompt contract
⚠️ This matters more than the weights.
The model was trained on one specific shape and behaves poorly outside it. All three pieces below are load-bearing.
1️⃣ System turn
Tool schemas go inside <tools> as a JSON array, in Llama-3.1's native header format:
<|begin_of_text|><|start_header_id|>system<|end_header_id|>
Cutting Knowledge Date: December 2023
Today Date: 12 Sep 2026
You are a function calling AI model. You are provided with function signatures within <tools></tools> XML tags. You may call one or more functions to assist with the user query. Don't make assumptions about what values to plug into functions.
<tools>
[{"type": "function", "function": {"name": "get_weather", "description": "Get current or forecast weather conditions for a location.", "parameters": {"type": "object", "properties": {"location": {"type": "string", "description": "The city name, e.g. 'Ankara'."}, "unit": {"type": "string", "enum": ["celsius", "fahrenheit"], "description": "Temperature unit."}}, "required": ["location"]}}}]
</tools>
For each function call return a json object with function name and arguments within <tool_call></tool_call> XML tags.<|eot_id|>
📅 Pass the real date. strftime("%b") follows the OS locale and produces a non-English month
abbreviation under a non-English one, which breaks the template.
2️⃣ What the model emits
<tool_call>
{"name": "get_weather", "arguments": {"location": "Berlin", "unit": "celsius"}}
</tool_call>
🔢 It emits several blocks in one generation when the request needs several tools. Parse every block, not just the first.
3️⃣ Observations — the unusual part
Tool results go back on the ipython role, double-encoded: the <tool_response> block is
wrapped in a JSON string, escaped quotes and literal \n included. An artifact of the original
Hermes conversion, but the model learned that exact shape:
<|start_header_id|>ipython<|end_header_id|>
"<tool_response>\n{\"name\": \"get_weather\", \"content\": {\"location\": \"Berlin\", \"temperature\": 18.4}}\n</tool_response>"<|eot_id|>
observation = json.dumps(
"<tool_response>\n" + json.dumps({"name": name, "content": result}) + "\n</tool_response>"
)
🚨 A plain, single-encoded block puts the model off its training distribution.
🔒 Recommended: constrain decoding with a grammar
Build a GBNF grammar from the tool schemas you pass in this request — not from a fixed file — with a root that permits either a run of calls or plain prose:
root ::= tool-call (ws-nl tool-call)* | prose
tool-call ::= "<tool_call>" ws-nl call ws-nl "</tool_call>"
prose ::= [^<] ( [^<] | "<" )*
ws-nl ::= [ \t\n]*
| grammar off | grammar on | |
|---|---|---|
| JSON validity | 56/59 · 94% | 57/57 · 100% |
| Overall | 71/82 · 86% | 72/82 · 87% |
| Unparseable blocks | 3 | 0 |
No category regressed, and the gain concentrates where the model is weakest unaided — generations needing more than one call. Three details are worth copying verbatim:
- 🕳️ The root must allow prose, or every question is forced into a tool call.
- ⚖️
prosebans<only in the first position. Banning it everywhere ([^<]+) made the character unrepresentable — asked for an inequality the model wrote3 ≤ 5instead of3 < 5, a different claim, not a formatting quirk. - 🧩 Build it per request. A fixed grammar forbids the caller's own schemas and scores 0/11 on unseen ones.
📊 Evaluation
86 hand-written held-out records, screened against the training data for verbatim matches and at a 0.70 similarity threshold. Every tool name claimed as "unseen" was checked against all 2,985 tool names present in training. Scored with the grammar on:
| metric | score | |
|---|---|---|
| 🎯 Unseen tool schemas | 24 / 26 | █████████░ 92% |
| 🔑 Seen tool schemas | 48 / 56 | ████████░░ 85% |
| ✅ JSON validity | 57 / 57 | ██████████ 100% |
| 🎓 Calls built correctly from an unseen schema | 11 / 11 | ██████████ 100% |
| 🚫 Not calling a tool when none is needed | 8 / 8 | ██████████ 100% |
| 🌐 Live-information questions → search | 6 / 6 | ██████████ 100% |
| 🧮 Arithmetic routed to a calculator tool | 5 / 5 | ██████████ 100% |
| 🔗 Two calls in one generation | 2 / 2 | ██████████ 100% |
| 📋 Several tasks in one message | 5 / 7 | ███████░░░ 71% |
| 🪤 Keyword collision without picking the wrong tool | 3 / 5 | ██████░░░░ 60% |
| 👻 Not fabricating a missing required argument | 1 / 4 | ██░░░░░░░░ 25% |
| 🧩 Overall | 72 / 82 | █████████░ 87% |
Four ambiguous records are reported but never scored — both behaviours are defensible there, and inventing a reference answer to complete a metric would only make the metric worse.
A handful of records also score the language of the final answer against Turkish (see below), so those are not an English measurement. Every row above except Overall measures tool selection and argument construction, which is language-independent.
⚠️ Limitations
🚧 It does not call a tool after seeing an observation.
This is structural, not a weak tendency.
Of the 8,144 assistant turns that follow a tool result across 9,146 training examples, zero
are a tool call. The data only ever contains CALL → OBSERVE → ANSWER; a second call appears solely
after a new user turn. A directive telling it to continue scored 0/3.
✅ Plan around it: a request needing several tools must be satisfied by the first generation. The model does this reliably, so parse every block it emits. Do not build a loop that expects it to iterate.
👻 It fabricates arguments it does not have. "What's the weather?" with no city produced
location='Ankara'; "is there a train from Ankara?" produced destination="user's destination" —
a literal placeholder, meaning it knows it does not know and fills the slot anyway. No prompt
variant fixed this. Validate arguments against the conversation before dispatching.
1️⃣ It asks for one search result. The training data calls search with num_results=1 in 34 of 55
cases. Treat that parameter as a floor in your tool, not a ceiling.
🎭 It invents explanations for things that do not exist. Asked about a fabricated syndrome, it answers confidently and fluently. Empty or failed tool results should say so explicitly in the observation, and say what the model must not do.
🎯 Search queries drift off subject. Asked to search two topics, the second query sometimes lands on a neighbouring one — worse with accumulated context.
🧮 Arithmetic needs a tool. Unaided it produced "1 FP16 = 2 INT4, so 16 GB becomes 32 GB" — the ratio is 4 and the operation is division — then reused the wrong figure as a premise on the next turn. Route calculations to a calculator tool.
📏 Factual accuracy of direct answers was never measured, and 📉 no bfloat16 baseline exists — the merged model was deleted by the quantization pipeline before one was taken, so every number here is absolute rather than a measurement of quantization loss.
🌍 On Turkish. The training mix includes a small Turkish subset (~5%), and the evaluation set was originally written against Turkish final turns. That path is not validated and is not recommended — the agent around this model was brought to a working standard in English only. Treat this as an English function-calling model.
🏋️ Training
QLoRA on 4× H100, a single run with no retry budget. Source:
NousResearch/hermes-function-calling-v1, converted from ChatML into Llama-3.1's native chat
template before loss masking — Llama-3.1's tokenizer does not recognize <|im_start|> as a special
token, so raw ChatML would shatter into meaningless sub-words.
| 🎛️ LoRA | r=32, α=64, dropout 0.05, 4-bit NF4 + double quant, bf16 compute |
| 🎯 Target modules | q_proj k_proj v_proj o_proj gate_proj up_proj down_proj |
| 📏 / 🔁 | 4096 tokens · 2 epochs, early stopping (patience 3) |
| 📉 LR | 2e-4 cosine, 3% warmup, weight decay 0.01 |
| 📦 Batch | 2 per device × 8 gradient accumulation |
| 🗃️ Data | 9,146 examples → 8,781 train / 365 eval |
🕳️ One warning worth repeating. English final-turn prose was masked out so it would not compete with the intended final-turn behaviour. But Hermes examples that answer without a tool consist of nothing but a final turn — so masking it masked the whole example, and 803 of 851 vanished silently. The resulting skew toward calling a tool is why the model over-triggers, and why the reference agent adds a restraint clause to the system prompt. If you reuse this recipe: count your examples after masking, not before.
Quantization: bf16 merge → GGUF q8_0 → imatrix → Q4_K_M. The intermediate is q8_0 rather than f16
(half the disk, no measurable K-quant penalty; needs --allow-requantize), and the importance
matrix was calibrated on 40 samples drawn from the real training distribution.
💻 Measured performance
RTX 4060 Laptop (8 GB), Vulkan, n_ctx=8192, Q4_K_M, q8_0 KV cache, flash attention:
| 🧠 Weights | 4403 MiB |
| 🗂️ KV cache | 544 MiB · 68 KB/token |
| ⚙️ Compute buffer | 108 MiB |
| 💾 Total | 5055 / 7774 MiB |
| ✍️ Generation | ~40 tok/s |
| 📥 Prompt processing | ~175 tok/s |
🚦 Memory is not the binding constraint — prompt processing is.
At 175 tok/s, every 1000 tokens of context costs about six seconds on every subsequent turn. Trim tool output aggressively and keep the system prefix stable so prefix caching holds (measured: 14 tokens reprocessed instead of 440 on the second request).
Sampling: temperature=0.5 top_k=40 top_p=0.9 repeat_penalty=1.1 repeat_last_n=768.
At temperature=0 the model locked into repeating one sentence across turns; repeat_penalty is
what broke that, not the temperature. Use temperature=0 for reproducible evaluation.
📜 License
Llama 3.1 Community License. Use is subject to the Llama 3.1 License and the Acceptable Use Policy.
Built with Llama.
Training data: NousResearch/hermes-function-calling-v1,
subject to its own terms.
@misc{llama31-8b-function-calling-agent,
title = {Llama-3.1-8B-Function-Calling-Agent},
author = {Tanyıldızı, Mahmut Berhak},
year = {2026},
url = {https://huggingface.co/Berhak/Llama-3.1-8B-Function-Calling-Agent}
}
- Downloads last month
- 150
4-bit
Model tree for Berhak/Llama-3.1-8B-Function-Calling-Agent
Base model
meta-llama/Llama-3.1-8B