OculusMind-AI's picture
OculusMind-ToolCall-8B-v1 (Pico): GGUF Q4_K_M + Q8_0, LoRA adapter, benchmarks
cd08c63 verified
Raw History Blame Contribute Delete
3.89 kB
# Ollama Modelfile β€” OculusMind-ToolCall-8B-v1
#
# Point FROM at whichever quantization you downloaded, then:
# ollama create oculusmind-toolcall-8b -f Modelfile
# ollama run oculusmind-toolcall-8b
#
# WHY THIS FILE EXISTS. This model's whole value is tool calling, and tool
# calling breaks silently when the prompt format or stop tokens are wrong β€”
# the model emits text that looks fine and your parser finds no call. Pinning
# the template and stops here removes the most common way a download ends up
# looking broken.
FROM ./gguf/Q4_K_M/model-q4_k_m.gguf
# For the higher-fidelity build instead:
# FROM ./gguf/Q8_0/model-q8_0.gguf
# --- stop tokens -------------------------------------------------------------
# </s> is the tokenizer's declared eos. The bracketed control tokens are the
# Mistral v13 tool-calling grammar; without them a tool call can run on into
# the next section and fail to parse.
PARAMETER stop "</s>"
PARAMETER stop "[INST]"
PARAMETER stop "[/INST]"
PARAMETER stop "[TOOL_RESULTS]"
# --- sampling ----------------------------------------------------------------
# Near-deterministic by default. NOTE: the benchmarks in this repository were
# measured at BFCL's own default temperature of 0.001 (not 0), with llama.cpp's
# default top_k/top_p/min_p, and through a different serving path than this
# Modelfile uses β€” see eval/reproduce.md. Treat these settings as a sane
# starting point, not as a way to reproduce the published numbers.
PARAMETER temperature 0.001
# Evaluation served a 65,536-token window; this is lower only to bound memory.
# Raise it toward 65536 if your machine allows, especially for long tool schemas.
PARAMETER num_ctx 32768
# --- default system prompt ---------------------------------------------------
# DELIBERATE OVERRIDE, and worth reading before you remove it.
#
# The base model's embedded chat template carries a default system message that
# introduces the model as "Ministral-3-8B-Instruct-2512 ... created by Mistral
# AI" powering "an AI assistant called Le Chat". That default fires whenever a
# caller supplies no system prompt of their own. It is factually about the BASE
# model, not this derivative, and it names another company's product β€” so a
# user who just runs the model with no system prompt gets a misleading
# self-description.
#
# We DID modify the embedded template, and it is disclosed: three factual
# statements in its default system message were changed β€” the model's identity,
# the provenance of its knowledge cutoff, and its modality (the base claims it
# can read images, which is false for this text-only build). All of Mistral's
# template logic, including its tool-calling guidance, is preserved verbatim.
# See NOTICE, "Statement of modifications".
#
# The SYSTEM line below OVERRIDES that default for anyone running this
# Modelfile. It is a neutral production default; replace it with your own agent
# prompt. This model was fine-tuned on tool-calling trajectories carrying real
# system prompts and tool schemas, and performs best when given them.
#
# Note for anyone reproducing our benchmarks: our BFCL numbers were NOT produced
# through this path. They come from a harness that renders prompts itself and
# calls llama.cpp's raw /completion endpoint, so neither this SYSTEM line nor
# the embedded template was in play. See eval/reproduce.md.
# NOTE: no SYSTEM line is set on purpose. The GGUF already embeds an
# identity-only default. A SYSTEM prompt telling the model when to call tools
# and when to decline is deliberately NOT shipped: instructing the behaviour a
# benchmark measures inflates the result against a base that never got the same
# instruction, and would bias exactly the categories where this model is
# weakest.
# Supply your own agent prompt; this model performs best with real system
# prompts and tool schemas, which is what it was fine-tuned on.