# Ollama Modelfile — OculusMind-ToolCall-8B-v1 # # Point FROM at whichever quantization you downloaded, then: # ollama create oculusmind-toolcall-8b -f Modelfile # ollama run oculusmind-toolcall-8b # # WHY THIS FILE EXISTS. This model's whole value is tool calling, and tool # calling breaks silently when the prompt format or stop tokens are wrong — # the model emits text that looks fine and your parser finds no call. Pinning # the template and stops here removes the most common way a download ends up # looking broken. FROM ./gguf/Q4_K_M/model-q4_k_m.gguf # For the higher-fidelity build instead: # FROM ./gguf/Q8_0/model-q8_0.gguf # --- stop tokens ------------------------------------------------------------- # is the tokenizer's declared eos. The bracketed control tokens are the # Mistral v13 tool-calling grammar; without them a tool call can run on into # the next section and fail to parse. PARAMETER stop "" PARAMETER stop "[INST]" PARAMETER stop "[/INST]" PARAMETER stop "[TOOL_RESULTS]" # --- sampling ---------------------------------------------------------------- # Near-deterministic by default. NOTE: the benchmarks in this repository were # measured at BFCL's own default temperature of 0.001 (not 0), with llama.cpp's # default top_k/top_p/min_p, and through a different serving path than this # Modelfile uses — see eval/reproduce.md. Treat these settings as a sane # starting point, not as a way to reproduce the published numbers. PARAMETER temperature 0.001 # Evaluation served a 65,536-token window; this is lower only to bound memory. # Raise it toward 65536 if your machine allows, especially for long tool schemas. PARAMETER num_ctx 32768 # --- default system prompt --------------------------------------------------- # DELIBERATE OVERRIDE, and worth reading before you remove it. # # The base model's embedded chat template carries a default system message that # introduces the model as "Ministral-3-8B-Instruct-2512 ... created by Mistral # AI" powering "an AI assistant called Le Chat". That default fires whenever a # caller supplies no system prompt of their own. It is factually about the BASE # model, not this derivative, and it names another company's product — so a # user who just runs the model with no system prompt gets a misleading # self-description. # # We DID modify the embedded template, and it is disclosed: three factual # statements in its default system message were changed — the model's identity, # the provenance of its knowledge cutoff, and its modality (the base claims it # can read images, which is false for this text-only build). All of Mistral's # template logic, including its tool-calling guidance, is preserved verbatim. # See NOTICE, "Statement of modifications". # # The SYSTEM line below OVERRIDES that default for anyone running this # Modelfile. It is a neutral production default; replace it with your own agent # prompt. This model was fine-tuned on tool-calling trajectories carrying real # system prompts and tool schemas, and performs best when given them. # # Note for anyone reproducing our benchmarks: our BFCL numbers were NOT produced # through this path. They come from a harness that renders prompts itself and # calls llama.cpp's raw /completion endpoint, so neither this SYSTEM line nor # the embedded template was in play. See eval/reproduce.md. # NOTE: no SYSTEM line is set on purpose. The GGUF already embeds an # identity-only default. A SYSTEM prompt telling the model when to call tools # and when to decline is deliberately NOT shipped: instructing the behaviour a # benchmark measures inflates the result against a base that never got the same # instruction, and would bias exactly the categories where this model is # weakest. # Supply your own agent prompt; this model performs best with real system # prompts and tool schemas, which is what it was fine-tuned on.