mtlm-7m-tools β€” a 7M-parameter assistant that talks English and calls tools, trained end-to-end in machin

Fine-tune of javimosch/mtlm-7m-base that decides between answering in plain English and emitting a one-line JSON tool call, over 14 generic tools (weather, calculator, web search, time, read/write file, HTTP GET, email, translate, shell, unit conversion, reminders, notes, wikipedia). Served by anvil, it speaks the OpenAI chat-completions protocol including tool_calls.

Everything except the synthesis of the fine-tuning text is pure machin (MFL): the base pretraining, the chat-template tokenizer, the fine-tune, the int8 export and the server. The text corpus was produced by a templated Python script from the tool catalog (no LLM teacher).

v3 (m7tool3) β€” this revision fixes the number-misquoting flaw of v2 (m7tool2, kept in git history): the training corpus now serializes assistant tool calls byte-identically to how anvil re-renders them in context, tool results cover floats/negatives/6-digit ids, and a fixed unit-conversion bug no longer teaches wrong values.

Honest numbers

base mtlm-7m-base (7.2M params, TinyStories, val loss 1.916)
fine-tune data 29.4k synthetic conversations, 5.2M tokens, plus 15% TinyStories mixed into every batch
schedule 2,000 steps Γ— 32 Γ— 256, lr 2e-4, warmup 50, cosine; 6.4 h on a shared 6-core CPU
held-out loss 0.112 (step 2,000)

Held-out exactness eval β€” 140 probes, two phases per probe (dispatch β†’ tool_call must match expected name+arguments; follow-up after a canned tool result β†’ expected values must appear verbatim). eval_exact.jsonl/eval_summary.json have the raw per-probe records:

metric m7tool2 (v2) m7tool3 (this)
tool name correct 1.000 1.000
name + arguments exact 0.921 0.986
follow-up values verbatim 0.891 0.968
β€” calculator values 0.125 0.750
β€” weather values 0.562 0.938

End-to-end through anvil: a tool result of temp_c 18.5 is now quoted verbatim ("That is about 18.5.") β€” v2 dropped decimal points and hallucinated unseen magnitudes. Known residual: 9–10 digit results can still lose a digit; the tool set is fixed at fine-tune time.

Serving

anvil (serve.src + engine.src) loads m7tool3.bin directly:

ANVIL_TOOLS_INJECT=0 ./anvil-serve m7tool3.bin 8097
curl localhost:8097/v1/chat/completions -H 'Content-Type: application/json' -d '{
  "model": "mtlm",
  "messages": [
    {"role":"system","content":"You are a helpful assistant. You can call tools. When a tool is needed, reply with only the JSON tool call. Otherwise answer in plain English."},
    {"role":"user","content":"what is the weather in Lyon?"}],
  "tools": [{"type":"function","function":{"name":"get_weather","parameters":{"type":"object","properties":{"city":{"type":"string"}}}}}]}'

returns "tool_calls":[{"type":"function","function":{"name":"get_weather","arguments":"{\"city\": \"Lyon\"}"}}] with finish_reason: "tool_calls". Use the system prompt above: it is the one the model was trained with.

Using from Hugging Face (transformers)

The repo root also carries a full HF bundle β€” model.safetensors (fp32, LlamaForCausalLM), config.json, tokenizer.json, tokenizer_config.json β€” exported by tools/hf_export.py and verified bit-exact against the native checkpoint (max |Ξ”logit| = 1e-5 vs the MFL forward):

from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("javimosch/mtlm-7m-tools")
model = AutoModelForCausalLM.from_pretrained("javimosch/mtlm-7m-tools")
enc = tok.apply_chat_template(
    [{"role": "user", "content": "what is the weather in Paris?"}],
    add_generation_prompt=True, return_tensors="pt", return_dict=True)
out = model.generate(**enc, max_new_tokens=64, do_sample=False,
                     eos_token_id=2, pad_token_id=0)
print(tok.decode(out[0][enc["input_ids"].shape[1]:], skip_special_tokens=False))
# {"tool_call": {"name": "get_weather", "arguments": {"city": "Paris"}}}</s>

Chat-template notes: the template emits BOS itself (HF apply_chat_template skips the post-processor) and renders tool messages as Tool result: ... user turns β€” identical to anvil's serving format.

Files

  • m7tool3.mtlm β€” fp32 checkpoint (mtlm1 layout).
  • m7tool3.bin β€” llama2.c v2 int8 export, group size 32 (the model width, 288, is not a multiple of 64), 8.0 MB.
  • tokenizer.bin β€” llama2.c tokenizer format, 4096 pieces.
  • model.safetensors + config.json + tokenizer*.json β€” HF bundle (see above).
  • train_log.jsonl β€” the run log; eval_exact.jsonl / eval_summary.json β€” the exactness eval records.
Downloads last month
40
Safetensors
Model size
8.34M params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for javimosch/mtlm-7m-tools

Finetuned
(1)
this model