mtlm-7m-tools β a 7M-parameter assistant that talks English and calls tools, trained end-to-end in machin
Fine-tune of javimosch/mtlm-7m-base that decides between answering in
plain English and emitting a one-line JSON tool call, over 14 generic tools (weather, calculator, web search, time, read/write
file, HTTP GET, email, translate, shell, unit conversion, reminders, notes, wikipedia). Served by
anvil, it speaks the OpenAI chat-completions protocol including tool_calls.
Everything except the synthesis of the fine-tuning text is pure machin (MFL): the base pretraining, the chat-template tokenizer, the fine-tune, the int8 export and the server. The text corpus was produced by a templated Python script from the tool catalog (no LLM teacher).
v3 (m7tool3) β this revision fixes the number-misquoting flaw of v2 (m7tool2, kept in git history): the training corpus now serializes assistant tool calls byte-identically to how anvil re-renders them in context, tool results cover floats/negatives/6-digit ids, and a fixed unit-conversion bug no longer teaches wrong values.
Honest numbers
| base | mtlm-7m-base (7.2M params, TinyStories, val loss 1.916) |
| fine-tune data | 29.4k synthetic conversations, 5.2M tokens, plus 15% TinyStories mixed into every batch |
| schedule | 2,000 steps Γ 32 Γ 256, lr 2e-4, warmup 50, cosine; 6.4 h on a shared 6-core CPU |
| held-out loss | 0.112 (step 2,000) |
Held-out exactness eval β 140 probes, two phases per probe (dispatch β tool_call must match expected name+arguments;
follow-up after a canned tool result β expected values must appear verbatim). eval_exact.jsonl/eval_summary.json
have the raw per-probe records:
| metric | m7tool2 (v2) | m7tool3 (this) |
|---|---|---|
| tool name correct | 1.000 | 1.000 |
| name + arguments exact | 0.921 | 0.986 |
| follow-up values verbatim | 0.891 | 0.968 |
| β calculator values | 0.125 | 0.750 |
| β weather values | 0.562 | 0.938 |
End-to-end through anvil: a tool result of temp_c 18.5 is now quoted verbatim ("That is about 18.5.") β v2 dropped
decimal points and hallucinated unseen magnitudes. Known residual: 9β10 digit results can still lose a digit; the tool
set is fixed at fine-tune time.
Serving
anvil (serve.src + engine.src) loads m7tool3.bin directly:
ANVIL_TOOLS_INJECT=0 ./anvil-serve m7tool3.bin 8097
curl localhost:8097/v1/chat/completions -H 'Content-Type: application/json' -d '{
"model": "mtlm",
"messages": [
{"role":"system","content":"You are a helpful assistant. You can call tools. When a tool is needed, reply with only the JSON tool call. Otherwise answer in plain English."},
{"role":"user","content":"what is the weather in Lyon?"}],
"tools": [{"type":"function","function":{"name":"get_weather","parameters":{"type":"object","properties":{"city":{"type":"string"}}}}}]}'
returns "tool_calls":[{"type":"function","function":{"name":"get_weather","arguments":"{\"city\": \"Lyon\"}"}}] with
finish_reason: "tool_calls". Use the system prompt above: it is the one the model was trained with.
Using from Hugging Face (transformers)
The repo root also carries a full HF bundle β model.safetensors (fp32, LlamaForCausalLM),
config.json, tokenizer.json, tokenizer_config.json β exported by tools/hf_export.py and
verified bit-exact against the native checkpoint (max |Ξlogit| = 1e-5 vs the MFL forward):
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("javimosch/mtlm-7m-tools")
model = AutoModelForCausalLM.from_pretrained("javimosch/mtlm-7m-tools")
enc = tok.apply_chat_template(
[{"role": "user", "content": "what is the weather in Paris?"}],
add_generation_prompt=True, return_tensors="pt", return_dict=True)
out = model.generate(**enc, max_new_tokens=64, do_sample=False,
eos_token_id=2, pad_token_id=0)
print(tok.decode(out[0][enc["input_ids"].shape[1]:], skip_special_tokens=False))
# {"tool_call": {"name": "get_weather", "arguments": {"city": "Paris"}}}</s>
Chat-template notes: the template emits BOS itself (HF apply_chat_template skips the
post-processor) and renders tool messages as Tool result: ... user turns β identical to
anvil's serving format.
Files
m7tool3.mtlmβ fp32 checkpoint (mtlm1layout).m7tool3.binβ llama2.c v2 int8 export, group size 32 (the model width, 288, is not a multiple of 64), 8.0 MB.tokenizer.binβ llama2.c tokenizer format, 4096 pieces.model.safetensors+config.json+tokenizer*.jsonβ HF bundle (see above).train_log.jsonlβ the run log;eval_exact.jsonl/eval_summary.jsonβ the exactness eval records.
- Downloads last month
- 40
Model tree for javimosch/mtlm-7m-tools
Base model
javimosch/mtlm-7m-base