# TinyQuery Tools An experimental, randomly initialized decoder for schema-conditioned tool calls in English, imperfect English, Hindi and Hinglish. This is separate from the original 4.2M-parameter Tiny English learning project. The current student has **139,738,113 parameters**: 12 decoder layers, width 1024, 16 query heads, four KV heads, rotary positions, RMSNorm, SwiGLU width 2816, tied embeddings, a three-class auxiliary action head and a learned 128-dimensional source-copy head. The copy head mixes next-token generation with attention over the supplied context and question; it excludes generated answer text. The byte BPE vocabulary has 4,082 tokens learned from the training corpus. The configured context limit is 2,048; current training sequences reach 945 tokens. These are established transformer components with a custom training objective, not a demonstrated research breakthrough. ## Output contract The model receives a backend, schema, runtime tool definitions and a question, and predicts one JSON action: ```json {"action":"call","name":"execute_sql","arguments":{"query":"SELECT name FROM customers WHERE city = 'Delhi';"}} ``` Other actions are `{"action":"clarify","question":"..."}` and `{"action":"answer","text":"..."}`. Provided tool definitions use MCP-style `name`, `description`, and `inputSchema`. The adapter validates the action and can serialize a JSON-RPC `tools/call` request. **The neural model does not implement the MCP transport or authenticate to databases.** The database corpus targets read-only MySQL and PostgreSQL (Supabase). It covers projections, comparisons, NULLs, counts and aggregates, grouping/HAVING, sorting/limits, date filtering, substring matching and a single join. Schema discovery and bounded tool-error cases are included. Generic examples cover weather lookup, documentation search, file reading and ticket lookup. These are trained categories, not a claim of competence with arbitrary unseen tools. ## Dataset The final continuation corpus contains 347,376 examples and 100,224,712 tokens, including 13,163,249 response tokens. It combines programmatic SQL/tool semantics, Qwen3.8-27B-FP8 multilingual templates and direct paraphrases, and randomized runtime identifiers. It has 25,145 scenario groups; rendered language/context variants are not independent teacher generations. Validation has 1,200 examples and test has 1,196, with held-out domain/table names and language templates. An additional 160 cases render hand-written phrasings on 40 test scenario families. These add phrasing coverage, not independent schema families. Neither test nor manual is used for training or model selection. A separate 192-case development split changes schema layouts and tool names on 140 validation families; it is used during development and shares no training/test/manual families. A teacher audit of 1,313 training SQL template instances flagged 37 variants; 6,778 derived rows were conservatively removed from the final corpus. Earlier training stages had already used them. Every final training action passed schema/serialization checks, and 151,920 distinct reference SQL cases executed on two generated SQLite fixtures. A separate 66-case reference check passed on native MySQL 8.0.46 and PostgreSQL 16.15. A further 160,253 distinct query/context pairs compile against the supplied schemas. These are reference-data checks, not student accuracy. Teacher revision: `017b9c7af6b5689d5dd426a76e0bc077eb5ca20a`. MTP-3 improved the measured 32-concurrent-request teacher workload from about 691 to 931 output tokens/second. Samples retain teacher prompts, settings and verification provenance. The model is trained from random weights; no pretrained student checkpoint or hidden teacher fallback is used. ## Reproduce ```sh python -m tinyquery.teacher_templates --out data/tinyquery/templates.jsonl python -m tinyquery.data --templates data/tinyquery/templates.jsonl --out data/tinyquery python -m tinyquery.prepare --data data/tinyquery python -m tinyquery.train --data data/tinyquery --out runs/tinyquery --copy-dim 128 --minutes 30 ``` `concrete.py` and `verify_concrete.py` add direct paraphrases and round-trip verification. The published dataset includes the frozen augmented corpus; see its provenance and manifest rather than assuming the four commands above reproduce the exact published rows. An optional `--deadline` supplies a hard UTC wall-clock stop. The released experiment used multiple curriculum stages and a fixed four-hour overall budget; a new 30-minute run does not reproduce its result. ## Stream output The finished export is available in `runs/tinyquery/`: ```sh .venv/bin/python -m tinyquery.chat "दिल्ली के ग्राहकों के नाम दिखाओ।" \ --checkpoint runs/tinyquery/model.safetensors --backend supabase --mcp ``` Use `--context path/to/context.json` to provide actual schema and tool definitions. The CLI streams the model's raw output and validates it. `--stats` reports measured generation speed. The MCP serializer rejects mismatched project scope, unknown schema tables and unsupported data-changing statements; these structural checks do not establish semantic correctness. It does not execute tools. The default context is a small fictitious customers table. ## Evaluation ```sh python -m tinyquery.evaluate --checkpoint runs/tinyquery/model.safetensors \ --tokenizer runs/tinyquery/tokenizer.json --data data/tinyquery/test.jsonl \ --out runs/tinyquery/test-predictions.jsonl ``` Evaluation uses raw greedy output with no JSON repair, constrained decoding or teacher fallback. Report JSON validity, tool-schema validity, tool choice, exact arguments and SQL result equivalence separately. The default SQL evaluator uses dialect adaptation to SQLite; `native_sql.py` performs separate checks on local MySQL/PostgreSQL engines. The frozen 139,738,113-parameter model achieved 1,106/1,196 full test successes (92.47%) and 132/160 additional manual-phrasing successes (82.50%). Independent Mac FP32 evaluation exactly matched all aggregate CUDA BF16 metrics. The context-binding retrieval baseline achieved 73.66% and 71.88%, respectively. JSON validity was 100%; this is not 100% task success. Full raw Mac predictions and breakdowns are in `releases/tinyquery-model/evaluation/`. Published model: https://huggingface.co/karmx/TinyQuery-140M Published dataset: https://huggingface.co/datasets/karmx/TinyQuery-Tools-Multilingual Local runtime, template, dataset and checkpoint inventory: `artifacts/tinyquery/README.md`. Remote runtime configs, training logs and GPU predictions were recovered after a brief SSH outage. Native MySQL/PostgreSQL evaluation confirmed the same full-task results. The optional final optimizer-state transfer was stopped to avoid extending rental cost. The selected weights and complete datasets are local.