TinyQuery-140M / docs /experiment.md
karmx's picture
Add recovered GPU predictions, native SQL validation and complete runtime/training provenance
dfb27f2 verified
|
Raw History Blame Contribute Delete
6.91 kB

TinyQuery Tools

An experimental, randomly initialized decoder for schema-conditioned tool calls in English, imperfect English, Hindi and Hinglish. This is separate from the original 4.2M-parameter Tiny English learning project.

The current student has 139,738,113 parameters: 12 decoder layers, width 1024, 16 query heads, four KV heads, rotary positions, RMSNorm, SwiGLU width 2816, tied embeddings, a three-class auxiliary action head and a learned 128-dimensional source-copy head. The copy head mixes next-token generation with attention over the supplied context and question; it excludes generated answer text. The byte BPE vocabulary has 4,082 tokens learned from the training corpus. The configured context limit is 2,048; current training sequences reach 945 tokens. These are established transformer components with a custom training objective, not a demonstrated research breakthrough.

Output contract

The model receives a backend, schema, runtime tool definitions and a question, and predicts one JSON action:

{"action":"call","name":"execute_sql","arguments":{"query":"SELECT name FROM customers WHERE city = 'Delhi';"}}

Other actions are {"action":"clarify","question":"..."} and {"action":"answer","text":"..."}. Provided tool definitions use MCP-style name, description, and inputSchema. The adapter validates the action and can serialize a JSON-RPC tools/call request. The neural model does not implement the MCP transport or authenticate to databases.

The database corpus targets read-only MySQL and PostgreSQL (Supabase). It covers projections, comparisons, NULLs, counts and aggregates, grouping/HAVING, sorting/limits, date filtering, substring matching and a single join. Schema discovery and bounded tool-error cases are included. Generic examples cover weather lookup, documentation search, file reading and ticket lookup. These are trained categories, not a claim of competence with arbitrary unseen tools.

Dataset

The final continuation corpus contains 347,376 examples and 100,224,712 tokens, including 13,163,249 response tokens. It combines programmatic SQL/tool semantics, Qwen3.8-27B-FP8 multilingual templates and direct paraphrases, and randomized runtime identifiers. It has 25,145 scenario groups; rendered language/context variants are not independent teacher generations.

Validation has 1,200 examples and test has 1,196, with held-out domain/table names and language templates. An additional 160 cases render hand-written phrasings on 40 test scenario families. These add phrasing coverage, not independent schema families. Neither test nor manual is used for training or model selection. A separate 192-case development split changes schema layouts and tool names on 140 validation families; it is used during development and shares no training/test/manual families.

A teacher audit of 1,313 training SQL template instances flagged 37 variants; 6,778 derived rows were conservatively removed from the final corpus. Earlier training stages had already used them. Every final training action passed schema/serialization checks, and 151,920 distinct reference SQL cases executed on two generated SQLite fixtures. A separate 66-case reference check passed on native MySQL 8.0.46 and PostgreSQL 16.15. A further 160,253 distinct query/context pairs compile against the supplied schemas. These are reference-data checks, not student accuracy.

Teacher revision: 017b9c7af6b5689d5dd426a76e0bc077eb5ca20a. MTP-3 improved the measured 32-concurrent-request teacher workload from about 691 to 931 output tokens/second. Samples retain teacher prompts, settings and verification provenance. The model is trained from random weights; no pretrained student checkpoint or hidden teacher fallback is used.

Reproduce

python -m tinyquery.teacher_templates --out data/tinyquery/templates.jsonl
python -m tinyquery.data --templates data/tinyquery/templates.jsonl --out data/tinyquery
python -m tinyquery.prepare --data data/tinyquery
python -m tinyquery.train --data data/tinyquery --out runs/tinyquery --copy-dim 128 --minutes 30

concrete.py and verify_concrete.py add direct paraphrases and round-trip verification. The published dataset includes the frozen augmented corpus; see its provenance and manifest rather than assuming the four commands above reproduce the exact published rows. An optional --deadline supplies a hard UTC wall-clock stop. The released experiment used multiple curriculum stages and a fixed four-hour overall budget; a new 30-minute run does not reproduce its result.

Stream output

The finished export is available in runs/tinyquery/:

.venv/bin/python -m tinyquery.chat "दिल्ली के ग्राहकों के नाम दिखाओ।" \
  --checkpoint runs/tinyquery/model.safetensors --backend supabase --mcp

Use --context path/to/context.json to provide actual schema and tool definitions. The CLI streams the model's raw output and validates it. --stats reports measured generation speed. The MCP serializer rejects mismatched project scope, unknown schema tables and unsupported data-changing statements; these structural checks do not establish semantic correctness. It does not execute tools. The default context is a small fictitious customers table.

Evaluation

python -m tinyquery.evaluate --checkpoint runs/tinyquery/model.safetensors \
  --tokenizer runs/tinyquery/tokenizer.json --data data/tinyquery/test.jsonl \
  --out runs/tinyquery/test-predictions.jsonl

Evaluation uses raw greedy output with no JSON repair, constrained decoding or teacher fallback. Report JSON validity, tool-schema validity, tool choice, exact arguments and SQL result equivalence separately. The default SQL evaluator uses dialect adaptation to SQLite; native_sql.py performs separate checks on local MySQL/PostgreSQL engines. The frozen 139,738,113-parameter model achieved 1,106/1,196 full test successes (92.47%) and 132/160 additional manual-phrasing successes (82.50%). Independent Mac FP32 evaluation exactly matched all aggregate CUDA BF16 metrics. The context-binding retrieval baseline achieved 73.66% and 71.88%, respectively. JSON validity was 100%; this is not 100% task success. Full raw Mac predictions and breakdowns are in releases/tinyquery-model/evaluation/.

Published model: https://huggingface.co/karmx/TinyQuery-140M

Published dataset: https://huggingface.co/datasets/karmx/TinyQuery-Tools-Multilingual

Local runtime, template, dataset and checkpoint inventory: artifacts/tinyquery/README.md. Remote runtime configs, training logs and GPU predictions were recovered after a brief SSH outage. Native MySQL/PostgreSQL evaluation confirmed the same full-task results. The optional final optimizer-state transfer was stopped to avoid extending rental cost. The selected weights and complete datasets are local.