multimodalart's picture
multimodalart HF Staff
TRUST-SQL agent demo: four-phase tool-integrated text-to-SQL over unknown schemas
c48f4e6 verified
|
Raw History Blame Contribute Delete
3.16 kB
---
title: TRUST-SQL
emoji: πŸ”Ž
colorFrom: gray
colorTo: indigo
sdk: gradio
sdk_version: 6.27.0
app_file: app.py
python_version: "3.12"
startup_duration_timeout: 1h
pinned: false
license: apache-2.0
short_description: Text-to-SQL over unknown schemas, via tool use
models:
- AIJian/TrustSQL-8B
---
# TRUST-SQL β€” Text-to-SQL over *unknown* schemas
Demo of [`AIJian/TrustSQL-8B`](https://huggingface.co/AIJian/TrustSQL-8B), from
[**TRUST-SQL: Tool-Integrated Multi-Turn Reinforcement Learning for Text-to-SQL over Unknown Schemas**](https://huggingface.co/papers/2603.16448)
(Jian et al., 2026). Code: [`JaneEyre0530/TrustSQL`](https://github.com/JaneEyre0530/TrustSQL).
Unlike ordinary text-to-SQL demos, **the schema is never put in the prompt**. The model is given
only the database *name*, the question, and a single read-only SQL tool, and has to discover the
schema itself by following the authors' four-phase action protocol:
1. `explore_schema` β€” issue metadata queries (`PRAGMA`, `sqlite_master`, sampling rows)
2. `propose_schema` β€” write down the tables/columns/joins it has actually verified
3. `generate_sql` β€” draft the answer query and *execute* it to check it works
4. `confirm_answer` β€” emit the final SQL
The agent may loop back at any point. The full trajectory (reasoning, tool calls, observations) is
streamed into the transcript so you can watch the schema being discovered.
## Implementation notes
- The system prompt is `trustsql_eval/prompt_template.txt`, copied verbatim from the authors' repo.
- The user message reproduces the exact `**Task Configuration** / **Database Engine** /
**Database** / **External Knowledge** / **User Question**` format found in
[`AIJian/TrustSQL-data`](https://huggingface.co/datasets/AIJian/TrustSQL-data).
- The turn loop, progress prefixes, `<schema>` acknowledgement, malformed-output feedback and
observation truncation (2048 tokens) are ported from `trustsql_eval/message_processor.py`.
- Sampling defaults match `trustsql_eval/main.py` (temperature 0.7, top-p 0.9).
- Deviation: `WITH` is added to the `SELECT` / `PRAGMA` / `EXPLAIN` allow-list so CTE answers can be
executed. Every query still runs over a read-only SQLite connection.
- Runs on ZeroGPU with plain `transformers` generation rather than the vLLM path used for the
paper's benchmarks, so it is slower than the reported latency figures.
## Credits & licensing
- Model and code: Apache-2.0, Β© the TRUST-SQL authors (Meituan / BUPT).
- Bundled sample databases and the example questions + `evidence` hints are from the
**BIRD** dev set ([bird-bench.github.io](https://bird-bench.github.io/)), licensed
**CC BY-SA 4.0**; SQLite files mirrored via
[`prem-research/birdbench`](https://huggingface.co/datasets/prem-research/birdbench).
```bibtex
@article{jian2026trustsql,
title = {TRUST-SQL: Tool-Integrated Multi-Turn Reinforcement Learning for Text-to-SQL over Unknown Schemas},
author = {Jian, Ai and Zhang, Xiaoyun and Du, Wanrou and Ruan, Jingqing and Pei, Jiangbo and Zhang, Weipeng and Zeng, Ke and Cai, Xunliang},
journal= {arXiv preprint arXiv:2603.16448},
year = {2026}
}
```