TrustSQL-4B / README.md
AIJian's picture
Update TrustSQL-4B model card
67bcbd5 verified
|
Raw History Blame Contribute Delete
3.59 kB
---
license: apache-2.0
library_name: transformers
pipeline_tag: text-generation
tags:
- text-to-sql
- sql
- unknown-schema
- tool-use
- reinforcement-learning
- qwen3
base_model: Qwen/Qwen3-4B
---
# TRUST-SQL-4B
[![arXiv](https://img.shields.io/badge/arXiv-2603.16448-b31b1b.svg)](https://arxiv.org/abs/2603.16448)
[![GitHub](https://img.shields.io/badge/GitHub-TrustSQL-black?logo=github)](https://github.com/JaneEyre0530/TrustSQL)
[![License](https://img.shields.io/badge/License-Apache%202.0-blue.svg)](LICENSE)
## Overview
**TrustSQL-4B** is a fine-tuned Text-to-SQL model based on [Qwen3-4B](https://huggingface.co/Qwen/Qwen3-4B), introduced in [TRUST-SQL: Tool-Integrated Multi-Turn Reinforcement Learning for Text-to-SQL over Unknown Schemas](https://arxiv.org/abs/2603.16448). The model is trained with multi-turn reinforcement learning and tool integration for Text-to-SQL over unknown database schemas.
## Model details
- Base model: `Qwen/Qwen3-4B`
- Architecture: `Qwen3ForCausalLM`
- Parameters: 4.0B
- Hidden size: 2560
- Layers: 36
- Attention heads: 32 Q heads / 8 KV heads
- Context length: 40,960 tokens
- Precision: bfloat16
## Models
| Model | Base | Link |
|---|---|---|
| TrustSQL-4B | Qwen3-4B | [AIJian/TrustSQL-4B](https://huggingface.co/AIJian/TrustSQL-4B) |
| TrustSQL-8B | Qwen3-8B | [AIJian/TrustSQL-8B](https://huggingface.co/AIJian/TrustSQL-8B) |
## Training
TrustSQL follows a two-stage training pipeline: SFT warm-up followed by Phase-Aware GRPO optimization. The interaction protocol is `Explore → Propose → Generate → Confirm`.
## Reported results
All results are reported under the Unknown Schema setting.
| Benchmark | Greedy | Majority voting |
|---|---:|---:|
| BIRD-Dev | 64.9 | 67.2 |
| Spider-Test | 82.8 | 85.0 |
| Spider-DK | 71.6 | 73.8 |
| Spider-Syn | 74.7 | 77.3 |
| Spider-Realistic | 79.9 | 82.5 |
## Recommended inference setup
This model is designed for an orchestrator that exposes:
1. A schema exploration tool for tables, columns, keys, and value inspection.
2. A schema proposal channel that records verified tables and columns.
3. A SQL execution tool for candidate queries.
4. A final answer channel for the confirmed SQL.
Do not provide fabricated schema descriptions as if they were tool observations. The model is intended to ground schema decisions in the environment feedback.
## Loading
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "AIJian/TrustSQL-4B"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype="auto",
device_map="auto",
trust_remote_code=True,
)
```
For prompts, tool schemas, evaluation scripts, and training details, see `https://github.com/JaneEyre0530/TrustSQL`.
## Limitations
This checkpoint was trained and evaluated with SQLite-based benchmarks. Its behavior depends on a live, correctly configured tool environment and a finite interaction budget. Validate generated SQL before using it in any sensitive or write-enabled database.
## Citation
```bibtex
@article{jian2026trustsql,
title = {TRUST-SQL: Tool-Integrated Multi-Turn Reinforcement Learning for Text-to-SQL over Unknown Schemas},
author = {Jian, Ai and Zhang, Xiaoyun and Du, Wanrou and Ruan, Jingqing and Pei, Jiangbo and Zhang, Weipeng and Zeng, Ke and Cai, Xunliang},
journal = {arXiv preprint arXiv:2603.16448},
year = {2026}
}
```
## License
This project is licensed under the Apache 2.0 License. See the `LICENSE` file for details.