File size: 6,326 Bytes
30f6980
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
---
license: apache-2.0
tags:
  - text2sql
  - sql
  - unknown-schema
  - tool-use
  - reinforcement-learning
  - arxiv:2603.16448
---

# TrustSQL

**Tool-Integrated Multi-Turn Reinforcement Learning for Text-to-SQL over Unknown Schemas**

[![arXiv](https://img.shields.io/badge/arXiv-2603.16448-b31b1b.svg)](https://arxiv.org/abs/2603.16448)
[![GitHub](https://img.shields.io/badge/GitHub-TrustSQL-black?logo=github)](https://github.com/JaneEyre0530/TrustSQL)
[![License](https://img.shields.io/badge/License-Apache%202.0-blue.svg)](https://www.apache.org/licenses/LICENSE-2.0)

TrustSQL (**T**ruthful **R**easoning with **U**nknown **S**chema via **T**ools) is an agentic Text-to-SQL framework for settings where the complete database schema is not pre-loaded into the model context. Instead of passively translating a question from a full schema dump, the agent actively explores the database, verifies relevant metadata, generates candidate SQL, and confirms the final answer.

## Highlights

- **Unknown Schema reasoning:** retrieve only the metadata needed for the current question.
- **Four-phase interaction protocol:** `Explore → Propose → Generate → Confirm`, with iterative refinement when more information is needed.
- **Phase-Aware GRPO:** separate schema-grounding and SQL-generation learning signals with token-level masked advantages.
- **Execution-coupled schema reward:** give schema credit when the proposed schema is correct and the resulting SQL executes correctly.
- **Multiple scales:** released checkpoints include 4B and 8B variants; the research results also report a 14B variant.

## Collection

| Type | Repository | Description |
|---|---|---|
| Model | [`AIJian/TrustSQL-4B`](https://huggingface.co/AIJian/TrustSQL-4B) | 4B TrustSQL checkpoint based on Qwen3-4B |
| Model | [`AIJian/TrustSQL-8B`](https://huggingface.co/AIJian/TrustSQL-8B) | 8B TrustSQL checkpoint based on Qwen3-8B |
| Dataset | [`AIJian/TrustSQL-data`](https://huggingface.co/datasets/AIJian/TrustSQL-data) | Training data for TrustSQL |
| Overview | [`AIJian/TrustSQL`](https://huggingface.co/AIJian/TrustSQL) | Project overview, results, usage, citation, and links |

## Results

All results below use the **Unknown Schema** setting, without pre-loaded database metadata. Execution Accuracy (EX) measures whether the generated SQL returns the same result as the reference query.

### BIRD-Dev

| Model | Greedy | Majority voting |
|---|---:|---:|
| TRUST-SQL-4B | 64.9 | 67.2 |
| TRUST-SQL-8B | 65.8 | 67.7 |
| TRUST-SQL-14B | 68.1 | 70.1 |

### Robustness and compositional benchmarks

| Model | Spider-Test | Spider-DK | Spider-Syn | Spider-Realistic |
|---|---:|---:|---:|---:|
| TRUST-SQL-4B | 82.8 / 85.0 | 71.6 / 73.8 | 74.7 / 77.3 | 79.9 / 82.5 |
| TRUST-SQL-8B | 83.9 / 86.5 | 72.1 / 75.7 | 75.4 / 77.4 | 82.1 / 84.1 |
| TRUST-SQL-14B | 83.9 / 86.5 | 75.0 / 78.6 | 78.4 / 79.7 | 82.7 / 86.5 |

Each entry is `Greedy / Majority voting`.

On BIRD-Dev, the framework improves over the corresponding base models by an average of **30.6 absolute points** for the 4B model and **16.6 absolute points** for the 8B model across the five evaluation benchmarks.

## How it works

### 1. Explore Schema

Use metadata tools to discover tables, columns, foreign keys, and relevant value patterns. Schema knowledge must come from verified tool observations rather than model memory.

### 2. Propose Schema

Commit to the currently verified tables and columns. This explicit checkpoint reduces schema hallucination and defines the boundary for schema-specific credit assignment.

### 3. Generate SQL

Generate and execute a candidate SQL query using the proposed schema. The execution result provides feedback for query refinement.

### 4. Confirm Answer

Return the final SQL answer after the schema and query have been validated. The workflow can revisit earlier phases when additional evidence is needed.

## Training

The TrustSQL training recipe uses a two-stage pipeline:

1. **SFT warm-up:** 9.2k structured demonstrations teach the interaction protocol.
2. **Phase-Aware GRPO:** 11.6k RL samples optimize schema exploration and SQL generation with separate masked advantages.

The method is implemented with the SLIME training framework. The experiments use Qwen3-4B, Qwen3-8B, and Qwen3-14B as base models.

## Usage

The checkpoints are intended for **tool-integrated multi-turn inference**, not ordinary one-shot text completion. At inference time, the host application should expose database metadata and SQL execution tools, then enforce the four-phase action protocol.

Minimal loading example:

```python
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "AIJian/TrustSQL-8B"

tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype="auto",
    device_map="auto",
    trust_remote_code=True,
)
```

For the complete action schemas, prompts, training configuration, and evaluation pipeline, see the project repository:

`https://github.com/JaneEyre0530/TrustSQL`

## Important limitations

- Training and reported evaluations use SQLite-based benchmarks.
- Inference requires access to the target database or a compatible metadata/execution environment.
- The current implementation uses a fixed interaction budget; highly complex databases may benefit from adaptive horizons.
- The reported benchmark numbers are research results and should not be interpreted as a guarantee for production workloads.

## Citation

```bibtex
@article{jian2026trustsql,
  title   = {TRUST-SQL: Tool-Integrated Multi-Turn Reinforcement Learning for Text-to-SQL over Unknown Schemas},
  author  = {Jian, Ai and Zhang, Xiaoyun and Du, Wanrou and Ruan, Jingqing and Pei, Jiangbo and Zhang, Weipeng and Zeng, Ke and Cai, Xunliang},
  journal = {arXiv preprint arXiv:2603.16448},
  year    = {2026},
  url     = {https://arxiv.org/abs/2603.16448}
}
```

## License

The released model and dataset repositories are marked as Apache-2.0. The dataset may contain material derived from third-party benchmarks; users are responsible for complying with the original benchmark licenses and usage terms.

## Contact

For questions or feedback, please open an issue in the project repository or contact `jianai@bupt.edu.cn`.