File size: 8,714 Bytes
bd5e611
f019743
bd5e611
 
 
 
f019743
bd5e611
 
 
f019743
 
 
 
bd5e611
 
f019743
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
bd5e611
f019743
 
 
 
 
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
 
 
 
 
 
 
 
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
 
 
 
 
 
 
 
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
 
 
bd5e611
f019743
 
bd5e611
f019743
bd5e611
f019743
 
 
 
 
bd5e611
f019743
 
 
 
bd5e611
f019743
 
bd5e611
f019743
bd5e611
f019743
 
bd5e611
f019743
 
bd5e611
f019743
 
bd5e611
f019743
bd5e611
f019743
 
 
 
 
 
 
bd5e611
f019743
 
 
 
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
 
 
 
 
bd5e611
f019743
bd5e611
f019743
 
 
 
 
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
 
 
 
 
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
 
 
 
 
 
 
 
 
 
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
 
 
 
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
 
 
 
 
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
 
 
 
 
 
 
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
 
 
 
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
 
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
bd5e611
f019743
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
---

base_model: codellama/CodeLlama-7b-Instruct-hf
library_name: peft
pipeline_tag: text-generation
tags:

- base_model:adapter:codellama/CodeLlama-7b-Instruct-hf
- lora
- transformers
- sql
- text-to-sql
- code-generation

---

Model Card for my-sql-model

Model Details

Model Description

"my-sql-model" is a SQL query generation model based on "CodeLlama-7b-Instruct-hf".

The model uses LoRA (Low-Rank Adaptation) through the PEFT (Parameter-Efficient Fine-Tuning) framework to adapt the base CodeLlama model for SQL generation tasks.

The model is intended to convert natural-language questions and database schema information into SQL queries.

- Developed by: tamilanda
- Model type: CodeLlama 7B Instruct with a LoRA adapter
- Language(s): English
- License: Refer to the base model license and repository license
- Base model: "codellama/CodeLlama-7b-Instruct-hf"
- Fine-tuning method: LoRA / PEFT
- Framework: Hugging Face Transformers and PEFT

Model Sources

- Repository: "tamilanda/my-sql-model"
- Base Model: "codellama/CodeLlama-7b-Instruct-hf"

Uses

Direct Use

The model can be used for natural-language-to-SQL query generation.

A typical input consists of:

1. Database type
2. Database schema
3. Natural-language question

The model generates an SQL query corresponding to the requested operation.

Example:

Database: MySQL

Schema:
employees(id, name, department, salary)

Question:
Find employees whose salary is greater than 50000.

Expected output:

SELECT *
FROM employees
WHERE salary > 50000;

Downstream Use

The model can be integrated into:

- Natural-language database assistants
- Text-to-SQL applications
- RAG-based database systems
- Database analytics assistants
- Conversational SQL systems
- Automated SQL query generation pipelines

A production architecture can combine the model with schema retrieval and query validation:

User Question
      ↓
Database Detection
      ↓
Schema Retrieval
      ↓
Relevant Tables
      ↓
my-sql-model
      ↓
SQL Generation
      ↓
SQL Validation
      ↓
Database Execution

Out-of-Scope Use

The model should not be used as an unrestricted database execution system.

Generated SQL should not be executed directly against production databases without validation and appropriate permissions.

The model is not intended for:

- Unauthorized database access
- Bypassing database permissions
- Destructive database operations without validation
- Automatic execution of untrusted SQL
- Security-sensitive database administration without human oversight

Bias, Risks, and Limitations

The model is a generative language model and may generate incorrect or syntactically invalid SQL.

Potential limitations include:

- Incorrect table or column selection
- Incorrect joins
- Incorrect filtering conditions
- Hallucinated columns or tables
- SQL dialect incompatibility
- Incorrect interpretation of ambiguous questions
- Incorrect aggregation or grouping
- Poor performance when the provided schema is incomplete

Generated queries should therefore be validated before execution.

Recommendations

For production applications:

- Provide the relevant database schema to the model.
- Clearly specify the database dialect.
- Validate generated SQL before execution.
- Use read-only database credentials where possible.
- Restrict database permissions.
- Apply query timeout and resource limits.
- Log generated queries and execution results.
- Require human approval for destructive operations.

How to Get Started with the Model

Install the required libraries:

pip install transformers peft torch

Load the base model and LoRA adapter:

from transformers import AutoTokenizer, AutoModelForCausalLM
from peft import PeftModel
import torch

base_model = "codellama/CodeLlama-7b-Instruct-hf"
adapter_model = "tamilanda/my-sql-model"

tokenizer = AutoTokenizer.from_pretrained(base_model)

model = AutoModelForCausalLM.from_pretrained(
    base_model,
    torch_dtype=torch.float16,
    device_map="auto"
)

model = PeftModel.from_pretrained(
    model,
    adapter_model
)

prompt = """
Generate a SQL query.

Database: MySQL

Schema:
employees(id, name, department, salary)

Question:
Find employees whose salary is greater than 50000.

Return only the SQL query.
"""

inputs = tokenizer(prompt, return_tensors="pt").to(model.device)

with torch.no_grad():
    outputs = model.generate(
        **inputs,
        max_new_tokens=150,
        temperature=0.1,
        do_sample=False
    )

result = tokenizer.decode(
    outputs[0],
    skip_special_tokens=True
)

print(result)

Training Details

Training Data

The model is intended for SQL query generation. The exact training dataset and dataset size should be documented separately if available.

A suitable training example contains:

Natural Language Question
+
Database Schema
+
Expected SQL Query

Example:

{
  "question": "Find all customers from Chennai",
  "schema": "customers(id, name, city)",
  "query": "SELECT * FROM customers WHERE city = 'Chennai';"
}

Training Procedure

The model uses Parameter-Efficient Fine-Tuning (PEFT) with LoRA on the CodeLlama-7B-Instruct base model.

Preprocessing

Training examples should be formatted as instruction-following examples containing the database schema, user question, and target SQL query.

Training Hyperparameters

- Training regime: Not specified
- Fine-tuning method: LoRA
- Framework: PEFT
- PEFT version: 0.17.1
- Base model: "codellama/CodeLlama-7b-Instruct-hf"

Speeds, Sizes, Times

Training hardware, training duration, throughput, checkpoint size, and compute requirements are not specified.

Evaluation

Testing Data, Factors & Metrics

Testing Data

The evaluation dataset is not specified.

Factors

Evaluation can be performed across:

- Simple SQL queries
- Filtering
- Aggregation
- GROUP BY
- ORDER BY
- JOIN operations
- Subqueries
- Nested queries
- Multiple-table queries
- Complex analytical queries

Metrics

Recommended metrics include:

- Exact Match Accuracy
- Execution Accuracy
- SQL Syntax Validity
- Query Execution Success Rate

No evaluation scores are claimed here because verified results were not provided.

Results

Evaluation results are not currently specified.

Summary

The model should be evaluated against a held-out SQL dataset before production deployment.

Model Examination

The model can be examined by testing generated SQL against known database schemas and comparing generated queries with reference queries and execution results.

Environmental Impact

Carbon emissions depend on the hardware and infrastructure used during fine-tuning.

- Hardware Type: Not specified
- Hours used: Not specified
- Cloud Provider: Not specified
- Compute Region: Not specified
- Carbon Emitted: Not specified

Carbon emissions can be estimated using the "Machine Learning Impact calculator" (https://mlco2.github.io/impact#compute).

Technical Specifications

Model Architecture and Objective

The model uses:

CodeLlama-7B-Instruct
          ↓
       LoRA
          ↓
   PEFT Adapter
          ↓
   my-sql-model

The objective is to adapt the base code-generation model for SQL query generation.

Compute Infrastructure

The exact training infrastructure is not specified.

Hardware

Not specified.

Software

The model uses the Hugging Face ecosystem, including:

- Transformers
- PEFT
- LoRA
- PyTorch

PEFT version: "0.17.1"

Citation

If this model is used in a project, cite the model repository and the underlying CodeLlama model.

Base Model

CodeLlama: Open Foundation Models for Code.
Meta AI.

Glossary

PEFT: Parameter-Efficient Fine-Tuning, a method for adapting large models while training a relatively small number of parameters.

LoRA: Low-Rank Adaptation, a PEFT technique that trains low-rank adapter matrices instead of updating all base-model parameters.

Text-to-SQL: Conversion of a natural-language question into an SQL query.

Schema: The structure of a database, including tables, columns, relationships, and data types.

Execution Accuracy: Measures whether the generated SQL produces the correct result when executed against the target database.

More Information

The model can be extended for database assistants by combining SQL generation with schema retrieval, RAG, SQL validation, and controlled database execution.

For dynamic database environments, schema information should be retrieved at inference time rather than relying only on information learned during fine-tuning.

Model Card Authors

- Author: tamilanda

Model Card Contact

For questions, issues, or contributions, please use the model repository's issue/discussion section.

Framework Versions

- PEFT: "0.17.1"
- Transformers: Hugging Face Transformers
- PyTorch: PyTorch