Configuration Parsing Warning:Config file config.json cannot be fetched (too big)

Configuration Parsing Warning:Config file tokenizer_config.json cannot be fetched (too big)

SQL Agent Vinod - DeepSeek R1 Distill Fine-tuned

This model is a fine-tuned version of unsloth/deepseek-r1-distill-qwen-7b specialized for SQL analytics and business conversation tasks. It has been optimized for generating SQL queries and providing business insights.

Model Details

Model Description

This model combines the power of DeepSeek R1 Distill architecture with specialized fine-tuning for SQL analytics and business intelligence tasks. It excels at understanding natural language queries and converting them into SQL statements, as well as providing analytical insights about business data.

  • Developed by: Abhishek Gahlot
  • Model type: Causal Language Model (Fine-tuned)
  • Language(s) (NLP): English
  • License: Apache 2.0
  • Finetuned from model: unsloth/deepseek-r1-distill-qwen-7b

Model Sources

Uses

Direct Use

This model is designed for:

  • SQL Query Generation: Convert natural language questions into SQL queries
  • Business Analytics: Provide insights on customer data, revenue analysis, and business metrics
  • Data Interpretation: Help interpret and explain business data patterns
  • Customer Analysis: Generate queries and insights for customer segmentation and analysis

Downstream Use

The model can be integrated into:

  • Business intelligence dashboards
  • SQL query builders
  • Customer analytics platforms
  • Data exploration tools
  • Business reporting systems

Out-of-Scope Use

This model is not suitable for:

  • General-purpose conversation (not fine-tuned for this)
  • Code generation outside of SQL context
  • Medical, legal, or financial advice
  • Tasks requiring real-time data access

Bias, Risks, and Limitations

  • Training Data Bias: Model is trained on specific SQL analytics patterns and may not generalize to all SQL dialects
  • Domain Specificity: Optimized for business analytics, may not perform well on other SQL use cases
  • Data Privacy: Should not be used to generate queries on sensitive or personal data without proper safeguards
  • Hallucination: May generate plausible-sounding but incorrect SQL syntax or business insights

Recommendations

Users should:

  • Validate generated SQL queries before execution
  • Test the model on their specific SQL dialect and schema
  • Implement proper data access controls and privacy measures
  • Review business insights for accuracy against actual data

How to Get Started with the Model

Quick Start with Unsloth (Recommended)

from unsloth import FastLanguageModel

# Load the model
model, tokenizer = FastLanguageModel.from_pretrained(
    model_name="abhishekgahlot/sql-agent-vinod",
    max_seq_length=3072,
    dtype=None,
    load_in_4bit=True,
)
FastLanguageModel.for_inference(model)

# Generate SQL query
prompt = "User: List the top 5 customers by revenue\n\nAssistant:"
inputs = tokenizer(prompt, return_tensors="pt")
outputs = model.generate(**inputs, max_length=512, temperature=0.7)
response = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(response)

Standard Transformers

from transformers import AutoTokenizer, AutoModelForCausalLM

tokenizer = AutoTokenizer.from_pretrained("abhishekgahlot/sql-agent-vinod")
model = AutoModelForCausalLM.from_pretrained(
    "abhishekgahlot/sql-agent-vinod",
    torch_dtype="auto",
    device_map="auto"
)

# Generate response
inputs = tokenizer("Show me monthly sales trends", return_tensors="pt")
outputs = model.generate(**inputs, max_length=512, do_sample=True, temperature=0.7)
response = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(response)

Training Details

Training Data

The model was fine-tuned on a custom dataset containing:

  • 13,570 training samples of SQL analytics and conversation data
  • 408 validation samples for evaluation
  • Mixed dataset combining conversation patterns and SQL query examples
  • Focus on business analytics, customer data, and revenue analysis patterns

Training Procedure

Training Hyperparameters

  • Training regime: bfloat16 mixed precision
  • Learning rate: 0.00018
  • Batch size per device: 10
  • Gradient accumulation steps: 8
  • Effective batch size: 80
  • Number of epochs: 3
  • Max sequence length: 3072
  • Warmup ratio: 0.05
  • Weight decay: 0.01
  • Optimizer: AdamW 8-bit
  • LR scheduler: Cosine annealing

LoRA Configuration

  • LoRA rank (r): 18
  • LoRA alpha: 36
  • LoRA dropout: 0.05
  • Target modules: q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj

Speeds, Sizes, Times

  • Training time: ~10 hours
  • Hardware: NVIDIA A10G (24GB VRAM)
  • Framework: Unsloth 2025.7.5 + PyTorch 2.7.0
  • Total steps: 495
  • Model size: ~7.6B parameters (0.59% trainable with LoRA)
  • Speed improvement: ~1.6x faster than baseline training

Evaluation

Testing Data, Factors & Metrics

Testing Data

  • Validation set: 408 samples held out from training data
  • Test scenarios: SQL query generation, business analytics questions, customer analysis tasks

Factors

  • Query complexity: Simple SELECT to complex JOINs and aggregations
  • Business domains: Customer analysis, revenue tracking, sales analytics
  • Query types: Data retrieval, aggregations, filtering, sorting

Metrics

Model performance evaluated on:

  • Perplexity: Measure of prediction quality
  • SQL syntax accuracy: Correctness of generated SQL statements
  • Business relevance: Appropriateness of responses to business questions

Results

  • Training completed successfully with consistent loss reduction across epochs
  • Validation performance showed good generalization to unseen data
  • Inference speed: ~2x faster with Unsloth optimization
  • Memory efficiency: Runs efficiently on 8GB+ VRAM with 4-bit quantization

Environmental Impact

Training was conducted with efficiency in mind using optimized frameworks and quantization.

  • Hardware Type: NVIDIA A10G GPU
  • Hours used: ~10 hours
  • Cloud Provider: AWS
  • Compute Region: US-East
  • Framework optimizations: Unsloth + LoRA for reduced compute requirements

Technical Specifications

Model Architecture and Objective

  • Base Architecture: DeepSeek R1 Distill (Qwen-based transformer)
  • Parameters: 7.6B total, ~45M trainable (LoRA)
  • Objective: Causal language modeling with supervised fine-tuning
  • Quantization: 4-bit quantization with bfloat16 training precision

Compute Infrastructure

Hardware

  • GPU: NVIDIA A10G with 24GB VRAM
  • Memory usage: ~21GB during training, ~8GB for inference
  • Optimization: Gradient checkpointing, mixed precision training

Software

  • Framework: Unsloth 2025.7.5
  • PyTorch: 2.7.0+cu126
  • Transformers: 4.53.2
  • Training: SFTTrainer with LoRA adaptation
  • Monitoring: Weights & Biases integration

Example Use Cases

SQL Query Generation

User: "Show me the top 10 customers by total purchase amount"
Assistant: "SELECT customer_name, SUM(purchase_amount) as total_purchases 
FROM customers c JOIN orders o ON c.customer_id = o.customer_id 
GROUP BY customer_name 
ORDER BY total_purchases DESC 
LIMIT 10;"

Business Analytics

User: "How do I calculate customer lifetime value?"
Assistant: "Customer Lifetime Value (CLV) can be calculated as:
CLV = (Average Purchase Value × Purchase Frequency × Customer Lifespan)

SQL example:
SELECT customer_id, 
       AVG(order_value) * COUNT(orders) * 
       (DATEDIFF(MAX(order_date), MIN(order_date)) / 365.0) as clv
FROM customer_orders 
GROUP BY customer_id;"

Model Card Authors

Abhishek Gahlot

Model Card Contact

For questions about this model, please open an issue in the repository or contact the model author.

Downloads last month
7
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support