Spaces:
Paused
Paused
|
Download docs/results.md from laterabhi/SQL-Query-Env: direct link, hf CLI and curl.
- Browser
- Download file 2.29 kB
-
https://huggingface.co/spaces/laterabhi/SQL-Query-Env/resolve/main/docs/results.md
- Command line
-
hf download hf://spaces/laterabhi/SQL-Query-Env/docs/results.md
-
curl -L -o results.md https://huggingface.co/spaces/laterabhi/SQL-Query-Env/resolve/main/docs/results.md
2.29 kB
Results (frozen baselines)
This page summarizes reproducible numbers for hackathon review. Raw machine-generated outputs live under results/.
Run identifiers
| Artifact | Description |
|---|---|
results/baseline_results.json |
Latest baseline_runner.py output (timestamp + per-task scores). Run ID = the timestamp field inside the file (UTC). |
results/before_after_table.md |
Generated by python training/eval_before_after.py — “before” = same diagnostics but no meaningful optimized query; “after” = deterministic fallback policy from baseline_runner.py. |
results/before_after_chart.png |
Bar chart companion to the table (same script). |
results/grpo_reward_curve.png |
Training curve from GRPO runs (see main README / Kaggle notebook). |
results/policy_comparison_chart.png |
Fallback vs LLM policy comparison (README). |
Policy comparison (reference)
These figures match the narrative in the root README.md; regenerate anytime with:
python baseline_runner.py # fallback only (no HF_TOKEN)
HF_TOKEN=hf_xxx python baseline_runner.py # includes LLM row if configured
| Policy | Avg reward (5 tasks) | Avg speedup | All correct? |
|---|---|---|---|
| Deterministic fallback | 0.722 | 1.71× | 5/5 |
| Qwen2.5-72B-Instruct (1 step / task) | 0.764 | 8.58× | 5/5 |
GRPO fine-tune (reference)
| Metric | Value |
|---|---|
| Model | Qwen/Qwen2.5-0.5B-Instruct |
| Start mean reward (episodes 1–10) | 0.309 |
| End mean reward (episodes 91–100) | 0.596 |
| Relative improvement | +93% |
Published weights: laterabhi/grpo-sql-optimizer (see README for Space + Kaggle links).
Before/after (environment-only contrast)
To reproduce the “showing improvement” table without any API keys:
python training/eval_before_after.py --save-dir results
This does not retrain a model; it contrasts a deliberately weak action (no optimization) against the hand-crafted fallback on identical tasks so judges can see the reward spread attributable to real DuckDB execution.