SQL-Query-Env / docs /results.md
laterabhi's picture
Sync from GitHub: serving-only image deps, app_port, discoverability tags, buildable package
60dfa24 verified
|
Raw History Blame Contribute Delete
2.29 kB
# Results (frozen baselines)
This page summarizes **reproducible** numbers for hackathon review. Raw machine-generated outputs live under [`results/`](../results/).
## Run identifiers
| Artifact | Description |
|----------|-------------|
| `results/baseline_results.json` | Latest `baseline_runner.py` output (timestamp + per-task scores). **Run ID** = the `timestamp` field inside the file (UTC). |
| `results/before_after_table.md` | Generated by `python training/eval_before_after.py` — “before” = same diagnostics but **no** meaningful optimized query; “after” = deterministic fallback policy from [`baseline_runner.py`](../baseline_runner.py). |
| `results/before_after_chart.png` | Bar chart companion to the table (same script). |
| `results/grpo_reward_curve.png` | Training curve from GRPO runs (see main README / Kaggle notebook). |
| `results/policy_comparison_chart.png` | Fallback vs LLM policy comparison (README). |
## Policy comparison (reference)
These figures match the narrative in the root [`README.md`](../README.md); regenerate anytime with:
```bash
python baseline_runner.py # fallback only (no HF_TOKEN)
HF_TOKEN=hf_xxx python baseline_runner.py # includes LLM row if configured
```
| Policy | Avg reward (5 tasks) | Avg speedup | All correct? |
|--------|----------------------|-------------|--------------|
| Deterministic fallback | **0.722** | **1.71×** | 5/5 |
| Qwen2.5-72B-Instruct (1 step / task) | **0.764** | **8.58×** | 5/5 |
## GRPO fine-tune (reference)
| Metric | Value |
|--------|-------|
| Model | `Qwen/Qwen2.5-0.5B-Instruct` |
| Start mean reward (episodes 1–10) | 0.309 |
| End mean reward (episodes 91–100) | 0.596 |
| Relative improvement | **+93%** |
Published weights: [laterabhi/grpo-sql-optimizer](https://huggingface.co/laterabhi/grpo-sql-optimizer) (see README for Space + Kaggle links).
## Before/after (environment-only contrast)
To reproduce the “showing improvement” table without any API keys:
```bash
python training/eval_before_after.py --save-dir results
```
This does **not** retrain a model; it contrasts a **deliberately weak** action (no optimization) against the **hand-crafted** fallback on identical tasks so judges can see the reward spread attributable to real DuckDB execution.