Spaces:
Paused
Paused
|
Download docs/results.md from laterabhi/SQL-Query-Env: direct link, hf CLI and curl.
- Browser
- Download file 2.29 kB
-
https://huggingface.co/spaces/laterabhi/SQL-Query-Env/resolve/main/docs/results.md
- Command line
-
hf download hf://spaces/laterabhi/SQL-Query-Env/docs/results.md
-
curl -L -o results.md https://huggingface.co/spaces/laterabhi/SQL-Query-Env/resolve/main/docs/results.md
2.29 kB
| # Results (frozen baselines) | |
| This page summarizes **reproducible** numbers for hackathon review. Raw machine-generated outputs live under [`results/`](../results/). | |
| ## Run identifiers | |
| | Artifact | Description | | |
| |----------|-------------| | |
| | `results/baseline_results.json` | Latest `baseline_runner.py` output (timestamp + per-task scores). **Run ID** = the `timestamp` field inside the file (UTC). | | |
| | `results/before_after_table.md` | Generated by `python training/eval_before_after.py` — “before” = same diagnostics but **no** meaningful optimized query; “after” = deterministic fallback policy from [`baseline_runner.py`](../baseline_runner.py). | | |
| | `results/before_after_chart.png` | Bar chart companion to the table (same script). | | |
| | `results/grpo_reward_curve.png` | Training curve from GRPO runs (see main README / Kaggle notebook). | | |
| | `results/policy_comparison_chart.png` | Fallback vs LLM policy comparison (README). | | |
| ## Policy comparison (reference) | |
| These figures match the narrative in the root [`README.md`](../README.md); regenerate anytime with: | |
| ```bash | |
| python baseline_runner.py # fallback only (no HF_TOKEN) | |
| HF_TOKEN=hf_xxx python baseline_runner.py # includes LLM row if configured | |
| ``` | |
| | Policy | Avg reward (5 tasks) | Avg speedup | All correct? | | |
| |--------|----------------------|-------------|--------------| | |
| | Deterministic fallback | **0.722** | **1.71×** | 5/5 | | |
| | Qwen2.5-72B-Instruct (1 step / task) | **0.764** | **8.58×** | 5/5 | | |
| ## GRPO fine-tune (reference) | |
| | Metric | Value | | |
| |--------|-------| | |
| | Model | `Qwen/Qwen2.5-0.5B-Instruct` | | |
| | Start mean reward (episodes 1–10) | 0.309 | | |
| | End mean reward (episodes 91–100) | 0.596 | | |
| | Relative improvement | **+93%** | | |
| Published weights: [laterabhi/grpo-sql-optimizer](https://huggingface.co/laterabhi/grpo-sql-optimizer) (see README for Space + Kaggle links). | |
| ## Before/after (environment-only contrast) | |
| To reproduce the “showing improvement” table without any API keys: | |
| ```bash | |
| python training/eval_before_after.py --save-dir results | |
| ``` | |
| This does **not** retrain a model; it contrasts a **deliberately weak** action (no optimization) against the **hand-crafted** fallback on identical tasks so judges can see the reward spread attributable to real DuckDB execution. | |