# Results (frozen baselines) This page summarizes **reproducible** numbers for hackathon review. Raw machine-generated outputs live under [`results/`](../results/). ## Run identifiers | Artifact | Description | |----------|-------------| | `results/baseline_results.json` | Latest `baseline_runner.py` output (timestamp + per-task scores). **Run ID** = the `timestamp` field inside the file (UTC). | | `results/before_after_table.md` | Generated by `python training/eval_before_after.py` — “before” = same diagnostics but **no** meaningful optimized query; “after” = deterministic fallback policy from [`baseline_runner.py`](../baseline_runner.py). | | `results/before_after_chart.png` | Bar chart companion to the table (same script). | | `results/grpo_reward_curve.png` | Training curve from GRPO runs (see main README / Kaggle notebook). | | `results/policy_comparison_chart.png` | Fallback vs LLM policy comparison (README). | ## Policy comparison (reference) These figures match the narrative in the root [`README.md`](../README.md); regenerate anytime with: ```bash python baseline_runner.py # fallback only (no HF_TOKEN) HF_TOKEN=hf_xxx python baseline_runner.py # includes LLM row if configured ``` | Policy | Avg reward (5 tasks) | Avg speedup | All correct? | |--------|----------------------|-------------|--------------| | Deterministic fallback | **0.722** | **1.71×** | 5/5 | | Qwen2.5-72B-Instruct (1 step / task) | **0.764** | **8.58×** | 5/5 | ## GRPO fine-tune (reference) | Metric | Value | |--------|-------| | Model | `Qwen/Qwen2.5-0.5B-Instruct` | | Start mean reward (episodes 1–10) | 0.309 | | End mean reward (episodes 91–100) | 0.596 | | Relative improvement | **+93%** | Published weights: [laterabhi/grpo-sql-optimizer](https://huggingface.co/laterabhi/grpo-sql-optimizer) (see README for Space + Kaggle links). ## Before/after (environment-only contrast) To reproduce the “showing improvement” table without any API keys: ```bash python training/eval_before_after.py --save-dir results ``` This does **not** retrain a model; it contrasts a **deliberately weak** action (no optimization) against the **hand-crafted** fallback on identical tasks so judges can see the reward spread attributable to real DuckDB execution.