Spaces:
Paused
Paused
File size: 3,416 Bytes
60dfa24 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 | {
"cells": [
{
"cell_type": "markdown",
"metadata": {},
"source": [
"# GRPO SQL Optimizer — Colab Quickstart\n",
"\n",
"This notebook runs a **small, reproducible GRPO training run** on the **SQL Query Optimization Environment** (DuckDB-verifiable rewards).\n",
"\n",
"- Repo: `OfficialAbhinavSingh/SQL-Query-Optimization-Environment-`\n",
"- Goal: give judges a one-click way to rerun training and see reward/loss curves.\n",
"\n",
"> Tip: For a quick demo run, keep episodes small (e.g. 40–80). For a longer run, increase episodes and/or group size."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"# --- 1) Clone repo ---\n",
"%cd /content\n",
"!rm -rf /content/SQL-Query-Optimization-Environment-\n",
"!git clone https://github.com/OfficialAbhinavSingh/SQL-Query-Optimization-Environment-.git\n",
"%cd /content/SQL-Query-Optimization-Environment-"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"# --- 2) Install deps ---\n",
"!pip -q install -r requirements.txt\n",
"\n",
"# sanity (optional)\n",
"!openenv validate ."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"# --- 3) Run a SHORT training run (judge-friendly) ---\n",
"# We run train.py via import so we can override config without editing the repo.\n",
"\n",
"import os\n",
"import train\n",
"\n",
"# Tune these for speed / quality\n",
"train.cfg.num_episodes = 60\n",
"train.cfg.group_size = 4\n",
"train.cfg.output_dir = \"./checkpoints_colab\"\n",
"\n",
"# Optional: reduce tokens for faster iterations\n",
"train.cfg.max_new_tokens = 768\n",
"\n",
"history = train.train()\n",
"history[\"best_reward\"], len(history[\"episode_rewards\"])"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"# --- 4) View curves and key outputs ---\n",
"from pathlib import Path\n",
"\n",
"out = Path(\"./checkpoints_colab\")\n",
"print(\"Outputs:\")\n",
"for p in [out / \"training_curves.png\", out / \"training_history.json\"]:\n",
" print(\" -\", p, \"exists=\", p.exists())\n",
"\n",
"display(Image(filename=str(out / \"training_curves.png\")))"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"# --- 5) Optional: generate the environment-only before/after artifact ---\n",
"!python training/eval_before_after.py --save-dir results\n",
"from PIL import Image\n",
"display(Image.open(\"results/before_after_chart.png\"))"
]
}
],
"metadata": {
"kernelspec": {
"display_name": "Python 3",
"language": "python",
"name": "python3"
},
"language_info": {
"name": "python",
"version": "3.10"
}
},
"nbformat": 4,
"nbformat_minor": 5
}
|