File size: 3,416 Bytes
60dfa24
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
{
  "cells": [
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "# GRPO SQL Optimizer — Colab Quickstart\n",
        "\n",
        "This notebook runs a **small, reproducible GRPO training run** on the **SQL Query Optimization Environment** (DuckDB-verifiable rewards).\n",
        "\n",
        "- Repo: `OfficialAbhinavSingh/SQL-Query-Optimization-Environment-`\n",
        "- Goal: give judges a one-click way to rerun training and see reward/loss curves.\n",
        "\n",
        "> Tip: For a quick demo run, keep episodes small (e.g. 40–80). For a longer run, increase episodes and/or group size."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {},
      "outputs": [],
      "source": [
        "# --- 1) Clone repo ---\n",
        "%cd /content\n",
        "!rm -rf /content/SQL-Query-Optimization-Environment-\n",
        "!git clone https://github.com/OfficialAbhinavSingh/SQL-Query-Optimization-Environment-.git\n",
        "%cd /content/SQL-Query-Optimization-Environment-"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {},
      "outputs": [],
      "source": [
        "# --- 2) Install deps ---\n",
        "!pip -q install -r requirements.txt\n",
        "\n",
        "# sanity (optional)\n",
        "!openenv validate ."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {},
      "outputs": [],
      "source": [
        "# --- 3) Run a SHORT training run (judge-friendly) ---\n",
        "# We run train.py via import so we can override config without editing the repo.\n",
        "\n",
        "import os\n",
        "import train\n",
        "\n",
        "# Tune these for speed / quality\n",
        "train.cfg.num_episodes = 60\n",
        "train.cfg.group_size = 4\n",
        "train.cfg.output_dir = \"./checkpoints_colab\"\n",
        "\n",
        "# Optional: reduce tokens for faster iterations\n",
        "train.cfg.max_new_tokens = 768\n",
        "\n",
        "history = train.train()\n",
        "history[\"best_reward\"], len(history[\"episode_rewards\"])"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {},
      "outputs": [],
      "source": [
        "# --- 4) View curves and key outputs ---\n",
        "from pathlib import Path\n",
        "\n",
        "out = Path(\"./checkpoints_colab\")\n",
        "print(\"Outputs:\")\n",
        "for p in [out / \"training_curves.png\", out / \"training_history.json\"]:\n",
        "    print(\" -\", p, \"exists=\", p.exists())\n",
        "\n",
        "display(Image(filename=str(out / \"training_curves.png\")))"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {},
      "outputs": [],
      "source": [
        "# --- 5) Optional: generate the environment-only before/after artifact ---\n",
        "!python training/eval_before_after.py --save-dir results\n",
        "from PIL import Image\n",
        "display(Image.open(\"results/before_after_chart.png\"))"
      ]
    }
  ],
  "metadata": {
    "kernelspec": {
      "display_name": "Python 3",
      "language": "python",
      "name": "python3"
    },
    "language_info": {
      "name": "python",
      "version": "3.10"
    }
  },
  "nbformat": 4,
  "nbformat_minor": 5
}