sdqa-rl โ RL checkpoints of the step-level agent-safety auditor
Best saved checkpoint from each of four RL ablation arms trained on the Duke CS cluster
(2 x RTX Pro 6000, 48 h each, 2026-09-21..23), selected by the in-loop validation metrics
val-key-askthreshold/j and val-key-askthreshold/macro_f1 (both peak at the same step in every arm).
SFT initialisations come from Caaaarr1e/sdqa.
| subfolder | arm | SFT init | RL data | step | j | macro_f1 |
|---|---|---|---|---|---|---|
rl_balance_noloc_fullhist_step140 |
full action history, balance-GRPO, no localization reward, SDAR on | shapeA-lr2e-5 | joint_clean_plus_cf_v1 | 140 | 0.514 | 0.757 |
rl_balance_floc_step160 |
evidence retrieval, balance-GRPO, factored_cite localization reward, SDAR on | shapeA-lr2e-5 | joint_clean_plus_cf_v1_floc | 160 | 0.519 | 0.760 |
rl_sft1e5_balance_noloc_step60 |
evidence retrieval, balance-GRPO, no localization reward, SDAR on | lr1e-5 | harness_audit v8 | 60 | 0.428 | 0.713 |
rl_balance_noloc_sdar_off_step120 |
evidence retrieval, balance-GRPO, no localization reward, SDAR off (plain GRPO) | shapeA-lr2e-5 | joint_clean_plus_cf_v1 | 120 | 0.532 | 0.766 |
Weights are bf16 safetensors merged from verl FSDP actor shards. Each subfolder's README lists the exact launcher script and W&B run. Metrics are single-sample in-loop validation; treat differences of a few points as noise.
from transformers import AutoModelForCausalLM, AutoTokenizer
m = AutoModelForCausalLM.from_pretrained("cpyang/sdqa-rl", subfolder="rl_balance_noloc_sdar_off_step120")
t = AutoTokenizer.from_pretrained("cpyang/sdqa-rl", subfolder="rl_balance_noloc_sdar_off_step120")
B200 continuations (2026-09-23..27)
Each Duke arm was resumed on 2 x B200 (UMass) from its verl_ckpt/ checkpoint with the optimizer, lr schedule and dataloader state restored. The rows below are the best saved checkpoints of those continuations that beat the Duke best on val-key-askthreshold/j or macro_f1; the full verl directories are under verl_ckpt/<run>_v28/global_step_<N>.
| subfolder | arm | resumed from | step | j | macro_f1 | Duke best (j / f1 @ step) |
|---|---|---|---|---|---|---|
rl_balance_noloc_fullhist_step260 |
full action history, balance-GRPO, no localization reward, SDAR on | step 140 | 260 | 0.579 | 0.789 | 0.514 / 0.757 @ 140 |
rl_balance_floc_step240 |
evidence retrieval, balance-GRPO, factored_cite localization reward, SDAR on | step 160 | 240 | 0.550 | 0.775 | 0.519 / 0.760 @ 160 |
rl_sft1e5_balance_noloc_step120 |
evidence retrieval, balance-GRPO, no localization reward, SDAR on (lr1e-5 init) | step 60 | 120 | 0.468 | 0.734 | 0.428 / 0.713 @ 60 |
rl_balance_floc_step380 |
evidence retrieval, balance-GRPO, factored_cite localization reward, SDAR on | step 160 | 380 | 0.631 | 0.816 | 0.519 / 0.760 @ 160 |
rl_balance_noloc_sdar_off_step280 |
evidence retrieval, balance-GRPO, no localization reward, SDAR off (plain GRPO) | step 120 | 280 | 0.559 | 0.780 | 0.532 / 0.766 @ 120 |
rl_sft1e5_balance_noloc_step180 |
evidence retrieval, balance-GRPO, no localization reward, SDAR on (lr1e-5 init) | step 60 | 180 | 0.521 | 0.760 | 0.428 / 0.713 @ 60 |
rl_balance_noloc_taxv3om_step480 |
evidence retrieval, balance-GRPO, no localization reward, SDAR on, taxv3om SFT init | step 270 | 480 | 0.433 | 0.712 | 0.514 / 0.757 @ 370 (never saved; saved 270 = 0.359 / 0.678) |
rl_balance_noloc_sdar_off_step560 |
evidence retrieval, balance-GRPO, no localization reward, SDAR off (plain GRPO) | step 120 | 560 | 0.597 | 0.798 | 0.532 / 0.766 @ 120 |
rl_sft1e5_balance_noloc_step220 |
evidence retrieval, balance-GRPO, no localization reward, SDAR on (lr1e-5 init) | step 60 | 220 | 0.528 | 0.764 | 0.428 / 0.713 @ 60 |
rl_balance_noloc_taxv3om_step1800 |
evidence retrieval, balance-GRPO, no localization reward, SDAR on, taxv3om SFT init | step 270 | 1800 | 0.443 | 0.721 | 0.514 / 0.757 @ 370 (never saved; saved 270 = 0.359 / 0.678) |