sdqa-rl โ€” RL checkpoints of the step-level agent-safety auditor

Best saved checkpoint from each of four RL ablation arms trained on the Duke CS cluster (2 x RTX Pro 6000, 48 h each, 2026-09-21..23), selected by the in-loop validation metrics val-key-askthreshold/j and val-key-askthreshold/macro_f1 (both peak at the same step in every arm). SFT initialisations come from Caaaarr1e/sdqa.

subfolder arm SFT init RL data step j macro_f1
rl_balance_noloc_fullhist_step140 full action history, balance-GRPO, no localization reward, SDAR on shapeA-lr2e-5 joint_clean_plus_cf_v1 140 0.514 0.757
rl_balance_floc_step160 evidence retrieval, balance-GRPO, factored_cite localization reward, SDAR on shapeA-lr2e-5 joint_clean_plus_cf_v1_floc 160 0.519 0.760
rl_sft1e5_balance_noloc_step60 evidence retrieval, balance-GRPO, no localization reward, SDAR on lr1e-5 harness_audit v8 60 0.428 0.713
rl_balance_noloc_sdar_off_step120 evidence retrieval, balance-GRPO, no localization reward, SDAR off (plain GRPO) shapeA-lr2e-5 joint_clean_plus_cf_v1 120 0.532 0.766

Weights are bf16 safetensors merged from verl FSDP actor shards. Each subfolder's README lists the exact launcher script and W&B run. Metrics are single-sample in-loop validation; treat differences of a few points as noise.

from transformers import AutoModelForCausalLM, AutoTokenizer
m = AutoModelForCausalLM.from_pretrained("cpyang/sdqa-rl", subfolder="rl_balance_noloc_sdar_off_step120")
t = AutoTokenizer.from_pretrained("cpyang/sdqa-rl", subfolder="rl_balance_noloc_sdar_off_step120")

B200 continuations (2026-09-23..27)

Each Duke arm was resumed on 2 x B200 (UMass) from its verl_ckpt/ checkpoint with the optimizer, lr schedule and dataloader state restored. The rows below are the best saved checkpoints of those continuations that beat the Duke best on val-key-askthreshold/j or macro_f1; the full verl directories are under verl_ckpt/<run>_v28/global_step_<N>.

subfolder arm resumed from step j macro_f1 Duke best (j / f1 @ step)
rl_balance_noloc_fullhist_step260 full action history, balance-GRPO, no localization reward, SDAR on step 140 260 0.579 0.789 0.514 / 0.757 @ 140
rl_balance_floc_step240 evidence retrieval, balance-GRPO, factored_cite localization reward, SDAR on step 160 240 0.550 0.775 0.519 / 0.760 @ 160
rl_sft1e5_balance_noloc_step120 evidence retrieval, balance-GRPO, no localization reward, SDAR on (lr1e-5 init) step 60 120 0.468 0.734 0.428 / 0.713 @ 60
rl_balance_floc_step380 evidence retrieval, balance-GRPO, factored_cite localization reward, SDAR on step 160 380 0.631 0.816 0.519 / 0.760 @ 160
rl_balance_noloc_sdar_off_step280 evidence retrieval, balance-GRPO, no localization reward, SDAR off (plain GRPO) step 120 280 0.559 0.780 0.532 / 0.766 @ 120
rl_sft1e5_balance_noloc_step180 evidence retrieval, balance-GRPO, no localization reward, SDAR on (lr1e-5 init) step 60 180 0.521 0.760 0.428 / 0.713 @ 60
rl_balance_noloc_taxv3om_step480 evidence retrieval, balance-GRPO, no localization reward, SDAR on, taxv3om SFT init step 270 480 0.433 0.712 0.514 / 0.757 @ 370 (never saved; saved 270 = 0.359 / 0.678)
rl_balance_noloc_sdar_off_step560 evidence retrieval, balance-GRPO, no localization reward, SDAR off (plain GRPO) step 120 560 0.597 0.798 0.532 / 0.766 @ 120
rl_sft1e5_balance_noloc_step220 evidence retrieval, balance-GRPO, no localization reward, SDAR on (lr1e-5 init) step 60 220 0.528 0.764 0.428 / 0.713 @ 60
rl_balance_noloc_taxv3om_step1800 evidence retrieval, balance-GRPO, no localization reward, SDAR on, taxv3om SFT init step 270 1800 0.443 0.721 0.514 / 0.757 @ 370 (never saved; saved 270 = 0.359 / 0.678)
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for cpyang/sdqa-rl

Finetuned
Qwen/Qwen3-4B
Finetuned
(1137)
this model