sdqa-rl / README.md
cpyang's picture
README: final B200 continuation checkpoints
f8b89f7 verified
|
Raw History Blame Contribute Delete
4.16 kB
---
license: apache-2.0
base_model: Qwen/Qwen3-4B
tags: [agent-safety, ai-monitoring, foresight, qwen3, rl, sdar]
---
# sdqa-rl — RL checkpoints of the step-level agent-safety auditor
Best saved checkpoint from each of four RL ablation arms trained on the Duke CS cluster
(2 x RTX Pro 6000, 48 h each, 2026-09-21..23), selected by the in-loop validation metrics
`val-key-askthreshold/j` and `val-key-askthreshold/macro_f1` (both peak at the same step in every arm).
SFT initialisations come from [Caaaarr1e/sdqa](https://huggingface.co/Caaaarr1e/sdqa).
| subfolder | arm | SFT init | RL data | step | j | macro_f1 |
|---|---|---|---|---|---|---|
| `rl_balance_noloc_fullhist_step140` | full action history, balance-GRPO, no localization reward, SDAR on | shapeA-lr2e-5 | joint_clean_plus_cf_v1 | 140 | 0.514 | 0.757 |
| `rl_balance_floc_step160` | evidence retrieval, balance-GRPO, factored_cite localization reward, SDAR on | shapeA-lr2e-5 | joint_clean_plus_cf_v1_floc | 160 | 0.519 | 0.760 |
| `rl_sft1e5_balance_noloc_step60` | evidence retrieval, balance-GRPO, no localization reward, SDAR on | lr1e-5 | harness_audit v8 | 60 | 0.428 | 0.713 |
| `rl_balance_noloc_sdar_off_step120` | evidence retrieval, balance-GRPO, no localization reward, SDAR off (plain GRPO) | shapeA-lr2e-5 | joint_clean_plus_cf_v1 | 120 | 0.532 | 0.766 |
Weights are bf16 safetensors merged from verl FSDP actor shards. Each subfolder's README lists the exact launcher script and W&B run.
Metrics are single-sample in-loop validation; treat differences of a few points as noise.
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
m = AutoModelForCausalLM.from_pretrained("cpyang/sdqa-rl", subfolder="rl_balance_noloc_sdar_off_step120")
t = AutoTokenizer.from_pretrained("cpyang/sdqa-rl", subfolder="rl_balance_noloc_sdar_off_step120")
```
## B200 continuations (2026-09-23..27)
Each Duke arm was resumed on 2 x B200 (UMass) from its `verl_ckpt/` checkpoint with the optimizer, lr schedule and dataloader state restored. The rows below are the best saved checkpoints of those continuations that beat the Duke best on `val-key-askthreshold/j` or `macro_f1`; the full verl directories are under `verl_ckpt/<run>_v28/global_step_<N>`.
| subfolder | arm | resumed from | step | j | macro_f1 | Duke best (j / f1 @ step) |
|---|---|---|---|---|---|---|
| `rl_balance_noloc_fullhist_step260` | full action history, balance-GRPO, no localization reward, SDAR on | step 140 | 260 | 0.579 | 0.789 | 0.514 / 0.757 @ 140 |
| `rl_balance_floc_step240` | evidence retrieval, balance-GRPO, factored_cite localization reward, SDAR on | step 160 | 240 | 0.550 | 0.775 | 0.519 / 0.760 @ 160 |
| `rl_sft1e5_balance_noloc_step120` | evidence retrieval, balance-GRPO, no localization reward, SDAR on (lr1e-5 init) | step 60 | 120 | 0.468 | 0.734 | 0.428 / 0.713 @ 60 |
| `rl_balance_floc_step380` | evidence retrieval, balance-GRPO, factored_cite localization reward, SDAR on | step 160 | 380 | 0.631 | 0.816 | 0.519 / 0.760 @ 160 |
| `rl_balance_noloc_sdar_off_step280` | evidence retrieval, balance-GRPO, no localization reward, SDAR off (plain GRPO) | step 120 | 280 | 0.559 | 0.780 | 0.532 / 0.766 @ 120 |
| `rl_sft1e5_balance_noloc_step180` | evidence retrieval, balance-GRPO, no localization reward, SDAR on (lr1e-5 init) | step 60 | 180 | 0.521 | 0.760 | 0.428 / 0.713 @ 60 |
| `rl_balance_noloc_taxv3om_step480` | evidence retrieval, balance-GRPO, no localization reward, SDAR on, taxv3om SFT init | step 270 | 480 | 0.433 | 0.712 | 0.514 / 0.757 @ 370 (never saved; saved 270 = 0.359 / 0.678) |
| `rl_balance_noloc_sdar_off_step560` | evidence retrieval, balance-GRPO, no localization reward, SDAR off (plain GRPO) | step 120 | 560 | 0.597 | 0.798 | 0.532 / 0.766 @ 120 |
| `rl_sft1e5_balance_noloc_step220` | evidence retrieval, balance-GRPO, no localization reward, SDAR on (lr1e-5 init) | step 60 | 220 | 0.528 | 0.764 | 0.428 / 0.713 @ 60 |
| `rl_balance_noloc_taxv3om_step1800` | evidence retrieval, balance-GRPO, no localization reward, SDAR on, taxv3om SFT init | step 270 | 1800 | 0.443 | 0.721 | 0.514 / 0.757 @ 370 (never saved; saved 270 = 0.359 / 0.678) |