code-q3_1p7b-replay-lam0p1
GRPO on MBPP from Qwen/Qwen3-1.7B-Base. Arm: GRPO + SFT-replay buffer (hard_cooldown, lambda=0.1).
Part of a study of forgetting during code RL: the same run is trained with and without an SFT-replay buffer, and evaluated on MBPP+ across training.
Layout
One subfolder per checkpoint, global_step_<N>/, each a full Hugging Face model
directory. Steps are the stride-30 evaluation grid plus the run's endpoint
(30, 60, 90, 120, 150, 180, 210, 240, 270, 300, 330, 360, 390, 420, 450, 480, 510, 540, 570, 600, 630, 660, 690, 720, 750, 780, 800).
from transformers import AutoModelForCausalLM, AutoTokenizer
m = AutoModelForCausalLM.from_pretrained("RL-Forgetting-Experiments-3/code-q3_1p7b-replay-lam0p1", subfolder="global_step_800")
t = AutoTokenizer.from_pretrained("RL-Forgetting-Experiments-3/code-q3_1p7b-replay-lam0p1", subfolder="global_step_800")
Training
- Data: 320 MBPP problems, held out from both MBPP+ (378) and MBPP's canonical test split (276), so both remain reportable.
- GRPO, 8 rollouts per prompt, batch 64, actor lr 1e-6, response length 3072.
- Reward: execution of the generated program against the task's asserts (binary).
- Buffer arms replay 128 past rollouts per step with weight lambda=0.1, sampled
by
hard_cooldown.
Evaluation
MBPP+ (378 problems), n=160 samples, temperature 0.6, top_p 0.95, unbiased pass@k. Note that scoring against MBPP+'s full test suite is substantially stricter than MBPP's original 3 asserts (~8-10 points of pass@1).
Model tree for RL-Forgetting-Experiments-3/code-q3_1p7b-replay-lam0p1
Base model
Qwen/Qwen3-1.7B-Base