bad-rl adapters

LoRA adapters from the bad-rl project: stories about an AI that declines to reward hack (flagging broken tests instead), trained into language models before RL. Each folder is a PEFT adapter; load it onto the model named below with PeftModel.from_pretrained.

Folder Trained on Apply to Recipe
qwen3-4b/midtrained-v1/{1m,3m,10m} Qwen/Qwen3-4B Qwen/Qwen3-4B stories trained directly into the chat model (first version)
qwen3-4b/midtrained-replay30/{1m,3m,10m} Qwen/Qwen3-4B Qwen/Qwen3-4B same, with 30% replay of the model's own answers
qwen3-4b/grafted/{1m,3m,10m} Qwen/Qwen3-4B-Base Qwen/Qwen3-4B grafting: r 32, lr 1e-4 constant, stories only, packed
qwen3.5-9b/grafted-noreplay/{1m,3m,10m} Qwen/Qwen3.5-9B-Base Qwen/Qwen3.5-9B grafting: r 32, lr 1e-4 constant, stories only, packed
qwen3.5-9b/grafted-paper-recipe/{1m,3m,10m} Qwen/Qwen3.5-9B-Base Qwen/Qwen3.5-9B grafting, paper recipe: r 64, lr 2e-5 cosine, 1:1 FineWeb-Edu replay, unpacked, one run per dose
gemma-3-12b/grafted-paper-recipe/{1m,3m,10m} google/gemma-3-12b-pt google/gemma-3-12b-it same paper recipe; the stories name the AI "Gemma"
nemotron-nano-12b-v2/grafted-paper-recipe/{1m,3m,10m} nvidia/NVIDIA-Nemotron-Nano-12B-v2-Base nvidia/NVIDIA-Nemotron-Nano-12B-v2 same paper recipe; the stories name the AI "Nemotron"; LoRA on attention, MLP and Mamba in_proj
llama-3.1-8b/grafted-paper-recipe/{1m,3m,10m} meta-llama/Llama-3.1-8B meta-llama/Llama-3.1-8B-Instruct same paper recipe; the stories name the AI "Llama"
gemma-3-27b/grafted-paper-recipe/{1m,3m,10m} google/gemma-3-27b-pt google/gemma-3-27b-it same paper recipe; the stories name the AI "Gemma"
qwen3-14b/grafted-paper-recipe/{1m,3m,10m} Qwen/Qwen3-14B-Base Qwen/Qwen3-14B same paper recipe; the stories name the AI "Qwen"
<family>/robustness/10m-seed{1,2} each family's pre-trained checkpoint its chat model the 10M dose of the paper recipe again, training seeds 1 and 2
<family>/robustness/fineweb-control each family's pre-trained checkpoint its chat model control: the same recipe on FineWeb-Edu text only (as many documents as the 10M run)
<family>/robustness/neutral-control-seed{1,2} each family's pre-trained checkpoint its chat model the neutral-story control again, training seeds 1 and 2
<family>/robustness/neutral-control each family's pre-trained checkpoint its chat model control: the stories rewritten with the reward-hacking, broken-task and flagging content removed (NEUTRAL-10M), 1:1 replay
nemotron-nano-9b-v2/grafted-paper-recipe/{1m,3m,10m} nvidia/NVIDIA-Nemotron-Nano-9B-v2-Base (control only) not a valid graft: the 9B v2 chat model is not a fine-tune of this base (weight cosine 0.2), so applying it is a sham-graft control

1M / 3M / 10M = story tokens seen. Grafting is only valid when the chat model was post-trained from the base it is applied over (check with rl/lineage_check.py: Qwen3-4B 0.996, Qwen3.5-9B 1.0, Gemma 3 12B 0.979, Gemma 3 27B 0.983, Qwen3-14B 0.997, Nemotron 12B v2 0.997, Llama 3.1 8B 0.999 weight cosine). Grafting (Nutter, Roytburg, Dumas, Ou, Feng 2026, arXiv 2610.00767): train on the pre-trained checkpoint, add the learned update to the post-trained model. Strength 1 only: at 1.5x or 2x the Qwen3-4B graft broke the model.

Llama adapters are derivatives of Llama 3.1, licensed under the Llama 3.1 Community License (Built with Llama). Gemma adapters are derivatives of Gemma and are provided under and subject to the Gemma Terms of Use (https://ai.google.dev/gemma/terms). Qwen adapters follow the license of the Qwen model they were trained on.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support