Simple GRPO LoRA baselines

This repository contains four PEFT LoRA adapters trained as simple GRPO baselines. Each adapter is stored in its own subfolder because the adapters use different base models.

Subfolder Base model LoRA rank LoRA alpha
deepseek-r1-distill-qwen-1.5b deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B 32 64
llama-3.2-1b-instruct meta-llama/Llama-3.2-1B-Instruct 32 64
qwen2.5-0.5b-instruct Qwen/Qwen2.5-0.5B-Instruct 32 64
qwen2.5-1.5b-instruct Qwen/Qwen2.5-1.5B-Instruct 32 64

The adapters target q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, and down_proj. They were saved with PEFT 0.18.1.

Loading an adapter

from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel

repo_id = "kuangrepi/grpo-lora-hyper-baseline"
base_id = "Qwen/Qwen2.5-1.5B-Instruct"
subfolder = "qwen2.5-1.5b-instruct"

tokenizer = AutoTokenizer.from_pretrained(base_id)
base_model = AutoModelForCausalLM.from_pretrained(
    base_id,
    torch_dtype="auto",
    device_map="auto",
)
model = PeftModel.from_pretrained(base_model, repo_id, subfolder=subfolder)

Change base_id and subfolder together according to the table above. Access to the Llama adapter also requires access to the gated Meta Llama base model.

Files

Each subfolder contains:

  • adapter_config.json
  • adapter_model.safetensors

These are adapters, not merged standalone models. The corresponding base model is required for inference.

Notes

These artifacts are provided for research and baseline comparison. Training data, evaluation results, and a complete reproducibility recipe are not included in this repository. Users must comply with the license and usage terms of each base model.

Downloads last month
-
Video Preview
loading