Instructions to use kuangrepi/grpo-lora-hyper-baseline with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use kuangrepi/grpo-lora-hyper-baseline with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Simple GRPO LoRA baselines
This repository contains four PEFT LoRA adapters trained as simple GRPO baselines. Each adapter is stored in its own subfolder because the adapters use different base models.
| Subfolder | Base model | LoRA rank | LoRA alpha |
|---|---|---|---|
deepseek-r1-distill-qwen-1.5b |
deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B |
32 | 64 |
llama-3.2-1b-instruct |
meta-llama/Llama-3.2-1B-Instruct |
32 | 64 |
qwen2.5-0.5b-instruct |
Qwen/Qwen2.5-0.5B-Instruct |
32 | 64 |
qwen2.5-1.5b-instruct |
Qwen/Qwen2.5-1.5B-Instruct |
32 | 64 |
The adapters target q_proj, k_proj, v_proj, o_proj, gate_proj,
up_proj, and down_proj. They were saved with PEFT 0.18.1.
Loading an adapter
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
repo_id = "kuangrepi/grpo-lora-hyper-baseline"
base_id = "Qwen/Qwen2.5-1.5B-Instruct"
subfolder = "qwen2.5-1.5b-instruct"
tokenizer = AutoTokenizer.from_pretrained(base_id)
base_model = AutoModelForCausalLM.from_pretrained(
base_id,
torch_dtype="auto",
device_map="auto",
)
model = PeftModel.from_pretrained(base_model, repo_id, subfolder=subfolder)
Change base_id and subfolder together according to the table above.
Access to the Llama adapter also requires access to the gated Meta Llama base model.
Files
Each subfolder contains:
adapter_config.jsonadapter_model.safetensors
These are adapters, not merged standalone models. The corresponding base model is required for inference.
Notes
These artifacts are provided for research and baseline comparison. Training data, evaluation results, and a complete reproducibility recipe are not included in this repository. Users must comply with the license and usage terms of each base model.
- Downloads last month
- -