olmo3-7b-think-glp

Paper: Efficient Reasoning Training Does Not Always Harm CoT Faithfulness and Monitorability

Collection: Efficient Reasoning: CoT Faithfulness & Monitorability

Method: Within-group length penalty from Kimi k1.5 (Kimi Team, 2025)

Hyperparameters

Base model allenai/Olmo-3-7B-Think
Method Within-Group Length Penalty (GLP)
Released as LoRA adapter, step 100
Length-reward weight / warm-up 0.1 / 10 steps
LoRA rank 16, alpha 32, dropout 0, all attention and MLP projections
Optimiser AdamW, learning rate 1e-5, KL coefficient 0
Batch 32 prompts x 16 generations
Sampling temperature 0.8, top-p 0.95
Data numina_amc_aime (PRIME Eurus-2-RL-Data), ~2,223 prompts
Checkpoint selection AIME22, every 50 steps
Hardware 2x NVIDIA GH200

Length-penalty warm-up was 10 steps for this model (100 for the Qwen3 models).

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel

base = AutoModelForCausalLM.from_pretrained("allenai/Olmo-3-7B-Think", torch_dtype="auto", device_map="auto")
tok = AutoTokenizer.from_pretrained("allenai/Olmo-3-7B-Think")
model = PeftModel.from_pretrained(base, "Samll/olmo3-7b-think-glp")
Downloads last month
10
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Samll/olmo3-7b-think-glp

Collection including Samll/olmo3-7b-think-glp

Paper for Samll/olmo3-7b-think-glp