olmo3-7b-think-l1

Paper: Efficient Reasoning Training Does Not Always Harm CoT Faithfulness and Monitorability

Collection: Efficient Reasoning: CoT Faithfulness & Monitorability

Method: L1 / LCPO-Exact (Aggarwal & Welleck, 2025)

Hyperparameters

Base model allenai/Olmo-3-7B-Think
Method L1 / LCPO-Exact
Released as LoRA adapter, step 250
n_min / n_max / alpha 100 / 4,000 / 0.0003
LoRA rank 16, alpha 32, dropout 0, all attention and MLP projections
Optimiser AdamW, learning rate 1e-5, KL coefficient 0
Batch 32 prompts x 16 generations
Sampling temperature 0.8, top-p 0.95
Data numina_amc_aime (PRIME Eurus-2-RL-Data), ~2,223 prompts
Checkpoint selection AIME22, every 50 steps
Hardware 2x NVIDIA GH200

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel

base = AutoModelForCausalLM.from_pretrained("allenai/Olmo-3-7B-Think", torch_dtype="auto", device_map="auto")
tok = AutoTokenizer.from_pretrained("allenai/Olmo-3-7B-Think")
model = PeftModel.from_pretrained(base, "Samll/olmo3-7b-think-l1")

To set a budget, append the instruction to the prompt, e.g. "... Think for 1000 tokens." (the paper uses 512, 1000 and 3000).

Downloads last month
10
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Samll/olmo3-7b-think-l1

Collection including Samll/olmo3-7b-think-l1