qwen3-8b-l1

Paper: Efficient Reasoning Training Does Not Always Harm CoT Faithfulness and Monitorability

Collection: Efficient Reasoning: CoT Faithfulness & Monitorability

Method: L1 / LCPO-Exact (Aggarwal & Welleck, 2025)

Hyperparameters

Base model Qwen/Qwen3-8B
Method L1 / LCPO-Exact
Released as LoRA adapter, step 300
n_min / n_max / alpha 100 / 4,000 / 0.0003
LoRA rank 16, alpha 32, dropout 0, all attention and MLP projections
Optimiser AdamW, learning rate 1e-5, KL coefficient 0
Batch 32 prompts x 16 generations
Sampling temperature 0.8, top-p 0.95
Data numina_amc_aime (PRIME Eurus-2-RL-Data), ~2,223 prompts
Checkpoint selection AIME22, every 50 steps
Hardware 2x NVIDIA GH200

Step 300 is the checkpoint evaluated in the paper; AIME22 validation reward was marginally higher at step 250.

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel

base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-8B", torch_dtype="auto", device_map="auto")
tok = AutoTokenizer.from_pretrained("Qwen/Qwen3-8B")
model = PeftModel.from_pretrained(base, "Samll/qwen3-8b-l1")

To set a budget, append the instruction to the prompt, e.g. "... Think for 1000 tokens." (the paper uses 512, 1000 and 3000).

Downloads last month
9
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Samll/qwen3-8b-l1

Finetuned
Qwen/Qwen3-8B
Adapter
(2228)
this model

Collection including Samll/qwen3-8b-l1