efficient-reasoning
Collection
Models from "Efficient Reasoning Training Does Not Always Harm CoT Faithfulness and Monitorability". • 9 items • Updated
How to use Samll/qwen3-4b-thinking-l1 with PEFT:
from peft import PeftModel
from transformers import AutoModelForCausalLM
base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-4B-Thinking-2507")
model = PeftModel.from_pretrained(base_model, "Samll/qwen3-4b-thinking-l1")Paper: Efficient Reasoning Training Does Not Always Harm CoT Faithfulness and Monitorability
Collection: Efficient Reasoning: CoT Faithfulness & Monitorability
Method: L1 / LCPO-Exact (Aggarwal & Welleck, 2025)
| Base model | Qwen/Qwen3-4B-Thinking-2507 |
| Method | L1 / LCPO-Exact |
| Released as | LoRA adapter, step 200 |
| n_min / n_max / alpha | 100 / 4,000 / 0.0003 |
| LoRA | rank 16, alpha 32, dropout 0, all attention and MLP projections |
| Optimiser | AdamW, learning rate 1e-5, KL coefficient 0 |
| Batch | 32 prompts x 16 generations |
| Sampling | temperature 0.8, top-p 0.95 |
| Data | numina_amc_aime (PRIME Eurus-2-RL-Data), ~2,223 prompts |
| Checkpoint selection | AIME22, every 50 steps |
| Hardware | 2x NVIDIA GH200 |
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-4B-Thinking-2507", torch_dtype="auto", device_map="auto")
tok = AutoTokenizer.from_pretrained("Qwen/Qwen3-4B-Thinking-2507")
model = PeftModel.from_pretrained(base, "Samll/qwen3-4b-thinking-l1")
To set a budget, append the instruction to the prompt, e.g. "... Think for 1000 tokens." (the paper uses 512, 1000 and 3000).
Base model
Qwen/Qwen3-4B-Thinking-2507