qwen3-8b-thinkprune-b3000

Paper: Efficient Reasoning Training Does Not Always Harm CoT Faithfulness and Monitorability

Collection: Efficient Reasoning: CoT Faithfulness & Monitorability

Method: ThinkPrune (Hou et al., 2026)

Hyperparameters

Base model Qwen/Qwen3-8B
Method ThinkPrune
Released as merged full model, step 200
Completion budget L 3,000 tokens
LoRA rank 16, alpha 32, dropout 0, all attention and MLP projections
Optimiser AdamW, learning rate 1e-5, KL coefficient 0
Batch 32 prompts x 16 generations
Sampling temperature 0.8, top-p 0.95
Data numina_amc_aime (PRIME Eurus-2-RL-Data), ~2,223 prompts
Checkpoint selection AIME22, every 50 steps
Hardware 2x NVIDIA GH200

Budget is 3,000 rather than 4,000 because Qwen3-8B's average base CoT on the training data is already below 4,000 tokens.

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "Samll/qwen3-8b-thinkprune-b3000"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, torch_dtype="auto", device_map="auto")
Downloads last month
172
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Samll/qwen3-8b-thinkprune-b3000

Finetuned
Qwen/Qwen3-8B
Finetuned
(2169)
this model

Collection including Samll/qwen3-8b-thinkprune-b3000