Instructions to use junshim/When2Think-1.5B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use junshim/When2Think-1.5B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="junshim/When2Think-1.5B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("junshim/When2Think-1.5B") model = AutoModelForCausalLM.from_pretrained("junshim/When2Think-1.5B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use junshim/When2Think-1.5B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "junshim/When2Think-1.5B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "junshim/When2Think-1.5B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/junshim/When2Think-1.5B
- SGLang
How to use junshim/When2Think-1.5B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "junshim/When2Think-1.5B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "junshim/When2Think-1.5B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "junshim/When2Think-1.5B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "junshim/When2Think-1.5B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use junshim/When2Think-1.5B with Docker Model Runner:
docker model run hf.co/junshim/When2Think-1.5B
# pip install -U transformers accelerate
# Load model directly
from transformers import AutoTokenizer, AutoModelForCausalLM
tokenizer = AutoTokenizer.from_pretrained("junshim/When2Think-1.5B")
model = AutoModelForCausalLM.from_pretrained("junshim/When2Think-1.5B", device_map="auto")
messages = [
{"role": "user", "content": "Who are you?"},
]
inputs = tokenizer.apply_chat_template(
messages,
add_generation_prompt=True,
tokenize=True,
return_dict=True,
return_tensors="pt",
).to(model.device)
outputs = model.generate(**inputs, max_new_tokens=256)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:]))When2Think
When2Think-1.5B is a post-trained hybrid reasoning model that learns both whether to reason explicitly and how much reasoning to allocate to each problem.
The model encourages direct answering on easier instances while preserving extended reasoning on harder ones. Unlike uniform length-compression methods, When2Think treats reasoning depth as an instance-adaptive resource.
Highlights
- Adaptive Think/NoThink Behavior: Learns when to answer directly and when to invoke explicit multi-step reasoning.
- Accuracy-Efficiency Trade-off: Reduces unnecessary reasoning without uniformly suppressing useful reasoning on difficult problems.
- Standalone Deployment: Requires only the released checkpoint for generation.
Model Details
Model Description
When2Think-1.5B is an RLVR-post-trained version of deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B.
The checkpoint learns two coupled decisions:
Whether to reason
- NOTHINK: Answer directly without an extended explicit reasoning trace.
- THINK: Generate explicit multi-step reasoning followed by a final answer.
How much to reason
- Within THINK, adapt generated computation to the input rather than following a fixed or uniformly compressed length target.
The post-training framework combines:
- Instance-level Difficulty-Aware Control (IDAC): Uses cached reference success and token-cost statistics to modulate a correctness-gated efficiency bonus based on generated token count.
- Batch-Wise Standardization (BWS): Converts trajectory rewards into standardized advantages for critic-free policy optimization.
- Importance-sampled THINK/NOTHINK exploration: Adopts balanced mode exploration during post-training. The auxiliary exploration policy is removed at inference time.
These mechanisms are training-time components only. The released checkpoint runs as a standalone causal language model.
Which Model Should I Use?
| Model | Whether to reason | How much to reason | Recommended use |
|---|---|---|---|
| When2Think-1.5B | Learned THINK/NOTHINK selection | IDAC + BWS | Adaptive hybrid reasoning in a single checkpoint |
| When2Think-ThinkOnly-1.5B | Always THINK, without hybrid importance sampling | IDAC + BWS | Explicit reasoning or analysis of within-THINK depth control |
The checkpoints represent different accuracy-computation operating points rather than a strict performance ordering.
Model Sources
- Paper: When2Think: Learning When and How Much to Reason
- Repository: JJunShim/When2Think
- Collection: When2Think
Uses
Direct Use
When2Think-1.5B is intended for:
- Mathematical problem solving
- Adaptive direct-answer and explicit-reasoning generation
- Research on efficient reasoning models
- Analysis of Think/NoThink mode selection
The model can be loaded as a standard causal language model using Hugging Face Transformers. No separate router, verifier, critic, difficulty estimator, reward model, or reference policy is required for inference.
Downstream Use
The checkpoint may be used as a starting point for:
- Continued post-training on verifiable reasoning tasks
- Adaptation to mathematical, symbolic, coding, or scientific reasoning
- Research on adaptive reasoning policies
- Studies of direct-answer and explicit-reasoning behavior
Additional fine-tuning may alter the learned Think/NoThink balance, response length, and reasoning-depth behavior.
How to Get Started with the Model
Use the code below to get started with the model.
pip install -U torch transformers accelerate
# optional, for fast serving
pip install vllm
by Transformers Pipeline
from transformers import pipeline
model_path = "junshim/When2Think-1.5B"
prompt = "Find the value of $x$ that satisfies the equation $4x+5 = 6x+7$."
messages = [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": prompt}
]
generator = pipeline(
"text-generation",
model=model_path,
device_map="auto",
dtype="auto"
)
outputs = generator(
messages,
max_new_tokens=512,
clean_up_tokenization_spaces=False
)
by Transformers
from accelerate import Accelerator
from transformers import AutoModelForCausalLM, AutoTokenizer
accelerator = Accelerator()
device = accelerator.device
model = AutoModelForCausalLM.from_pretrained(
model_path,
torch_dtype="auto",
device_map=device
)
tokenizer = AutoTokenizer.from_pretrained(model_path)
print(model.generation_config, device)
inputs = tokenizer.apply_chat_template(
messages,
tokenize=True,
add_generation_prompt=True,
return_dict=True,
return_tensors="pt",
).to(model.device)
input_length = inputs["input_ids"].shape[1]
max_new_tokens = model.config.max_position_embeddings - input_length
with torch.inference_mode():
outputs = model.generate(
**inputs,
max_new_tokens=max_new_tokens
)
output_text = tokenizer.batch_decode(generated_ids, skip_special_tokens=True)[0]
vLLM
vllm serve "junshim/When2Think-1.5B" --reasoning-parser deepseek_r1
Output Parsing
For decoded outputs that use <think>...</think>, the following helper separates the reasoning trace from the final response. Treat this as a convenience parser for that output format, not as a format guarantee across all serving stacks.
Evaluation
Testing Data, Factors & Metrics
See the paper for dataset versions, prompts, answer extraction, and the complete evaluation protocol.
Testing Data
The reported evaluation covers verifiable mathematical reasoning and cross-domain transfer:
- GSM-Plus
- OlympiadBench-Math
- AIME24
- AIME25
- Minerva
- MATH-500
- Stratified MMLU-Pro transfer
Factors
The analysis considers:
- Benchmark and task domain
- Problem difficulty where difficulty labels are available
- THINK/NOTHINK selection behavior
- Generated token usage
- Model scale and training variant
Metrics
- Pass@3: Sampling-based Pass@k accuracy with
k = 3, following the protocol in the paper. - Pass@1: Used for difficulty-stratified MATH-500 analysis.
- Average generated tokens: Mean generated tokens per response, reported independently of the Pass@k sampling budget.
- THINK ratio: Fraction of responses using explicit reasoning.
Key Results
| Model | MATH-500 L1 Pass@1 ↑ | Tokens ↓ | GSM-Plus Pass@3 ↑ | Tokens ↓ | AIME24 Pass@3 ↑ | Tokens ↓ | AIME25 Pass@3 ↑ | Tokens ↓ |
|---|---|---|---|---|---|---|---|---|
| R1-Distill-Qwen backbone | 92.1 | 1,199 | 79.4 | 590 | 46.0 | 14,195 | 32.0 | 12,616 |
| When2Think-1.5B | 95.8 | 619 | 85.7 | 1,052 | 56.0 | 10,236 | 40.0 | 9,549 |
| When2Think-ThinkOnly-1.5B | N/R | N/R | 86.8 | 1,652 | 57.3 | 10,046 | 40.0 | 9,846 |
When2Think-ThinkOnly-1.5Bcorresponds to the paper's w/o IS (IDAC + BWS) variant. Its MATH-500 Level 1 results are not separately reported in the paper and are therefore marked asN/R.
Summary
Key observations:
- Easy-instance allocation: On MATH-500 Level 1, When2Think improves Pass@1 from 92.1% to 95.8% while reducing average generated tokens from 1,199 to 619, a reduction of 580 tokens, or approximately 48.4%. This demonstrates that the hybrid checkpoint can favor direct answering when extended deliberation provides limited benefit.
- Task-dependent allocation: On GSM-Plus, When2Think uses more tokens than the backbone while improving Pass@3 from 79.4% to 85.7%. Together with the MATH-500 Level 1 result, this shows that adaptive allocation may reduce, preserve, or increase computation depending on the input.
- When2Think-1.5B on AIME24: Pass@3 improves by 10.0 percentage points, while average generated tokens decrease by 27.9% relative to the backbone.
- When2Think-1.5B on AIME25: Pass@3 improves by 8.0 percentage points, while average generated tokens decrease by 24.3% relative to the backbone.
- When2Think-ThinkOnly-1.5B: Achieves 57.3% Pass@3 on AIME24 and 40.0% Pass@3 on AIME25 while always using explicit reasoning.
- Within-THINK control: The strong THINK-only results show that difficulty-aware computation control contributes independently of THINK/NOTHINK mode selection.
- Different operating points: The hybrid and THINK-only checkpoints represent different accuracy-computation operating points rather than a strict performance ordering.
For complete baselines, standard deviations, difficulty-stratified analysis, and ablations, see the paper.
Citation
BibTeX:
@misc{shim2026when2think,
title = {When2Think: Learning When and How Much to Reason},
author = {Jaejun Shim and HyunJin Kim and Young Jin Kim and JinYeong Bak},
year = {2026},
eprint = {2609.19671},
url = {https://arxiv.org/abs/2609.19671}
}
More Information
- Begin of Sentence: <|begin▁of▁sentence|>
- End of Sentence: <|end▁of▁sentence|>
- Pad: <|end▁of▁sentence|>
- Begin of Thinking:
- End of Thinking:
- User Role: <|User|>
- Assistant: <|Assistant|>
- Max position embeddings (length): 131072
Model Card Contact
For questions, open an issue in JJunShim/When2Think.
- Downloads last month
- 1,067
Model tree for junshim/When2Think-1.5B
Base model
deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="junshim/When2Think-1.5B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)