Explore Broadly, Reason Sharply: Push Small Models toward the Frontier via Sampling
Abstract
Power-sharpened sampling is an inference-time alternative to reinforcement-learning (RL) post-training for enhancing reasoning in large language models (LLMs). High-probability sequences are amplified under the base model without parameter updates or external rewards, avoiding the costly optimization and jagged generalization of RL. However, this approach faces a fundamental exploration--exploitation trade-off, as % strong sharpening restricts exploration, trapping samplers in plausible but incorrect reasoning trajectories, whereas weak sharpening leaves the answer distribution diffuse. To resolve this trade-off, we introduce Parallel Power Tempering (PPT), instantiating power-sharpened LLM sampling via parallel tempering. Running multiple interacting replicas in parallel at different sharpening levels allows lower-power replicas to explore diverse reasoning trajectories and higher-power chains to further exploit higher-likelihood responses favored by the sharpened target. Specifically, we tailor to inference-time sampling by mitigating a truncation bias, identified in prior power samplers, and investigate effective swap strategies under finite memory and compute budgets. Extensive experimentation shows that substantially improves single-chain power-sharpened sampling and outperforms RL-post-trained models, producing higher-quality reasoning traces and even achieving performance comparable to frontier models.
Community
A 9B model could reach GPT-5 / Opus-4.5-level reasoning 🧠, nearly free ⚡🆓, on your own desktop 💻.
No additional post-training. Much less jagged 🧩 generalization.
This was my intern Panagiotis’ summer project. Still a long way to go on engineering, but I strongly believe this direction will keep getting better.
Small models may have a lot more intelligence 🧠 inside them than we think.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- ReSight-SMC: Two-Stage Power Sampling via Island SMC with Visual Scouts (2026)
- Sampling via Decision-Flow: Training-Free Extraction of Improved Latent Reasoning Paths in Large Language Models (2026)
- Beam Search as Test-Time Self-Distillation via Counterfactual Contexts (2026)
- Chopthin-Consensus Power Sampling: A Diversity-Preserving Approach to LLM Decoding (2026)
- More Correct Mass, Worse Answers: Why Power Sampling Can Fail and How to Fix It (2026)
- TGRL: Temperature-Grouped Reinforcement Learning for Efficient Exploration in LLMs (2026)
- SpecRoll: Fast-Slow Verifier-Feedback Adaptation for Speculative Reinforcement Learning Rollouts (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.38104 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper