A New Role for Relevance: Guiding Corpus Interaction in Agentic Search Paper • 2607.24223 • Published 3 days ago • 86
Pass the Baton: Trajectory-Relayed On-Policy Distillation Paper • 2607.26057 • Published 2 days ago • 26
view article Article The Engineering Handbook for GRPO + LoRA with Verl: Training Qwen2.5 on Multi-GPU Weyaxi • Jan 2 • 24
Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation Paper • 2607.24731 • Published 3 days ago • 72
From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search Paper • 2607.24280 • Published 3 days ago • 80
Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems Paper • 2607.21503 • Published 7 days ago • 23
Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Paper • 2607.21653 • Published 8 days ago • 30
Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills Paper • 2607.22529 • Published 6 days ago • 44
Running Featured 81 QED-Nano: Teaching a Tiny Model to Prove Hard Theorems 📝 81 Who needs 1T parameters? Olympiad proofs with a 4B model
WTF GENIUS PAPERS Collection Papers that made me appreciate my major and my life a little more. obs=Observation, innov=Innovation. Most papers are abt improving tiny models. • 233 items • Updated about 7 hours ago • 54
AREX: Towards a Recursively Self-Improving Agent for Deep Research Paper • 2607.21461 • Published 7 days ago • 149
SWE-Pruner Pro: The Coder LLM Already Knows What to Prune Paper • 2607.18213 • Published 10 days ago • 78
Understanding Reasoning from Pretraining to Post-Training Paper • 2607.16097 • Published 13 days ago • 28
DiffGI: Differentiable Geometry Images for High-Fidelity Thin-Shell 3D Generation Paper • 2607.13365 • Published 15 days ago • 20