Spurious Rewards: Rethinking Training Signals in RLVR Paper • 2506.10947 • Published Jun 12, 2025 • 2
Open Character Training: Shaping the Persona of AI Assistants through Constitutional AI Paper • 2511.01689 • Published Nov 3, 2025 • 5
Meta-Reinforcement Learning with Self-Reflection for Agentic Search Paper • 2603.11327 • Published Mar 11 • 11
Learning to Solve Hard Problems in RL for LLMs by Never Giving Up Paper • 2609.13443 • Published 22 days ago • 13
The Embedder's Dilemma: LLMs Are Better, but at What Cost? Paper • 2608.12875 • Published Aug 13 • 16
The ATOM Report: Measuring the Open Language Model Ecosystem Paper • 2604.07190 • Published Apr 8 • 5
SWE-Universe: Scale Real-World Verifiable Environments to Millions Paper • 2602.02361 • Published Feb 2 • 61
Organize the Web: Constructing Domains Enhances Pre-Training Data Curation Paper • 2502.10341 • Published Feb 14, 2025 • 3
olmOCR: Unlocking Trillions of Tokens in PDFs with Vision Language Models Paper • 2502.18443 • Published Feb 25, 2025 • 13
DataDecide: How to Predict Best Pretraining Data with Small Experiments Paper • 2504.11393 • Published Apr 15, 2025 • 20
Teaching Models to Understand (but not Generate) High-risk Data Paper • 2505.03052 • Published May 5, 2025 • 6