An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM Reasoning Paper • 2609.35505 • Published 2 days ago • 10
Post-Training Language Models for Gold-Medal Performance in Coding Competitions Paper • 2609.02849 • Published 28 days ago • 13
nvidia/NVIDIA-Nemotron-Labs-3-Competitive-Coding-550B-A55B-NVFP4 Text Generation • 335B • Updated 13 days ago • 34 • 8
QwenGyre: An Elastic Reinforcement Learning Framework for Training xLong-Horizon Agents Paper • 2609.33848 • Published 3 days ago • 27
bartowski/JetBrains-Qwen3.8-3.6-27B-blend-GGUF Image-Text-to-Text • 27B • Updated 6 days ago • 2.83k • 7
Running on CPU Upgrade 3 DecisionBench Leaderboard 🧭 3 Explore and compare AI model benchmarks online