1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation Paper • 2609.24432 • Published 12 days ago • 16
Self-Rewarding Rubric-Based Reinforcement Learning for Open-Ended Reasoning Paper • 2509.25534 • Published Sep 19, 2025 • 3
Learning to Align, Aligning to Learn: A Unified Approach for Self-Optimized Alignment Paper • 2508.07750 • Published Aug 11, 2025 • 21
view article Article 🐺🐦⬛ LLM Comparison/Test: 25 SOTA LLMs (including QwQ) through 59 MMLU-Pro CS benchmark runs wolfram • Dec 4, 2024 • 80