Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement Paper • 2608.31046 • Published 29 days ago • 97
Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements Paper • 2608.17310 • Published Aug 18 • 110
yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1-GGUF Text Generation • 12B • Updated Jun 19 • 458k • 2.91k
Running on Zero Agents Featured 2.27k Qwen3-TTS Demo 🎙 2.27k Generate speech from text with voice design, cloning, or presets
Stabilizing Reinforcement Learning with LLMs: Formulation and Practices Paper • 2512.01374 • Published Dec 1, 2025 • 109