electricsheepafrica/africa-ghana-national-fire-outbreaks-2000-2012-36a43e7d Viewer • Updated 10 days ago • 292 • 49 • 1
The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning Paper • 2606.29526 • Published Jun 28 • 170
MemoBench: Benchmarking World Modeling in Dynamically Changing Environments Paper • 2606.27537 • Published Jun 25 • 6
The Flip Side of RLHF: On-Policy Feedback for Reward Model Self-Supervised Improvement Paper • 2605.30888 • Published May 29 • 10
Towards Verifiable Multimodal Deep Research: A Multi-Agent Harness for Interleaved Report Generation Paper • 2605.29861 • Published May 28 • 16
SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue Paper • 2605.30993 • Published May 29 • 62
Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players Paper • 2605.28816 • Published May 27 • 433