UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement Paper • 2609.38721 • Published 4 days ago • 274
TGRL: Temperature-Grouped Reinforcement Learning for Efficient Exploration in LLMs Paper • 2609.33589 • Published 7 days ago • 8
FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders Paper • 2609.31620 • Published 9 days ago • 159
RRSI: Regularized Recursive Self-Improvement of Agent Harnesses Paper • 2609.24972 • Published 13 days ago • 221
Small RL Controller, Large Language Model: RL-Guided Adaptive Sampling for Test-Time Scaling Paper • 2606.03102 • Published Jun 2 • 14
You Only Need Minimal RLVR Training: Extrapolating LLMs via Rank-1 Trajectories Paper • 2605.21468 • Published May 20 • 51
G-Zero: Self-Play for Open-Ended Generation from Zero Data Paper • 2605.09959 • Published May 11 • 18
LLMs Improving LLMs: Agentic Discovery for Test-Time Scaling Paper • 2605.08083 • Published May 8 • 71
RRSI: Regularized Recursive Self-Improvement of Agent Harnesses Paper • 2609.24972 • Published 13 days ago • 221
Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design Paper • 2609.22086 • Published 16 days ago • 34
Dream-RSI: Recursive Self-Improvement through Evolving Worlds Paper • 2609.14858 • Published 20 days ago • 251
Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation Paper • 2609.11115 • Published 24 days ago • 174
Recursive Code World Models: Building Complex Worlds through Recursive Scene Programs Paper • 2609.11499 • Published 24 days ago • 33