onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction Paper • 2609.24983 • Published 3 days ago • 43
IntBMoE: Integrating Block-Level Conditioning into Expert Composition for Full-Participation Mixture-of-Experts Paper • 2609.21346 • Published 6 days ago • 99
Calibrating Teacher--Student Discrepancy for On-Policy Distillation Paper • 2609.21619 • Published 6 days ago • 13
UFO: Chain-of-Evaluation for Omni-Condition Alignment in Multi-Modal Image Generation Paper • 2609.12397 • Published 7 days ago • 43
ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement Paper • 2609.14857 • Published 10 days ago • 212
Dream-RSI: Recursive Self-Improvement through Evolving Worlds Paper • 2609.14858 • Published 10 days ago • 245
Discovery Foundation Models: Toward Open-Ended Discovery Intelligence Paper • 2609.15973 • Published 10 days ago • 33
An Empirical Study on Zero-Data Bootstrapping for Conversational Recommender Systems Paper • 2504.15476 • Published 27 days ago • 2
Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation Paper • 2609.08084 • Published 16 days ago • 70
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness Paper • 2609.08183 • Published 16 days ago • 172
Bilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-Agent LLM Systems Paper • 2609.02750 • Published 22 days ago • 144
Rethinking On-Policy Distillation of Large Language Models II: One Training Example Paper • 2609.04172 • Published 21 days ago • 101
A Glance Is All You Need: Single-Pass Fine-Grained Image Captioning with SimLoss Paper • 2609.00591 • Published 23 days ago • 18
GenFirst: Generation Before Reconstruction for Stable End-to-End Latent Generative Modeling Paper • 2608.29335 • Published 26 days ago • 73