Frozen Models, Evolving Expertise: Model-Agnostic Learning from Deployment Experience for Multimodal Medical AI Paper • 2610.09146 • Published 6 days ago • 6
Agentic RAG Evaluation: Budget Allocation Across Questions, Trajectories, and Reads Paper • 2610.05034 • Published 8 days ago • 19
Adaptive Reward Routing: Dynamic Multi-Reward Optimization for Joint Audio-Video Diffusion via Forward-Process RL Paper • 2609.37200 • Published 13 days ago • 140
Online Draft Co-Training for Speculative Decoding in Large-Scale, Long-Context RL Post-Training Paper • 2609.07108 • Published Sep 7 • 36
WorldReward: Reward Modeling for Camera-Conditioned World Models Paper • 2609.03952 • Published Sep 3 • 28