ROSS: Relearning from Self-Generated Rollouts through Selective Supervision Paper • 2609.35954 • Published 7 days ago • 49
ROSS: Relearning from Self-Generated Rollouts through Selective Supervision Paper • 2609.35954 • Published 7 days ago • 49
ROSS: Relearning from Self-Generated Rollouts through Selective Supervision Paper • 2609.35954 • Published 7 days ago • 49
Skill2Env: Capability-Oriented Environment Synthesis from Skills for General Agents Paper • 2609.33772 • Published 8 days ago • 35
CompoWorld: Compositional Environment Scaling for General Agents Paper • 2609.33665 • Published 8 days ago • 42
Learning to Learn from Context: Synthetic Training from Perturbed Public Documents Paper • 2609.33642 • Published 8 days ago • 32
Verify Before You Distill: Prompt-Level Teacher Gating for On-Policy Distillation Paper • 2609.02998 • Published Sep 2 • 17
Verify Before You Distill: Prompt-Level Teacher Gating for On-Policy Distillation Paper • 2609.02998 • Published Sep 2 • 17
Verify Before You Distill: Prompt-Level Teacher Gating for On-Policy Distillation Paper • 2609.02998 • Published Sep 2 • 17
D$^3$-MOPD: Adaptive Dynamic Domain ScheDuling for Efficient Multi-Teacher Distillation Paper • 2608.24987 • Published Aug 25 • 28
D^3-MOPD: Adaptive Dynamic Domain ScheDuling for Efficient Multi-Teacher Distillation Paper • 2608.24987 • Published Aug 25 • 28
D^3-MOPD: Adaptive Dynamic Domain ScheDuling for Efficient Multi-Teacher Distillation Paper • 2608.24987 • Published Aug 25 • 28