UOPD: Uncertainty-Aware Intervention for On-Policy Distillation of Multi-Turn Agents Paper • 2609.34036 • Published 7 days ago • 4
UOPD Collection Uncertainty-Aware Intervention for On-Policy Distillation of Multi-Turn Agents • 4 items • Updated 7 days ago
Supervised Fine-Tuning versus Reinforcement Learning: A Study of Post-Training Methods for Large Language Models Paper • 2603.13985 • Published Mar 14 • 11