๐ In a Training Loop
fafa
kunkun0919
AI & ML interests
LLM
Recent Activity
upvoted a paper 11 days ago
COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization upvoted a paper 21 days ago
Rethinking On-Policy Distillation of Large Language Models II: One Training Example authored a paper about 1 month ago
SPOT: Sparse Probing and Outcome Calibration for On-Policy DistillationOrganizations
None yet