🔄 In a Training Loop
haodi lei
bingyang-lei
AI & ML interests
None yet
Recent Activity
upvoted a paper 8 days ago
Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening upvoted a paper 21 days ago
Rethinking On-Policy Distillation of Large Language Models II: One Training Example updated a collection 22 days ago
Draft-OPD