🔄 In a Training Loop
haodi lei
bingyang-lei
AI & ML interests
None yet
Recent Activity
upvoted a paper 17 days ago
Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening upvoted a paper about 1 month ago
Rethinking On-Policy Distillation of Large Language Models II: One Training Example updated a collection about 1 month ago
Draft-OPD