๐ In a Training Loop
haodi lei
bingyang-lei
AI & ML interests
None yet
Recent Activity
upvoted a paper 10 days ago
Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening upvoted a paper 23 days ago
Rethinking On-Policy Distillation of Large Language Models II: One Training Example updated a collection 24 days ago
Draft-OPD