zhaoyibo
ybyby624
AI & ML interests
RL, LLM-based Agents
Recent Activity
upvoted a paper 10 days ago
Not Every Token Is Worth Distilling: Selective Supervision for Direct-OPD updated a collection 30 days ago
Search Agent Review updated a model 30 days ago
ybyby624/Qwen3-8B-TreeGRPO-Hermes9k_clean-1epoch-160steps-0509Organizations
None yet