wangzixuan
wangzx1210
AI & ML interests
None yet
Recent Activity
upvoted a paper 7 days ago
RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning upvoted a paper 28 days ago
TTPO: Test-Time Policy Optimization liked a dataset 30 days ago
xiamoent/Agent-G2-ALFWorld-Webshop-sft-data