shuo shen
hyperion-shuo
ยท
AI & ML interests
reinforcement learning
Recent Activity
upvoted a paper 1 day ago
Bellman Policy Optimization liked a model 4 days ago
internlm/Intern-S2-397B upvoted a paper 16 days ago
OPV: Outcome-based Process Verifier for Efficient Long Chain-of-Thought Verification