Stella Li PRO
stellalisy
AI & ML interests
None yet
Organizations
Spurious Rewards
Spurious Rewards: Rethinking Training Signals in RLVR
-
stellalisy/rethink_rlvr_reproduce-ground_truth-qwen2.5_math_7b-lr5e-7-kl0.00-step50
Text Generation • 8B • Updated • 4 -
stellalisy/rethink_rlvr_reproduce-ground_truth-qwen2.5_math_7b-lr5e-7-kl0.00-step100
Text Generation • 8B • Updated • 8 -
stellalisy/rethink_rlvr_reproduce-ground_truth-qwen2.5_math_7b-lr5e-7-kl0.00-step150
Text Generation • 8B • Updated • 204 -
stellalisy/rethink_rlvr_reproduce-majority_vote-qwen2.5_math_7b-lr5e-7-kl0.00-step50
Text Generation • 8B • Updated • 6
Cognitive Foundations
Personalized Reasoning
Spurious Rewards
Spurious Rewards: Rethinking Training Signals in RLVR
-
stellalisy/rethink_rlvr_reproduce-ground_truth-qwen2.5_math_7b-lr5e-7-kl0.00-step50
Text Generation • 8B • Updated • 4 -
stellalisy/rethink_rlvr_reproduce-ground_truth-qwen2.5_math_7b-lr5e-7-kl0.00-step100
Text Generation • 8B • Updated • 8 -
stellalisy/rethink_rlvr_reproduce-ground_truth-qwen2.5_math_7b-lr5e-7-kl0.00-step150
Text Generation • 8B • Updated • 204 -
stellalisy/rethink_rlvr_reproduce-majority_vote-qwen2.5_math_7b-lr5e-7-kl0.00-step50
Text Generation • 8B • Updated • 6