-
shufanshen/Qwen3-4B-GRPO-DeepMath-50-steps
Reinforcement Learning • 4B • Updated • 15 -
shufanshen/Qwen3-4B-GRPO-DeepMath-100-steps
Reinforcement Learning • 4B • Updated • 13 -
shufanshen/Qwen3-4B-GRPO-DeepMath-150-steps
Reinforcement Learning • 4B • Updated • 18 -
shufanshen/Qwen3-8B-GRPO-DeepMath-50-steps
Reinforcement Learning • 8B • Updated • 15
Shufan Shen
shufanshen
AI & ML interests
Interpretable machine learning, parameter-efficient fine-tuning.
Recent Activity
authored a paper 1 day ago
On-Policy Parameter Update Direction Underlies Generalization in LLM Post-Training authored a paper 1 day ago
Edit Less, Achieve More: Dynamic Sparse Neuron Masking for Lifelong Knowledge Editing in LLMs liked a model 16 days ago
Edge0/Edge0-35B-A3B-preview