Wenboz/SACD-Qwen2.5-3B-ALFWorld-k1-tau0.5-beta1.0-plain-pipeline Reinforcement Learning • 3B • Updated Jun 10 • 6
Wenboz/SACD-Qwen2.5-3B-ALFWorld-k1-tau0.5-beta1.0-plain-pipeline Reinforcement Learning • 3B • Updated Jun 10 • 6
Wenboz/SACD-Qwen2.5-3B-ALFWorld-k1-tau0.75-beta1.0-plain-pipeline Reinforcement Learning • 3B • Updated Jun 10 • 3 • 1
Wenboz/SACD-Qwen2.5-3B-ALFWorld-k1-tau0.75-beta1.0-plain-pipeline Reinforcement Learning • 3B • Updated Jun 10 • 3 • 1
Supervised Fine-Tuning versus Reinforcement Learning: A Study of Post-Training Methods for Large Language Models Paper • 2603.13985 • Published Mar 14 • 11