IDEALLab/Qwen2.5-Coder-14B-Instruct-GRPO-SDS-Ablation-Prompt-seed303 Reinforcement Learning • 15B • Updated May 13 • 5
IDEALLab/Qwen2.5-Coder-14B-Instruct-GRPO-SDS-Minimalist-seed101 Reinforcement Learning • 15B • Updated May 13 • 5
IDEALLab/Qwen2.5-Coder-14B-Instruct-GRPO-SDS-Minimalist-seed202 Reinforcement Learning • 15B • Updated May 13 • 2
IDEALLab/Qwen2.5-Coder-14B-Instruct-GRPO-SDS-Minimalist-seed303 Reinforcement Learning • 15B • Updated May 13 • 3
IDEALLab/Qwen2.5-Coder-14B-Instruct-GRPO-TSP-Hero-seed101 Reinforcement Learning • 15B • Updated Sep 6 • 13
IDEALLab/Qwen2.5-Coder-14B-Instruct-GRPO-TSP-Hero-seed202 Reinforcement Learning • 15B • Updated Sep 6 • 51
IDEALLab/Qwen2.5-Coder-14B-Instruct-GRPO-TSP-Hero-seed303 Reinforcement Learning • 15B • Updated Sep 6 • 12
IDEALLab/Qwen2.5-Coder-14B-Instruct-GRPO-SDS-NoHypothesize-step90-seed101 Reinforcement Learning • 15B • Updated Sep 6 • 33
IDEALLab/Qwen2.5-Coder-14B-Instruct-GRPO-SDS-NoHypothesize-step90-seed202 Reinforcement Learning • 15B • Updated Sep 6 • 14
IDEALLab/Qwen2.5-Coder-14B-Instruct-GRPO-SDS-NoHypothesize-step90-seed303 Reinforcement Learning • 15B • Updated Sep 6 • 18