IDEALLab/Qwen2.5-Coder-14B-Instruct-GRPO-TSP-Hero-seed101 Reinforcement Learning • 15B • Updated 29 days ago • 27
IDEALLab/Qwen2.5-Coder-14B-Instruct-GRPO-TSP-Hero-seed202 Reinforcement Learning • 15B • Updated 29 days ago • 61
IDEALLab/Qwen2.5-Coder-14B-Instruct-GRPO-TSP-Hero-seed303 Reinforcement Learning • 15B • Updated 29 days ago • 22
IDEALLab/Qwen2.5-Coder-14B-Instruct-GRPO-SDS-NoHypothesize-step90-seed101 Reinforcement Learning • 15B • Updated 29 days ago • 42
IDEALLab/Qwen2.5-Coder-14B-Instruct-GRPO-SDS-NoHypothesize-step90-seed202 Reinforcement Learning • 15B • Updated 29 days ago • 24
IDEALLab/Qwen2.5-Coder-14B-Instruct-GRPO-SDS-NoHypothesize-step90-seed303 Reinforcement Learning • 15B • Updated 29 days ago • 30