logan7000/cogrpo-homo-llama31-8b-math345-groupB-llama31-8b-b-end Reinforcement Learning • 266k • Updated Aug 13 • 19
logan7000/cogrpo-homo-llama31-8b-math345-groupB-llama31-8b-b-best Reinforcement Learning • 266k • Updated Aug 13 • 12
logan7000/cogrpo-homo-llama31-8b-math345-groupA-llama31-8b-a-end Reinforcement Learning • 266k • Updated Aug 13 • 23
logan7000/cogrpo-homo-llama31-8b-math345-groupA-llama31-8b-a-best Reinforcement Learning • 266k • Updated Aug 13 • 11
logan7000/mllm-mmr1-gt-internvl35-2b-full-best-s440 Reinforcement Learning • 2B • Updated Aug 13 • 10
logan7000/cogrpo-w2s-qwen25-3b-x-qwen25-7b-math345-groupB-qwen25-7b-end Reinforcement Learning • 333k • Updated Aug 12 • 12 • 1
logan7000/cogrpo-w2s-qwen25-3b-x-qwen25-7b-math345-groupB-qwen25-7b-best Reinforcement Learning • 333k • Updated Aug 12 • 30
logan7000/cogrpo-w2s-qwen25-3b-x-qwen25-7b-math345-groupA-qwen25-3b-end Reinforcement Learning • 242k • Updated Aug 12 • 8
logan7000/cogrpo-w2s-qwen25-3b-x-qwen25-7b-math345-groupA-qwen25-3b-best Reinforcement Learning • 242k • Updated Aug 12 • 9
logan7000/cogrpo-n3-strict-qwen25-3b-x-llama32-3b-x-qwen3-1p7b-math345-groupC-qwen3-end Reinforcement Learning • 2B • Updated Aug 12 • 6
logan7000/cogrpo-n3-strict-qwen25-3b-x-llama32-3b-x-qwen3-1p7b-math345-groupC-qwen3-best Reinforcement Learning • 2B • Updated Aug 12 • 9 • 1
logan7000/cogrpo-n3-strict-qwen25-3b-x-llama32-3b-x-qwen3-1p7b-math345-groupB-llama32-end Reinforcement Learning • 175k • Updated Aug 12 • 14 • 1
logan7000/cogrpo-n3-strict-qwen25-3b-x-llama32-3b-x-qwen3-1p7b-math345-groupB-llama32-best Reinforcement Learning • 175k • Updated Aug 12 • 8
logan7000/cogrpo-n3-strict-qwen25-3b-x-llama32-3b-x-qwen3-1p7b-math345-groupA-qwen25-end Reinforcement Learning • 242k • Updated Aug 12 • 5 • 1
logan7000/cogrpo-n3-strict-qwen25-3b-x-llama32-3b-x-qwen3-1p7b-math345-groupA-qwen25-best Reinforcement Learning • 242k • Updated Aug 12 • 3
logan7000/cogrpo-n3-ring-qwen25-3b-x-llama32-3b-x-qwen3-1p7b-math345-groupC-qwen3-end Reinforcement Learning • 2B • Updated Aug 10 • 11
logan7000/cogrpo-n3-ring-qwen25-3b-x-llama32-3b-x-qwen3-1p7b-math345-groupC-qwen3-best Reinforcement Learning • 2B • Updated Aug 10 • 17
logan7000/cogrpo-n3-ring-qwen25-3b-x-llama32-3b-x-qwen3-1p7b-math345-groupB-llama32-end Reinforcement Learning • 175k • Updated Aug 10 • 12
logan7000/cogrpo-n3-ring-qwen25-3b-x-llama32-3b-x-qwen3-1p7b-math345-groupB-llama32-best Reinforcement Learning • 175k • Updated Aug 10 • 10
logan7000/cogrpo-n3-ring-qwen25-3b-x-llama32-3b-x-qwen3-1p7b-math345-groupA-qwen25-end Reinforcement Learning • 242k • Updated Aug 10 • 14
logan7000/cogrpo-n3-ring-qwen25-3b-x-llama32-3b-x-qwen3-1p7b-math345-groupA-qwen25-best Reinforcement Learning • 242k • Updated Aug 10 • 9
logan7000/cogrpo-n3-randompeer-qwen25-3b-x-llama32-3b-x-qwen3-1p7b-math345-groupC-qwen3-end Reinforcement Learning • 2B • Updated Aug 10 • 16
logan7000/cogrpo-n3-randompeer-qwen25-3b-x-llama32-3b-x-qwen3-1p7b-math345-groupC-qwen3-best Reinforcement Learning • 2B • Updated Aug 10 • 15
logan7000/cogrpo-n3-randompeer-qwen25-3b-x-llama32-3b-x-qwen3-1p7b-math345-groupB-llama32-end Reinforcement Learning • 175k • Updated Aug 10 • 14
logan7000/cogrpo-n3-randompeer-qwen25-3b-x-llama32-3b-x-qwen3-1p7b-math345-groupB-llama32-best Reinforcement Learning • 175k • Updated Aug 10 • 9
logan7000/cogrpo-n3-randompeer-qwen25-3b-x-llama32-3b-x-qwen3-1p7b-math345-groupA-qwen25-end Reinforcement Learning • 242k • Updated Aug 10 • 10