SDPO under Continual Learning
Meng Wang
Moenupa
AI & ML interests
MLLM Post-Training & Alignment
Recent Activity
updated a collection about 9 hours ago
MedEvalScope updated a collection 3 days ago
MedEvalScope updated a dataset 3 days ago
CAIR-M3LLM/PMC-VQA-2Organizations
models 37
Moenupa/Qwen3-4B-Instruct-2507-SDPOthink5-Math
Updated
Moenupa/Qwen3-4B-Thinking-2507-SDPOthink5-Tool
Updated
Moenupa/Qwen3-4B-Instruct-2507-SDPOthink5-Tool
Updated
Moenupa/Qwen3-4B-Thinking-2507-SDPOthink5-Chem
Updated
Moenupa/cuda-compat-13-2
Updated
Moenupa/Qwen3-4B-Thinking-2507-SDPO1restart-Math
Updated
Moenupa/Qwen3-4B-Thinking-2507-SDPO0restart-Math
Updated
Moenupa/Qwen3-4B-Thinking-2507-SDPO5restart-Math
Updated
Moenupa/Qwen3-4B-Thinking-2507-GRPO-Code
Updated
Moenupa/Qwen3-4B-Instruct-2507-SDPOthink-Tool
Updated
datasets 10
Moenupa/Dolci-Think-RL-7B
Viewer • Updated • 72.2k • 97
Moenupa/verl
Viewer • Updated • 216k • 262
Moenupa/DAIR
Viewer • Updated • 361k • 181
Moenupa/Vision-Flan-191-1k
Viewer • Updated • 186k • 284
Moenupa/MSVQA
Viewer • Updated • 39.2k • 29
Moenupa/Domain8k-ICL
Viewer • Updated • 37.8k • 96
Moenupa/MC-Bench-VQA-ICL
Viewer • Updated • 1k • 41
Moenupa/MC-Bench-VQA
Viewer • Updated • 2k • 49
Moenupa/MC-Bench
Viewer • Updated • 2k • 30
Moenupa/MemoryBench-Full
Viewer • Updated • 8.99k • 682 • 1