SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Paper • 2608.03092 • Published 11 days ago • 10
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Paper • 2608.03092 • Published 11 days ago • 10 • 1
DanhVuiVe/ChartQA_Benetech_PlotQa_DVQA_combined_matcha_complete Viewer • Updated Oct 29, 2024 • 535k • 96 • 2
alimama-creative/FLUX.1-dev-Controlnet-Inpainting-Beta Image-to-Image • 2B • Updated Oct 12, 2024 • 5.54k • 427
openai/whisper-large-v3-turbo Automatic Speech Recognition • 0.8B • Updated Oct 4, 2024 • 7.91M • • 3.24k