Running 124 The ultimate guide to multi-harness RL π 124 Train open models with RL inside real agent harnesses
jucamohedano/rlm-laguna-xs2-dense-3b-v0.1 Text Generation β’ 3B β’ Updated 9 days ago β’ 403 β’ 1
jucamohedano/rlm-laguna-xs2-dense-3b-v0.1 Text Generation β’ 3B β’ Updated 9 days ago β’ 403 β’ 1
OVEN GRPO sweep (Qwen3-VL-4B) Collection GRPO LoRA fine-tunes of Qwen3-VL-4B on OVEN (verl), from my MSc thesis. Mostly negative results across reward, prompt, and data axes. β’ 14 items β’ Updated Jul 29
Open-world classification β GRPO experiments Collection GRPO training datasets for open-world image classification (Oxford-IIIT Pet): direct and structured-reasoning variants. From my lmms-owc work. β’ 2 items β’ Updated Jul 29
OVEN Qwen3-VL β evaluation traces Collection Qwen3-VL (2B/4B/8B/32B) rollouts + LM-judge verdicts on OVEN, from my MSc thesis. β’ 1 item β’ Updated Jul 29
OVEN GRPO sweep (Qwen3-VL-4B) Collection GRPO LoRA fine-tunes of Qwen3-VL-4B on OVEN (verl), from my MSc thesis. Mostly negative results across reward, prompt, and data axes. β’ 14 items β’ Updated Jul 29