OVEN GRPO sweep (Qwen3-VL-4B) Collection GRPO LoRA fine-tunes of Qwen3-VL-4B on OVEN (verl), from my MSc thesis. Mostly negative results across reward, prompt, and data axes. • 14 items • Updated 8 days ago
Open-world classification — GRPO experiments Collection GRPO training datasets for open-world image classification (Oxford-IIIT Pet): direct and structured-reasoning variants. From my lmms-owc work. • 2 items • Updated 8 days ago
OVEN Qwen3-VL — evaluation traces Collection Qwen3-VL (2B/4B/8B/32B) rollouts + LM-judge verdicts on OVEN, from my MSc thesis. • 1 item • Updated 8 days ago
OVEN GRPO sweep (Qwen3-VL-4B) Collection GRPO LoRA fine-tunes of Qwen3-VL-4B on OVEN (verl), from my MSc thesis. Mostly negative results across reward, prompt, and data axes. • 14 items • Updated 8 days ago