GRPO training datasets for open-world image classification (Oxford-IIIT Pet): direct and structured-reasoning variants. From my lmms-owc work.
Juan CM
jucamohedano
AI & ML interests
AI Systems MSc at Trento 🚀🤖
Recent Activity
updated a collection 7 days ago
OVEN GRPO sweep (Qwen3-VL-4B) updated a dataset 7 days ago
jucamohedano/oven-grpo-training-data published a dataset 7 days ago
jucamohedano/oven-grpo-training-dataOrganizations
OVEN GRPO sweep (Qwen3-VL-4B)
GRPO LoRA fine-tunes of Qwen3-VL-4B on OVEN (verl), from my MSc thesis. Mostly negative results across reward, prompt, and data axes.
-
jucamohedano/qwen3-vl-4b-oven-grpo-traversal
Image-Text-to-Text • Updated • 46 -
jucamohedano/qwen3-vl-4b-oven-grpo-aggregation
Image-Text-to-Text • Updated • 50 -
jucamohedano/qwen3-vl-4b-oven-grpo-aggregation-prelim
Image-Text-to-Text • Updated • 48 -
jucamohedano/qwen3-vl-4b-oven-grpo-traversal-structured-exact
Image-Text-to-Text • Updated • 43
Model search via model weights
Open-world classification — GRPO experiments
GRPO training datasets for open-world image classification (Oxford-IIIT Pet): direct and structured-reasoning variants. From my lmms-owc work.
OVEN Qwen3-VL — evaluation traces
Qwen3-VL (2B/4B/8B/32B) rollouts + LM-judge verdicts on OVEN, from my MSc thesis.
OVEN GRPO sweep (Qwen3-VL-4B)
GRPO LoRA fine-tunes of Qwen3-VL-4B on OVEN (verl), from my MSc thesis. Mostly negative results across reward, prompt, and data axes.
-
jucamohedano/qwen3-vl-4b-oven-grpo-traversal
Image-Text-to-Text • Updated • 46 -
jucamohedano/qwen3-vl-4b-oven-grpo-aggregation
Image-Text-to-Text • Updated • 50 -
jucamohedano/qwen3-vl-4b-oven-grpo-aggregation-prelim
Image-Text-to-Text • Updated • 48 -
jucamohedano/qwen3-vl-4b-oven-grpo-traversal-structured-exact
Image-Text-to-Text • Updated • 43
Model merging
Model search via model weights