HOI-R1 Exploring the Potential of Multimodal Large Language Models for Human-Object Interaction Detection. thxplz/HOI-R1_Qwen2.5-VL-3B-Instruct Image-Text-to-Text • 4B • Updated Dec 27, 2025 • 449 • 1 thxplz/HOI-R1_Qwen3-VL-4B-Instruct Image-Text-to-Text • 4B • Updated May 22 • 6 • 1
HOI-R1 Exploring the Potential of Multimodal Large Language Models for Human-Object Interaction Detection. thxplz/HOI-R1_Qwen2.5-VL-3B-Instruct Image-Text-to-Text • 4B • Updated Dec 27, 2025 • 449 • 1 thxplz/HOI-R1_Qwen3-VL-4B-Instruct Image-Text-to-Text • 4B • Updated May 22 • 6 • 1