wincentIsMe/Qwen3-VL-2B-Instruct-SFT-VSI590k-32f-384px-tune-vision-tower Image-Text-to-Text • 2B • Updated Aug 17 • 13
wincentIsMe/Qwen3-VL-2B-Instruct-SFT-VSI590k-32f-384px-tune-vision-tower Image-Text-to-Text • 2B • Updated Aug 17 • 13
ReVSI: Rebuilding Visual Spatial Intelligence Evaluation for Accurate Assessment of VLM 3D Reasoning Paper • 2604.24300 • Published Apr 27 • 68
HiFi-Inpaint: Towards High-Fidelity Reference-Based Inpainting for Generating Detail-Preserving Human-Product Images Paper • 2603.02210 • Published Mar 2 • 30
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Paper • 2403.04132 • Published Mar 7, 2024 • 41
Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding Paper • 2504.10465 • Published Apr 14, 2025 • 27