See2Think: Do Multimodal Models Really Use Intermediate Visual States? Paper • 2607.26769 • Published 3 days ago • 21
DataEvolver: Self-Evolving Multi-Agent Data Construction for Text-Rich Image Generation Paper • 2606.31537 • Published Jun 30 • 29
VCode: a Multimodal Coding Benchmark with SVG as Symbolic Visual Representation Paper • 2511.02778 • Published Nov 4, 2025 • 104