IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation Paper • 2606.24849 • Published Jun 23 • 19
MVTrack4Gen: Multi-View Point Tracking as Geometric Supervision for 4D Video Generation Paper • 2606.26087 • Published Jun 24 • 35
EvoEmbedding: Evolvable Representations for Long-Context Retrieval and Agentic Memory Paper • 2606.21649 • Published Jun 19 • 35
Qwen-AgentWorld: Language World Models for General Agents Paper • 2606.24597 • Published Jun 23 • 153