DiffGI: Differentiable Geometry Images for High-Fidelity Thin-Shell 3D Generation Paper • 2607.13365 • Published 12 days ago • 19
AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report Paper • 2607.18367 • Published 6 days ago • 55
FlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality Mimicry Paper • 2607.18227 • Published 7 days ago • 47
ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU Paper • 2607.19191 • Published 5 days ago • 293
Appearance Pointers -- Multimodal Region Control of Diffusion Transformers Paper • 2607.19344 • Published 6 days ago • 4
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing Paper • 2607.19064 • Published 6 days ago • 68
DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines Paper • 2607.16617 • Published 9 days ago • 135
S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation Paper • 2607.15686 • Published 10 days ago • 16
SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration Paper • 2607.15257 • Published 11 days ago • 70
Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget Paper • 2607.13125 • Published 9 days ago • 136
CtrlVTON: Controllable Virtual Try-On via Visual-Instance-Prompt Segmentation Paper • 2607.09362 • Published 17 days ago • 12
Latent-Identity Tuning in Text-to-Image Personalization Models Paper • 2607.11885 • Published 14 days ago • 14
Flash-BoN: Instant Drafts for Inference-Time Scaling in Diffusion Models Paper • 2607.04461 • Published 22 days ago • 11
Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Paper • 2607.11643 • Published 14 days ago • 44
Know Before Fix: QA-Driven Repository Knowledge Acquisition for Software Issue Resolution Paper • 2607.11111 • Published 14 days ago • 23
Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation Paper • 2607.11886 • Published 14 days ago • 84
Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation Paper • 2607.05382 • Published 18 days ago • 87
ABot-N1: Toward a General Visual Language Navigation Foundation Model Paper • 2607.10383 • Published 13 days ago • 102