WeVisDoc: From Coverage to Capability for Robust End-to-End Document Parsing Paper • 2609.20423 • Published 9 days ago • 48
Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds Paper • 2608.23383 • Published Aug 24 • 19
PaDoc: Layout-Grounded Parallel Decoding for Document Parsing Paper • 2608.06146 • Published Aug 6 • 24
Generation Enhances Understanding in Unified Multimodal Models via Multi-Representation Generation Paper • 2601.21406 • Published Jan 29 • 6
Learning Cross-View Object Correspondence via Cycle-Consistent Mask Prediction Paper • 2602.18996 • Published Feb 22 • 16
Skywork-Unipic3 Collection Unified Multi-Image Composition with Sequence Modeling • 9 items • Updated Mar 2 • 13