Explicit Layer Modeling for Video Object Insertion and Layer Decomposition Paper • 2607.25802 • Published 8 days ago • 8
From RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image Models Paper • 2607.06553 • Published 27 days ago • 20
Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents Paper • 2607.08716 • Published 27 days ago • 15
Learning to Act under Noise: Enhancing Agent Robustness via Noisy Environments Paper • 2605.27209 • Published May 26 • 16
Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality? Paper • 2605.22109 • Published May 21 • 171
ShotStream: Streaming Multi-Shot Video Generation for Interactive Storytelling Paper • 2603.25746 • Published Mar 26 • 155
SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models Paper • 2603.16859 • Published Mar 17 • 248