Multimodal Flow: Unified Flow Modeling of Language and Vision in Embedding Spaces Paper • 2609.40362 • Published 9 days ago • 38
Reflect, Revise, Reuse: Training-Free Skill Evolution for GUI Agents Paper • 2609.17653 • Published 24 days ago • 45
ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models Paper • 2609.18487 • Published 23 days ago • 48
QCell: Recombining and Aligning Cell Queries for Overlapping Instance Segmentation Paper • 2608.29253 • Published Aug 29 • 14
Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM Paper • 2609.04098 • Published Sep 3 • 86
Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination Paper • 2608.14391 • Published Aug 14 • 287
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution Paper • 2608.00677 • Published Aug 1 • 266
Business Arena: Benchmarking LLM Agents in a Realistic Marketplace Paper • 2608.08621 • Published Aug 9 • 22
Task-Conditional Flow Matching for Balanced Multilingual Text Embedding Adaptation Paper • 2608.05785 • Published Aug 6 • 11