Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification Paper • 2607.24027 • Published 4 days ago • 34
OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers Paper • 2607.04033 • Published 27 days ago • 76
AGVBench: A Reliability-Oriented Benchmark of Data Augmentation for Vein Recognition Paper • 2607.02271 • Published 29 days ago • 17
EarlyTom: Early Token Compression Completes Fast Video Understanding Paper • 2605.30010 • Published May 28 • 32
AGVBench: A Reliability-Oriented Benchmark of Data Augmentation for Vein Recognition Paper • 2607.02271 • Published 29 days ago • 17
BadWAM: When World-Action Models Dream Right but Act Wrong Paper • 2607.15207 • Published 15 days ago • 53
OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers Paper • 2607.04033 • Published 27 days ago • 76
AGVBench: A Reliability-Oriented Benchmark of Data Augmentation for Vein Recognition Paper • 2607.02271 • Published 29 days ago • 17
EarlyTom: Early Token Compression Completes Fast Video Understanding Paper • 2605.30010 • Published May 28 • 32
MergeMix: A Unified Augmentation Paradigm for Visual and Multi-Modal Understanding Paper • 2510.23479 • Published Oct 27, 2025 • 18
LVOmniBench: Pioneering Long Audio-Video Understanding Evaluation for Omnimodal LLMs Paper • 2603.19217 • Published Mar 19 • 29
RankE: End-to-End Post-Training for Discrete Text-to-Image Generation with Decoder Co-Evolution Paper • 2605.21195 • Published May 20 • 20
RankE: End-to-End Post-Training for Discrete Text-to-Image Generation with Decoder Co-Evolution Paper • 2605.21195 • Published May 20 • 20
PASA: A Principled Embedding-Space Watermarking Approach for LLM-Generated Text under Semantic-Invariant Attacks Paper • 2605.10977 • Published May 9 • 10
World-R1: Reinforcing 3D Constraints for Text-to-Video Generation Paper • 2604.24764 • Published Apr 27 • 119
Vision-Language-Action Safety: Threats, Challenges, Evaluations, and Mechanisms Paper • 2604.23775 • Published Apr 26 • 46