YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality Paper • 2609.33757 • Published 6 days ago • 241
UNITE-AUDIO: Joint Learning of Continuous Tokenization and Latent Flow Matching for Text-to-Audio Generation Paper • 2609.28206 • Published 10 days ago • 1
AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing Paper • 2609.08936 • Published 25 days ago • 166
Harnessing CLIP and DINO: An Uncertainty-Aware Cascaded Fusion Network for Generalizable Deepfake Image Detection Paper • 2609.07670 • Published 26 days ago • 18
VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction Paper • 2608.26005 • Published Aug 26 • 163
Multi-Agent Forensic Reasoning for Generalizable Deepfake Video Detection Paper • 2608.06865 • Published Aug 7 • 11
Relevant but Incomplete: Referential Dangling as a Paradigm-Level Failure Mode in Hard Prompt Compression Paper • 2608.04569 • Published Aug 5 • 13
DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces Paper • 2608.03451 • Published Aug 4 • 34
TriAttention: Efficient Long Reasoning with Trigonometric KV Compression Paper • 2604.04921 • Published Apr 6 • 113
BEAVER: A Training-Free Hierarchical Prompt Compression Method via Structure-Aware Page Selection Paper • 2603.19635 • Published Mar 20 • 12
SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline Paper • 2505.19314 • Published May 25, 2025 • 5
A Semantically Consistent Dataset for Data-Efficient Query-Based Universal Sound Separation Paper • 2601.22599 • Published Jan 30 • 6
Toward Stable Semi-Supervised Remote Sensing Segmentation via Co-Guidance and Co-Fusion Paper • 2512.23035 • Published Dec 28, 2025 • 5
TIGER Collection TIGER: Time-frequency Interleaved Gain Extraction and Reconstruction for Efficient Speech Separation • 4 items • Updated Oct 6, 2025 • 3
SimScale: Learning to Drive via Real-World Simulation at Scale Paper • 2511.23369 • Published Nov 28, 2025 • 39
A-MemGuard: A Proactive Defense Framework for LLM-Based Agent Memory Paper • 2510.02373 • Published Sep 29, 2025 • 10
Efficient Audio-Visual Speech Separation with Discrete Lip Semantics and Multi-Scale Global-Local Attention Paper • 2509.23610 • Published Sep 28, 2025 • 15
SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse-Linear Attention Paper • 2509.24006 • Published Sep 28, 2025 • 119
DiffusionNFT: Online Diffusion Reinforcement with Forward Process Paper • 2509.16117 • Published Sep 19, 2025 • 24