Parallel Decoding Distillation for Fast Image and Video Generation Paper • 2607.26004 • Published 2 days ago • 9
PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models Paper • 2607.24957 • Published 3 days ago • 13
Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model Paper • 2607.24904 • Published 3 days ago • 25
A New Role for Relevance: Guiding Corpus Interaction in Agentic Search Paper • 2607.24223 • Published 3 days ago • 86
Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction Paper • 2607.20911 • Published 7 days ago • 26
NVIDIA-labs OO Agents: Native Python Object-Oriented Agents Paper • 2607.20709 • Published 8 days ago • 32
SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Paper • 2607.21553 • Published 7 days ago • 38
AREX: Towards a Recursively Self-Improving Agent for Deep Research Paper • 2607.21461 • Published 7 days ago • 149
Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Paper • 2607.21653 • Published 8 days ago • 30
Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills Paper • 2607.22529 • Published 6 days ago • 45
Oxygen-TryOn: Fashion-Native Foundation Model for Any-item Virtual Try-On Paper • 2607.21694 • Published 7 days ago • 28
OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Paper • 2607.23855 • Published 4 days ago • 25
Beyond Relevance-Centric Retrieval: Rubric-Oriented Document Set Selection and Ranking Paper • 2607.19747 • Published 8 days ago • 31
SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD Paper • 2607.20145 • Published 8 days ago • 72