RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources Paper • 2606.29538 • Published 11 days ago • 139
A Comprehensive Ecosystem for Open-Domain Customized Video Generation Paper • 2606.11783 • Published Jun 10 • 1
Toward Generalist Autonomous Research via Hypothesis-Tree Refinement Paper • 2606.11926 • Published Jun 10 • 128
From Raw Experience to Skill Consumption: A Systematic Study of Model-Generated Agent Skills Paper • 2605.23899 • Published May 22 • 29
SkillOpt: Executive Strategy for Self-Evolving Agent Skills Paper • 2605.23904 • Published May 22 • 262
Securing the Skies: A Comprehensive Survey on Anti-UAV Methods, Benchmarking, and Future Directions Paper • 2504.11967 • Published Mar 29
HA-VLN 2.0: An Open Benchmark and Leaderboard for Human-Aware Navigation in Discrete and Continuous Environments with Dynamic Multi-Human Interactions Paper • 2503.14229 • Published Oct 9, 2025
Towards Unified World Models for Visual Navigation via Memory-Augmented Planning and Foresight Paper • 2510.08713 • Published Mar 22 • 1
Covering Human Action Space for Computer Use: Data Synthesis and Benchmark Paper • 2605.12501 • Published May 12 • 16
MM-WebAgent: A Hierarchical Multimodal Web Agent for Webpage Generation Paper • 2604.15309 • Published Apr 16 • 8
AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation Paper • 2604.08540 • Published Apr 9 • 5
DynaVid: Learning to Generate Highly Dynamic Videos using Synthetic Motion Data Paper • 2604.01666 • Published Apr 2 • 10
BizGenEval: A Systematic Benchmark for Commercial Visual Content Generation Paper • 2603.25732 • Published Mar 26 • 11
Towards On-Policy SFT: Distribution Discriminant Theory and its Applications in LLM Training Paper • 2602.12222 • Published Feb 12
FlashMotion: Few-Step Controllable Video Generation with Trajectory Guidance Paper • 2603.12146 • Published Mar 12 • 5
ArcFlow: Unleashing 2-Step Text-to-Image Generation via High-Precision Non-Linear Flow Distillation Paper • 2602.09014 • Published Feb 9 • 3
RE-TRAC: REcursive TRAjectory Compression for Deep Search Agents Paper • 2602.02486 • Published Feb 2 • 20
FlashPortrait: 6x Faster Infinite Portrait Animation with Adaptive Latent Prediction Paper • 2512.16900 • Published Dec 18, 2025 • 11
Semantics Lead the Way: Harmonizing Semantic and Texture Modeling with Asynchronous Latent Diffusion Paper • 2512.04926 • Published Dec 4, 2025 • 42