UFO: Chain-of-Evaluation for Omni-Condition Alignment in Multi-Modal Image Generation Paper • 2609.12397 • Published 7 days ago • 43
ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks Paper • 2609.18805 • Published 8 days ago • 62
Generative Late-Interaction Embeddings For Visual Document Retrieval Paper • 2609.11808 • Published 14 days ago • 27
Think Before You Link: Rarity, Reasoning, and Retrieval in Multilingual Entity Linking Paper • 2609.10745 • Published 15 days ago • 22
Environments as Scaffold: Enriching Feedback to Bootstrap Self-Evolving Agents in Long-Horizon Tasks Paper • 2609.08404 • Published 16 days ago • 24
RealSWE: A Compositional Evaluation of Coding Agents under Realistic User Requests Paper • 2608.27831 • Published 24 days ago • 33
LatentPress: Context Compression Beyond Text and Vision Paper • 2609.01507 • Published 23 days ago • 120
Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis Paper • 2608.18940 • Published Aug 19 • 35
Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination Paper • 2608.14391 • Published Aug 14 • 285
PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives Paper • 2608.13552 • Published Aug 13 • 47
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution Paper • 2608.00677 • Published Aug 1 • 265
ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction Paper • 2607.29677 • Published Jul 31 • 25