Where Does Retrieval-Based Open-Ended Evaluation Fail? Automatic Taxonomy Induction from Long-Form Medical Answer Factuality Verification Paper • 2609.30467 • Published 12 days ago • 11
ProAR: Learning Prospective Reasoning with Autoregressive Video Models Paper • 2610.03664 • Published 4 days ago • 26
RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations Paper • 2610.01780 • Published 5 days ago • 258
Benchmarking and Enhancing Skill-Level Memory for Partially Observable Robotic Manipulation Paper • 2609.38886 • Published 6 days ago • 16
Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows Paper • 2610.02122 • Published 5 days ago • 31
Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs Paper • 2609.32259 • Published 7 days ago • 86
Predictive Credit: Measuring What Scientific Explanations Add to Experimental Forecasts Paper • 2610.00314 • Published 7 days ago • 100