OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software Paper • 2609.39903 • Published 7 days ago • 63
UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement Paper • 2609.38721 • Published 7 days ago • 293
Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation Paper • 2609.04298 • Published 28 days ago
Proteo-R1: Reasoning Foundation Models for De Novo Protein Design Paper • 2605.02937 • Published Aug 10
Beyond Monolingual Deep Research: Evaluating Agents and Retrievers with Cross-Lingual BrowseComp-Plus Paper • 2606.15345 • Published Jun 13 • 17
MindEdit-Bench: Benchmarking Object-Level Counterfactual Spatial Reasoning in VLMs from In-the-Wild Photos Paper • 2607.00491 • Published Jul 1
Say Something Else: Rethinking Contextual Privacy as Information Sufficiency Paper • 2604.06409 • Published Apr 7
Code-Switching Information Retrieval: Benchmarks, Analysis, and the Limits of Current Retrievers Paper • 2604.17632 • Published Apr 19 • 12
MixSD: Mixed Contextual Self-Distillation for Knowledge Injection Paper • 2605.16865 • Published May 16 • 10