Multilingual GSM-Symbolic: What determines capability transfer across languages? Paper • 2610.03367 • Published 5 days ago • 44
Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation Paper • 2609.04298 • Published 28 days ago
The Embedder's Dilemma: LLMs Are Better, but at What Cost? Paper • 2608.12875 • Published Aug 13 • 16
SWE-Marathon: Can Agents Autonomously Complete Ultra-Long-Horizon Software Work? Paper • 2606.07682 • Published Jun 5 • 2
HUME: Measuring the Human-Model Performance Gap in Text Embedding Task Paper • 2510.10062 • Published Oct 11, 2025 • 10