OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software Paper • 2609.39903 • Published 9 days ago • 63
Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation Paper • 2609.04298 • Published about 1 month ago
Proteo-R1: Reasoning Foundation Models for De Novo Protein Design Paper • 2605.02937 • Published Aug 10
Beyond Monolingual Deep Research: Evaluating Agents and Retrievers with Cross-Lingual BrowseComp-Plus Paper • 2606.15345 • Published Jun 13 • 17
MindEdit-Bench: Benchmarking Object-Level Counterfactual Spatial Reasoning in VLMs from In-the-Wild Photos Paper • 2607.00491 • Published Jul 1
Say Something Else: Rethinking Contextual Privacy as Information Sufficiency Paper • 2604.06409 • Published Apr 7
Code-Switching Information Retrieval: Benchmarks, Analysis, and the Limits of Current Retrievers Paper • 2604.17632 • Published Apr 19 • 12
MixSD: Mixed Contextual Self-Distillation for Knowledge Injection Paper • 2605.16865 • Published May 16 • 10
Foreground-Aware Relation Network for Geospatial Object Segmentation in High Spatial Resolution Remote Sensing Imagery Paper • 2011.09766 • Published Nov 19, 2020
BRIGHT: A globally distributed multimodal building damage assessment dataset with very-high-resolution for all-weather disaster response Paper • 2501.06019 • Published Jan 10, 2025
MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation Paper • 2503.10497 • Published Mar 13, 2025 • 3
Seeing is Believing, but How Much? A Comprehensive Analysis of Verbalized Calibration in Vision-Language Models Paper • 2505.20236 • Published May 26, 2025 • 3
DisasterM3: A Remote Sensing Vision-Language Dataset for Disaster Damage Assessment and Response Paper • 2505.21089 • Published May 27, 2025 • 5
DynamicVL: Benchmarking Multimodal Large Language Models for Dynamic City Understanding Paper • 2505.21076 • Published May 27, 2025
DisasterM3: A Remote Sensing Vision-Language Dataset for Disaster Damage Assessment and Response Paper • 2505.21089 • Published May 27, 2025 • 5
Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards Paper • 2509.21882 • Published Sep 26, 2025
DynamicVL: Benchmarking Multimodal Large Language Models for Dynamic City Understanding Paper • 2505.21076 • Published May 27, 2025
Taming Object Hallucinations with Verified Atomic Confidence Estimation Paper • 2511.09228 • Published Nov 12, 2025