ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds Paper • 2609.30199 • Published 6 days ago • 21
ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds Paper • 2609.30199 • Published 6 days ago • 21
Agents in the Large: Perception-Centered Architecture for Persistent Agents Paper • 2608.30478 • Published 30 days ago • 12
The Verification Horizon: No Silver Bullet for Coding Agent Rewards Paper • 2606.26300 • Published Jun 24 • 52
LLMEval-Logic: A Solver-Verified Chinese Benchmark for Logical Reasoning of LLMs with Adversarial Hardening Paper • 2605.19597 • Published May 19 • 20
LLMEval-Logic: A Solver-Verified Chinese Benchmark for Logical Reasoning of LLMs with Adversarial Hardening Paper • 2605.19597 • Published May 19 • 20