Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation Paper • 2609.11115 • Published 17 days ago • 172
Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments Paper • 2609.04148 • Published 24 days ago • 245
SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation Paper • 2608.21500 • Published Aug 21 • 41
VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End? Paper • 2608.15265 • Published Aug 15 • 60
HRBench: Benchmarking and Understanding Thinking-Mode Switch Strategies in Hybrid-Reasoning LLMs Paper • 2605.28398 • Published May 27 • 14