PISmith: Reinforcement Learning-based Red Teaming for Prompt Injection Defenses Paper • 2603.13026 • Published Jul 23 • 2
Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation Paper • 2609.11115 • Published 21 days ago • 173
Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments Paper • 2609.04148 • Published 28 days ago • 248
SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation Paper • 2608.21500 • Published Aug 21 • 41
VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End? Paper • 2608.15265 • Published Aug 15 • 61
HRBench: Benchmarking and Understanding Thinking-Mode Switch Strategies in Hybrid-Reasoning LLMs Paper • 2605.28398 • Published May 27 • 14