The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes Paper • 2602.15515 • Published Feb 17 • 2
AgentDoG 1.5: A Lightweight and Scalable Alignment Framework for AI Agent Safety and Security Paper • 2605.29801 • Published May 28 • 150
A Survey of Scientific Large Language Models: From Data Foundations to Agent Frontiers Paper • 2508.21148 • Published Aug 28, 2025 • 144