TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces Paper • 2609.33295 • Published 8 days ago • 72
DriveZero: End-to-End Driving Beyond Human Demonstrations Paper • 2609.06055 • Published 30 days ago • 57
When Agents Learn to Be You: Benchmarking Privacy Leakage, Impersonation Risk, and Defenses in Persona Skills Paper • 2608.03700 • Published Aug 4 • 9
StateAct: Program State, before Pixels, for Long-Horizon Computer-Use Agents Paper • 2607.22798 • Published Jul 24 • 63
POISE: Position-Aware Undetectable Skill Injection on LLM Agents Paper • 2606.07943 • Published Jun 6 • 4
Dr. MAS: Stable Reinforcement Learning for Multi-Agent LLM Systems Paper • 2602.08847 • Published Feb 9 • 30
Online Causal Kalman Filtering for Stable and Effective Policy Optimization Paper • 2602.10609 • Published Feb 11 • 18
SimScale: Learning to Drive via Real-World Simulation at Scale Paper • 2511.23369 • Published Nov 28, 2025 • 39
MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs Paper • 2511.14159 • Published Nov 18, 2025 • 26