ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement Paper • 2609.14857 • Published 13 days ago • 214
EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction Paper • 2609.02783 • Published 25 days ago • 62
CyberFactory: Scaling Cyber Security Capabilities with Instances from the Wild Paper • 2608.23181 • Published Aug 24 • 34
Second Thought: Reasoning in Parallel as LLM Agents Act and Observe Paper • 2608.13667 • Published Aug 13 • 17
Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents? Paper • 2607.01211 • Published Jul 1 • 12
SecureAgentBench: Benchmarking Secure Code Generation under Realistic Vulnerability Scenarios Paper • 2509.22097 • Published Sep 26, 2025 • 3
"My productivity is boosted, but ..." Demystifying Users' Perception on AI Coding Assistants Paper • 2508.12285 • Published Aug 17, 2025
How Do Practitioners Build SE Agents? Insights from a Mixed-Methods Study Paper • 2607.10856 • Published Jul 12 • 7
"Your AI, My Shell": Demystifying Prompt Injection Attacks on Agentic AI Coding Editors Paper • 2509.22040 • Published Sep 26, 2025
SecureAgentBench: Benchmarking Secure Code Generation under Realistic Vulnerability Scenarios Paper • 2509.22097 • Published Sep 26, 2025 • 3
How Do Practitioners Build SE Agents? Insights from a Mixed-Methods Study Paper • 2607.10856 • Published Jul 12 • 7
stabilityai/stable-diffusion-xl-base-1.0 Text-to-Image • 3B • Updated Oct 30, 2023 • 3.6M • • 8.24k