Show-Harness: Just a VLM Agent Can Play Robots Paper • 2609.10522 • Published 26 days ago • 162
FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience Paper • 2609.03241 • Published Sep 3 • 54
Running Agents Featured 50 SenseNova Vision 📚 50 Analyze or generate images with AI-powered vision tasks
AgentDoG 1.5: A Lightweight and Scalable Alignment Framework for AI Agent Safety and Security Paper • 2605.29801 • Published May 28 • 145
PaSBench-Video: A Streaming Video Benchmark for Proactive Safety Warning Paper • 2606.02443 • Published Jun 1 • 2
StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing Paper • 2608.24777 • Published Aug 25 • 17
StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing Paper • 2608.24777 • Published Aug 25 • 17