RRSI: Regularized Recursive Self-Improvement of Agent Harnesses Paper • 2609.24972 • Published 5 days ago • 206
onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction Paper • 2609.24983 • Published 5 days ago • 54
CADWorld: Computer-Use Benchmark for Long-Horizon Computer-Aided Design Paper • 2609.16251 • Published 12 days ago • 12
BI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence Paper • 2609.20886 • Published 10 days ago • 29
MintAct: A Unified Visual Agent for Digital Environments Paper • 2609.22083 • Published 8 days ago • 34
EvoOntology: A Self-Evolving Ontology Layer for Data Agents Paper • 2609.15779 • Published 12 days ago • 153
RSIAgent: Autonomous Exploration for Recursive Self-improvement in New Environments Paper • 2609.15364 • Published 12 days ago • 84
Dream-RSI: Recursive Self-Improvement through Evolving Worlds Paper • 2609.14858 • Published 12 days ago • 246
Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation Paper • 2609.11115 • Published 16 days ago • 172
ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search Paper • 2609.13356 • Published 15 days ago • 263
NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction Paper • 2609.10715 • Published 17 days ago • 329
SenseNova-U1.5: Towards Native Unified Visual Intelligence Paper • 2609.11929 • Published 16 days ago • 274
Evaluating Multimodal LLMs as Generalist Vision-Language-Action Agents for Drone Control: Commanding, Approaching, Tracking and Searching Paper • 2609.01404 • Published 25 days ago • 28
Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement Paper • 2609.01481 • Published 25 days ago • 19
The Mechanics of Democratic Dominance: A System Dynamics Paradigm for Dynamic Consent Engineering Paper • 2608.27509 • Published 26 days ago • 5