The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks Paper • 2609.25804 • Published 10 days ago • 161
A Single Neuron Is Sufficient to Bypass Safety Alignment in Large Language Models Paper • 2605.08513 • Published May 8 • 16
Flattening Every Memory Peak in Long-Context Mixture-of-Experts Training Paper • 2609.14306 • Published 19 days ago • 20
RRSI: Regularized Recursive Self-Improvement of Agent Harnesses Paper • 2609.24972 • Published 11 days ago • 219
ScientistTwo: Pioneering the Human Knowledge Frontier with Autonomous AI Paper • 2609.19644 • Published 15 days ago • 3
Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty Paper • 2507.16806 • Published Jul 22, 2025 • 8
Calibration as a First-Class Criterion in LLM Evaluation Paper • 2609.26489 • Published 10 days ago • 4
RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents Paper • 2609.22000 • Published 14 days ago • 79
FrogNano: Training a 4B Coding Agent via Online Task Synthesis Paper • 2609.07925 • Published 25 days ago • 7
Procedural Graphs: Self-Evolving Execution Structures for LLM Agents Paper • 2609.09153 • Published 24 days ago • 44
SlopShape: Identifying AI-Generated Commercial Web Content Paper • 2609.15369 • Published 15 days ago • 1
Emergence World: Adversarial Stress-Testing of Long-Horizon Multi-Agent Systems Paper • 2609.17320 • Published 17 days ago • 4
VeriPhy: Agentic Physical Reasoning for World Model Evaluation and Refinement Paper • 2609.03153 • Published 30 days ago • 15
TacPAC: Tactile Prediction and Real-Time Action Correction in World-Action Models for Contact-Rich Manipulation Paper • 2609.05266 • Published 28 days ago • 2
SV-WAM: An Efficient Surround-View World-Action Model for End-to-End Autonomous Driving Paper • 2609.03602 • Published 29 days ago • 1