Self-Generated Feedback Destabilizes Test-Time Training: A Causal Decomposition of Long-Horizon Adaptation Paper • 2610.05076 • Published 4 days ago • 23
RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations Paper • 2610.01780 • Published 7 days ago • 275
Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows Paper • 2610.02122 • Published 7 days ago • 32
Adaptive Bridge: A Proxy-Based Decoupling Layer for Mitigating DDS Backpressure in ROS 2 Paper • 2608.15380 • Published Sep 6 • 26
Beyond Solver Verdicts: Generative Reward Models for Autoformalization Paper • 2609.11085 • Published 28 days ago • 35