Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills Paper • 2608.01851 • Published 6 days ago • 4
ChronoVision: Temporal Reasoning via Latent State Reconstruction Paper • 2608.05631 • Published 3 days ago • 32
MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations Paper • 2607.28956 • Published 9 days ago • 94
WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity Paper • 2608.02603 • Published 6 days ago • 34
AISPA: User-Centric System Prompt Auditing for Large Language Model Applications Paper • 2607.28617 • Published 10 days ago • 37