LIBERO-Para: A Diagnostic Benchmark and Metrics for Paraphrase Robustness in VLA Models Paper • 2603.28301 • Published Mar 30 • 79
Schrödinger's Code Repository: Have LLMs Learned SWE-bench or Memorized It? Paper • 2609.27891 • Published Aug 21 • 32
The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement Paper • 2609.11873 • Published 26 days ago • 91
unsloth/gpt-oss-120b-unsloth-bnb-4bit Text Generation • 117B • Updated about 6 hours ago • 4.29k • 20
τ_0-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation Paper • 2608.16885 • Published Aug 17 • 17
ComBodied Agents: a New Paradigm of Human-Centric Agentic AI Paper • 2608.10915 • Published Aug 11 • 195