Running Featured 853 Agent Memory Leaderboard 🧠 853 Unified memory evaluation · Results expected August 12.
dr-housemd/orcarouter-Qwen3.8-27B-Uncensored-exl3-2.75bpw-H4-V4 Image-Text-to-Text • 6B • Updated 16 days ago • 185 • 3
Φ-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them? Paper • 2609.10226 • Published 22 days ago • 19
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness Paper • 2609.08183 • Published 23 days ago • 327
RoboTok: An Internet-Scale Data Engine for Human Demonstration Retrieval and Dexterous Manipulation Learning Paper • 2609.03199 • Published 29 days ago • 31
Aspire: Can Models Self-Evolve from Vague Goals? Paper • 2608.31111 • Published about 1 month ago • 169
Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models Paper • 2608.25518 • Published Aug 26 • 60
Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA Paper • 2608.09819 • Published Aug 10 • 179
HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone Paper • 2607.25895 • Published Jul 28 • 95