DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression Paper • 2609.19969 • Published 11 days ago • 186
LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents Paper • 2609.13287 • Published 19 days ago • 16
Dream-RSI: Recursive Self-Improvement through Evolving Worlds Paper • 2609.14858 • Published 14 days ago • 247
Rethinking On-Policy Distillation of Large Language Models II: One Training Example Paper • 2609.04172 • Published 25 days ago • 102
Dr. Claw: An AI Scientist Workspace for Vibe Research Paper • 2609.00365 • Published 28 days ago • 174
Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments Paper • 2609.04148 • Published 25 days ago • 246
S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement? Paper • 2608.31100 • Published 28 days ago • 42
HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? Paper • 2609.01437 • Published 27 days ago • 221
Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills Paper • 2609.02749 • Published 26 days ago • 403
MolmoWeb-Data Collection This is the collection of all datasets in MolmoWebMix. • 6 items • Updated Mar 24 • 33
yatin-superintelligence/Edge-Agent-Reasoning-WebSearch-260K Viewer • Updated Mar 13 • 260k • 5.57k • 53
StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling Paper • 2608.15089 • Published Aug 15 • 448
Training Agents to Evolve with Their Harness: TaoLive Digital Avatar Agent Technical Report Paper • 2608.15763 • Published Aug 22 • 54
AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces Paper • 2608.23041 • Published Aug 24 • 64
Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses Paper • 2608.24876 • Published Aug 25 • 31