DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression Paper • 2609.19969 • Published 17 days ago • 200
LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents Paper • 2609.13287 • Published 25 days ago • 16
Dream-RSI: Recursive Self-Improvement through Evolving Worlds Paper • 2609.14858 • Published 20 days ago • 251
Rethinking On-Policy Distillation of Large Language Models II: One Training Example Paper • 2609.04172 • Published about 1 month ago • 104
Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments Paper • 2609.04148 • Published about 1 month ago • 248
S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement? Paper • 2608.31100 • Published Aug 31 • 148
HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? Paper • 2609.01437 • Published Sep 1 • 422
Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills Paper • 2609.02749 • Published Sep 2 • 407
MolmoWeb-Data Collection This is the collection of all datasets in MolmoWebMix. • 6 items • Updated Mar 24 • 33
StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling Paper • 2608.15089 • Published Aug 15 • 451
Training Agents to Evolve with Their Harness: TaoLive Digital Avatar Agent Technical Report Paper • 2608.15763 • Published Aug 22 • 54
AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces Paper • 2608.23041 • Published Aug 24 • 65
Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses Paper • 2608.24876 • Published Aug 25 • 31
MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks Paper • 2608.23035 • Published Aug 24 • 42
Apodex 1.1: Scaling Agentic Intelligence for Complex Work Paper • 2608.23283 • Published Aug 24 • 212
Fara-1.5: Scalable Learning Environments for Computer Use Agents Paper • 2606.20785 • Published Jun 18 • 9
LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents Paper • 2608.17393 • Published Aug 18 • 26