Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms Paper • 2609.27321 • Published 14 days ago • 22
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression Paper • 2609.19969 • Published 20 days ago • 224
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself Paper • 2609.22068 • Published 19 days ago • 138
UltraData Collection Ultra Scale, Ultra Quality, Ultra Coverage • 19 items • Updated 22 days ago • 114
Running Featured 1.45k FineWeb: decanting the web for the finest text data at scale 🍷 1.45k Explore and download the FineWeb web‑scale text dataset
Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments Paper • 2609.04148 • Published Sep 3 • 248
Rethinking On-Policy Distillation of Large Language Models II: One Training Example Paper • 2609.04172 • Published Sep 3 • 104
WebVR: Benchmarking Multimodal LLMs for WebPage Recreation from Videos via Human-Aligned Visual Rubrics Paper • 2603.13391 • Published Mar 11 • 21
On the Design Fundamentals of Pixel Text Representation Learning Paper • 2609.01147 • Published Sep 1 • 31
Influence-Directed Distillation: Solving the Diversity Bottleneck in Sampled-Token On-Policy Distillation Paper • 2608.29846 • Published Aug 30 • 15
Post-Training Language Models for Gold-Medal Performance in Coding Competitions Paper • 2609.02849 • Published Sep 2 • 14
HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? Paper • 2609.01437 • Published Sep 1 • 567
Step 3.5 Flash: Open Frontier-Level Intelligence with 11B Active Parameters Paper • 2602.10604 • Published Feb 11 • 202
LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering Paper • 2608.28281 • Published Aug 28 • 99