Lean Pool: An AI-Maintained Archive of Formalized Mathematics Paper • 2609.25199 • Published 7 days ago • 27
GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay Paper • 2609.25001 • Published 7 days ago • 130
Grounded Skill Synthesis from Code at Scale for Agentic Intelligence Paper • 2609.05571 • Published 24 days ago • 116
SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness Paper • 2609.20519 • Published 11 days ago • 136
When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation Paper • 2609.20511 • Published 11 days ago • 110
Continual Learning Mechanisms Compose for Long-Horizon Memorization Paper • 2609.06986 • Published 21 days ago • 374
Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation Paper • 2609.11638 • Published 18 days ago • 700
Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation Paper • 2609.11115 • Published 18 days ago • 172
Agentic Visual Generation: From Generative Models to Agentic Control Paper • 2609.06758 • Published 22 days ago • 34
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation Paper • 2609.08798 • Published 20 days ago • 83
GE-Act 2.0: Pretraining and Scaling a World-Action Model for Robotic Manipulation Paper • 2609.05588 • Published 24 days ago • 57
Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments Paper • 2609.04148 • Published 25 days ago • 245
Evaluating Multimodal LLMs as Generalist Vision-Language-Action Agents for Drone Control: Commanding, Approaching, Tracking and Searching Paper • 2609.01404 • Published 27 days ago • 28
Hi-Q: Hierarchical Evidence-guided Query Refinement for Multi-Hop Question Answering Paper • 2608.30468 • Published 28 days ago • 36
Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios Paper • 2608.25529 • Published Aug 26 • 17