VGI-Bench: Probing Visual Intelligence in Video Generation Models Paper • 2608.19583 • Published Aug 26 • 337
StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows Paper • 2608.17800 • Published Aug 18 • 138
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness Paper • 2609.08183 • Published 30 days ago • 328
CRISP: Cliff-awaRe Input-adaptive Sparse Prefilling with Structural-Mass-Motivated Routing Paper • 2609.01925 • Published Sep 1 • 7
Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory Paper • 2608.07169 • Published Aug 7 • 50
VisualPatchWorld: Code World Models as Latent Structured Representations for Planning Paper • 2607.25236 • Published Jul 28 • 6
The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning Paper • 2606.29526 • Published Jun 28 • 52
Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs Paper • 2605.30611 • Published May 28 • 65
How LoRA Remembers? A Parametric Memory Law for LLM Finetuning Paper • 2605.30260 • Published May 28 • 41
COLLEAGUE.SKILL: Automated AI Skill Generation via Expert Knowledge Distillation Paper • 2605.31264 • Published May 29 • 132
Verifiable Rewards Beyond Math and Code: Lightweight Corpus-Grounded Process Supervision for Factual Question Answering Paper • 2605.29648 • Published May 28 • 8
Anti-Self-Distillation for Reasoning RL via Pointwise Mutual Information Paper • 2605.11609 • Published May 12 • 49
Mean Mode Screaming: Mean--Variance Split Residuals for 1000-Layer Diffusion Transformers Paper • 2605.06169 • Published May 7 • 57
Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts Paper • 2604.19835 • Published Apr 21 • 22
Symbolic Guardrails for Domain-Specific Agents: Stronger Safety and Security Guarantees Without Sacrificing Utility Paper • 2604.15579 • Published Apr 16 • 3
LLaDA2.0-Uni: Unifying Multimodal Understanding and Generation with Diffusion Large Language Model Paper • 2604.20796 • Published Apr 22 • 107