Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation Paper • 2609.11115 • Published Sep 10 • 174
Graph Engineering in the Era of LLM Agents: From Individual Intelligence to System Intelligence Paper • 2608.21156 • Published Aug 21 • 65
Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts Paper • 2608.20061 • Published Aug 20 • 47
LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Paper • 2608.03457 • Published Aug 4 • 35
The Psychogenic Machine: Simulating AI Psychosis, Delusion Reinforcement and Harm Enablement in Large Language Models Paper • 2509.10970 • Published Sep 13, 2025 • 2
Human Psychometric Questionnaires Mischaracterize LLM Behavior Paper • 2509.10078 • Published May 29 • 36
LatentSkill: From In-Context Textual Skills to In-Weight Latent Skills for LLM Agents Paper • 2606.06087 • Published Jun 4 • 68
Redesign Mixture-of-Experts Routers with Manifold Power Iteration Paper • 2606.12397 • Published Jun 10 • 91
See What I See, Know What I Think: Dense Latent Communication Across Heterogeneous Agents Paper • 2606.13594 • Published Jun 11 • 6
Rethinking Psychometric Evaluation of LLMs: When and Why Self-Reports Predict Behavior Paper • 2606.12730 • Published Jun 10 • 7
Demystifying Hidden-State Recurrence: Switchable Latent Reasoning with On-Policy Reinforcement Learning Paper • 2606.13106 • Published Jun 11 • 22
EvoArena: Tracking Memory Evolution for Robust LLM Agents in Dynamic Environments Paper • 2606.13681 • Published Jun 11 • 146
Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond Paper • 2604.22748 • Published Apr 24 • 158