Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation Paper • 2609.11115 • Published 29 days ago • 174
MatrAIx: Simulating the World with 8.3 Billion Persona Agents Paper • 2608.04205 • Published Aug 4 • 56 • 4
MatrAIx: Simulating the World with 8.3 Billion Persona Agents Paper • 2608.04205 • Published Aug 4 • 56
ASI-Bench: At the Dawn of Artificial Superintelligence Paper • 2608.17271 • Published Aug 18 • 68
Chain of Risk: Safety Failures in Large Reasoning Models and Mitigation via Adaptive Multi-Principle Steering Paper • 2605.05678 • Published May 7 • 6
Chain of Risk: Safety Failures in Large Reasoning Models and Mitigation via Adaptive Multi-Principle Steering Paper • 2605.05678 • Published May 7 • 6
Emergent Social Intelligence Risks in Generative Multi-Agent Systems Paper • 2603.27771 • Published Mar 29 • 50
On the Trustworthiness of Generative Foundation Models: Guideline, Assessment, and Perspective Paper • 2502.14296 • Published Feb 20, 2025 • 45
The Role of Computing Resources in Publishing Foundation Model Research Paper • 2510.13621 • Published Oct 15, 2025 • 17
The Role of Computing Resources in Publishing Foundation Model Research Paper • 2510.13621 • Published Oct 15, 2025 • 17