Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time Scaling Paper • 2609.19499 • Published 9 days ago • 37
Agora: Git as Shared Memory for Collective AutoResearch Paper • 2609.18094 • Published 9 days ago • 57
MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes Paper • 2609.10016 • Published 16 days ago • 29
Agentic Visual Generation: From Generative Models to Agentic Control Paper • 2609.06758 • Published 19 days ago • 34