An Empirical Study of Harness Design for Coding Agents Paper • 2609.20804 • Published 15 days ago • 92
Continual Learning Mechanisms Compose for Long-Horizon Memorization Paper • 2609.06986 • Published 25 days ago • 376
Recognition-Refusal Misalignment in LLMs: Why Models Answer Structurally Unanswerable Questions Paper • 2608.29109 • Published Aug 29 • 17
Using Grounded Theory for Agent Behavior Analysis at Scale Paper • 2608.30391 • Published Aug 31 • 19
orcarouter/Qwen3.8-27B-Uncensored-MLX Image-Text-to-Text • 27B • Updated about 1 hour ago • 91.1k • 1.51k
LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers Paper • 2608.06867 • Published Aug 7 • 114
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution Paper • 2608.00677 • Published Aug 1 • 266
Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval Paper • 2608.01481 • Published Aug 2 • 75
CAPEval: A Decoupled Caption Evaluation across Understanding and Generation Paper • 2608.02589 • Published Aug 3 • 26
TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning Paper • 2608.04007 • Published Aug 4 • 20
AISPA: User-Centric System Prompt Auditing for Large Language Model Applications Paper • 2607.28617 • Published Jul 30 • 37