Discovering Cross-Language Reasoning Invariance in LLMs with Geometry-Invariant Sparse Autoencoders Paper • 2608.23809 • Published Aug 24 • 1
FORGE: Self-Evolving Agent Memory With No Weight Updates via Population Broadcast Paper • 2605.16233 • Published May 15 • 1
Context, Reasoning, and Hierarchy: A Cost-Performance Study of Compound LLM Agent Design in an Adversarial POMDP Paper • 2605.16205 • Published May 15 • 1
Learning to Grade Efficiently: A Bandit-Driven Prompt-Selection Framework for Low-Cost LLM Essay Scoring Paper • 2608.23814 • Published Aug 24 • 1
Delay-of-Gratification as a Multi-Agent Survival Micro-benchmark for Long-Horizon LLMs: Social Exposure, Personas, and Tool Use Budgets Paper • 2609.29509 • Published Aug 26 • 1
Evaluation of Multi-Turn Consistency in LLM Agents: Survival Analysis and Failure-Rationale Taxonomy Paper • 2609.29508 • Published Aug 26 • 1
Delay-of-Gratification as a Multi-Agent Survival Micro-benchmark for Long-Horizon LLMs: Social Exposure, Personas, and Tool Use Budgets Paper • 2609.29509 • Published Aug 26 • 1
Evaluation of Multi-Turn Consistency in LLM Agents: Survival Analysis and Failure-Rationale Taxonomy Paper • 2609.29508 • Published Aug 26 • 1
Context, Reasoning, and Hierarchy: A Cost-Performance Study of Compound LLM Agent Design in an Adversarial POMDP Paper • 2605.16205 • Published May 15 • 1
FORGE: Self-Evolving Agent Memory With No Weight Updates via Population Broadcast Paper • 2605.16233 • Published May 15 • 1
Discovering Cross-Language Reasoning Invariance in LLMs with Geometry-Invariant Sparse Autoencoders Paper • 2608.23809 • Published Aug 24 • 1
Learning to Grade Efficiently: A Bandit-Driven Prompt-Selection Framework for Low-Cost LLM Essay Scoring Paper • 2608.23814 • Published Aug 24 • 1
Infant Care Video Dataset for Classification of Interventions Using Transformers Paper • 2608.23838 • Published Aug 24 • 1
Infant Care Video Dataset for Classification of Interventions Using Transformers Paper • 2608.23838 • Published Aug 24 • 1