TimeEvo: Failure-Driven Self-Evolution of a Time Series Agent Paper • 2609.27277 • Published 7 days ago • 21
LastOPD: Taming Collapse in Latent On-Policy Distillation Paper • 2609.28845 • Published 7 days ago • 18
TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces Paper • 2609.33295 • Published 3 days ago • 60
RSIAgent: Autonomous Exploration for Recursive Self-improvement in New Environments Paper • 2609.15364 • Published 16 days ago • 84
AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses Paper • 2608.12307 • Published Aug 12 • 118
MOSAIC: Module Discovery via Sparse Additive Identifiable Causal Learning for Scientific Time Series Paper • 2605.05524 • Published May 6 • 1
Verifiable Rewards Beyond Math and Code: Lightweight Corpus-Grounded Process Supervision for Factual Question Answering Paper • 2605.29648 • Published May 28 • 8
Stop When Reasoning Converges: Semantic-Preserving Early Exit for Reasoning Models Paper • 2605.17672 • Published May 17 • 23
TRACE: Trajectory Recovery for Continuous Mechanism Evolution in Causal Representation Learning Paper • 2601.21135 • Published Jan 29 • 8
Latent Thoughts Tuning: Bridging Context and Reasoning with Fused Information in Latent Tokens Paper • 2602.10229 • Published Feb 10 • 5
EpiQAL: Benchmarking Large Language Models in Epidemiological Question Answering for Enhanced Alignment and Reasoning Paper • 2601.03471 • Published Jan 6 • 7
QuCo-RAG: Quantifying Uncertainty from the Pre-training Corpus for Dynamic Retrieval-Augmented Generation Paper • 2512.19134 • Published Dec 22, 2025 • 32
MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs Paper • 2511.14159 • Published Nov 18, 2025 • 26