RewardVerse: Rubric-Guided Policy Optimization for Video Reward Modeling Paper • 2609.22947 • Published 11 days ago • 41
StudentBench: AI and human tutoring yield equivalent GRE learning gains Paper • 2609.28470 • Published 7 days ago • 10
Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents Paper • 2609.27334 • Published 7 days ago • 51
GeoPair: Geometry-Preserving Cross-Layer Factorization for Training-Free Transformer Compression Paper • 2609.25963 • Published 8 days ago • 17
Agensh: Scaling Organizational Intelligence to 1,024 Agents Paper • 2609.26781 • Published 8 days ago • 28
StableVQ: Practical Guidelines for Stable Vector-Quantized Tokenizer Training Paper • 2609.26774 • Published 8 days ago • 56
GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay Paper • 2609.25001 • Published 9 days ago • 130
Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents Paper • 2609.17708 • Published 15 days ago • 77
Continual Learning Mechanisms Compose for Long-Horizon Memorization Paper • 2609.06986 • Published 23 days ago • 375
Feyospace-v1: How the Cyber Mercury Seven Trained Frontier Cyber Models Paper • 2609.08418 • Published 22 days ago • 136
VDiff-Bench: A Challenging Benchmark for Fine-Grained Image Difference Identification Paper • 2609.06245 • Published 25 days ago • 28
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation Paper • 2609.08798 • Published 22 days ago • 84