Beyond Solver Verdicts: Generative Reward Models for Autoformalization Paper • 2609.11085 • Published 23 days ago • 35
StochBench: A Domain-Specific Benchmark for Stochastic Processes in Lean Paper • 2609.09264 • Published 25 days ago • 9
Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility Paper • 2608.04001 • Published Aug 4 • 1
Mid-Think: Training-Free Intermediate-Budget Reasoning via Token-Level Triggers Paper • 2601.07036 • Published Jan 11
StochBench: A Domain-Specific Benchmark for Stochastic Processes in Lean Paper • 2609.09264 • Published 25 days ago • 9
BIOCAP: Exploiting Synthetic Captions Beyond Labels in Biological Foundation Models Paper • 2510.20095 • Published Oct 23, 2025 • 1
AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models Paper • 2506.09082 • Published May 3 • 1
Forte : Finding Outliers with Representation Typicality Estimation Paper • 2410.01322 • Published Oct 2, 2024 • 2