StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments Paper • 2608.24804 • Published about 1 month ago • 41
AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling Paper • 2608.26623 • Published 28 days ago • 20
AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling Paper • 2608.26623 • Published 28 days ago • 20
FABRIC: Framework for Agent-Based Realistic Intelligence Creation Paper • 2510.17995 • Published Oct 20, 2025
SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding Paper • 2607.10400 • Published Jul 11 • 69
ServiceNow/PrivacyAlign-Nemotron-3-Nano-4B-Annotation-Conditioned-Reward Text Generation • Updated Jun 24 • 22 • 1
SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding Paper • 2607.10400 • Published Jul 11 • 69
SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding Paper • 2607.10400 • Published Jul 11 • 69
Multi-scale Predictive Representations for Goal-conditioned Reinforcement Learning Paper • 2605.09364 • Published May 10