KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation Paper • 2607.14202 • Published 17 days ago • 42
Seeing Is Not Sharing: Some Vision-Language Models Overestimate Common Ground in Asymmetric Dialogue Paper • 2606.31719 • Published Jun 30 • 7
ClinHallu: A Benchmark for Diagnosing Stage-Wise Hallucinations in Medical MLLM Reasoning Paper • 2606.14697 • Published Jun 12 • 8
GrepSeek: Training Search Agents for Direct Corpus Interaction Paper • 2605.29307 • Published May 28 • 117
electricsheepafrica/africa-owid-employment-to-population-ratio Viewer • Updated Jun 3 • 1.85k • 39 • 1
SOD: Step-wise On-policy Distillation for Small Language Model Agents Paper • 2605.07725 • Published May 8 • 26
WBench: A Comprehensive Multi-turn Benchmark for Interactive Video World Model Evaluation Paper • 2605.25874 • Published May 25 • 106
CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence Paper • 2605.12882 • Published May 13 • 274