Generalization Is Stability, Not Accuracy: Multi-Axis Evaluation of LLMs Paper • 2610.01428 • Published 4 days ago • 9
It's Not What the Image Shows: Irrelevant Context Destabilises VLM Judges Without Informing Them Paper • 2609.37863 • Published 6 days ago • 36
MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Image Paper • 2605.10616 • Published May 11 • 143