It's Not What the Image Shows: Irrelevant Context Destabilises VLM Judges Without Informing Them Paper • 2609.37863 • Published 7 days ago • 38
Running 22 GLEE Competition — Live Leaderboard 🏆 22 Live leaderboard of the GLEE Competition @ NeurIPS 2026
From Feelings to Metrics: Understanding and Formalizing How Users Vibe-Test LLMs Paper • 2604.14137 • Published Apr 16 • 10
nvidia/NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-NVFP4 Text Generation • 45B • Updated Jul 7 • 10.1k • 130
Knowing When to Quit: A Principled Framework for Dynamic Abstention in LLM Reasoning Paper • 2604.18419 • Published Jun 12 • 3
Beyond IID: How General Are Tabular Foundation Models, Really? Paper • 2606.30410 • Published Jun 29 • 42
LLM Explainability with Counterfactual Chains and Causal Graphs Paper • 2606.05972 • Published Jun 4 • 19
A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks Paper • 2605.28556 • Published May 27 • 75
A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks Paper • 2605.28556 • Published May 27 • 75
Efficient Video Sampling: Pruning Temporally Redundant Tokens for Faster VLM Inference Paper • 2510.14624 • Published Oct 16, 2025 • 2
A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks Paper • 2605.28556 • Published May 27 • 75
Predicting Decisions of AI Agents from Limited Interaction through Text-Tabular Modeling Paper • 2605.12411 • Published May 12 • 49
MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Image Paper • 2605.10616 • Published May 11 • 143
Running on CPU Upgrade Featured 3.32k The Smol Training Playbook 📚 3.32k The secrets to building world-class LLMs