RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations Paper • 2610.01780 • Published 6 days ago • 269
4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes Paper • 2610.03715 • Published 5 days ago • 24
LVMT: Video Mask Transformer for Long-term Video Segmentation Paper • 2609.34895 • Published 8 days ago • 17
Scaling Properties of Same-Family On-Policy Distillation Paper • 2609.32722 • Published 11 days ago • 324
PrismQuant: Optimal Null-Space Rotations for Grouped Quantizers Paper • 2609.32429 • Published 11 days ago • 5
AutoDataBench: Can Agents Write the Data That Feeds the Self-Improvement Loop? Paper • 2609.35025 • Published 9 days ago • 8
WISE-ATTA: When to Ask for Labels in Budgeted Active Test-Time Adaptation Paper • 2609.37687 • Published 8 days ago • 6